mirror of
https://github.com/Egonex-AI/Understand-Anything.git
synced 2026-10-07 16:28:04 +08:00
The domain-context scanner and a merge test read files via Path.read_text(errors="replace") with no explicit encoding. pathlib's read_text/write_text fall back to locale.getpreferredencoding(False) when encoding is omitted; on Windows, absent Python UTF-8 mode (PEP 540), that is the ANSI codepage (cp1252 / cp936 / cp932), not UTF-8. extract-domain-context.py (lines 126, 211, 270, 315, 418) reads source files, .gitignore, and metadata, then writes domain-context.json. On Windows the reads decode UTF-8 source against the wrong codepage, silently mojibaking CJK comments, accented identifiers, em-dashes, and translated README text into the context fed to the domain-analyzer agent. errors="replace" does not catch this: legacy codepages decode nearly every byte to a *wrong* character rather than raising, so corruption is silent. test_merge_batch_graphs.py (lines 986, 1105) reads assembled-graph.json back for assertions; merge-batch-graphs.py writes it with ensure_ascii=False, so it can contain raw non-ASCII that the unencoded read would mojibake on Windows. Both now pass encoding="utf-8", matching every other Python script in the repo (merge-batch-graphs.py, merge-subdomain-graphs.py, parse-knowledge-base.py, merge-knowledge-graph.py). No behavior change on Linux/macOS, where the locale default is already UTF-8. Relates to the Windows-compat reports #262, #340. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>