mirror of
https://github.com/ZhuLinsen/daily_stock_analysis.git
synced 2026-10-07 12:58:23 +08:00
fix(#2063): close PR #2129 review blocker OR-COR-us-prefix-nonalpha-guard-gap by extending us-prefix guard to all non-canonical bases
OpenReview Bot 在 PR #2129 head49e3da6e上重新复核后给出 1 个未关闭的高置信度 correctness blocker(OR-COR-us-prefix-nonalpha-guard-gap),同时关闭前轮的 4 个 blocker(OR-COR-9c3d2c44 / 2f0d1a7e / 7b45f5c1 三个 round-1/2 blocker 已关闭,本轮只闭合 OR-COR-us-prefix-nonalpha-guard-gap)。本 commit 闭环该剩余 blocker。 == OR-COR-us-prefix-nonalpha-guard-gap root cause == src/services/stock_list_parser.py:526-545 的 us-prefix reject guard 用 raw.isalpha() 作为前置条件: if ( raw.isalpha() # ❌ 前置 isalpha 过滤 and len(raw) > 2 and raw[:2].lower() == "us" and not raw[2:].isupper() ): return AnalysisTarget(..., asset_type=UNSUPPORTED, ...) 这意味着含标点或数字的 us-prefix 输入走不到 reject 路径,被 silently rewrite 为不同的 stock: - parse_analysis_target("usbrk.b") -> stock canonical="BRK.B" (lowecase base+标点) - parse_analysis_target("usshop.us") -> stock canonical="SHOP.US" (lowercase base+标点) - parse_analysis_target("us1") -> stock canonical="1" (lowercase prefix+数字 base) reviewer 指出这些 us-prefixed 输入应被 surfacing 为 unsupported,以免 callers 收到误导性的 canonical stock id(特别是 canonical="1" 不是合法 US symbol shape)。 reviewer 同时给出非阻断建议:补 dotted/numeric us-prefixed inputs 回归测试覆盖。 == 修法 == 把 guard 从「raw.isalpha() AND base 含 lowercase letter」改为「base 不匹配 canonical US ticker regex」: _US_TICKER_SHAPE_RE = re.compile(r"^[A-Z]{1,5}(\.[A-Z]{1,2})?$") if ( len(raw) > 2 and raw[:2].lower() == "us" and _US_TICKER_SHAPE_RE.match(raw[2:]) is None # ❌ 改为 regex match ): return AnalysisTarget(..., asset_type=UNSUPPORTED, ...) 新 _US_TICKER_SHAPE_RE 模块级常量与 data_provider/us_index_mapping.py:16-17 和 stock_code_utils._normalize_code_and_exchange 用的同一 regex 一致——canonical US symbol shape:1-5 个大写字母可选跟一个 . + 1-2 个大写字母(covers AAPL/BRK.B/SHOP.US/HKD/USFD 等)。 这个 guard 比 raw.isalpha() + lowercase-letter 检查更严格: - usbrk.b:base "brk.b" 不 match regex(含 lowercase)→ unsupported ✓ - usshop.us:base "shop.us" 不 match → unsupported ✓ - us1:base "1" 不 match(不是 1-5 大写字母)→ unsupported ✓ - US1:base "1" 不 match → unsupported ✓(含数字的 US-prefix 也被 reject,与 reviewer 期望一致) - usfd / usibm / Usaapl:base 含 lowercase → 不 match → unsupported ✓(保持原 reject) - usAAPL / usBRK.B / usSHOP.US:base match → 不 reject → 走原 split-prefix 路径 ✓ - USFD / USBRK.B / AAPL / BRK.B:raw[:2].lower()=="us" false 或 base match → 不 reject → 走原路径 ✓ 测试覆盖: - tests/test_stock_list_parser.py::test_lowercase_us_prefix_is_unsupported parametrize list 加 7 个 new case: * usbrk.b、usshop.us(lowercase base + punctuation) * us1、us1a、us12a(lowercase base + digits) * US1、US12345(all-uppercase but 含数字 invalid US shape) - 所有 case 断言 asset_type==UNSUPPORTED、exchange=="US"、canonical_id==raw、unsupported_reason 含 "uppercase" - docstring 与 parametrize 注释同步更新加 OR-COR-us-prefix-nonalpha-guard-gap 解释 == 验证 == 本地: - 111 个 stock_list_parser 测试全过(104 已有 + 7 新增 parametrize case) - PYTHONPATH=src python3 -c "from services.stock_list_parser import parse_analysis_target; for s in [...]: ..." 26 个 case 手动验证全部预期通过 - python3 -m flake8 src/services/stock_list_parser.py tests/test_stock_list_parser.py 我的改动无新增 lint 错误(pre-existing F401 'json'/'Path'/'Optional'/'AnalysisTarget' 不在本 commit 范围) - backend-gate local run 因 1.8G 内存机 OOM killed(pre-existing limitation,与 PR 2140 解决的 issue #2131 同源)留给 CI web-gate 跑 CI 状态留给 push 后看。 == 真实路径 == - PR #2129 review 在 head49e3da6e收到 OpenReview Bot OR-COR-us-prefix-nonalpha-guard-gap blocker - 本 commit 在分支 fix/pr-2122-blockers-r3 上修复并 push - CI 全绿后请 maintainer 在新 head 复审
This commit is contained in:
+1
-1
@@ -20,7 +20,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/).
|
||||
- [新功能] STOCK_LIST 解析新增 `parse_analysis_target()` 单条目解析契约,支持 sh/sz/bj/hk/us 前缀校验、裸码默认归股票、未命中前缀降级为股票三段语义;保留现有 `split_stock_list()`/`serialize_stock_list()` 行为不变,并对外暴露 `IndexRegistry`、`AnalysisTarget`、`ParseStatus`、`default_index_registry()` 以便上层注入自定义指数白名单(关联 issue #2063 Phase 1)
|
||||
- [修复] `parse_analysis_target()` 在显式交易所后缀输入被规范化层拒绝时(如 `600519.BJ`、`600000.HK`、`1234567.SH`、`abc.SH`),不再静默改写为 `sh<digits>` 或继续走裸码分类导致误判为 US,而是直接返回 `unsupported` 并携带可定位原因;保留 `000300.SH` / `sh000300.SH` 等已知 INDEX alias 的索引命中路径(关闭 PR #2122 review blocker OR-COR-607f1395 / OR-COR-26596201 / OR-COR-d6afd0d6)
|
||||
- [修复] `parse_analysis_target()` 收敛显式交易所后缀输入的 3 个新 correctness blocker:畸形混合 alias(如 `sh0x00300.SH`)不再被数字过滤后重建为已注册指数;dotted-prefix 形态(如 `SH.000999`)在 invalid base 时不被 strict-suffix 误拒,亦不静默降级为畸形 canonical stock,统一返回 `unsupported`(白名单交易所的合法 alias 通过 `sz399001.SZ` / `sz399006.SZ` 等命中 INDEX);外盘半显式 suffix(.T/.KS/.KQ/.TW/.TWO)非法 base 不再静默回退到 US stock,而是返回 `unsupported` 并标记外盘 suffix(关闭 PR #2129 review blocker OR-COR-d83a3580 / OR-COR-b3e32200 / OR-COR-e21e9de5)
|
||||
- [修复] `parse_analysis_target()` 在 `us` 前缀分支按 Phase 1 contract(issue #2063 maintainer clarification 2026-08-01)统一收紧:`us` 前缀本身大小写不敏感(`us`/`US`/`uS`/`Us` 均可),但 ticker base 必须为 canonical uppercase US symbol 形态。`us` 前缀输入只要 base 含任意小写字母(`usfd` / `usm` / `usibm` / `usamd` / `usge` / `usbk` / `usaapl` / `usshop` 全小写、`Usfd` / `USibm` / `Usaapl` / `uSfd` / `USaapl` mixed-case prefix + lowercase base)一律直接返回 `unsupported` 并提示用户使用 uppercase base(如 `usAAPL`)或 bare 大写形态(如 `USFD`)。docstring 同步更新明确「exchange 前缀大小写不敏感,但 `us` 后 base 必须 uppercase」契约。该 contract 一致规则同时关闭 PR #2129 review blocker OR-COR-9c3d2c44(`usfd`/`usm` 全 lower 静默大写)、OR-COR-2f0d1a7e(`usibm`/`usge`/`usbk`/`usaapl` 长度依赖剥前缀分叉)与 OR-COR-7b45f5c1(`Usfd`/`USibm`/`Usaapl` mixed-case prefix + lowercase base 绕过 lowercase-only guard),无需 US ticker 白名单。mixed-case `usBRK` / `UsBRK` / `uSBRK` / `usFD` / `usM`(prefix any case + upper base)与全 upper `USAAPL` 等仍按显式前缀剥前缀;`USFD`/`USM` 等 ≤5 字母全 upper 仍按 bare ticker 保留
|
||||
- [修复] `parse_analysis_target()` 在 `us` 前缀分支按 Phase 1 contract(issue #2063 maintainer clarification 2026-08-01)统一收紧:`us` 前缀本身大小写不敏感(`us`/`US`/`uS`/`Us` 均可),但 ticker base 必须为 canonical uppercase US symbol 形态(regex `^[A-Z]{1,5}(\.[A-Z]{1,2})?$`,与 `data_provider/us_index_mapping.py` 和 `stock_code_utils._normalize_code_and_exchange` 一致)。`us` 前缀输入只要 base 不 match 该 regex——含任意小写字母(`usfd` / `usm` / `usibm` / `usamd` / `usge` / `usbk` / `usaapl` / `usshop` 全小写、`Usfd` / `USibm` / `Usaapl` / `uSfd` / `USaapl` mixed-case prefix + lowercase base)、含标点(`usbrk.b` / `usshop.us` lowercase base with punctuation)、含数字(`us1` / `us1a` / `us12a` lowercase base with digits、`US1` / `US12345` all-uppercase but invalid US shape)一律直接返回 `unsupported` 并提示用户使用 uppercase base(如 `usAAPL` / `usBRK.B` / `usSHOP.US`)或 bare 大写形态(如 `USFD`)。docstring 同步更新明确「exchange 前缀大小写不敏感,但 `us` 后 base 必须 match canonical US symbol shape」契约。该 contract 一致规则同时关闭 PR #2129 review blocker OR-COR-9c3d2c44(`usfd`/`usm` 全 lower 静默大写)、OR-COR-2f0d1a7e(`usibm`/`usge`/`usbk`/`usaapl` 长度依赖剥前缀分叉)、OR-COR-7b45f5c1(`Usfd`/`USibm`/`Usaapl` mixed-case prefix + lowercase base 绕过 lowercase-only guard)与 OR-COR-us-prefix-nonalpha-guard-gap(`usbrk.b`/`usshop.us`/`us1` lowercase/non-alphabetic base 绕过早期 `raw.isalpha()`-gated guard 走到 `_split_prefix` 异常剥前缀),无需 US ticker 白名单。mixed-case `usBRK` / `UsBRK` / `uSBRK` / `usFD` / `usM`(prefix any case + upper base)与全 upper `USAAPL` 等仍按显式前缀剥前缀;`USFD`/`USM` 等 ≤5 字母全 upper 仍按 bare ticker 保留
|
||||
<!-- 新条目格式:- [类型] 描述(类型取值:新功能/改进/修复/文档/测试/chore)-->
|
||||
<!-- 每条独立一行追加到本段末尾,无需分类标题,合并时冲突最小 -->
|
||||
- [修复] 本地 CLI 的 `stdout_preview` / `stderr_preview` 按环境变量、JSON、YAML/日志标量与 URL 的独立契约脱敏短凭证,避免小于 32 字符的 API key、secret 或 token 进入诊断;普通字段仅按敏感名称判定,未加引号的 YAML 敏感标量则 fail-closed 脱敏至行尾(refs #1784)。
|
||||
|
||||
@@ -127,6 +127,16 @@ _KNOWN_PREFIXES_SORTED = tuple(
|
||||
sorted(_EXCHANGE_PREFIX_TO_CODE.keys(), key=lambda p: -len(p))
|
||||
)
|
||||
|
||||
# Canonical US ticker shape: 1-5 uppercase letters optionally followed by
|
||||
# a single dot + 1-2 uppercase letters (covers ``AAPL`` / ``BRK.B`` /
|
||||
# ``SHOP.US`` / ``HKD`` / ``USFD`` etc.). This mirrors the regex used by
|
||||
# ``data_provider/us_index_mapping.py:16-17`` and
|
||||
# ``stock_code_utils._normalize_code_and_exchange`` for the US branch.
|
||||
# Used by ``parse_analysis_target`` to enforce the Phase 1 contract that
|
||||
# ``us``-prefixed tokens must use a valid uppercase US ticker base — see
|
||||
# the guard in ``parse_analysis_target`` for the full rationale.
|
||||
_US_TICKER_SHAPE_RE = re.compile(r"^[A-Z]{1,5}(\.[A-Z]{1,2})?$")
|
||||
|
||||
|
||||
def _split_prefix(token: str) -> Tuple[Optional[str], str]:
|
||||
"""Split ``token`` into ``(prefix, bare_code)`` when the leader matches
|
||||
@@ -507,27 +517,39 @@ def parse_analysis_target(
|
||||
# Phase 1 contract (issue #2063, maintainer clarification 2026-08-01):
|
||||
# the ``us`` exchange prefix is case-insensitive on the *prefix* itself
|
||||
# but the ticker base must arrive in canonical uppercase US symbol
|
||||
# shape. ``us``-prefixed tokens whose base contains any lowercase
|
||||
# letter — e.g. ``usfd`` / ``usm`` / ``usibm`` / ``usaapl`` /
|
||||
# ``usshop`` (all-lowercase) AND ``Usfd`` / ``USibm`` / ``Usaapl``
|
||||
# (mixed-case prefix with lowercase base) — are neither bare US
|
||||
# tickers (silently upper-casing would synthesise ``USFD`` / ``USIBM``,
|
||||
# which are different real or non-existent securities) nor explicit
|
||||
# ``us``-prefix stock symbols (silently splitting would rewrite them
|
||||
# to ``FD`` / ``IBM`` / ``AAPL``, losing the user's original intent).
|
||||
# shape. ``us``-prefixed tokens whose base does NOT match the canonical
|
||||
# US ticker shape ``^[A-Z]{1,5}(\.[A-Z]{1,2})?$`` (the same regex used by
|
||||
# ``data_provider/us_index_mapping.py:16-17`` and
|
||||
# ``stock_code_utils._normalize_code_and_exchange``) — e.g. ``usfd`` /
|
||||
# ``usm`` / ``usibm`` / ``usaapl`` / ``usshop`` (all-lowercase) AND
|
||||
# ``Usfd`` / ``USibm`` / ``Usaapl`` (mixed-case prefix with lowercase
|
||||
# base) AND ``usbrk.b`` / ``usshop.us`` (lowercase base with punctuation)
|
||||
# AND ``us1`` / ``us12345`` (lowercase prefix with non-letter base —
|
||||
# pure digit bases never match the US ticker regex) — are neither bare
|
||||
# US tickers (silently upper-casing would synthesise ``USFD`` / ``USIBM``
|
||||
# / ``USBRK.B`` / ``US1``, which are different real or non-existent
|
||||
# securities, and ``US1`` is not even a valid US symbol shape) nor
|
||||
# explicit ``us``-prefix stock symbols (silently splitting would
|
||||
# rewrite them to ``FD`` / ``IBM`` / ``BRK.B`` / ``1``, losing the
|
||||
# user's original intent, and ``1`` is also not a valid US symbol).
|
||||
# Reject them up-front as ``unsupported`` regardless of what the
|
||||
# normalizer computed for ``norm_code``, so the caller can prompt
|
||||
# the user to retype the ticker base in canonical uppercase form.
|
||||
# This closes OR-COR-9c3d2c44 (``usfd``/``usm`` silent bare rewrite),
|
||||
# OR-COR-2f0d1a7e (``usibm``/``usge``/``usbk``/``usaapl`` length-dependent
|
||||
# split bifurcation) and OR-COR-7b45f5c1 (``Usfd``/``USibm``/
|
||||
# ``Usaapl`` mixed-case prefix with lowercase base) under one
|
||||
# consistent contract rule.
|
||||
# split bifurcation), OR-COR-7b45f5c1 (``Usfd``/``USibm``/
|
||||
# ``Usaapl`` mixed-case prefix with lowercase base) and
|
||||
# OR-COR-us-prefix-nonalpha-guard-gap (``usbrk.b``/``usshop.us``/
|
||||
# ``us1`` lowercase/non-alphabetic bases bypassed the earlier
|
||||
# ``raw.isalpha()``-gated guard) under one consistent contract rule:
|
||||
# the guard checks ``raw[2:]`` against the canonical US ticker regex
|
||||
# ``_US_TICKER_SHAPE_RE``, so any base that isn't a valid uppercase
|
||||
# US symbol shape is rejected up-front, regardless of case or
|
||||
# character class.
|
||||
if (
|
||||
raw.isalpha()
|
||||
and len(raw) > 2
|
||||
len(raw) > 2
|
||||
and raw[:2].lower() == "us"
|
||||
and not raw[2:].isupper()
|
||||
and _US_TICKER_SHAPE_RE.match(raw[2:]) is None
|
||||
):
|
||||
return AnalysisTarget(
|
||||
raw_input=raw_input,
|
||||
@@ -536,9 +558,10 @@ def parse_analysis_target(
|
||||
display_code=raw,
|
||||
exchange="US",
|
||||
unsupported_reason=(
|
||||
f"us-prefixed token {raw!r} must use an uppercase ticker "
|
||||
f"base (e.g. 'us{raw[2:].upper()}') or be a bare US "
|
||||
f"ticker in canonical uppercase form"
|
||||
f"us-prefixed token {raw!r} must use a canonical uppercase "
|
||||
f"US ticker base matching '^[A-Z]{{1,5}}(.[A-Z]{{1,2}})?$' "
|
||||
f"(e.g. 'usAAPL' / 'usBRK.B' / 'usSHOP.US'), or be a bare "
|
||||
f"US ticker in canonical uppercase form"
|
||||
),
|
||||
normalized_prefix=None,
|
||||
normalized_code=raw,
|
||||
|
||||
@@ -241,8 +241,9 @@ class TestContract3PrefixedUnknownDegradesToStock:
|
||||
# Phase 1 contract (issue #2063, maintainer clarification
|
||||
# 2026-08-01): the ``us`` exchange prefix is case-insensitive
|
||||
# on the prefix itself, but the ticker base must arrive in the
|
||||
# canonical uppercase US symbol shape. ``us``-prefixed tokens
|
||||
# whose base contains any lowercase letter are surfaced as
|
||||
# canonical uppercase US symbol shape (regex
|
||||
# ``^[A-Z]{1,5}(\.[A-Z]{1,2})?$``). ``us``-prefixed tokens
|
||||
# whose base does NOT match that shape are surfaced as
|
||||
# ``unsupported`` so callers can prompt the user to retype in
|
||||
# mixed/upper case. This uniformly rejects:
|
||||
# - bare-US collisions (``usfd``/``usm``, previously
|
||||
@@ -251,6 +252,9 @@ class TestContract3PrefixedUnknownDegradesToStock:
|
||||
# /``usbk``/``usaapl``/``usshop``, previously OR-COR-2f0d1a7e)
|
||||
# - mixed-case prefix with lowercase base (``Usfd``/``USibm``/
|
||||
# ``Usaapl``/``uSfd``/``USaapl``, previously OR-COR-7b45f5c1)
|
||||
# - lowercase/non-alphabetic base bypassing the earlier
|
||||
# ``raw.isalpha()``-gated guard (``usbrk.b``/``usshop.us``/
|
||||
# ``us1``, previously OR-COR-us-prefix-nonalpha-guard-gap)
|
||||
# under one consistent contract rule — no US ticker whitelist
|
||||
# or length-dependent heuristic needed. ``canonical_id`` carries
|
||||
# the raw token verbatim so the caller can echo it back to the
|
||||
@@ -273,6 +277,26 @@ class TestContract3PrefixedUnknownDegradesToStock:
|
||||
("Usaapl", "Usaapl"),
|
||||
("uSfd", "uSfd"),
|
||||
("USaapl", "USaapl"),
|
||||
# Lowercase base with punctuation/digits (OR-COR-us-prefix-
|
||||
# nonalpha-guard-gap): the earlier ``raw.isalpha()``-gated
|
||||
# guard let ``usbrk.b`` / ``usshop.us`` / ``us1`` slip through
|
||||
# to the normalizer, which silently rewrote them to ``BRK.B``
|
||||
# / ``SHOP.US`` / ``1``. The new regex-based guard catches
|
||||
# these regardless of character class — digit-only bases,
|
||||
# lowercase+dotted bases, lowercase+digit bases alike.
|
||||
("usbrk.b", "usbrk.b"),
|
||||
("usshop.us", "usshop.us"),
|
||||
("us1", "us1"),
|
||||
("us1a", "us1a"),
|
||||
("us12a", "us12a"),
|
||||
# All-uppercase but invalid US shape (digits in base): ``US1``
|
||||
# contains a digit so it doesn't match ``^[A-Z]{1,5}(\.[A-Z]{1,2})?$``.
|
||||
# Previously ``_split_prefix`` would strip ``US`` and the
|
||||
# normalizer would accept ``1`` as the canonical_id — surfacing
|
||||
# an invalid US symbol to callers. The regex-based guard
|
||||
# rejects it up-front.
|
||||
("US1", "US1"),
|
||||
("US12345", "US12345"),
|
||||
],
|
||||
)
|
||||
def test_lowercase_us_prefix_is_unsupported(
|
||||
@@ -280,15 +304,18 @@ class TestContract3PrefixedUnknownDegradesToStock:
|
||||
ticker: str,
|
||||
expected_canonical_id: str,
|
||||
) -> None:
|
||||
"""Regression for PR #2129 review blockers OR-COR-9c3d2c44 (closed),
|
||||
OR-COR-2f0d1a7e (closed), and OR-COR-7b45f5c1: ``us``-prefixed
|
||||
tokens whose ticker base contains any lowercase letter — whether
|
||||
the prefix itself is lowercase, mixed-case, or uppercase — are
|
||||
neither bare US tickers nor explicit-prefix stock symbols under
|
||||
the Phase 1 contract from issue #2063 (maintainer clarification
|
||||
2026-08-01). They are surfaced as ``unsupported`` so callers can
|
||||
prompt the user to retype the ticker base in canonical uppercase
|
||||
form (``usAAPL``/``usBRK``/``USFD``). The contract closes three
|
||||
r"""Regression for PR #2129 review blockers OR-COR-9c3d2c44 (closed),
|
||||
OR-COR-2f0d1a7e (closed), OR-COR-7b45f5c1 (closed), and
|
||||
OR-COR-us-prefix-nonalpha-guard-gap: ``us``-prefixed tokens whose
|
||||
ticker base does NOT match the canonical US symbol shape regex
|
||||
``^[A-Z]{1,5}(\.[A-Z]{1,2})?$`` — whether the prefix itself is
|
||||
lowercase, mixed-case, or uppercase, and whether the base
|
||||
contains lowercase letters, digits, or punctuation — are neither
|
||||
bare US tickers nor explicit-prefix stock symbols under the Phase
|
||||
1 contract from issue #2063 (maintainer clarification 2026-08-01).
|
||||
They are surfaced as ``unsupported`` so callers can prompt the
|
||||
user to retype the ticker base in canonical uppercase form
|
||||
(``usAAPL``/``usBRK.B``/``USFD``). The contract closes four
|
||||
prior blockers under one uniform rule — no US ticker whitelist
|
||||
required:
|
||||
* OR-COR-9c3d2c44: ``usfd``/``usm`` silent bare-rewrite bloom
|
||||
@@ -296,6 +323,9 @@ class TestContract3PrefixedUnknownDegradesToStock:
|
||||
length-dependent explicit-prefix split bifurcation
|
||||
* OR-COR-7b45f5c1: ``Usfd``/``USibm``/``Usaapl`` mixed-case
|
||||
prefix with lowercase base bypassing the lowercase-only guard
|
||||
* OR-COR-us-prefix-nonalpha-guard-gap: ``usbrk.b``/``usshop.us``/
|
||||
``us1`` lowercase base with punctuation/digits bypassing the
|
||||
earlier ``raw.isalpha()``-gated guard
|
||||
"""
|
||||
target = parse_analysis_target(ticker)
|
||||
assert target.asset_type == ParseStatus.UNSUPPORTED
|
||||
|
||||
Reference in New Issue
Block a user