fix(#2063): close PR #2129 review blocker OR-COR-us-prefix-nonalpha-guard-gap by extending us-prefix guard to all non-canonical bases

OpenReview Bot 在 PR #2129 head 49e3da6e 上重新复核后给出 1 个未关闭的高置信度 correctness blocker(OR-COR-us-prefix-nonalpha-guard-gap),同时关闭前轮的 4 个 blocker(OR-COR-9c3d2c44 / 2f0d1a7e / 7b45f5c1 三个 round-1/2 blocker 已关闭,本轮只闭合 OR-COR-us-prefix-nonalpha-guard-gap)。本 commit 闭环该剩余 blocker。

== OR-COR-us-prefix-nonalpha-guard-gap root cause ==

src/services/stock_list_parser.py:526-545 的 us-prefix reject guard 用 raw.isalpha() 作为前置条件:

    if (
        raw.isalpha()                                  # ❌ 前置 isalpha 过滤
        and len(raw) > 2
        and raw[:2].lower() == "us"
        and not raw[2:].isupper()
    ):
        return AnalysisTarget(..., asset_type=UNSUPPORTED, ...)

这意味着含标点或数字的 us-prefix 输入走不到 reject 路径,被 silently rewrite 为不同的 stock:
- parse_analysis_target("usbrk.b") -> stock canonical="BRK.B"  (lowecase base+标点)
- parse_analysis_target("usshop.us") -> stock canonical="SHOP.US"  (lowercase base+标点)
- parse_analysis_target("us1") -> stock canonical="1"  (lowercase prefix+数字 base)

reviewer 指出这些 us-prefixed 输入应被 surfacing 为 unsupported,以免 callers 收到误导性的 canonical stock id(特别是 canonical="1" 不是合法 US symbol shape)。

reviewer 同时给出非阻断建议:补 dotted/numeric us-prefixed inputs 回归测试覆盖。

== 修法 ==

把 guard 从「raw.isalpha() AND base 含 lowercase letter」改为「base 不匹配 canonical US ticker regex」:

    _US_TICKER_SHAPE_RE = re.compile(r"^[A-Z]{1,5}(\.[A-Z]{1,2})?$")
    if (
        len(raw) > 2
        and raw[:2].lower() == "us"
        and _US_TICKER_SHAPE_RE.match(raw[2:]) is None           # ❌ 改为 regex match
    ):
        return AnalysisTarget(..., asset_type=UNSUPPORTED, ...)

新 _US_TICKER_SHAPE_RE 模块级常量与 data_provider/us_index_mapping.py:16-17 和 stock_code_utils._normalize_code_and_exchange 用的同一 regex 一致——canonical US symbol shape:1-5 个大写字母可选跟一个 . + 1-2 个大写字母(covers AAPL/BRK.B/SHOP.US/HKD/USFD 等)。

这个 guard 比 raw.isalpha() + lowercase-letter 检查更严格:
- usbrk.b:base "brk.b" 不 match regex(含 lowercase)→ unsupported ✓
- usshop.us:base "shop.us" 不 match → unsupported ✓
- us1:base "1" 不 match(不是 1-5 大写字母)→ unsupported ✓
- US1:base "1" 不 match → unsupported ✓(含数字的 US-prefix 也被 reject,与 reviewer 期望一致)
- usfd / usibm / Usaapl:base 含 lowercase → 不 match → unsupported ✓(保持原 reject)
- usAAPL / usBRK.B / usSHOP.US:base match → 不 reject → 走原 split-prefix 路径 ✓
- USFD / USBRK.B / AAPL / BRK.B:raw[:2].lower()=="us" false 或 base match → 不 reject → 走原路径 ✓

测试覆盖:
- tests/test_stock_list_parser.py::test_lowercase_us_prefix_is_unsupported parametrize list 加 7 个 new case:
  * usbrk.b、usshop.us(lowercase base + punctuation)
  * us1、us1a、us12a(lowercase base + digits)
  * US1、US12345(all-uppercase but 含数字 invalid US shape)
- 所有 case 断言 asset_type==UNSUPPORTED、exchange=="US"、canonical_id==raw、unsupported_reason 含 "uppercase"
- docstring 与 parametrize 注释同步更新加 OR-COR-us-prefix-nonalpha-guard-gap 解释

== 验证 ==

本地:
- 111 个 stock_list_parser 测试全过(104 已有 + 7 新增 parametrize case)
- PYTHONPATH=src python3 -c "from services.stock_list_parser import parse_analysis_target; for s in [...]: ..." 26 个 case 手动验证全部预期通过
- python3 -m flake8 src/services/stock_list_parser.py tests/test_stock_list_parser.py 我的改动无新增 lint 错误(pre-existing F401 'json'/'Path'/'Optional'/'AnalysisTarget' 不在本 commit 范围)
- backend-gate local run 因 1.8G 内存机 OOM killed(pre-existing limitation,与 PR 2140 解决的 issue #2131 同源)留给 CI web-gate 跑

CI 状态留给 push 后看。

== 真实路径 ==

- PR #2129 review 在 head 49e3da6e 收到 OpenReview Bot OR-COR-us-prefix-nonalpha-guard-gap blocker
- 本 commit 在分支 fix/pr-2122-blockers-r3 上修复并 push
- CI 全绿后请 maintainer 在新 head 复审
This commit is contained in:
xxiaoxiong
2026-08-01 18:20:22 +08:00
parent 49e3da6eeb
commit 04d86b1181
3 changed files with 82 additions and 29 deletions
+1 -1
View File
@@ -20,7 +20,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/).
- [新功能] STOCK_LIST 解析新增 `parse_analysis_target()` 单条目解析契约,支持 sh/sz/bj/hk/us 前缀校验、裸码默认归股票、未命中前缀降级为股票三段语义;保留现有 `split_stock_list()`/`serialize_stock_list()` 行为不变,并对外暴露 `IndexRegistry`、`AnalysisTarget`、`ParseStatus`、`default_index_registry()` 以便上层注入自定义指数白名单(关联 issue #2063 Phase 1)
- [修复] `parse_analysis_target()` 在显式交易所后缀输入被规范化层拒绝时(如 `600519.BJ`、`600000.HK`、`1234567.SH`、`abc.SH`),不再静默改写为 `sh<digits>` 或继续走裸码分类导致误判为 US,而是直接返回 `unsupported` 并携带可定位原因;保留 `000300.SH` / `sh000300.SH` 等已知 INDEX alias 的索引命中路径(关闭 PR #2122 review blocker OR-COR-607f1395 / OR-COR-26596201 / OR-COR-d6afd0d6)
- [修复] `parse_analysis_target()` 收敛显式交易所后缀输入的 3 个新 correctness blocker:畸形混合 alias(如 `sh0x00300.SH`)不再被数字过滤后重建为已注册指数;dotted-prefix 形态(如 `SH.000999`)在 invalid base 时不被 strict-suffix 误拒,亦不静默降级为畸形 canonical stock,统一返回 `unsupported`(白名单交易所的合法 alias 通过 `sz399001.SZ` / `sz399006.SZ` 等命中 INDEX);外盘半显式 suffix(.T/.KS/.KQ/.TW/.TWO)非法 base 不再静默回退到 US stock,而是返回 `unsupported` 并标记外盘 suffix(关闭 PR #2129 review blocker OR-COR-d83a3580 / OR-COR-b3e32200 / OR-COR-e21e9de5)
- [修复] `parse_analysis_target()` 在 `us` 前缀分支按 Phase 1 contract(issue #2063 maintainer clarification 2026-08-01)统一收紧:`us` 前缀本身大小写不敏感(`us`/`US`/`uS`/`Us` 均可),但 ticker base 必须为 canonical uppercase US symbol 形态。`us` 前缀输入只要 base 含任意小写字母(`usfd` / `usm` / `usibm` / `usamd` / `usge` / `usbk` / `usaapl` / `usshop` 全小写、`Usfd` / `USibm` / `Usaapl` / `uSfd` / `USaapl` mixed-case prefix + lowercase base)一律直接返回 `unsupported` 并提示用户使用 uppercase base(如 `usAAPL`)或 bare 大写形态(如 `USFD`)。docstring 同步更新明确「exchange 前缀大小写不敏感,但 `us` 后 base 必须 uppercase」契约。该 contract 一致规则同时关闭 PR #2129 review blocker OR-COR-9c3d2c44(`usfd`/`usm` 全 lower 静默大写)、OR-COR-2f0d1a7e(`usibm`/`usge`/`usbk`/`usaapl` 长度依赖剥前缀分叉)与 OR-COR-7b45f5c1(`Usfd`/`USibm`/`Usaapl` mixed-case prefix + lowercase base 绕过 lowercase-only guard),无需 US ticker 白名单。mixed-case `usBRK` / `UsBRK` / `uSBRK` / `usFD` / `usM`(prefix any case + upper base)与全 upper `USAAPL` 等仍按显式前缀剥前缀;`USFD`/`USM` 等 ≤5 字母全 upper 仍按 bare ticker 保留
- [修复] `parse_analysis_target()` 在 `us` 前缀分支按 Phase 1 contract(issue #2063 maintainer clarification 2026-08-01)统一收紧:`us` 前缀本身大小写不敏感(`us`/`US`/`uS`/`Us` 均可),但 ticker base 必须为 canonical uppercase US symbol 形态(regex `^[A-Z]{1,5}(\.[A-Z]{1,2})?$`,与 `data_provider/us_index_mapping.py` 和 `stock_code_utils._normalize_code_and_exchange` 一致)。`us` 前缀输入只要 base 不 match 该 regex——含任意小写字母(`usfd` / `usm` / `usibm` / `usamd` / `usge` / `usbk` / `usaapl` / `usshop` 全小写、`Usfd` / `USibm` / `Usaapl` / `uSfd` / `USaapl` mixed-case prefix + lowercase base)、含标点(`usbrk.b` / `usshop.us` lowercase base with punctuation)、含数字(`us1` / `us1a` / `us12a` lowercase base with digits、`US1` / `US12345` all-uppercase but invalid US shape)一律直接返回 `unsupported` 并提示用户使用 uppercase base(如 `usAAPL` / `usBRK.B` / `usSHOP.US`)或 bare 大写形态(如 `USFD`)。docstring 同步更新明确「exchange 前缀大小写不敏感,但 `us` 后 base 必须 match canonical US symbol shape」契约。该 contract 一致规则同时关闭 PR #2129 review blocker OR-COR-9c3d2c44(`usfd`/`usm` 全 lower 静默大写)、OR-COR-2f0d1a7e(`usibm`/`usge`/`usbk`/`usaapl` 长度依赖剥前缀分叉)、OR-COR-7b45f5c1(`Usfd`/`USibm`/`Usaapl` mixed-case prefix + lowercase base 绕过 lowercase-only guard)与 OR-COR-us-prefix-nonalpha-guard-gap(`usbrk.b`/`usshop.us`/`us1` lowercase/non-alphabetic base 绕过早期 `raw.isalpha()`-gated guard 走到 `_split_prefix` 异常剥前缀),无需 US ticker 白名单。mixed-case `usBRK` / `UsBRK` / `uSBRK` / `usFD` / `usM`(prefix any case + upper base)与全 upper `USAAPL` 等仍按显式前缀剥前缀;`USFD`/`USM` 等 ≤5 字母全 upper 仍按 bare ticker 保留
<!-- 新条目格式:- [类型] 描述(类型取值:新功能/改进/修复/文档/测试/chore)-->
<!-- 每条独立一行追加到本段末尾,无需分类标题,合并时冲突最小 -->
- [修复] 本地 CLI 的 `stdout_preview` / `stderr_preview` 按环境变量、JSON、YAML/日志标量与 URL 的独立契约脱敏短凭证,避免小于 32 字符的 API key、secret 或 token 进入诊断;普通字段仅按敏感名称判定,未加引号的 YAML 敏感标量则 fail-closed 脱敏至行尾(refs #1784)。
+40 -17
View File
@@ -127,6 +127,16 @@ _KNOWN_PREFIXES_SORTED = tuple(
sorted(_EXCHANGE_PREFIX_TO_CODE.keys(), key=lambda p: -len(p))
)
# Canonical US ticker shape: 1-5 uppercase letters optionally followed by
# a single dot + 1-2 uppercase letters (covers ``AAPL`` / ``BRK.B`` /
# ``SHOP.US`` / ``HKD`` / ``USFD`` etc.). This mirrors the regex used by
# ``data_provider/us_index_mapping.py:16-17`` and
# ``stock_code_utils._normalize_code_and_exchange`` for the US branch.
# Used by ``parse_analysis_target`` to enforce the Phase 1 contract that
# ``us``-prefixed tokens must use a valid uppercase US ticker base — see
# the guard in ``parse_analysis_target`` for the full rationale.
_US_TICKER_SHAPE_RE = re.compile(r"^[A-Z]{1,5}(\.[A-Z]{1,2})?$")
def _split_prefix(token: str) -> Tuple[Optional[str], str]:
"""Split ``token`` into ``(prefix, bare_code)`` when the leader matches
@@ -507,27 +517,39 @@ def parse_analysis_target(
# Phase 1 contract (issue #2063, maintainer clarification 2026-08-01):
# the ``us`` exchange prefix is case-insensitive on the *prefix* itself
# but the ticker base must arrive in canonical uppercase US symbol
# shape. ``us``-prefixed tokens whose base contains any lowercase
# letter — e.g. ``usfd`` / ``usm`` / ``usibm`` / ``usaapl`` /
# ``usshop`` (all-lowercase) AND ``Usfd`` / ``USibm`` / ``Usaapl``
# (mixed-case prefix with lowercase base) — are neither bare US
# tickers (silently upper-casing would synthesise ``USFD`` / ``USIBM``,
# which are different real or non-existent securities) nor explicit
# ``us``-prefix stock symbols (silently splitting would rewrite them
# to ``FD`` / ``IBM`` / ``AAPL``, losing the user's original intent).
# shape. ``us``-prefixed tokens whose base does NOT match the canonical
# US ticker shape ``^[A-Z]{1,5}(\.[A-Z]{1,2})?$`` (the same regex used by
# ``data_provider/us_index_mapping.py:16-17`` and
# ``stock_code_utils._normalize_code_and_exchange``) — e.g. ``usfd`` /
# ``usm`` / ``usibm`` / ``usaapl`` / ``usshop`` (all-lowercase) AND
# ``Usfd`` / ``USibm`` / ``Usaapl`` (mixed-case prefix with lowercase
# base) AND ``usbrk.b`` / ``usshop.us`` (lowercase base with punctuation)
# AND ``us1`` / ``us12345`` (lowercase prefix with non-letter base —
# pure digit bases never match the US ticker regex) — are neither bare
# US tickers (silently upper-casing would synthesise ``USFD`` / ``USIBM``
# / ``USBRK.B`` / ``US1``, which are different real or non-existent
# securities, and ``US1`` is not even a valid US symbol shape) nor
# explicit ``us``-prefix stock symbols (silently splitting would
# rewrite them to ``FD`` / ``IBM`` / ``BRK.B`` / ``1``, losing the
# user's original intent, and ``1`` is also not a valid US symbol).
# Reject them up-front as ``unsupported`` regardless of what the
# normalizer computed for ``norm_code``, so the caller can prompt
# the user to retype the ticker base in canonical uppercase form.
# This closes OR-COR-9c3d2c44 (``usfd``/``usm`` silent bare rewrite),
# OR-COR-2f0d1a7e (``usibm``/``usge``/``usbk``/``usaapl`` length-dependent
# split bifurcation) and OR-COR-7b45f5c1 (``Usfd``/``USibm``/
# ``Usaapl`` mixed-case prefix with lowercase base) under one
# consistent contract rule.
# split bifurcation), OR-COR-7b45f5c1 (``Usfd``/``USibm``/
# ``Usaapl`` mixed-case prefix with lowercase base) and
# OR-COR-us-prefix-nonalpha-guard-gap (``usbrk.b``/``usshop.us``/
# ``us1`` lowercase/non-alphabetic bases bypassed the earlier
# ``raw.isalpha()``-gated guard) under one consistent contract rule:
# the guard checks ``raw[2:]`` against the canonical US ticker regex
# ``_US_TICKER_SHAPE_RE``, so any base that isn't a valid uppercase
# US symbol shape is rejected up-front, regardless of case or
# character class.
if (
raw.isalpha()
and len(raw) > 2
len(raw) > 2
and raw[:2].lower() == "us"
and not raw[2:].isupper()
and _US_TICKER_SHAPE_RE.match(raw[2:]) is None
):
return AnalysisTarget(
raw_input=raw_input,
@@ -536,9 +558,10 @@ def parse_analysis_target(
display_code=raw,
exchange="US",
unsupported_reason=(
f"us-prefixed token {raw!r} must use an uppercase ticker "
f"base (e.g. 'us{raw[2:].upper()}') or be a bare US "
f"ticker in canonical uppercase form"
f"us-prefixed token {raw!r} must use a canonical uppercase "
f"US ticker base matching '^[A-Z]{{1,5}}(.[A-Z]{{1,2}})?$' "
f"(e.g. 'usAAPL' / 'usBRK.B' / 'usSHOP.US'), or be a bare "
f"US ticker in canonical uppercase form"
),
normalized_prefix=None,
normalized_code=raw,
+41 -11
View File
@@ -241,8 +241,9 @@ class TestContract3PrefixedUnknownDegradesToStock:
# Phase 1 contract (issue #2063, maintainer clarification
# 2026-08-01): the ``us`` exchange prefix is case-insensitive
# on the prefix itself, but the ticker base must arrive in the
# canonical uppercase US symbol shape. ``us``-prefixed tokens
# whose base contains any lowercase letter are surfaced as
# canonical uppercase US symbol shape (regex
# ``^[A-Z]{1,5}(\.[A-Z]{1,2})?$``). ``us``-prefixed tokens
# whose base does NOT match that shape are surfaced as
# ``unsupported`` so callers can prompt the user to retype in
# mixed/upper case. This uniformly rejects:
# - bare-US collisions (``usfd``/``usm``, previously
@@ -251,6 +252,9 @@ class TestContract3PrefixedUnknownDegradesToStock:
# /``usbk``/``usaapl``/``usshop``, previously OR-COR-2f0d1a7e)
# - mixed-case prefix with lowercase base (``Usfd``/``USibm``/
# ``Usaapl``/``uSfd``/``USaapl``, previously OR-COR-7b45f5c1)
# - lowercase/non-alphabetic base bypassing the earlier
# ``raw.isalpha()``-gated guard (``usbrk.b``/``usshop.us``/
# ``us1``, previously OR-COR-us-prefix-nonalpha-guard-gap)
# under one consistent contract rule — no US ticker whitelist
# or length-dependent heuristic needed. ``canonical_id`` carries
# the raw token verbatim so the caller can echo it back to the
@@ -273,6 +277,26 @@ class TestContract3PrefixedUnknownDegradesToStock:
("Usaapl", "Usaapl"),
("uSfd", "uSfd"),
("USaapl", "USaapl"),
# Lowercase base with punctuation/digits (OR-COR-us-prefix-
# nonalpha-guard-gap): the earlier ``raw.isalpha()``-gated
# guard let ``usbrk.b`` / ``usshop.us`` / ``us1`` slip through
# to the normalizer, which silently rewrote them to ``BRK.B``
# / ``SHOP.US`` / ``1``. The new regex-based guard catches
# these regardless of character class — digit-only bases,
# lowercase+dotted bases, lowercase+digit bases alike.
("usbrk.b", "usbrk.b"),
("usshop.us", "usshop.us"),
("us1", "us1"),
("us1a", "us1a"),
("us12a", "us12a"),
# All-uppercase but invalid US shape (digits in base): ``US1``
# contains a digit so it doesn't match ``^[A-Z]{1,5}(\.[A-Z]{1,2})?$``.
# Previously ``_split_prefix`` would strip ``US`` and the
# normalizer would accept ``1`` as the canonical_id — surfacing
# an invalid US symbol to callers. The regex-based guard
# rejects it up-front.
("US1", "US1"),
("US12345", "US12345"),
],
)
def test_lowercase_us_prefix_is_unsupported(
@@ -280,15 +304,18 @@ class TestContract3PrefixedUnknownDegradesToStock:
ticker: str,
expected_canonical_id: str,
) -> None:
"""Regression for PR #2129 review blockers OR-COR-9c3d2c44 (closed),
OR-COR-2f0d1a7e (closed), and OR-COR-7b45f5c1: ``us``-prefixed
tokens whose ticker base contains any lowercase letter — whether
the prefix itself is lowercase, mixed-case, or uppercase — are
neither bare US tickers nor explicit-prefix stock symbols under
the Phase 1 contract from issue #2063 (maintainer clarification
2026-08-01). They are surfaced as ``unsupported`` so callers can
prompt the user to retype the ticker base in canonical uppercase
form (``usAAPL``/``usBRK``/``USFD``). The contract closes three
r"""Regression for PR #2129 review blockers OR-COR-9c3d2c44 (closed),
OR-COR-2f0d1a7e (closed), OR-COR-7b45f5c1 (closed), and
OR-COR-us-prefix-nonalpha-guard-gap: ``us``-prefixed tokens whose
ticker base does NOT match the canonical US symbol shape regex
``^[A-Z]{1,5}(\.[A-Z]{1,2})?$`` — whether the prefix itself is
lowercase, mixed-case, or uppercase, and whether the base
contains lowercase letters, digits, or punctuation — are neither
bare US tickers nor explicit-prefix stock symbols under the Phase
1 contract from issue #2063 (maintainer clarification 2026-08-01).
They are surfaced as ``unsupported`` so callers can prompt the
user to retype the ticker base in canonical uppercase form
(``usAAPL``/``usBRK.B``/``USFD``). The contract closes four
prior blockers under one uniform rule — no US ticker whitelist
required:
* OR-COR-9c3d2c44: ``usfd``/``usm`` silent bare-rewrite bloom
@@ -296,6 +323,9 @@ class TestContract3PrefixedUnknownDegradesToStock:
length-dependent explicit-prefix split bifurcation
* OR-COR-7b45f5c1: ``Usfd``/``USibm``/``Usaapl`` mixed-case
prefix with lowercase base bypassing the lowercase-only guard
* OR-COR-us-prefix-nonalpha-guard-gap: ``usbrk.b``/``usshop.us``/
``us1`` lowercase base with punctuation/digits bypassing the
earlier ``raw.isalpha()``-gated guard
"""
target = parse_analysis_target(ticker)
assert target.asset_type == ParseStatus.UNSUPPORTED