Commit Graph
5470 Commits
Author SHA1 Message Date
Wesley Liddick 954d44c19c Merge pull request #5198 from Wei-Shaw/fix/codex-originator-load-shed
fix(codex): 归一化降载 originator,缓解 Codex 账号频繁过载不可用
2026-08-03 09:59:34 +08:00
shaw e1b76e2245 fix(codex): normalize load-shed originators to avoid upstream capacity shedding
上游 /backend-api/codex 按 Originator 头分桶调度容量:落在降载桶的请求即使返回
HTTP 200,也会立刻推 SSE `event: error`(code=server_is_overloaded)并以
response.failed 收尾。2026-07-29 起 codex-tui 落入降载桶,codex_cli_rs 正常——
判定因子是 originator 而非 User-Agent(codex_cli_rs 配 curl UA 亦可正常返回)。

网关会把该错误判定为瞬时上游故障并冷却账号,对外表现为 Codex 账号频繁过载不可用:
server_is_overloaded → isOpenAITransientProcessingError →
shouldCooldownOpenAITransientUpstreamError → 账号冷却 → 客户端 503。

本项目有三处降载身份来源:浏览器 UA 兜底的默认 UA、客户端透传的真实 TUI 身份、
以及指纹缓存注入探针的 UA。

修复收口在 enforceCodexIdentityHeaders——HTTP / 透传 / WS 握手 / compat 桥接 /
探针 / PAT / 模型列表 / alpha-search 八条出站路径共用的唯一纯函数收口点:

- 新增 NormalizeCodexClientIdentityToCLI,把降载桶身份改写为 codex_cli_rs,
  只替换身份段并裁掉尾部 (name; version) 客户端标识组,保留版本 / OS / 架构 /
  终端指纹;改写后 originator 与 UA 首段仍然配套,不破坏 #3901 的配对不变式,
  且改写幂等。
- DefaultOpenAICodexUserAgent 从 TUI 身份改为 CLI 身份(浏览器兜底路径上最大的
  降载身份来源)。
- 管理端 Codex UA 的 placeholder / hint 原本在把管理员往降载桶引导,一并修正。

新增 gateway.disable_codex_originator_normalization(默认 false,即归一化开启),
供上游调整分桶后回滚。该开关经 NewOpenAIGatewayService 发布为进程级快照,故必须
保持反义命名:正向命名的 Go 零值 false 会让未经 viper 加载而手工构造的 Config
静默关掉全局保护,viper.SetDefault 救不了这条路径。已加用例钉住该属性。

降载桶集合是上游容量策略快照而非协议常量,上游调整分桶后需同步修订。
2026-08-02 23:00:12 +08:00
github-actions[bot] 7e2e9ba050 chore: sync VERSION to 0.1.170 [skip ci] 2026-08-02 10:46:21 +00:00
Wesley Liddick c043c24774 Merge pull request #4925 from Brisbanehuang/feat/group-profit-control
feat(scheduler): 分组级利润控制——按账号倍率过滤 token 调度候选
v0.1.170
2026-08-02 18:19:13 +08:00
Wesley Liddick 11c1e944b9 Merge pull request #4911 from Brisbanehuang/feat/upstream-billing-rate-writeback
feat(billing-probe): 探测成功后可选将上游声明倍率自动同步为账号倍率
2026-08-02 18:18:50 +08:00
Wesley Liddick d99ee72911 Merge pull request #4896 from Brisbanehuang/feat/upstream-billing-probe-multi-platform
feat(billing-probe): 上游计费倍率探测放宽到全部 API-key 平台账号
2026-08-02 18:18:37 +08:00
shaw dec47e8fae fix(profit-control): stop leaking profit policy, close veto livelock, restore passthrough turn pricing
审计修复,逐条如下。

H1 利润策略泄露给所有普通用户
  profit_control_enabled / profit_min_margin / profit_safety_buffer 从
  dto.Group 移到 dto.AdminGroup(后者内嵌前者),赋值相应从
  groupFromServiceBase 移到 GroupFromServiceAdmin;前端 TS 同步从 Group 移到
  AdminGroup。dto.Group 是 GET /api/v1/groups/available 的响应体,该响应本就带
  rate_multiplier,相乘即可反推运营方上游采购成本上限。
  api_contract_test.go 的 /groups/available golden JSON 回滚这三个字段,并把
  fixture 改成非零值(require.JSONEq 是精确比对,缺字段即失败)。
  新增 dto 层边界测试:普通用户 DTO 不含三字段、管理员 DTO 仍含。

M1 利润终检 continue 与 failover 503 退避互动产生活锁
  FailoverState 新增 profitVetoedAccountIDs / profitVetoCount 与
  RecordProfitVeto():加入排除集 + 计数,达 maxProfitVetoAttempts(10) 返回
  FailoverExhausted。HandleSelectionExhausted 的 503 清空分支改为清空后把利润
  否决的账号放回排除集;若排除集已全部由利润否决贡献,清空不会带来任何新候选,
  直接判定耗尽(否则 SwitchCount 永不前进、退避条件永远成立,每 2s 空转一轮)。
  五个 handler 否决点(gateway_handler ×2 / responses / chat_completions /
  gemini_v1beta)改为经 RecordProfitVeto 决策,耗尽时按无可用账号终止。
  回归测试钉死:503 之后持续利润否决必须有限步终止且不 spin;未启用利润控制的
  请求退避语义完全不变。

M2 排队等槽后才终检,延迟可放大到 N × WaitPlan.Timeout
  OpenAI 侧选号循环(自有 failedAccountIDs map,非 FailoverState)新增
  recordOpenAIProfitVeto + handleOpenAIProfitVetoExhausted,共用同一上限语义。
  覆盖 responses / messages-dispatch / chat_completions / alpha_search /
  embeddings / images / grok_media 七处,以及 WS 两处否决分支。

M4 ws_v2 透传 ingress 绕过 per-turn 重定价(选方案 B:最小止血)
  透传 relay 只回调 AfterTurn、没有任何 turn 起始回调,hooks.BeforeTurn 永远
  不触发,而 handler 把 turnPricingAt 初始化成建连时刻 ⇒ 透传连接全部 turn 按
  建连时刻的高峰因子结算,客户端峰前建连保活即可全程谷价——正是本 PR 想堵的
  漏洞。改为 openAIWSTurnPricing 零值起步、只由 BeforeTurn 冻结;透传路径保持
  零值,RecordUsage 回退记录时刻,与引入利润控制前的基线一致。
  未选方案 A(给透传补 turn 起始回调):passthrough_relay.go 是 #5167 刚修过的
  取消传播/close frame 时序敏感区;且 BeforeTurn 还承担 turn>1 的并发槽位抢占,
  接进去等于给透传连接引入 per-turn 抢槽,风险远超本次修复范围。透传仍有建连时
  的准入门,只是没有 turn 级复核,已在两处注释写明。
  测试:service 层钉死透传 ingress 不触发 BeforeTurn(含失败时的复核指引),
  handler 层钉死零值语义与逐 turn 覆盖。

M5 装门读分组走了带账号计数聚合的 GetByID
  SchedulerSnapshotService 新增 GetGroupByIDLite,openai/gateway 两处装门改用
  之。门只需要平台/倍率/利润/高峰字段,且该查询发生在「是否启用利润控制」判定
  之前,未启用的分组同样付代价。两个测试 stub 的 GetByID 改成 panic 守卫。

M6 认证快照注释与真实读取路径相反
  门解析优先取 ctxkey.Group,而它就是本快照物化出来的对象,直连流量走的正是这
  条路。改正注释,与 api_key_repo.go 投影处的说明对齐,避免后人照旧注释删列。

M3 rate_multiplier 为 nil 时利润门 fail-closed(不改行为,加护栏)
  保留 fail-closed。补 repository 层测试钉死账号调度快照的 full/metadata 两份
  payload 都必须保留 RateMultiplier(含 0 值),漏列在 CI 就红。

L1 迁移号注释 191 / 191-192 改为实际的 192/193。
L2 admin group Create 的利润配置预校验改用与 CreateGroup 一致的归一化平台
   (新增 service.NormalizeGroupPlatform,两边共用)。保留预校验而非删除:
   service 层返回的是无类型 error,经 ErrorFrom 会变成 500,删掉会把合法的
   400 降级成 500。
L3 前端利润校验的上界改为判定换算后的小数(后端按小数校验 [0,1)),
   99.999% 会四舍五入进位成 1.0 而被后端 400;i18n en/zh 同步改为 0-99.99。
L4 clampProfitControlThreshold / profitControlOverThreshold 抽为共用函数,
   线上装门/否决点与 profit-preview 不再各自实现,附边界语义测试。
L5 profit-preview 补「默认 D 有账号但最低有效 D 归零」的告警(两档都为 0 由
   既有告警覆盖,不重复)。
2026-08-01 22:39:33 +08:00
Brisbanehuang fad2f215e8 fix(profit-control): decouple gate from image intent, propagate gate via selection, restore eager sticky fallback
Review fixes for the profit-control feature commit:

- Image intent no longer disables the profit gate. The shared /v1/responses
  handler previously skipped the pricing context (and therefore the gate)
  whenever the platform-wide image intent predicate matched, which includes
  Codex's passive image_gen namespace declaration: any client could disable
  admission control for anthropic/gemini/antigravity groups by declaring a
  namespace tool in the request body. Both /v1/responses paths now always
  install the token pricing context; image intent only drives capability
  routing and image billing. Mixed token+image requests stay token-gated;
  only dedicated media endpoints remain out of scope.
- Out-of-scope paths are now explicitly suppressed instead of implicitly
  ungated: Grok media (billed by media multipliers; also prevents in-flight
  video lookups from turning into spurious 404s), OpenAI-group count_tokens
  (unbilled), and Live calls (duration billed) carry a suppress marker that
  every install point honors, including the defensive scheduler-entry
  install.
- The gate resolved during selection now travels back to handlers on the
  AccountSelectionResult. The shared gateway installed the gate only on a
  scheduler-local context, so handler-side post-slot terminal rechecks and
  post-admission sticky binding were no-ops for anthropic/gemini/
  antigravity/shared-grok requests (and for composite-routed member groups
  on the OpenAI path). Handlers re-apply the carried gate via
  ContextWithSelectionProfitGate before the terminal recheck and binding;
  the WS acquired-selection branch gained the previously missing recheck.
- Sticky binding semantics restored for ungated traffic:
  BindStickySessionAfterProfitAdmission falls back to the official eager
  bind when no gate is installed (wait paths lost their only binding point
  otherwise), reads the pre-existing binding at bind time only when gated
  (removes the unconditional per-request Redis read the feature added to
  the shared handlers), and the legacy engine's three selection-time
  binding writes are skipped under a gate so a terminally vetoed account
  can no longer become the new sticky target.
- Responses WS connections re-freeze pricingAt and re-resolve the gate at
  every turn (BeforeTurn) and bill each turn with its own instant, closing
  the connect-at-valley/bill-at-valley window; a turn that fails the
  recheck closes the connection so the client reselects on reconnect.
- Terminal recheck no longer swaps a DB-fresh selected account for a
  staler snapshot object (UpdatedAt guard), and observer counters are
  documented as per-evaluation.
- Migrations renumbered 191/192 -> 192/193 after the passkey migration
  landed upstream as 191.

New regressions: selection-carried gate propagation (control group proves
the pre-fix no-op), image intent not disabling the shared gate, eager
binding fallback without a gate (both services), gated
read-failure/sentinel-miss binding semantics, legacy-engine deferred
binding under a gate with eager behavior preserved ungated, turn-level
pricing refresh (config re-resolution, scheduled-group precedence,
suppress, mid-connection disable), and suppress-marker coverage for the
request pricing context.
2026-08-01 22:39:32 +08:00
Brisbanehuang 20ad5ec506 feat(scheduler): per-group profit control for token account admission
Group pricing (rate multiplier, peak windows, per-user overrides) and
account cost (accounts.rate_multiplier) already live side by side, but
nothing stops the scheduler from handing a request to an account whose
cost multiplier exceeds what the group's pricing can profitably serve.
Add an opt-in per-group profit gate that filters scheduling candidates
by a margin rule, while ordering, scoring, stickiness and breakers keep
working unchanged among qualified accounts.

Admission rule: an account qualifies iff U <= D * (1 - min_margin -
safety_buffer) within a small relative epsilon, where U is
accounts.rate_multiplier (0 is legal; missing/negative/NaN/Inf are
conservatively rejected as invalid) and D is the requester's effective
downstream multiplier (user-group override ?? group default, times the
group peak factor) frozen at the request's pricing instant.

- groups gain profit_control_enabled / profit_min_margin /
  profit_safety_buffer (migration 191); the durable auth-cache
  invalidation trigger additionally watches the profit and pricing
  columns (migration 192) so out-of-band group edits cannot leave
  stale auth snapshots; GetByKeyForAuth explicitly projects the new
  columns and the API-key auth snapshot version is bumped to force a
  refresh of pre-existing snapshots
- request-level pricing instant: token entry points install pricingAt
  into ctx; the profit threshold D and the RecordUsage peak factor
  read the same instant, so one request never changes price mid-flight
  across waits/retries/failover (media and unwired paths keep the
  existing record-time semantics)
- the gate covers token requests on openai, anthropic, gemini, grok
  and antigravity groups: OpenAI-family handlers via
  WithOpenAIRequestPricingContext (responses incl. WS bridge, chat
  completions, messages, embeddings, alpha search), the shared gateway
  via WithGatewayTokenRequestPricing (messages, chat completions,
  responses, gemini model actions); composite groups cannot enable it
  directly; image/video/models/usage/count_tokens stay ungated and an
  explicit image-generation intent suppresses the gate end to end
- post-slot recheck: after a slot is acquired the account is re-read
  via SchedulerSnapshotService.GetAccount (scheduler cache, then DB;
  only when both fail the check fails open with WARN + metric); a
  vetoed account releases its slot and joins the request's exclusion
  set for reselection; sticky bindings are written only after the
  final check passes, and an over-threshold sticky account is skipped,
  not deleted, so it comes back once its rate recovers
- sticky-session cache contract: GatewayCache.GetSessionAccountID now
  returns ErrStickySessionNotFound on a miss (mapped from redis.Nil in
  the repository implementation, mirroring ErrRefreshTokenNotFound) so
  the profit sticky path can distinguish "no binding yet" from a real
  read failure without importing the cache driver in service code
- cross-group re-entry (composite parent -> member group) resolves the
  gate against the member group and clears a stale parent gate instead
  of letting a foreign threshold veto accounts
- per-platform/group activity counters (installs, threshold vetoes,
  invalid-rate vetoes, refresh failures) for observability
- admin UI: profit-control section on the five platforms' group forms
  with percent input, validation and platform-switch reset; group
  create/update/duplicate normalize and validate the config at a
  single choke point
- cmd/profit-preview: offline what-if tool that replays the production
  admission semantics over an exported config/account/override/model
  dump, reports per-model admitted-account counts under the default
  and the worst-case (lowest user override) D, and surfaces probe-sync
  staleness as warnings without affecting admission

Tests: service unit coverage for gate resolution/veto/threshold
epsilon/pricing instant/suppress marker/scheduler filtering and
post-slot recheck (incl. -race on the profit surface), unit-tagged
handler slot-recheck and capability-mapping regressions, sqlmock and
real-PostgreSQL integration regressions for the GetByKeyForAuth
projection and the migration-192 trigger watch list, API contract
update, and frontend specs for the five-platform form helpers.
2026-08-01 22:39:31 +08:00
shaw 0b6b4ea956 fix(billing-probe): govern the automatic account rate write-back
- 自动写回值域治理与留痕:上游声明值必须 > 0 且 <= 100 才写回。0 会让
  accountCost 恒为 0(账号总额/日/周配额与成本告警全部静默失效),极大值可
  一次打爆配额并污染成本报表;越界时保持原倍率、记 WARN,探测快照照常记 ok。
  写回成功时记结构化 slog(account_id / 旧值 / 新值 / source),并在快照中
  新增 synced_rate_multiplier 记录本次写回值——后台任务裸 SQL 不产生
  audit_logs,这两处是唯一可追溯来源。管理员手工设 0 不受影响。
- 写回改用 resolved_rate_multiplier(不含高峰的基准倍率):effective 含探测
  那一刻的高峰系数,写回会把一个探测周期的峰值/谷值冻结进静态列,而展示与
  调度用 upstreamBillingRateAt 按当前时间重算高峰,两者会持续不一致。
- 倍率解析失败不再污染公共探测路径:账号级值域/精度只在该账号已开启同步、
  真要写回时才有影响;未开同步的账号照常记 ok 快照,不累计 failure_count、
  不进入指数退避。
- 单账号编辑补 service 层守卫:同步开启时拒绝手工倍率(新增
  UPSTREAM_BILLING_RATE_SYNC_CONFLICT,与批量路径同族),此前只有前端
  disabled 挡人,直接 PUT /admin/accounts/{id} 可写入并活到下次成功探测。
  判断的是本次请求生效后的状态,"关同步 + 改倍率"同请求仍然放行。
- 修正开关反推方向:不再由 rate_sync=true 推出 probe=true,否则一条"同步开、
  探测键缺失"的僵尸记录会在任意一次无关编辑时静默打开周期性外呼;改为探测
  关闭/缺失一律把同步归零。
- 删除死代码 UpdateWithUpstreamBillingProbeEnabled(PR 删接口后生产已无调用
  方),其回滚测试改为直接覆盖生产路径 UpdateWithAccountBillingSettings。
  upstreamBillingRateSyncEnabled 不再是只服务测试的假门控,现为写回前置过滤,
  SQL CAS 仍是权威门控,两侧均加注释说明分工。
- en/zh 文案补充:同步的是不含高峰的基准倍率;开启同步会连带打开自动探测。
2026-08-01 22:11:10 +08:00
Brisbanehuang b0f5007f04 feat(billing-probe): optionally sync account rate from upstream declared rate
Successful upstream billing probes already persist the upstream-declared
rate as a display-only snapshot. Add a per-account opt-in that writes
that declared rate back to the account's rate_multiplier, so the account
cost basis follows upstream repricing automatically instead of drifting
until an operator notices.

- new per-account flag upstream_billing_rate_sync_enabled stored next to
  the probe flag in account extra: enabling sync force-enables the
  probe, disabling the probe cascades sync off, and eligibility follows
  IsUpstreamBillingProbeIdentity (tightened from any non-empty platform
  to an explicit whitelist of the five supported API-key platforms so
  future platforms do not silently inherit probe/sync semantics)
- only a successful probe whose declared rate survives validation
  (finite, within bounds, not rounded to zero at the rate_multiplier
  decimal(10,4) scale) writes back; failed/unsupported/invalid probes
  leave rate_multiplier unchanged
- the writeback rides the existing snapshot CAS transaction:
  UpdateUpstreamBillingProbeSnapshot takes an optional rateMultiplier
  and applies it atomically with the snapshot under the same
  identity/snapshot compare-and-swap, so a probe result observed on a
  stale account cannot clobber a concurrent admin edit
- admin edit goes through UpdateWithAccountBillingSettings, which
  applies the form without overwriting a rate that a probe synchronized
  after the edit form was loaded (nil rateMultiplier = not edited);
  once sync is enabled the edit form shows the rate as managed
- bulk update rejects a manual rate_multiplier change when any target
  account has rate sync enabled (whole batch fails with a dedicated
  error so partial writes cannot bypass the sync ownership)
- frontend: sync toggle with hints in the edit modal (probe/sync
  enable/disable coupling enforced in the form), synced-rate tooltip on
  the rate cell, bulk edit modal warns and blocks rate edits that hit
  sync-enabled accounts; en/zh copy updated
- tests: service unit tests for sync gating/validation/cascade, sqlmock
  repo tests for the extended CAS, real-PostgreSQL integration tests
  (rate written only for successful+enabled accounts, manual rate
  protected after sync disabled, admin edit preserved across concurrent
  probe sync), handler/API contract updates, frontend specs for modal
  coupling, bulk rejection and rate cell
2026-08-01 22:11:09 +08:00
shaw 56f3d3c9b0 fix(billing-probe): 收口探测资格放宽后的抑制清单与调度信任面
#4896 把上游计费探测从 OpenAI API-key 放宽到全部 API-key 账号后,
遗留了四处需要收口的问题:

- M1 官方域抑制清单补 ollama.com。Ollama Cloud 是本仓一等支持配置
  (platform openai/anthropic + type apikey + base_url
  https://ollama.com/v1),放宽后 anthropic 侧这类账号会每个探测周期
  拿 Ollama Key 请求 ollama.com/v1/sub2api/billing 并恒定落空,正是该
  守卫注释声明要防的行为。后缀匹配已覆盖 www.ollama.com 等子域,
  notollama.com 等形似域不受影响。
- M2 legacy 低倍率优先排序补平台门控。newOpenAILegacyUpstreamRateOrder
  遍历全部候选且无平台门控,与 openAIUpstreamCostFactors 的门控不对称;
  放宽后 grok 账号的上游自报倍率开始影响 legacy 调度排序,而实际结算走
  本地倍率,中转方自报低价即可吸流量。现补上同一道门控,使调度侧信任面
  回到 PR 前状态(探测资格的放宽保持不变)。
- L1 修正 IsUpstreamBillingProbeIdentity 注释。原注释称类型限制的依据是
  "OAuth/Bedrock 没有静态 API key",但 AccountTypeUpstream(antigravity
  中转账号)同样是 base_url + 静态 api_key 却也被排除。仅改注释如实说明
  取舍,不改行为。
- L2 批量探测空选文案去掉 OpenAI 限定(en/zh 成对)。
- L3 unsupported 状态改用加长退避(interval 的 8 倍,仍按 24h 封顶)。
  放宽后大量官方域账号会落 unsupported 并按常规 interval 重排,占满每周期
  20 个名额,把真正接入 sub2api 的中转账号挤到后面。封顶保证上游后来接入时
  最迟一天内会被重新发现;Retry-After 更长时原样保留不被缩短;手动探测不受
  退避影响。

测试:ollama.com 官方域行为级与 host 匹配矩阵用例、legacy 排序平台门控
(含混合候选集)用例、unsupported 退避上下界与 runner 跳过/手动探测放行
用例;同步更新既有 unsupported 的 next_probe_at 断言。
2026-08-01 21:54:11 +08:00
Wesley Liddick b74024c786 Merge pull request #5167 from Wei-Shaw/fix/openai-ws-passthrough-close-frame-race
fix(openai-ws): keep downstream writes off the relay cancellation context
2026-08-01 20:47:42 +08:00
shaw 21aacde0b3 fix(openai-ws): keep downstream writes off the relay cancellation context
coder/websocket arms a context.AfterFunc that hard-closes the connection
when a write context is canceled, and AfterFunc stop does not wait for a
callback that already started. An external cancellation (e.g. ingress
lease loss) landing inside the disarm window of an already-successful
downstream write could therefore kill the TCP connection before the
retry close frame (1013) was written, leaving the client with a bare
EOF. Mirror the read side: bound downstream writes with the write
timeout only, and rely on the explicit Close/CloseNow performed by every
relay exit path for teardown.

Fixes the flaky TestPassthroughLifecycle_LeaseLossSendsRetryClose.
2026-08-01 20:32:10 +08:00
Brisbanehuang f3a3d86845 feat(billing-probe): extend upstream billing probe to all API-key platforms
/v1/sub2api/billing is a key-scoped sub2api convention: any API-key
account whose base_url points at a sub2api-compatible upstream answers
it regardless of the account platform. Widen probe eligibility from
platform=openai to every API-key account (OAuth/Bedrock stay excluded:
no static key to present).

- central predicate exported as IsUpstreamBillingProbeIdentity; runner,
  manual probe, SetAccountEnabled, admin create/update/bulk validation
  and CRS reconcile all follow it
- probe target resolution reads credentials.api_key/base_url directly.
  OpenAI keeps its official-default base URL and openai transport
  profile. Other platforms whose base_url is empty or points at an
  official provider API domain (anthropic.com, googleapis.com, x.ai,
  grok.com, openai.com - matched on the normalized hostname, port and
  trailing dot stripped, as the exact host or any subdomain) persist
  "unsupported" without sending a request: the create form fills empty
  base_url with official defaults and offers official regional presets
  (e.g. us-east-1.api.x.ai), and official APIs cannot answer
  /v1/sub2api/billing, so probing would only send the account key to a
  nonexistent official path
- due-scan SQL, BulkUpdate probe WHERE, UpdateCredentials stale-snapshot
  CASE and proxy-change invalidation drop their platform filters
  (type='apikey' retained)
- frontend: rate cell, edit/create/bulk modals gate on type==='apikey';
  the antigravity upstream create flow (its own helper and form section)
  shows the auto-probe toggle and passes upstream_billing_probe_enabled;
  bulk WS-mode section stays OpenAI-only; settings copy de-scoped
- tests: multiplatform service unit tests (relay success, official/empty
  base_url unsupported without request, normalized official-host matrix,
  OpenAI defaults preserved), real-PostgreSQL due-scan coverage for
  openai/anthropic/grok plus oauth/disabled exclusion, updated
  sqlmock/CRS/frontend specs; probe stays opt-in per account
2026-08-01 02:31:04 -04:00
Wesley Liddick b22f73e725 Merge pull request #5154 from Wei-Shaw/fix/issue-5148-stream-partial-usage-billing
fix(gateway): 流中断时保留已观测 usage 入账,修复 newapi 类上游大面积漏记(#5148)
2026-08-01 13:51:06 +08:00
shaw d6d53052f8 chore: update sponsors 2026-08-01 11:32:31 +08:00
shaw bd52e5d770 fix(gateway): record observed usage when anthropic stream is interrupted
Fixes #5148: with aggregator upstreams (e.g. newapi) that end SSE
streams without a proper terminal event, every such request was
silently missing from usage logs and billing.

Root cause (tracked via the nested audit issue): the low-level
Anthropic SSE readers already return the partially collected usage
together with the stream error (missing terminal event, read error,
interval timeout), but Forward converted every such result to
(nil, err) and the handler returned before submitting RecordUsage.

Changes:
- Add partialStreamUsageResult: on stream errors, wrap observed usage
  into a ForwardResult and return it alongside the error, for both the
  regular Anthropic path and the API-key passthrough path. Invariants:
  UpstreamFailoverError always keeps result=nil (failover retries are
  billed as the successful attempt, never twice), and zero observed
  usage returns no partial result (no phantom zero-usage records).
- Messages handler: hoist the usage submission block into a closure
  shared by the success path and the new partial-result error path.
- Usage record worker pool: distinguish pool-stopped drops
  (dropped_stopped) from operator-configured drop/sample overflow
  drops; billing tasks now fall back to inline synchronous execution
  only during the shutdown window, while explicit drop/sample overflow
  semantics are preserved. Image usage keeps its mandatory fallback
  for both drop kinds via the new mode.Dropped() helper.

Tests: Forward-level regressions for missing-terminal / read-error /
no-usage / failover-invariant on both paths, plus handler-level
stopped-pool sync fallback and drop-policy preservation tests.
2026-08-01 11:29:58 +08:00
Wesley Liddick d4cada3b6b Merge pull request #5089 from feeeei/fix/openai_sse_rate_limit
fix(openai): retry SSE rate limits as HTTP 429
2026-08-01 10:50:57 +08:00
Wesley Liddick 8f5caef78e Merge pull request #5153 from Wei-Shaw/fix/issue-5152-classifier-multi-system-entries
fix(anthropic): recognize classifier requests with extra system entries
2026-08-01 10:50:27 +08:00
shaw 2ef1246295 fix(anthropic): recognize classifier requests with extra system entries
Real claude-cli/2.1.220 auto-mode classifier requests carry two system
entries: the security-monitor prompt plus an appended session-context
block. The previous len(systemEntries) != 1 guard rejected them before
any content check ran, so claude_code_only groups kept refusing the
classifier (#5152, follow-up to #5041/#5048).

Scan every entry for the monitor prompt instead of requiring exactly
one. Discrimination is unchanged: the matching entry still needs the
10k-char minimum, the fixed prefix, and all eight markers.
2026-08-01 10:32:43 +08:00
shaw dd9a177a62 chore: update sponsors 2026-08-01 10:01:52 +08:00
feeeei 85a27fae39 fix(openai): retry SSE rate limits as HTTP 429
OpenAI Responses may return code=rate_limit_exceeded in a response.failed SSE event while the HTTP status remains 200. Classify these failures as 429 so configured pool-mode retries and account failover are applied. No account-level rate-limit state is written on this path: the 200-stream response headers carry normal quota snapshots, and retry semantics stay owned by the failover engine.

Ignore compact keepalive bytes when determining whether semantic output has started, preserving safe retries before real output.
2026-08-01 10:01:47 +08:00
Wesley Liddick eb1c5c7ee8 Merge pull request #5146 from tudoujunha/codex/fix-responses-tool-output-media
fix(apicompat): preserve images in Responses tool outputs
2026-08-01 08:54:54 +08:00
Wesley Liddick d9fba8fe78 Merge pull request #5101 from Tongzai123/feat/admin-select-all-filtered-results
feat(admin): 支持按筛选结果全选账号
2026-08-01 08:54:38 +08:00
Wesley Liddick 15b3c0c5aa Merge pull request #5145 from zvensmoluya/codex/update-auto-review-pricing
[codex] update Codex Auto-review pricing
2026-08-01 08:54:26 +08:00
Wesley Liddick 2e338af822 Merge pull request #5085 from feeeei/main
feat(model-plaza): Filter bar line alignment, model sorting, and table spacing
2026-08-01 08:54:11 +08:00
Wesley Liddick 682c4fe0e6 Merge pull request #5147 from Wei-Shaw/feat/moderation-proxy-and-smtp-starttls
feat(moderation): proxy support for content audit; fix(email): SMTP STARTTLS test/send parity
2026-07-31 23:24:44 +08:00
shaw 948b63c9ca feat(moderation): route content moderation through configurable proxy server
Implements #2646: the risk-control content audit can now send OpenAI
Moderations requests through a proxy from IP Management - Proxy Servers.

Backend:
- ContentModerationConfig gains proxy_id (nil = direct, unchanged default)
- update semantics: null keeps, 0 clears, >0 selects (validated to exist)
- moderation calls build the client via the shared httpclient pool; proxy
  resolution failure surfaces as a moderation error and never silently
  falls back to direct connection
- proxy_id -> URL resolution cached 60s (single-entry, invalidated on
  config save) so the pre-block hot path does not hit the DB per request
- test-key endpoint accepts proxy_id too (null = saved config's proxy,
  0 = force direct), so input-key/saved-key tests exercise the same path
- proxy usage/inactivity logged (content_moderation.proxy_enabled /
  proxy_not_active) without leaking credentials

Frontend:
- ProxySelector in the risk-control basic settings tab, proxy list loaded
  non-blockingly; save and test payloads carry proxy_id; zh/en i18n
2026-07-31 23:12:22 +08:00
shaw 4c80d160dd fix(email): unify SMTP connection path between send and test-connection
- UseTLS now tries implicit TLS first (port 465 semantics) and, when the
  server answers in plaintext (tls.RecordHeaderError, e.g. port 587
  submission), automatically retries with mandatory STARTTLS; encryption
  is never silently downgraded (fixes #1470, supersedes #1488)
- TestSMTPConnectionWithConfig now shares connectSMTP with the send path,
  adding the opportunistic STARTTLS upgrade the send path gained in
  b402c367d; this removes the 'test connection fails but test email
  sends' mismatch reported in #1488
- test-connection now also honors dial/IO timeouts and ignores
  non-standard QUIT responses, matching the send path
2026-07-31 23:12:02 +08:00
Wesley Liddick 570ea74d12 Merge pull request #5117 from gaoren002/feat/prompt-audit-blocking-latest-input
feat(security-audit): add optional narrow blocking audit scope
2026-07-31 22:32:00 +08:00
Wesley Liddick 2980ff3850 Merge pull request #5094 from wucm667/feat/issue-5065-compact-homepage
feat(home): add compact home page preset to avoid abuse classification
2026-07-31 21:51:33 +08:00
Wesley Liddick 04c96a2015 Merge pull request #4981 from INKCR0W/fix/openai-preserve-codex-namespaces
fix(openai): OAuth 原生 Responses 默认保留 Codex namespace,修复 code_mode_only 模型无法派发子代理
2026-07-31 21:49:44 +08:00
Wesley Liddick 07f980b99f Merge pull request #5084 from apple-ouyang/codex/fix-openai-compaction-encrypted-retry
fix(openai): recover stale encrypted compaction
2026-07-31 21:48:10 +08:00
Ouyang Xingyuan fe21725865 fix(openai): recover stale encrypted compaction
Reason:
- Responses retries can carry account-bound encrypted compaction items that OpenAI rejects with invalid_encrypted_content.

Changes:
- Drop encrypted compaction and compaction_summary items only during the existing recovery retry.
- Preserve unencrypted compaction items and cover HTTP and WebSocket recovery paths.
2026-07-31 21:21:13 +08:00
Zven 698547418f fix(pricing): keep Auto-review rates evidence-based 2026-07-31 21:20:32 +08:00
Zven f54e9827a0 fix(pricing): update Codex Auto-review rates 2026-07-31 21:16:06 +08:00
Wesley Liddick d29acc29a5 Merge pull request #5066 from wucm667/fix/issue-5051-subscription-quota-window
fix(subscription): align quota windows with subscription term
2026-07-31 20:40:18 +08:00
Wesley Liddick 66998918b6 Merge pull request #5143 from wucm667/fix/issue-5138-codex-instructions
fix(openai): default missing passthrough instructions
2026-07-31 20:39:57 +08:00
Wesley Liddick da6194c1c3 Merge pull request #5112 from chenty2333/fix/openai-stream-capacity-pool-retry
fix(openai): retry streamed capacity errors in pool mode
2026-07-31 20:38:15 +08:00
Wesley Liddick 132d446ca9 Merge pull request #5133 from dawnx/fix/payment-visible-method-wipe
fix(payment): 保存系统设置时不再清空可见支付方式配置
2026-07-31 20:37:58 +08:00
Wesley Liddick 0eac363e67 Merge pull request #5120 from Vibeone/fix/grok-pool-mode-cooldown-bypass
fix(grok): 公共池模式跳过所有默认冷却路径
2026-07-31 20:27:46 +08:00
Wesley Liddick 796313e993 Merge pull request #5131 from wucm667/fix/issue-5125-image-data-url-offload
fix(images): decode data URLs during task offload
2026-07-31 20:27:35 +08:00
Wesley Liddick c772d18666 Merge pull request #5130 from moonfunjohn/codex/fix-epay-method-selector-overflow
fix(payment): prevent EasyPay method selector overflow
2026-07-31 20:27:25 +08:00
tudoujun 2bf9c6d56b 修复工具输出图片桥接 2026-07-31 19:36:47 +08:00
Wesley Liddick 94df1fffc2 Merge pull request #5124 from wucm667/fix/issue-5105-filter-grok-billing-ping
fix(grok): filter billing ping response events
2026-07-31 19:20:16 +08:00
shaw 30967d5d9a fix(grok): ping 帧统一改写为 SSE 注释并限制过滤缓冲
Responses 事件类型对严格客户端是闭合枚举,任何 event: ping 帧都会令
grok CLI / Codex CLI 整轮失败。原实现只精确匹配 inference-cost 标记帧
和 cost=="0" 帧,上游尾帧携带非零 cost 或格式微调即复发 #5105;且对
其它 ping 变体的保守放行同样会炸掉严格解析器。现改为:event: ping 帧
(data 声明的 type 与事件名不冲突时)一律改写为 SSE 注释 ": ping",
所有解析器安全忽略且保留保活效果。

同时把整帧缓冲改为增量状态机:非 ping 帧首行即判定、逐行零拷贝直通,
不再累积;仅 ping 候选帧缓冲,并设 16 行 / 16KB 上限,超限回放原文
转直通,杜绝上游用永不结束的帧撑爆网关内存。
2026-07-31 18:57:43 +08:00
wucm667 dfdbc27709 fix(openai): default missing passthrough instructions 2026-07-31 18:55:15 +08:00
wucm667 beeb2f989b test(settings): include compact home in API contracts 2026-07-31 18:34:13 +08:00
wucm667 77d4df9544 test(grok): check filter body close error 2026-07-31 18:34:11 +08:00