Commit Graph
5492 Commits
Author SHA1 Message Date
shaw c899c8cf37 fix(codex): 管理员配置的 UA 只贡献指纹,版本段一律用生效版本重建
面板「OpenAI Codex UA」此前只在客户端 UA 是浏览器型时才生效,本分支把它提升为
所有 OAuth 出站的规范身份来源,作用域扩大了一个数量级。而它的旧 placeholder 原文
就是 codex_cli_rs/0.144.1 (...):照抄填写过的存量部署会被永久钉在 0.144.1 上——
该值不低于上游门槛 0.144.0,version 头也随之变成 0.144.1,绕过版本自动同步,
稳定落在上游优先降载的那一侧,且面板、日志、审计都不提示。账号级自定义 UA
(credentials.user_agent)有完全相同的陷阱。

改为:管理员配置的 UA 只贡献客户端名与 OS / 架构 / 终端指纹(这是该输入框唯一
不可替代的价值),版本段一律用当前生效版本重建。存量的陈旧配置无需迁移即自愈,
UA 与 version 头从此由构造保证同源,原先「覆写 UA 陈旧时二者不一致」的次优解消失。

- 新增 openai.SetCodexUserAgentVersion:重建首段版本声明,OS / 架构 / 终端指纹
  原样保留;尾部官方客户端标识组 (name; version) 与首段同源,一并更新,避免拼出
  首段声明新版本、尾部仍是旧版本的自相矛盾身份;非官方括号组(如 OS 组)不动。
- resolveCodexOutboundIdentity 收敛为单一funnel:生效版本只从规范身份取,
  候选 UA 的版本段不再参与判定。
- GetOpenAICodexCanonicalUserAgent 重建面板 UA 的版本段;原「面板值等于兜底常量
  视同未填」的特例随之成为恒等变换,删除。
- 补齐 4 处静默丢弃账号级自定义 UA 的出站路径(WS 握手、账号测试 x2、用量探针):
  它们先写 customUA 再调不带 override 的收口,赋值随即被覆盖,既是行为不一致
  也是死代码;账号测试尤其要紧——注释写着「与真实转发一致」而实际不一致。
- 补测试:SetCodexUserAgentVersion / CodexUserAgentVersion 包内用例、陈旧面板 UA
  与账号 UA 的版本重建回归、WS 握手尊重账号级 UA。
2026-08-03 21:50:49 +08:00
zhiyu 4c4ff36380 perf(codex): 版本同步间隔改为 6 小时并加启动防抖
- 同步间隔 3h → 6h:客户端版本是天级变化,6 小时足够跟上,
  对 GitHub 的调用降到每天 4 次。
- 新增启动防抖:借同步设置行自身的 UpdatedAt 判断,若同步值仍在一个周期内
  则跳过启动同步。频繁重启、滚动发布或崩溃重启原本会把「启动即同步」
  放大成对 GitHub 的连续请求;首次部署尚无同步值时不受影响。
- 补测试:启动防抖两个方向,以及版本比较按段取数字的回归
  (字典序会把 0.99.0 判为大于 0.146.0,同时让「取最大值」与
  「只向前推进」两处判错)。
2026-08-03 20:55:37 +08:00
zhiyu 1e08c4c56f test(server): 补齐 settings 契约用例的 Codex 版本号字段
新增的三个设置键会出现在 GET /api/v1/admin/settings 响应里,
契约用例的期望 JSON 需同步,否则 -tags=unit 下断言失败。
2026-08-03 20:26:50 +08:00
zhiyu 2eb24814fe fix(codex): 强制统一出站身份并让客户端版本号跟随官方发布
上游 /backend-api/codex 在容量紧张时按客户端身份分优先级降载,被降载的请求
HTTP 200 后立刻推流内 server_is_overloaded。此前网关对配不出官方身份的客户端
整体回退到硬编码的 codex_cli_rs/0.144.1(落后官方 4 个发布),这些请求稳定
落在被优先丢弃的一侧。

- 强制统一出口:所有 OAuth 出站的 User-Agent / originator / version 一律改写
  为网关规范身份,客户端自报身份不参与构造;HTTP / 透传 / WS / alpha-search /
  探针全覆盖。compat 桥接故意删除 originator 的路径保持 no-op。
- 版本号收敛为单一来源,运行时优先级为面板覆写 → 自动同步值 → 内置常量;
  UA 与 version 头同源派生,不再各自硬编码。
- 新增 3 小时自动同步官方客户端最新稳定版,面板可关闭,无需为跟版本而发版。
- 流内 server_is_overloaded / slow_down 改为先在同账号有界重试再切号,并标记为
  请求级瞬时故障,不再据此临时封禁账号。
- 移除被取代的降载身份黑名单、浏览器 UA 兜底及其辅助函数。
2026-08-03 20:14:58 +08:00
Wesley Liddick 825ca7b1fc Merge pull request #5183 from rick147/codex/feat-openai-reset-credit-cache
feat(openai): refresh reset credit state after quota reset
2026-08-03 16:01:10 +08:00
shaw 54a2bcfd15 fix(openai): harden reset-credit refresh and account recovery
Review follow-ups on the reset-credit caching flow:

- Recover account state BEFORE (and independently of) the reset-credit
  display cache. A failed cache refresh could previously abort the run and
  leave the account rate-limited — the very reason the credit was spent
  (#3672 / #3740). The recovered account row is now returned even when the
  cache refresh fails.
- Run the post-reset bookkeeping on a detached, time-boxed context and give
  the panel reset call a larger timeout. A client abort no longer strands a
  consumed (non-refundable) credit with an unrecovered account, and the
  chained upstream calls can no longer exceed the client timeout and invite a
  retry that spends a second credit.
- Persist the reset-credit snapshot through POST /accounts/:id/quota/refresh
  instead of a side-effecting GET flag, so the write is covered by the audit
  middleware. A rejected snapshot write now degrades to cache_persisted=false
  instead of turning a successful upstream read into a 502 that left the card
  without a credit count and the reset button permanently disabled.
- Reject snapshots whose positive count carries no expiration entries, and
  drop expired credits (clamping the count) when rehydrating, so a stale
  cache can no longer light up the reset button.
- Keep nil quota / rate-limit services nil in the handler's interface fields;
  storing a nil *Service made the "not enabled" guards non-nil.
- Time-box the usage-refresh suppression and reuse handleAccountUpdated so the
  patched row also enters the auto-refresh silent window.
2026-08-03 14:40:55 +08:00
Wesley Liddick 27e8f69a9e Merge pull request #5171 from heathermhuang/codex/composite-reasoning-policy
feat(composite): enforce reasoning effort policy
2026-08-03 11:28:38 +08:00
Wesley Liddick 684ab20a0b Merge pull request #5164 from wucm667/fix/issue-5099-messages-temp-failover
fix(openai): fail over Messages temporary account errors
2026-08-03 11:28:16 +08:00
Wesley Liddick 0173830df6 Merge pull request #5193 from wucm667/fix/issue-5187-stripe-refund-idempotency
fix(payment): make Stripe refunds idempotent
2026-08-03 11:28:05 +08:00
Wesley Liddick 724565e4aa Merge pull request #5199 from wucm667/fix/issue-5191-refund-balance-force
fix(payment): require force for insufficient refund balance
2026-08-03 11:27:52 +08:00
Wesley Liddick a03de418c8 Merge pull request #5192 from luckydududu/up/admin-refund-require-force
fix(admin-ui): 退款 require_force 前端接线,补上无法完成的退款场景
2026-08-03 11:27:37 +08:00
Wesley Liddick 61ebdbdd43 Merge pull request #5200 from mrlitong/fix/auth-refresh-race
fix(auth): prevent refresh token races across tabs
2026-08-03 10:46:36 +08:00
Wesley Liddick a1f1a0cc6b Merge pull request #5194 from wucm667/fix/issue-5189-persist-unsettled-usage
fix(billing): retain usage logs on billing failure
2026-08-03 10:27:30 +08:00
Wesley Liddick 954d44c19c Merge pull request #5198 from Wei-Shaw/fix/codex-originator-load-shed
fix(codex): 归一化降载 originator,缓解 Codex 账号频繁过载不可用
2026-08-03 09:59:34 +08:00
litongtongxue@gmail.com 38081ef72e fix(auth): prevent refresh token rotation races 2026-08-02 08:13:15 -07:00
wucm667 3c20f9a666 [verified] fix(payment): require force for insufficient refund balance 2026-08-02 23:11:17 +08:00
shaw e1b76e2245 fix(codex): normalize load-shed originators to avoid upstream capacity shedding
上游 /backend-api/codex 按 Originator 头分桶调度容量:落在降载桶的请求即使返回
HTTP 200,也会立刻推 SSE `event: error`(code=server_is_overloaded)并以
response.failed 收尾。2026-07-29 起 codex-tui 落入降载桶,codex_cli_rs 正常——
判定因子是 originator 而非 User-Agent(codex_cli_rs 配 curl UA 亦可正常返回)。

网关会把该错误判定为瞬时上游故障并冷却账号,对外表现为 Codex 账号频繁过载不可用:
server_is_overloaded → isOpenAITransientProcessingError →
shouldCooldownOpenAITransientUpstreamError → 账号冷却 → 客户端 503。

本项目有三处降载身份来源:浏览器 UA 兜底的默认 UA、客户端透传的真实 TUI 身份、
以及指纹缓存注入探针的 UA。

修复收口在 enforceCodexIdentityHeaders——HTTP / 透传 / WS 握手 / compat 桥接 /
探针 / PAT / 模型列表 / alpha-search 八条出站路径共用的唯一纯函数收口点:

- 新增 NormalizeCodexClientIdentityToCLI,把降载桶身份改写为 codex_cli_rs,
  只替换身份段并裁掉尾部 (name; version) 客户端标识组,保留版本 / OS / 架构 /
  终端指纹;改写后 originator 与 UA 首段仍然配套,不破坏 #3901 的配对不变式,
  且改写幂等。
- DefaultOpenAICodexUserAgent 从 TUI 身份改为 CLI 身份(浏览器兜底路径上最大的
  降载身份来源)。
- 管理端 Codex UA 的 placeholder / hint 原本在把管理员往降载桶引导,一并修正。

新增 gateway.disable_codex_originator_normalization(默认 false,即归一化开启),
供上游调整分桶后回滚。该开关经 NewOpenAIGatewayService 发布为进程级快照,故必须
保持反义命名:正向命名的 Go 零值 false 会让未经 viper 加载而手工构造的 Config
静默关掉全局保护,viper.SetDefault 救不了这条路径。已加用例钉住该属性。

降载桶集合是上游容量策略快照而非协议常量,上游调整分桶后需同步修订。
2026-08-02 23:00:12 +08:00
rick147 f970bd48c9 fix: stop scheduler work after request cancellation 2026-08-02 21:32:31 +08:00
rick147 a0802f00b6 feat: cache OpenAI reset credit details 2026-08-02 21:32:31 +08:00
wucm667 0b9f40e230 fix(billing): retain usage logs on billing failure 2026-08-02 20:50:13 +08:00
wucm667 0b26acac07 fix(payment): make Stripe refunds idempotent 2026-08-02 20:27:44 +08:00
github-actions[bot] 7e2e9ba050 chore: sync VERSION to 0.1.170 [skip ci] 2026-08-02 10:46:21 +00:00
Lucky 2a42833c4c fix(admin-ui): wire require_force into the refund dialog
AdminRefundDialog already implements a `requireForce` prop that renders the
force checkbox and blocks submit until it is checked, but AdminOrdersView
never consumed `require_force` from the refund response nor bound the prop,
so the checkbox could never render. Admins hitting a require-force case (for
example a user who spent their balance after requesting the refund) only saw
an error toast and had no way to complete the refund from the UI.

Keep the dialog open when the backend answers `require_force`, surface the
checkbox and the backend warning, and reset the flag whenever the dialog is
opened or closed.

Verified with `npm run typecheck` (vue-tsc --noEmit, clean).
2026-08-02 10:31:15 +00:00
Wesley Liddick c043c24774 Merge pull request #4925 from Brisbanehuang/feat/group-profit-control
feat(scheduler): 分组级利润控制——按账号倍率过滤 token 调度候选
v0.1.170
2026-08-02 18:19:13 +08:00
Wesley Liddick 11c1e944b9 Merge pull request #4911 from Brisbanehuang/feat/upstream-billing-rate-writeback
feat(billing-probe): 探测成功后可选将上游声明倍率自动同步为账号倍率
2026-08-02 18:18:50 +08:00
Wesley Liddick d99ee72911 Merge pull request #4896 from Brisbanehuang/feat/upstream-billing-probe-multi-platform
feat(billing-probe): 上游计费倍率探测放宽到全部 API-key 平台账号
2026-08-02 18:18:37 +08:00
shaw dec47e8fae fix(profit-control): stop leaking profit policy, close veto livelock, restore passthrough turn pricing
审计修复,逐条如下。

H1 利润策略泄露给所有普通用户
  profit_control_enabled / profit_min_margin / profit_safety_buffer 从
  dto.Group 移到 dto.AdminGroup(后者内嵌前者),赋值相应从
  groupFromServiceBase 移到 GroupFromServiceAdmin;前端 TS 同步从 Group 移到
  AdminGroup。dto.Group 是 GET /api/v1/groups/available 的响应体,该响应本就带
  rate_multiplier,相乘即可反推运营方上游采购成本上限。
  api_contract_test.go 的 /groups/available golden JSON 回滚这三个字段,并把
  fixture 改成非零值(require.JSONEq 是精确比对,缺字段即失败)。
  新增 dto 层边界测试:普通用户 DTO 不含三字段、管理员 DTO 仍含。

M1 利润终检 continue 与 failover 503 退避互动产生活锁
  FailoverState 新增 profitVetoedAccountIDs / profitVetoCount 与
  RecordProfitVeto():加入排除集 + 计数,达 maxProfitVetoAttempts(10) 返回
  FailoverExhausted。HandleSelectionExhausted 的 503 清空分支改为清空后把利润
  否决的账号放回排除集;若排除集已全部由利润否决贡献,清空不会带来任何新候选,
  直接判定耗尽(否则 SwitchCount 永不前进、退避条件永远成立,每 2s 空转一轮)。
  五个 handler 否决点(gateway_handler ×2 / responses / chat_completions /
  gemini_v1beta)改为经 RecordProfitVeto 决策,耗尽时按无可用账号终止。
  回归测试钉死:503 之后持续利润否决必须有限步终止且不 spin;未启用利润控制的
  请求退避语义完全不变。

M2 排队等槽后才终检,延迟可放大到 N × WaitPlan.Timeout
  OpenAI 侧选号循环(自有 failedAccountIDs map,非 FailoverState)新增
  recordOpenAIProfitVeto + handleOpenAIProfitVetoExhausted,共用同一上限语义。
  覆盖 responses / messages-dispatch / chat_completions / alpha_search /
  embeddings / images / grok_media 七处,以及 WS 两处否决分支。

M4 ws_v2 透传 ingress 绕过 per-turn 重定价(选方案 B:最小止血)
  透传 relay 只回调 AfterTurn、没有任何 turn 起始回调,hooks.BeforeTurn 永远
  不触发,而 handler 把 turnPricingAt 初始化成建连时刻 ⇒ 透传连接全部 turn 按
  建连时刻的高峰因子结算,客户端峰前建连保活即可全程谷价——正是本 PR 想堵的
  漏洞。改为 openAIWSTurnPricing 零值起步、只由 BeforeTurn 冻结;透传路径保持
  零值,RecordUsage 回退记录时刻,与引入利润控制前的基线一致。
  未选方案 A(给透传补 turn 起始回调):passthrough_relay.go 是 #5167 刚修过的
  取消传播/close frame 时序敏感区;且 BeforeTurn 还承担 turn>1 的并发槽位抢占,
  接进去等于给透传连接引入 per-turn 抢槽,风险远超本次修复范围。透传仍有建连时
  的准入门,只是没有 turn 级复核,已在两处注释写明。
  测试:service 层钉死透传 ingress 不触发 BeforeTurn(含失败时的复核指引),
  handler 层钉死零值语义与逐 turn 覆盖。

M5 装门读分组走了带账号计数聚合的 GetByID
  SchedulerSnapshotService 新增 GetGroupByIDLite,openai/gateway 两处装门改用
  之。门只需要平台/倍率/利润/高峰字段,且该查询发生在「是否启用利润控制」判定
  之前,未启用的分组同样付代价。两个测试 stub 的 GetByID 改成 panic 守卫。

M6 认证快照注释与真实读取路径相反
  门解析优先取 ctxkey.Group,而它就是本快照物化出来的对象,直连流量走的正是这
  条路。改正注释,与 api_key_repo.go 投影处的说明对齐,避免后人照旧注释删列。

M3 rate_multiplier 为 nil 时利润门 fail-closed(不改行为,加护栏)
  保留 fail-closed。补 repository 层测试钉死账号调度快照的 full/metadata 两份
  payload 都必须保留 RateMultiplier(含 0 值),漏列在 CI 就红。

L1 迁移号注释 191 / 191-192 改为实际的 192/193。
L2 admin group Create 的利润配置预校验改用与 CreateGroup 一致的归一化平台
   (新增 service.NormalizeGroupPlatform,两边共用)。保留预校验而非删除:
   service 层返回的是无类型 error,经 ErrorFrom 会变成 500,删掉会把合法的
   400 降级成 500。
L3 前端利润校验的上界改为判定换算后的小数(后端按小数校验 [0,1)),
   99.999% 会四舍五入进位成 1.0 而被后端 400;i18n en/zh 同步改为 0-99.99。
L4 clampProfitControlThreshold / profitControlOverThreshold 抽为共用函数,
   线上装门/否决点与 profit-preview 不再各自实现,附边界语义测试。
L5 profit-preview 补「默认 D 有账号但最低有效 D 归零」的告警(两档都为 0 由
   既有告警覆盖,不重复)。
2026-08-01 22:39:33 +08:00
Brisbanehuang fad2f215e8 fix(profit-control): decouple gate from image intent, propagate gate via selection, restore eager sticky fallback
Review fixes for the profit-control feature commit:

- Image intent no longer disables the profit gate. The shared /v1/responses
  handler previously skipped the pricing context (and therefore the gate)
  whenever the platform-wide image intent predicate matched, which includes
  Codex's passive image_gen namespace declaration: any client could disable
  admission control for anthropic/gemini/antigravity groups by declaring a
  namespace tool in the request body. Both /v1/responses paths now always
  install the token pricing context; image intent only drives capability
  routing and image billing. Mixed token+image requests stay token-gated;
  only dedicated media endpoints remain out of scope.
- Out-of-scope paths are now explicitly suppressed instead of implicitly
  ungated: Grok media (billed by media multipliers; also prevents in-flight
  video lookups from turning into spurious 404s), OpenAI-group count_tokens
  (unbilled), and Live calls (duration billed) carry a suppress marker that
  every install point honors, including the defensive scheduler-entry
  install.
- The gate resolved during selection now travels back to handlers on the
  AccountSelectionResult. The shared gateway installed the gate only on a
  scheduler-local context, so handler-side post-slot terminal rechecks and
  post-admission sticky binding were no-ops for anthropic/gemini/
  antigravity/shared-grok requests (and for composite-routed member groups
  on the OpenAI path). Handlers re-apply the carried gate via
  ContextWithSelectionProfitGate before the terminal recheck and binding;
  the WS acquired-selection branch gained the previously missing recheck.
- Sticky binding semantics restored for ungated traffic:
  BindStickySessionAfterProfitAdmission falls back to the official eager
  bind when no gate is installed (wait paths lost their only binding point
  otherwise), reads the pre-existing binding at bind time only when gated
  (removes the unconditional per-request Redis read the feature added to
  the shared handlers), and the legacy engine's three selection-time
  binding writes are skipped under a gate so a terminally vetoed account
  can no longer become the new sticky target.
- Responses WS connections re-freeze pricingAt and re-resolve the gate at
  every turn (BeforeTurn) and bill each turn with its own instant, closing
  the connect-at-valley/bill-at-valley window; a turn that fails the
  recheck closes the connection so the client reselects on reconnect.
- Terminal recheck no longer swaps a DB-fresh selected account for a
  staler snapshot object (UpdatedAt guard), and observer counters are
  documented as per-evaluation.
- Migrations renumbered 191/192 -> 192/193 after the passkey migration
  landed upstream as 191.

New regressions: selection-carried gate propagation (control group proves
the pre-fix no-op), image intent not disabling the shared gate, eager
binding fallback without a gate (both services), gated
read-failure/sentinel-miss binding semantics, legacy-engine deferred
binding under a gate with eager behavior preserved ungated, turn-level
pricing refresh (config re-resolution, scheduled-group precedence,
suppress, mid-connection disable), and suppress-marker coverage for the
request pricing context.
2026-08-01 22:39:32 +08:00
Brisbanehuang 20ad5ec506 feat(scheduler): per-group profit control for token account admission
Group pricing (rate multiplier, peak windows, per-user overrides) and
account cost (accounts.rate_multiplier) already live side by side, but
nothing stops the scheduler from handing a request to an account whose
cost multiplier exceeds what the group's pricing can profitably serve.
Add an opt-in per-group profit gate that filters scheduling candidates
by a margin rule, while ordering, scoring, stickiness and breakers keep
working unchanged among qualified accounts.

Admission rule: an account qualifies iff U <= D * (1 - min_margin -
safety_buffer) within a small relative epsilon, where U is
accounts.rate_multiplier (0 is legal; missing/negative/NaN/Inf are
conservatively rejected as invalid) and D is the requester's effective
downstream multiplier (user-group override ?? group default, times the
group peak factor) frozen at the request's pricing instant.

- groups gain profit_control_enabled / profit_min_margin /
  profit_safety_buffer (migration 191); the durable auth-cache
  invalidation trigger additionally watches the profit and pricing
  columns (migration 192) so out-of-band group edits cannot leave
  stale auth snapshots; GetByKeyForAuth explicitly projects the new
  columns and the API-key auth snapshot version is bumped to force a
  refresh of pre-existing snapshots
- request-level pricing instant: token entry points install pricingAt
  into ctx; the profit threshold D and the RecordUsage peak factor
  read the same instant, so one request never changes price mid-flight
  across waits/retries/failover (media and unwired paths keep the
  existing record-time semantics)
- the gate covers token requests on openai, anthropic, gemini, grok
  and antigravity groups: OpenAI-family handlers via
  WithOpenAIRequestPricingContext (responses incl. WS bridge, chat
  completions, messages, embeddings, alpha search), the shared gateway
  via WithGatewayTokenRequestPricing (messages, chat completions,
  responses, gemini model actions); composite groups cannot enable it
  directly; image/video/models/usage/count_tokens stay ungated and an
  explicit image-generation intent suppresses the gate end to end
- post-slot recheck: after a slot is acquired the account is re-read
  via SchedulerSnapshotService.GetAccount (scheduler cache, then DB;
  only when both fail the check fails open with WARN + metric); a
  vetoed account releases its slot and joins the request's exclusion
  set for reselection; sticky bindings are written only after the
  final check passes, and an over-threshold sticky account is skipped,
  not deleted, so it comes back once its rate recovers
- sticky-session cache contract: GatewayCache.GetSessionAccountID now
  returns ErrStickySessionNotFound on a miss (mapped from redis.Nil in
  the repository implementation, mirroring ErrRefreshTokenNotFound) so
  the profit sticky path can distinguish "no binding yet" from a real
  read failure without importing the cache driver in service code
- cross-group re-entry (composite parent -> member group) resolves the
  gate against the member group and clears a stale parent gate instead
  of letting a foreign threshold veto accounts
- per-platform/group activity counters (installs, threshold vetoes,
  invalid-rate vetoes, refresh failures) for observability
- admin UI: profit-control section on the five platforms' group forms
  with percent input, validation and platform-switch reset; group
  create/update/duplicate normalize and validate the config at a
  single choke point
- cmd/profit-preview: offline what-if tool that replays the production
  admission semantics over an exported config/account/override/model
  dump, reports per-model admitted-account counts under the default
  and the worst-case (lowest user override) D, and surfaces probe-sync
  staleness as warnings without affecting admission

Tests: service unit coverage for gate resolution/veto/threshold
epsilon/pricing instant/suppress marker/scheduler filtering and
post-slot recheck (incl. -race on the profit surface), unit-tagged
handler slot-recheck and capability-mapping regressions, sqlmock and
real-PostgreSQL integration regressions for the GetByKeyForAuth
projection and the migration-192 trigger watch list, API contract
update, and frontend specs for the five-platform form helpers.
2026-08-01 22:39:31 +08:00
shaw 0b6b4ea956 fix(billing-probe): govern the automatic account rate write-back
- 自动写回值域治理与留痕:上游声明值必须 > 0 且 <= 100 才写回。0 会让
  accountCost 恒为 0(账号总额/日/周配额与成本告警全部静默失效),极大值可
  一次打爆配额并污染成本报表;越界时保持原倍率、记 WARN,探测快照照常记 ok。
  写回成功时记结构化 slog(account_id / 旧值 / 新值 / source),并在快照中
  新增 synced_rate_multiplier 记录本次写回值——后台任务裸 SQL 不产生
  audit_logs,这两处是唯一可追溯来源。管理员手工设 0 不受影响。
- 写回改用 resolved_rate_multiplier(不含高峰的基准倍率):effective 含探测
  那一刻的高峰系数,写回会把一个探测周期的峰值/谷值冻结进静态列,而展示与
  调度用 upstreamBillingRateAt 按当前时间重算高峰,两者会持续不一致。
- 倍率解析失败不再污染公共探测路径:账号级值域/精度只在该账号已开启同步、
  真要写回时才有影响;未开同步的账号照常记 ok 快照,不累计 failure_count、
  不进入指数退避。
- 单账号编辑补 service 层守卫:同步开启时拒绝手工倍率(新增
  UPSTREAM_BILLING_RATE_SYNC_CONFLICT,与批量路径同族),此前只有前端
  disabled 挡人,直接 PUT /admin/accounts/{id} 可写入并活到下次成功探测。
  判断的是本次请求生效后的状态,"关同步 + 改倍率"同请求仍然放行。
- 修正开关反推方向:不再由 rate_sync=true 推出 probe=true,否则一条"同步开、
  探测键缺失"的僵尸记录会在任意一次无关编辑时静默打开周期性外呼;改为探测
  关闭/缺失一律把同步归零。
- 删除死代码 UpdateWithUpstreamBillingProbeEnabled(PR 删接口后生产已无调用
  方),其回滚测试改为直接覆盖生产路径 UpdateWithAccountBillingSettings。
  upstreamBillingRateSyncEnabled 不再是只服务测试的假门控,现为写回前置过滤,
  SQL CAS 仍是权威门控,两侧均加注释说明分工。
- en/zh 文案补充:同步的是不含高峰的基准倍率;开启同步会连带打开自动探测。
2026-08-01 22:11:10 +08:00
Brisbanehuang b0f5007f04 feat(billing-probe): optionally sync account rate from upstream declared rate
Successful upstream billing probes already persist the upstream-declared
rate as a display-only snapshot. Add a per-account opt-in that writes
that declared rate back to the account's rate_multiplier, so the account
cost basis follows upstream repricing automatically instead of drifting
until an operator notices.

- new per-account flag upstream_billing_rate_sync_enabled stored next to
  the probe flag in account extra: enabling sync force-enables the
  probe, disabling the probe cascades sync off, and eligibility follows
  IsUpstreamBillingProbeIdentity (tightened from any non-empty platform
  to an explicit whitelist of the five supported API-key platforms so
  future platforms do not silently inherit probe/sync semantics)
- only a successful probe whose declared rate survives validation
  (finite, within bounds, not rounded to zero at the rate_multiplier
  decimal(10,4) scale) writes back; failed/unsupported/invalid probes
  leave rate_multiplier unchanged
- the writeback rides the existing snapshot CAS transaction:
  UpdateUpstreamBillingProbeSnapshot takes an optional rateMultiplier
  and applies it atomically with the snapshot under the same
  identity/snapshot compare-and-swap, so a probe result observed on a
  stale account cannot clobber a concurrent admin edit
- admin edit goes through UpdateWithAccountBillingSettings, which
  applies the form without overwriting a rate that a probe synchronized
  after the edit form was loaded (nil rateMultiplier = not edited);
  once sync is enabled the edit form shows the rate as managed
- bulk update rejects a manual rate_multiplier change when any target
  account has rate sync enabled (whole batch fails with a dedicated
  error so partial writes cannot bypass the sync ownership)
- frontend: sync toggle with hints in the edit modal (probe/sync
  enable/disable coupling enforced in the form), synced-rate tooltip on
  the rate cell, bulk edit modal warns and blocks rate edits that hit
  sync-enabled accounts; en/zh copy updated
- tests: service unit tests for sync gating/validation/cascade, sqlmock
  repo tests for the extended CAS, real-PostgreSQL integration tests
  (rate written only for successful+enabled accounts, manual rate
  protected after sync disabled, admin edit preserved across concurrent
  probe sync), handler/API contract updates, frontend specs for modal
  coupling, bulk rejection and rate cell
2026-08-01 22:11:09 +08:00
shaw 56f3d3c9b0 fix(billing-probe): 收口探测资格放宽后的抑制清单与调度信任面
#4896 把上游计费探测从 OpenAI API-key 放宽到全部 API-key 账号后,
遗留了四处需要收口的问题:

- M1 官方域抑制清单补 ollama.com。Ollama Cloud 是本仓一等支持配置
  (platform openai/anthropic + type apikey + base_url
  https://ollama.com/v1),放宽后 anthropic 侧这类账号会每个探测周期
  拿 Ollama Key 请求 ollama.com/v1/sub2api/billing 并恒定落空,正是该
  守卫注释声明要防的行为。后缀匹配已覆盖 www.ollama.com 等子域,
  notollama.com 等形似域不受影响。
- M2 legacy 低倍率优先排序补平台门控。newOpenAILegacyUpstreamRateOrder
  遍历全部候选且无平台门控,与 openAIUpstreamCostFactors 的门控不对称;
  放宽后 grok 账号的上游自报倍率开始影响 legacy 调度排序,而实际结算走
  本地倍率,中转方自报低价即可吸流量。现补上同一道门控,使调度侧信任面
  回到 PR 前状态(探测资格的放宽保持不变)。
- L1 修正 IsUpstreamBillingProbeIdentity 注释。原注释称类型限制的依据是
  "OAuth/Bedrock 没有静态 API key",但 AccountTypeUpstream(antigravity
  中转账号)同样是 base_url + 静态 api_key 却也被排除。仅改注释如实说明
  取舍,不改行为。
- L2 批量探测空选文案去掉 OpenAI 限定(en/zh 成对)。
- L3 unsupported 状态改用加长退避(interval 的 8 倍,仍按 24h 封顶)。
  放宽后大量官方域账号会落 unsupported 并按常规 interval 重排,占满每周期
  20 个名额,把真正接入 sub2api 的中转账号挤到后面。封顶保证上游后来接入时
  最迟一天内会被重新发现;Retry-After 更长时原样保留不被缩短;手动探测不受
  退避影响。

测试:ollama.com 官方域行为级与 host 匹配矩阵用例、legacy 排序平台门控
(含混合候选集)用例、unsupported 退避上下界与 runner 跳过/手动探测放行
用例;同步更新既有 unsupported 的 next_probe_at 断言。
2026-08-01 21:54:11 +08:00
Heatherm Huang d60f4e442b feat(composite): enforce reasoning effort policy 2026-08-01 21:15:44 +08:00
Wesley Liddick b74024c786 Merge pull request #5167 from Wei-Shaw/fix/openai-ws-passthrough-close-frame-race
fix(openai-ws): keep downstream writes off the relay cancellation context
2026-08-01 20:47:42 +08:00
shaw 21aacde0b3 fix(openai-ws): keep downstream writes off the relay cancellation context
coder/websocket arms a context.AfterFunc that hard-closes the connection
when a write context is canceled, and AfterFunc stop does not wait for a
callback that already started. An external cancellation (e.g. ingress
lease loss) landing inside the disarm window of an already-successful
downstream write could therefore kill the TCP connection before the
retry close frame (1013) was written, leaving the client with a bare
EOF. Mirror the read side: bound downstream writes with the write
timeout only, and rely on the explicit Close/CloseNow performed by every
relay exit path for teardown.

Fixes the flaky TestPassthroughLifecycle_LeaseLossSendsRetryClose.
2026-08-01 20:32:10 +08:00
wucm667 ddf4c6fd81 fix(openai): fail over Messages temporary account errors 2026-08-01 18:45:33 +08:00
Brisbanehuang f3a3d86845 feat(billing-probe): extend upstream billing probe to all API-key platforms
/v1/sub2api/billing is a key-scoped sub2api convention: any API-key
account whose base_url points at a sub2api-compatible upstream answers
it regardless of the account platform. Widen probe eligibility from
platform=openai to every API-key account (OAuth/Bedrock stay excluded:
no static key to present).

- central predicate exported as IsUpstreamBillingProbeIdentity; runner,
  manual probe, SetAccountEnabled, admin create/update/bulk validation
  and CRS reconcile all follow it
- probe target resolution reads credentials.api_key/base_url directly.
  OpenAI keeps its official-default base URL and openai transport
  profile. Other platforms whose base_url is empty or points at an
  official provider API domain (anthropic.com, googleapis.com, x.ai,
  grok.com, openai.com - matched on the normalized hostname, port and
  trailing dot stripped, as the exact host or any subdomain) persist
  "unsupported" without sending a request: the create form fills empty
  base_url with official defaults and offers official regional presets
  (e.g. us-east-1.api.x.ai), and official APIs cannot answer
  /v1/sub2api/billing, so probing would only send the account key to a
  nonexistent official path
- due-scan SQL, BulkUpdate probe WHERE, UpdateCredentials stale-snapshot
  CASE and proxy-change invalidation drop their platform filters
  (type='apikey' retained)
- frontend: rate cell, edit/create/bulk modals gate on type==='apikey';
  the antigravity upstream create flow (its own helper and form section)
  shows the auto-probe toggle and passes upstream_billing_probe_enabled;
  bulk WS-mode section stays OpenAI-only; settings copy de-scoped
- tests: multiplatform service unit tests (relay success, official/empty
  base_url unsupported without request, normalized official-host matrix,
  OpenAI defaults preserved), real-PostgreSQL due-scan coverage for
  openai/anthropic/grok plus oauth/disabled exclusion, updated
  sqlmock/CRS/frontend specs; probe stays opt-in per account
2026-08-01 02:31:04 -04:00
Wesley Liddick b22f73e725 Merge pull request #5154 from Wei-Shaw/fix/issue-5148-stream-partial-usage-billing
fix(gateway): 流中断时保留已观测 usage 入账,修复 newapi 类上游大面积漏记(#5148)
2026-08-01 13:51:06 +08:00
shaw d6d53052f8 chore: update sponsors 2026-08-01 11:32:31 +08:00
shaw bd52e5d770 fix(gateway): record observed usage when anthropic stream is interrupted
Fixes #5148: with aggregator upstreams (e.g. newapi) that end SSE
streams without a proper terminal event, every such request was
silently missing from usage logs and billing.

Root cause (tracked via the nested audit issue): the low-level
Anthropic SSE readers already return the partially collected usage
together with the stream error (missing terminal event, read error,
interval timeout), but Forward converted every such result to
(nil, err) and the handler returned before submitting RecordUsage.

Changes:
- Add partialStreamUsageResult: on stream errors, wrap observed usage
  into a ForwardResult and return it alongside the error, for both the
  regular Anthropic path and the API-key passthrough path. Invariants:
  UpstreamFailoverError always keeps result=nil (failover retries are
  billed as the successful attempt, never twice), and zero observed
  usage returns no partial result (no phantom zero-usage records).
- Messages handler: hoist the usage submission block into a closure
  shared by the success path and the new partial-result error path.
- Usage record worker pool: distinguish pool-stopped drops
  (dropped_stopped) from operator-configured drop/sample overflow
  drops; billing tasks now fall back to inline synchronous execution
  only during the shutdown window, while explicit drop/sample overflow
  semantics are preserved. Image usage keeps its mandatory fallback
  for both drop kinds via the new mode.Dropped() helper.

Tests: Forward-level regressions for missing-terminal / read-error /
no-usage / failover-invariant on both paths, plus handler-level
stopped-pool sync fallback and drop-policy preservation tests.
2026-08-01 11:29:58 +08:00
Wesley Liddick d4cada3b6b Merge pull request #5089 from feeeei/fix/openai_sse_rate_limit
fix(openai): retry SSE rate limits as HTTP 429
2026-08-01 10:50:57 +08:00
Wesley Liddick 8f5caef78e Merge pull request #5153 from Wei-Shaw/fix/issue-5152-classifier-multi-system-entries
fix(anthropic): recognize classifier requests with extra system entries
2026-08-01 10:50:27 +08:00
shaw 2ef1246295 fix(anthropic): recognize classifier requests with extra system entries
Real claude-cli/2.1.220 auto-mode classifier requests carry two system
entries: the security-monitor prompt plus an appended session-context
block. The previous len(systemEntries) != 1 guard rejected them before
any content check ran, so claude_code_only groups kept refusing the
classifier (#5152, follow-up to #5041/#5048).

Scan every entry for the monitor prompt instead of requiring exactly
one. Discrimination is unchanged: the matching entry still needs the
10k-char minimum, the fixed prefix, and all eight markers.
2026-08-01 10:32:43 +08:00
shaw dd9a177a62 chore: update sponsors 2026-08-01 10:01:52 +08:00
feeeei 85a27fae39 fix(openai): retry SSE rate limits as HTTP 429
OpenAI Responses may return code=rate_limit_exceeded in a response.failed SSE event while the HTTP status remains 200. Classify these failures as 429 so configured pool-mode retries and account failover are applied. No account-level rate-limit state is written on this path: the 200-stream response headers carry normal quota snapshots, and retry semantics stay owned by the failover engine.

Ignore compact keepalive bytes when determining whether semantic output has started, preserving safe retries before real output.
2026-08-01 10:01:47 +08:00
Wesley Liddick eb1c5c7ee8 Merge pull request #5146 from tudoujunha/codex/fix-responses-tool-output-media
fix(apicompat): preserve images in Responses tool outputs
2026-08-01 08:54:54 +08:00
Wesley Liddick d9fba8fe78 Merge pull request #5101 from Tongzai123/feat/admin-select-all-filtered-results
feat(admin): 支持按筛选结果全选账号
2026-08-01 08:54:38 +08:00
Wesley Liddick 15b3c0c5aa Merge pull request #5145 from zvensmoluya/codex/update-auto-review-pricing
[codex] update Codex Auto-review pricing
2026-08-01 08:54:26 +08:00
Wesley Liddick 2e338af822 Merge pull request #5085 from feeeei/main
feat(model-plaza): Filter bar line alignment, model sorting, and table spacing
2026-08-01 08:54:11 +08:00
Wesley Liddick 682c4fe0e6 Merge pull request #5147 from Wei-Shaw/feat/moderation-proxy-and-smtp-starttls
feat(moderation): proxy support for content audit; fix(email): SMTP STARTTLS test/send parity
2026-07-31 23:24:44 +08:00