Commit Graph
4284 Commits
Author SHA1 Message Date
wucm667 5ea03c178d fix(lint): resolve remaining nil dereference warnings 2026-08-11 16:57:39 +08:00
wucm667 1afb8264e5 fix(lint): use require.NotNil for staticcheck SA5011
Replace if-then-t.Fatal nil checks with testify require.NotNil so
staticcheck can prove the pointer is non-nil before field access.
Fixes SA5011 in payment_handler_test, parse_test, and
openai_images_incomplete_test.
2026-08-11 16:57:38 +08:00
wucm667 0b35370a7a fix(claude): support top-level deferred tools 2026-08-11 16:57:38 +08:00
wucm667 9c36b75a7d fix(claude): strip cache control from deferred tools 2026-08-11 16:57:38 +08:00
Wesley Liddick 1e618dbc29 Merge pull request #5054 from wucm667/fix/issue-5029-openai-passthrough-pool-auth-retry
fix(openai): retry pool auth failures before failover
2026-08-11 14:21:45 +08:00
shaw a3bbf35cbd Merge branch 'main' into fix/issue-5029-openai-passthrough-pool-auth-retry
Resolve conflict in backend/internal/handler/openai_gateway_handler_test.go.

main and this branch each appended a passthrough upstream stub plus a test at
the same two insertion points:

  main   openAIHTTPPassthroughSSERateLimitUpstream
         TestOpenAIResponses_APIKeyPassthroughSSERateLimitUsesConfiguredPoolRetry
  branch openAIHTTPPassthroughAuthFailoverUpstream
         TestOpenAIResponses_APIKeyPassthroughPoolAuthFailureRetriesThenSwitchesToHealthyAccount

Both sides are kept verbatim; the only edit is giving each stub its own
calls() body instead of sharing the trailing one. No assertion was changed.

openai_gateway_passthrough.go and openai_oauth_passthrough_test.go merged
automatically.
2026-08-11 14:09:07 +08:00
Wesley Liddick caa1abb13a Merge pull request #5404 from wucm667/fix/issue-5400-oauth-image-stream-error
fix(openai): fail over OAuth image stream errors
2026-08-11 14:04:42 +08:00
Wesley Liddick 574dfad2dc Merge pull request #5488 from wucm667/fix/issue-5482-stale-codex-threshold
fix(openai): skip stale and reset Codex snapshots in scheduling threshold evaluator
2026-08-11 14:00:33 +08:00
Wesley Liddick 6876477371 Merge pull request #5304 from wucm667/fix/issue-5302-chat-reasoning-alias
fix(apicompat): accept chat reasoning alias
2026-08-11 13:59:49 +08:00
Wesley Liddick b918874f81 Merge pull request #5403 from cyhhao/fix/codex-capacity-exponential-backoff
fix(openai): back off capacity retries exponentially
2026-08-11 13:58:06 +08:00
Wesley Liddick 20a2d12dde Merge pull request #5415 from Yuxin-Qiao/fix/openai-responses-empty-completed-failover
fix(openai): fail over empty response.completed streams instead of recording 0/0 success
2026-08-11 13:53:43 +08:00
Wesley Liddick ca9e2b48ee Merge pull request #5413 from Yuxin-Qiao/fix/openai-responses-reasoning-item-id
fix(openai): strip invalid reasoning item IDs in API key passthrough
2026-08-11 13:53:05 +08:00
Wesley Liddick 0f73203e35 Merge pull request #5277 from Pluviobyte/agent/fix-grok-missing-usage-billing
fix: reject Grok responses without billable usage
2026-08-11 09:28:43 +08:00
Wesley Liddick e2d9034d3d Merge pull request #5327 from fengshao1227/fix/fingerprint-user-agent-validation
fix(identity): validate user-agent before persisting account fingerprint
2026-08-11 09:28:27 +08:00
Wesley Liddick ba4bd0c0b4 Merge pull request #5481 from fengshao1227/fix/responses-deterministic-400-passthrough
fix(openai): 原生 Responses 路径不再把上游确定性 400 归一成可重试 502
2026-08-11 09:27:34 +08:00
wucm667 3e1674a060 fix(settings): cache unset scheduling thresholds 2026-08-11 04:34:49 +08:00
anya e5b325e481 fix(billing): harden response-model billing admission
Three guards on the response_model billing basis, all scoped to the opt-in
channel mode so existing channels are unaffected.

1. Per-unit billing gate was stale. Audio (AudioUsage) and the search
   surcharge (SearchCount) reached the billing paths after this branch was
   cut; both are priced per unit rather than per token, so they must be
   excluded like image/video/web-search already are. Audio pricing ignores
   the model entirely, so the previous code "adopted" a basis switch that
   changed nothing and emitted a misleading audit log for it.

2. Never zero out a billable request. A catalog entry whose token prices are
   explicitly 0 still passes the identified-pricing gate (TokenPricingAbsent
   only means both prices are missing), so an upstream could declare a free
   model name and drop the bill to zero. Reject a zero (or negative)
   recomputation whenever the baseline was billable; an already-zero baseline
   is unaffected.

3. Never cross from channel pricing to the global table. Channel pricing
   matches exact keys and prefix wildcards and does not strip date suffixes,
   while the global table's identified lookup does. Upstreams routinely
   declare dated model IDs (claude-opus-4-5-20251101), so allowing a
   cross-source comparison would silently bypass an administrator's channel
   markup on essentially every request. Admins who want a downgrade target
   discounted can price it explicitly on the channel.

Also skip the recomputation entirely when the declared model equals the
baseline: it is provably the same cost and only burned a pricing resolve.

The identified-pricing helpers now return whether the model resolved to
channel pricing so the third guard costs no extra resolve.
2026-08-10 19:11:47 +08:00
shaw 33351c7bc7 fix(billing): gofmt channel.go and drop the redundant response-model hint
- channel.go: 常量块里插入注释后 gofmt 会把 BillingModelSourceResponse 单独成组,
  原写法沿用了上一组的对齐空格,CI 的 gofmt 检查因此失败
  (internal/service/channel.go:45: File is not properly formatted)。
- 撤掉渠道表单里新加的那条提示:说明本就多余,且"只降不升"只在"相对基线收费"
  这个口径下成立,容易被读成"上游返回更贵的模型也不会多收",反而误导。
2026-08-10 18:45:16 +08:00
shaw b689e5b401 fix(billing): harden response-model billing and repair its test fixtures
按上游响应模型计费的准入过宽、且自带用例必然失败,本次一并修复。

严格化准入
- 新增 PricingService.GetIdentifiedModelPricing / BillingService.HasIdentifiedTokenPricing:
  只接受价格表中能被确定性识别的条目(精确名、已知拼写变体、去掉日期版本后缀),
  不再接受 getFallbackPricing / matchByModelFamily 按子串猜出的系列兜底价。
  此前上游只要自报一个含 "haiku" 的编造名字就能被判定"已定价",把账单压到最便宜的
  系列价(实测 claude-opus-4.8 基线 $0.0019250 → $0.0000963,20 倍少收)。
  GetModelPricing 的对外行为不变,仅把前三步查找抽成共用函数。
- 图片 / 视频 / 网页搜索请求不再走响应模型覆盖:这些路径按张、按秒、按次定价,
  与准入检查所验的 token 价不是同一套价格表。
- 准入判断抽成 responseModelBillingDeclaration,两条计费主干共用同一套规则。

正确性与可观测性
- 去掉 recordUsageCore 中 billingModel 的无效赋值(ineffassign 已启用,会让
  golangci-lint 直接失败),改由日志表达实际生效的计费基准。
- 补 cost != nil 守卫,与 OpenAI 侧及本文件既有写法对齐。
- 每次实际生效的基准切换记一条 billing.response_model_applied,少收可审计。
- 修正 upstreamResponseModelObserver 上"冲突仅用于诊断、永不影响计费"的过期注释。

测试
- gpt-5.1 与 gpt-5.5 实际共用同一条 gpt-5.4 价格,夹具"价格必须不同"的前置断言
  必然失败,OpenAI 侧 3 个用例(含 4 个子用例)从未跑通;改用 gpt-5.4-nano /
  gpt-5.5。Anthropic 侧 claude-opus-4 不是价格表精确条目,改用 claude-opus-4.8。
- 夹具增加"必须可被确定性识别"的前置断言,避免用例被更靠前的门挡掉而失去判别力。
- 新增:准入规则表驱动用例、可识别性判定用例、编造家族名在两条主干上均被拒的用例。

前端
- 选择该模式时提示"计费基准以上游自报模型为准,只降不升,仅对可信上游启用"。
2026-08-10 18:45:15 +08:00
pigzwy 9096492b55 feat(billing): support safe upstream response model billing 2026-08-10 18:45:14 +08:00
wucm667 3d3aee2e72 fix(openai): skip stale and reset Codex snapshots in scheduling threshold evaluator 2026-08-10 14:27:04 +08:00
li 591d47fb9b fix(openai): 原生 Responses 路径不再把上游确定性 400 归一成可重试 502
上游返回 400(invalid_function_parameters、missing_required_parameter 等)时,
`/v1/responses` 原生路径把它包成 `502 upstream_error / "Upstream request failed"`,
真实的 message/code/param 全部丢弃。

502 属于可重试类,下游网关会把这个永远不会成功的请求反复重放(issue #5479 实测
30 个失败请求被放大成 60 次上游调用),客户端也拿不到「哪个字段非法」的线索。

同一个 service 上的兄弟路径早已做对:
- `handleCompatErrorResponse`(ChatCompletions / Anthropic)回真实状态码 +
  invalid_request_error + 真实 message;
- `/v1/images/generations` 还额外透传 code/param。

原生 Responses 是唯一漏掉的一条,本次对齐。

改动只放行 400:走到该分支说明 `shouldFailoverOpenAIUpstreamResponse` 已判定这个
400 不可 failover,即 server_is_overloaded / at capacity 这类可重试的 400 不会到达。
401/402/403(运营方凭据问题)、404/405(可能是 base_url 配错)、429 一律维持原状。

Fixes #5479
2026-08-10 11:25:46 +08:00
Wesley Liddick 10a4c6e3ad Merge pull request #5234 from wucm667/fix/issue-5230-deduplicate-latest-turn-audit
fix(security-audit): deduplicate websocket turn audits
2026-08-10 10:53:17 +08:00
Wesley Liddick f3c7a1a8c4 Merge pull request #5295 from wucm667/fix/issue-5289-streaming-upstream-error
fix: emit response.failed when compact keepalive commits headers but no SSE payload
2026-08-10 10:52:49 +08:00
Wesley Liddick f19e3816d0 Merge pull request #5316 from wucm667/fix/issue-5313-legacy-scheduler-diagnostics
fix(scheduler): diagnose legacy OpenAI exclusions
2026-08-10 10:52:14 +08:00
Wesley Liddick 30d0405388 Merge pull request #5464 from wucm667/fix/issue-5455-api-key-input-validation
fix(api-key): validate quota and expiry inputs
2026-08-10 10:52:04 +08:00
Wesley Liddick 895e8247af Merge pull request #5472 from wucm667/fix/issue-5468-openai-threshold-percent
fix(openai): preserve Codex usage percentages for scheduling thresholds
2026-08-10 10:51:43 +08:00
Wesley Liddick 2d0976ac2c Merge pull request #5475 from fengshao1227/fix/oauth-personal-subscription-expiry
fix(openai-oauth): 个人订阅到期时间不再被 POID workspace 的 entitlement 覆盖
2026-08-10 10:23:05 +08:00
Wesley Liddick af6928a268 Revert "fix(risk-control): 风控后端异常时阻断提示词请求" 2026-08-10 09:59:42 +08:00
li 358e4a89a1 fix(openai-oauth): 个人订阅到期时间不再被 POID workspace 的 entitlement 覆盖
enrichTokenInfo 里 plan_type 与 subscription_expires_at 的取值口径不一致:

- plan_type 有 shouldApplyChatGPTAccountInfoPlanType 护着(#3641),
  id_token 里的个人套餐优先,accounts/check 只在个人值为空时补位;
- subscription_expires_at 却是无条件覆盖。

accounts/check 是多账号/工作区端点,命中的记录由 access_token JWT 的 poid 决定。
当 poid 指向默认 Personal workspace、而 chatgpt_account_id 指向个人 ChatGPT 账号时
(两者是不同标识),拿到的 entitlement.expires_at 描述的是 workspace 权益。
配上被保护下来的个人 plan_type,账号页就显示成「个人 Pro + workspace 到期时间」。

代码里本来已经有正确的来源——用个人 chatgpt_account_id 查
/backend-api/subscriptions 的 active_until——但它被
`if TrimSpace(SubscriptionExpiresAt) == ""` 挡着,workspace 值一旦写进去就永远不触发。

改为让两个字段始终描述同一份订阅:

- 套餐取自 accounts/check 时(id_token 没带 chatgpt_plan_type),到期时间跟着取
  同一条记录,行为不变;
- 套餐保留了 id_token 的个人值时,只有该记录确实属于个人账号才用它的
  entitlement.expires_at;不属于就跳过,并强制回落到个人订阅端点。

为此给 ChatGPTAccountInfo 补 AccountID:优先读 account.account_id,缺失时退回
accounts 的 map key(key 可能是 "default" 这类别名)。两侧任一缺 ID 时判定为
无法区分并沿用旧行为,poid == chatgpt_account_id 的单人账号完全不受影响。

Fixes #5459
2026-08-10 09:29:14 +08:00
Wesley Liddick b5e83156d9 Merge pull request #5394 from wucm667/fix/issue-5388-risk-control-fail-closed
fix(risk-control): 风控后端异常时阻断提示词请求
2026-08-10 08:51:46 +08:00
wucm667 99b31067f7 fix(openai): preserve Codex usage percentages in scheduling threshold 2026-08-10 04:29:08 +08:00
wucm667 f5c108c836 fix(api-key): validate quota and expiry inputs 2026-08-09 22:42:51 +08:00
lyen1688 bbc8b6e906 完善大文件备份分卷上传与恢复 2026-08-09 20:58:07 +08:00
github-actions[bot] 48eb3766d2 chore: sync VERSION to 0.1.173 [skip ci] 2026-08-09 08:26:22 +00:00
shaw 563a72ca73 feat: add default-off switch for email domain registration quota
PR #5423 relaxed the email suffix whitelist: once a whitelist is
configured, non-whitelisted registrable domains are each allowed to
register one account. That behavior activated unconditionally.

Add registration_email_domain_quota_enabled (default false) to gate it:

- Off (default): restore pre-#5423 strict whitelist semantics — with a
  non-empty whitelist, non-whitelisted domains are rejected with
  EMAIL_SUFFIX_NOT_ALLOWED; the register/verify views restore the
  client-side whitelist pre-check and allowed-domain hint.
- On: keep #5423 behavior — one account per non-whitelisted registrable
  domain (EMAIL_DOMAIN_REGISTRATION_LIMIT).
- Empty whitelist keeps allowing all domains in both states.

Gating lives in validateRegistrationEmailQuota and (as a race-safety
backstop) createUserWithRegistrationEmailGuard; the repository-level
domain lock + in-tx recheck is unchanged. The admin update field is
*bool (omitted = keep current) so stale full-payload saves cannot
silently flip the switch. Email binding and OAuth auto-signup keep
their strict policy, and pending-OAuth bind-login for existing
accounts is unaffected because the handler resolves existing emails
before the quota check.

Frontend adds the toggle to admin settings (zh/en copy; whitelist hint
restored to strict wording, quota wording moved to the new toggle) and
exposes the flag via public settings + SSR injection payload.

Tests: #5423 quota tests now enable the switch explicitly; new
default-off regression tests cover register/send-code/async/pending
OAuth/OIDC create-account plus both register views; API contract JSON
and the injection drift guard are updated.
2026-08-09 15:53:40 +08:00
Wesley Liddick f2da30bcd9 Merge pull request #5423 from lyen1688/feat/email-domain-registration-quota
完善邮箱域名注册额度策略
2026-08-09 15:16:26 +08:00
Wesley Liddick 7821c4005e Merge pull request #5424 from fengshao1227/fix/gemini-native-image-billing
fix(gemini): 原生生图按上游实际回吐的图片张数计费,修复自定义模型名下生图记 $0
2026-08-09 15:02:08 +08:00
Wesley Liddick c5bda8b8e4 Merge pull request #5416 from fengshao1227/fix/images-apikey-detach-upstream-context
fix(openai): 非流式生图脱钩上游 context,客户端断开不再导致图已出却不扣费
2026-08-09 14:55:28 +08:00
Wesley Liddick a1073843ac Merge pull request #5437 from Brisbanehuang/codex/upstream-response-model-audit-followup
perf(usage): 优化上游响应模型观察热路径
2026-08-09 14:55:15 +08:00
Wesley Liddick 7a113fb7a3 Merge pull request #5352 from feeeei/main
fix(gemini): 修复Gemini池模式时,依然被 429 response 触发账户限流问题
2026-08-09 14:55:02 +08:00
shaw d92edc01be Merge origin/main into feat/channel-monitor-v2-ops-ui
Resolves three conflicts, all of the "both branches appended to the same
block" shape. Every one is resolved as a union of both sides; nothing from
either parent is dropped.

- handler/admin/setting_handler_update.go: keep ChannelMonitorHideThroughput
  (V2) alongside GrokDefaultTextModel / GrokCrossClientModelMapEnabled /
  GrokDefaultBaseURLMode (#5408). UpdateSettings writes every key on each
  save, so dropping either side would reset those settings to zero values.
- service/domain_constants.go: keep SettingKeyChannelMonitorHideThroughput
  and the three SettingKeyGrok* constants.
- repository/migrations_runner.go: keep the 195 checksum rule (V2) and the
  218/219/220 rules (#5408).
2026-08-09 12:11:35 +08:00
Brisbanehuang 6e34fb09c9 perf(usage): optimize upstream response model observation 2026-08-08 22:04:49 -04:00
IanShaw027 5315896b30 fix(grok): align free soft-gate tests with async fail-open cache
Treat cacheTTL=0 as non-expiring known entries, always store negative
refresh markers, and update sticky getSchedulableAccount tests to expect
first-hit fail-open then block after background stats warm.
2026-08-08 23:38:50 +08:00
IanShaw027 a54a4b674b fix(grok): clear golangci errcheck and gofmt on free-quota path
Check the sync.Map type assertion in free-quota refresh coalescing, and
gofmt migration checksum rules plus prompt-audit route map alignment.
2026-08-08 22:52:31 +08:00
li b6eb6c1efa fix(gemini): 原生生图按上游实际回吐的图片张数计费
/v1beta/models/{model}:generateContent 与 Anthropic→Gemini 兼容路径的
ImageCount 只由 isImageGenerationModel(originalModel) 决定,而该白名单是
按 Google 官方模型名精确/前缀匹配写死的(antigravity_image_test.go 里
明确断言 my-gemini-3-pro-image-test 这类自定义名返回 false)。

GeminiMessagesCompatService 服务的却主要是 API Key + 自定义模型映射的账号:
客户端请求名和 GetMappedModel 后的上游名都可能是站长自取的别名,白名单必然
判不出来 → ImageCount=0 → calculateRecordUsageCost 里 `if result.ImageCount > 0`
的按次计费分支整条不触发 → 生图请求全部记 $0(issue #5358)。

改为优先按上游响应里真实的 inlineData 图片 part 计数:
- 新增请求级计数器,挂在 gin.Context 上,与既有的
  upstreamResponseModelObserver 同一批调用点取解包后的响应体;
- 取「单个 payload 内的最大值」而非累加:Gemini 兼容上游的 SSE 分片可能是
  累积式的(computeGeminiTextDelta 正是为此存在),逐 chunk 累加会把同一张图
  重复计费。max 保证累积式流与非流式都得到真实张数,增量式多图流最差退化到 1,
  与改动前同值,不构成回退;
- 每次 Forward 开头重置计数器,避免 failover 复用同一个 gin.Context 时
  把失败账号已回吐的图叠加到成功账号账单上;
- 响应里数不出图时(fileData 引用式回图等)退回原有模型名启发式,并额外认
  映射后的上游模型名,与 shouldSkipCodexPlanGatedImageModelCooldown 同时取
  requestedModel / modelKey 的口径一致。

inlineData / inline_data 两种字段风格都认(官方 SDK 与部分中转回 snake_case),
只统计带 base64 数据且 MIME 为图片的 part。

Fixes #5358
2026-08-08 21:36:18 +08:00
lyen1688 4999231d61 修复邮箱域名注册额度策略 2026-08-08 21:05:20 +08:00
feeeei cbc2a3dd46 修复池模式 Gemini 账号被 429 打上账号级限流
- 429 的标记点在重试循环内,先于 CheckErrorPolicy 执行,池模式豁免只能落在
  handleGeminiUpstreamError 自身;否则一次上游 429 会把账号锁到 PST 午夜,
  即便重试已经成功返回客户端
- 判定条件与 HandleUpstreamError 对齐:自定义错误码优先级高于池模式;
  401/403/529 仍委派给 RateLimitService,临时不可调度规则不受影响
- chat completions 路径的策略分发改为只有 None / Matched 才处理账号状态,
  与 messages 兼容层的 switch 一致,ErrorPolicySkipped 不再漏进来
2026-08-08 16:51:48 +08:00
li cbf2be05a3 fix(openai): 非流式生图脱钩上游 context,客户端断开不再导致图已出却不扣费
forwardOpenAIImagesAPIKey 走的是 detachStreamUpstreamContext(ctx, parsed.Stream),
该函数在非流式时原样返回请求 context(gateway_usage_billing.go:488)。于是客户端
中途断开会连带取消已经在出图的上游调用:上游那边图已生成并计费,网关这边拿到
context canceled,记 502,result 为 nil,handler 的 result.ImageCount > 0 兜底
无法命中,本次请求不产生扣费。

生图是长耗时(数十秒)且上游侧已产生实际成本的操作,与普通非流式 chat 不同:
提前取消并不能省下上游开销,只会丢掉已付费的产出。

同一端点的 OAuth 分支 forwardOpenAIImagesOAuth(openai_images_responses.go:1690)
以及同属媒体生成的 grok_media.go 本来就用无条件脱钩的 detachUpstreamContext,
OpenAI 网关侧 13 个转发点里只有这一处是例外。这里对齐。

Fixes #5411
2026-08-08 15:24:34 +08:00
IanShaw027 e91b494168 fix(test): stabilize OpenAI streaming preamble keepalive assertion
Keepalive is gated on downstream idle, not upstream tick cadence. The old
fixture wrote progress events every 250ms and could finish without a true
1s idle window on loaded CI runners, so ":\n\n" never appeared. Pause after
preamble long enough for the keepalive ticker before completing the stream.
2026-08-08 15:00:33 +08:00