Commit Graph
4315 Commits
Author SHA1 Message Date
Wesley Liddick 4853662f99 Merge pull request #5540 from luckydududu/fix/channel-pricing-model-name-normalization
fix(channel): 定价冲突检测与定价缓存键的归一化对齐(#4754 修复残留)
2026-08-13 09:12:41 +08:00
IanShaw027 b61e4bcc45 fix: 对齐 JWT 常量 gofmt,并补全 available groups 契约字段
golangci-lint 要求 const 块按类型对齐;分组 DTO 新增
long_context_pricing_enabled 后,契约夹具未同步导致比较失败。
2026-08-13 08:55:56 +08:00
IanShaw027 0ae151a238 fix: 徽章改读 grok_usage_snapshot,增量刷新比较 Grok 快照
后端写入 grok_usage_snapshot,列表却读无 writer 的 grok_quota_snapshot;
增量刷新不比 Grok 快照,档位更新后行不替换。

- 正式快照优先 usage,quota 仅兜底;计划字段只收非空字符串
- buildGrokUsageRefreshKey 接入自动刷新
- 注明网关 Extra 写入依赖每请求独立 Account 副本
2026-08-13 08:38:10 +08:00
IanShaw027 b830bc14d6 fix: 长上下文默认保持开启,并与 OpenAI 账号开关取交集
迁移默认 false 会让存量分组静默丢掉 ≥200k 阶梯与渠道多档价。
OpenAI 路径传入 Group 后,账号级 openai_long_context_billing_enabled 被顶掉。

- 列默认 TRUE,并对存量行回填
- 分组价卡不再手填 token 区间,勾选走官方/预设阶梯
- 分组开关与账号开关 AND;价卡 JSON 损坏时打 warn
2026-08-13 08:38:10 +08:00
IanShaw027 8c4c3c09ce fix: 未登记的 Grok 文本模型回退 grok-4.5 价卡
新文本 ID 在闭集价卡上 fail-closed,请求成功而用量记 0。

- grok / grok-N* / build / composer 未登记别名继承 4.5 价
- imagine/voice/web-search/speech 等非文本族不回退
2026-08-13 08:38:10 +08:00
IanShaw027 678eb22a40 fix: Realtime 仅在观察到音频后计费,并修正标志位求值顺序
原先 elapsed>0 就出账,握手失败也会扣费;随后用 audioObserved,
但又在同一个 return 里先 Load 再收 errCh,标志位恒为 false,会话全部漏计。

- 先等中继结束再读 audioObserved
- 无音频或零时长不出账;每次连接独立 request id
2026-08-13 08:38:10 +08:00
IanShaw027 c4d883b8da feat: Chat 与 Responses 往返保留 x_search,并补 sources 抽取
Chat Completions 的 {"type":"x_search"} 会被 apicompat 丢掉,独立
/x_search 又没带 include 与结构化提示,主路径容易返回空结果仍计费。

- Chat↔Responses 保留 x_search 过滤字段与 tool_choice
- declared 只注册实际存活的 x_search,web_search 选择项仍丢弃
- 上游补 include x_search_call.action.sources 与结构化输出提示
2026-08-13 08:37:57 +08:00
IanShaw027 0de6d7e9ba feat: 新增独立 /x_search,走原生 x_search 并沿用搜索计费
Grok 分组原先只有 /web_search,无法带 handle/日期过滤,也无法走
xAI 的 x_search tool。

- POST /x_search(仅 Grok 分组),复用 web_search 的审计、failover 与按次计费
- 上游 Responses 强制 x_search;计费模型记为 grok-x-search
2026-08-13 07:49:19 +08:00
IanShaw027 363cc4994b fix: SuperGrokPro 用 4.5 窗口区分 Heavy,容量抖动只封单模型
JWT 的 SuperGrokPro 同时覆盖 SuperGrok 与 Heavy,账单月额度又滞后,
账号会误升/误降。multi-agent 容量抖动还会把整号提出调度。

- CanonicalGrokPlan:明确 JWT 优先;模糊档仅采新鲜的 grok-4.5 Responses 窗口
- 配额快照写入 plan_from_45_responses 时间戳,过期信号不用
- engine_overloaded 对 multi-agent 只封当前模型 0.5s–5min
2026-08-13 07:49:14 +08:00
IanShaw027 f3d9491071 feat: 分组支持逐模型定价,并可关闭长上下文阶梯
运营需要按分组覆盖渠道/内置价,且部分套餐不应自动吃 200k 倍率。
原先只能改渠道价卡,分组侧只剩 Voice 三列。

- groups 新增 model_pricing / long_context_pricing_enabled,解析链改为 Group → Channel → 内置
- 关闭长上下文时 token 模型只取最低档;video 按秒计费可写进同一价卡
- 回退价对齐官方卡:4.5 缓存 $0.30、4.3/imagine/audio/search 默认值一并校正
2026-08-13 07:49:08 +08:00
IanShaw027 a04ce49016 feat: 新增 grok-4.6 目录、官方定价与请求路径支持
官方目录价:<200k 为 $2 / 缓存读取 $0.50 / $6,≥200k 全部 2 倍。
原先没有独立价卡,/models 宣称可配 reasoning 但请求路径会剥掉 effort。

- 目录与别名接入 grok-4.6 / grok-4.6-latest,默认文本模型仍为 grok-4.5
- 独立回退价卡含缓存读取价与 200k 长上下文倍率
- 保留 reasoning_effort;Chat→Responses 的 cache/vision 桥接同样识别 4.6
2026-08-13 01:23:45 +08:00
IanShaw027 bb9e74285e feat: 从 JWT tier 识别 Grok 订阅档位,刷新后覆盖失效订阅
Grok Build access token 带有数字 tier claim(0=free、1=supergrok、5=heavy、6=lite),
原先只解 email/sub/team_id,refresh 还会把旧档位抄回去,订阅失效后一直显示 Heavy。

- 解码 JWT tier,并归一化 display / header 别名(含 free-tier、SuperGrok Lite)
- 仅从 access token 读取档位;新 JWT 覆盖已存凭证,AT 无 claim 时保留旧值
- 用量与 free-cache 判定以当前 AT 为准,账单月额度仅作降级
2026-08-13 01:23:37 +08:00
github-actions[bot] ef4f99f292 chore: sync VERSION to 0.1.175 [skip ci] 2026-08-12 11:07:43 +00:00
shaw 04f8cdb194 fix: 修复 golangci-lint errcheck 和 gofmt 格式问题 2026-08-12 18:20:54 +08:00
shaw c0ab3a00ea feat: Codex OAuth 设备指纹收敛,减少上游可见的设备数和会话数
多人共享同一 OAuth 账号时,各用户 Codex 客户端携带各自不同的 installation_id/session_id/thread_id,
上游据此判定设备数和会话数并限制配额。本功能将这些标识改写为账号级恒定值。

四档策略(账号级 extra 字段 codex_fingerprint_mode):
- off: 不做任何收敛,原样透传
- device: 仅收敛 installation_id
- session(默认): 收敛 installation_id + session_id,thread_id 按客户端原始 session 派生
- full: 收敛所有标识(installation_id + session_id + thread_id)

改写覆盖 6 个指纹载体:x-codex-turn-metadata 头(JSON 内部字段)、x-codex-window-id、
x-codex-installation-id、x-client-request-id、session-id/session_id/thread-id、
请求体 client_metadata。头和体共享同一份预计算 IDs 确保 turn_id 等随机字段一致。
2026-08-12 17:57:50 +08:00
Lucky bd404c16f2 fix(channel): align pricing conflict detection with the pricing cache key
`validateNoConflictingModels` used `toModelEntry`, which only lowercases,
while `expandPricingToCache` keys the cache with
`normalizeChannelPricingModelName` (lowercase + TrimSpace + `.` -> `-`
for `claude-*`). Two pricing entries that the validator considers distinct
therefore collapse onto the same cache key, and the one written later
silently overwrites the other.

Add `toPricingModelEntry`, which reuses `normalizeChannelPricingModelName`,
and use it for pricing conflict detection. `toModelEntry` is left untouched
for model *mapping* validation, because `expandMappingToCache` keys the
mapping cache with plain `strings.ToLower` -- normalizing mappings the same
way would reject configurations that the cache keeps separate.

This completes the fix for #4754, which added the normalization to the
lookup side only.
2026-08-12 02:17:09 +00:00
Wesley Liddick 46cbb7187b Merge pull request #5531 from SamizuHM/fix/openai-compat-nested-data-usage
fix(openai-compat): parse usage from nested data envelopes
2026-08-12 10:01:09 +08:00
Wesley Liddick 80acae16fa Merge pull request #5525 from Fool0ntheHill/codex/fix-openai-visible-ttft
fix(openai): record Responses TTFT on visible output
2026-08-12 09:59:52 +08:00
Wesley Liddick 19c6007a22 Merge pull request #5342 from wucm667/fix/issue-5340-ws-v2-terminal-ttft
fix(openai-ws): exclude terminal events from TTFT
2026-08-12 09:59:43 +08:00
Wesley Liddick 0ed1a9f22a Merge pull request #5514 from wucm667/fix/issue-5510-cyber-policy-audit-scope
fix(audit): scope cyber policy events
2026-08-12 09:58:32 +08:00
Wesley Liddick a29fce4a61 Merge pull request #5511 from wucm667/fix/pr-5234-ws-audit-logging
fix(security-audit): restore websocket audit logs
2026-08-12 09:58:24 +08:00
Wesley Liddick 1225437099 Merge pull request #5502 from pyt111/codex/fix-account-stats-service-tier
fix: 修复 service tier 账号成本统计
2026-08-12 09:58:16 +08:00
Wesley Liddick 5192abb6b5 Merge pull request #5503 from fengshao1227/fix/openai-html-403-not-account-penalty
fix(openai): 上游 HTML 403 不再被当成账号级错误处罚账号
2026-08-12 09:58:08 +08:00
Wesley Liddick 40aae11888 Merge pull request #5513 from wucm667/fix/issue-5506-gemini-exclusive-minimum
fix(gemini): normalize exclusive minimum tool schemas
2026-08-12 09:57:52 +08:00
SamizuHM a163742fc9 fix(openai): preserve usage path precedence 2026-08-11 22:01:47 +08:00
SamizuHM 04dc540b23 fix(openai): parse nested data usage envelopes 2026-08-11 18:15:31 +08:00
Fool0ntheHill 900194fab2 fix(openai): 修正 Responses 可见输出 TTFT 2026-08-11 16:31:09 +08:00
wucm667 662444774f test(gemini): check cleaned schema assertions 2026-08-11 16:19:18 +08:00
wucm667 e24cb99b79 fix(openai-ws): retain no-delta TTFT fallback 2026-08-11 16:01:34 +08:00
Wesley Liddick 1e618dbc29 Merge pull request #5054 from wucm667/fix/issue-5029-openai-passthrough-pool-auth-retry
fix(openai): retry pool auth failures before failover
2026-08-11 14:21:45 +08:00
shaw a3bbf35cbd Merge branch 'main' into fix/issue-5029-openai-passthrough-pool-auth-retry
Resolve conflict in backend/internal/handler/openai_gateway_handler_test.go.

main and this branch each appended a passthrough upstream stub plus a test at
the same two insertion points:

  main   openAIHTTPPassthroughSSERateLimitUpstream
         TestOpenAIResponses_APIKeyPassthroughSSERateLimitUsesConfiguredPoolRetry
  branch openAIHTTPPassthroughAuthFailoverUpstream
         TestOpenAIResponses_APIKeyPassthroughPoolAuthFailureRetriesThenSwitchesToHealthyAccount

Both sides are kept verbatim; the only edit is giving each stub its own
calls() body instead of sharing the trailing one. No assertion was changed.

openai_gateway_passthrough.go and openai_oauth_passthrough_test.go merged
automatically.
2026-08-11 14:09:07 +08:00
Wesley Liddick caa1abb13a Merge pull request #5404 from wucm667/fix/issue-5400-oauth-image-stream-error
fix(openai): fail over OAuth image stream errors
2026-08-11 14:04:42 +08:00
Wesley Liddick 574dfad2dc Merge pull request #5488 from wucm667/fix/issue-5482-stale-codex-threshold
fix(openai): skip stale and reset Codex snapshots in scheduling threshold evaluator
2026-08-11 14:00:33 +08:00
Wesley Liddick 6876477371 Merge pull request #5304 from wucm667/fix/issue-5302-chat-reasoning-alias
fix(apicompat): accept chat reasoning alias
2026-08-11 13:59:49 +08:00
Wesley Liddick b918874f81 Merge pull request #5403 from cyhhao/fix/codex-capacity-exponential-backoff
fix(openai): back off capacity retries exponentially
2026-08-11 13:58:06 +08:00
Wesley Liddick 20a2d12dde Merge pull request #5415 from Yuxin-Qiao/fix/openai-responses-empty-completed-failover
fix(openai): fail over empty response.completed streams instead of recording 0/0 success
2026-08-11 13:53:43 +08:00
Wesley Liddick ca9e2b48ee Merge pull request #5413 from Yuxin-Qiao/fix/openai-responses-reasoning-item-id
fix(openai): strip invalid reasoning item IDs in API key passthrough
2026-08-11 13:53:05 +08:00
wucm667 6564d376e5 fix(audit): scope cyber policy events 2026-08-11 13:05:56 +08:00
wucm667 c8d9af6ce1 fix(gemini): normalize exclusive minimum tool schemas 2026-08-11 12:34:51 +08:00
wucm667 2d9920ba7d fix(security-audit): restore websocket audit logs 2026-08-11 11:56:46 +08:00
li 12abb54700 fix(openai): 上游 HTML 403 不再被当成账号级错误处罚账号
上游代理 / CDN 在请求到达 OpenAI API 之前拦下时,回的是 HTML 403 页面而不是
{"error":{...}} 结构化错误。这类响应描述的是「这条链路 / 这个端点被挡了」,
不构成账号凭据或权限失效的证据。

但 handleOpenAI403 不区分响应形态,一律按账号级 403 处理:

- 首次即 SetTempUnschedulable,账号 10 分钟不可调度;
- 连续 openAI403DisableThreshold(3) 次直接 SetError 永久禁用账号;
- 403 又在 failover 状态集里,同一个坏请求会被逐个账号重放。

结果是一个请求级错误被放大成整组账号下线。issue #5334 给出的触发方式是
POST /v1/responses/not-exist —— 该路径能通过只做结构校验的 guardResponsesSubpath,
转发后拿回 HTML 403,任何持有 API Key 的调用方重复几次即可打穿一个分组。
附带问题:整张 HTML 页面会被拼进账号错误信息写库并显示在管理端。

仓库对这类响应早有明确口径,只是没有应用到这条路径:

- openai_gateway_count_tokens.go 的 isOpenAIOAuthInputTokensUnsupported 已把
  「HTML 403 page without a structured error」按端点级响应处理;
- shouldApplyOpenAIAlphaSearchAccountErrorSideEffects 的既定不变式是
  端点级错误只换号、不写账号错误状态。

本次复用现成的 isHTMLResponse,在 handleOpenAI403 入口跳过账号处罚:不递增
连续 403 计数、不设临时不可调度、不永久禁用。failover 行为不变 —— 换一个走
不同代理的账号仍有可能成功。

Fixes #5334
2026-08-11 10:18:20 +08:00
pyt111 9261dd7734 [verified] fix: apply service-tier pricing to account cost 2026-08-11 09:55:19 +08:00
Wesley Liddick 0f73203e35 Merge pull request #5277 from Pluviobyte/agent/fix-grok-missing-usage-billing
fix: reject Grok responses without billable usage
2026-08-11 09:28:43 +08:00
Wesley Liddick e2d9034d3d Merge pull request #5327 from fengshao1227/fix/fingerprint-user-agent-validation
fix(identity): validate user-agent before persisting account fingerprint
2026-08-11 09:28:27 +08:00
Wesley Liddick ba4bd0c0b4 Merge pull request #5481 from fengshao1227/fix/responses-deterministic-400-passthrough
fix(openai): 原生 Responses 路径不再把上游确定性 400 归一成可重试 502
2026-08-11 09:27:34 +08:00
wucm667 3e1674a060 fix(settings): cache unset scheduling thresholds 2026-08-11 04:34:49 +08:00
anya e5b325e481 fix(billing): harden response-model billing admission
Three guards on the response_model billing basis, all scoped to the opt-in
channel mode so existing channels are unaffected.

1. Per-unit billing gate was stale. Audio (AudioUsage) and the search
   surcharge (SearchCount) reached the billing paths after this branch was
   cut; both are priced per unit rather than per token, so they must be
   excluded like image/video/web-search already are. Audio pricing ignores
   the model entirely, so the previous code "adopted" a basis switch that
   changed nothing and emitted a misleading audit log for it.

2. Never zero out a billable request. A catalog entry whose token prices are
   explicitly 0 still passes the identified-pricing gate (TokenPricingAbsent
   only means both prices are missing), so an upstream could declare a free
   model name and drop the bill to zero. Reject a zero (or negative)
   recomputation whenever the baseline was billable; an already-zero baseline
   is unaffected.

3. Never cross from channel pricing to the global table. Channel pricing
   matches exact keys and prefix wildcards and does not strip date suffixes,
   while the global table's identified lookup does. Upstreams routinely
   declare dated model IDs (claude-opus-4-5-20251101), so allowing a
   cross-source comparison would silently bypass an administrator's channel
   markup on essentially every request. Admins who want a downgrade target
   discounted can price it explicitly on the channel.

Also skip the recomputation entirely when the declared model equals the
baseline: it is provably the same cost and only burned a pricing resolve.

The identified-pricing helpers now return whether the model resolved to
channel pricing so the third guard costs no extra resolve.
2026-08-10 19:11:47 +08:00
shaw 33351c7bc7 fix(billing): gofmt channel.go and drop the redundant response-model hint
- channel.go: 常量块里插入注释后 gofmt 会把 BillingModelSourceResponse 单独成组,
  原写法沿用了上一组的对齐空格,CI 的 gofmt 检查因此失败
  (internal/service/channel.go:45: File is not properly formatted)。
- 撤掉渠道表单里新加的那条提示:说明本就多余,且"只降不升"只在"相对基线收费"
  这个口径下成立,容易被读成"上游返回更贵的模型也不会多收",反而误导。
2026-08-10 18:45:16 +08:00
shaw b689e5b401 fix(billing): harden response-model billing and repair its test fixtures
按上游响应模型计费的准入过宽、且自带用例必然失败,本次一并修复。

严格化准入
- 新增 PricingService.GetIdentifiedModelPricing / BillingService.HasIdentifiedTokenPricing:
  只接受价格表中能被确定性识别的条目(精确名、已知拼写变体、去掉日期版本后缀),
  不再接受 getFallbackPricing / matchByModelFamily 按子串猜出的系列兜底价。
  此前上游只要自报一个含 "haiku" 的编造名字就能被判定"已定价",把账单压到最便宜的
  系列价(实测 claude-opus-4.8 基线 $0.0019250 → $0.0000963,20 倍少收)。
  GetModelPricing 的对外行为不变,仅把前三步查找抽成共用函数。
- 图片 / 视频 / 网页搜索请求不再走响应模型覆盖:这些路径按张、按秒、按次定价,
  与准入检查所验的 token 价不是同一套价格表。
- 准入判断抽成 responseModelBillingDeclaration,两条计费主干共用同一套规则。

正确性与可观测性
- 去掉 recordUsageCore 中 billingModel 的无效赋值(ineffassign 已启用,会让
  golangci-lint 直接失败),改由日志表达实际生效的计费基准。
- 补 cost != nil 守卫,与 OpenAI 侧及本文件既有写法对齐。
- 每次实际生效的基准切换记一条 billing.response_model_applied,少收可审计。
- 修正 upstreamResponseModelObserver 上"冲突仅用于诊断、永不影响计费"的过期注释。

测试
- gpt-5.1 与 gpt-5.5 实际共用同一条 gpt-5.4 价格,夹具"价格必须不同"的前置断言
  必然失败,OpenAI 侧 3 个用例(含 4 个子用例)从未跑通;改用 gpt-5.4-nano /
  gpt-5.5。Anthropic 侧 claude-opus-4 不是价格表精确条目,改用 claude-opus-4.8。
- 夹具增加"必须可被确定性识别"的前置断言,避免用例被更靠前的门挡掉而失去判别力。
- 新增:准入规则表驱动用例、可识别性判定用例、编造家族名在两条主干上均被拒的用例。

前端
- 选择该模式时提示"计费基准以上游自报模型为准,只降不升,仅对可信上游启用"。
2026-08-10 18:45:15 +08:00
pigzwy 9096492b55 feat(billing): support safe upstream response model billing 2026-08-10 18:45:14 +08:00