Commit Graph
4326 Commits
Author SHA1 Message Date
wucm667 674570ca17 fix: preserve group pricing in auth snapshots 2026-08-14 00:25:55 +08:00
Wesley Liddick fbfdcef818 Merge pull request #5573 from IanShaw027/fix/grok-long-context-and-media-fallback
fix(grok): 长上下文阶梯不再被 OpenAI 账号开关否决,媒体 ID 不继承文本价
2026-08-13 10:32:07 +08:00
github-actions[bot] 0e82efe489 chore: sync VERSION to 0.1.176 [skip ci] 2026-08-13 01:46:47 +00:00
Wesley Liddick e803e3851c Merge pull request #5559 from seng1e/fix/scheduled-backup-leader-lock
fix(backup): 定时备份加 leader 锁,避免多实例重复备份
2026-08-13 09:24:11 +08:00
Wesley Liddick 5912ae960c Merge pull request #5543 from luckydududu/fix/invalidate-channel-cache-on-group-platform-change
fix(group): 改分组 platform 后失效渠道缓存(否则最长 10 分钟按旧平台计价)
2026-08-13 09:24:00 +08:00
Wesley Liddick e77b4b260e Merge pull request #5504 from fengshao1227/fix/responses-probe-inconclusive-keeps-unknown
fix(openai): 探测响应未跑完时不再落标为「上游不支持 Responses」
2026-08-13 09:23:50 +08:00
IanShaw027 e29b93a1fb fix: 未知 Grok 文本兜底排除 image/video/audio 等媒体族
带版本号的媒体 ID(grok-2-image-1212、grok-5-video)会被
grok-N* 白名单当成文本,错吃 4.5 token 价。vision 多模态
仍按 token 计;HasIdentifiedTokenPricing 也不放行这些 ID。
2026-08-13 09:22:38 +08:00
IanShaw027 fd82dfd52d fix: Grok 长上下文只跟分组开关,不受 OpenAI 账号开关否决
openai_long_context_billing_enabled 是 OpenAI 账号设置,Grok
账号无法打开。按账号硬传 false 会让分组开关失效,官方 200k
阶梯和渠道多档区间全部塌到第一档。非 OpenAI 不再传账号门闩。
2026-08-13 09:22:34 +08:00
Wesley Liddick 4853662f99 Merge pull request #5540 from luckydududu/fix/channel-pricing-model-name-normalization
fix(channel): 定价冲突检测与定价缓存键的归一化对齐(#4754 修复残留)
2026-08-13 09:12:41 +08:00
IanShaw027 b61e4bcc45 fix: 对齐 JWT 常量 gofmt,并补全 available groups 契约字段
golangci-lint 要求 const 块按类型对齐;分组 DTO 新增
long_context_pricing_enabled 后,契约夹具未同步导致比较失败。
2026-08-13 08:55:56 +08:00
IanShaw027 0ae151a238 fix: 徽章改读 grok_usage_snapshot,增量刷新比较 Grok 快照
后端写入 grok_usage_snapshot,列表却读无 writer 的 grok_quota_snapshot;
增量刷新不比 Grok 快照,档位更新后行不替换。

- 正式快照优先 usage,quota 仅兜底;计划字段只收非空字符串
- buildGrokUsageRefreshKey 接入自动刷新
- 注明网关 Extra 写入依赖每请求独立 Account 副本
2026-08-13 08:38:10 +08:00
IanShaw027 b830bc14d6 fix: 长上下文默认保持开启,并与 OpenAI 账号开关取交集
迁移默认 false 会让存量分组静默丢掉 ≥200k 阶梯与渠道多档价。
OpenAI 路径传入 Group 后,账号级 openai_long_context_billing_enabled 被顶掉。

- 列默认 TRUE,并对存量行回填
- 分组价卡不再手填 token 区间,勾选走官方/预设阶梯
- 分组开关与账号开关 AND;价卡 JSON 损坏时打 warn
2026-08-13 08:38:10 +08:00
IanShaw027 8c4c3c09ce fix: 未登记的 Grok 文本模型回退 grok-4.5 价卡
新文本 ID 在闭集价卡上 fail-closed,请求成功而用量记 0。

- grok / grok-N* / build / composer 未登记别名继承 4.5 价
- imagine/voice/web-search/speech 等非文本族不回退
2026-08-13 08:38:10 +08:00
IanShaw027 678eb22a40 fix: Realtime 仅在观察到音频后计费,并修正标志位求值顺序
原先 elapsed>0 就出账,握手失败也会扣费;随后用 audioObserved,
但又在同一个 return 里先 Load 再收 errCh,标志位恒为 false,会话全部漏计。

- 先等中继结束再读 audioObserved
- 无音频或零时长不出账;每次连接独立 request id
2026-08-13 08:38:10 +08:00
IanShaw027 c4d883b8da feat: Chat 与 Responses 往返保留 x_search,并补 sources 抽取
Chat Completions 的 {"type":"x_search"} 会被 apicompat 丢掉,独立
/x_search 又没带 include 与结构化提示,主路径容易返回空结果仍计费。

- Chat↔Responses 保留 x_search 过滤字段与 tool_choice
- declared 只注册实际存活的 x_search,web_search 选择项仍丢弃
- 上游补 include x_search_call.action.sources 与结构化输出提示
2026-08-13 08:37:57 +08:00
IanShaw027 0de6d7e9ba feat: 新增独立 /x_search,走原生 x_search 并沿用搜索计费
Grok 分组原先只有 /web_search,无法带 handle/日期过滤,也无法走
xAI 的 x_search tool。

- POST /x_search(仅 Grok 分组),复用 web_search 的审计、failover 与按次计费
- 上游 Responses 强制 x_search;计费模型记为 grok-x-search
2026-08-13 07:49:19 +08:00
IanShaw027 363cc4994b fix: SuperGrokPro 用 4.5 窗口区分 Heavy,容量抖动只封单模型
JWT 的 SuperGrokPro 同时覆盖 SuperGrok 与 Heavy,账单月额度又滞后,
账号会误升/误降。multi-agent 容量抖动还会把整号提出调度。

- CanonicalGrokPlan:明确 JWT 优先;模糊档仅采新鲜的 grok-4.5 Responses 窗口
- 配额快照写入 plan_from_45_responses 时间戳,过期信号不用
- engine_overloaded 对 multi-agent 只封当前模型 0.5s–5min
2026-08-13 07:49:14 +08:00
IanShaw027 f3d9491071 feat: 分组支持逐模型定价,并可关闭长上下文阶梯
运营需要按分组覆盖渠道/内置价,且部分套餐不应自动吃 200k 倍率。
原先只能改渠道价卡,分组侧只剩 Voice 三列。

- groups 新增 model_pricing / long_context_pricing_enabled,解析链改为 Group → Channel → 内置
- 关闭长上下文时 token 模型只取最低档;video 按秒计费可写进同一价卡
- 回退价对齐官方卡:4.5 缓存 $0.30、4.3/imagine/audio/search 默认值一并校正
2026-08-13 07:49:08 +08:00
IanShaw027 a04ce49016 feat: 新增 grok-4.6 目录、官方定价与请求路径支持
官方目录价:<200k 为 $2 / 缓存读取 $0.50 / $6,≥200k 全部 2 倍。
原先没有独立价卡,/models 宣称可配 reasoning 但请求路径会剥掉 effort。

- 目录与别名接入 grok-4.6 / grok-4.6-latest,默认文本模型仍为 grok-4.5
- 独立回退价卡含缓存读取价与 200k 长上下文倍率
- 保留 reasoning_effort;Chat→Responses 的 cache/vision 桥接同样识别 4.6
2026-08-13 01:23:45 +08:00
IanShaw027 bb9e74285e feat: 从 JWT tier 识别 Grok 订阅档位,刷新后覆盖失效订阅
Grok Build access token 带有数字 tier claim(0=free、1=supergrok、5=heavy、6=lite),
原先只解 email/sub/team_id,refresh 还会把旧档位抄回去,订阅失效后一直显示 Heavy。

- 解码 JWT tier,并归一化 display / header 别名(含 free-tier、SuperGrok Lite)
- 仅从 access token 读取档位;新 JWT 覆盖已存凭证,AT 无 claim 时保留旧值
- 用量与 free-cache 判定以当前 AT 为准,账单月额度仅作降级
2026-08-13 01:23:37 +08:00
seng1eandClaude Opus 4.8 bba6a55e0f fix(backup): 定时备份加 leader 锁,避免多实例重复备份
仓库里所有周期任务都用 tryAcquireSingletonLeaderLock 选主、只让一个实例跑
(ops:*、dashboard:aggregation、subscription:expiry、payment:order:expiry 等),
唯独定时备份 runScheduledBackup 没接这套锁,每个实例都自己跑一遍。

多实例部署下同一时刻:
- 同一个库被 N 次 pg_dump;
- 内存峰值 ×N —— 归档整个读进内存再上传,内存吃紧的节点会直接 OOM;
- 上传的是同一个带时间戳的 key,N 份互相覆盖,白干一场连多余副本都留不下。

实测 3 节点:每天 backup_records 里三条一模一样的记录(同 key、同大小)。

改动:把 BackupService 接进和其它任务一样的 tryAcquireSingletonLeaderLock
(key=backup:scheduled:leader)。只锁定时备份,手动备份(CreateBackup/
StartBackup)不锁;无协调后端(cache/db 均 nil)时不加锁照常跑,单机/单测
行为不变;锁 TTL 35m > 备份自身 30m context 上限,防止大库 dump 中途锁过期。
wire_gen.go 由 wire 重新生成。

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-12 21:36:41 +08:00
github-actions[bot] ef4f99f292 chore: sync VERSION to 0.1.175 [skip ci] 2026-08-12 11:07:43 +00:00
shaw 04f8cdb194 fix: 修复 golangci-lint errcheck 和 gofmt 格式问题 2026-08-12 18:20:54 +08:00
shaw c0ab3a00ea feat: Codex OAuth 设备指纹收敛,减少上游可见的设备数和会话数
多人共享同一 OAuth 账号时,各用户 Codex 客户端携带各自不同的 installation_id/session_id/thread_id,
上游据此判定设备数和会话数并限制配额。本功能将这些标识改写为账号级恒定值。

四档策略(账号级 extra 字段 codex_fingerprint_mode):
- off: 不做任何收敛,原样透传
- device: 仅收敛 installation_id
- session(默认): 收敛 installation_id + session_id,thread_id 按客户端原始 session 派生
- full: 收敛所有标识(installation_id + session_id + thread_id)

改写覆盖 6 个指纹载体:x-codex-turn-metadata 头(JSON 内部字段)、x-codex-window-id、
x-codex-installation-id、x-client-request-id、session-id/session_id/thread-id、
请求体 client_metadata。头和体共享同一份预计算 IDs 确保 turn_id 等随机字段一致。
2026-08-12 17:57:50 +08:00
Lucky 814ecfba7c fix(group): invalidate the channel cache when a group's platform changes
The channel cache holds a groupID -> platform map with a 10 minute TTL, and
only channel Create/Update/Delete call invalidateCache(). Changing a group's
platform through the admin API therefore leaves the cache pointing at the old
platform for up to 10 minutes.

Channel pricing, model mapping and the model whitelist are all matched per
platform, so during that window the lookups silently miss: pricing falls back
to the global LiteLLM price list, renames stop applying and the whitelist
stops restricting. Nothing is logged.

Inject a narrow ChannelCacheInvalidator into the admin service (same shape as
the existing APIKeyAuthCacheInvalidator) and call it from UpdateGroup only when
the platform actually changed. The dependency is optional -- when it is nil the
cache simply rebuilds on TTL expiry, as before.
2026-08-12 03:22:06 +00:00
Lucky bd404c16f2 fix(channel): align pricing conflict detection with the pricing cache key
`validateNoConflictingModels` used `toModelEntry`, which only lowercases,
while `expandPricingToCache` keys the cache with
`normalizeChannelPricingModelName` (lowercase + TrimSpace + `.` -> `-`
for `claude-*`). Two pricing entries that the validator considers distinct
therefore collapse onto the same cache key, and the one written later
silently overwrites the other.

Add `toPricingModelEntry`, which reuses `normalizeChannelPricingModelName`,
and use it for pricing conflict detection. `toModelEntry` is left untouched
for model *mapping* validation, because `expandMappingToCache` keys the
mapping cache with plain `strings.ToLower` -- normalizing mappings the same
way would reject configurations that the cache keeps separate.

This completes the fix for #4754, which added the normalization to the
lookup side only.
2026-08-12 02:17:09 +00:00
Wesley Liddick 46cbb7187b Merge pull request #5531 from SamizuHM/fix/openai-compat-nested-data-usage
fix(openai-compat): parse usage from nested data envelopes
2026-08-12 10:01:09 +08:00
Wesley Liddick 80acae16fa Merge pull request #5525 from Fool0ntheHill/codex/fix-openai-visible-ttft
fix(openai): record Responses TTFT on visible output
2026-08-12 09:59:52 +08:00
Wesley Liddick 19c6007a22 Merge pull request #5342 from wucm667/fix/issue-5340-ws-v2-terminal-ttft
fix(openai-ws): exclude terminal events from TTFT
2026-08-12 09:59:43 +08:00
Wesley Liddick 0ed1a9f22a Merge pull request #5514 from wucm667/fix/issue-5510-cyber-policy-audit-scope
fix(audit): scope cyber policy events
2026-08-12 09:58:32 +08:00
Wesley Liddick a29fce4a61 Merge pull request #5511 from wucm667/fix/pr-5234-ws-audit-logging
fix(security-audit): restore websocket audit logs
2026-08-12 09:58:24 +08:00
Wesley Liddick 1225437099 Merge pull request #5502 from pyt111/codex/fix-account-stats-service-tier
fix: 修复 service tier 账号成本统计
2026-08-12 09:58:16 +08:00
Wesley Liddick 5192abb6b5 Merge pull request #5503 from fengshao1227/fix/openai-html-403-not-account-penalty
fix(openai): 上游 HTML 403 不再被当成账号级错误处罚账号
2026-08-12 09:58:08 +08:00
Wesley Liddick 40aae11888 Merge pull request #5513 from wucm667/fix/issue-5506-gemini-exclusive-minimum
fix(gemini): normalize exclusive minimum tool schemas
2026-08-12 09:57:52 +08:00
SamizuHM a163742fc9 fix(openai): preserve usage path precedence 2026-08-11 22:01:47 +08:00
SamizuHM 04dc540b23 fix(openai): parse nested data usage envelopes 2026-08-11 18:15:31 +08:00
Fool0ntheHill 900194fab2 fix(openai): 修正 Responses 可见输出 TTFT 2026-08-11 16:31:09 +08:00
wucm667 662444774f test(gemini): check cleaned schema assertions 2026-08-11 16:19:18 +08:00
wucm667 e24cb99b79 fix(openai-ws): retain no-delta TTFT fallback 2026-08-11 16:01:34 +08:00
Wesley Liddick 1e618dbc29 Merge pull request #5054 from wucm667/fix/issue-5029-openai-passthrough-pool-auth-retry
fix(openai): retry pool auth failures before failover
2026-08-11 14:21:45 +08:00
shaw a3bbf35cbd Merge branch 'main' into fix/issue-5029-openai-passthrough-pool-auth-retry
Resolve conflict in backend/internal/handler/openai_gateway_handler_test.go.

main and this branch each appended a passthrough upstream stub plus a test at
the same two insertion points:

  main   openAIHTTPPassthroughSSERateLimitUpstream
         TestOpenAIResponses_APIKeyPassthroughSSERateLimitUsesConfiguredPoolRetry
  branch openAIHTTPPassthroughAuthFailoverUpstream
         TestOpenAIResponses_APIKeyPassthroughPoolAuthFailureRetriesThenSwitchesToHealthyAccount

Both sides are kept verbatim; the only edit is giving each stub its own
calls() body instead of sharing the trailing one. No assertion was changed.

openai_gateway_passthrough.go and openai_oauth_passthrough_test.go merged
automatically.
2026-08-11 14:09:07 +08:00
Wesley Liddick caa1abb13a Merge pull request #5404 from wucm667/fix/issue-5400-oauth-image-stream-error
fix(openai): fail over OAuth image stream errors
2026-08-11 14:04:42 +08:00
Wesley Liddick 574dfad2dc Merge pull request #5488 from wucm667/fix/issue-5482-stale-codex-threshold
fix(openai): skip stale and reset Codex snapshots in scheduling threshold evaluator
2026-08-11 14:00:33 +08:00
Wesley Liddick 6876477371 Merge pull request #5304 from wucm667/fix/issue-5302-chat-reasoning-alias
fix(apicompat): accept chat reasoning alias
2026-08-11 13:59:49 +08:00
Wesley Liddick b918874f81 Merge pull request #5403 from cyhhao/fix/codex-capacity-exponential-backoff
fix(openai): back off capacity retries exponentially
2026-08-11 13:58:06 +08:00
Wesley Liddick 20a2d12dde Merge pull request #5415 from Yuxin-Qiao/fix/openai-responses-empty-completed-failover
fix(openai): fail over empty response.completed streams instead of recording 0/0 success
2026-08-11 13:53:43 +08:00
Wesley Liddick ca9e2b48ee Merge pull request #5413 from Yuxin-Qiao/fix/openai-responses-reasoning-item-id
fix(openai): strip invalid reasoning item IDs in API key passthrough
2026-08-11 13:53:05 +08:00
wucm667 6564d376e5 fix(audit): scope cyber policy events 2026-08-11 13:05:56 +08:00
wucm667 c8d9af6ce1 fix(gemini): normalize exclusive minimum tool schemas 2026-08-11 12:34:51 +08:00
wucm667 2d9920ba7d fix(security-audit): restore websocket audit logs 2026-08-11 11:56:46 +08:00