Commit Graph
5816 Commits
Author SHA1 Message Date
InCerryGit 22fc0cdbf5 fix(frontend): clarify OpenAI Fast/Flex policy rules 2026-08-16 23:49:21 +08:00
github-actions[bot] baeac1f3de chore: sync VERSION to 0.1.177 [skip ci] 2026-08-15 13:40:21 +00:00
Wesley Liddick 073e92d171 Merge pull request #5668 from Wei-Shaw/fix/codex-turn-state-and-fingerprint-optin
fix(openai): Codex 回合状态回传、v2 压缩探测与指纹收敛改为 opt-in
v0.1.177
2026-08-15 17:10:18 +08:00
shaw 4d9fedee20 fix(test): check type assertions in turn-state provenance tests
errcheck runs with check-type-assertions enabled, so the single-value
assertions on the provenance map values fail lint.
2026-08-15 16:44:30 +08:00
shaw fce41e318f fix(openai): make Codex fingerprint convergence opt-in and cover passthrough
Default codex_fingerprint_mode to off. v0.1.175 treated a missing key as
"session", so upgrading silently rewrote installation/session/thread/turn/
window identifiers for every existing OAuth account that had never configured
this field. The quota regressions in #5555, #5556 and #5582 line up with that
version boundary, with A/B reports that rolling back to v0.1.173 restores
quota. Convergence is now explicit opt-in (#5610).

Only accounts that never set the field change behaviour; explicit off /
device / session / full keep working exactly as configured. That required
flipping the persistence condition in all three account modals from
"!== 'session'" to "!== 'off'": the old rule deleted the key when it equalled
the default, which after the flip would have silently discarded an
administrator's explicit opt-in to session.

Also extend convergence to the passthrough path, which previously left client
identifiers untouched:

- resolve the ids once in forwardOpenAIPassthrough and rewrite
  client_metadata on the raw bytes (gjson extract + sjson splice) because
  passthrough is a hot path that must not fully unmarshal multi-MB bodies;
  a shared core keeps the raw and map variants from drifting
- both request builders apply the staged ids at the same relative position
  (after session isolation, before identity enforcement) so headers and body
  share one id set and turn_id stays consistent
- stage the ids unconditionally, including nil: a failover from a converged
  account to an off account must not leave the previous account's ids behind
2026-08-15 16:35:26 +08:00
shaw 8ae6d8f67e fix(openai): send session-level beta features and probe native compaction v2
OpenAI sunset the legacy unary /responses/compact endpoint (404, #5598,
#5624), so the account "compact probe" in the admin UI kept failing even for
healthy accounts, and the beta-feature negotiation header was only attached
to compaction turns.

Beta features (codex-rs session/mod.rs build_model_client_beta_features_header
+ client.rs build_responses_headers): the header is a session-level constant
attached to every /responses request, the WS handshake and /responses/compact.
Enumerating FEATURES shows no Experimental feature is enabled by default, so a
default install sends exactly "remote_compaction_v2". Mirror that:

- OAuth requests without a client-declared header get the default shape, so we
  no longer produce a "header only on compaction turns" pattern real Codex
  never emits (#5586 chains that strip the header)
- a client-declared header is preserved as-is: non-empty without v2 means the
  user disabled the feature and the gateway must not rewrite that
- native v2 turns (compaction_trigger in body) always ensure v2 is present
- non-OAuth upstreams keep the compaction-turn-only behaviour
- the WS injection sits outside the client-header copy block so prewarm and
  turn handshakes cannot land in different pool compatibility buckets

Compact probe now exercises native v2 (streaming /responses +
compaction_trigger) instead of the dead endpoint. Success requires an actual
compaction output item — scanning output_item.done/added, the terminal
response.output[] and the whole-JSON fallback — so a 2xx that silently drops
the trigger is reported as unsupported (the "got 0 items" class, #5478,
#5648). Probe identity is now UUID-shaped and applies the account's
convergence, matching real traffic on the same endpoint.
2026-08-15 16:35:08 +08:00
shaw 8219dcfc87 fix(openai): relay x-codex-turn-state and guard cross-account echo
Codex captures x-codex-turn-state from /responses SSE, /responses/compact
JSON and the WS handshake (codex-api sse/responses.rs, endpoint/compact.rs),
then echoes it back on later requests of the same turn. The HTTP path dropped
it because the header is not in the generic response allowlist, while the WS
path already relayed it — an inconsistency that broke the protocol chain.

Relay it explicitly at every commit point instead of widening the global
allowlist (which would leak it into Anthropic/Gemini responses):

- streaming, non-streaming and SSE-to-JSON handlers relay it, clearing any
  value left over by a previous failover attempt when upstream sends none
- under the first-output guard the header is only staged; provenance is
  recorded when applyAttemptResponseHeaders actually writes it, because a
  first-output timeout discards the staged headers and the client never
  receives that blob
- record (api key + client session) -> minting account, TTL-bounded with an
  opportunistic sweep, and strip echoes known to come from another account
  before they go upstream. Stripping only: injection is the Claude bridge's
  job. Same-account or unknown provenance passes through unchanged.
2026-08-15 16:34:47 +08:00
Wesley Liddick 1d3b9665c8 Merge pull request #5641 from InCerryGit/fix/issue-5624-remote-compaction-v2
fix(openai): preserve remote compaction v2 responses endpoint
2026-08-15 13:46:31 +08:00
Wesley Liddick c204d33b09 Merge pull request #5649 from lyen1688/feat/group-usage-daily-rollups
feat: 优化分组用量统计
2026-08-15 09:00:16 +08:00
lyen1688 45dcce0e49 修复:隔离分组用量仓储测试时区
分组用量仓储测试原先固定恢复上海时区,覆盖集成测试包初始化的 UTC,并导致后续今日统计多算。统一保存并恢复测试进入前的时区,避免 unit 和 integration 用例相互污染。
2026-08-14 23:48:03 +08:00
lyen1688 89d826be29 修复:解决分组用量 PR 的 CI 失败
恢复分组时区测试进入前的全局时区,避免污染高峰计费和支付租约测试。同步升级 Go 1.26.6 与 nanoid 3.3.18,清除后端和前端安全门禁。
2026-08-14 23:25:31 +08:00
lyen1688 cb7b03795d feat: 优化分组用量统计 2026-08-14 22:34:43 +08:00
InCerryGit a8b9ea22b7 fix(openai): separate native and legacy compaction routing 2026-08-14 18:18:53 +08:00
InCerryGit 9662cff2e7 fix(openai): preserve remote compaction v2 responses endpoint 2026-08-14 16:23:52 +08:00
Wesley Liddick fbfdcef818 Merge pull request #5573 from IanShaw027/fix/grok-long-context-and-media-fallback
fix(grok): 长上下文阶梯不再被 OpenAI 账号开关否决,媒体 ID 不继承文本价
2026-08-13 10:32:07 +08:00
github-actions[bot] 0e82efe489 chore: sync VERSION to 0.1.176 [skip ci] 2026-08-13 01:46:47 +00:00
Wesley Liddick e803e3851c Merge pull request #5559 from seng1e/fix/scheduled-backup-leader-lock
fix(backup): 定时备份加 leader 锁,避免多实例重复备份
v0.1.176
2026-08-13 09:24:11 +08:00
Wesley Liddick 5912ae960c Merge pull request #5543 from luckydududu/fix/invalidate-channel-cache-on-group-platform-change
fix(group): 改分组 platform 后失效渠道缓存(否则最长 10 分钟按旧平台计价)
2026-08-13 09:24:00 +08:00
Wesley Liddick e77b4b260e Merge pull request #5504 from fengshao1227/fix/responses-probe-inconclusive-keeps-unknown
fix(openai): 探测响应未跑完时不再落标为「上游不支持 Responses」
2026-08-13 09:23:50 +08:00
IanShaw027 e215c98c2c fix: 账号页自动刷新偏好改为模块初始化时恢复
onMounted 再读一次会覆盖已在 setup 阶段恢复的偏好。
2026-08-13 09:22:41 +08:00
IanShaw027 e29b93a1fb fix: 未知 Grok 文本兜底排除 image/video/audio 等媒体族
带版本号的媒体 ID(grok-2-image-1212、grok-5-video)会被
grok-N* 白名单当成文本,错吃 4.5 token 价。vision 多模态
仍按 token 计;HasIdentifiedTokenPricing 也不放行这些 ID。
2026-08-13 09:22:38 +08:00
IanShaw027 fd82dfd52d fix: Grok 长上下文只跟分组开关,不受 OpenAI 账号开关否决
openai_long_context_billing_enabled 是 OpenAI 账号设置,Grok
账号无法打开。按账号硬传 false 会让分组开关失效,官方 200k
阶梯和渠道多档区间全部塌到第一档。非 OpenAI 不再传账号门闩。
2026-08-13 09:22:34 +08:00
Wesley Liddick 4853662f99 Merge pull request #5540 from luckydududu/fix/channel-pricing-model-name-normalization
fix(channel): 定价冲突检测与定价缓存键的归一化对齐(#4754 修复残留)
2026-08-13 09:12:41 +08:00
Wesley Liddick 0fa577a19c Merge pull request #5571 from IanShaw027/feat/grok-jwt-tier-and-4.6
feat(grok): 以 JWT tier 识别订阅档位并接入 grok-4.6,修正徽章滞后、未知模型零计费与搜索/Voice 边界
2026-08-13 09:08:35 +08:00
IanShaw027 b61e4bcc45 fix: 对齐 JWT 常量 gofmt,并补全 available groups 契约字段
golangci-lint 要求 const 块按类型对齐;分组 DTO 新增
long_context_pricing_enabled 后,契约夹具未同步导致比较失败。
2026-08-13 08:55:56 +08:00
IanShaw027 0ae151a238 fix: 徽章改读 grok_usage_snapshot,增量刷新比较 Grok 快照
后端写入 grok_usage_snapshot,列表却读无 writer 的 grok_quota_snapshot;
增量刷新不比 Grok 快照,档位更新后行不替换。

- 正式快照优先 usage,quota 仅兜底;计划字段只收非空字符串
- buildGrokUsageRefreshKey 接入自动刷新
- 注明网关 Extra 写入依赖每请求独立 Account 副本
2026-08-13 08:38:10 +08:00
IanShaw027 b830bc14d6 fix: 长上下文默认保持开启,并与 OpenAI 账号开关取交集
迁移默认 false 会让存量分组静默丢掉 ≥200k 阶梯与渠道多档价。
OpenAI 路径传入 Group 后,账号级 openai_long_context_billing_enabled 被顶掉。

- 列默认 TRUE,并对存量行回填
- 分组价卡不再手填 token 区间,勾选走官方/预设阶梯
- 分组开关与账号开关 AND;价卡 JSON 损坏时打 warn
2026-08-13 08:38:10 +08:00
IanShaw027 8c4c3c09ce fix: 未登记的 Grok 文本模型回退 grok-4.5 价卡
新文本 ID 在闭集价卡上 fail-closed,请求成功而用量记 0。

- grok / grok-N* / build / composer 未登记别名继承 4.5 价
- imagine/voice/web-search/speech 等非文本族不回退
2026-08-13 08:38:10 +08:00
IanShaw027 678eb22a40 fix: Realtime 仅在观察到音频后计费,并修正标志位求值顺序
原先 elapsed>0 就出账,握手失败也会扣费;随后用 audioObserved,
但又在同一个 return 里先 Load 再收 errCh,标志位恒为 false,会话全部漏计。

- 先等中继结束再读 audioObserved
- 无音频或零时长不出账;每次连接独立 request id
2026-08-13 08:38:10 +08:00
IanShaw027 c4d883b8da feat: Chat 与 Responses 往返保留 x_search,并补 sources 抽取
Chat Completions 的 {"type":"x_search"} 会被 apicompat 丢掉,独立
/x_search 又没带 include 与结构化提示,主路径容易返回空结果仍计费。

- Chat↔Responses 保留 x_search 过滤字段与 tool_choice
- declared 只注册实际存活的 x_search,web_search 选择项仍丢弃
- 上游补 include x_search_call.action.sources 与结构化输出提示
2026-08-13 08:37:57 +08:00
IanShaw027 0de6d7e9ba feat: 新增独立 /x_search,走原生 x_search 并沿用搜索计费
Grok 分组原先只有 /web_search,无法带 handle/日期过滤,也无法走
xAI 的 x_search tool。

- POST /x_search(仅 Grok 分组),复用 web_search 的审计、failover 与按次计费
- 上游 Responses 强制 x_search;计费模型记为 grok-x-search
2026-08-13 07:49:19 +08:00
IanShaw027 363cc4994b fix: SuperGrokPro 用 4.5 窗口区分 Heavy,容量抖动只封单模型
JWT 的 SuperGrokPro 同时覆盖 SuperGrok 与 Heavy,账单月额度又滞后,
账号会误升/误降。multi-agent 容量抖动还会把整号提出调度。

- CanonicalGrokPlan:明确 JWT 优先;模糊档仅采新鲜的 grok-4.5 Responses 窗口
- 配额快照写入 plan_from_45_responses 时间戳,过期信号不用
- engine_overloaded 对 multi-agent 只封当前模型 0.5s–5min
2026-08-13 07:49:14 +08:00
IanShaw027 f3d9491071 feat: 分组支持逐模型定价,并可关闭长上下文阶梯
运营需要按分组覆盖渠道/内置价,且部分套餐不应自动吃 200k 倍率。
原先只能改渠道价卡,分组侧只剩 Voice 三列。

- groups 新增 model_pricing / long_context_pricing_enabled,解析链改为 Group → Channel → 内置
- 关闭长上下文时 token 模型只取最低档;video 按秒计费可写进同一价卡
- 回退价对齐官方卡:4.5 缓存 $0.30、4.3/imagine/audio/search 默认值一并校正
2026-08-13 07:49:08 +08:00
IanShaw027 69648476d4 fix: 账号徽章与用量格按实时档位展示,避免账单滞后误判
订阅失效刷新后凭证已是 free,但残留的 monthly_limit / usage_percent
仍会把用量格锁在付费 7d/30d,徽章也优先信滞后的 billing.plan。

- 凭证 subscription_tier 优先于账单/配额快照
- 用量格先看实时档位,再回退账单指标;Lite 仍走付费条
- 徽章补齐 SuperGrok Lite / Plus / X Basic
2026-08-13 01:23:54 +08:00
IanShaw027 a04ce49016 feat: 新增 grok-4.6 目录、官方定价与请求路径支持
官方目录价:<200k 为 $2 / 缓存读取 $0.50 / $6,≥200k 全部 2 倍。
原先没有独立价卡,/models 宣称可配 reasoning 但请求路径会剥掉 effort。

- 目录与别名接入 grok-4.6 / grok-4.6-latest,默认文本模型仍为 grok-4.5
- 独立回退价卡含缓存读取价与 200k 长上下文倍率
- 保留 reasoning_effort;Chat→Responses 的 cache/vision 桥接同样识别 4.6
2026-08-13 01:23:45 +08:00
IanShaw027 bb9e74285e feat: 从 JWT tier 识别 Grok 订阅档位,刷新后覆盖失效订阅
Grok Build access token 带有数字 tier claim(0=free、1=supergrok、5=heavy、6=lite),
原先只解 email/sub/team_id,refresh 还会把旧档位抄回去,订阅失效后一直显示 Heavy。

- 解码 JWT tier,并归一化 display / header 别名(含 free-tier、SuperGrok Lite)
- 仅从 access token 读取档位;新 JWT 覆盖已存凭证,AT 无 claim 时保留旧值
- 用量与 free-cache 判定以当前 AT 为准,账单月额度仅作降级
2026-08-13 01:23:37 +08:00
seng1eandClaude Opus 4.8 bba6a55e0f fix(backup): 定时备份加 leader 锁,避免多实例重复备份
仓库里所有周期任务都用 tryAcquireSingletonLeaderLock 选主、只让一个实例跑
(ops:*、dashboard:aggregation、subscription:expiry、payment:order:expiry 等),
唯独定时备份 runScheduledBackup 没接这套锁,每个实例都自己跑一遍。

多实例部署下同一时刻:
- 同一个库被 N 次 pg_dump;
- 内存峰值 ×N —— 归档整个读进内存再上传,内存吃紧的节点会直接 OOM;
- 上传的是同一个带时间戳的 key,N 份互相覆盖,白干一场连多余副本都留不下。

实测 3 节点:每天 backup_records 里三条一模一样的记录(同 key、同大小)。

改动:把 BackupService 接进和其它任务一样的 tryAcquireSingletonLeaderLock
(key=backup:scheduled:leader)。只锁定时备份,手动备份(CreateBackup/
StartBackup)不锁;无协调后端(cache/db 均 nil)时不加锁照常跑,单机/单测
行为不变;锁 TTL 35m > 备份自身 30m context 上限,防止大库 dump 中途锁过期。
wire_gen.go 由 wire 重新生成。

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-12 21:36:41 +08:00
shaw 5935e674a8 chore: update sponsors 2026-08-12 19:17:04 +08:00
github-actions[bot] ef4f99f292 chore: sync VERSION to 0.1.175 [skip ci] 2026-08-12 11:07:43 +00:00
Wesley Liddick 93c32fa1a2 Merge pull request #5553 from Wei-Shaw/feat/codex-fingerprint-convergence
feat: Codex OAuth 设备指纹收敛
v0.1.175
2026-08-12 18:52:01 +08:00
shaw 04f8cdb194 fix: 修复 golangci-lint errcheck 和 gofmt 格式问题 2026-08-12 18:20:54 +08:00
shaw c0ab3a00ea feat: Codex OAuth 设备指纹收敛,减少上游可见的设备数和会话数
多人共享同一 OAuth 账号时,各用户 Codex 客户端携带各自不同的 installation_id/session_id/thread_id,
上游据此判定设备数和会话数并限制配额。本功能将这些标识改写为账号级恒定值。

四档策略(账号级 extra 字段 codex_fingerprint_mode):
- off: 不做任何收敛,原样透传
- device: 仅收敛 installation_id
- session(默认): 收敛 installation_id + session_id,thread_id 按客户端原始 session 派生
- full: 收敛所有标识(installation_id + session_id + thread_id)

改写覆盖 6 个指纹载体:x-codex-turn-metadata 头(JSON 内部字段)、x-codex-window-id、
x-codex-installation-id、x-client-request-id、session-id/session_id/thread-id、
请求体 client_metadata。头和体共享同一份预计算 IDs 确保 turn_id 等随机字段一致。
2026-08-12 17:57:50 +08:00
Lucky 814ecfba7c fix(group): invalidate the channel cache when a group's platform changes
The channel cache holds a groupID -> platform map with a 10 minute TTL, and
only channel Create/Update/Delete call invalidateCache(). Changing a group's
platform through the admin API therefore leaves the cache pointing at the old
platform for up to 10 minutes.

Channel pricing, model mapping and the model whitelist are all matched per
platform, so during that window the lookups silently miss: pricing falls back
to the global LiteLLM price list, renames stop applying and the whitelist
stops restricting. Nothing is logged.

Inject a narrow ChannelCacheInvalidator into the admin service (same shape as
the existing APIKeyAuthCacheInvalidator) and call it from UpdateGroup only when
the platform actually changed. The dependency is optional -- when it is nil the
cache simply rebuilds on TTL expiry, as before.
2026-08-12 03:22:06 +00:00
Lucky bd404c16f2 fix(channel): align pricing conflict detection with the pricing cache key
`validateNoConflictingModels` used `toModelEntry`, which only lowercases,
while `expandPricingToCache` keys the cache with
`normalizeChannelPricingModelName` (lowercase + TrimSpace + `.` -> `-`
for `claude-*`). Two pricing entries that the validator considers distinct
therefore collapse onto the same cache key, and the one written later
silently overwrites the other.

Add `toPricingModelEntry`, which reuses `normalizeChannelPricingModelName`,
and use it for pricing conflict detection. `toModelEntry` is left untouched
for model *mapping* validation, because `expandMappingToCache` keys the
mapping cache with plain `strings.ToLower` -- normalizing mappings the same
way would reject configurations that the cache keeps separate.

This completes the fix for #4754, which added the normalization to the
lookup side only.
2026-08-12 02:17:09 +00:00
Wesley Liddick 4ec9ceec4a Merge pull request #5508 from pcmid/fix/show-security-audit-menu-in-simple-mode
fix(frontend): show security audit menu in simple mode
2026-08-12 10:01:20 +08:00
Wesley Liddick 46cbb7187b Merge pull request #5531 from SamizuHM/fix/openai-compat-nested-data-usage
fix(openai-compat): parse usage from nested data envelopes
2026-08-12 10:01:09 +08:00
Wesley Liddick 80acae16fa Merge pull request #5525 from Fool0ntheHill/codex/fix-openai-visible-ttft
fix(openai): record Responses TTFT on visible output
2026-08-12 09:59:52 +08:00
Wesley Liddick 19c6007a22 Merge pull request #5342 from wucm667/fix/issue-5340-ws-v2-terminal-ttft
fix(openai-ws): exclude terminal events from TTFT
2026-08-12 09:59:43 +08:00
Wesley Liddick 0ed1a9f22a Merge pull request #5514 from wucm667/fix/issue-5510-cyber-policy-audit-scope
fix(audit): scope cyber policy events
2026-08-12 09:58:32 +08:00
Wesley Liddick a29fce4a61 Merge pull request #5511 from wucm667/fix/pr-5234-ws-audit-logging
fix(security-audit): restore websocket audit logs
2026-08-12 09:58:24 +08:00