Commit Graph
4394 Commits
Author SHA1 Message Date
Wesley Liddick 58ccea4eaa Merge pull request #5767 from hansnow/fix/ws-http-bridge-custom-tools
fix(openai): 补齐客户端工具终止事件恢复
2026-08-18 16:21:50 +08:00
hansnow c253bd2c72 fix(openai): restore client tools in terminal events 2026-08-18 15:48:02 +08:00
Wesley Liddick 26cb59df05 Merge pull request #5764 from hansnow/fix/ws-http-bridge-custom-tools
fix(openai): 补齐 WS HTTP bridge 的客户端工具适配
2026-08-18 15:47:59 +08:00
Wesley Liddick 58ea46e894 Merge pull request #5661 from wucm667/fix/issue-5659-openai-custom-tools
fix(openai): restore API-key custom tool calls
2026-08-18 15:47:50 +08:00
Wesley Liddick 1ed3b6aef6 Merge pull request #5760 from spongehah/feature/unify-codex-outbound-identity
fix(Fingerprint): 将 Codex 非推理出站身份统一到推理解析链
2026-08-18 15:35:34 +08:00
Wesley Liddick f211a630c8 Merge pull request #5720 from tamseno/fix/invitation-code-toctou-race
fix(auth): make invitation code consumption atomic with user creation
2026-08-18 15:35:16 +08:00
Wesley Liddick 37732dcd34 Merge pull request #5725 from tamseno/fix/gemini-include-server-side-tool-invocations
fix(gemini): support includeServerSideToolInvocations in GeminiToolConfig
2026-08-18 15:35:01 +08:00
yaxin a341239596 fix(fingerprint): align credential-face identity with the real client and de-drift models version
- Replace ApplyCodexCanonicalIdentity with CodexCanonicalAuthIdentity /
  ApplyCodexCanonicalAuthIdentity: the credential face (auth.openai.com
  token exchange / refresh / PAT whoami) now sends the originator +
  canonical User-Agent pair and no version header, matching codex-rs
  default_headers(); the version gate (#3901) only exists on the
  /backend-api/codex inference face. whoami keeps its original header
  shape (originator + UA) with the canonical UA source.
- Token exchange and refresh send the full pair instead of a bare UA,
  eliminating the half-identity (UA without originator) combination no
  real client ever emits.
- Codex models manifest: the Version header now follows the client's
  own client_version when it is valid and >= the upstream floor (same
  source as the query param, restoring the pre-refactor consistency),
  falling back to the canonical version otherwise; the query param
  keeps its verbatim passthrough contract.
- Drop the now-unreferenced openAICodexProbeVersion constant and its
  vacuous consistency assertions; probes resolve their version through
  resolveCodexOutboundIdentity at runtime.
2026-08-18 15:20:08 +08:00
yaxin 1ba92449c7 fix(gemini): wire includeServerSideToolInvocations into the typed transform path
The struct field alone never reached the wire: the raw passthrough
pipeline is covered by enableMixedGeminiToolInvocations (#5711), but
TransformClaudeToGeminiWithOptions builds GeminiToolConfig from scratch
and never set the flag, so gemini-* models entering through the Claude
format gateway could still hit the upstream 400 from issue #5709.

- Set IncludeServerSideToolInvocations=true when the built tool
  declarations mix functionDeclarations with googleSearch, matching the
  raw-path injection semantics.
- Replace the marshal-roundtrip-only test with behavior tests that
  drive TransformClaudeToGeminiWithOptions: mixed tools set the flag,
  function-only and web-search-only requests leave it unset.
2026-08-18 15:05:54 +08:00
yaxin 9617775f9a fix(repo): tolerate ErrTxStarted for tx-bound clients and harden test stubs
- user_repo.create(): keep the TxFromContext fast path, but restore
  tolerance for dbent.ErrTxStarted in the self-owned-transaction branch.
  ent's Client.Tx only inspects the driver type, so a repository built
  from a tx-bound client (client-injected transactions, e.g. the
  integration fixture testEntTx + tx.Client()) hits ErrTxStarted; reuse
  that client instead of failing. Fixes the two red integration tests in
  allowed_groups_contract_integration_test.go.
- createUserAndClaimInvitation: roll back via defer (matching the OAuth
  registration precedent) so a panic inside the transaction cannot leak
  the connection.
- settingRepoStub: guard call counters and state with a mutex; the new
  concurrency regression test exercises it from multiple goroutines and
  the unsynchronized counters were flagged by -race.
2026-08-18 15:02:25 +08:00
Wesley Liddick 1870b58c1d Merge pull request #5721 from lyy0709/codex/bulk-openai-settings
fix(openai): complete bulk account settings
2026-08-18 14:53:55 +08:00
hansnow 7e579cb28d fix(openai): adapt client tools in WS HTTP bridge 2026-08-18 14:42:30 +08:00
Wesley Liddick 938f1868ae Merge pull request #5714 from wucm667/fix/issue-2733-current-main
fix(ops): avoid single-insert fallback after batch failure
2026-08-18 14:11:42 +08:00
Wesley Liddick 4d19836189 Merge pull request #5716 from wucm667/fix/issue-2695-current-main
fix: skip expiry reminders without SMTP config
2026-08-18 14:06:32 +08:00
Wesley Liddick 1ea4150bf0 Merge pull request #5581 from wucm667/fix/issue-5574-passthrough-model-discovery
fix(gateway): align passthrough model discovery
2026-08-18 13:56:02 +08:00
Wesley Liddick ed2da82396 Merge pull request #5711 from wucm667/fix/issue-5709-antigravity-tool-config
fix(antigravity): preserve mixed Gemini tool config
2026-08-18 13:55:24 +08:00
Wesley Liddick 6259940ef2 Merge pull request #5669 from feeeei/main
feat(openai): OpenAI Team 联动熔断
2026-08-18 13:54:32 +08:00
Wesley Liddick baaf59d417 Merge pull request #5755 from feeeei/feat/gemini
fix(gemini): Skipped 错误策略对齐 OpenAI,上游 4xx 不再硬改 500
2026-08-18 13:54:06 +08:00
Wesley Liddick c0325d24f9 Merge pull request #5759 from o2e/codex/fix-codex-usage-probe-model-clean
[codex] 修复部分账号 Codex 额度查询 400
2026-08-18 13:53:49 +08:00
Wesley Liddick 1a3ecd2b94 Merge pull request #5004 from wucm667/fix/issue-4990-deferred-tool-cache-control
fix(claude): strip cache control from deferred tools
2026-08-18 13:50:57 +08:00
Wesley Liddick cddb03c0f1 Merge pull request #5609 from wucm667/fix/issue-5607-auth-pricing-snapshot
fix: preserve group pricing in auth snapshots
2026-08-18 13:50:13 +08:00
Wesley Liddick a7a321232f Merge pull request #5567 from wucm667/fix/issue-5563-anthropic-sse-overload
fix(gateway): handle Anthropic SSE overload errors
2026-08-18 13:49:59 +08:00
o2e 16e4f7ecc3 修复 Codex 额度探针模型兼容性 2026-08-18 13:08:28 +08:00
spongehah bb6c3b4f6a fix: unify Codex OAuth outbound identity onto the inference resolver
Token exchange, PAT whoami, models, probes, and pre-writes now follow
the same UA/version chain as Codex inference instead of hardcoded
codex-cli/0.91.0 or compile-time constants.
2026-08-18 13:03:22 +08:00
feeeei ab0fcd1a0e fix(gemini): Skipped 错误策略对齐 OpenAI,上游 4xx 不再硬改 500
ErrorPolicySkipped(池模式、或自定义错误码未命中)原来在响应写出上
自成一派:v1beta 原生把上游 4xx 硬改 500 后原文透传,/v1/messages
硬传 500 进映射(客户端拿到 502)。下游网关据此把请求级错误当可重
试的服务端故障反复换号,耗尽后改写成 All available accounts
exhausted(2026-08-17 gemini 生产事故链)。现对齐 OpenAI 路径语义:

- Skipped 只豁免账号状态标记,不豁免换号:可 failover 状态码一律
  返回 UpstreamFailoverError(poolModeSkippedFailoverError 泛化为
  skippedErrorPolicyFailoverError,同账号重试标记仍仅池模式携带)
- 池模式的不可 failover 4xx 保真:v1beta 原码+原文透传(新
  writeGeminiNativeUpstreamError 与 ErrorPolicyNone 共用同一写出,
  并补记 ops 事件),/v1/messages 与 chat completions 按真实状态
  码映射
- 自定义错误码未命中且不可 failover:三路径统一 500 + "Upstream
  gateway error" 固定文案,上游细节仅记 ops 错误日志
- 400 属确定性请求错误:mapped 写出回传脱敏后的上游 message,客户
  端可据此定位非法字段
2026-08-18 11:24:39 +08:00
Wesley Liddick 7d633f5fc3 Merge pull request #5742 from heathermhuang/codex/fix-grok-response-model-audit
fix: normalize Grok response model audit aliases
2026-08-18 10:03:26 +08:00
Wesley Liddick b2d1c3859a Merge pull request #5738 from okbexx/fix/codex-identity-snapshot
fix(openai): make Codex convergence identity consistent
2026-08-18 09:57:05 +08:00
shaw 6bf335965a merge main 并修复与 #5730 的语义冲突
main 侧 #5730 新增的 openai_gateway_cn_fixes_test.go 按旧 11 参签名调用
calculateOpenAIRecordUsageCost;本分支为该函数新增了第 12 个参数
pricingAt。文本无冲突但 test build 会失败,此处按本分支对同类测试
调用点的既有处理方式补传 time.Time{}。
2026-08-17 22:27:22 +08:00
Heatherm Huang c46d07ca07 fix: normalize Grok response model audit aliases 2026-08-17 21:15:12 +08:00
Jarl 6793d5ac85 fix(openai): make Codex convergence identity consistent 2026-08-17 20:25:28 +08:00
lyen1688 9f24a55305 功能:支持渠道模型分时倍率定价 2026-08-17 19:45:07 +08:00
shaw 10c8b70203 fix(cn-providers): 修复 CN 分组五项功能缺陷(调度闸门/计费/断开漏记/count_tokens/403)
对已合并 PR #5666 + 分组入口放行后的全量功能审计发现的 P0/P1 修复,
全部先经代码与厂商文档实证再实施:

1. /v1/messages 调度闸门(P0):sanitizeGroupMessagesDispatchFields 对非
   openai 平台恒置 AllowMessagesDispatch=false,而闸门豁免名单只有 grok,
   CN 分组经正常途径创建后恒 403——原生 Anthropic 直通(Claude Code 主用例)
   完全不可达。修复:闸门对 CN 与 grok 同语义豁免;count_tokens 处的内联
   裸检查统一走同一 helper;ResolveMessagesDispatchModel 对 CN 早退,避免
   openai 专属的 gpt-5.x 默认映射发给 CN 上游。

2. 计费候选链(P0):候选链兜底含客户端原始模型名,配合 getFallbackPricing
   的 claude→Sonnet 统一兜底,映射的 CN 模型无价时 CN 流量会按 Claude 原价
   (数倍~数十倍)静默误计,且 usage 日志显示 claude-* 名无从察觉。修复:
   CN 账号的 claude-* 候选仅在显式分组/渠道定价时放行;候选全滤空时按
   ErrModelPricingUnavailable 走零成本+告警落账(顺带修复原空候选错误会
   丢弃整条 usage 记录的次生问题)。

3. 断开/中断漏记(P0,#5148 对齐,惠及 openai 平台):messages/responses/
   chat_completions 三个 handler 的错误路径此前在 err!=nil 时丢弃携带的
   部分 result——客户端断开排水后的完整 usage 被丢,payg 上游照常计费而
   平台漏记(anthropic 网关早有同修复,openai 网关缺失)。修复:错误路径
   result 非空时照常提交 usage;failover 错误恒 result=nil 无重复计费。
   Responses×anthropic 流式转换器同时改为断开后继续排水至流自然结束
   (末尾 message_delta 的 output_tokens 不再丢),finalize 帧补工具名反转
   与客户端工具还原、仅在客户端仍连接时写出。

4. count_tokens(P1,证据修正):经实证三家 Anthropic 兼容层均无
   /v1/messages/count_tokens(DeepSeek 官方文档无此端点且注明
   anthropic-version 被忽略;OpenModel 标注该端点 Anthropic only),
   anthropic 协议转发上游=常态 404,且错误处置缺模型上下文会把不计费的
   探测放大成整账号停调。修复:CN 全协议一律本地 tiktoken 估算(与 Grok
   同方案),删除上游转发死代码。

5. 403 处置(P1):CN 此前落入通用 handleAuthError,单次 HTML 403(CDN/
   代理拦截页)即永久禁用,且 403 在 failover 集里会逐账号重放连环禁用
   整组。修复:CN 与 openai 同口径——HTML 豁免 + 3 次累计 + 临时冷却。

新增回归测试 8 项:闸门豁免(含 openai 仍受控断言)、CN 调度映射空返回、
候选过滤三态、空候选零成本落账、断开排水 usage 完整性、HTML-403 零处罚、
结构化 403 首次临时停调。handler/service 全包测试通过。
2026-08-17 17:28:53 +08:00
shaw 7cdca9e495 feat(groups): 放行 kimi/zhipu/deepseek 平台分组创建入口
PR #5666 引入 CN 平台后,路由/调度/前端类型均已支持 CN 平台分组,但分组
创建入口两头缺失:后端 Create/UpdateGroupRequest 的 platform oneof 白名单
与前端 GroupsView 平台选项都没有三平台,导致 CN 账号「无可用分组」、整条
流量链路不通(composite 不能作为替代:CN 不可为 composite 路由目标)。

- group_handler.go: 两处 oneof 加 kimi/zhipu/deepseek;composite 路由目标
  白名单有意不动(DetectModelPlatform/isConcreteRequestPlatform 均无 CN 分支)
- GroupsView: platformOptions/platformFilterOptions 补三项;两处徽章配色链
  按 platformColors.ts 色系补 CN 分支
- i18n: admin.groups.platforms 补 kimi/zhipu/deepseek 键(zh/en),缺键时
  分组徽章/GroupRPM/RateMultipliers 弹窗/ChannelsView 会渲染原始 key
- GroupBadge: badgeClass/labelClass 补 CN 配色
- 新增表驱动测试:9 平台 Create/Update 全放行、非法值(别名/大小写/空格)
  全拒绝、composite target 对 CN 保持拒绝的守卫
2026-08-17 17:28:51 +08:00
wucm667 971544570d test(antigravity): check tool config assertions 2026-08-17 17:11:33 +08:00
Tamseno 3c3bb2fa19 fix(gemini): support includeServerSideToolInvocations in GeminiToolConfig
- Add IncludeServerSideToolInvocations field to GeminiToolConfig to prevent dropping client tool settings.
- Fix HTTP 400 error when mixing built-in tools (e.g. Google Search) with function calling on Gemini 3.6/3.7 models.
- Add serialization/deserialization unit test TestGeminiToolConfig_IncludeServerSideToolInvocations.

Fixes #5709
2026-08-17 15:23:49 +08:00
Wesley Liddick e330c243a8 Merge pull request #5666 from Randark-JMT/feat/cn-providers-kimi-zhipu-deepseek
feat: 国产供应商多协议支持(Kimi/Zhipu/DeepSeek 原生 Anthropic 直通 + DeepSeek Responses)与配额/余额监控
2026-08-17 14:55:20 +08:00
Tamseno b8642ef674 fix(auth): make invitation code consumption atomic with user creation
RegisterWithVerification checked CanUse() and then marked the code used in
two separate, non-transactional steps; the second step's failure was
swallowed ("invitation code mark failure does not affect registration").
Concurrent registrations with the same invitation code could all pass the
check and each create an account, turning a one-time invitation code into
an unlimited account factory (TOCTOU race).

Fix:
- AuthService: create user and claim the invitation code inside one DB
  transaction (createUserAndClaimInvitation). The claim reuses
  redeemRepo.Use's conditional UPDATE (WHERE status='unused'); losers are
  rejected with INVITATION_CODE_INVALID and their transaction (including
  the user insert) is rolled back. No-code registration path unchanged.
- userRepository.create: explicitly join an outer ent transaction via
  TxFromContext instead of relying on Client.Tx returning ErrTxStarted
  (ent's Tx never inspects the context, so the old reuse branch was dead
  code and user inserts always committed in their own transaction,
  leaving orphan users behind when the outer transaction rolled back).

Regression tests:
- unit: concurrent register with one invitation code must succeed exactly
  once (8 goroutines -> 1 success, 7 x INVITATION_CODE_INVALID)
- integration: outer-tx rollback removes user and releases the claim;
  commit persists both atomically
2026-08-17 13:49:46 +08:00
lyy0709 76b70b1685 fix(openai): validate bulk account settings 2026-08-17 13:34:41 +08:00
wucm667 cb5e03a720 fix(antigravity): preserve mixed Gemini tool config 2026-08-17 12:36:58 +08:00
Randark e728545382 style: gofmt 迁移测试注释(CI golangci-lint gofmt) 2026-08-17 04:26:22 +00:00
Randark 4b667ccd45 fix(review): 处理 PR #5666 四个阻断项(B1-B4)
- B1: 移除根目录 docker-compose.yml(个人镜像测试产物混入)
- B2: 新增迁移 224 放宽 user_platform_quotas CHECK 至 8 平台,补
  BulkInsertInitial 国产平台集成测试与迁移内容断言测试
- B3: CC/Responses×anthropic 四个上游读循环接入间隔超时泵
  (gateway.stream_data_interval_timeout,默认 180s):上游挂住 SSE
  不发数据也不断连时结束排水/组装、关闭 resp.Body 归还连接池位,
  排水超时仍按已累计 usage 返回(与 messages 主路径同语义)
- B4: 配额/余额探测发起前过 cnValidateProbeURL(与网关转发同一套
  security.url_allowlist 策略),被拒时零出站、API key 不离开本机
2026-08-17 04:07:06 +00:00
wucm667 79c2eb5020 fix: skip expiry reminders without SMTP config 2026-08-17 10:59:43 +08:00
wucm667 5f19433103 fix(ops): avoid single-insert fallback after batch failure 2026-08-17 10:52:57 +08:00
Lucky 11e1e22882 fix(docker): bump Go builder image to 1.26.6 to match go.mod
89d826be2 raised backend/go.mod to `go 1.26.6` and updated the three CI
workflows' version assertions, but left the Go builder image in all three
Dockerfiles pinned at 1.26.5. Since the official golang images set
GOTOOLCHAIN=local, the toolchain is not auto-downloaded and any image build
fails hard at `go mod download`.

CI does not catch this: the workflows build with actions/setup-go, not with
these Dockerfiles.

Also extend the Go-upgrade checklist in DEV_GUIDE.md, which listed only the
CI files -- that omission is why the Dockerfiles were missed.
2026-08-16 06:17:43 +00:00
github-actions[bot] baeac1f3de chore: sync VERSION to 0.1.177 [skip ci] 2026-08-15 13:40:21 +00:00
Randark 901a0439f1 feat: 国产供应商一等支持(Kimi/Zhipu/DeepSeek 多协议 + 配额/余额监控)
后端:
- 协议凭证维度 credentials[api_protocol] ∈ chat_completions(默认)/anthropic/responses(deepseek)
- /v1/messages 零转换直通原生 Anthropic 端点(kimi/zhipu/deepseek),CC/Responses
  入站交叉组合走 apicompat 双向转换链(responses/chat_completions anthropic-native 转发器)
- count_tokens:anthropic 协议透传原生端点;其余 CN 协议本地 tiktoken 估算
- Coding Plan 额度探测(5h/weekly 滚动窗口)+ payg 余额探测(kimi/deepseek),
  deepseek 双币种 CNY+USD 明细,任一币种达标不停调
- 周期任务 [CNBalance] 并发探测 + 预算随工作量放大;响应式 429 冷却到最早窗口
  重置点;余额不足可恢复临时停调;智谱 CREDIT_LIMIT 不污染窗口解析
- CC→anthropic 流式客户端断开后继续排水上游保住 usage 计量

前端:
- 创建/编辑弹窗 account_mode + api_protocol + base_url 联动预设(含 watcher 竞态防护)
- 用量单元格:kimi/zhipu coding 显示 5h/weekly 窗口,kimi/deepseek payg 显示余额,
  多币种并列展示;探测失败保留快照;挂载自动探测 5min 去抖
- 调度阈值设置面板补 kimi/zhipu 平台(对齐后端 AllowedSchedulingThresholdPlatforms)
2026-08-15 10:37:51 +00:00
feeeei d677d67dda feat: OpenAI Team 联动熔断
OpenAI OAuth 账户收到 402 deactivated_workspace(ChatGPT Team 工作区被停用)时,
将同一 Team(credentials.chatgpt_account_id 相同)的其余 active 账户一并置为
error 并进程内立即熔断,让调度瞬间切走,避免流量继续打到已死的兄弟账户。

- 触发条件仅限 402 + detail.code == deactivated_workspace;通用 402 与 401 行为不变
- 调用点两处:fastpath 各类早退之前 + HandleUpstreamError 顶部,进程内 60s 按 Team 去重
- 逐账户 SetError(错误信息标注触发账户),单账户失败不中断其余
- 默认生效,无配置开关
2026-08-15 16:48:40 +08:00
shaw 4d9fedee20 fix(test): check type assertions in turn-state provenance tests
errcheck runs with check-type-assertions enabled, so the single-value
assertions on the provenance map values fail lint.
2026-08-15 16:44:30 +08:00
shaw fce41e318f fix(openai): make Codex fingerprint convergence opt-in and cover passthrough
Default codex_fingerprint_mode to off. v0.1.175 treated a missing key as
"session", so upgrading silently rewrote installation/session/thread/turn/
window identifiers for every existing OAuth account that had never configured
this field. The quota regressions in #5555, #5556 and #5582 line up with that
version boundary, with A/B reports that rolling back to v0.1.173 restores
quota. Convergence is now explicit opt-in (#5610).

Only accounts that never set the field change behaviour; explicit off /
device / session / full keep working exactly as configured. That required
flipping the persistence condition in all three account modals from
"!== 'session'" to "!== 'off'": the old rule deleted the key when it equalled
the default, which after the flip would have silently discarded an
administrator's explicit opt-in to session.

Also extend convergence to the passthrough path, which previously left client
identifiers untouched:

- resolve the ids once in forwardOpenAIPassthrough and rewrite
  client_metadata on the raw bytes (gjson extract + sjson splice) because
  passthrough is a hot path that must not fully unmarshal multi-MB bodies;
  a shared core keeps the raw and map variants from drifting
- both request builders apply the staged ids at the same relative position
  (after session isolation, before identity enforcement) so headers and body
  share one id set and turn_id stays consistent
- stage the ids unconditionally, including nil: a failover from a converged
  account to an off account must not leave the previous account's ids behind
2026-08-15 16:35:26 +08:00
shaw 8ae6d8f67e fix(openai): send session-level beta features and probe native compaction v2
OpenAI sunset the legacy unary /responses/compact endpoint (404, #5598,
#5624), so the account "compact probe" in the admin UI kept failing even for
healthy accounts, and the beta-feature negotiation header was only attached
to compaction turns.

Beta features (codex-rs session/mod.rs build_model_client_beta_features_header
+ client.rs build_responses_headers): the header is a session-level constant
attached to every /responses request, the WS handshake and /responses/compact.
Enumerating FEATURES shows no Experimental feature is enabled by default, so a
default install sends exactly "remote_compaction_v2". Mirror that:

- OAuth requests without a client-declared header get the default shape, so we
  no longer produce a "header only on compaction turns" pattern real Codex
  never emits (#5586 chains that strip the header)
- a client-declared header is preserved as-is: non-empty without v2 means the
  user disabled the feature and the gateway must not rewrite that
- native v2 turns (compaction_trigger in body) always ensure v2 is present
- non-OAuth upstreams keep the compaction-turn-only behaviour
- the WS injection sits outside the client-header copy block so prewarm and
  turn handshakes cannot land in different pool compatibility buckets

Compact probe now exercises native v2 (streaming /responses +
compaction_trigger) instead of the dead endpoint. Success requires an actual
compaction output item — scanning output_item.done/added, the terminal
response.output[] and the whole-JSON fallback — so a 2xx that silently drops
the trigger is reported as unsupported (the "got 0 items" class, #5478,
#5648). Probe identity is now UUID-shaped and applies the account's
convergence, matching real traffic on the same endpoint.
2026-08-15 16:35:08 +08:00