Commit Graph
3930 Commits
Author SHA1 Message Date
Wesley Liddick 031c83b7e0 Merge pull request #4875 from StarryKira/codex/fix-4859-gemini-36-flash-billing
fix(billing): price Antigravity Gemini 3.6 Flash
2026-07-27 11:43:49 +08:00
Wesley Liddick 7a3fda57c8 Merge pull request #4820 from feeeei/main
fix(gemini): 完善gemini号池模式时retryable失效问题
2026-07-27 11:43:13 +08:00
Wesley Liddick 16365199aa Merge pull request #4884 from Brisbanehuang/fix/probe-scheduling-nanosecond-timestamps
fix(repository): 修复上游计费倍率探测因纳秒时间戳解析失败导致的调度饿死
2026-07-27 11:42:59 +08:00
Wesley Liddick 91a2281c7a Merge pull request #4861 from coo1white/fix-flaky-concurrency-tests
test: stop four concurrency tests from failing on a busy machine
2026-07-27 11:42:34 +08:00
Wesley Liddick ece9517091 Merge pull request #4930 from wucm667/fix/issue-4928-config-file-path
fix(config): honor explicit CONFIG_FILE path
2026-07-27 11:42:08 +08:00
Wesley Liddick 4cc88e27b6 Merge pull request #4873 from wey-gu/fix/admin-usage-request-id-filter
fix(admin): filter usage logs by request id
2026-07-27 11:41:09 +08:00
Wesley Liddick b468e428e9 Merge pull request #4926 from Vibeone/fix/oauth-mimicry-cache-prefix-break
fix(gateway): 识别被代理的 Claude Code 流量,避免 mimicry 重写破坏 prompt cache
2026-07-27 11:40:43 +08:00
Wesley Liddick a40d6de12e Merge pull request #4907 from feitianbubu/fix/bump-claude-cli-version-2.1.220
fix(claude): 伪装的 Claude Code CLI 版本号升级到 2.1.220
2026-07-27 11:40:18 +08:00
Wesley Liddick eb6e3d1f1d Merge pull request #4787 from KtzeAbyss/fix/4760-ws-turn-model-billing
fix(openai): track WebSocket models per turn
2026-07-27 10:25:57 +08:00
Wesley Liddick 8f47bd5fa0 Merge pull request #4893 from wucm667/fix/issue-4887-prompt-audit-config-load
fix(security-audit): reject unavailable prompt config
2026-07-27 10:20:14 +08:00
wucm667 5c471485ab fix(config): honor explicit CONFIG_FILE path
Make CONFIG_FILE select an explicit config for both full loading and lightweight address lookup, with regression tests.
2026-07-27 06:20:11 +08:00
eyre 7b3ed2a961 fix(gateway): detect proxied Claude Code traffic by body to preserve prompt cache
When an upstream API gateway (e.g. new-api) relays real Claude Code
requests, the User-Agent becomes Go-http-client while the body retains
the full Claude Code fingerprint (billing attribution block +
metadata.user_id + cache_control breakpoints).

Previously, the OAuth mimicry path relied solely on UA matching to
detect Claude Code clients. Without a matching UA, the gateway would
rewrite the system prompt — replacing the client's carefully structured
system blocks and cache_control breakpoints with its own injection.
This breaks Anthropic's prefix-based prompt cache: since the cache key
evaluates tools → system → messages in order, a changed system
invalidates all downstream message caching.

Symptoms observed:
- cache_read permanently locked at ~25K (only system prompt cached)
- cache_creation growing monotonically every turn (full messages rewrite)
- Single-request costs $17-27 instead of normal $1-2

Fix: when UA does not match but the body contains a valid billing
attribution block (x-anthropic-billing-header with cc_entrypoint=),
treat the request as proxied Claude Code traffic and skip mimicry.
This preserves the client's original system structure and cache_control
breakpoints, allowing Anthropic's prompt cache to function correctly.
2026-07-26 17:54:56 +00:00
yan9651688 2db0cbd292 fix(grok): pause accounts after manual test payment failure
Manual Grok connection tests previously surfaced upstream HTTP 402 errors without changing account availability. Persist the same 30-minute payment-required cooldown used by the live forwarding path so refreshed account lists no longer present the account as schedulable.

Constraint: Keep manual-test HTTP 402 handling aligned with existing Grok forwarding semantics.
Rejected: Mark the account permanently error | payment state can recover and the forwarding path intentionally uses a bounded cooldown.
Confidence: high
Scope-risk: narrow
Directive: Keep the manual-test cooldown reason and duration aligned with handleGrokAccountUpstreamError.
Tested: go test -tags=unit ./internal/service -count=1; go vet -tags=unit ./internal/service; production package compile check
Not-tested: Live xAI account with an exhausted subscription
Related: #4794
2026-07-26 21:04:18 +08:00
feitianbubu 7af28ca843 fix(claude): 伪装的 Claude Code CLI 版本号升级到 2.1.220 2026-07-26 19:20:11 +08:00
wucm667 56b5f0df68 fix(security-audit): reject unavailable prompt config 2026-07-26 14:33:59 +08:00
KtzeAbyss 7ce6e8d652 fix(openai): track websocket models per turn 2026-07-26 13:43:06 +08:00
Brisbanehuang 2447c44f85 fix(repository): parse nanosecond next_probe_at in due probe scheduling
Go persists upstream_billing_probe.next_probe_at via RFC3339Nano, but
jsonpath datetime() parses at most 6 fractional digits, so every stored
timestamp failed to parse and was treated as malformed: fail-open due,
ordered into the invalid bucket by id ASC. With more enabled accounts
than the per-cycle limit, the same lowest IDs monopolized every cycle
and higher IDs were never probed again. Trim the fraction to
microseconds before datetime(), mirroring ListDueOllamaCloudUsageAccounts,
and pin the behavior with integration regressions: nanosecond parsing,
due-time ordering beyond the limit, preserved fail-open for truly
invalid dates.
2026-07-25 23:04:21 -04:00
feeeei fd7e2039d3 fix(gemini): 完善gemini号池模式时retryable失效问题
此前 Gemini 三条转发路径对池模式账号命中 ErrorPolicySkipped 时直接把
上游错误体透传给客户端(强制标 500),既不同账号重试也不换号,
pool_mode_retry_count 配置完全不生效;而 Anthropic/OpenAI 等路径
会构造 UpstreamFailoverError 交给 handler 层按池模式配置重试后换号。

- 新增 poolModeSkippedFailoverError:池模式 + 可 failover 状态码时
  返回 UpstreamFailoverError,RetryableOnSameAccount 按
  pool_mode_retry_status_codes(默认 401/403/429)判定
- 原生 v1beta(ForwardNative)与 Claude messages 兼容路径的
  Skipped 分支接入;非池模式的自定义错误码透传行为不变
- chat completions 兼容路径的 failover 错误补上
  RetryableOnSameAccount 池模式判定
2026-07-26 09:52:47 +08:00
haruka 6c7625800a fix(billing): price Antigravity Gemini 3.6 Flash 2026-07-26 01:56:10 +08:00
Wey Gu 1850e00955 fix(admin): filter usage logs by request id 2026-07-26 00:53:13 +08:00
Nick 291a737422 test: stop four concurrency tests from failing on a busy machine
All four pass on a quiet box and fail on a loaded one, each for its own
reason. None of them is testing the clock, so none of them should be
failing on it.

- ollama_cloud_usage_test.go: the second caller was released with
  `close(release)` right after its goroutine was started, not after it
  had reached the singleflight group. When the first refresh won that
  race, the second became a new singleflight execution, re-read the
  account, saw the LastAttemptAt the first one had just written, and
  came back with the 30-second manual-refresh 429 at line 685. It now
  counts account loads and waits for the second caller's own load,
  which happens right before it joins the group. Adds a counting
  GetByID to the test repo.

- gateway_hotpath_optimization_test.go: a 20ms sleep was meant to let
  all 12 callers reach the cache before the loader was released. A
  caller that arrived after the load had finished got a hit, not a
  miss, so the miss count came out 11 of 12. It now waits on the miss
  counter itself, which is the value the test asserts on.

- token_refresh_pool_health_test.go: the floor was `configuredSpacing`
  minus 10ms, i.e. 40ms out of 50ms. Each start timestamp is taken
  after the rate gate releases the goroutine, so scheduler delay can
  compress one observed gap with the gate behaving correctly — seen at
  37ms and again at 13ms. The floor is now a tenth of the configured
  spacing. Measured with providerQPS=20, 8 attempts, concurrency 2:
  gate at 50ms gives a minimum gap of 49.97ms, gate at 0 gives 22µs.
  So an unpaced gate sits three orders of magnitude under the 5ms floor
  and is still caught, while jitter has room to move. A comment warns
  against replacing this with an assertion on the total span of the
  starts: the span is set by how long each attempt takes under the
  concurrency limit, not by the gate — 471ms paced against 241ms
  unpaced — so a span check passes with the gate disabled.

- prompt_guard_test.go: the bound only has to show the failover shared
  the first endpoint's 70ms deadline and did not take the second
  endpoint's own 500ms one. An unshared deadline lands near 535ms, so
  350ms still fails loudly (seen: 224ms against a 180ms bound).

Tests only; no product code is touched. Each fix was checked in both
directions: it passes with the behaviour intact, and it still fails when
the behaviour is broken on purpose (for the QPS one, by swapping the
shared rate gate for a zero-interval one).
2026-07-25 21:59:20 +07:00
github-actions[bot] 2730c1c43b chore: sync VERSION to 0.1.165 [skip ci] 2026-07-25 13:59:09 +00:00
shaw ef0ca5bdf5 style: 修正 dotStrippedEmailExpr 注释以满足 gofmt 文档注释规则 2026-07-25 19:51:19 +08:00
shaw bc3acd6e28 fix(auth): 收紧注册别名查重(根点绕过 / 误拒 / 无界扫描 / 并发竞态)
对 #4814 的审计跟进修复:

- 域名尾随点绕过:user@gmail.com. 的域名不在 gmail 家族名单内,点号折叠与
  googlemail 归一被整体跳过,别名刷号原样可复现。归一化入口统一去掉 FQDN 根点。
- 误拒合法用户:剥 "+后缀" 缺空串守卫,+alice@ 与 +bob@ 都折叠成 @domain,
  该域后续 "+x@" 注册会永久 EMAIL_EXISTS 且无自助恢复。改为仅当 "+" 不在首位时剥离。
- 无界不可索引全表扫描:原实现按 LOWER(email) LIKE '%@domain' 把整域邮箱读进内存,
  且挂在公开未鉴权的 send-verify-code 上。改为按去点邮箱
  REPLACE(LOWER(TRIM(email)), '.', '') 做等值 + "local+%@domain" 前缀探针并带 LIMIT,
  新增同表达式的部分索引(migrations/190)。TRIM 口径与既有精确匹配一致,
  历史带首尾空白的行同样命中;LIKE 元字符转义,% 与 _ 不会扩大匹配面。
- 并发竞态:注册改走 CreateWithEmailAliasGuard,在邮箱唯一性锁上追加收件箱身份锁并在
  锁内复查,避免同一收件箱的多个别名变体同时通过服务层前置查重。管理员建号仍走
  Create,不受别名限制。
- 能力断言静默 fail-open:别名查重方法上提到 UserRepository 端口(编译期强制),
  移除可选接口类型断言与静默降级分支。
- OAuth 邮箱注册的两条建号路径(同样发放注册赠额)纳入同一查重口径;邮箱换绑/绑定
  不纳入,否则用户把邮箱改成自己收件箱的别名会被误拒。
2026-07-25 19:40:53 +08:00
yianandClaude Fable 5 b6f9277515 fix(auth): normalize email aliases in registration dedup
Registration duplicate checks compared the email verbatim (lowercase +
trim only), so a single inbox could spawn unlimited accounts via provider
alias features: plus addressing (user+tag@gmail.com) and the Gmail dot
trick (u.s.e.r@gmail.com) both deliver to the same mailbox but were seen
as distinct, verifiable addresses. This lets abusers bulk-register to farm
signup grants while the domain whitelist and email verification pass.

Add NormalizeEmailForAliasDedup to collapse these variants to a single
"inbox identity" and use it (via existsByEmailOrAlias) on the three local
email-registration paths: register, send-verify-code, and its async
variant. The exact ExistsByEmail check runs first; only on a miss do we
scan the candidate domains (gmail.com/googlemail.com are mutual aliases)
and compare normalized forms. Lookup errors fail closed, matching the
existing check, so the path cannot be bypassed by inducing errors.

Scoped to registration only — email storage, display, login, and delivery
are unchanged. OAuth-bound emails (already provider-verified) are out of
scope.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-25 19:19:15 +08:00
shaw 1763db3a2d fix(ollama): 修复 PG<=16 上 due 判定失效并恢复抓取下限
#4850 的用量刷新调度有三个问题,本次一并修复。

jsonpath .datetime() 直到 PostgreSQL 17 才接受 ISO-8601 的 Z 标识符,
而服务写入的快照时间戳全部是 UTC(即 Z 形式)。在 PG 14/15/16 上
ollamaCloudUsageParseRFC3339SQL 因此把 fetched_at / last_attempt_at /
next_refresh_at 全部解析成 NULL,due 判定整条退化为 fail-open 分支:
每轮 20 个刷新额度被 id 最小的非 due 组占满,id 更大的组永远不会刷新——
正是该 PR 声称修复的饥饿场景。项目文档声明支持 PG 14/15/16,而集成测试
harness 固定使用 postgres:18.1,因此现有测试无法发现。

修法是在送入 jsonpath 前把结尾的 Z 改写为 +00:00。仍保留 jsonpath 而不
直接 ::timestamptz,因为通过形状正则但日历非法的值(如 2026-02-30)需要
fail-open 成 NULL 而非中断整条查询。已在真实 PG 14/15/16/17/18 上验证
五个版本行为一致,且非法日历与垃圾输入仍正确 fail-open。

成功路径不再查阅 next_refresh_at,而 nextOllamaCloudUsageDelay 的
15 分钟下限正作用于该字段,导致同组对 ollama.com 的抓取下限从 15 分钟
降到一个 runner 周期。请求间隔略大于 debounce 的交互式流量(典型 Claude
Code 用法)会把单组 24 小时抓取次数从 24-96 抬高到数百次。改为对成功路径
显式施加 fetched_at + OllamaCloudUsageMinFetchInterval 的下限,Go 与 SQL
两侧同步;空闲账号不再轮询这一主要收益不受影响。

debounce_minutes 与 interval_minutes 此前各自独立校验,因此
debounce >= interval 是合法组合;此时 min(lastUsed+debounce,
fetchedAt+maxWait) 中的 debounce 项恒为死项,管理员配置被静默忽略。
改为在写入时拒绝该组合。

另修复 refreshAccount 中 ListOllamaCloudUsageGroupAccounts 的错误被
静默吞掉(违反 CLAUDE.md 禁止忽略错误):失败时回退到更窄的活动信号会
改变 due 语义,现在记录日志。

测试:
- 恢复被删除的 7/8/9 位小数秒解析覆盖,并改写为可判别形式(断言"不应
  返回",解析失败会落入 fail-open 而被捕获)。已验证该测试在 PG 15 上
  无此修复时失败、有修复时通过。
- 新增 min fetch interval 下限与 debounce/interval 交叉校验的单元测试。
- integration harness 新增 SUB2API_TEST_POSTGRES_IMAGE 覆盖,使套件可
  针对最低受支持版本运行。
2026-07-25 18:34:37 +08:00
alfadb b403f88f51 fix(ollama): 避免刷新候选饥饿
ListDue 在 LIMIT 前用与 service 纯函数一致的 debounce/max-wait/backoff
规则筛真正 due 组,防止有活动但未到期的组占满每轮 20 名额。
2026-07-25 15:37:06 +08:00
alfadb f1eaed31ed fix(ollama): 写入设置时缺省 debounce 补 1
兼容旧客户端省略 debounce_minutes 的 PUT 请求。
2026-07-25 15:37:06 +08:00
alfadb 0f72b7dca7 feat(ollama): 按模型请求刷新云端用量
将 Ollama Cloud settings HTML 自动抓取从固定间隔轮询改为请求驱动的
trailing debounce + max-wait:无新请求不再抓取,连续请求最晚在
max-wait 强制刷新;失败退避仍优先于活动 due。新增 debounce_minutes
(默认 1),interval_minutes 保留为最长等待兼容字段。
2026-07-25 15:37:05 +08:00
shaw 5374ce2a0f Merge remote-tracking branch 'origin/main' into feature/openai-live-gateway 2026-07-25 15:16:05 +08:00
shaw 7acc13a29b fix(live): 租约丢失终止会话并补齐过期时的 usage log
三处功能修复(审计发现):

1. 租约续租失败按会话终结处理。RefreshLiveLease 的 Lua 在 leaseID 被 GC 后不会
   重新 ZADD,重连拿不回并发槽;原先把 ErrLiveUnavailable 当临时错误交给 observer
   重连,会让会话以约 1 秒一轮的节奏空转到 ExpiresAt(最长 60 分钟),期间持着
   上游 WS 连接却不计入账号/用户/API Key 的任何并发限制。

2. observer 因过期放弃时补写 usage log。waitForLiveObserverRetry 原先把过期判定
   混在重试条件里并返回 false,直接 return 绕过了 observeLiveCall 循环顶部的过期
   分支,导致该路径既不写 usage log 也不释放租约(静默结束)。改为只判控制权归属,
   过期交回循环顶部 finalize。

3. 把两处重复的三项终止判据抽成 liveSessionEnded,消除 ProxyLiveSideband 与
   observeLiveCall 之间的判据漂移风险。

迁移改名 186/187 -> 188/189:原编号会成为第三个 186,且 187 与已合入 main 的
#4801 的 187_add_usage_log_session_id.sql 撞号。文件名即迁移主键,功能无害但易误读。

新增两个测试均经变异验证(去掉修复后会失败)。
2026-07-25 15:15:59 +08:00
Wesley Liddick e6a3030beb Merge pull request #4803 from wucm667/fix/issue-4798-gemini-image-chat-content
fix(gemini): preserve image output in chat completions
2026-07-25 14:43:30 +08:00
shaw f06c8c5698 Merge branch 'main' into fix/issue-4798-gemini-image-chat-content
解决 frontend/pnpm-lock.yaml 冲突:采用 main 的 frontend/package.json 与
frontend/pnpm-lock.yaml,回退本 PR 附带的 postcss ^8.4.32 -> ^8.5.12 改动。
main 已通过 pnpm.overrides ("postcss@<8.5.18": ">=8.5.18") 强制安全下限,
本 PR 的直接依赖下限低于该值,保留会与 overrides 冲突且无安全收益。
本 PR 的功能改动(gemini_* compat service 及其测试)未受影响。
2026-07-25 14:33:06 +08:00
shaw 8bed40c67b Merge branch 'main' into fix/issue-4819-drop-orphan-tool-choice
解决 frontend/pnpm-lock.yaml 冲突:采用 main 的 frontend/package.json 与
frontend/pnpm-lock.yaml,回退本 PR 附带的 postcss ^8.4.32 -> ^8.5.12 改动。
main 已通过 pnpm.overrides ("postcss@<8.5.18": ">=8.5.18") 强制安全下限,
本 PR 的直接依赖下限低于该值,保留会与 overrides 冲突且无安全收益。
本 PR 的功能改动(openai_gateway_grok.go 及其测试)未受影响。
2026-07-25 14:32:18 +08:00
Wesley Liddick a29a62112b Merge pull request #4842 from Cynicismcart/fix/openai-pool-retry-transient-cooldown
fix(openai): preserve pool-mode same-account retries
2026-07-25 14:30:04 +08:00
Wesley Liddick eb03be0dbf Merge pull request #4788 from wucm667/fix/issue-4782-pool-5xx-cooldown
fix(grok): avoid pool cooldown on 5xx
2026-07-25 14:29:36 +08:00
Wesley Liddick 6d956bdc20 Merge pull request #4801 from SemonCat/feat/persist-session-id
feat(usage): persist client session identifiers
2026-07-25 14:03:32 +08:00
Wesley Liddick 333acde7d1 Merge pull request #4832 from ListenCodes/fix/openai-apikey-responses-item-id-v2
fix(openai): sanitize API-key responses item IDs
2026-07-25 13:42:35 +08:00
Wesley Liddick 9994eaa70f Merge pull request #4804 from scp-planet/fix/openai-strip-input-namespace
fix(openai): strip input item namespaces before HTTP forwarding / HTTP 转发前移除 input 项 namespace
2026-07-25 13:42:17 +08:00
Wesley Liddick 0e39e21fa2 Merge pull request #4796 from 404QAQ/codex/fix-image-request-diagnostics
fix(images): log requested quality and size
2026-07-25 13:41:52 +08:00
Wesley Liddick 95cc6ee2c0 Merge pull request #4793 from 404QAQ/codex/fix-pricing-empty-remote-url
fix(pricing): skip remote scheduler without URL
2026-07-25 13:41:40 +08:00
song db6fbdbf29 fix(openai): satisfy Live CI checks 2026-07-25 12:51:20 +08:00
song 988d4b577e feat(openai): add macOS Live attestation 2026-07-25 12:50:46 +08:00
song ec23716ee4 fix(openai): align Live request metadata 2026-07-25 12:50:46 +08:00
song e6eb23eaac feat(openai): add Live gateway support 2026-07-25 12:50:46 +08:00
shaw 6c9b84cc7a feat: 适配 Anthropic 新模型 claude-opus-5
模型登记:/v1/models 清单、Bedrock 默认映射(us.anthropic.claude-opus-5-v1)、
定价条目($5/$25 per MTok、1M 上下文、128K 输出)、前端模型清单与
Anthropic/Bedrock 预设映射、限流 scope 简称。

同时修复两个会静默出错的问题:

- 定价家族兜底 3 倍超收:定价数据缺 claude-opus-5 时,matchByModelFamily
  的 Phase 2 关键字兜底会落到 opus-4 系列、getFallbackPricing 会落到
  claude-3-opus,两条路都按 $15/$75 计费(官方 $5/$25),输入输出双双
  3 倍超收且无任何报错。两处补 opus-5 家族并回退到同价的 4.8;判断用
  opus-5/opus5 子串而非裸 "5",避免误伤 claude-opus-4-5。顺带补齐兜底表
  缺失的 claude-opus-4.8(此前同样会掉到 claude-3-opus)。

- Bedrock 版本闸门降级:claudeVersionRe 强制要求 major-minor 两段版本号,
  只有主版本号的 claude-opus-5 / claude-sonnet-5 完全不匹配,被当成旧模型:
  isBedrockOpus47OrNewer 假导致 thinking.enabled 不转 adaptive(Opus 5 上游
  已移除 budget_tokens,透传直接 400)、isBedrockClaude45OrNewer 假导致
  cache_control.ttl 被剥离、bedrockModelSupportsToolSearch 假导致 tool search
  被过滤。改为 minor 可选(缺省 minor=0),claude-sonnet-5 的同一问题一并修复。

Vertex 无需改动:normalizeVertexAnthropicModelID 只处理 -YYYYMMDD→@YYYYMMDD,
无日期后缀的裸 ID 原样透传即正确。context-1m-2025-08-07 白名单不动:Opus 系
上游不接受该 beta,且 Opus 5 的 1M 上下文是默认能力。

Antigravity 暂不接入:无上游支持证据,mapAntigravityModel 对未映射模型返回
空字符串即"该账号不支持",fails closed 安全。

回归测试 internal/service/claude_opus5_test.go 覆盖定价两层兜底、Bedrock
三个闸门、thinking 转换与模型清单;逐个回退上述修复已确认测试会红。
2026-07-25 11:22:42 +08:00
Cynicismcart 521db6869e fix(openai): preserve pool-mode same-account retries
Avoid recording the generic account-model transient cooldown for pool-mode statuses explicitly configured for same-account retry. This lets the bounded request-local retry budget complete while preserving cooldowns for other statuses and non-pool accounts.
2026-07-25 05:31:31 +08:00
LIULIXING 1891faa68d fix(openai): keep item ID sanitization linear 2026-07-25 01:40:26 +08:00
ListenCodes c5d9d57940 fix(openai): sanitize API-key responses item IDs 2026-07-24 20:15:58 +08:00
wucm667 41f12e874d fix(grok): drop orphaned Responses tool choice 2026-07-24 16:25:53 +08:00