Commit Graph
5314 Commits
Author SHA1 Message Date
Wesley Liddick ad34f89152 Merge pull request #4933 from visa2/fix/usage-model-mapping-statistics
fix(usage): report channel-mapped requests under their real upstream model
2026-07-27 11:45:58 +08:00
Wesley Liddick de6b189a6b Merge pull request #4879 from wey-gu/fix/security-deps-20260726
fix(deps): update image and telemetry packages
2026-07-27 11:45:06 +08:00
Wesley Liddick a74e11c26a Merge pull request #4868 from visa2/fix/settings-partial-update-clobber
fix(settings): keep fields a settings PUT never sent at their stored value
2026-07-27 11:44:40 +08:00
Wesley Liddick 131d42d25d Merge pull request #4839 from visa2/fix/composite-route-prefix-passthrough
fix(composite): pass the requested model through when a prefix route leaves upstream_model empty
2026-07-27 11:44:15 +08:00
Wesley Liddick 031c83b7e0 Merge pull request #4875 from StarryKira/codex/fix-4859-gemini-36-flash-billing
fix(billing): price Antigravity Gemini 3.6 Flash
2026-07-27 11:43:49 +08:00
Wesley Liddick 7a3fda57c8 Merge pull request #4820 from feeeei/main
fix(gemini): 完善gemini号池模式时retryable失效问题
2026-07-27 11:43:13 +08:00
Wesley Liddick 16365199aa Merge pull request #4884 from Brisbanehuang/fix/probe-scheduling-nanosecond-timestamps
fix(repository): 修复上游计费倍率探测因纳秒时间戳解析失败导致的调度饿死
2026-07-27 11:42:59 +08:00
Wesley Liddick 91a2281c7a Merge pull request #4861 from coo1white/fix-flaky-concurrency-tests
test: stop four concurrency tests from failing on a busy machine
2026-07-27 11:42:34 +08:00
Wesley Liddick ece9517091 Merge pull request #4930 from wucm667/fix/issue-4928-config-file-path
fix(config): honor explicit CONFIG_FILE path
2026-07-27 11:42:08 +08:00
Wesley Liddick 4cc88e27b6 Merge pull request #4873 from wey-gu/fix/admin-usage-request-id-filter
fix(admin): filter usage logs by request id
2026-07-27 11:41:09 +08:00
Wesley Liddick b468e428e9 Merge pull request #4926 from Vibeone/fix/oauth-mimicry-cache-prefix-break
fix(gateway): 识别被代理的 Claude Code 流量,避免 mimicry 重写破坏 prompt cache
2026-07-27 11:40:43 +08:00
Wesley Liddick a40d6de12e Merge pull request #4907 from feitianbubu/fix/bump-claude-cli-version-2.1.220
fix(claude): 伪装的 Claude Code CLI 版本号升级到 2.1.220
2026-07-27 11:40:18 +08:00
Wesley Liddick bc9173be15 Merge pull request #4934 from OG-Wang/fix/monitor-timeline-overflow
fix(frontend): 修复渠道监控时间线在窄卡片下溢出
2026-07-27 11:39:34 +08:00
Wesley Liddick eb6e3d1f1d Merge pull request #4787 from KtzeAbyss/fix/4760-ws-turn-model-billing
fix(openai): track WebSocket models per turn
2026-07-27 10:25:57 +08:00
Wesley Liddick a93bfb6623 Merge pull request #4757 from lucas-ward/codex/fix-4691-caddy-sse-buffering
fix(deploy): prevent Caddy compression from buffering SSE
2026-07-27 10:22:58 +08:00
Wesley Liddick 8f47bd5fa0 Merge pull request #4893 from wucm667/fix/issue-4887-prompt-audit-config-load
fix(security-audit): reject unavailable prompt config
2026-07-27 10:20:14 +08:00
Wesley Liddick beeb4b84ed Merge pull request #4900 from wucm667/fix/issue-4889-mobile-available-channels
fix(frontend): adapt available channels for mobile
2026-07-27 10:19:44 +08:00
Wesley Liddick aac44473aa Merge pull request #4876 from wucm667/fix/issue-4846-show-usage-user
fix: show routed user in usage filters
2026-07-27 10:19:31 +08:00
Wesley Liddick 465362e1af Merge pull request #4877 from wucm667/fix/issue-4863-turnstile-invite-overlap
fix: show optional affiliate code on registration
2026-07-27 10:19:16 +08:00
Wesley Liddick 6ee2304dcd Merge pull request #4912 from yan9651688/fix/issue-4794-grok-test-402
fix(grok): pause accounts after manual test payment failure
2026-07-27 10:19:01 +08:00
Rick e94383a4c4 fix(frontend): 修复渠道监控时间线在窄卡片下溢出
MonitorTimeline 每根柱子设置了 min-w-[3px],60 根柱子加 2px 间距的
最小总宽度为 298px。当卡片内容区宽度低于该值时(如 100% 缩放下的
部分布局),时间线整体溢出卡片边缘。改为 min-w-0 让柱子随容器等分
压缩,任意宽度下均不再溢出。
2026-07-27 09:08:56 +08:00
shaw 7d3a896fcd chore: update sponsors 2026-07-27 08:59:12 +08:00
wucm667 5c471485ab fix(config): honor explicit CONFIG_FILE path
Make CONFIG_FILE select an explicit config for both full loading and lightweight address lookup, with regression tests.
2026-07-27 06:20:11 +08:00
eyre 7b3ed2a961 fix(gateway): detect proxied Claude Code traffic by body to preserve prompt cache
When an upstream API gateway (e.g. new-api) relays real Claude Code
requests, the User-Agent becomes Go-http-client while the body retains
the full Claude Code fingerprint (billing attribution block +
metadata.user_id + cache_control breakpoints).

Previously, the OAuth mimicry path relied solely on UA matching to
detect Claude Code clients. Without a matching UA, the gateway would
rewrite the system prompt — replacing the client's carefully structured
system blocks and cache_control breakpoints with its own injection.
This breaks Anthropic's prefix-based prompt cache: since the cache key
evaluates tools → system → messages in order, a changed system
invalidates all downstream message caching.

Symptoms observed:
- cache_read permanently locked at ~25K (only system prompt cached)
- cache_creation growing monotonically every turn (full messages rewrite)
- Single-request costs $17-27 instead of normal $1-2

Fix: when UA does not match but the body contains a valid billing
attribution block (x-anthropic-billing-header with cc_entrypoint=),
treat the request as proxied Claude Code traffic and skip mimicry.
This preserves the client's original system structure and cache_control
breakpoints, allowing Anthropic's prompt cache to function correctly.
2026-07-26 17:54:56 +00:00
visa2 be65c713ff fix(usage): preserve final upstream model 2026-07-27 00:43:53 +08:00
visa2 1f45c99de7 fix(usage): correct mapped model statistics 2026-07-26 23:56:34 +08:00
yan9651688 2db0cbd292 fix(grok): pause accounts after manual test payment failure
Manual Grok connection tests previously surfaced upstream HTTP 402 errors without changing account availability. Persist the same 30-minute payment-required cooldown used by the live forwarding path so refreshed account lists no longer present the account as schedulable.

Constraint: Keep manual-test HTTP 402 handling aligned with existing Grok forwarding semantics.
Rejected: Mark the account permanently error | payment state can recover and the forwarding path intentionally uses a bounded cooldown.
Confidence: high
Scope-risk: narrow
Directive: Keep the manual-test cooldown reason and duration aligned with handleGrokAccountUpstreamError.
Tested: go test -tags=unit ./internal/service -count=1; go vet -tags=unit ./internal/service; production package compile check
Not-tested: Live xAI account with an exhausted subscription
Related: #4794
2026-07-26 21:04:18 +08:00
feitianbubu 7af28ca843 fix(claude): 伪装的 Claude Code CLI 版本号升级到 2.1.220 2026-07-26 19:20:11 +08:00
wucm667 16dd3d8ee6 fix(frontend): adapt available channels for mobile 2026-07-26 16:26:25 +08:00
wucm667 56b5f0df68 fix(security-audit): reject unavailable prompt config 2026-07-26 14:33:59 +08:00
KtzeAbyss 7ce6e8d652 fix(openai): track websocket models per turn 2026-07-26 13:43:06 +08:00
Brisbanehuang 2447c44f85 fix(repository): parse nanosecond next_probe_at in due probe scheduling
Go persists upstream_billing_probe.next_probe_at via RFC3339Nano, but
jsonpath datetime() parses at most 6 fractional digits, so every stored
timestamp failed to parse and was treated as malformed: fail-open due,
ordered into the invalid bucket by id ASC. With more enabled accounts
than the per-cycle limit, the same lowest IDs monopolized every cycle
and higher IDs were never probed again. Trim the fraction to
microseconds before datetime(), mirroring ListDueOllamaCloudUsageAccounts,
and pin the behavior with integration regressions: nanosecond parsing,
due-time ordering beyond the limit, preserved fail-open for truly
invalid dates.
2026-07-25 23:04:21 -04:00
feeeei fd7e2039d3 fix(gemini): 完善gemini号池模式时retryable失效问题
此前 Gemini 三条转发路径对池模式账号命中 ErrorPolicySkipped 时直接把
上游错误体透传给客户端(强制标 500),既不同账号重试也不换号,
pool_mode_retry_count 配置完全不生效;而 Anthropic/OpenAI 等路径
会构造 UpstreamFailoverError 交给 handler 层按池模式配置重试后换号。

- 新增 poolModeSkippedFailoverError:池模式 + 可 failover 状态码时
  返回 UpstreamFailoverError,RetryableOnSameAccount 按
  pool_mode_retry_status_codes(默认 401/403/429)判定
- 原生 v1beta(ForwardNative)与 Claude messages 兼容路径的
  Skipped 分支接入;非池模式的自定义错误码透传行为不变
- chat completions 兼容路径的 failover 错误补上
  RetryableOnSameAccount 池模式判定
2026-07-26 09:52:47 +08:00
Wey Gu ba5fa6a38f fix(deps): update image and telemetry packages 2026-07-26 04:23:02 +08:00
wucm667 0875143d98 fix: show optional affiliate code on registration 2026-07-26 02:47:35 +08:00
wucm667 d11b838702 fix: show routed user in usage filters 2026-07-26 02:25:54 +08:00
haruka 6c7625800a fix(billing): price Antigravity Gemini 3.6 Flash 2026-07-26 01:56:10 +08:00
Wey Gu 1850e00955 fix(admin): filter usage logs by request id 2026-07-26 00:53:13 +08:00
Nick 291a737422 test: stop four concurrency tests from failing on a busy machine
All four pass on a quiet box and fail on a loaded one, each for its own
reason. None of them is testing the clock, so none of them should be
failing on it.

- ollama_cloud_usage_test.go: the second caller was released with
  `close(release)` right after its goroutine was started, not after it
  had reached the singleflight group. When the first refresh won that
  race, the second became a new singleflight execution, re-read the
  account, saw the LastAttemptAt the first one had just written, and
  came back with the 30-second manual-refresh 429 at line 685. It now
  counts account loads and waits for the second caller's own load,
  which happens right before it joins the group. Adds a counting
  GetByID to the test repo.

- gateway_hotpath_optimization_test.go: a 20ms sleep was meant to let
  all 12 callers reach the cache before the loader was released. A
  caller that arrived after the load had finished got a hit, not a
  miss, so the miss count came out 11 of 12. It now waits on the miss
  counter itself, which is the value the test asserts on.

- token_refresh_pool_health_test.go: the floor was `configuredSpacing`
  minus 10ms, i.e. 40ms out of 50ms. Each start timestamp is taken
  after the rate gate releases the goroutine, so scheduler delay can
  compress one observed gap with the gate behaving correctly — seen at
  37ms and again at 13ms. The floor is now a tenth of the configured
  spacing. Measured with providerQPS=20, 8 attempts, concurrency 2:
  gate at 50ms gives a minimum gap of 49.97ms, gate at 0 gives 22µs.
  So an unpaced gate sits three orders of magnitude under the 5ms floor
  and is still caught, while jitter has room to move. A comment warns
  against replacing this with an assertion on the total span of the
  starts: the span is set by how long each attempt takes under the
  concurrency limit, not by the gate — 471ms paced against 241ms
  unpaced — so a span check passes with the gate disabled.

- prompt_guard_test.go: the bound only has to show the failover shared
  the first endpoint's 70ms deadline and did not take the second
  endpoint's own 500ms one. An unshared deadline lands near 535ms, so
  350ms still fails loudly (seen: 224ms against a 180ms bound).

Tests only; no product code is touched. Each fix was checked in both
directions: it passes with the behaviour intact, and it still fails when
the behaviour is broken on purpose (for the QPS one, by swapping the
shared rate gate for a zero-interval one).
2026-07-25 21:59:20 +07:00
visa2 1614ae9c99 Merge remote-tracking branch 'origin/main' into fix/composite-route-prefix-passthrough 2026-07-25 22:20:27 +08:00
github-actions[bot] 2730c1c43b chore: sync VERSION to 0.1.165 [skip ci] 2026-07-25 13:59:09 +00:00
visa2andClaude Opus 5 0b5903d458 fix(settings): keep fields a settings PUT never sent at their stored value
PUT /api/v1/admin/settings is a whole-document write. The admin UI always sends
the complete document, so saving from the settings page is unaffected both
before and after this change. The bug is only reachable when an API client calls
the endpoint directly and sends just the fields it wants to change, which is the
natural assumption for a PUT on a settings resource.

Value-typed fields of UpdateSettingsRequest bind to their zero value when the
payload omits them, and buildSystemSettingsUpdates writes every key
unconditionally, so such a caller has no way to say "leave this one alone".
Omitting a field and explicitly clearing it are indistinguishable on the wire.
A caller that sends only the field it wants to change, e.g.

    {"risk_control_enabled": true}

sets that flag and clears every other unguarded field in the same request.
Measured against a fully configured store, one such call empties site_name,
site_subtitle, api_base_url, contact_info and doc_url, and turns
registration_enabled, email_verify_enabled, invitation_code_enabled and
turnstile_enabled off. turnstile_enabled alone gates the captcha on login,
register, forgot-password and both verify-code endpoints, and
email_verify_enabled is a precondition of IsPasswordResetEnabled.

The damage is easy to miss. site_name has a built-in fallback, so
getStringOrDefault renders the cleared value as the default product name and the
login page visibly changes, while the toggles just go quiet. Reopening the
settings page reads the already-cleared state back into the form, so correcting
the one visible field and saving persists the rest of the damage.

Fields that grew their own guard already survive this: the SMTP block falls back
to the previous values when smtp_host arrives empty, secret fields are written
only when non-empty, and 132 request fields are pointers whose handler merges an
omitted field with the stored value. This generalizes that pattern rather than
adding a fourth ad-hoc guard.

The handler now decodes the payload a second time as a raw field map, resolves
the setting key each absent field would have written, and hands that set to the
service, which drops those keys before SetMultiple, so the stored value is never
touched. Fields the payload does carry are written as before, giving the caller
the partial-update semantics it was already assuming. The mapping is reflected off
the request's json tags so new fields are covered without maintaining a list;
smtp_from_email is the only field whose json name differs from its setting key
and is aliased explicitly.

Only value-typed fields are filtered. Pointer fields keep whole-document
behaviour on purpose: forwarded_client_ip_headers and
api_key_acl_trust_forwarded_ip depend on being rewritten on every save to
re-normalize fail-closed state, which the malformed forwarded-client-IP header
test pins down.

An explicitly sent empty value is still a deliberate clear; only absent fields
are preserved. A partial write refreshes the in-process caches from storage
instead of from the request struct, which holds zero values for whatever the
caller omitted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 21:03:00 +08:00
Wesley Liddick e9a58c1cb8 Merge pull request #4814 from yiancode/fix/email-alias-registration-dedup
fix(auth): 注册查重归一化邮箱别名,防止单收件箱批量注册
v0.1.165
2026-07-25 20:54:37 +08:00
shaw ef0ca5bdf5 style: 修正 dotStrippedEmailExpr 注释以满足 gofmt 文档注释规则 2026-07-25 19:51:19 +08:00
shaw bc3acd6e28 fix(auth): 收紧注册别名查重(根点绕过 / 误拒 / 无界扫描 / 并发竞态)
对 #4814 的审计跟进修复:

- 域名尾随点绕过:user@gmail.com. 的域名不在 gmail 家族名单内,点号折叠与
  googlemail 归一被整体跳过,别名刷号原样可复现。归一化入口统一去掉 FQDN 根点。
- 误拒合法用户:剥 "+后缀" 缺空串守卫,+alice@ 与 +bob@ 都折叠成 @domain,
  该域后续 "+x@" 注册会永久 EMAIL_EXISTS 且无自助恢复。改为仅当 "+" 不在首位时剥离。
- 无界不可索引全表扫描:原实现按 LOWER(email) LIKE '%@domain' 把整域邮箱读进内存,
  且挂在公开未鉴权的 send-verify-code 上。改为按去点邮箱
  REPLACE(LOWER(TRIM(email)), '.', '') 做等值 + "local+%@domain" 前缀探针并带 LIMIT,
  新增同表达式的部分索引(migrations/190)。TRIM 口径与既有精确匹配一致,
  历史带首尾空白的行同样命中;LIKE 元字符转义,% 与 _ 不会扩大匹配面。
- 并发竞态:注册改走 CreateWithEmailAliasGuard,在邮箱唯一性锁上追加收件箱身份锁并在
  锁内复查,避免同一收件箱的多个别名变体同时通过服务层前置查重。管理员建号仍走
  Create,不受别名限制。
- 能力断言静默 fail-open:别名查重方法上提到 UserRepository 端口(编译期强制),
  移除可选接口类型断言与静默降级分支。
- OAuth 邮箱注册的两条建号路径(同样发放注册赠额)纳入同一查重口径;邮箱换绑/绑定
  不纳入,否则用户把邮箱改成自己收件箱的别名会被误拒。
2026-07-25 19:40:53 +08:00
yianandClaude Fable 5 b6f9277515 fix(auth): normalize email aliases in registration dedup
Registration duplicate checks compared the email verbatim (lowercase +
trim only), so a single inbox could spawn unlimited accounts via provider
alias features: plus addressing (user+tag@gmail.com) and the Gmail dot
trick (u.s.e.r@gmail.com) both deliver to the same mailbox but were seen
as distinct, verifiable addresses. This lets abusers bulk-register to farm
signup grants while the domain whitelist and email verification pass.

Add NormalizeEmailForAliasDedup to collapse these variants to a single
"inbox identity" and use it (via existsByEmailOrAlias) on the three local
email-registration paths: register, send-verify-code, and its async
variant. The exact ExistsByEmail check runs first; only on a miss do we
scan the candidate domains (gmail.com/googlemail.com are mutual aliases)
and compare normalized forms. Lookup errors fail closed, matching the
existing check, so the path cannot be bypassed by inducing errors.

Scoped to registration only — email storage, display, login, and delivery
are unchanged. OAuth-bound emails (already provider-verified) are out of
scope.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-25 19:19:15 +08:00
Wesley Liddick 2e2638c01d Merge pull request #4864 from Wei-Shaw/fix/ollama-usage-pg-compat-and-fetch-floor
fix(ollama): 修复 PG<=16 上 due 判定失效并恢复抓取下限
2026-07-25 18:58:06 +08:00
shaw 1763db3a2d fix(ollama): 修复 PG<=16 上 due 判定失效并恢复抓取下限
#4850 的用量刷新调度有三个问题,本次一并修复。

jsonpath .datetime() 直到 PostgreSQL 17 才接受 ISO-8601 的 Z 标识符,
而服务写入的快照时间戳全部是 UTC(即 Z 形式)。在 PG 14/15/16 上
ollamaCloudUsageParseRFC3339SQL 因此把 fetched_at / last_attempt_at /
next_refresh_at 全部解析成 NULL,due 判定整条退化为 fail-open 分支:
每轮 20 个刷新额度被 id 最小的非 due 组占满,id 更大的组永远不会刷新——
正是该 PR 声称修复的饥饿场景。项目文档声明支持 PG 14/15/16,而集成测试
harness 固定使用 postgres:18.1,因此现有测试无法发现。

修法是在送入 jsonpath 前把结尾的 Z 改写为 +00:00。仍保留 jsonpath 而不
直接 ::timestamptz,因为通过形状正则但日历非法的值(如 2026-02-30)需要
fail-open 成 NULL 而非中断整条查询。已在真实 PG 14/15/16/17/18 上验证
五个版本行为一致,且非法日历与垃圾输入仍正确 fail-open。

成功路径不再查阅 next_refresh_at,而 nextOllamaCloudUsageDelay 的
15 分钟下限正作用于该字段,导致同组对 ollama.com 的抓取下限从 15 分钟
降到一个 runner 周期。请求间隔略大于 debounce 的交互式流量(典型 Claude
Code 用法)会把单组 24 小时抓取次数从 24-96 抬高到数百次。改为对成功路径
显式施加 fetched_at + OllamaCloudUsageMinFetchInterval 的下限,Go 与 SQL
两侧同步;空闲账号不再轮询这一主要收益不受影响。

debounce_minutes 与 interval_minutes 此前各自独立校验,因此
debounce >= interval 是合法组合;此时 min(lastUsed+debounce,
fetchedAt+maxWait) 中的 debounce 项恒为死项,管理员配置被静默忽略。
改为在写入时拒绝该组合。

另修复 refreshAccount 中 ListOllamaCloudUsageGroupAccounts 的错误被
静默吞掉(违反 CLAUDE.md 禁止忽略错误):失败时回退到更窄的活动信号会
改变 due 语义,现在记录日志。

测试:
- 恢复被删除的 7/8/9 位小数秒解析覆盖,并改写为可判别形式(断言"不应
  返回",解析失败会落入 fail-open 而被捕获)。已验证该测试在 PG 15 上
  无此修复时失败、有修复时通过。
- 新增 min fetch interval 下限与 debounce/interval 交叉校验的单元测试。
- integration harness 新增 SUB2API_TEST_POSTGRES_IMAGE 覆盖,使套件可
  针对最低受支持版本运行。
2026-07-25 18:34:37 +08:00
Wesley Liddick bb0c38306f Merge pull request #4850 from alfadb/feat/ollama-usage-request-debounce
feat(ollama): 按模型请求刷新云端用量
2026-07-25 18:09:21 +08:00
alfadb b403f88f51 fix(ollama): 避免刷新候选饥饿
ListDue 在 LIMIT 前用与 service 纯函数一致的 debounce/max-wait/backoff
规则筛真正 due 组,防止有活动但未到期的组占满每轮 20 名额。
2026-07-25 15:37:06 +08:00