Commit Graph
3913 Commits
Author SHA1 Message Date
Rain ba92d70422 fix grok usage guard for compatible accounts 2026-08-05 17:59:20 +08:00
Rain 5c52fa93d5 fix grok missing usage billing bypass 2026-08-05 09:49:25 +08:00
feeeei 26e0a89323 人机验证增加阿里云验证码 2.0
沿用腾讯天御验证码引入的多服务商模型:aliyun_captcha_enabled 作为独立
开关,与 Cloudflare Turnstile、腾讯天御三方互斥(保存校验 + 运行时
CAPTCHA_PROVIDER_CONFLICT)。后台「安全与认证」合并为单张人机验证卡片:
总开关 + 服务商单选(Turnstile / 腾讯天御 / 阿里云),选中即启用该家并
关闭其它,落库仍是三个独立开关键,由前端映射保证互斥。

阿里云侧同时支持 aliyun 中国站与国际站(alibabacloud.com):两站前端脚本、
region 取值与服务端 API 完全一致,仅账号与 AccessKey 相互独立,因此由
「服务地域」决定线路即可——中国内地走 captcha.cn-shanghai.aliyuncs.com,
非中国内地(新加坡)走 captcha.ap-southeast-1.aliyuncs.com,AccessKey
取自持有该实例的账号,无需在配置中区分站点。

- AliyunCaptchaService 对称 TencentCaptchaService:服务端校验走官方 SDK
  VerifyIntelligentCaptcha,调用异常按 fail-closed 拦截,与 Turnstile
  网络错误行为对称;保存设置时真实探测 AK/SK 有效性
- 保护面对齐腾讯扩展入口:VerifyTencentCaptchaIfEnabled 通用化为
  VerifyActionCaptchaIfEnabled,OAuth 登录启动、passkey 登录在阿里云
  启用时同样拦截;Turnstile 维持既有覆盖不扩大
- 前端 AliyunCaptchaWidget 为表单内预验证按钮(popup 模式),同时暴露
  verify() 供 OAuth 启动、passkey 等动作入口程序化弹窗;未预验证直接
  提交时弹窗兜底。SDK 按钮绑定异步完成,弹窗未出现前按 tick 重试触发,
  并轮询弹窗可见性识别用户关闭
- captchaVerifyParam 复用 turnstile_token 请求字段提交;公开设置下发
  aliyun_captcha_enabled / scene_id / prefix / region
- CSP 放行验证码 CDN:script-src/style-src 加 *.alicdn.com
2026-08-04 20:57:15 +08:00
Wesley Liddick 8b3fe664dc Merge pull request #5261 from lyen1688/feat/tencent-captcha-gate
新增腾讯天御验证码认证门禁
2026-08-04 16:39:55 +08:00
Wesley Liddick a4d263f62f Merge pull request #5224 from wucm667/fix/issue-5190-lock-subscription-renewal
fix(subscription): serialize concurrent renewals
2026-08-04 16:28:52 +08:00
Wesley Liddick ae81dfd933 Merge pull request #5177 from r266-tech/fix-nonpassthrough-write-context
fix(openai-ws): preserve terminal event on lease loss
2026-08-04 16:28:22 +08:00
Wesley Liddick 35cab3c814 Merge pull request #5258 from feeeei/fix/model_plaza
fix(model-plaza): Model Plaza image model price display is inconsistent with the actual price
2026-08-04 16:27:53 +08:00
Wesley Liddick 770e35b474 Merge pull request #5040 from coo1white/upstream-pr/grok-cli-0.2.114
fix(grok): bump pinned Grok CLI version to 0.2.114
2026-08-04 16:16:20 +08:00
Wesley Liddick 846dd310a3 Merge pull request #5158 from neboyang/feat/claude-oauth-authorize-url
fix(oauth): use claude.com/cai authorize endpoint
2026-08-04 16:15:50 +08:00
Wesley Liddick 9b4575e434 Merge pull request #5243 from wucm667/feat/issue-5240-dashboard-username
fix(dashboard): show usernames in spending ranking
2026-08-04 16:15:21 +08:00
Wesley Liddick 1f4cfc44c1 Merge pull request #5226 from spongehah/codex/fix-prompt-audit-output-text
fix(prompt-audit): parse responses output text
2026-08-04 16:15:04 +08:00
lyen1688 e592c5f9e0 新增腾讯天御验证码认证门禁 2026-08-04 15:09:29 +08:00
feeeei 785b61d424 修复模型广场图片模型价格展示与实收口径不一致
图片计费模型的广场展示价按实收口径计算:档位单价取
分组图片价 > 渠道档位价 > 渠道默认按次价;分组开启生图
独立倍率时实付倍率取独立倍率,不取分组/专属倍率。
2026-08-04 13:48:19 +08:00
shaw 2d3e845205 perf(codex): 版本同步主路径改用 /releases/latest 并保留列表回退
自动同步原本每次拉 `/releases?per_page=30`,实测响应 10,017,686 字节——该仓库
预发布极密集,30 条里只有 2 条稳定版(0.145.0 已排到第 26 位),页大小是为了
「窗口里至少有一条稳定版」而定的,不能靠调小来省流量。

改为主路径走 `/releases/latest`(实测约 0.3MB):该端点本身排除 draft 与
prerelease,直接给出最新正式发布,不受预发布密度影响。复用端口上已有的
FetchLatestRelease,不新增任何 HTTP 代码,代理配置、User-Agent、可选 token
与跨主机重定向剥离 Authorization 的行为全部沿用。

保留列表扫描作为回退:latest 是跨 tag 家族按 published_at 取的,若同仓库其他
组件(如 rusty-v8-*)某天发布正式 release 而成为 latest,主路径会被 rust-v
前缀过滤挡掉,此时必须扫一页才能继续跟随,否则版本号会静默停更。

两条路径共用 latestCodexStableReleaseVersion 的同一套过滤(前缀 / draft /
prerelease / 版本号形态),语义不会分叉;单条 latest 也要过版本号校验——仓库里
确实存在 rust-vv0.99.0-alpha.8 这类畸形 tag。只向前推进、抓取失败保持既有值
两条不变式未改动。

测试覆盖:主路径命中时不再拉列表页、四种拿不到(异家族 / 预发布 / 抓取失败 /
空对象)都回退、畸形 tag 被主路径拒绝后由回退兜底、两路径皆失败时保值。
2026-08-04 10:51:04 +08:00
wucm667 80af44cad7 [verified] fix(dashboard): show usernames in spending ranking 2026-08-04 00:43:05 +08:00
shaw c899c8cf37 fix(codex): 管理员配置的 UA 只贡献指纹,版本段一律用生效版本重建
面板「OpenAI Codex UA」此前只在客户端 UA 是浏览器型时才生效,本分支把它提升为
所有 OAuth 出站的规范身份来源,作用域扩大了一个数量级。而它的旧 placeholder 原文
就是 codex_cli_rs/0.144.1 (...):照抄填写过的存量部署会被永久钉在 0.144.1 上——
该值不低于上游门槛 0.144.0,version 头也随之变成 0.144.1,绕过版本自动同步,
稳定落在上游优先降载的那一侧,且面板、日志、审计都不提示。账号级自定义 UA
(credentials.user_agent)有完全相同的陷阱。

改为:管理员配置的 UA 只贡献客户端名与 OS / 架构 / 终端指纹(这是该输入框唯一
不可替代的价值),版本段一律用当前生效版本重建。存量的陈旧配置无需迁移即自愈,
UA 与 version 头从此由构造保证同源,原先「覆写 UA 陈旧时二者不一致」的次优解消失。

- 新增 openai.SetCodexUserAgentVersion:重建首段版本声明,OS / 架构 / 终端指纹
  原样保留;尾部官方客户端标识组 (name; version) 与首段同源,一并更新,避免拼出
  首段声明新版本、尾部仍是旧版本的自相矛盾身份;非官方括号组(如 OS 组)不动。
- resolveCodexOutboundIdentity 收敛为单一funnel:生效版本只从规范身份取,
  候选 UA 的版本段不再参与判定。
- GetOpenAICodexCanonicalUserAgent 重建面板 UA 的版本段;原「面板值等于兜底常量
  视同未填」的特例随之成为恒等变换,删除。
- 补齐 4 处静默丢弃账号级自定义 UA 的出站路径(WS 握手、账号测试 x2、用量探针):
  它们先写 customUA 再调不带 override 的收口,赋值随即被覆盖,既是行为不一致
  也是死代码;账号测试尤其要紧——注释写着「与真实转发一致」而实际不一致。
- 补测试:SetCodexUserAgentVersion / CodexUserAgentVersion 包内用例、陈旧面板 UA
  与账号 UA 的版本重建回归、WS 握手尊重账号级 UA。
2026-08-03 21:50:49 +08:00
zhiyu 4c4ff36380 perf(codex): 版本同步间隔改为 6 小时并加启动防抖
- 同步间隔 3h → 6h:客户端版本是天级变化,6 小时足够跟上,
  对 GitHub 的调用降到每天 4 次。
- 新增启动防抖:借同步设置行自身的 UpdatedAt 判断,若同步值仍在一个周期内
  则跳过启动同步。频繁重启、滚动发布或崩溃重启原本会把「启动即同步」
  放大成对 GitHub 的连续请求;首次部署尚无同步值时不受影响。
- 补测试:启动防抖两个方向,以及版本比较按段取数字的回归
  (字典序会把 0.99.0 判为大于 0.146.0,同时让「取最大值」与
  「只向前推进」两处判错)。
2026-08-03 20:55:37 +08:00
zhiyu 1e08c4c56f test(server): 补齐 settings 契约用例的 Codex 版本号字段
新增的三个设置键会出现在 GET /api/v1/admin/settings 响应里,
契约用例的期望 JSON 需同步,否则 -tags=unit 下断言失败。
2026-08-03 20:26:50 +08:00
zhiyu 2eb24814fe fix(codex): 强制统一出站身份并让客户端版本号跟随官方发布
上游 /backend-api/codex 在容量紧张时按客户端身份分优先级降载,被降载的请求
HTTP 200 后立刻推流内 server_is_overloaded。此前网关对配不出官方身份的客户端
整体回退到硬编码的 codex_cli_rs/0.144.1(落后官方 4 个发布),这些请求稳定
落在被优先丢弃的一侧。

- 强制统一出口:所有 OAuth 出站的 User-Agent / originator / version 一律改写
  为网关规范身份,客户端自报身份不参与构造;HTTP / 透传 / WS / alpha-search /
  探针全覆盖。compat 桥接故意删除 originator 的路径保持 no-op。
- 版本号收敛为单一来源,运行时优先级为面板覆写 → 自动同步值 → 内置常量;
  UA 与 version 头同源派生,不再各自硬编码。
- 新增 3 小时自动同步官方客户端最新稳定版,面板可关闭,无需为跟版本而发版。
- 流内 server_is_overloaded / slow_down 改为先在同账号有界重试再切号,并标记为
  请求级瞬时故障,不再据此临时封禁账号。
- 移除被取代的降载身份黑名单、浏览器 UA 兜底及其辅助函数。
2026-08-03 20:14:58 +08:00
spongehah 1b04e03cc4 fix(prompt-audit): parse responses output text 2026-08-03 17:38:04 +08:00
wucm667 2be047d1cd [verified] test(subscription): update unit repository stubs 2026-08-03 17:22:18 +08:00
wucm667 db725a775a [verified] fix(subscription): serialize concurrent renewals 2026-08-03 16:59:05 +08:00
Wesley Liddick 825ca7b1fc Merge pull request #5183 from rick147/codex/feat-openai-reset-credit-cache
feat(openai): refresh reset credit state after quota reset
2026-08-03 16:01:10 +08:00
shaw 54a2bcfd15 fix(openai): harden reset-credit refresh and account recovery
Review follow-ups on the reset-credit caching flow:

- Recover account state BEFORE (and independently of) the reset-credit
  display cache. A failed cache refresh could previously abort the run and
  leave the account rate-limited — the very reason the credit was spent
  (#3672 / #3740). The recovered account row is now returned even when the
  cache refresh fails.
- Run the post-reset bookkeeping on a detached, time-boxed context and give
  the panel reset call a larger timeout. A client abort no longer strands a
  consumed (non-refundable) credit with an unrecovered account, and the
  chained upstream calls can no longer exceed the client timeout and invite a
  retry that spends a second credit.
- Persist the reset-credit snapshot through POST /accounts/:id/quota/refresh
  instead of a side-effecting GET flag, so the write is covered by the audit
  middleware. A rejected snapshot write now degrades to cache_persisted=false
  instead of turning a successful upstream read into a 502 that left the card
  without a credit count and the reset button permanently disabled.
- Reject snapshots whose positive count carries no expiration entries, and
  drop expired credits (clamping the count) when rehydrating, so a stale
  cache can no longer light up the reset button.
- Keep nil quota / rate-limit services nil in the handler's interface fields;
  storing a nil *Service made the "not enabled" guards non-nil.
- Time-box the usage-refresh suppression and reuse handleAccountUpdated so the
  patched row also enters the auto-refresh silent window.
2026-08-03 14:40:55 +08:00
Wesley Liddick 27e8f69a9e Merge pull request #5171 from heathermhuang/codex/composite-reasoning-policy
feat(composite): enforce reasoning effort policy
2026-08-03 11:28:38 +08:00
Wesley Liddick 684ab20a0b Merge pull request #5164 from wucm667/fix/issue-5099-messages-temp-failover
fix(openai): fail over Messages temporary account errors
2026-08-03 11:28:16 +08:00
Wesley Liddick 0173830df6 Merge pull request #5193 from wucm667/fix/issue-5187-stripe-refund-idempotency
fix(payment): make Stripe refunds idempotent
2026-08-03 11:28:05 +08:00
Wesley Liddick 724565e4aa Merge pull request #5199 from wucm667/fix/issue-5191-refund-balance-force
fix(payment): require force for insufficient refund balance
2026-08-03 11:27:52 +08:00
Wesley Liddick a1f1a0cc6b Merge pull request #5194 from wucm667/fix/issue-5189-persist-unsettled-usage
fix(billing): retain usage logs on billing failure
2026-08-03 10:27:30 +08:00
wucm667 3c20f9a666 [verified] fix(payment): require force for insufficient refund balance 2026-08-02 23:11:17 +08:00
shaw e1b76e2245 fix(codex): normalize load-shed originators to avoid upstream capacity shedding
上游 /backend-api/codex 按 Originator 头分桶调度容量:落在降载桶的请求即使返回
HTTP 200,也会立刻推 SSE `event: error`(code=server_is_overloaded)并以
response.failed 收尾。2026-07-29 起 codex-tui 落入降载桶,codex_cli_rs 正常——
判定因子是 originator 而非 User-Agent(codex_cli_rs 配 curl UA 亦可正常返回)。

网关会把该错误判定为瞬时上游故障并冷却账号,对外表现为 Codex 账号频繁过载不可用:
server_is_overloaded → isOpenAITransientProcessingError →
shouldCooldownOpenAITransientUpstreamError → 账号冷却 → 客户端 503。

本项目有三处降载身份来源:浏览器 UA 兜底的默认 UA、客户端透传的真实 TUI 身份、
以及指纹缓存注入探针的 UA。

修复收口在 enforceCodexIdentityHeaders——HTTP / 透传 / WS 握手 / compat 桥接 /
探针 / PAT / 模型列表 / alpha-search 八条出站路径共用的唯一纯函数收口点:

- 新增 NormalizeCodexClientIdentityToCLI,把降载桶身份改写为 codex_cli_rs,
  只替换身份段并裁掉尾部 (name; version) 客户端标识组,保留版本 / OS / 架构 /
  终端指纹;改写后 originator 与 UA 首段仍然配套,不破坏 #3901 的配对不变式,
  且改写幂等。
- DefaultOpenAICodexUserAgent 从 TUI 身份改为 CLI 身份(浏览器兜底路径上最大的
  降载身份来源)。
- 管理端 Codex UA 的 placeholder / hint 原本在把管理员往降载桶引导,一并修正。

新增 gateway.disable_codex_originator_normalization(默认 false,即归一化开启),
供上游调整分桶后回滚。该开关经 NewOpenAIGatewayService 发布为进程级快照,故必须
保持反义命名:正向命名的 Go 零值 false 会让未经 viper 加载而手工构造的 Config
静默关掉全局保护,viper.SetDefault 救不了这条路径。已加用例钉住该属性。

降载桶集合是上游容量策略快照而非协议常量,上游调整分桶后需同步修订。
2026-08-02 23:00:12 +08:00
rick147 f970bd48c9 fix: stop scheduler work after request cancellation 2026-08-02 21:32:31 +08:00
rick147 a0802f00b6 feat: cache OpenAI reset credit details 2026-08-02 21:32:31 +08:00
wucm667 0b9f40e230 fix(billing): retain usage logs on billing failure 2026-08-02 20:50:13 +08:00
wucm667 0b26acac07 fix(payment): make Stripe refunds idempotent 2026-08-02 20:27:44 +08:00
Wesley Liddick c043c24774 Merge pull request #4925 from Brisbanehuang/feat/group-profit-control
feat(scheduler): 分组级利润控制——按账号倍率过滤 token 调度候选
2026-08-02 18:19:13 +08:00
Wesley Liddick 11c1e944b9 Merge pull request #4911 from Brisbanehuang/feat/upstream-billing-rate-writeback
feat(billing-probe): 探测成功后可选将上游声明倍率自动同步为账号倍率
2026-08-02 18:18:50 +08:00
Wesley Liddick d99ee72911 Merge pull request #4896 from Brisbanehuang/feat/upstream-billing-probe-multi-platform
feat(billing-probe): 上游计费倍率探测放宽到全部 API-key 平台账号
2026-08-02 18:18:37 +08:00
r266-tech 30d2589ef0 fix(openai-ws): preserve terminal event on lease loss 2026-08-02 11:47:59 +08:00
shaw dec47e8fae fix(profit-control): stop leaking profit policy, close veto livelock, restore passthrough turn pricing
审计修复,逐条如下。

H1 利润策略泄露给所有普通用户
  profit_control_enabled / profit_min_margin / profit_safety_buffer 从
  dto.Group 移到 dto.AdminGroup(后者内嵌前者),赋值相应从
  groupFromServiceBase 移到 GroupFromServiceAdmin;前端 TS 同步从 Group 移到
  AdminGroup。dto.Group 是 GET /api/v1/groups/available 的响应体,该响应本就带
  rate_multiplier,相乘即可反推运营方上游采购成本上限。
  api_contract_test.go 的 /groups/available golden JSON 回滚这三个字段,并把
  fixture 改成非零值(require.JSONEq 是精确比对,缺字段即失败)。
  新增 dto 层边界测试:普通用户 DTO 不含三字段、管理员 DTO 仍含。

M1 利润终检 continue 与 failover 503 退避互动产生活锁
  FailoverState 新增 profitVetoedAccountIDs / profitVetoCount 与
  RecordProfitVeto():加入排除集 + 计数,达 maxProfitVetoAttempts(10) 返回
  FailoverExhausted。HandleSelectionExhausted 的 503 清空分支改为清空后把利润
  否决的账号放回排除集;若排除集已全部由利润否决贡献,清空不会带来任何新候选,
  直接判定耗尽(否则 SwitchCount 永不前进、退避条件永远成立,每 2s 空转一轮)。
  五个 handler 否决点(gateway_handler ×2 / responses / chat_completions /
  gemini_v1beta)改为经 RecordProfitVeto 决策,耗尽时按无可用账号终止。
  回归测试钉死:503 之后持续利润否决必须有限步终止且不 spin;未启用利润控制的
  请求退避语义完全不变。

M2 排队等槽后才终检,延迟可放大到 N × WaitPlan.Timeout
  OpenAI 侧选号循环(自有 failedAccountIDs map,非 FailoverState)新增
  recordOpenAIProfitVeto + handleOpenAIProfitVetoExhausted,共用同一上限语义。
  覆盖 responses / messages-dispatch / chat_completions / alpha_search /
  embeddings / images / grok_media 七处,以及 WS 两处否决分支。

M4 ws_v2 透传 ingress 绕过 per-turn 重定价(选方案 B:最小止血)
  透传 relay 只回调 AfterTurn、没有任何 turn 起始回调,hooks.BeforeTurn 永远
  不触发,而 handler 把 turnPricingAt 初始化成建连时刻 ⇒ 透传连接全部 turn 按
  建连时刻的高峰因子结算,客户端峰前建连保活即可全程谷价——正是本 PR 想堵的
  漏洞。改为 openAIWSTurnPricing 零值起步、只由 BeforeTurn 冻结;透传路径保持
  零值,RecordUsage 回退记录时刻,与引入利润控制前的基线一致。
  未选方案 A(给透传补 turn 起始回调):passthrough_relay.go 是 #5167 刚修过的
  取消传播/close frame 时序敏感区;且 BeforeTurn 还承担 turn>1 的并发槽位抢占,
  接进去等于给透传连接引入 per-turn 抢槽,风险远超本次修复范围。透传仍有建连时
  的准入门,只是没有 turn 级复核,已在两处注释写明。
  测试:service 层钉死透传 ingress 不触发 BeforeTurn(含失败时的复核指引),
  handler 层钉死零值语义与逐 turn 覆盖。

M5 装门读分组走了带账号计数聚合的 GetByID
  SchedulerSnapshotService 新增 GetGroupByIDLite,openai/gateway 两处装门改用
  之。门只需要平台/倍率/利润/高峰字段,且该查询发生在「是否启用利润控制」判定
  之前,未启用的分组同样付代价。两个测试 stub 的 GetByID 改成 panic 守卫。

M6 认证快照注释与真实读取路径相反
  门解析优先取 ctxkey.Group,而它就是本快照物化出来的对象,直连流量走的正是这
  条路。改正注释,与 api_key_repo.go 投影处的说明对齐,避免后人照旧注释删列。

M3 rate_multiplier 为 nil 时利润门 fail-closed(不改行为,加护栏)
  保留 fail-closed。补 repository 层测试钉死账号调度快照的 full/metadata 两份
  payload 都必须保留 RateMultiplier(含 0 值),漏列在 CI 就红。

L1 迁移号注释 191 / 191-192 改为实际的 192/193。
L2 admin group Create 的利润配置预校验改用与 CreateGroup 一致的归一化平台
   (新增 service.NormalizeGroupPlatform,两边共用)。保留预校验而非删除:
   service 层返回的是无类型 error,经 ErrorFrom 会变成 500,删掉会把合法的
   400 降级成 500。
L3 前端利润校验的上界改为判定换算后的小数(后端按小数校验 [0,1)),
   99.999% 会四舍五入进位成 1.0 而被后端 400;i18n en/zh 同步改为 0-99.99。
L4 clampProfitControlThreshold / profitControlOverThreshold 抽为共用函数,
   线上装门/否决点与 profit-preview 不再各自实现,附边界语义测试。
L5 profit-preview 补「默认 D 有账号但最低有效 D 归零」的告警(两档都为 0 由
   既有告警覆盖,不重复)。
2026-08-01 22:39:33 +08:00
Brisbanehuang fad2f215e8 fix(profit-control): decouple gate from image intent, propagate gate via selection, restore eager sticky fallback
Review fixes for the profit-control feature commit:

- Image intent no longer disables the profit gate. The shared /v1/responses
  handler previously skipped the pricing context (and therefore the gate)
  whenever the platform-wide image intent predicate matched, which includes
  Codex's passive image_gen namespace declaration: any client could disable
  admission control for anthropic/gemini/antigravity groups by declaring a
  namespace tool in the request body. Both /v1/responses paths now always
  install the token pricing context; image intent only drives capability
  routing and image billing. Mixed token+image requests stay token-gated;
  only dedicated media endpoints remain out of scope.
- Out-of-scope paths are now explicitly suppressed instead of implicitly
  ungated: Grok media (billed by media multipliers; also prevents in-flight
  video lookups from turning into spurious 404s), OpenAI-group count_tokens
  (unbilled), and Live calls (duration billed) carry a suppress marker that
  every install point honors, including the defensive scheduler-entry
  install.
- The gate resolved during selection now travels back to handlers on the
  AccountSelectionResult. The shared gateway installed the gate only on a
  scheduler-local context, so handler-side post-slot terminal rechecks and
  post-admission sticky binding were no-ops for anthropic/gemini/
  antigravity/shared-grok requests (and for composite-routed member groups
  on the OpenAI path). Handlers re-apply the carried gate via
  ContextWithSelectionProfitGate before the terminal recheck and binding;
  the WS acquired-selection branch gained the previously missing recheck.
- Sticky binding semantics restored for ungated traffic:
  BindStickySessionAfterProfitAdmission falls back to the official eager
  bind when no gate is installed (wait paths lost their only binding point
  otherwise), reads the pre-existing binding at bind time only when gated
  (removes the unconditional per-request Redis read the feature added to
  the shared handlers), and the legacy engine's three selection-time
  binding writes are skipped under a gate so a terminally vetoed account
  can no longer become the new sticky target.
- Responses WS connections re-freeze pricingAt and re-resolve the gate at
  every turn (BeforeTurn) and bill each turn with its own instant, closing
  the connect-at-valley/bill-at-valley window; a turn that fails the
  recheck closes the connection so the client reselects on reconnect.
- Terminal recheck no longer swaps a DB-fresh selected account for a
  staler snapshot object (UpdatedAt guard), and observer counters are
  documented as per-evaluation.
- Migrations renumbered 191/192 -> 192/193 after the passkey migration
  landed upstream as 191.

New regressions: selection-carried gate propagation (control group proves
the pre-fix no-op), image intent not disabling the shared gate, eager
binding fallback without a gate (both services), gated
read-failure/sentinel-miss binding semantics, legacy-engine deferred
binding under a gate with eager behavior preserved ungated, turn-level
pricing refresh (config re-resolution, scheduled-group precedence,
suppress, mid-connection disable), and suppress-marker coverage for the
request pricing context.
2026-08-01 22:39:32 +08:00
Brisbanehuang 20ad5ec506 feat(scheduler): per-group profit control for token account admission
Group pricing (rate multiplier, peak windows, per-user overrides) and
account cost (accounts.rate_multiplier) already live side by side, but
nothing stops the scheduler from handing a request to an account whose
cost multiplier exceeds what the group's pricing can profitably serve.
Add an opt-in per-group profit gate that filters scheduling candidates
by a margin rule, while ordering, scoring, stickiness and breakers keep
working unchanged among qualified accounts.

Admission rule: an account qualifies iff U <= D * (1 - min_margin -
safety_buffer) within a small relative epsilon, where U is
accounts.rate_multiplier (0 is legal; missing/negative/NaN/Inf are
conservatively rejected as invalid) and D is the requester's effective
downstream multiplier (user-group override ?? group default, times the
group peak factor) frozen at the request's pricing instant.

- groups gain profit_control_enabled / profit_min_margin /
  profit_safety_buffer (migration 191); the durable auth-cache
  invalidation trigger additionally watches the profit and pricing
  columns (migration 192) so out-of-band group edits cannot leave
  stale auth snapshots; GetByKeyForAuth explicitly projects the new
  columns and the API-key auth snapshot version is bumped to force a
  refresh of pre-existing snapshots
- request-level pricing instant: token entry points install pricingAt
  into ctx; the profit threshold D and the RecordUsage peak factor
  read the same instant, so one request never changes price mid-flight
  across waits/retries/failover (media and unwired paths keep the
  existing record-time semantics)
- the gate covers token requests on openai, anthropic, gemini, grok
  and antigravity groups: OpenAI-family handlers via
  WithOpenAIRequestPricingContext (responses incl. WS bridge, chat
  completions, messages, embeddings, alpha search), the shared gateway
  via WithGatewayTokenRequestPricing (messages, chat completions,
  responses, gemini model actions); composite groups cannot enable it
  directly; image/video/models/usage/count_tokens stay ungated and an
  explicit image-generation intent suppresses the gate end to end
- post-slot recheck: after a slot is acquired the account is re-read
  via SchedulerSnapshotService.GetAccount (scheduler cache, then DB;
  only when both fail the check fails open with WARN + metric); a
  vetoed account releases its slot and joins the request's exclusion
  set for reselection; sticky bindings are written only after the
  final check passes, and an over-threshold sticky account is skipped,
  not deleted, so it comes back once its rate recovers
- sticky-session cache contract: GatewayCache.GetSessionAccountID now
  returns ErrStickySessionNotFound on a miss (mapped from redis.Nil in
  the repository implementation, mirroring ErrRefreshTokenNotFound) so
  the profit sticky path can distinguish "no binding yet" from a real
  read failure without importing the cache driver in service code
- cross-group re-entry (composite parent -> member group) resolves the
  gate against the member group and clears a stale parent gate instead
  of letting a foreign threshold veto accounts
- per-platform/group activity counters (installs, threshold vetoes,
  invalid-rate vetoes, refresh failures) for observability
- admin UI: profit-control section on the five platforms' group forms
  with percent input, validation and platform-switch reset; group
  create/update/duplicate normalize and validate the config at a
  single choke point
- cmd/profit-preview: offline what-if tool that replays the production
  admission semantics over an exported config/account/override/model
  dump, reports per-model admitted-account counts under the default
  and the worst-case (lowest user override) D, and surfaces probe-sync
  staleness as warnings without affecting admission

Tests: service unit coverage for gate resolution/veto/threshold
epsilon/pricing instant/suppress marker/scheduler filtering and
post-slot recheck (incl. -race on the profit surface), unit-tagged
handler slot-recheck and capability-mapping regressions, sqlmock and
real-PostgreSQL integration regressions for the GetByKeyForAuth
projection and the migration-192 trigger watch list, API contract
update, and frontend specs for the five-platform form helpers.
2026-08-01 22:39:31 +08:00
shaw 0b6b4ea956 fix(billing-probe): govern the automatic account rate write-back
- 自动写回值域治理与留痕:上游声明值必须 > 0 且 <= 100 才写回。0 会让
  accountCost 恒为 0(账号总额/日/周配额与成本告警全部静默失效),极大值可
  一次打爆配额并污染成本报表;越界时保持原倍率、记 WARN,探测快照照常记 ok。
  写回成功时记结构化 slog(account_id / 旧值 / 新值 / source),并在快照中
  新增 synced_rate_multiplier 记录本次写回值——后台任务裸 SQL 不产生
  audit_logs,这两处是唯一可追溯来源。管理员手工设 0 不受影响。
- 写回改用 resolved_rate_multiplier(不含高峰的基准倍率):effective 含探测
  那一刻的高峰系数,写回会把一个探测周期的峰值/谷值冻结进静态列,而展示与
  调度用 upstreamBillingRateAt 按当前时间重算高峰,两者会持续不一致。
- 倍率解析失败不再污染公共探测路径:账号级值域/精度只在该账号已开启同步、
  真要写回时才有影响;未开同步的账号照常记 ok 快照,不累计 failure_count、
  不进入指数退避。
- 单账号编辑补 service 层守卫:同步开启时拒绝手工倍率(新增
  UPSTREAM_BILLING_RATE_SYNC_CONFLICT,与批量路径同族),此前只有前端
  disabled 挡人,直接 PUT /admin/accounts/{id} 可写入并活到下次成功探测。
  判断的是本次请求生效后的状态,"关同步 + 改倍率"同请求仍然放行。
- 修正开关反推方向:不再由 rate_sync=true 推出 probe=true,否则一条"同步开、
  探测键缺失"的僵尸记录会在任意一次无关编辑时静默打开周期性外呼;改为探测
  关闭/缺失一律把同步归零。
- 删除死代码 UpdateWithUpstreamBillingProbeEnabled(PR 删接口后生产已无调用
  方),其回滚测试改为直接覆盖生产路径 UpdateWithAccountBillingSettings。
  upstreamBillingRateSyncEnabled 不再是只服务测试的假门控,现为写回前置过滤,
  SQL CAS 仍是权威门控,两侧均加注释说明分工。
- en/zh 文案补充:同步的是不含高峰的基准倍率;开启同步会连带打开自动探测。
2026-08-01 22:11:10 +08:00
Brisbanehuang b0f5007f04 feat(billing-probe): optionally sync account rate from upstream declared rate
Successful upstream billing probes already persist the upstream-declared
rate as a display-only snapshot. Add a per-account opt-in that writes
that declared rate back to the account's rate_multiplier, so the account
cost basis follows upstream repricing automatically instead of drifting
until an operator notices.

- new per-account flag upstream_billing_rate_sync_enabled stored next to
  the probe flag in account extra: enabling sync force-enables the
  probe, disabling the probe cascades sync off, and eligibility follows
  IsUpstreamBillingProbeIdentity (tightened from any non-empty platform
  to an explicit whitelist of the five supported API-key platforms so
  future platforms do not silently inherit probe/sync semantics)
- only a successful probe whose declared rate survives validation
  (finite, within bounds, not rounded to zero at the rate_multiplier
  decimal(10,4) scale) writes back; failed/unsupported/invalid probes
  leave rate_multiplier unchanged
- the writeback rides the existing snapshot CAS transaction:
  UpdateUpstreamBillingProbeSnapshot takes an optional rateMultiplier
  and applies it atomically with the snapshot under the same
  identity/snapshot compare-and-swap, so a probe result observed on a
  stale account cannot clobber a concurrent admin edit
- admin edit goes through UpdateWithAccountBillingSettings, which
  applies the form without overwriting a rate that a probe synchronized
  after the edit form was loaded (nil rateMultiplier = not edited);
  once sync is enabled the edit form shows the rate as managed
- bulk update rejects a manual rate_multiplier change when any target
  account has rate sync enabled (whole batch fails with a dedicated
  error so partial writes cannot bypass the sync ownership)
- frontend: sync toggle with hints in the edit modal (probe/sync
  enable/disable coupling enforced in the form), synced-rate tooltip on
  the rate cell, bulk edit modal warns and blocks rate edits that hit
  sync-enabled accounts; en/zh copy updated
- tests: service unit tests for sync gating/validation/cascade, sqlmock
  repo tests for the extended CAS, real-PostgreSQL integration tests
  (rate written only for successful+enabled accounts, manual rate
  protected after sync disabled, admin edit preserved across concurrent
  probe sync), handler/API contract updates, frontend specs for modal
  coupling, bulk rejection and rate cell
2026-08-01 22:11:09 +08:00
shaw 56f3d3c9b0 fix(billing-probe): 收口探测资格放宽后的抑制清单与调度信任面
#4896 把上游计费探测从 OpenAI API-key 放宽到全部 API-key 账号后,
遗留了四处需要收口的问题:

- M1 官方域抑制清单补 ollama.com。Ollama Cloud 是本仓一等支持配置
  (platform openai/anthropic + type apikey + base_url
  https://ollama.com/v1),放宽后 anthropic 侧这类账号会每个探测周期
  拿 Ollama Key 请求 ollama.com/v1/sub2api/billing 并恒定落空,正是该
  守卫注释声明要防的行为。后缀匹配已覆盖 www.ollama.com 等子域,
  notollama.com 等形似域不受影响。
- M2 legacy 低倍率优先排序补平台门控。newOpenAILegacyUpstreamRateOrder
  遍历全部候选且无平台门控,与 openAIUpstreamCostFactors 的门控不对称;
  放宽后 grok 账号的上游自报倍率开始影响 legacy 调度排序,而实际结算走
  本地倍率,中转方自报低价即可吸流量。现补上同一道门控,使调度侧信任面
  回到 PR 前状态(探测资格的放宽保持不变)。
- L1 修正 IsUpstreamBillingProbeIdentity 注释。原注释称类型限制的依据是
  "OAuth/Bedrock 没有静态 API key",但 AccountTypeUpstream(antigravity
  中转账号)同样是 base_url + 静态 api_key 却也被排除。仅改注释如实说明
  取舍,不改行为。
- L2 批量探测空选文案去掉 OpenAI 限定(en/zh 成对)。
- L3 unsupported 状态改用加长退避(interval 的 8 倍,仍按 24h 封顶)。
  放宽后大量官方域账号会落 unsupported 并按常规 interval 重排,占满每周期
  20 个名额,把真正接入 sub2api 的中转账号挤到后面。封顶保证上游后来接入时
  最迟一天内会被重新发现;Retry-After 更长时原样保留不被缩短;手动探测不受
  退避影响。

测试:ollama.com 官方域行为级与 host 匹配矩阵用例、legacy 排序平台门控
(含混合候选集)用例、unsupported 退避上下界与 runner 跳过/手动探测放行
用例;同步更新既有 unsupported 的 next_probe_at 断言。
2026-08-01 21:54:11 +08:00
Heatherm Huang d60f4e442b feat(composite): enforce reasoning effort policy 2026-08-01 21:15:44 +08:00
shaw 21aacde0b3 fix(openai-ws): keep downstream writes off the relay cancellation context
coder/websocket arms a context.AfterFunc that hard-closes the connection
when a write context is canceled, and AfterFunc stop does not wait for a
callback that already started. An external cancellation (e.g. ingress
lease loss) landing inside the disarm window of an already-successful
downstream write could therefore kill the TCP connection before the
retry close frame (1013) was written, leaving the client with a bare
EOF. Mirror the read side: bound downstream writes with the write
timeout only, and rely on the explicit Close/CloseNow performed by every
relay exit path for teardown.

Fixes the flaky TestPassthroughLifecycle_LeaseLossSendsRetryClose.
2026-08-01 20:32:10 +08:00
wucm667 ddf4c6fd81 fix(openai): fail over Messages temporary account errors 2026-08-01 18:45:33 +08:00
neboyang ff7ab2308d test(oauth): lock AuthorizeURL to claude.com/cai endpoint 2026-08-01 15:03:59 +08:00
neboyang 0204ce285e fix(oauth): use claude.com/cai authorize endpoint
Align Claude OAuth authorize URL with Claude Code CLI
(https://claude.com/cai/oauth/authorize).
2026-08-01 14:55:09 +08:00