Commit Graph
4106 Commits
Author SHA1 Message Date
shaw dbb42881c0 fix(openai): default OAuth identity to codex-tui 2026-08-07 09:56:03 +08:00
shaw 287a9f386b fix: 修复腾讯验证码票据过期与区域切换 2026-08-06 21:27:06 +08:00
shaw 8e102b3a0f fix: 完善腾讯验证码区域适配与 CSP 白名单
修复国内站和国际站 SDK 构造、验证容器、票据重置及动态资源加载问题,并补充认证流程回归测试。
2026-08-06 20:34:37 +08:00
Wesley Liddick e08aee49ed Merge pull request #5266 from shentry/fix/transient-streak-rate-dependence
fix(openai): keep transient failure streak from resetting on sparse traffic
2026-08-06 14:11:49 +08:00
Wesley Liddick c9e60d1f26 Merge pull request #5031 from keaipiao/fix/easypay-error-utf8
fix(payment): preserve UTF-8 in EasyPay errors
2026-08-06 14:10:41 +08:00
Wesley Liddick 47c03c75d8 Merge pull request #5232 from fengshao1227/fix/billing-quantize-monetary-scale
fix(billing): quantize usage billing amounts to the NUMERIC(20,8) scale
2026-08-06 14:00:04 +08:00
github-actions[bot] aac53afe0e chore: sync VERSION to 0.1.171 [skip ci] 2026-08-04 13:41:47 +00:00
feeeei 26e0a89323 人机验证增加阿里云验证码 2.0
沿用腾讯天御验证码引入的多服务商模型:aliyun_captcha_enabled 作为独立
开关,与 Cloudflare Turnstile、腾讯天御三方互斥(保存校验 + 运行时
CAPTCHA_PROVIDER_CONFLICT)。后台「安全与认证」合并为单张人机验证卡片:
总开关 + 服务商单选(Turnstile / 腾讯天御 / 阿里云),选中即启用该家并
关闭其它,落库仍是三个独立开关键,由前端映射保证互斥。

阿里云侧同时支持 aliyun 中国站与国际站(alibabacloud.com):两站前端脚本、
region 取值与服务端 API 完全一致,仅账号与 AccessKey 相互独立,因此由
「服务地域」决定线路即可——中国内地走 captcha.cn-shanghai.aliyuncs.com,
非中国内地(新加坡)走 captcha.ap-southeast-1.aliyuncs.com,AccessKey
取自持有该实例的账号,无需在配置中区分站点。

- AliyunCaptchaService 对称 TencentCaptchaService:服务端校验走官方 SDK
  VerifyIntelligentCaptcha,调用异常按 fail-closed 拦截,与 Turnstile
  网络错误行为对称;保存设置时真实探测 AK/SK 有效性
- 保护面对齐腾讯扩展入口:VerifyTencentCaptchaIfEnabled 通用化为
  VerifyActionCaptchaIfEnabled,OAuth 登录启动、passkey 登录在阿里云
  启用时同样拦截;Turnstile 维持既有覆盖不扩大
- 前端 AliyunCaptchaWidget 为表单内预验证按钮(popup 模式),同时暴露
  verify() 供 OAuth 启动、passkey 等动作入口程序化弹窗;未预验证直接
  提交时弹窗兜底。SDK 按钮绑定异步完成,弹窗未出现前按 tick 重试触发,
  并轮询弹窗可见性识别用户关闭
- captchaVerifyParam 复用 turnstile_token 请求字段提交;公开设置下发
  aliyun_captcha_enabled / scene_id / prefix / region
- CSP 放行验证码 CDN:script-src/style-src 加 *.alicdn.com
2026-08-04 20:57:15 +08:00
shentryandClaude Opus 5 7d38e67120 fix(openai): keep transient failure streak from resetting on sparse traffic
The account+model transient breaker reset its failure streak whenever the
gap since the previous failure exceeded a one-minute window. That made the
breaker's sensitivity a function of request rate rather than upstream
health: a gateway called less often than once a minute never advanced past
streak 1, where the cooldown is zero, so a consistently broken account was
never blocked. Every request re-selected it, paid a full upstream attempt,
and only then failed over to a healthy account.

Observed on a low-traffic deployment: two accounts returning 500 and 503
stayed in rotation indefinitely, logging `failure_streak: 1, cooldown_ms: 0`
on every request and adding ~750ms to each one before a working account was
reached.

The streak is already cleared on success — recordSuccess deletes the entry,
and every OpenAI handler reports the schedule result — so the time-based
reset is not needed to recover a healthy account. Keep a TTL purely to bound
the map for account+model pairs that stopped being used, and raise it well
above the cooldowns so it no longer doubles as a streak reset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 17:59:16 +08:00
Wesley Liddick 8b3fe664dc Merge pull request #5261 from lyen1688/feat/tencent-captcha-gate
新增腾讯天御验证码认证门禁
2026-08-04 16:39:55 +08:00
Wesley Liddick a4d263f62f Merge pull request #5224 from wucm667/fix/issue-5190-lock-subscription-renewal
fix(subscription): serialize concurrent renewals
2026-08-04 16:28:52 +08:00
Wesley Liddick ae81dfd933 Merge pull request #5177 from r266-tech/fix-nonpassthrough-write-context
fix(openai-ws): preserve terminal event on lease loss
2026-08-04 16:28:22 +08:00
Wesley Liddick 35cab3c814 Merge pull request #5258 from feeeei/fix/model_plaza
fix(model-plaza): Model Plaza image model price display is inconsistent with the actual price
2026-08-04 16:27:53 +08:00
Wesley Liddick 770e35b474 Merge pull request #5040 from coo1white/upstream-pr/grok-cli-0.2.114
fix(grok): bump pinned Grok CLI version to 0.2.114
2026-08-04 16:16:20 +08:00
Wesley Liddick 846dd310a3 Merge pull request #5158 from neboyang/feat/claude-oauth-authorize-url
fix(oauth): use claude.com/cai authorize endpoint
2026-08-04 16:15:50 +08:00
Wesley Liddick 9b4575e434 Merge pull request #5243 from wucm667/feat/issue-5240-dashboard-username
fix(dashboard): show usernames in spending ranking
2026-08-04 16:15:21 +08:00
Wesley Liddick 1f4cfc44c1 Merge pull request #5226 from spongehah/codex/fix-prompt-audit-output-text
fix(prompt-audit): parse responses output text
2026-08-04 16:15:04 +08:00
lyen1688 e592c5f9e0 新增腾讯天御验证码认证门禁 2026-08-04 15:09:29 +08:00
feeeei 785b61d424 修复模型广场图片模型价格展示与实收口径不一致
图片计费模型的广场展示价按实收口径计算:档位单价取
分组图片价 > 渠道档位价 > 渠道默认按次价;分组开启生图
独立倍率时实付倍率取独立倍率,不取分组/专属倍率。
2026-08-04 13:48:19 +08:00
shaw 2d3e845205 perf(codex): 版本同步主路径改用 /releases/latest 并保留列表回退
自动同步原本每次拉 `/releases?per_page=30`,实测响应 10,017,686 字节——该仓库
预发布极密集,30 条里只有 2 条稳定版(0.145.0 已排到第 26 位),页大小是为了
「窗口里至少有一条稳定版」而定的,不能靠调小来省流量。

改为主路径走 `/releases/latest`(实测约 0.3MB):该端点本身排除 draft 与
prerelease,直接给出最新正式发布,不受预发布密度影响。复用端口上已有的
FetchLatestRelease,不新增任何 HTTP 代码,代理配置、User-Agent、可选 token
与跨主机重定向剥离 Authorization 的行为全部沿用。

保留列表扫描作为回退:latest 是跨 tag 家族按 published_at 取的,若同仓库其他
组件(如 rusty-v8-*)某天发布正式 release 而成为 latest,主路径会被 rust-v
前缀过滤挡掉,此时必须扫一页才能继续跟随,否则版本号会静默停更。

两条路径共用 latestCodexStableReleaseVersion 的同一套过滤(前缀 / draft /
prerelease / 版本号形态),语义不会分叉;单条 latest 也要过版本号校验——仓库里
确实存在 rust-vv0.99.0-alpha.8 这类畸形 tag。只向前推进、抓取失败保持既有值
两条不变式未改动。

测试覆盖:主路径命中时不再拉列表页、四种拿不到(异家族 / 预发布 / 抓取失败 /
空对象)都回退、畸形 tag 被主路径拒绝后由回退兜底、两路径皆失败时保值。
2026-08-04 10:51:04 +08:00
wucm667 80af44cad7 [verified] fix(dashboard): show usernames in spending ranking 2026-08-04 00:43:05 +08:00
shaw c899c8cf37 fix(codex): 管理员配置的 UA 只贡献指纹,版本段一律用生效版本重建
面板「OpenAI Codex UA」此前只在客户端 UA 是浏览器型时才生效,本分支把它提升为
所有 OAuth 出站的规范身份来源,作用域扩大了一个数量级。而它的旧 placeholder 原文
就是 codex_cli_rs/0.144.1 (...):照抄填写过的存量部署会被永久钉在 0.144.1 上——
该值不低于上游门槛 0.144.0,version 头也随之变成 0.144.1,绕过版本自动同步,
稳定落在上游优先降载的那一侧,且面板、日志、审计都不提示。账号级自定义 UA
(credentials.user_agent)有完全相同的陷阱。

改为:管理员配置的 UA 只贡献客户端名与 OS / 架构 / 终端指纹(这是该输入框唯一
不可替代的价值),版本段一律用当前生效版本重建。存量的陈旧配置无需迁移即自愈,
UA 与 version 头从此由构造保证同源,原先「覆写 UA 陈旧时二者不一致」的次优解消失。

- 新增 openai.SetCodexUserAgentVersion:重建首段版本声明,OS / 架构 / 终端指纹
  原样保留;尾部官方客户端标识组 (name; version) 与首段同源,一并更新,避免拼出
  首段声明新版本、尾部仍是旧版本的自相矛盾身份;非官方括号组(如 OS 组)不动。
- resolveCodexOutboundIdentity 收敛为单一funnel:生效版本只从规范身份取,
  候选 UA 的版本段不再参与判定。
- GetOpenAICodexCanonicalUserAgent 重建面板 UA 的版本段;原「面板值等于兜底常量
  视同未填」的特例随之成为恒等变换,删除。
- 补齐 4 处静默丢弃账号级自定义 UA 的出站路径(WS 握手、账号测试 x2、用量探针):
  它们先写 customUA 再调不带 override 的收口,赋值随即被覆盖,既是行为不一致
  也是死代码;账号测试尤其要紧——注释写着「与真实转发一致」而实际不一致。
- 补测试:SetCodexUserAgentVersion / CodexUserAgentVersion 包内用例、陈旧面板 UA
  与账号 UA 的版本重建回归、WS 握手尊重账号级 UA。
2026-08-03 21:50:49 +08:00
zhiyu 4c4ff36380 perf(codex): 版本同步间隔改为 6 小时并加启动防抖
- 同步间隔 3h → 6h:客户端版本是天级变化,6 小时足够跟上,
  对 GitHub 的调用降到每天 4 次。
- 新增启动防抖:借同步设置行自身的 UpdatedAt 判断,若同步值仍在一个周期内
  则跳过启动同步。频繁重启、滚动发布或崩溃重启原本会把「启动即同步」
  放大成对 GitHub 的连续请求;首次部署尚无同步值时不受影响。
- 补测试:启动防抖两个方向,以及版本比较按段取数字的回归
  (字典序会把 0.99.0 判为大于 0.146.0,同时让「取最大值」与
  「只向前推进」两处判错)。
2026-08-03 20:55:37 +08:00
zhiyu 1e08c4c56f test(server): 补齐 settings 契约用例的 Codex 版本号字段
新增的三个设置键会出现在 GET /api/v1/admin/settings 响应里,
契约用例的期望 JSON 需同步,否则 -tags=unit 下断言失败。
2026-08-03 20:26:50 +08:00
zhiyu 2eb24814fe fix(codex): 强制统一出站身份并让客户端版本号跟随官方发布
上游 /backend-api/codex 在容量紧张时按客户端身份分优先级降载,被降载的请求
HTTP 200 后立刻推流内 server_is_overloaded。此前网关对配不出官方身份的客户端
整体回退到硬编码的 codex_cli_rs/0.144.1(落后官方 4 个发布),这些请求稳定
落在被优先丢弃的一侧。

- 强制统一出口:所有 OAuth 出站的 User-Agent / originator / version 一律改写
  为网关规范身份,客户端自报身份不参与构造;HTTP / 透传 / WS / alpha-search /
  探针全覆盖。compat 桥接故意删除 originator 的路径保持 no-op。
- 版本号收敛为单一来源,运行时优先级为面板覆写 → 自动同步值 → 内置常量;
  UA 与 version 头同源派生,不再各自硬编码。
- 新增 3 小时自动同步官方客户端最新稳定版,面板可关闭,无需为跟版本而发版。
- 流内 server_is_overloaded / slow_down 改为先在同账号有界重试再切号,并标记为
  请求级瞬时故障,不再据此临时封禁账号。
- 移除被取代的降载身份黑名单、浏览器 UA 兜底及其辅助函数。
2026-08-03 20:14:58 +08:00
li e2652eb853 fix(billing): quantize usage billing amounts to the NUMERIC(20,8) scale
同一笔 ActualCost 会被分别写入两条方向相反的 SQL:

    balance    = balance - $1      -- 存剩余额度,舍入的是"减法结果"
    quota_used = quota_used + $1   -- 存累计用量,舍入的是"加法结果"

两列都是 NUMERIC(20,8),PostgreSQL 按 half-away-from-zero 舍入运算结果。
金额在第 9 位落到 half 边界时,两侧朝相反方向舍入:

    10 输入 token × 0.00000125 + 5 输出 token × 0.00001000 = 0.0000625
    × 1.25(分组倍率) = 0.000078125

    balance:    10000 - 0.000078125 = 9999.999921875 → 9999.99992188  delta 0.00007812
    quota_used:     0 + 0.000078125 =     0.000078125 →     0.00007813  delta 0.00007813

余额少扣、API Key 配额多记,单次相差 1e-8 且随请求量线性累积(1000 次 ≈ 1e-5 USD),
余额、Key 配额与用量记录无法精确对账,只能靠 epsilon 比较勉强吻合。

修复:在参数进入 SQL 之前,把命令中的全部金额统一量化到 8 位小数
(half-away-from-zero,与 PostgreSQL NUMERIC 一致)。两条语句拿到的是
同一个已落在 8 位刻度上的金额,存储阶段不再发生舍入,delta 精确相等。

- 走 shopspring/decimal 而非 math.Round(v*1e8)/1e8:后者在乘除中引入额外
  二进制误差,边界值可能被推到错误的一侧
- 量化排在指纹计算之后:指纹是请求幂等键,保持由原始金额派生,
  避免升级前后同一 request_id 的重试算出不同指纹而被判为 fingerprint conflict

Fixes #5229
2026-08-03 20:00:32 +08:00
spongehah 1b04e03cc4 fix(prompt-audit): parse responses output text 2026-08-03 17:38:04 +08:00
wucm667 2be047d1cd [verified] test(subscription): update unit repository stubs 2026-08-03 17:22:18 +08:00
wucm667 db725a775a [verified] fix(subscription): serialize concurrent renewals 2026-08-03 16:59:05 +08:00
Wesley Liddick 825ca7b1fc Merge pull request #5183 from rick147/codex/feat-openai-reset-credit-cache
feat(openai): refresh reset credit state after quota reset
2026-08-03 16:01:10 +08:00
shaw 54a2bcfd15 fix(openai): harden reset-credit refresh and account recovery
Review follow-ups on the reset-credit caching flow:

- Recover account state BEFORE (and independently of) the reset-credit
  display cache. A failed cache refresh could previously abort the run and
  leave the account rate-limited — the very reason the credit was spent
  (#3672 / #3740). The recovered account row is now returned even when the
  cache refresh fails.
- Run the post-reset bookkeeping on a detached, time-boxed context and give
  the panel reset call a larger timeout. A client abort no longer strands a
  consumed (non-refundable) credit with an unrecovered account, and the
  chained upstream calls can no longer exceed the client timeout and invite a
  retry that spends a second credit.
- Persist the reset-credit snapshot through POST /accounts/:id/quota/refresh
  instead of a side-effecting GET flag, so the write is covered by the audit
  middleware. A rejected snapshot write now degrades to cache_persisted=false
  instead of turning a successful upstream read into a 502 that left the card
  without a credit count and the reset button permanently disabled.
- Reject snapshots whose positive count carries no expiration entries, and
  drop expired credits (clamping the count) when rehydrating, so a stale
  cache can no longer light up the reset button.
- Keep nil quota / rate-limit services nil in the handler's interface fields;
  storing a nil *Service made the "not enabled" guards non-nil.
- Time-box the usage-refresh suppression and reuse handleAccountUpdated so the
  patched row also enters the auto-refresh silent window.
2026-08-03 14:40:55 +08:00
Wesley Liddick 27e8f69a9e Merge pull request #5171 from heathermhuang/codex/composite-reasoning-policy
feat(composite): enforce reasoning effort policy
2026-08-03 11:28:38 +08:00
Wesley Liddick 684ab20a0b Merge pull request #5164 from wucm667/fix/issue-5099-messages-temp-failover
fix(openai): fail over Messages temporary account errors
2026-08-03 11:28:16 +08:00
Wesley Liddick 0173830df6 Merge pull request #5193 from wucm667/fix/issue-5187-stripe-refund-idempotency
fix(payment): make Stripe refunds idempotent
2026-08-03 11:28:05 +08:00
Wesley Liddick 724565e4aa Merge pull request #5199 from wucm667/fix/issue-5191-refund-balance-force
fix(payment): require force for insufficient refund balance
2026-08-03 11:27:52 +08:00
Wesley Liddick a1f1a0cc6b Merge pull request #5194 from wucm667/fix/issue-5189-persist-unsettled-usage
fix(billing): retain usage logs on billing failure
2026-08-03 10:27:30 +08:00
wucm667 3c20f9a666 [verified] fix(payment): require force for insufficient refund balance 2026-08-02 23:11:17 +08:00
shaw e1b76e2245 fix(codex): normalize load-shed originators to avoid upstream capacity shedding
上游 /backend-api/codex 按 Originator 头分桶调度容量:落在降载桶的请求即使返回
HTTP 200,也会立刻推 SSE `event: error`(code=server_is_overloaded)并以
response.failed 收尾。2026-07-29 起 codex-tui 落入降载桶,codex_cli_rs 正常——
判定因子是 originator 而非 User-Agent(codex_cli_rs 配 curl UA 亦可正常返回)。

网关会把该错误判定为瞬时上游故障并冷却账号,对外表现为 Codex 账号频繁过载不可用:
server_is_overloaded → isOpenAITransientProcessingError →
shouldCooldownOpenAITransientUpstreamError → 账号冷却 → 客户端 503。

本项目有三处降载身份来源:浏览器 UA 兜底的默认 UA、客户端透传的真实 TUI 身份、
以及指纹缓存注入探针的 UA。

修复收口在 enforceCodexIdentityHeaders——HTTP / 透传 / WS 握手 / compat 桥接 /
探针 / PAT / 模型列表 / alpha-search 八条出站路径共用的唯一纯函数收口点:

- 新增 NormalizeCodexClientIdentityToCLI,把降载桶身份改写为 codex_cli_rs,
  只替换身份段并裁掉尾部 (name; version) 客户端标识组,保留版本 / OS / 架构 /
  终端指纹;改写后 originator 与 UA 首段仍然配套,不破坏 #3901 的配对不变式,
  且改写幂等。
- DefaultOpenAICodexUserAgent 从 TUI 身份改为 CLI 身份(浏览器兜底路径上最大的
  降载身份来源)。
- 管理端 Codex UA 的 placeholder / hint 原本在把管理员往降载桶引导,一并修正。

新增 gateway.disable_codex_originator_normalization(默认 false,即归一化开启),
供上游调整分桶后回滚。该开关经 NewOpenAIGatewayService 发布为进程级快照,故必须
保持反义命名:正向命名的 Go 零值 false 会让未经 viper 加载而手工构造的 Config
静默关掉全局保护,viper.SetDefault 救不了这条路径。已加用例钉住该属性。

降载桶集合是上游容量策略快照而非协议常量,上游调整分桶后需同步修订。
2026-08-02 23:00:12 +08:00
rick147 f970bd48c9 fix: stop scheduler work after request cancellation 2026-08-02 21:32:31 +08:00
rick147 a0802f00b6 feat: cache OpenAI reset credit details 2026-08-02 21:32:31 +08:00
wucm667 0b9f40e230 fix(billing): retain usage logs on billing failure 2026-08-02 20:50:13 +08:00
wucm667 0b26acac07 fix(payment): make Stripe refunds idempotent 2026-08-02 20:27:44 +08:00
github-actions[bot] 7e2e9ba050 chore: sync VERSION to 0.1.170 [skip ci] 2026-08-02 10:46:21 +00:00
Wesley Liddick c043c24774 Merge pull request #4925 from Brisbanehuang/feat/group-profit-control
feat(scheduler): 分组级利润控制——按账号倍率过滤 token 调度候选
2026-08-02 18:19:13 +08:00
Wesley Liddick 11c1e944b9 Merge pull request #4911 from Brisbanehuang/feat/upstream-billing-rate-writeback
feat(billing-probe): 探测成功后可选将上游声明倍率自动同步为账号倍率
2026-08-02 18:18:50 +08:00
Wesley Liddick d99ee72911 Merge pull request #4896 from Brisbanehuang/feat/upstream-billing-probe-multi-platform
feat(billing-probe): 上游计费倍率探测放宽到全部 API-key 平台账号
2026-08-02 18:18:37 +08:00
r266-tech 30d2589ef0 fix(openai-ws): preserve terminal event on lease loss 2026-08-02 11:47:59 +08:00
shaw dec47e8fae fix(profit-control): stop leaking profit policy, close veto livelock, restore passthrough turn pricing
审计修复,逐条如下。

H1 利润策略泄露给所有普通用户
  profit_control_enabled / profit_min_margin / profit_safety_buffer 从
  dto.Group 移到 dto.AdminGroup(后者内嵌前者),赋值相应从
  groupFromServiceBase 移到 GroupFromServiceAdmin;前端 TS 同步从 Group 移到
  AdminGroup。dto.Group 是 GET /api/v1/groups/available 的响应体,该响应本就带
  rate_multiplier,相乘即可反推运营方上游采购成本上限。
  api_contract_test.go 的 /groups/available golden JSON 回滚这三个字段,并把
  fixture 改成非零值(require.JSONEq 是精确比对,缺字段即失败)。
  新增 dto 层边界测试:普通用户 DTO 不含三字段、管理员 DTO 仍含。

M1 利润终检 continue 与 failover 503 退避互动产生活锁
  FailoverState 新增 profitVetoedAccountIDs / profitVetoCount 与
  RecordProfitVeto():加入排除集 + 计数,达 maxProfitVetoAttempts(10) 返回
  FailoverExhausted。HandleSelectionExhausted 的 503 清空分支改为清空后把利润
  否决的账号放回排除集;若排除集已全部由利润否决贡献,清空不会带来任何新候选,
  直接判定耗尽(否则 SwitchCount 永不前进、退避条件永远成立,每 2s 空转一轮)。
  五个 handler 否决点(gateway_handler ×2 / responses / chat_completions /
  gemini_v1beta)改为经 RecordProfitVeto 决策,耗尽时按无可用账号终止。
  回归测试钉死:503 之后持续利润否决必须有限步终止且不 spin;未启用利润控制的
  请求退避语义完全不变。

M2 排队等槽后才终检,延迟可放大到 N × WaitPlan.Timeout
  OpenAI 侧选号循环(自有 failedAccountIDs map,非 FailoverState)新增
  recordOpenAIProfitVeto + handleOpenAIProfitVetoExhausted,共用同一上限语义。
  覆盖 responses / messages-dispatch / chat_completions / alpha_search /
  embeddings / images / grok_media 七处,以及 WS 两处否决分支。

M4 ws_v2 透传 ingress 绕过 per-turn 重定价(选方案 B:最小止血)
  透传 relay 只回调 AfterTurn、没有任何 turn 起始回调,hooks.BeforeTurn 永远
  不触发,而 handler 把 turnPricingAt 初始化成建连时刻 ⇒ 透传连接全部 turn 按
  建连时刻的高峰因子结算,客户端峰前建连保活即可全程谷价——正是本 PR 想堵的
  漏洞。改为 openAIWSTurnPricing 零值起步、只由 BeforeTurn 冻结;透传路径保持
  零值,RecordUsage 回退记录时刻,与引入利润控制前的基线一致。
  未选方案 A(给透传补 turn 起始回调):passthrough_relay.go 是 #5167 刚修过的
  取消传播/close frame 时序敏感区;且 BeforeTurn 还承担 turn>1 的并发槽位抢占,
  接进去等于给透传连接引入 per-turn 抢槽,风险远超本次修复范围。透传仍有建连时
  的准入门,只是没有 turn 级复核,已在两处注释写明。
  测试:service 层钉死透传 ingress 不触发 BeforeTurn(含失败时的复核指引),
  handler 层钉死零值语义与逐 turn 覆盖。

M5 装门读分组走了带账号计数聚合的 GetByID
  SchedulerSnapshotService 新增 GetGroupByIDLite,openai/gateway 两处装门改用
  之。门只需要平台/倍率/利润/高峰字段,且该查询发生在「是否启用利润控制」判定
  之前,未启用的分组同样付代价。两个测试 stub 的 GetByID 改成 panic 守卫。

M6 认证快照注释与真实读取路径相反
  门解析优先取 ctxkey.Group,而它就是本快照物化出来的对象,直连流量走的正是这
  条路。改正注释,与 api_key_repo.go 投影处的说明对齐,避免后人照旧注释删列。

M3 rate_multiplier 为 nil 时利润门 fail-closed(不改行为,加护栏)
  保留 fail-closed。补 repository 层测试钉死账号调度快照的 full/metadata 两份
  payload 都必须保留 RateMultiplier(含 0 值),漏列在 CI 就红。

L1 迁移号注释 191 / 191-192 改为实际的 192/193。
L2 admin group Create 的利润配置预校验改用与 CreateGroup 一致的归一化平台
   (新增 service.NormalizeGroupPlatform,两边共用)。保留预校验而非删除:
   service 层返回的是无类型 error,经 ErrorFrom 会变成 500,删掉会把合法的
   400 降级成 500。
L3 前端利润校验的上界改为判定换算后的小数(后端按小数校验 [0,1)),
   99.999% 会四舍五入进位成 1.0 而被后端 400;i18n en/zh 同步改为 0-99.99。
L4 clampProfitControlThreshold / profitControlOverThreshold 抽为共用函数,
   线上装门/否决点与 profit-preview 不再各自实现,附边界语义测试。
L5 profit-preview 补「默认 D 有账号但最低有效 D 归零」的告警(两档都为 0 由
   既有告警覆盖,不重复)。
2026-08-01 22:39:33 +08:00
Brisbanehuang fad2f215e8 fix(profit-control): decouple gate from image intent, propagate gate via selection, restore eager sticky fallback
Review fixes for the profit-control feature commit:

- Image intent no longer disables the profit gate. The shared /v1/responses
  handler previously skipped the pricing context (and therefore the gate)
  whenever the platform-wide image intent predicate matched, which includes
  Codex's passive image_gen namespace declaration: any client could disable
  admission control for anthropic/gemini/antigravity groups by declaring a
  namespace tool in the request body. Both /v1/responses paths now always
  install the token pricing context; image intent only drives capability
  routing and image billing. Mixed token+image requests stay token-gated;
  only dedicated media endpoints remain out of scope.
- Out-of-scope paths are now explicitly suppressed instead of implicitly
  ungated: Grok media (billed by media multipliers; also prevents in-flight
  video lookups from turning into spurious 404s), OpenAI-group count_tokens
  (unbilled), and Live calls (duration billed) carry a suppress marker that
  every install point honors, including the defensive scheduler-entry
  install.
- The gate resolved during selection now travels back to handlers on the
  AccountSelectionResult. The shared gateway installed the gate only on a
  scheduler-local context, so handler-side post-slot terminal rechecks and
  post-admission sticky binding were no-ops for anthropic/gemini/
  antigravity/shared-grok requests (and for composite-routed member groups
  on the OpenAI path). Handlers re-apply the carried gate via
  ContextWithSelectionProfitGate before the terminal recheck and binding;
  the WS acquired-selection branch gained the previously missing recheck.
- Sticky binding semantics restored for ungated traffic:
  BindStickySessionAfterProfitAdmission falls back to the official eager
  bind when no gate is installed (wait paths lost their only binding point
  otherwise), reads the pre-existing binding at bind time only when gated
  (removes the unconditional per-request Redis read the feature added to
  the shared handlers), and the legacy engine's three selection-time
  binding writes are skipped under a gate so a terminally vetoed account
  can no longer become the new sticky target.
- Responses WS connections re-freeze pricingAt and re-resolve the gate at
  every turn (BeforeTurn) and bill each turn with its own instant, closing
  the connect-at-valley/bill-at-valley window; a turn that fails the
  recheck closes the connection so the client reselects on reconnect.
- Terminal recheck no longer swaps a DB-fresh selected account for a
  staler snapshot object (UpdatedAt guard), and observer counters are
  documented as per-evaluation.
- Migrations renumbered 191/192 -> 192/193 after the passkey migration
  landed upstream as 191.

New regressions: selection-carried gate propagation (control group proves
the pre-fix no-op), image intent not disabling the shared gate, eager
binding fallback without a gate (both services), gated
read-failure/sentinel-miss binding semantics, legacy-engine deferred
binding under a gate with eager behavior preserved ungated, turn-level
pricing refresh (config re-resolution, scheduled-group precedence,
suppress, mid-connection disable), and suppress-marker coverage for the
request pricing context.
2026-08-01 22:39:32 +08:00
Brisbanehuang 20ad5ec506 feat(scheduler): per-group profit control for token account admission
Group pricing (rate multiplier, peak windows, per-user overrides) and
account cost (accounts.rate_multiplier) already live side by side, but
nothing stops the scheduler from handing a request to an account whose
cost multiplier exceeds what the group's pricing can profitably serve.
Add an opt-in per-group profit gate that filters scheduling candidates
by a margin rule, while ordering, scoring, stickiness and breakers keep
working unchanged among qualified accounts.

Admission rule: an account qualifies iff U <= D * (1 - min_margin -
safety_buffer) within a small relative epsilon, where U is
accounts.rate_multiplier (0 is legal; missing/negative/NaN/Inf are
conservatively rejected as invalid) and D is the requester's effective
downstream multiplier (user-group override ?? group default, times the
group peak factor) frozen at the request's pricing instant.

- groups gain profit_control_enabled / profit_min_margin /
  profit_safety_buffer (migration 191); the durable auth-cache
  invalidation trigger additionally watches the profit and pricing
  columns (migration 192) so out-of-band group edits cannot leave
  stale auth snapshots; GetByKeyForAuth explicitly projects the new
  columns and the API-key auth snapshot version is bumped to force a
  refresh of pre-existing snapshots
- request-level pricing instant: token entry points install pricingAt
  into ctx; the profit threshold D and the RecordUsage peak factor
  read the same instant, so one request never changes price mid-flight
  across waits/retries/failover (media and unwired paths keep the
  existing record-time semantics)
- the gate covers token requests on openai, anthropic, gemini, grok
  and antigravity groups: OpenAI-family handlers via
  WithOpenAIRequestPricingContext (responses incl. WS bridge, chat
  completions, messages, embeddings, alpha search), the shared gateway
  via WithGatewayTokenRequestPricing (messages, chat completions,
  responses, gemini model actions); composite groups cannot enable it
  directly; image/video/models/usage/count_tokens stay ungated and an
  explicit image-generation intent suppresses the gate end to end
- post-slot recheck: after a slot is acquired the account is re-read
  via SchedulerSnapshotService.GetAccount (scheduler cache, then DB;
  only when both fail the check fails open with WARN + metric); a
  vetoed account releases its slot and joins the request's exclusion
  set for reselection; sticky bindings are written only after the
  final check passes, and an over-threshold sticky account is skipped,
  not deleted, so it comes back once its rate recovers
- sticky-session cache contract: GatewayCache.GetSessionAccountID now
  returns ErrStickySessionNotFound on a miss (mapped from redis.Nil in
  the repository implementation, mirroring ErrRefreshTokenNotFound) so
  the profit sticky path can distinguish "no binding yet" from a real
  read failure without importing the cache driver in service code
- cross-group re-entry (composite parent -> member group) resolves the
  gate against the member group and clears a stale parent gate instead
  of letting a foreign threshold veto accounts
- per-platform/group activity counters (installs, threshold vetoes,
  invalid-rate vetoes, refresh failures) for observability
- admin UI: profit-control section on the five platforms' group forms
  with percent input, validation and platform-switch reset; group
  create/update/duplicate normalize and validate the config at a
  single choke point
- cmd/profit-preview: offline what-if tool that replays the production
  admission semantics over an exported config/account/override/model
  dump, reports per-model admitted-account counts under the default
  and the worst-case (lowest user override) D, and surfaces probe-sync
  staleness as warnings without affecting admission

Tests: service unit coverage for gate resolution/veto/threshold
epsilon/pricing instant/suppress marker/scheduler filtering and
post-slot recheck (incl. -race on the profit surface), unit-tagged
handler slot-recheck and capability-mapping regressions, sqlmock and
real-PostgreSQL integration regressions for the GetByKeyForAuth
projection and the migration-192 trigger watch list, API contract
update, and frontend specs for the five-platform form helpers.
2026-08-01 22:39:31 +08:00