Commit Graph
5718 Commits
Author SHA1 Message Date
Wesley Liddick 61acf29125 Merge pull request #5501 from wucm667/fix/issue-5500-unset-scheduling-thresholds
fix(settings): cache unset scheduling thresholds
2026-08-11 09:27:22 +08:00
wucm667 3e1674a060 fix(settings): cache unset scheduling thresholds 2026-08-11 04:34:49 +08:00
Wesley Liddick 0b3fe95afd Merge pull request #5439 from pigzwy/feat/response-model-billing
feat(billing): support safe billing by upstream response model
2026-08-10 19:47:22 +08:00
anya e5b325e481 fix(billing): harden response-model billing admission
Three guards on the response_model billing basis, all scoped to the opt-in
channel mode so existing channels are unaffected.

1. Per-unit billing gate was stale. Audio (AudioUsage) and the search
   surcharge (SearchCount) reached the billing paths after this branch was
   cut; both are priced per unit rather than per token, so they must be
   excluded like image/video/web-search already are. Audio pricing ignores
   the model entirely, so the previous code "adopted" a basis switch that
   changed nothing and emitted a misleading audit log for it.

2. Never zero out a billable request. A catalog entry whose token prices are
   explicitly 0 still passes the identified-pricing gate (TokenPricingAbsent
   only means both prices are missing), so an upstream could declare a free
   model name and drop the bill to zero. Reject a zero (or negative)
   recomputation whenever the baseline was billable; an already-zero baseline
   is unaffected.

3. Never cross from channel pricing to the global table. Channel pricing
   matches exact keys and prefix wildcards and does not strip date suffixes,
   while the global table's identified lookup does. Upstreams routinely
   declare dated model IDs (claude-opus-4-5-20251101), so allowing a
   cross-source comparison would silently bypass an administrator's channel
   markup on essentially every request. Admins who want a downgrade target
   discounted can price it explicitly on the channel.

Also skip the recomputation entirely when the declared model equals the
baseline: it is provably the same cost and only burned a pricing resolve.

The identified-pricing helpers now return whether the model resolved to
channel pricing so the third guard costs no extra resolve.
2026-08-10 19:11:47 +08:00
shaw 33351c7bc7 fix(billing): gofmt channel.go and drop the redundant response-model hint
- channel.go: 常量块里插入注释后 gofmt 会把 BillingModelSourceResponse 单独成组,
  原写法沿用了上一组的对齐空格,CI 的 gofmt 检查因此失败
  (internal/service/channel.go:45: File is not properly formatted)。
- 撤掉渠道表单里新加的那条提示:说明本就多余,且"只降不升"只在"相对基线收费"
  这个口径下成立,容易被读成"上游返回更贵的模型也不会多收",反而误导。
2026-08-10 18:45:16 +08:00
shaw b689e5b401 fix(billing): harden response-model billing and repair its test fixtures
按上游响应模型计费的准入过宽、且自带用例必然失败,本次一并修复。

严格化准入
- 新增 PricingService.GetIdentifiedModelPricing / BillingService.HasIdentifiedTokenPricing:
  只接受价格表中能被确定性识别的条目(精确名、已知拼写变体、去掉日期版本后缀),
  不再接受 getFallbackPricing / matchByModelFamily 按子串猜出的系列兜底价。
  此前上游只要自报一个含 "haiku" 的编造名字就能被判定"已定价",把账单压到最便宜的
  系列价(实测 claude-opus-4.8 基线 $0.0019250 → $0.0000963,20 倍少收)。
  GetModelPricing 的对外行为不变,仅把前三步查找抽成共用函数。
- 图片 / 视频 / 网页搜索请求不再走响应模型覆盖:这些路径按张、按秒、按次定价,
  与准入检查所验的 token 价不是同一套价格表。
- 准入判断抽成 responseModelBillingDeclaration,两条计费主干共用同一套规则。

正确性与可观测性
- 去掉 recordUsageCore 中 billingModel 的无效赋值(ineffassign 已启用,会让
  golangci-lint 直接失败),改由日志表达实际生效的计费基准。
- 补 cost != nil 守卫,与 OpenAI 侧及本文件既有写法对齐。
- 每次实际生效的基准切换记一条 billing.response_model_applied,少收可审计。
- 修正 upstreamResponseModelObserver 上"冲突仅用于诊断、永不影响计费"的过期注释。

测试
- gpt-5.1 与 gpt-5.5 实际共用同一条 gpt-5.4 价格,夹具"价格必须不同"的前置断言
  必然失败,OpenAI 侧 3 个用例(含 4 个子用例)从未跑通;改用 gpt-5.4-nano /
  gpt-5.5。Anthropic 侧 claude-opus-4 不是价格表精确条目,改用 claude-opus-4.8。
- 夹具增加"必须可被确定性识别"的前置断言,避免用例被更靠前的门挡掉而失去判别力。
- 新增:准入规则表驱动用例、可识别性判定用例、编造家族名在两条主干上均被拒的用例。

前端
- 选择该模式时提示"计费基准以上游自报模型为准,只降不升,仅对可信上游启用"。
2026-08-10 18:45:15 +08:00
pigzwy 9096492b55 feat(billing): support safe upstream response model billing 2026-08-10 18:45:14 +08:00
Wesley Liddick 10a4c6e3ad Merge pull request #5234 from wucm667/fix/issue-5230-deduplicate-latest-turn-audit
fix(security-audit): deduplicate websocket turn audits
2026-08-10 10:53:17 +08:00
Wesley Liddick f3c7a1a8c4 Merge pull request #5295 from wucm667/fix/issue-5289-streaming-upstream-error
fix: emit response.failed when compact keepalive commits headers but no SSE payload
2026-08-10 10:52:49 +08:00
Wesley Liddick f19e3816d0 Merge pull request #5316 from wucm667/fix/issue-5313-legacy-scheduler-diagnostics
fix(scheduler): diagnose legacy OpenAI exclusions
2026-08-10 10:52:14 +08:00
Wesley Liddick 30d0405388 Merge pull request #5464 from wucm667/fix/issue-5455-api-key-input-validation
fix(api-key): validate quota and expiry inputs
2026-08-10 10:52:04 +08:00
Wesley Liddick 5deeb7ef15 Merge pull request #5376 from wucm667/fix/issue-5369-composite-image-permission
fix(admin): enable image generation permission for Composite groups
2026-08-10 10:51:54 +08:00
Wesley Liddick 895e8247af Merge pull request #5472 from wucm667/fix/issue-5468-openai-threshold-percent
fix(openai): preserve Codex usage percentages for scheduling thresholds
2026-08-10 10:51:43 +08:00
Wesley Liddick 2d0976ac2c Merge pull request #5475 from fengshao1227/fix/oauth-personal-subscription-expiry
fix(openai-oauth): 个人订阅到期时间不再被 POID workspace 的 entitlement 覆盖
2026-08-10 10:23:05 +08:00
Wesley Liddick 4667f9012f Merge pull request #5477 from Wei-Shaw/revert-5394-fix/issue-5388-risk-control-fail-closed
Revert "fix(risk-control): 风控后端异常时阻断提示词请求"
2026-08-10 10:00:02 +08:00
Wesley Liddick af6928a268 Revert "fix(risk-control): 风控后端异常时阻断提示词请求" 2026-08-10 09:59:42 +08:00
li 358e4a89a1 fix(openai-oauth): 个人订阅到期时间不再被 POID workspace 的 entitlement 覆盖
enrichTokenInfo 里 plan_type 与 subscription_expires_at 的取值口径不一致:

- plan_type 有 shouldApplyChatGPTAccountInfoPlanType 护着(#3641),
  id_token 里的个人套餐优先,accounts/check 只在个人值为空时补位;
- subscription_expires_at 却是无条件覆盖。

accounts/check 是多账号/工作区端点,命中的记录由 access_token JWT 的 poid 决定。
当 poid 指向默认 Personal workspace、而 chatgpt_account_id 指向个人 ChatGPT 账号时
(两者是不同标识),拿到的 entitlement.expires_at 描述的是 workspace 权益。
配上被保护下来的个人 plan_type,账号页就显示成「个人 Pro + workspace 到期时间」。

代码里本来已经有正确的来源——用个人 chatgpt_account_id 查
/backend-api/subscriptions 的 active_until——但它被
`if TrimSpace(SubscriptionExpiresAt) == ""` 挡着,workspace 值一旦写进去就永远不触发。

改为让两个字段始终描述同一份订阅:

- 套餐取自 accounts/check 时(id_token 没带 chatgpt_plan_type),到期时间跟着取
  同一条记录,行为不变;
- 套餐保留了 id_token 的个人值时,只有该记录确实属于个人账号才用它的
  entitlement.expires_at;不属于就跳过,并强制回落到个人订阅端点。

为此给 ChatGPTAccountInfo 补 AccountID:优先读 account.account_id,缺失时退回
accounts 的 map key(key 可能是 "default" 这类别名)。两侧任一缺 ID 时判定为
无法区分并沿用旧行为,poid == chatgpt_account_id 的单人账号完全不受影响。

Fixes #5459
2026-08-10 09:29:14 +08:00
Wesley Liddick b5e83156d9 Merge pull request #5394 from wucm667/fix/issue-5388-risk-control-fail-closed
fix(risk-control): 风控后端异常时阻断提示词请求
2026-08-10 08:51:46 +08:00
Wesley Liddick fddc806db2 Merge pull request #5458 from lyen1688/feat/backup-large-file-parts
完善大文件备份分卷上传与恢复
2026-08-10 08:36:08 +08:00
wucm667 99b31067f7 fix(openai): preserve Codex usage percentages in scheduling threshold 2026-08-10 04:29:08 +08:00
wucm667 f5c108c836 fix(api-key): validate quota and expiry inputs 2026-08-09 22:42:51 +08:00
lyen1688 bbc8b6e906 完善大文件备份分卷上传与恢复 2026-08-09 20:58:07 +08:00
github-actions[bot] 48eb3766d2 chore: sync VERSION to 0.1.173 [skip ci] 2026-08-09 08:26:22 +00:00
Wesley Liddick 29009f0b2e Merge pull request #5446 from Wei-Shaw/feat/email-domain-quota-switch
feat: 邮箱域名限量注册增加独立开关(默认关闭)
v0.1.173
2026-08-09 16:09:17 +08:00
shaw 563a72ca73 feat: add default-off switch for email domain registration quota
PR #5423 relaxed the email suffix whitelist: once a whitelist is
configured, non-whitelisted registrable domains are each allowed to
register one account. That behavior activated unconditionally.

Add registration_email_domain_quota_enabled (default false) to gate it:

- Off (default): restore pre-#5423 strict whitelist semantics — with a
  non-empty whitelist, non-whitelisted domains are rejected with
  EMAIL_SUFFIX_NOT_ALLOWED; the register/verify views restore the
  client-side whitelist pre-check and allowed-domain hint.
- On: keep #5423 behavior — one account per non-whitelisted registrable
  domain (EMAIL_DOMAIN_REGISTRATION_LIMIT).
- Empty whitelist keeps allowing all domains in both states.

Gating lives in validateRegistrationEmailQuota and (as a race-safety
backstop) createUserWithRegistrationEmailGuard; the repository-level
domain lock + in-tx recheck is unchanged. The admin update field is
*bool (omitted = keep current) so stale full-payload saves cannot
silently flip the switch. Email binding and OAuth auto-signup keep
their strict policy, and pending-OAuth bind-login for existing
accounts is unaffected because the handler resolves existing emails
before the quota check.

Frontend adds the toggle to admin settings (zh/en copy; whitelist hint
restored to strict wording, quota wording moved to the new toggle) and
exposes the flag via public settings + SSR injection payload.

Tests: #5423 quota tests now enable the switch explicitly; new
default-off regression tests cover register/send-code/async/pending
OAuth/OIDC create-account plus both register views; API contract JSON
and the injection drift guard are updated.
2026-08-09 15:53:40 +08:00
Wesley Liddick f2da30bcd9 Merge pull request #5423 from lyen1688/feat/email-domain-registration-quota
完善邮箱域名注册额度策略
2026-08-09 15:16:26 +08:00
Wesley Liddick 7821c4005e Merge pull request #5424 from fengshao1227/fix/gemini-native-image-billing
fix(gemini): 原生生图按上游实际回吐的图片张数计费,修复自定义模型名下生图记 $0
2026-08-09 15:02:08 +08:00
Wesley Liddick c5bda8b8e4 Merge pull request #5416 from fengshao1227/fix/images-apikey-detach-upstream-context
fix(openai): 非流式生图脱钩上游 context,客户端断开不再导致图已出却不扣费
2026-08-09 14:55:28 +08:00
Wesley Liddick a1073843ac Merge pull request #5437 from Brisbanehuang/codex/upstream-response-model-audit-followup
perf(usage): 优化上游响应模型观察热路径
2026-08-09 14:55:15 +08:00
Wesley Liddick 7a113fb7a3 Merge pull request #5352 from feeeei/main
fix(gemini): 修复Gemini池模式时,依然被 429 response 触发账户限流问题
2026-08-09 14:55:02 +08:00
Wesley Liddick 909cbb9b12 Merge pull request #5356 from IanShaw027/feat/channel-monitor-v2-ops-ui
feat(channel-monitor-v2): 新增被动渠道监控 V2
2026-08-09 12:25:37 +08:00
shaw d92edc01be Merge origin/main into feat/channel-monitor-v2-ops-ui
Resolves three conflicts, all of the "both branches appended to the same
block" shape. Every one is resolved as a union of both sides; nothing from
either parent is dropped.

- handler/admin/setting_handler_update.go: keep ChannelMonitorHideThroughput
  (V2) alongside GrokDefaultTextModel / GrokCrossClientModelMapEnabled /
  GrokDefaultBaseURLMode (#5408). UpdateSettings writes every key on each
  save, so dropping either side would reset those settings to zero values.
- service/domain_constants.go: keep SettingKeyChannelMonitorHideThroughput
  and the three SettingKeyGrok* constants.
- repository/migrations_runner.go: keep the 195 checksum rule (V2) and the
  218/219/220 rules (#5408).
2026-08-09 12:11:35 +08:00
Wesley Liddick fb0475656c Merge pull request #5408 from IanShaw027/feat/grok-complete-integration
feat(grok): 完善 Grok 平台集成 — 授权、模型映射、媒体/Voice/搜索计费与调度门禁
2026-08-09 11:31:09 +08:00
Brisbanehuang 6e34fb09c9 perf(usage): optimize upstream response model observation 2026-08-08 22:04:49 -04:00
IanShaw027 5315896b30 fix(grok): align free soft-gate tests with async fail-open cache
Treat cacheTTL=0 as non-expiring known entries, always store negative
refresh markers, and update sticky getSchedulableAccount tests to expect
first-hit fail-open then block after background stats warm.
2026-08-08 23:38:50 +08:00
IanShaw027 a54a4b674b fix(grok): clear golangci errcheck and gofmt on free-quota path
Check the sync.Map type assertion in free-quota refresh coalescing, and
gofmt migration checksum rules plus prompt-audit route map alignment.
2026-08-08 22:52:31 +08:00
li b6eb6c1efa fix(gemini): 原生生图按上游实际回吐的图片张数计费
/v1beta/models/{model}:generateContent 与 Anthropic→Gemini 兼容路径的
ImageCount 只由 isImageGenerationModel(originalModel) 决定,而该白名单是
按 Google 官方模型名精确/前缀匹配写死的(antigravity_image_test.go 里
明确断言 my-gemini-3-pro-image-test 这类自定义名返回 false)。

GeminiMessagesCompatService 服务的却主要是 API Key + 自定义模型映射的账号:
客户端请求名和 GetMappedModel 后的上游名都可能是站长自取的别名,白名单必然
判不出来 → ImageCount=0 → calculateRecordUsageCost 里 `if result.ImageCount > 0`
的按次计费分支整条不触发 → 生图请求全部记 $0(issue #5358)。

改为优先按上游响应里真实的 inlineData 图片 part 计数:
- 新增请求级计数器,挂在 gin.Context 上,与既有的
  upstreamResponseModelObserver 同一批调用点取解包后的响应体;
- 取「单个 payload 内的最大值」而非累加:Gemini 兼容上游的 SSE 分片可能是
  累积式的(computeGeminiTextDelta 正是为此存在),逐 chunk 累加会把同一张图
  重复计费。max 保证累积式流与非流式都得到真实张数,增量式多图流最差退化到 1,
  与改动前同值,不构成回退;
- 每次 Forward 开头重置计数器,避免 failover 复用同一个 gin.Context 时
  把失败账号已回吐的图叠加到成功账号账单上;
- 响应里数不出图时(fileData 引用式回图等)退回原有模型名启发式,并额外认
  映射后的上游模型名,与 shouldSkipCodexPlanGatedImageModelCooldown 同时取
  requestedModel / modelKey 的口径一致。

inlineData / inline_data 两种字段风格都认(官方 SDK 与部分中转回 snake_case),
只统计带 base64 数据且 MIME 为图片的 part。

Fixes #5358
2026-08-08 21:36:18 +08:00
lyen1688 4999231d61 修复邮箱域名注册额度策略 2026-08-08 21:05:20 +08:00
feeeei cbc2a3dd46 修复池模式 Gemini 账号被 429 打上账号级限流
- 429 的标记点在重试循环内,先于 CheckErrorPolicy 执行,池模式豁免只能落在
  handleGeminiUpstreamError 自身;否则一次上游 429 会把账号锁到 PST 午夜,
  即便重试已经成功返回客户端
- 判定条件与 HandleUpstreamError 对齐:自定义错误码优先级高于池模式;
  401/403/529 仍委派给 RateLimitService,临时不可调度规则不受影响
- chat completions 路径的策略分发改为只有 None / Matched 才处理账号状态,
  与 messages 兼容层的 switch 一致,ErrorPolicySkipped 不再漏进来
2026-08-08 16:51:48 +08:00
li cbf2be05a3 fix(openai): 非流式生图脱钩上游 context,客户端断开不再导致图已出却不扣费
forwardOpenAIImagesAPIKey 走的是 detachStreamUpstreamContext(ctx, parsed.Stream),
该函数在非流式时原样返回请求 context(gateway_usage_billing.go:488)。于是客户端
中途断开会连带取消已经在出图的上游调用:上游那边图已生成并计费,网关这边拿到
context canceled,记 502,result 为 nil,handler 的 result.ImageCount > 0 兜底
无法命中,本次请求不产生扣费。

生图是长耗时(数十秒)且上游侧已产生实际成本的操作,与普通非流式 chat 不同:
提前取消并不能省下上游开销,只会丢掉已付费的产出。

同一端点的 OAuth 分支 forwardOpenAIImagesOAuth(openai_images_responses.go:1690)
以及同属媒体生成的 grok_media.go 本来就用无条件脱钩的 detachUpstreamContext,
OpenAI 网关侧 13 个转发点里只有这一处是例外。这里对齐。

Fixes #5411
2026-08-08 15:24:34 +08:00
IanShaw027 e91b494168 fix(test): stabilize OpenAI streaming preamble keepalive assertion
Keepalive is gated on downstream idle, not upstream tick cadence. The old
fixture wrote progress events every 250ms and could finish without a true
1s idle window on loaded CI runners, so ":\n\n" never appeared. Pause after
preamble long enough for the keepalive ticker before completing the stream.
2026-08-08 15:00:33 +08:00
IanShaw027 04d9eeaf07 fix(channel-monitor-v2): clear CI lint and align trend axis with range zoom
Fix golangci unused/gofmt on the gentle-backfill path. Plot matrix and
line-chart X axes on the selected [requested_start, requested_end)
window (empty slots while backfill lags), and let plain mouse-wheel zoom
narrow the visible interval so pulse blocks grow wider.
2026-08-08 14:46:35 +08:00
IanShaw027 7eb1310701 fix(grok): close free-by-default billing and related review blockers
H1/H2: bill search and voice with code defaults when group prices are
nil (explicit 0 remains free); bump API key auth snapshot to v19 and
refresh incomplete media/search/audio projections.

M1–M6: free-quota soft gate fails open on cache miss with background
refresh and 60s default TTL; correct password_auth config docs; default
cross-client model map to true (→ grok-4.5); audit /tts and /web_search;
exclude composite from migration 220 video-price clears; never let a
search surcharge mask token pricing failures.
2026-08-08 14:39:22 +08:00
IanShaw027 825f9c78d5 fix(channel-monitor-v2): default mode v1 and gentle adaptive backfill
Default channel_monitor_mode to v1 (opt-in V2) so upgrades keep active
probes; existing explicit v2 rows are left alone via ON CONFLICT DO NOTHING
plus migration checksum compatibility for already-applied 195.

V2 first-enable backfill no longer compresses ticks to 5s or uses 24h
chunks. Each tick does recent overlap plus at most one historical chunk
with depth-based ceilings (2h/4h/6h), adaptive grow/shrink, and failure
backoff. Error request_id dedup is bounded by a 90-minute lookback so
ops_error_logs is not scanned for full history.

Also align hide_throughput parse default with privacy-preserving public
runtime (missing key → true).
2026-08-08 13:59:11 +08:00
IanShaw027 cec922d335 fix(grok): clear golangci-lint findings on complete-integration branch
Check Close/CloseNow errors, drop unused helpers and dead constants,
lowercase ST1005 error strings, and stop discarding unwrap status as an
unused assignment so CI golangci-lint passes.
2026-08-08 12:56:40 +08:00
IanShaw027 a4faa015a7 fix(deps): bump nanoid 3.3.16 -> 3.3.17 for GHSA-2v37-7h3g-55p8
Match upstream lockfile pin so frontend-security audit gate passes.
2026-08-08 12:50:53 +08:00
IanShaw027 3c22aeeb3d fix(channel-monitor-v2): default hide_throughput true for privacy parity
Align InitializeDefaultSettings and parseSettings with the public-settings
path and migration 206 so admin and user views agree when the key is unset.
2026-08-08 12:49:39 +08:00
IanShaw027 1f58e25ab3 Merge upstream/main into feat/grok-complete-integration
冲突集中在 chat completions / messages 两条 Responses 转发路径:
upstream 给 OpenAIForwardResult 增加了 UpstreamResponseModel 与
UpstreamResponseModelConflict(配套 beginUpstreamResponseModelObservation
观测器),本分支在同样位置把返回值改成了具名变量以便挂 Grok 原生搜索计数。
两侧不互斥,合并结果同时保留上游的响应模型观测字段与 Grok SearchCount 逻辑。

frontend/pnpm-lock.yaml 取 upstream 版本:package.json 与 upstream 完全一致,
本地差异只是 pnpm install 的重解析噪音。
2026-08-08 11:12:56 +08:00
IanShaw027 07b46e93e4 fix(grok): 修正 voice 路由推导、视频价归一化与门禁缓存驻留
- custom-voices endpoint 改由匹配到的路由模板推导(c.FullPath())。
  原实现按请求 URL 字面后缀判 /audio,voice_id 恰为 "audio" 时
  GET /custom-voices/audio 会被改写成 custom-voices/audio/audio,
  把档案查询变成音频下载。补 voice_id="audio" 的路由推导测试。

- NormalizeVideoModelPrices 不再把无法识别的分辨率静默折算成 480p:
  新增 LookupVideoBillingResolution 报告未知档位,配置解析路径丢弃并告警,
  运行时计费仍走 OrDefault 兜底。model/tier 两层遍历改为排序遍历,
  多个别名收敛到同一 family 时结果不再随 Go map 顺序漂移,冲突单价告警。

- grok free quota 软门禁缓存新增过期淘汰。条目按 account_id 键控且只写不删,
  账号下线后会驻留至进程结束;淘汰只挂在已受 TTL 约束的查询路径上,
  不影响缓存命中热路径。

- 迁移 220 清空非 Grok 分组视频价前先落快照表 groups_video_price_backup_220,
  原 UPDATE 不可回滚。

- 简化 grok_search_count 两处等价冗余的事件类型分支。
2026-08-08 11:02:22 +08:00
Wesley Liddick cc67b1aca1 Merge pull request #5406 from bestony/fix/openai-oauth-routing-hints
fix(openai): forward OAuth routing hints
2026-08-08 10:58:53 +08:00