Commit Graph
1371 Commits
Author SHA1 Message Date
Wesley Liddick 8b3fe664dc Merge pull request #5261 from lyen1688/feat/tencent-captcha-gate
新增腾讯天御验证码认证门禁
2026-08-04 16:39:55 +08:00
Wesley Liddick ae81dfd933 Merge pull request #5177 from r266-tech/fix-nonpassthrough-write-context
fix(openai-ws): preserve terminal event on lease loss
2026-08-04 16:28:22 +08:00
Wesley Liddick 35cab3c814 Merge pull request #5258 from feeeei/fix/model_plaza
fix(model-plaza): Model Plaza image model price display is inconsistent with the actual price
2026-08-04 16:27:53 +08:00
lyen1688 e592c5f9e0 新增腾讯天御验证码认证门禁 2026-08-04 15:09:29 +08:00
feeeei 785b61d424 修复模型广场图片模型价格展示与实收口径不一致
图片计费模型的广场展示价按实收口径计算:档位单价取
分组图片价 > 渠道档位价 > 渠道默认按次价;分组开启生图
独立倍率时实付倍率取独立倍率,不取分组/专属倍率。
2026-08-04 13:48:19 +08:00
zhiyu 2eb24814fe fix(codex): 强制统一出站身份并让客户端版本号跟随官方发布
上游 /backend-api/codex 在容量紧张时按客户端身份分优先级降载,被降载的请求
HTTP 200 后立刻推流内 server_is_overloaded。此前网关对配不出官方身份的客户端
整体回退到硬编码的 codex_cli_rs/0.144.1(落后官方 4 个发布),这些请求稳定
落在被优先丢弃的一侧。

- 强制统一出口:所有 OAuth 出站的 User-Agent / originator / version 一律改写
  为网关规范身份,客户端自报身份不参与构造;HTTP / 透传 / WS / alpha-search /
  探针全覆盖。compat 桥接故意删除 originator 的路径保持 no-op。
- 版本号收敛为单一来源,运行时优先级为面板覆写 → 自动同步值 → 内置常量;
  UA 与 version 头同源派生,不再各自硬编码。
- 新增 3 小时自动同步官方客户端最新稳定版,面板可关闭,无需为跟版本而发版。
- 流内 server_is_overloaded / slow_down 改为先在同账号有界重试再切号,并标记为
  请求级瞬时故障,不再据此临时封禁账号。
- 移除被取代的降载身份黑名单、浏览器 UA 兜底及其辅助函数。
2026-08-03 20:14:58 +08:00
Wesley Liddick 825ca7b1fc Merge pull request #5183 from rick147/codex/feat-openai-reset-credit-cache
feat(openai): refresh reset credit state after quota reset
2026-08-03 16:01:10 +08:00
shaw 54a2bcfd15 fix(openai): harden reset-credit refresh and account recovery
Review follow-ups on the reset-credit caching flow:

- Recover account state BEFORE (and independently of) the reset-credit
  display cache. A failed cache refresh could previously abort the run and
  leave the account rate-limited — the very reason the credit was spent
  (#3672 / #3740). The recovered account row is now returned even when the
  cache refresh fails.
- Run the post-reset bookkeeping on a detached, time-boxed context and give
  the panel reset call a larger timeout. A client abort no longer strands a
  consumed (non-refundable) credit with an unrecovered account, and the
  chained upstream calls can no longer exceed the client timeout and invite a
  retry that spends a second credit.
- Persist the reset-credit snapshot through POST /accounts/:id/quota/refresh
  instead of a side-effecting GET flag, so the write is covered by the audit
  middleware. A rejected snapshot write now degrades to cache_persisted=false
  instead of turning a successful upstream read into a 502 that left the card
  without a credit count and the reset button permanently disabled.
- Reject snapshots whose positive count carries no expiration entries, and
  drop expired credits (clamping the count) when rehydrating, so a stale
  cache can no longer light up the reset button.
- Keep nil quota / rate-limit services nil in the handler's interface fields;
  storing a nil *Service made the "not enabled" guards non-nil.
- Time-box the usage-refresh suppression and reuse handleAccountUpdated so the
  patched row also enters the auto-refresh silent window.
2026-08-03 14:40:55 +08:00
Wesley Liddick 27e8f69a9e Merge pull request #5171 from heathermhuang/codex/composite-reasoning-policy
feat(composite): enforce reasoning effort policy
2026-08-03 11:28:38 +08:00
rick147 a0802f00b6 feat: cache OpenAI reset credit details 2026-08-02 21:32:31 +08:00
r266-tech 30d2589ef0 fix(openai-ws): preserve terminal event on lease loss 2026-08-02 11:47:59 +08:00
shaw dec47e8fae fix(profit-control): stop leaking profit policy, close veto livelock, restore passthrough turn pricing
审计修复,逐条如下。

H1 利润策略泄露给所有普通用户
  profit_control_enabled / profit_min_margin / profit_safety_buffer 从
  dto.Group 移到 dto.AdminGroup(后者内嵌前者),赋值相应从
  groupFromServiceBase 移到 GroupFromServiceAdmin;前端 TS 同步从 Group 移到
  AdminGroup。dto.Group 是 GET /api/v1/groups/available 的响应体,该响应本就带
  rate_multiplier,相乘即可反推运营方上游采购成本上限。
  api_contract_test.go 的 /groups/available golden JSON 回滚这三个字段,并把
  fixture 改成非零值(require.JSONEq 是精确比对,缺字段即失败)。
  新增 dto 层边界测试:普通用户 DTO 不含三字段、管理员 DTO 仍含。

M1 利润终检 continue 与 failover 503 退避互动产生活锁
  FailoverState 新增 profitVetoedAccountIDs / profitVetoCount 与
  RecordProfitVeto():加入排除集 + 计数,达 maxProfitVetoAttempts(10) 返回
  FailoverExhausted。HandleSelectionExhausted 的 503 清空分支改为清空后把利润
  否决的账号放回排除集;若排除集已全部由利润否决贡献,清空不会带来任何新候选,
  直接判定耗尽(否则 SwitchCount 永不前进、退避条件永远成立,每 2s 空转一轮)。
  五个 handler 否决点(gateway_handler ×2 / responses / chat_completions /
  gemini_v1beta)改为经 RecordProfitVeto 决策,耗尽时按无可用账号终止。
  回归测试钉死:503 之后持续利润否决必须有限步终止且不 spin;未启用利润控制的
  请求退避语义完全不变。

M2 排队等槽后才终检,延迟可放大到 N × WaitPlan.Timeout
  OpenAI 侧选号循环(自有 failedAccountIDs map,非 FailoverState)新增
  recordOpenAIProfitVeto + handleOpenAIProfitVetoExhausted,共用同一上限语义。
  覆盖 responses / messages-dispatch / chat_completions / alpha_search /
  embeddings / images / grok_media 七处,以及 WS 两处否决分支。

M4 ws_v2 透传 ingress 绕过 per-turn 重定价(选方案 B:最小止血)
  透传 relay 只回调 AfterTurn、没有任何 turn 起始回调,hooks.BeforeTurn 永远
  不触发,而 handler 把 turnPricingAt 初始化成建连时刻 ⇒ 透传连接全部 turn 按
  建连时刻的高峰因子结算,客户端峰前建连保活即可全程谷价——正是本 PR 想堵的
  漏洞。改为 openAIWSTurnPricing 零值起步、只由 BeforeTurn 冻结;透传路径保持
  零值,RecordUsage 回退记录时刻,与引入利润控制前的基线一致。
  未选方案 A(给透传补 turn 起始回调):passthrough_relay.go 是 #5167 刚修过的
  取消传播/close frame 时序敏感区;且 BeforeTurn 还承担 turn>1 的并发槽位抢占,
  接进去等于给透传连接引入 per-turn 抢槽,风险远超本次修复范围。透传仍有建连时
  的准入门,只是没有 turn 级复核,已在两处注释写明。
  测试:service 层钉死透传 ingress 不触发 BeforeTurn(含失败时的复核指引),
  handler 层钉死零值语义与逐 turn 覆盖。

M5 装门读分组走了带账号计数聚合的 GetByID
  SchedulerSnapshotService 新增 GetGroupByIDLite,openai/gateway 两处装门改用
  之。门只需要平台/倍率/利润/高峰字段,且该查询发生在「是否启用利润控制」判定
  之前,未启用的分组同样付代价。两个测试 stub 的 GetByID 改成 panic 守卫。

M6 认证快照注释与真实读取路径相反
  门解析优先取 ctxkey.Group,而它就是本快照物化出来的对象,直连流量走的正是这
  条路。改正注释,与 api_key_repo.go 投影处的说明对齐,避免后人照旧注释删列。

M3 rate_multiplier 为 nil 时利润门 fail-closed(不改行为,加护栏)
  保留 fail-closed。补 repository 层测试钉死账号调度快照的 full/metadata 两份
  payload 都必须保留 RateMultiplier(含 0 值),漏列在 CI 就红。

L1 迁移号注释 191 / 191-192 改为实际的 192/193。
L2 admin group Create 的利润配置预校验改用与 CreateGroup 一致的归一化平台
   (新增 service.NormalizeGroupPlatform,两边共用)。保留预校验而非删除:
   service 层返回的是无类型 error,经 ErrorFrom 会变成 500,删掉会把合法的
   400 降级成 500。
L3 前端利润校验的上界改为判定换算后的小数(后端按小数校验 [0,1)),
   99.999% 会四舍五入进位成 1.0 而被后端 400;i18n en/zh 同步改为 0-99.99。
L4 clampProfitControlThreshold / profitControlOverThreshold 抽为共用函数,
   线上装门/否决点与 profit-preview 不再各自实现,附边界语义测试。
L5 profit-preview 补「默认 D 有账号但最低有效 D 归零」的告警(两档都为 0 由
   既有告警覆盖,不重复)。
2026-08-01 22:39:33 +08:00
Brisbanehuang fad2f215e8 fix(profit-control): decouple gate from image intent, propagate gate via selection, restore eager sticky fallback
Review fixes for the profit-control feature commit:

- Image intent no longer disables the profit gate. The shared /v1/responses
  handler previously skipped the pricing context (and therefore the gate)
  whenever the platform-wide image intent predicate matched, which includes
  Codex's passive image_gen namespace declaration: any client could disable
  admission control for anthropic/gemini/antigravity groups by declaring a
  namespace tool in the request body. Both /v1/responses paths now always
  install the token pricing context; image intent only drives capability
  routing and image billing. Mixed token+image requests stay token-gated;
  only dedicated media endpoints remain out of scope.
- Out-of-scope paths are now explicitly suppressed instead of implicitly
  ungated: Grok media (billed by media multipliers; also prevents in-flight
  video lookups from turning into spurious 404s), OpenAI-group count_tokens
  (unbilled), and Live calls (duration billed) carry a suppress marker that
  every install point honors, including the defensive scheduler-entry
  install.
- The gate resolved during selection now travels back to handlers on the
  AccountSelectionResult. The shared gateway installed the gate only on a
  scheduler-local context, so handler-side post-slot terminal rechecks and
  post-admission sticky binding were no-ops for anthropic/gemini/
  antigravity/shared-grok requests (and for composite-routed member groups
  on the OpenAI path). Handlers re-apply the carried gate via
  ContextWithSelectionProfitGate before the terminal recheck and binding;
  the WS acquired-selection branch gained the previously missing recheck.
- Sticky binding semantics restored for ungated traffic:
  BindStickySessionAfterProfitAdmission falls back to the official eager
  bind when no gate is installed (wait paths lost their only binding point
  otherwise), reads the pre-existing binding at bind time only when gated
  (removes the unconditional per-request Redis read the feature added to
  the shared handlers), and the legacy engine's three selection-time
  binding writes are skipped under a gate so a terminally vetoed account
  can no longer become the new sticky target.
- Responses WS connections re-freeze pricingAt and re-resolve the gate at
  every turn (BeforeTurn) and bill each turn with its own instant, closing
  the connect-at-valley/bill-at-valley window; a turn that fails the
  recheck closes the connection so the client reselects on reconnect.
- Terminal recheck no longer swaps a DB-fresh selected account for a
  staler snapshot object (UpdatedAt guard), and observer counters are
  documented as per-evaluation.
- Migrations renumbered 191/192 -> 192/193 after the passkey migration
  landed upstream as 191.

New regressions: selection-carried gate propagation (control group proves
the pre-fix no-op), image intent not disabling the shared gate, eager
binding fallback without a gate (both services), gated
read-failure/sentinel-miss binding semantics, legacy-engine deferred
binding under a gate with eager behavior preserved ungated, turn-level
pricing refresh (config re-resolution, scheduled-group precedence,
suppress, mid-connection disable), and suppress-marker coverage for the
request pricing context.
2026-08-01 22:39:32 +08:00
Brisbanehuang 20ad5ec506 feat(scheduler): per-group profit control for token account admission
Group pricing (rate multiplier, peak windows, per-user overrides) and
account cost (accounts.rate_multiplier) already live side by side, but
nothing stops the scheduler from handing a request to an account whose
cost multiplier exceeds what the group's pricing can profitably serve.
Add an opt-in per-group profit gate that filters scheduling candidates
by a margin rule, while ordering, scoring, stickiness and breakers keep
working unchanged among qualified accounts.

Admission rule: an account qualifies iff U <= D * (1 - min_margin -
safety_buffer) within a small relative epsilon, where U is
accounts.rate_multiplier (0 is legal; missing/negative/NaN/Inf are
conservatively rejected as invalid) and D is the requester's effective
downstream multiplier (user-group override ?? group default, times the
group peak factor) frozen at the request's pricing instant.

- groups gain profit_control_enabled / profit_min_margin /
  profit_safety_buffer (migration 191); the durable auth-cache
  invalidation trigger additionally watches the profit and pricing
  columns (migration 192) so out-of-band group edits cannot leave
  stale auth snapshots; GetByKeyForAuth explicitly projects the new
  columns and the API-key auth snapshot version is bumped to force a
  refresh of pre-existing snapshots
- request-level pricing instant: token entry points install pricingAt
  into ctx; the profit threshold D and the RecordUsage peak factor
  read the same instant, so one request never changes price mid-flight
  across waits/retries/failover (media and unwired paths keep the
  existing record-time semantics)
- the gate covers token requests on openai, anthropic, gemini, grok
  and antigravity groups: OpenAI-family handlers via
  WithOpenAIRequestPricingContext (responses incl. WS bridge, chat
  completions, messages, embeddings, alpha search), the shared gateway
  via WithGatewayTokenRequestPricing (messages, chat completions,
  responses, gemini model actions); composite groups cannot enable it
  directly; image/video/models/usage/count_tokens stay ungated and an
  explicit image-generation intent suppresses the gate end to end
- post-slot recheck: after a slot is acquired the account is re-read
  via SchedulerSnapshotService.GetAccount (scheduler cache, then DB;
  only when both fail the check fails open with WARN + metric); a
  vetoed account releases its slot and joins the request's exclusion
  set for reselection; sticky bindings are written only after the
  final check passes, and an over-threshold sticky account is skipped,
  not deleted, so it comes back once its rate recovers
- sticky-session cache contract: GatewayCache.GetSessionAccountID now
  returns ErrStickySessionNotFound on a miss (mapped from redis.Nil in
  the repository implementation, mirroring ErrRefreshTokenNotFound) so
  the profit sticky path can distinguish "no binding yet" from a real
  read failure without importing the cache driver in service code
- cross-group re-entry (composite parent -> member group) resolves the
  gate against the member group and clears a stale parent gate instead
  of letting a foreign threshold veto accounts
- per-platform/group activity counters (installs, threshold vetoes,
  invalid-rate vetoes, refresh failures) for observability
- admin UI: profit-control section on the five platforms' group forms
  with percent input, validation and platform-switch reset; group
  create/update/duplicate normalize and validate the config at a
  single choke point
- cmd/profit-preview: offline what-if tool that replays the production
  admission semantics over an exported config/account/override/model
  dump, reports per-model admitted-account counts under the default
  and the worst-case (lowest user override) D, and surfaces probe-sync
  staleness as warnings without affecting admission

Tests: service unit coverage for gate resolution/veto/threshold
epsilon/pricing instant/suppress marker/scheduler filtering and
post-slot recheck (incl. -race on the profit surface), unit-tagged
handler slot-recheck and capability-mapping regressions, sqlmock and
real-PostgreSQL integration regressions for the GetByKeyForAuth
projection and the migration-192 trigger watch list, API contract
update, and frontend specs for the five-platform form helpers.
2026-08-01 22:39:31 +08:00
Brisbanehuang b0f5007f04 feat(billing-probe): optionally sync account rate from upstream declared rate
Successful upstream billing probes already persist the upstream-declared
rate as a display-only snapshot. Add a per-account opt-in that writes
that declared rate back to the account's rate_multiplier, so the account
cost basis follows upstream repricing automatically instead of drifting
until an operator notices.

- new per-account flag upstream_billing_rate_sync_enabled stored next to
  the probe flag in account extra: enabling sync force-enables the
  probe, disabling the probe cascades sync off, and eligibility follows
  IsUpstreamBillingProbeIdentity (tightened from any non-empty platform
  to an explicit whitelist of the five supported API-key platforms so
  future platforms do not silently inherit probe/sync semantics)
- only a successful probe whose declared rate survives validation
  (finite, within bounds, not rounded to zero at the rate_multiplier
  decimal(10,4) scale) writes back; failed/unsupported/invalid probes
  leave rate_multiplier unchanged
- the writeback rides the existing snapshot CAS transaction:
  UpdateUpstreamBillingProbeSnapshot takes an optional rateMultiplier
  and applies it atomically with the snapshot under the same
  identity/snapshot compare-and-swap, so a probe result observed on a
  stale account cannot clobber a concurrent admin edit
- admin edit goes through UpdateWithAccountBillingSettings, which
  applies the form without overwriting a rate that a probe synchronized
  after the edit form was loaded (nil rateMultiplier = not edited);
  once sync is enabled the edit form shows the rate as managed
- bulk update rejects a manual rate_multiplier change when any target
  account has rate sync enabled (whole batch fails with a dedicated
  error so partial writes cannot bypass the sync ownership)
- frontend: sync toggle with hints in the edit modal (probe/sync
  enable/disable coupling enforced in the form), synced-rate tooltip on
  the rate cell, bulk edit modal warns and blocks rate edits that hit
  sync-enabled accounts; en/zh copy updated
- tests: service unit tests for sync gating/validation/cascade, sqlmock
  repo tests for the extended CAS, real-PostgreSQL integration tests
  (rate written only for successful+enabled accounts, manual rate
  protected after sync disabled, admin edit preserved across concurrent
  probe sync), handler/API contract updates, frontend specs for modal
  coupling, bulk rejection and rate cell
2026-08-01 22:11:09 +08:00
Heatherm Huang d60f4e442b feat(composite): enforce reasoning effort policy 2026-08-01 21:15:44 +08:00
shaw bd52e5d770 fix(gateway): record observed usage when anthropic stream is interrupted
Fixes #5148: with aggregator upstreams (e.g. newapi) that end SSE
streams without a proper terminal event, every such request was
silently missing from usage logs and billing.

Root cause (tracked via the nested audit issue): the low-level
Anthropic SSE readers already return the partially collected usage
together with the stream error (missing terminal event, read error,
interval timeout), but Forward converted every such result to
(nil, err) and the handler returned before submitting RecordUsage.

Changes:
- Add partialStreamUsageResult: on stream errors, wrap observed usage
  into a ForwardResult and return it alongside the error, for both the
  regular Anthropic path and the API-key passthrough path. Invariants:
  UpstreamFailoverError always keeps result=nil (failover retries are
  billed as the successful attempt, never twice), and zero observed
  usage returns no partial result (no phantom zero-usage records).
- Messages handler: hoist the usage submission block into a closure
  shared by the success path and the new partial-result error path.
- Usage record worker pool: distinguish pool-stopped drops
  (dropped_stopped) from operator-configured drop/sample overflow
  drops; billing tasks now fall back to inline synchronous execution
  only during the shutdown window, while explicit drop/sample overflow
  semantics are preserved. Image usage keeps its mandatory fallback
  for both drop kinds via the new mode.Dropped() helper.

Tests: Forward-level regressions for missing-terminal / read-error /
no-usage / failover-invariant on both paths, plus handler-level
stopped-pool sync fallback and drop-policy preservation tests.
2026-08-01 11:29:58 +08:00
feeeei 85a27fae39 fix(openai): retry SSE rate limits as HTTP 429
OpenAI Responses may return code=rate_limit_exceeded in a response.failed SSE event while the HTTP status remains 200. Classify these failures as 429 so configured pool-mode retries and account failover are applied. No account-level rate-limit state is written on this path: the 200-stream response headers carry normal quota snapshots, and retry semantics stay owned by the failover engine.

Ignore compact keepalive bytes when determining whether semantic output has started, preserving safe retries before real output.
2026-08-01 10:01:47 +08:00
Wesley Liddick d9fba8fe78 Merge pull request #5101 from Tongzai123/feat/admin-select-all-filtered-results
feat(admin): 支持按筛选结果全选账号
2026-08-01 08:54:38 +08:00
shaw 948b63c9ca feat(moderation): route content moderation through configurable proxy server
Implements #2646: the risk-control content audit can now send OpenAI
Moderations requests through a proxy from IP Management - Proxy Servers.

Backend:
- ContentModerationConfig gains proxy_id (nil = direct, unchanged default)
- update semantics: null keeps, 0 clears, >0 selects (validated to exist)
- moderation calls build the client via the shared httpclient pool; proxy
  resolution failure surfaces as a moderation error and never silently
  falls back to direct connection
- proxy_id -> URL resolution cached 60s (single-entry, invalidated on
  config save) so the pre-block hot path does not hit the DB per request
- test-key endpoint accepts proxy_id too (null = saved config's proxy,
  0 = force direct), so input-key/saved-key tests exercise the same path
- proxy usage/inactivity logged (content_moderation.proxy_enabled /
  proxy_not_active) without leaking credentials

Frontend:
- ProxySelector in the risk-control basic settings tab, proxy list loaded
  non-blockingly; save and test payloads carry proxy_id; zh/en i18n
2026-07-31 23:12:22 +08:00
Wesley Liddick 2980ff3850 Merge pull request #5094 from wucm667/feat/issue-5065-compact-homepage
feat(home): add compact home page preset to avoid abuse classification
2026-07-31 21:51:33 +08:00
shaw 017f6bbd5e fix(gateway): 收紧上游 URL 路径片段校验
网关有若干位置会把客户端可控的字符串拼进上游请求的 URL path(Responses
子路径、Gemini 模型名)。此前这些字符串未经校验直接参与拼接,可能改变上游
请求的路径结构,使实际发出的请求与客户端意图不一致。

- 新增 internal/service/upstream_path_guard.go:路径片段闭集允许清单
  (\w + `-` + `.`),拒绝空片段、纯点片段、超长片段与过深后缀
- /responses/*subpath 三条路由入口新增守卫,不可转发的子路径直接 404;
  service 层同时保证不产出不合规后缀,拼接函数再兜底一层
- Gemini AI Studio 原先 5 处重复的 URL 拼接收敛为唯一构造点
  buildGeminiAIStudioModelActionURL(校验模型片段 + action 白名单)
- Gemini native / GetModel handler 增加入口校验;ForwardAIStudioGET 逐片段校验
- Grok video 端点的 request_id 增加片段合规校验

合法子路径(/compact、/compact/detail、/{id}/cancel 形态)与既有模型名行为
不变,通配路由保留。
2026-07-31 16:11:45 +08:00
Wesley Liddick 60f6dc91cf Merge pull request #5115 from zvensmoluya/codex/update-gpt56-luna-terra-pricing
[codex] update GPT-5.6 Luna and Terra pricing
2026-07-31 11:44:21 +08:00
Zven 313121f3f7 test(pricing): update requested Terra cost 2026-07-31 10:02:04 +08:00
Zven 488d3b09ec test(pricing): update Terra billing ratios 2026-07-31 09:55:16 +08:00
Litong a35ff9613e feat(admin): 为账号批量删除增加并发限制 2026-07-30 18:03:54 +08:00
wucm667 739c0ff9c5 feat(home): add compact home page preset to avoid abuse classification
- Add compact_home_enabled setting to provide a minimal landing page
- Preserves custom home_content priority over compact mode
- Renders only site identity, navigation, and login/dashboard link
- Avoids marketing copy that triggers anti-fraud systems
- Includes focused backend/frontend tests and i18n support

Fixes #5065
2026-07-30 16:28:19 +08:00
Heatherm Huang 92dc61d401 fix(channels): show composite models by platform
Expand visible Composite groups into each configured concrete model platform while preserving ordinary group isolation and empty-state behavior.\n\nFixes #4985
2026-07-29 12:32:44 +08:00
Wesley Liddick 8fd01c2814 Merge pull request #5003 from feeeei/main
feat(模型广场): add model plaza with group-scoped pricing showcase
2026-07-28 20:02:52 +08:00
shaw 86fb4781f4 refactor(repository): scope user/api-key updates to declared columns
UserRepository.Update and APIKeyRepository.Update rewrote the whole row on
every call, regardless of which fields the caller meant to change. Several
columns on those tables are maintained by dedicated atomic paths (balance
deduction, quota and rate-limit counters, limit adjustments, activity
timestamps), so a caller holding a slightly older snapshot could silently
roll them back - a lost update.

Both methods now take an explicit column mask and persist only the columns
the caller declares; everything else keeps its current database value.

- All user and API-key call sites declare exactly what they mutate, which
  turns admin edits and profile saves into genuine partial updates.
- Email uniqueness locking/lookup and allowed_groups sync only run when
  those fields are part of the update.
- UserUpdateFields deliberately has no balance/total_recharged members, so
  Update cannot touch them. New AdjustBalance/SetBalance apply the change in
  a single statement and return before/after values; admin balance
  adjustment uses them instead of read-modify-write.
- promo_codes.used_count is no longer written by Update; it is only ever
  incremented by the redemption path.
- The billing hot path that marks an API key quota-exhausted writes only
  status.
- Dropped a no-op row write in RevokeAllUserTokens: users has no
  token_version column, so it persisted nothing while still overwriting
  concurrently-updated columns.

Adds integration coverage that a stale snapshot cannot revert concurrent
atomic writes, and unit coverage pinning the column set each entry point
declares.
2026-07-28 17:21:32 +08:00
feeeei 720c405e35 feat: add model plaza with group-scoped pricing showcase
- public /model-plaza page (standalone + admin-embedded) listing groups
  with discounted effective prices alongside LiteLLM official reference
- faceted platform/group/rate filters: cross-dimension options gray out
  instead of disappearing, platform-tinted chips via accent color-mix
- paid-price columns highlighted with per-platform tint band
- OptionalJWT middleware so anonymous and signed-in users share one route
- admin settings: enable switch, require-auth switch, markdown description
2026-07-28 16:19:41 +08:00
Wesley Liddick 2e432173f7 Merge pull request #4920 from alexj11324/feat/passkey-auth
feat: add passkey authentication
2026-07-28 14:58:37 +08:00
shaw 38ef8dc069 feat: require account password for passkey enrollment and revocation
A hijacked session must not be able to silently add a passkey as a
persistent backdoor or remove the victim's credentials. Registration
(begin) and deletion now verify the account password server-side,
reusing the existing PASSWORD_REQUIRED / PASSWORD_INCORRECT errors.

The password is used instead of TOTP step-up so the guard also protects
deployments that never configured a TOTP encryption key. The password
key in both request bodies is covered by the audit middleware's
key-substring redaction, so no credential material reaches audit_logs.

Frontend: the add-passkey form gains a current-password field, and the
delete confirmation is now a dialog with a password input (replacing
window.confirm), mirroring the TOTP disable dialog. Backend error
messages (e.g. wrong password) are surfaced instead of the generic
failure toast. Rename remains password-free as it is cosmetic.
2026-07-28 14:12:46 +08:00
Wesley Liddick a12d88ed4b Merge pull request #4957 from wucm667/fix/issue-4948-codex-web-search-manifest
fix(openai): preserve web search for API-key Codex clients
2026-07-28 11:07:16 +08:00
eyre 248236ce6d fix(gateway): 修复模拟响应使用 Bedrock msg_bdrk_ 格式,改为正宗 Anthropic msg_01 格式
问题:
探针拦截(suggestion mode / warmup / max_tokens=1 haiku)的模拟响应以及
Gemini/Antigravity 兼容层生成的 message ID 不符合 Anthropic 官方 API 格式,
容易被客户端识别为非正宗响应。

修复:
1. generateRealisticMsgID():msg_bdrk_ + 24字符 → msg_01 + 22位 Base62
   (与官方 API 返回的 msg_011CdS6b8gAhoKWdW9jE87Zs 格式一致)
2. 去掉固定的 msg_mock_suggestion / msg_mock_warmup,统一使用随机 ID
3. 流式响应格式对齐官方:
   - message_start 增加 stop_details/cache token 字段
   - content_block_start 字段顺序修正
   - message_delta.usage 只含 output_tokens
4. 非流式响应:增加 stop_details:null,移除非标准 total_tokens
5. Gemini Messages/ChatCompletions 兼容层:msg_ + hex → msg_01 + Base62
6. Antigravity response/stream transformer:msg_ + 12位 → msg_01 + 22位 Base62

验证方式:对照 Anthropic 官方 API 实际响应格式确认。
2026-07-27 15:20:03 +00:00
wucm667 3c62b5ca85 fix(openai): preserve web search for API-key Codex clients 2026-07-27 18:42:34 +08:00
shaw fead4c7ec3 feat(security): add panel API rate limiting to protect DB from high-frequency requests
用户可高频刷面板接口(usage/dashboard 等重聚合查询)直接打爆数据库:
现有限流器只覆盖登录/注册等公开认证入口,登录后的全部面板端点无任何限流。

三层防护(阈值均可在后台可视化配置,panel_rate_limit_settings):

1. 认证面板接口按「用户 ID」限流,与来源 IP 无关——反向代理/NAT 共享出口
   (所有请求源地址坍缩为 127.0.0.1 等)不会互相误伤:
   - Global 档(默认 240 rpm/账号):user/auth/payment/admin 全部登录后路由
   - Heavy 档(默认 60 rpm/账号):/usage、/usage/dashboard/*、
     /user/api-keys/:id/usage/daily 等重 SQL 聚合端点叠加计数
   - 管理员默认豁免(可关闭)

2. 无认证公开接口(/api/v1/settings/*,每次请求都查 DB)按安全客户端 IP
   限流(默认 300 rpm/IP);回环/私网/链路本地地址(反代内部转发地址)
   一律跳过计数,杜绝把整条反代链路合并进同一个桶造成大面积误拦截。

3. 修复既有隐患:auth 入口限流的 IP 取值从 c.ClientIP() 切换到与审计日志/
   会话绑定/API Key ACL 同源的安全客户端 IP 解析(尊重后台「信任反代转发
   IP」开关快照)。原实现下默认反代部署(未配置 server.trusted_proxies)
   所有用户共享同一个登录限流桶,既会全员误拦也可被单人恶意占满形成登录
   DoS;开关关闭时行为与原来完全一致。

工程约束:
- 配置热路径走进程内缓存(atomic.Value + singleflight,60s TTL),
  限流中间件零 DB 访问;保存后当前节点立即生效
- 面板限流 Redis 故障 fail-open(auth 入口保持原有 fail-close)
- 429 响应携带 Retry-After;错误码 RATE_LIMITED
- 支付 webhook / 公开支付回调有意不挂限流
- 新增 GET/PUT /api/v1/admin/settings/panel-rate-limit;设置页安全 tab
  新增「面板接口限流」卡片(zh/en i18n 全量)

测试:rate_limiter/panel_rate_limit/setting_panel_rate_limit 单测全绿;
routes、handler/admin、-tags unit 契约测试通过;前端 vue-tsc/ESLint/
SettingsView spec(26/26,含新增交互用例)/i18n 守卫全部通过。
2026-07-27 15:12:51 +08:00
Senn Chinn 3ce8efc125 Merge branch 'Wei-Shaw:main' into fix/antigravity-openai-compat 2026-07-27 13:40:05 +09:00
Wesley Liddick b765a7f9f6 Merge pull request #4890 from SemonCat/fix/openai-cross-mode-reasoning-failover
fix(openai): strip foreign reasoning on account failover
2026-07-27 11:46:49 +08:00
Wesley Liddick a74e11c26a Merge pull request #4868 from visa2/fix/settings-partial-update-clobber
fix(settings): keep fields a settings PUT never sent at their stored value
2026-07-27 11:44:40 +08:00
Wesley Liddick 4cc88e27b6 Merge pull request #4873 from wey-gu/fix/admin-usage-request-id-filter
fix(admin): filter usage logs by request id
2026-07-27 11:41:09 +08:00
Zhixuan Jiang 357c5b917b feat: add passkey sign-in settings control 2026-07-26 11:07:12 -04:00
Zhixuan Jiang 4158e73b3f fix: harden passkey deployment readiness 2026-07-26 10:14:55 -04:00
Zhixuan Jiang cc62979aa7 feat: add passkey authentication 2026-07-26 09:50:28 -04:00
chinnsenn 71d7f86883 fix(antigravity):
1. harden OpenAI compatibility forwarding
 2. reject usage-only non-stream responses
2026-07-26 20:23:50 +09:00
KtzeAbyss 7ce6e8d652 fix(openai): track websocket models per turn 2026-07-26 13:43:06 +08:00
Edison42andClaude Opus 4.8 a36222748d fix(openai): strip foreign reasoning on account failover
When a /responses forwarding loop has attempted an OpenAI passthrough
account, sanitize every subsequent non-passthrough attempt by deriving its
body from the immutable canonical request and dropping provider-specific
encrypted reasoning input items in full.

This prevents Bedrock-compatible accounts from rejecting Kiro reasoning
IDs and encrypted_content, including on same-account retries and later
non-passthrough failovers. Preserve JSON numbers exactly while sanitizing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-26 13:33:56 +08:00
Wey Gu 1850e00955 fix(admin): filter usage logs by request id 2026-07-26 00:53:13 +08:00
visa2andClaude Opus 5 0b5903d458 fix(settings): keep fields a settings PUT never sent at their stored value
PUT /api/v1/admin/settings is a whole-document write. The admin UI always sends
the complete document, so saving from the settings page is unaffected both
before and after this change. The bug is only reachable when an API client calls
the endpoint directly and sends just the fields it wants to change, which is the
natural assumption for a PUT on a settings resource.

Value-typed fields of UpdateSettingsRequest bind to their zero value when the
payload omits them, and buildSystemSettingsUpdates writes every key
unconditionally, so such a caller has no way to say "leave this one alone".
Omitting a field and explicitly clearing it are indistinguishable on the wire.
A caller that sends only the field it wants to change, e.g.

    {"risk_control_enabled": true}

sets that flag and clears every other unguarded field in the same request.
Measured against a fully configured store, one such call empties site_name,
site_subtitle, api_base_url, contact_info and doc_url, and turns
registration_enabled, email_verify_enabled, invitation_code_enabled and
turnstile_enabled off. turnstile_enabled alone gates the captcha on login,
register, forgot-password and both verify-code endpoints, and
email_verify_enabled is a precondition of IsPasswordResetEnabled.

The damage is easy to miss. site_name has a built-in fallback, so
getStringOrDefault renders the cleared value as the default product name and the
login page visibly changes, while the toggles just go quiet. Reopening the
settings page reads the already-cleared state back into the form, so correcting
the one visible field and saving persists the rest of the damage.

Fields that grew their own guard already survive this: the SMTP block falls back
to the previous values when smtp_host arrives empty, secret fields are written
only when non-empty, and 132 request fields are pointers whose handler merges an
omitted field with the stored value. This generalizes that pattern rather than
adding a fourth ad-hoc guard.

The handler now decodes the payload a second time as a raw field map, resolves
the setting key each absent field would have written, and hands that set to the
service, which drops those keys before SetMultiple, so the stored value is never
touched. Fields the payload does carry are written as before, giving the caller
the partial-update semantics it was already assuming. The mapping is reflected off
the request's json tags so new fields are covered without maintaining a list;
smtp_from_email is the only field whose json name differs from its setting key
and is aliased explicitly.

Only value-typed fields are filtered. Pointer fields keep whole-document
behaviour on purpose: forwarded_client_ip_headers and
api_key_acl_trust_forwarded_ip depend on being rewritten on every save to
re-normalize fail-closed state, which the malformed forwarded-client-IP header
test pins down.

An explicitly sent empty value is still a deliberate clear; only absent fields
are preserved. A partial write refreshes the in-process caches from storage
instead of from the request struct, which holds zero values for whatever the
caller omitted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 21:03:00 +08:00
shaw bc3acd6e28 fix(auth): 收紧注册别名查重(根点绕过 / 误拒 / 无界扫描 / 并发竞态)
对 #4814 的审计跟进修复:

- 域名尾随点绕过:user@gmail.com. 的域名不在 gmail 家族名单内,点号折叠与
  googlemail 归一被整体跳过,别名刷号原样可复现。归一化入口统一去掉 FQDN 根点。
- 误拒合法用户:剥 "+后缀" 缺空串守卫,+alice@ 与 +bob@ 都折叠成 @domain,
  该域后续 "+x@" 注册会永久 EMAIL_EXISTS 且无自助恢复。改为仅当 "+" 不在首位时剥离。
- 无界不可索引全表扫描:原实现按 LOWER(email) LIKE '%@domain' 把整域邮箱读进内存,
  且挂在公开未鉴权的 send-verify-code 上。改为按去点邮箱
  REPLACE(LOWER(TRIM(email)), '.', '') 做等值 + "local+%@domain" 前缀探针并带 LIMIT,
  新增同表达式的部分索引(migrations/190)。TRIM 口径与既有精确匹配一致,
  历史带首尾空白的行同样命中;LIKE 元字符转义,% 与 _ 不会扩大匹配面。
- 并发竞态:注册改走 CreateWithEmailAliasGuard,在邮箱唯一性锁上追加收件箱身份锁并在
  锁内复查,避免同一收件箱的多个别名变体同时通过服务层前置查重。管理员建号仍走
  Create,不受别名限制。
- 能力断言静默 fail-open:别名查重方法上提到 UserRepository 端口(编译期强制),
  移除可选接口类型断言与静默降级分支。
- OAuth 邮箱注册的两条建号路径(同样发放注册赠额)纳入同一查重口径;邮箱换绑/绑定
  不纳入,否则用户把邮箱改成自己收件箱的别名会被误拒。
2026-07-25 19:40:53 +08:00