Review follow-ups on the reset-credit caching flow:
- Recover account state BEFORE (and independently of) the reset-credit
display cache. A failed cache refresh could previously abort the run and
leave the account rate-limited — the very reason the credit was spent
(#3672 / #3740). The recovered account row is now returned even when the
cache refresh fails.
- Run the post-reset bookkeeping on a detached, time-boxed context and give
the panel reset call a larger timeout. A client abort no longer strands a
consumed (non-refundable) credit with an unrecovered account, and the
chained upstream calls can no longer exceed the client timeout and invite a
retry that spends a second credit.
- Persist the reset-credit snapshot through POST /accounts/:id/quota/refresh
instead of a side-effecting GET flag, so the write is covered by the audit
middleware. A rejected snapshot write now degrades to cache_persisted=false
instead of turning a successful upstream read into a 502 that left the card
without a credit count and the reset button permanently disabled.
- Reject snapshots whose positive count carries no expiration entries, and
drop expired credits (clamping the count) when rehydrating, so a stale
cache can no longer light up the reset button.
- Keep nil quota / rate-limit services nil in the handler's interface fields;
storing a nil *Service made the "not enabled" guards non-nil.
- Time-box the usage-refresh suppression and reuse handleAccountUpdated so the
patched row also enters the auto-refresh silent window.
Review fixes for the profit-control feature commit:
- Image intent no longer disables the profit gate. The shared /v1/responses
handler previously skipped the pricing context (and therefore the gate)
whenever the platform-wide image intent predicate matched, which includes
Codex's passive image_gen namespace declaration: any client could disable
admission control for anthropic/gemini/antigravity groups by declaring a
namespace tool in the request body. Both /v1/responses paths now always
install the token pricing context; image intent only drives capability
routing and image billing. Mixed token+image requests stay token-gated;
only dedicated media endpoints remain out of scope.
- Out-of-scope paths are now explicitly suppressed instead of implicitly
ungated: Grok media (billed by media multipliers; also prevents in-flight
video lookups from turning into spurious 404s), OpenAI-group count_tokens
(unbilled), and Live calls (duration billed) carry a suppress marker that
every install point honors, including the defensive scheduler-entry
install.
- The gate resolved during selection now travels back to handlers on the
AccountSelectionResult. The shared gateway installed the gate only on a
scheduler-local context, so handler-side post-slot terminal rechecks and
post-admission sticky binding were no-ops for anthropic/gemini/
antigravity/shared-grok requests (and for composite-routed member groups
on the OpenAI path). Handlers re-apply the carried gate via
ContextWithSelectionProfitGate before the terminal recheck and binding;
the WS acquired-selection branch gained the previously missing recheck.
- Sticky binding semantics restored for ungated traffic:
BindStickySessionAfterProfitAdmission falls back to the official eager
bind when no gate is installed (wait paths lost their only binding point
otherwise), reads the pre-existing binding at bind time only when gated
(removes the unconditional per-request Redis read the feature added to
the shared handlers), and the legacy engine's three selection-time
binding writes are skipped under a gate so a terminally vetoed account
can no longer become the new sticky target.
- Responses WS connections re-freeze pricingAt and re-resolve the gate at
every turn (BeforeTurn) and bill each turn with its own instant, closing
the connect-at-valley/bill-at-valley window; a turn that fails the
recheck closes the connection so the client reselects on reconnect.
- Terminal recheck no longer swaps a DB-fresh selected account for a
staler snapshot object (UpdatedAt guard), and observer counters are
documented as per-evaluation.
- Migrations renumbered 191/192 -> 192/193 after the passkey migration
landed upstream as 191.
New regressions: selection-carried gate propagation (control group proves
the pre-fix no-op), image intent not disabling the shared gate, eager
binding fallback without a gate (both services), gated
read-failure/sentinel-miss binding semantics, legacy-engine deferred
binding under a gate with eager behavior preserved ungated, turn-level
pricing refresh (config re-resolution, scheduled-group precedence,
suppress, mid-connection disable), and suppress-marker coverage for the
request pricing context.
Group pricing (rate multiplier, peak windows, per-user overrides) and
account cost (accounts.rate_multiplier) already live side by side, but
nothing stops the scheduler from handing a request to an account whose
cost multiplier exceeds what the group's pricing can profitably serve.
Add an opt-in per-group profit gate that filters scheduling candidates
by a margin rule, while ordering, scoring, stickiness and breakers keep
working unchanged among qualified accounts.
Admission rule: an account qualifies iff U <= D * (1 - min_margin -
safety_buffer) within a small relative epsilon, where U is
accounts.rate_multiplier (0 is legal; missing/negative/NaN/Inf are
conservatively rejected as invalid) and D is the requester's effective
downstream multiplier (user-group override ?? group default, times the
group peak factor) frozen at the request's pricing instant.
- groups gain profit_control_enabled / profit_min_margin /
profit_safety_buffer (migration 191); the durable auth-cache
invalidation trigger additionally watches the profit and pricing
columns (migration 192) so out-of-band group edits cannot leave
stale auth snapshots; GetByKeyForAuth explicitly projects the new
columns and the API-key auth snapshot version is bumped to force a
refresh of pre-existing snapshots
- request-level pricing instant: token entry points install pricingAt
into ctx; the profit threshold D and the RecordUsage peak factor
read the same instant, so one request never changes price mid-flight
across waits/retries/failover (media and unwired paths keep the
existing record-time semantics)
- the gate covers token requests on openai, anthropic, gemini, grok
and antigravity groups: OpenAI-family handlers via
WithOpenAIRequestPricingContext (responses incl. WS bridge, chat
completions, messages, embeddings, alpha search), the shared gateway
via WithGatewayTokenRequestPricing (messages, chat completions,
responses, gemini model actions); composite groups cannot enable it
directly; image/video/models/usage/count_tokens stay ungated and an
explicit image-generation intent suppresses the gate end to end
- post-slot recheck: after a slot is acquired the account is re-read
via SchedulerSnapshotService.GetAccount (scheduler cache, then DB;
only when both fail the check fails open with WARN + metric); a
vetoed account releases its slot and joins the request's exclusion
set for reselection; sticky bindings are written only after the
final check passes, and an over-threshold sticky account is skipped,
not deleted, so it comes back once its rate recovers
- sticky-session cache contract: GatewayCache.GetSessionAccountID now
returns ErrStickySessionNotFound on a miss (mapped from redis.Nil in
the repository implementation, mirroring ErrRefreshTokenNotFound) so
the profit sticky path can distinguish "no binding yet" from a real
read failure without importing the cache driver in service code
- cross-group re-entry (composite parent -> member group) resolves the
gate against the member group and clears a stale parent gate instead
of letting a foreign threshold veto accounts
- per-platform/group activity counters (installs, threshold vetoes,
invalid-rate vetoes, refresh failures) for observability
- admin UI: profit-control section on the five platforms' group forms
with percent input, validation and platform-switch reset; group
create/update/duplicate normalize and validate the config at a
single choke point
- cmd/profit-preview: offline what-if tool that replays the production
admission semantics over an exported config/account/override/model
dump, reports per-model admitted-account counts under the default
and the worst-case (lowest user override) D, and surfaces probe-sync
staleness as warnings without affecting admission
Tests: service unit coverage for gate resolution/veto/threshold
epsilon/pricing instant/suppress marker/scheduler filtering and
post-slot recheck (incl. -race on the profit surface), unit-tagged
handler slot-recheck and capability-mapping regressions, sqlmock and
real-PostgreSQL integration regressions for the GetByKeyForAuth
projection and the migration-192 trigger watch list, API contract
update, and frontend specs for the five-platform form helpers.
Successful upstream billing probes already persist the upstream-declared
rate as a display-only snapshot. Add a per-account opt-in that writes
that declared rate back to the account's rate_multiplier, so the account
cost basis follows upstream repricing automatically instead of drifting
until an operator notices.
- new per-account flag upstream_billing_rate_sync_enabled stored next to
the probe flag in account extra: enabling sync force-enables the
probe, disabling the probe cascades sync off, and eligibility follows
IsUpstreamBillingProbeIdentity (tightened from any non-empty platform
to an explicit whitelist of the five supported API-key platforms so
future platforms do not silently inherit probe/sync semantics)
- only a successful probe whose declared rate survives validation
(finite, within bounds, not rounded to zero at the rate_multiplier
decimal(10,4) scale) writes back; failed/unsupported/invalid probes
leave rate_multiplier unchanged
- the writeback rides the existing snapshot CAS transaction:
UpdateUpstreamBillingProbeSnapshot takes an optional rateMultiplier
and applies it atomically with the snapshot under the same
identity/snapshot compare-and-swap, so a probe result observed on a
stale account cannot clobber a concurrent admin edit
- admin edit goes through UpdateWithAccountBillingSettings, which
applies the form without overwriting a rate that a probe synchronized
after the edit form was loaded (nil rateMultiplier = not edited);
once sync is enabled the edit form shows the rate as managed
- bulk update rejects a manual rate_multiplier change when any target
account has rate sync enabled (whole batch fails with a dedicated
error so partial writes cannot bypass the sync ownership)
- frontend: sync toggle with hints in the edit modal (probe/sync
enable/disable coupling enforced in the form), synced-rate tooltip on
the rate cell, bulk edit modal warns and blocks rate edits that hit
sync-enabled accounts; en/zh copy updated
- tests: service unit tests for sync gating/validation/cascade, sqlmock
repo tests for the extended CAS, real-PostgreSQL integration tests
(rate written only for successful+enabled accounts, manual rate
protected after sync disabled, admin edit preserved across concurrent
probe sync), handler/API contract updates, frontend specs for modal
coupling, bulk rejection and rate cell
Fixes#5148: with aggregator upstreams (e.g. newapi) that end SSE
streams without a proper terminal event, every such request was
silently missing from usage logs and billing.
Root cause (tracked via the nested audit issue): the low-level
Anthropic SSE readers already return the partially collected usage
together with the stream error (missing terminal event, read error,
interval timeout), but Forward converted every such result to
(nil, err) and the handler returned before submitting RecordUsage.
Changes:
- Add partialStreamUsageResult: on stream errors, wrap observed usage
into a ForwardResult and return it alongside the error, for both the
regular Anthropic path and the API-key passthrough path. Invariants:
UpstreamFailoverError always keeps result=nil (failover retries are
billed as the successful attempt, never twice), and zero observed
usage returns no partial result (no phantom zero-usage records).
- Messages handler: hoist the usage submission block into a closure
shared by the success path and the new partial-result error path.
- Usage record worker pool: distinguish pool-stopped drops
(dropped_stopped) from operator-configured drop/sample overflow
drops; billing tasks now fall back to inline synchronous execution
only during the shutdown window, while explicit drop/sample overflow
semantics are preserved. Image usage keeps its mandatory fallback
for both drop kinds via the new mode.Dropped() helper.
Tests: Forward-level regressions for missing-terminal / read-error /
no-usage / failover-invariant on both paths, plus handler-level
stopped-pool sync fallback and drop-policy preservation tests.
OpenAI Responses may return code=rate_limit_exceeded in a response.failed SSE event while the HTTP status remains 200. Classify these failures as 429 so configured pool-mode retries and account failover are applied. No account-level rate-limit state is written on this path: the 200-stream response headers carry normal quota snapshots, and retry semantics stay owned by the failover engine.
Ignore compact keepalive bytes when determining whether semantic output has started, preserving safe retries before real output.
Implements #2646: the risk-control content audit can now send OpenAI
Moderations requests through a proxy from IP Management - Proxy Servers.
Backend:
- ContentModerationConfig gains proxy_id (nil = direct, unchanged default)
- update semantics: null keeps, 0 clears, >0 selects (validated to exist)
- moderation calls build the client via the shared httpclient pool; proxy
resolution failure surfaces as a moderation error and never silently
falls back to direct connection
- proxy_id -> URL resolution cached 60s (single-entry, invalidated on
config save) so the pre-block hot path does not hit the DB per request
- test-key endpoint accepts proxy_id too (null = saved config's proxy,
0 = force direct), so input-key/saved-key tests exercise the same path
- proxy usage/inactivity logged (content_moderation.proxy_enabled /
proxy_not_active) without leaking credentials
Frontend:
- ProxySelector in the risk-control basic settings tab, proxy list loaded
non-blockingly; save and test payloads carry proxy_id; zh/en i18n
- Add compact_home_enabled setting to provide a minimal landing page
- Preserves custom home_content priority over compact mode
- Renders only site identity, navigation, and login/dashboard link
- Avoids marketing copy that triggers anti-fraud systems
- Includes focused backend/frontend tests and i18n support
Fixes#5065
Expand visible Composite groups into each configured concrete model platform while preserving ordinary group isolation and empty-state behavior.\n\nFixes #4985
UserRepository.Update and APIKeyRepository.Update rewrote the whole row on
every call, regardless of which fields the caller meant to change. Several
columns on those tables are maintained by dedicated atomic paths (balance
deduction, quota and rate-limit counters, limit adjustments, activity
timestamps), so a caller holding a slightly older snapshot could silently
roll them back - a lost update.
Both methods now take an explicit column mask and persist only the columns
the caller declares; everything else keeps its current database value.
- All user and API-key call sites declare exactly what they mutate, which
turns admin edits and profile saves into genuine partial updates.
- Email uniqueness locking/lookup and allowed_groups sync only run when
those fields are part of the update.
- UserUpdateFields deliberately has no balance/total_recharged members, so
Update cannot touch them. New AdjustBalance/SetBalance apply the change in
a single statement and return before/after values; admin balance
adjustment uses them instead of read-modify-write.
- promo_codes.used_count is no longer written by Update; it is only ever
incremented by the redemption path.
- The billing hot path that marks an API key quota-exhausted writes only
status.
- Dropped a no-op row write in RevokeAllUserTokens: users has no
token_version column, so it persisted nothing while still overwriting
concurrently-updated columns.
Adds integration coverage that a stale snapshot cannot revert concurrent
atomic writes, and unit coverage pinning the column set each entry point
declares.
A hijacked session must not be able to silently add a passkey as a
persistent backdoor or remove the victim's credentials. Registration
(begin) and deletion now verify the account password server-side,
reusing the existing PASSWORD_REQUIRED / PASSWORD_INCORRECT errors.
The password is used instead of TOTP step-up so the guard also protects
deployments that never configured a TOTP encryption key. The password
key in both request bodies is covered by the audit middleware's
key-substring redaction, so no credential material reaches audit_logs.
Frontend: the add-passkey form gains a current-password field, and the
delete confirmation is now a dialog with a password input (replacing
window.confirm), mirroring the TOTP disable dialog. Backend error
messages (e.g. wrong password) are surfaced instead of the generic
failure toast. Rename remains password-free as it is cosmetic.
When a /responses forwarding loop has attempted an OpenAI passthrough
account, sanitize every subsequent non-passthrough attempt by deriving its
body from the immutable canonical request and dropping provider-specific
encrypted reasoning input items in full.
This prevents Bedrock-compatible accounts from rejecting Kiro reasoning
IDs and encrypted_content, including on same-account retries and later
non-passthrough failovers. Preserve JSON numbers exactly while sanitizing.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
PUT /api/v1/admin/settings is a whole-document write. The admin UI always sends
the complete document, so saving from the settings page is unaffected both
before and after this change. The bug is only reachable when an API client calls
the endpoint directly and sends just the fields it wants to change, which is the
natural assumption for a PUT on a settings resource.
Value-typed fields of UpdateSettingsRequest bind to their zero value when the
payload omits them, and buildSystemSettingsUpdates writes every key
unconditionally, so such a caller has no way to say "leave this one alone".
Omitting a field and explicitly clearing it are indistinguishable on the wire.
A caller that sends only the field it wants to change, e.g.
{"risk_control_enabled": true}
sets that flag and clears every other unguarded field in the same request.
Measured against a fully configured store, one such call empties site_name,
site_subtitle, api_base_url, contact_info and doc_url, and turns
registration_enabled, email_verify_enabled, invitation_code_enabled and
turnstile_enabled off. turnstile_enabled alone gates the captcha on login,
register, forgot-password and both verify-code endpoints, and
email_verify_enabled is a precondition of IsPasswordResetEnabled.
The damage is easy to miss. site_name has a built-in fallback, so
getStringOrDefault renders the cleared value as the default product name and the
login page visibly changes, while the toggles just go quiet. Reopening the
settings page reads the already-cleared state back into the form, so correcting
the one visible field and saving persists the rest of the damage.
Fields that grew their own guard already survive this: the SMTP block falls back
to the previous values when smtp_host arrives empty, secret fields are written
only when non-empty, and 132 request fields are pointers whose handler merges an
omitted field with the stored value. This generalizes that pattern rather than
adding a fourth ad-hoc guard.
The handler now decodes the payload a second time as a raw field map, resolves
the setting key each absent field would have written, and hands that set to the
service, which drops those keys before SetMultiple, so the stored value is never
touched. Fields the payload does carry are written as before, giving the caller
the partial-update semantics it was already assuming. The mapping is reflected off
the request's json tags so new fields are covered without maintaining a list;
smtp_from_email is the only field whose json name differs from its setting key
and is aliased explicitly.
Only value-typed fields are filtered. Pointer fields keep whole-document
behaviour on purpose: forwarded_client_ip_headers and
api_key_acl_trust_forwarded_ip depend on being rewritten on every save to
re-normalize fail-closed state, which the malformed forwarded-client-IP header
test pins down.
An explicitly sent empty value is still a deliberate clear; only absent fields
are preserved. A partial write refreshes the in-process caches from storage
instead of from the request struct, which holds zero values for whatever the
caller omitted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>