Review follow-ups on the reset-credit caching flow:
- Recover account state BEFORE (and independently of) the reset-credit
display cache. A failed cache refresh could previously abort the run and
leave the account rate-limited — the very reason the credit was spent
(#3672 / #3740). The recovered account row is now returned even when the
cache refresh fails.
- Run the post-reset bookkeeping on a detached, time-boxed context and give
the panel reset call a larger timeout. A client abort no longer strands a
consumed (non-refundable) credit with an unrecovered account, and the
chained upstream calls can no longer exceed the client timeout and invite a
retry that spends a second credit.
- Persist the reset-credit snapshot through POST /accounts/:id/quota/refresh
instead of a side-effecting GET flag, so the write is covered by the audit
middleware. A rejected snapshot write now degrades to cache_persisted=false
instead of turning a successful upstream read into a 502 that left the card
without a credit count and the reset button permanently disabled.
- Reject snapshots whose positive count carries no expiration entries, and
drop expired credits (clamping the count) when rehydrating, so a stale
cache can no longer light up the reset button.
- Keep nil quota / rate-limit services nil in the handler's interface fields;
storing a nil *Service made the "not enabled" guards non-nil.
- Time-box the usage-refresh suppression and reuse handleAccountUpdated so the
patched row also enters the auto-refresh silent window.
AdminRefundDialog already implements a `requireForce` prop that renders the
force checkbox and blocks submit until it is checked, but AdminOrdersView
never consumed `require_force` from the refund response nor bound the prop,
so the checkbox could never render. Admins hitting a require-force case (for
example a user who spent their balance after requesting the refund) only saw
an error toast and had no way to complete the refund from the UI.
Keep the dialog open when the backend answers `require_force`, surface the
checkbox and the backend warning, and reset the flag whenever the dialog is
opened or closed.
Verified with `npm run typecheck` (vue-tsc --noEmit, clean).
Review fixes for the profit-control feature commit:
- Image intent no longer disables the profit gate. The shared /v1/responses
handler previously skipped the pricing context (and therefore the gate)
whenever the platform-wide image intent predicate matched, which includes
Codex's passive image_gen namespace declaration: any client could disable
admission control for anthropic/gemini/antigravity groups by declaring a
namespace tool in the request body. Both /v1/responses paths now always
install the token pricing context; image intent only drives capability
routing and image billing. Mixed token+image requests stay token-gated;
only dedicated media endpoints remain out of scope.
- Out-of-scope paths are now explicitly suppressed instead of implicitly
ungated: Grok media (billed by media multipliers; also prevents in-flight
video lookups from turning into spurious 404s), OpenAI-group count_tokens
(unbilled), and Live calls (duration billed) carry a suppress marker that
every install point honors, including the defensive scheduler-entry
install.
- The gate resolved during selection now travels back to handlers on the
AccountSelectionResult. The shared gateway installed the gate only on a
scheduler-local context, so handler-side post-slot terminal rechecks and
post-admission sticky binding were no-ops for anthropic/gemini/
antigravity/shared-grok requests (and for composite-routed member groups
on the OpenAI path). Handlers re-apply the carried gate via
ContextWithSelectionProfitGate before the terminal recheck and binding;
the WS acquired-selection branch gained the previously missing recheck.
- Sticky binding semantics restored for ungated traffic:
BindStickySessionAfterProfitAdmission falls back to the official eager
bind when no gate is installed (wait paths lost their only binding point
otherwise), reads the pre-existing binding at bind time only when gated
(removes the unconditional per-request Redis read the feature added to
the shared handlers), and the legacy engine's three selection-time
binding writes are skipped under a gate so a terminally vetoed account
can no longer become the new sticky target.
- Responses WS connections re-freeze pricingAt and re-resolve the gate at
every turn (BeforeTurn) and bill each turn with its own instant, closing
the connect-at-valley/bill-at-valley window; a turn that fails the
recheck closes the connection so the client reselects on reconnect.
- Terminal recheck no longer swaps a DB-fresh selected account for a
staler snapshot object (UpdatedAt guard), and observer counters are
documented as per-evaluation.
- Migrations renumbered 191/192 -> 192/193 after the passkey migration
landed upstream as 191.
New regressions: selection-carried gate propagation (control group proves
the pre-fix no-op), image intent not disabling the shared gate, eager
binding fallback without a gate (both services), gated
read-failure/sentinel-miss binding semantics, legacy-engine deferred
binding under a gate with eager behavior preserved ungated, turn-level
pricing refresh (config re-resolution, scheduled-group precedence,
suppress, mid-connection disable), and suppress-marker coverage for the
request pricing context.
Group pricing (rate multiplier, peak windows, per-user overrides) and
account cost (accounts.rate_multiplier) already live side by side, but
nothing stops the scheduler from handing a request to an account whose
cost multiplier exceeds what the group's pricing can profitably serve.
Add an opt-in per-group profit gate that filters scheduling candidates
by a margin rule, while ordering, scoring, stickiness and breakers keep
working unchanged among qualified accounts.
Admission rule: an account qualifies iff U <= D * (1 - min_margin -
safety_buffer) within a small relative epsilon, where U is
accounts.rate_multiplier (0 is legal; missing/negative/NaN/Inf are
conservatively rejected as invalid) and D is the requester's effective
downstream multiplier (user-group override ?? group default, times the
group peak factor) frozen at the request's pricing instant.
- groups gain profit_control_enabled / profit_min_margin /
profit_safety_buffer (migration 191); the durable auth-cache
invalidation trigger additionally watches the profit and pricing
columns (migration 192) so out-of-band group edits cannot leave
stale auth snapshots; GetByKeyForAuth explicitly projects the new
columns and the API-key auth snapshot version is bumped to force a
refresh of pre-existing snapshots
- request-level pricing instant: token entry points install pricingAt
into ctx; the profit threshold D and the RecordUsage peak factor
read the same instant, so one request never changes price mid-flight
across waits/retries/failover (media and unwired paths keep the
existing record-time semantics)
- the gate covers token requests on openai, anthropic, gemini, grok
and antigravity groups: OpenAI-family handlers via
WithOpenAIRequestPricingContext (responses incl. WS bridge, chat
completions, messages, embeddings, alpha search), the shared gateway
via WithGatewayTokenRequestPricing (messages, chat completions,
responses, gemini model actions); composite groups cannot enable it
directly; image/video/models/usage/count_tokens stay ungated and an
explicit image-generation intent suppresses the gate end to end
- post-slot recheck: after a slot is acquired the account is re-read
via SchedulerSnapshotService.GetAccount (scheduler cache, then DB;
only when both fail the check fails open with WARN + metric); a
vetoed account releases its slot and joins the request's exclusion
set for reselection; sticky bindings are written only after the
final check passes, and an over-threshold sticky account is skipped,
not deleted, so it comes back once its rate recovers
- sticky-session cache contract: GatewayCache.GetSessionAccountID now
returns ErrStickySessionNotFound on a miss (mapped from redis.Nil in
the repository implementation, mirroring ErrRefreshTokenNotFound) so
the profit sticky path can distinguish "no binding yet" from a real
read failure without importing the cache driver in service code
- cross-group re-entry (composite parent -> member group) resolves the
gate against the member group and clears a stale parent gate instead
of letting a foreign threshold veto accounts
- per-platform/group activity counters (installs, threshold vetoes,
invalid-rate vetoes, refresh failures) for observability
- admin UI: profit-control section on the five platforms' group forms
with percent input, validation and platform-switch reset; group
create/update/duplicate normalize and validate the config at a
single choke point
- cmd/profit-preview: offline what-if tool that replays the production
admission semantics over an exported config/account/override/model
dump, reports per-model admitted-account counts under the default
and the worst-case (lowest user override) D, and surfaces probe-sync
staleness as warnings without affecting admission
Tests: service unit coverage for gate resolution/veto/threshold
epsilon/pricing instant/suppress marker/scheduler filtering and
post-slot recheck (incl. -race on the profit surface), unit-tagged
handler slot-recheck and capability-mapping regressions, sqlmock and
real-PostgreSQL integration regressions for the GetByKeyForAuth
projection and the migration-192 trigger watch list, API contract
update, and frontend specs for the five-platform form helpers.
Successful upstream billing probes already persist the upstream-declared
rate as a display-only snapshot. Add a per-account opt-in that writes
that declared rate back to the account's rate_multiplier, so the account
cost basis follows upstream repricing automatically instead of drifting
until an operator notices.
- new per-account flag upstream_billing_rate_sync_enabled stored next to
the probe flag in account extra: enabling sync force-enables the
probe, disabling the probe cascades sync off, and eligibility follows
IsUpstreamBillingProbeIdentity (tightened from any non-empty platform
to an explicit whitelist of the five supported API-key platforms so
future platforms do not silently inherit probe/sync semantics)
- only a successful probe whose declared rate survives validation
(finite, within bounds, not rounded to zero at the rate_multiplier
decimal(10,4) scale) writes back; failed/unsupported/invalid probes
leave rate_multiplier unchanged
- the writeback rides the existing snapshot CAS transaction:
UpdateUpstreamBillingProbeSnapshot takes an optional rateMultiplier
and applies it atomically with the snapshot under the same
identity/snapshot compare-and-swap, so a probe result observed on a
stale account cannot clobber a concurrent admin edit
- admin edit goes through UpdateWithAccountBillingSettings, which
applies the form without overwriting a rate that a probe synchronized
after the edit form was loaded (nil rateMultiplier = not edited);
once sync is enabled the edit form shows the rate as managed
- bulk update rejects a manual rate_multiplier change when any target
account has rate sync enabled (whole batch fails with a dedicated
error so partial writes cannot bypass the sync ownership)
- frontend: sync toggle with hints in the edit modal (probe/sync
enable/disable coupling enforced in the form), synced-rate tooltip on
the rate cell, bulk edit modal warns and blocks rate edits that hit
sync-enabled accounts; en/zh copy updated
- tests: service unit tests for sync gating/validation/cascade, sqlmock
repo tests for the extended CAS, real-PostgreSQL integration tests
(rate written only for successful+enabled accounts, manual rate
protected after sync disabled, admin edit preserved across concurrent
probe sync), handler/API contract updates, frontend specs for modal
coupling, bulk rejection and rate cell
coder/websocket arms a context.AfterFunc that hard-closes the connection
when a write context is canceled, and AfterFunc stop does not wait for a
callback that already started. An external cancellation (e.g. ingress
lease loss) landing inside the disarm window of an already-successful
downstream write could therefore kill the TCP connection before the
retry close frame (1013) was written, leaving the client with a bare
EOF. Mirror the read side: bound downstream writes with the write
timeout only, and rely on the explicit Close/CloseNow performed by every
relay exit path for teardown.
Fixes the flaky TestPassthroughLifecycle_LeaseLossSendsRetryClose.
/v1/sub2api/billing is a key-scoped sub2api convention: any API-key
account whose base_url points at a sub2api-compatible upstream answers
it regardless of the account platform. Widen probe eligibility from
platform=openai to every API-key account (OAuth/Bedrock stay excluded:
no static key to present).
- central predicate exported as IsUpstreamBillingProbeIdentity; runner,
manual probe, SetAccountEnabled, admin create/update/bulk validation
and CRS reconcile all follow it
- probe target resolution reads credentials.api_key/base_url directly.
OpenAI keeps its official-default base URL and openai transport
profile. Other platforms whose base_url is empty or points at an
official provider API domain (anthropic.com, googleapis.com, x.ai,
grok.com, openai.com - matched on the normalized hostname, port and
trailing dot stripped, as the exact host or any subdomain) persist
"unsupported" without sending a request: the create form fills empty
base_url with official defaults and offers official regional presets
(e.g. us-east-1.api.x.ai), and official APIs cannot answer
/v1/sub2api/billing, so probing would only send the account key to a
nonexistent official path
- due-scan SQL, BulkUpdate probe WHERE, UpdateCredentials stale-snapshot
CASE and proxy-change invalidation drop their platform filters
(type='apikey' retained)
- frontend: rate cell, edit/create/bulk modals gate on type==='apikey';
the antigravity upstream create flow (its own helper and form section)
shows the auto-probe toggle and passes upstream_billing_probe_enabled;
bulk WS-mode section stays OpenAI-only; settings copy de-scoped
- tests: multiplatform service unit tests (relay success, official/empty
base_url unsupported without request, normalized official-host matrix,
OpenAI defaults preserved), real-PostgreSQL due-scan coverage for
openai/anthropic/grok plus oauth/disabled exclusion, updated
sqlmock/CRS/frontend specs; probe stays opt-in per account
Fixes#5148: with aggregator upstreams (e.g. newapi) that end SSE
streams without a proper terminal event, every such request was
silently missing from usage logs and billing.
Root cause (tracked via the nested audit issue): the low-level
Anthropic SSE readers already return the partially collected usage
together with the stream error (missing terminal event, read error,
interval timeout), but Forward converted every such result to
(nil, err) and the handler returned before submitting RecordUsage.
Changes:
- Add partialStreamUsageResult: on stream errors, wrap observed usage
into a ForwardResult and return it alongside the error, for both the
regular Anthropic path and the API-key passthrough path. Invariants:
UpstreamFailoverError always keeps result=nil (failover retries are
billed as the successful attempt, never twice), and zero observed
usage returns no partial result (no phantom zero-usage records).
- Messages handler: hoist the usage submission block into a closure
shared by the success path and the new partial-result error path.
- Usage record worker pool: distinguish pool-stopped drops
(dropped_stopped) from operator-configured drop/sample overflow
drops; billing tasks now fall back to inline synchronous execution
only during the shutdown window, while explicit drop/sample overflow
semantics are preserved. Image usage keeps its mandatory fallback
for both drop kinds via the new mode.Dropped() helper.
Tests: Forward-level regressions for missing-terminal / read-error /
no-usage / failover-invariant on both paths, plus handler-level
stopped-pool sync fallback and drop-policy preservation tests.