The account+model transient breaker reset its failure streak whenever the
gap since the previous failure exceeded a one-minute window. That made the
breaker's sensitivity a function of request rate rather than upstream
health: a gateway called less often than once a minute never advanced past
streak 1, where the cooldown is zero, so a consistently broken account was
never blocked. Every request re-selected it, paid a full upstream attempt,
and only then failed over to a healthy account.
Observed on a low-traffic deployment: two accounts returning 500 and 503
stayed in rotation indefinitely, logging `failure_streak: 1, cooldown_ms: 0`
on every request and adding ~750ms to each one before a working account was
reached.
The streak is already cleared on success — recordSuccess deletes the entry,
and every OpenAI handler reports the schedule result — so the time-based
reset is not needed to recover a healthy account. Keep a TTL purely to bound
the map for account+model pairs that stopped being used, and raise it well
above the cooldowns so it no longer doubles as a streak reset.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review follow-ups on the reset-credit caching flow:
- Recover account state BEFORE (and independently of) the reset-credit
display cache. A failed cache refresh could previously abort the run and
leave the account rate-limited — the very reason the credit was spent
(#3672 / #3740). The recovered account row is now returned even when the
cache refresh fails.
- Run the post-reset bookkeeping on a detached, time-boxed context and give
the panel reset call a larger timeout. A client abort no longer strands a
consumed (non-refundable) credit with an unrecovered account, and the
chained upstream calls can no longer exceed the client timeout and invite a
retry that spends a second credit.
- Persist the reset-credit snapshot through POST /accounts/:id/quota/refresh
instead of a side-effecting GET flag, so the write is covered by the audit
middleware. A rejected snapshot write now degrades to cache_persisted=false
instead of turning a successful upstream read into a 502 that left the card
without a credit count and the reset button permanently disabled.
- Reject snapshots whose positive count carries no expiration entries, and
drop expired credits (clamping the count) when rehydrating, so a stale
cache can no longer light up the reset button.
- Keep nil quota / rate-limit services nil in the handler's interface fields;
storing a nil *Service made the "not enabled" guards non-nil.
- Time-box the usage-refresh suppression and reuse handleAccountUpdated so the
patched row also enters the auto-refresh silent window.
Review fixes for the profit-control feature commit:
- Image intent no longer disables the profit gate. The shared /v1/responses
handler previously skipped the pricing context (and therefore the gate)
whenever the platform-wide image intent predicate matched, which includes
Codex's passive image_gen namespace declaration: any client could disable
admission control for anthropic/gemini/antigravity groups by declaring a
namespace tool in the request body. Both /v1/responses paths now always
install the token pricing context; image intent only drives capability
routing and image billing. Mixed token+image requests stay token-gated;
only dedicated media endpoints remain out of scope.
- Out-of-scope paths are now explicitly suppressed instead of implicitly
ungated: Grok media (billed by media multipliers; also prevents in-flight
video lookups from turning into spurious 404s), OpenAI-group count_tokens
(unbilled), and Live calls (duration billed) carry a suppress marker that
every install point honors, including the defensive scheduler-entry
install.
- The gate resolved during selection now travels back to handlers on the
AccountSelectionResult. The shared gateway installed the gate only on a
scheduler-local context, so handler-side post-slot terminal rechecks and
post-admission sticky binding were no-ops for anthropic/gemini/
antigravity/shared-grok requests (and for composite-routed member groups
on the OpenAI path). Handlers re-apply the carried gate via
ContextWithSelectionProfitGate before the terminal recheck and binding;
the WS acquired-selection branch gained the previously missing recheck.
- Sticky binding semantics restored for ungated traffic:
BindStickySessionAfterProfitAdmission falls back to the official eager
bind when no gate is installed (wait paths lost their only binding point
otherwise), reads the pre-existing binding at bind time only when gated
(removes the unconditional per-request Redis read the feature added to
the shared handlers), and the legacy engine's three selection-time
binding writes are skipped under a gate so a terminally vetoed account
can no longer become the new sticky target.
- Responses WS connections re-freeze pricingAt and re-resolve the gate at
every turn (BeforeTurn) and bill each turn with its own instant, closing
the connect-at-valley/bill-at-valley window; a turn that fails the
recheck closes the connection so the client reselects on reconnect.
- Terminal recheck no longer swaps a DB-fresh selected account for a
staler snapshot object (UpdatedAt guard), and observer counters are
documented as per-evaluation.
- Migrations renumbered 191/192 -> 192/193 after the passkey migration
landed upstream as 191.
New regressions: selection-carried gate propagation (control group proves
the pre-fix no-op), image intent not disabling the shared gate, eager
binding fallback without a gate (both services), gated
read-failure/sentinel-miss binding semantics, legacy-engine deferred
binding under a gate with eager behavior preserved ungated, turn-level
pricing refresh (config re-resolution, scheduled-group precedence,
suppress, mid-connection disable), and suppress-marker coverage for the
request pricing context.
Group pricing (rate multiplier, peak windows, per-user overrides) and
account cost (accounts.rate_multiplier) already live side by side, but
nothing stops the scheduler from handing a request to an account whose
cost multiplier exceeds what the group's pricing can profitably serve.
Add an opt-in per-group profit gate that filters scheduling candidates
by a margin rule, while ordering, scoring, stickiness and breakers keep
working unchanged among qualified accounts.
Admission rule: an account qualifies iff U <= D * (1 - min_margin -
safety_buffer) within a small relative epsilon, where U is
accounts.rate_multiplier (0 is legal; missing/negative/NaN/Inf are
conservatively rejected as invalid) and D is the requester's effective
downstream multiplier (user-group override ?? group default, times the
group peak factor) frozen at the request's pricing instant.
- groups gain profit_control_enabled / profit_min_margin /
profit_safety_buffer (migration 191); the durable auth-cache
invalidation trigger additionally watches the profit and pricing
columns (migration 192) so out-of-band group edits cannot leave
stale auth snapshots; GetByKeyForAuth explicitly projects the new
columns and the API-key auth snapshot version is bumped to force a
refresh of pre-existing snapshots
- request-level pricing instant: token entry points install pricingAt
into ctx; the profit threshold D and the RecordUsage peak factor
read the same instant, so one request never changes price mid-flight
across waits/retries/failover (media and unwired paths keep the
existing record-time semantics)
- the gate covers token requests on openai, anthropic, gemini, grok
and antigravity groups: OpenAI-family handlers via
WithOpenAIRequestPricingContext (responses incl. WS bridge, chat
completions, messages, embeddings, alpha search), the shared gateway
via WithGatewayTokenRequestPricing (messages, chat completions,
responses, gemini model actions); composite groups cannot enable it
directly; image/video/models/usage/count_tokens stay ungated and an
explicit image-generation intent suppresses the gate end to end
- post-slot recheck: after a slot is acquired the account is re-read
via SchedulerSnapshotService.GetAccount (scheduler cache, then DB;
only when both fail the check fails open with WARN + metric); a
vetoed account releases its slot and joins the request's exclusion
set for reselection; sticky bindings are written only after the
final check passes, and an over-threshold sticky account is skipped,
not deleted, so it comes back once its rate recovers
- sticky-session cache contract: GatewayCache.GetSessionAccountID now
returns ErrStickySessionNotFound on a miss (mapped from redis.Nil in
the repository implementation, mirroring ErrRefreshTokenNotFound) so
the profit sticky path can distinguish "no binding yet" from a real
read failure without importing the cache driver in service code
- cross-group re-entry (composite parent -> member group) resolves the
gate against the member group and clears a stale parent gate instead
of letting a foreign threshold veto accounts
- per-platform/group activity counters (installs, threshold vetoes,
invalid-rate vetoes, refresh failures) for observability
- admin UI: profit-control section on the five platforms' group forms
with percent input, validation and platform-switch reset; group
create/update/duplicate normalize and validate the config at a
single choke point
- cmd/profit-preview: offline what-if tool that replays the production
admission semantics over an exported config/account/override/model
dump, reports per-model admitted-account counts under the default
and the worst-case (lowest user override) D, and surfaces probe-sync
staleness as warnings without affecting admission
Tests: service unit coverage for gate resolution/veto/threshold
epsilon/pricing instant/suppress marker/scheduler filtering and
post-slot recheck (incl. -race on the profit surface), unit-tagged
handler slot-recheck and capability-mapping regressions, sqlmock and
real-PostgreSQL integration regressions for the GetByKeyForAuth
projection and the migration-192 trigger watch list, API contract
update, and frontend specs for the five-platform form helpers.