PR #5423 relaxed the email suffix whitelist: once a whitelist is
configured, non-whitelisted registrable domains are each allowed to
register one account. That behavior activated unconditionally.
Add registration_email_domain_quota_enabled (default false) to gate it:
- Off (default): restore pre-#5423 strict whitelist semantics — with a
non-empty whitelist, non-whitelisted domains are rejected with
EMAIL_SUFFIX_NOT_ALLOWED; the register/verify views restore the
client-side whitelist pre-check and allowed-domain hint.
- On: keep #5423 behavior — one account per non-whitelisted registrable
domain (EMAIL_DOMAIN_REGISTRATION_LIMIT).
- Empty whitelist keeps allowing all domains in both states.
Gating lives in validateRegistrationEmailQuota and (as a race-safety
backstop) createUserWithRegistrationEmailGuard; the repository-level
domain lock + in-tx recheck is unchanged. The admin update field is
*bool (omitted = keep current) so stale full-payload saves cannot
silently flip the switch. Email binding and OAuth auto-signup keep
their strict policy, and pending-OAuth bind-login for existing
accounts is unaffected because the handler resolves existing emails
before the quota check.
Frontend adds the toggle to admin settings (zh/en copy; whitelist hint
restored to strict wording, quota wording moved to the new toggle) and
exposes the flag via public settings + SSR injection payload.
Tests: #5423 quota tests now enable the switch explicitly; new
default-off regression tests cover register/send-code/async/pending
OAuth/OIDC create-account plus both register views; API contract JSON
and the injection drift guard are updated.
Resolves three conflicts, all of the "both branches appended to the same
block" shape. Every one is resolved as a union of both sides; nothing from
either parent is dropped.
- handler/admin/setting_handler_update.go: keep ChannelMonitorHideThroughput
(V2) alongside GrokDefaultTextModel / GrokCrossClientModelMapEnabled /
GrokDefaultBaseURLMode (#5408). UpdateSettings writes every key on each
save, so dropping either side would reset those settings to zero values.
- service/domain_constants.go: keep SettingKeyChannelMonitorHideThroughput
and the three SettingKeyGrok* constants.
- repository/migrations_runner.go: keep the 195 checksum rule (V2) and the
218/219/220 rules (#5408).
H1/H2: bill search and voice with code defaults when group prices are
nil (explicit 0 remains free); bump API key auth snapshot to v19 and
refresh incomplete media/search/audio projections.
M1–M6: free-quota soft gate fails open on cache miss with background
refresh and 60s default TTL; correct password_auth config docs; default
cross-client model map to true (→ grok-4.5); audit /tts and /web_search;
exclude composite from migration 220 video-price clears; never let a
search surcharge mask token pricing failures.
Default channel_monitor_mode to v1 (opt-in V2) so upgrades keep active
probes; existing explicit v2 rows are left alone via ON CONFLICT DO NOTHING
plus migration checksum compatibility for already-applied 195.
V2 first-enable backfill no longer compresses ticks to 5s or uses 24h
chunks. Each tick does recent overlap plus at most one historical chunk
with depth-based ceilings (2h/4h/6h), adaptive grow/shrink, and failure
backoff. Error request_id dedup is bounded by a 90-minute lookback so
ops_error_logs is not scanned for full history.
Also align hide_throughput parse default with privacy-preserving public
runtime (missing key → true).
Align gateway Grok media/voice paths and model lists, harden upstream failure
and quota handling, clear non-Grok video generation config migration, and polish
temp-unsched/status indicators with model whitelist updates.
Review follow-ups on the reset-credit caching flow:
- Recover account state BEFORE (and independently of) the reset-credit
display cache. A failed cache refresh could previously abort the run and
leave the account rate-limited — the very reason the credit was spent
(#3672 / #3740). The recovered account row is now returned even when the
cache refresh fails.
- Run the post-reset bookkeeping on a detached, time-boxed context and give
the panel reset call a larger timeout. A client abort no longer strands a
consumed (non-refundable) credit with an unrecovered account, and the
chained upstream calls can no longer exceed the client timeout and invite a
retry that spends a second credit.
- Persist the reset-credit snapshot through POST /accounts/:id/quota/refresh
instead of a side-effecting GET flag, so the write is covered by the audit
middleware. A rejected snapshot write now degrades to cache_persisted=false
instead of turning a successful upstream read into a 502 that left the card
without a credit count and the reset button permanently disabled.
- Reject snapshots whose positive count carries no expiration entries, and
drop expired credits (clamping the count) when rehydrating, so a stale
cache can no longer light up the reset button.
- Keep nil quota / rate-limit services nil in the handler's interface fields;
storing a nil *Service made the "not enabled" guards non-nil.
- Time-box the usage-refresh suppression and reuse handleAccountUpdated so the
patched row also enters the auto-refresh silent window.
Group pricing (rate multiplier, peak windows, per-user overrides) and
account cost (accounts.rate_multiplier) already live side by side, but
nothing stops the scheduler from handing a request to an account whose
cost multiplier exceeds what the group's pricing can profitably serve.
Add an opt-in per-group profit gate that filters scheduling candidates
by a margin rule, while ordering, scoring, stickiness and breakers keep
working unchanged among qualified accounts.
Admission rule: an account qualifies iff U <= D * (1 - min_margin -
safety_buffer) within a small relative epsilon, where U is
accounts.rate_multiplier (0 is legal; missing/negative/NaN/Inf are
conservatively rejected as invalid) and D is the requester's effective
downstream multiplier (user-group override ?? group default, times the
group peak factor) frozen at the request's pricing instant.
- groups gain profit_control_enabled / profit_min_margin /
profit_safety_buffer (migration 191); the durable auth-cache
invalidation trigger additionally watches the profit and pricing
columns (migration 192) so out-of-band group edits cannot leave
stale auth snapshots; GetByKeyForAuth explicitly projects the new
columns and the API-key auth snapshot version is bumped to force a
refresh of pre-existing snapshots
- request-level pricing instant: token entry points install pricingAt
into ctx; the profit threshold D and the RecordUsage peak factor
read the same instant, so one request never changes price mid-flight
across waits/retries/failover (media and unwired paths keep the
existing record-time semantics)
- the gate covers token requests on openai, anthropic, gemini, grok
and antigravity groups: OpenAI-family handlers via
WithOpenAIRequestPricingContext (responses incl. WS bridge, chat
completions, messages, embeddings, alpha search), the shared gateway
via WithGatewayTokenRequestPricing (messages, chat completions,
responses, gemini model actions); composite groups cannot enable it
directly; image/video/models/usage/count_tokens stay ungated and an
explicit image-generation intent suppresses the gate end to end
- post-slot recheck: after a slot is acquired the account is re-read
via SchedulerSnapshotService.GetAccount (scheduler cache, then DB;
only when both fail the check fails open with WARN + metric); a
vetoed account releases its slot and joins the request's exclusion
set for reselection; sticky bindings are written only after the
final check passes, and an over-threshold sticky account is skipped,
not deleted, so it comes back once its rate recovers
- sticky-session cache contract: GatewayCache.GetSessionAccountID now
returns ErrStickySessionNotFound on a miss (mapped from redis.Nil in
the repository implementation, mirroring ErrRefreshTokenNotFound) so
the profit sticky path can distinguish "no binding yet" from a real
read failure without importing the cache driver in service code
- cross-group re-entry (composite parent -> member group) resolves the
gate against the member group and clears a stale parent gate instead
of letting a foreign threshold veto accounts
- per-platform/group activity counters (installs, threshold vetoes,
invalid-rate vetoes, refresh failures) for observability
- admin UI: profit-control section on the five platforms' group forms
with percent input, validation and platform-switch reset; group
create/update/duplicate normalize and validate the config at a
single choke point
- cmd/profit-preview: offline what-if tool that replays the production
admission semantics over an exported config/account/override/model
dump, reports per-model admitted-account counts under the default
and the worst-case (lowest user override) D, and surfaces probe-sync
staleness as warnings without affecting admission
Tests: service unit coverage for gate resolution/veto/threshold
epsilon/pricing instant/suppress marker/scheduler filtering and
post-slot recheck (incl. -race on the profit surface), unit-tagged
handler slot-recheck and capability-mapping regressions, sqlmock and
real-PostgreSQL integration regressions for the GetByKeyForAuth
projection and the migration-192 trigger watch list, API contract
update, and frontend specs for the five-platform form helpers.
Successful upstream billing probes already persist the upstream-declared
rate as a display-only snapshot. Add a per-account opt-in that writes
that declared rate back to the account's rate_multiplier, so the account
cost basis follows upstream repricing automatically instead of drifting
until an operator notices.
- new per-account flag upstream_billing_rate_sync_enabled stored next to
the probe flag in account extra: enabling sync force-enables the
probe, disabling the probe cascades sync off, and eligibility follows
IsUpstreamBillingProbeIdentity (tightened from any non-empty platform
to an explicit whitelist of the five supported API-key platforms so
future platforms do not silently inherit probe/sync semantics)
- only a successful probe whose declared rate survives validation
(finite, within bounds, not rounded to zero at the rate_multiplier
decimal(10,4) scale) writes back; failed/unsupported/invalid probes
leave rate_multiplier unchanged
- the writeback rides the existing snapshot CAS transaction:
UpdateUpstreamBillingProbeSnapshot takes an optional rateMultiplier
and applies it atomically with the snapshot under the same
identity/snapshot compare-and-swap, so a probe result observed on a
stale account cannot clobber a concurrent admin edit
- admin edit goes through UpdateWithAccountBillingSettings, which
applies the form without overwriting a rate that a probe synchronized
after the edit form was loaded (nil rateMultiplier = not edited);
once sync is enabled the edit form shows the rate as managed
- bulk update rejects a manual rate_multiplier change when any target
account has rate sync enabled (whole batch fails with a dedicated
error so partial writes cannot bypass the sync ownership)
- frontend: sync toggle with hints in the edit modal (probe/sync
enable/disable coupling enforced in the form), synced-rate tooltip on
the rate cell, bulk edit modal warns and blocks rate edits that hit
sync-enabled accounts; en/zh copy updated
- tests: service unit tests for sync gating/validation/cascade, sqlmock
repo tests for the extended CAS, real-PostgreSQL integration tests
(rate written only for successful+enabled accounts, manual rate
protected after sync disabled, admin edit preserved across concurrent
probe sync), handler/API contract updates, frontend specs for modal
coupling, bulk rejection and rate cell
UserRepository.Update and APIKeyRepository.Update rewrote the whole row on
every call, regardless of which fields the caller meant to change. Several
columns on those tables are maintained by dedicated atomic paths (balance
deduction, quota and rate-limit counters, limit adjustments, activity
timestamps), so a caller holding a slightly older snapshot could silently
roll them back - a lost update.
Both methods now take an explicit column mask and persist only the columns
the caller declares; everything else keeps its current database value.
- All user and API-key call sites declare exactly what they mutate, which
turns admin edits and profile saves into genuine partial updates.
- Email uniqueness locking/lookup and allowed_groups sync only run when
those fields are part of the update.
- UserUpdateFields deliberately has no balance/total_recharged members, so
Update cannot touch them. New AdjustBalance/SetBalance apply the change in
a single statement and return before/after values; admin balance
adjustment uses them instead of read-modify-write.
- promo_codes.used_count is no longer written by Update; it is only ever
incremented by the redemption path.
- The billing hot path that marks an API key quota-exhausted writes only
status.
- Dropped a no-op row write in RevokeAllUserTokens: users has no
token_version column, so it persisted nothing while still overwriting
concurrently-updated columns.
Adds integration coverage that a stale snapshot cannot revert concurrent
atomic writes, and unit coverage pinning the column set each entry point
declares.
- parseSettings now reports passkey_enabled=false whenever the WebAuthn
deployment config is absent: a stale "true" row left behind after the
config is removed previously made the admin update gate reject every
settings save while the UI toggle was disabled, leaving no recovery
path from the admin panel. Added a regression test.
- update the admin settings API contract goldens with the new
passkey_enabled/passkey_configured/passkey_rp_id/passkey_rp_origins
fields.
- errcheck: check rows.Close in passkey repository (repo convention).
- staticcheck QF1001: apply De Morgan's law in WebAuthn origin scheme
validation.