Resolves three conflicts, all of the "both branches appended to the same
block" shape. Every one is resolved as a union of both sides; nothing from
either parent is dropped.
- handler/admin/setting_handler_update.go: keep ChannelMonitorHideThroughput
(V2) alongside GrokDefaultTextModel / GrokCrossClientModelMapEnabled /
GrokDefaultBaseURLMode (#5408). UpdateSettings writes every key on each
save, so dropping either side would reset those settings to zero values.
- service/domain_constants.go: keep SettingKeyChannelMonitorHideThroughput
and the three SettingKeyGrok* constants.
- repository/migrations_runner.go: keep the 195 checksum rule (V2) and the
218/219/220 rules (#5408).
H1/H2: bill search and voice with code defaults when group prices are
nil (explicit 0 remains free); bump API key auth snapshot to v19 and
refresh incomplete media/search/audio projections.
M1–M6: free-quota soft gate fails open on cache miss with background
refresh and 60s default TTL; correct password_auth config docs; default
cross-client model map to true (→ grok-4.5); audit /tts and /web_search;
exclude composite from migration 220 video-price clears; never let a
search surcharge mask token pricing failures.
Default channel_monitor_mode to v1 (opt-in V2) so upgrades keep active
probes; existing explicit v2 rows are left alone via ON CONFLICT DO NOTHING
plus migration checksum compatibility for already-applied 195.
V2 first-enable backfill no longer compresses ticks to 5s or uses 24h
chunks. Each tick does recent overlap plus at most one historical chunk
with depth-based ceilings (2h/4h/6h), adaptive grow/shrink, and failure
backoff. Error request_id dedup is bounded by a 90-minute lookback so
ops_error_logs is not scanned for full history.
Also align hide_throughput parse default with privacy-preserving public
runtime (missing key → true).
Align gateway Grok media/voice paths and model lists, harden upstream failure
and quota handling, clear non-Grok video generation config migration, and polish
temp-unsched/status indicators with model whitelist updates.
Group pricing (rate multiplier, peak windows, per-user overrides) and
account cost (accounts.rate_multiplier) already live side by side, but
nothing stops the scheduler from handing a request to an account whose
cost multiplier exceeds what the group's pricing can profitably serve.
Add an opt-in per-group profit gate that filters scheduling candidates
by a margin rule, while ordering, scoring, stickiness and breakers keep
working unchanged among qualified accounts.
Admission rule: an account qualifies iff U <= D * (1 - min_margin -
safety_buffer) within a small relative epsilon, where U is
accounts.rate_multiplier (0 is legal; missing/negative/NaN/Inf are
conservatively rejected as invalid) and D is the requester's effective
downstream multiplier (user-group override ?? group default, times the
group peak factor) frozen at the request's pricing instant.
- groups gain profit_control_enabled / profit_min_margin /
profit_safety_buffer (migration 191); the durable auth-cache
invalidation trigger additionally watches the profit and pricing
columns (migration 192) so out-of-band group edits cannot leave
stale auth snapshots; GetByKeyForAuth explicitly projects the new
columns and the API-key auth snapshot version is bumped to force a
refresh of pre-existing snapshots
- request-level pricing instant: token entry points install pricingAt
into ctx; the profit threshold D and the RecordUsage peak factor
read the same instant, so one request never changes price mid-flight
across waits/retries/failover (media and unwired paths keep the
existing record-time semantics)
- the gate covers token requests on openai, anthropic, gemini, grok
and antigravity groups: OpenAI-family handlers via
WithOpenAIRequestPricingContext (responses incl. WS bridge, chat
completions, messages, embeddings, alpha search), the shared gateway
via WithGatewayTokenRequestPricing (messages, chat completions,
responses, gemini model actions); composite groups cannot enable it
directly; image/video/models/usage/count_tokens stay ungated and an
explicit image-generation intent suppresses the gate end to end
- post-slot recheck: after a slot is acquired the account is re-read
via SchedulerSnapshotService.GetAccount (scheduler cache, then DB;
only when both fail the check fails open with WARN + metric); a
vetoed account releases its slot and joins the request's exclusion
set for reselection; sticky bindings are written only after the
final check passes, and an over-threshold sticky account is skipped,
not deleted, so it comes back once its rate recovers
- sticky-session cache contract: GatewayCache.GetSessionAccountID now
returns ErrStickySessionNotFound on a miss (mapped from redis.Nil in
the repository implementation, mirroring ErrRefreshTokenNotFound) so
the profit sticky path can distinguish "no binding yet" from a real
read failure without importing the cache driver in service code
- cross-group re-entry (composite parent -> member group) resolves the
gate against the member group and clears a stale parent gate instead
of letting a foreign threshold veto accounts
- per-platform/group activity counters (installs, threshold vetoes,
invalid-rate vetoes, refresh failures) for observability
- admin UI: profit-control section on the five platforms' group forms
with percent input, validation and platform-switch reset; group
create/update/duplicate normalize and validate the config at a
single choke point
- cmd/profit-preview: offline what-if tool that replays the production
admission semantics over an exported config/account/override/model
dump, reports per-model admitted-account counts under the default
and the worst-case (lowest user override) D, and surfaces probe-sync
staleness as warnings without affecting admission
Tests: service unit coverage for gate resolution/veto/threshold
epsilon/pricing instant/suppress marker/scheduler filtering and
post-slot recheck (incl. -race on the profit surface), unit-tagged
handler slot-recheck and capability-mapping regressions, sqlmock and
real-PostgreSQL integration regressions for the GetByKeyForAuth
projection and the migration-192 trigger watch list, API contract
update, and frontend specs for the five-platform form helpers.
Remove unsupported none handling, align the auth cache version with the current baseline, consolidate the unshipped migrations, and update the public API contract.
Persist reasoning ceilings and exact mappings for OpenAI groups, enforce them across HTTP and WebSocket forwarding, and invalidate cached auth snapshots.
- Add prompt_audit_events.full_prompt (migration 182) so admins can review
the exact unredacted prompt that triggered a finding; blocking mode writes
it from the snapshot, async mode reconstructs it from the Redis scan
payload so jobs rows stay redaction-only
- Event detail API returns full_prompt (list endpoint stays lean); text is
NUL-stripped and capped at 65536 runes
- Detail dialog shows the full prompt in a scrollable pane with fallback to
the legacy redacted preview; page copy updated to match the new behavior
- Rework filter deletion into a dedicated dialog with time-range presets and
criteria-change preview invalidation; localize decision/risk/category
labels across the events workspace
- Fix pre-existing i18n message-compile spec by declaring the
@intlify/message-compiler dev dependency
Admins often recreate groups with the same pricing, routing, and account membership. A server-side duplicate creates an inactive copy for review, preserves eligible account priorities, and recovers ambiguous retries without creating extra groups.
Constraint: Group has no neutral JSON metadata field for durable operation recovery
Constraint: Model routing references account IDs, so copied configuration requires matching bindings
Rejected: Rebuild from the list response | it omits configuration and account priority details
Rejected: Store operation identity in business configuration | it would pollute real group settings
Confidence: high
Scope-risk: moderate
Reversibility: clean
Directive: Keep duplicated groups inactive until an administrator reviews the copied configuration
Tested: Go unit and full tests, go vet, integration-tag compile, frontend Vitest, lint, typecheck, production build, and Playwright duplicate flow
Not-tested: PostgreSQL container integration locally because Docker is unavailable; CI will execute the database-backed suite