Add gateway.models_list_read_max_bytes with the existing 8 MiB behavior as its default, and apply it consistently to generic, Codex, and Antigravity model-list reads.
Read one sentinel byte for Codex manifests so oversized responses return an explicit bounded upstream error instead of malformed JSON.
PR #1463 removed the Sora platform, but some references survived:
- README/README_CN/README_JA kept the 'Sora status (temporarily
unavailable)' sections and gateway.sora_* docs that the removal PR
never touched.
- deploy/config.example.yaml still documented ~130 lines of sora_*
gateway keys, the top-level sora: direct-client/storage block, and
token_refresh.sync_linked_sora_accounts - none of which map to any
field in the config structs anymore.
- The OIDC login PR (02a66a01c, branched off pre-removal main and
merged 4 days after #1463) re-added the dead
PublicSettings.SoraClientEnabled field, which no code ever sets.
- A release sync (748a84d87) re-introduced sora i18n keys that the
later i18n split (d9e514f98) faithfully carried into
locales/{zh,en}/admin/{overview,settings}.ts. No component references
any of these keys.
This drops all of the above. Pure deletions, no behavior change.
H1/H2: bill search and voice with code defaults when group prices are
nil (explicit 0 remains free); bump API key auth snapshot to v19 and
refresh incomplete media/search/audio projections.
M1–M6: free-quota soft gate fails open on cache miss with background
refresh and 60s default TTL; correct password_auth config docs; default
cross-client model map to true (→ grok-4.5); audit /tts and /web_search;
exclude composite from migration 220 video-price clears; never let a
search surcharge mask token pricing failures.
Ship Use Key templates that match Grok Build / Codex best practice: env
vars + multi-model config.toml with api_backend=responses, env_key over
hardcoded secrets, and clearer shell/path guidance for Claude/Codex/OpenCode.
Also set free_quota_token_limit default to 500k (24h soft-gate), clean up
personal-dev-only comments, and keep billing test fixtures aligned.
The proxy stream circuit introduced in v0.1.164 (#4749) removes every
account behind a quarantined proxy from scheduling. When all schedulable
accounts share one proxy (a common deployment), two mid-stream
disconnects within a minute zeroed out capacity for 10 minutes and every
request failed with 502. One HTTP/2 connection loss also killed all
multiplexed streams at once, tripping the threshold from a single event.
- Quarantine now degrades to a preference: when the only reason no
account is available is proxy quarantine, selection retries once with
the quarantine bypassed, so capacity can never reach zero.
- Disconnects within 3s per proxy collapse into one failure event.
- Add gateway.openai_proxy_stream_circuit.disabled escape hatch.
- A completed stream still clears the quarantine immediately; TTL,
thresholds and recording guards are unchanged.
Add an opt-in first semantic output budget for native HTTP Responses, including response-header wait. Keep preamble and keepalive bytes non-semantic so a stalled account can fail over once without replaying its response IDs. Defaults remain disabled.
Related to #4201, #4185, and #4248. Complements the HTTP/2 dead-connection fix in #4207.
Add a dedicated client first-message timeout while preserving the legacy 30-second default.
Use the resolved value for both the WebSocket read deadline and structured timeout logs, and document tuning for large requests or slow links.
Add configuration, validation, handler, and resolver regression coverage.
Refs #4158
Adds a use-it-or-lose-it scheduling strategy: prefer accounts whose
session window resets soonest, so near-reset accounts get drained first
instead of accounts whose reset is still far away.
Both schedulers, opt-in, default behavior unchanged:
- Anthropic (gateway_service.go): new GatewaySchedulingConfig
.PreferSoonestReset flag. When on, the layered load-aware selection
inserts a filterBySoonestReset stage (priority -> soonest-reset ->
load -> LRU). Accounts with no active SessionWindowEnd are treated as
lowest priority; ties fall through to LRU.
- OpenAI/Codex (openai_account_scheduler.go): new "reset" score weight
in GatewayOpenAIWSSchedulerScoreWeights. Soonest-reset accounts score
higher; weight defaults to 0 (no effect).
SessionWindowEnd (upstream 5h/quota ResetsAt) is already carried in the
scheduler snapshot, so no snapshot changes are needed.
Documented in deploy/config.example.yaml. Adds unit tests for the
Anthropic filter and the OpenAI reset factor.