89d826be2 raised backend/go.mod to `go 1.26.6` and updated the three CI
workflows' version assertions, but left the Go builder image in all three
Dockerfiles pinned at 1.26.5. Since the official golang images set
GOTOOLCHAIN=local, the toolchain is not auto-downloaded and any image build
fails hard at `go mod download`.
CI does not catch this: the workflows build with actions/setup-go, not with
these Dockerfiles.
Also extend the Go-upgrade checklist in DEV_GUIDE.md, which listed only the
CI files -- that omission is why the Dockerfiles were missed.
Default codex_fingerprint_mode to off. v0.1.175 treated a missing key as
"session", so upgrading silently rewrote installation/session/thread/turn/
window identifiers for every existing OAuth account that had never configured
this field. The quota regressions in #5555, #5556 and #5582 line up with that
version boundary, with A/B reports that rolling back to v0.1.173 restores
quota. Convergence is now explicit opt-in (#5610).
Only accounts that never set the field change behaviour; explicit off /
device / session / full keep working exactly as configured. That required
flipping the persistence condition in all three account modals from
"!== 'session'" to "!== 'off'": the old rule deleted the key when it equalled
the default, which after the flip would have silently discarded an
administrator's explicit opt-in to session.
Also extend convergence to the passthrough path, which previously left client
identifiers untouched:
- resolve the ids once in forwardOpenAIPassthrough and rewrite
client_metadata on the raw bytes (gjson extract + sjson splice) because
passthrough is a hot path that must not fully unmarshal multi-MB bodies;
a shared core keeps the raw and map variants from drifting
- both request builders apply the staged ids at the same relative position
(after session isolation, before identity enforcement) so headers and body
share one id set and turn_id stays consistent
- stage the ids unconditionally, including nil: a failover from a converged
account to an off account must not leave the previous account's ids behind
OpenAI sunset the legacy unary /responses/compact endpoint (404, #5598,
#5624), so the account "compact probe" in the admin UI kept failing even for
healthy accounts, and the beta-feature negotiation header was only attached
to compaction turns.
Beta features (codex-rs session/mod.rs build_model_client_beta_features_header
+ client.rs build_responses_headers): the header is a session-level constant
attached to every /responses request, the WS handshake and /responses/compact.
Enumerating FEATURES shows no Experimental feature is enabled by default, so a
default install sends exactly "remote_compaction_v2". Mirror that:
- OAuth requests without a client-declared header get the default shape, so we
no longer produce a "header only on compaction turns" pattern real Codex
never emits (#5586 chains that strip the header)
- a client-declared header is preserved as-is: non-empty without v2 means the
user disabled the feature and the gateway must not rewrite that
- native v2 turns (compaction_trigger in body) always ensure v2 is present
- non-OAuth upstreams keep the compaction-turn-only behaviour
- the WS injection sits outside the client-header copy block so prewarm and
turn handshakes cannot land in different pool compatibility buckets
Compact probe now exercises native v2 (streaming /responses +
compaction_trigger) instead of the dead endpoint. Success requires an actual
compaction output item — scanning output_item.done/added, the terminal
response.output[] and the whole-JSON fallback — so a 2xx that silently drops
the trigger is reported as unsupported (the "got 0 items" class, #5478,
#5648). Probe identity is now UUID-shaped and applies the account's
convergence, matching real traffic on the same endpoint.
Codex captures x-codex-turn-state from /responses SSE, /responses/compact
JSON and the WS handshake (codex-api sse/responses.rs, endpoint/compact.rs),
then echoes it back on later requests of the same turn. The HTTP path dropped
it because the header is not in the generic response allowlist, while the WS
path already relayed it — an inconsistency that broke the protocol chain.
Relay it explicitly at every commit point instead of widening the global
allowlist (which would leak it into Anthropic/Gemini responses):
- streaming, non-streaming and SSE-to-JSON handlers relay it, clearing any
value left over by a previous failover attempt when upstream sends none
- under the first-output guard the header is only staged; provenance is
recorded when applyAttemptResponseHeaders actually writes it, because a
first-output timeout discards the staged headers and the client never
receives that blob
- record (api key + client session) -> minting account, TTL-bounded with an
opportunistic sweep, and strip echoes known to come from another account
before they go upstream. Stripping only: injection is the Claude bridge's
job. Same-account or unknown provenance passes through unchanged.
The channel cache holds a groupID -> platform map with a 10 minute TTL, and
only channel Create/Update/Delete call invalidateCache(). Changing a group's
platform through the admin API therefore leaves the cache pointing at the old
platform for up to 10 minutes.
Channel pricing, model mapping and the model whitelist are all matched per
platform, so during that window the lookups silently miss: pricing falls back
to the global LiteLLM price list, renames stop applying and the whitelist
stops restricting. Nothing is logged.
Inject a narrow ChannelCacheInvalidator into the admin service (same shape as
the existing APIKeyAuthCacheInvalidator) and call it from UpdateGroup only when
the platform actually changed. The dependency is optional -- when it is nil the
cache simply rebuilds on TTL expiry, as before.