Genuine Claude Code CLI sub-requests (e.g. the security monitor) carry
no identity system prompt, so claude_code_only groups wrongly rejected
them with "this group only allows Claude Code clients". Detect the
x-anthropic-billing-header block with cc_entrypoint=cli as a stable
client signal, while keeping the existing header/metadata checks.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
TTFT (first_token_ms) is only recorded for streaming requests, but the
ops dashboard weighted merged TTFT percentiles by success_count (all
successful requests, streaming + non-streaming). When non-streaming
traffic was present this diluted/skewed the merged TTFT figures shown
for longer (pre-aggregated) time ranges; the realtime path was exact.
Add a per-bucket ttft_sample_count (rows that actually recorded
first_token_ms) to ops_metrics_hourly / ops_metrics_daily and weight all
TTFT percentile merges by it instead of success_count:
- hourly/daily pre-agg upserts populate and propagate ttft_sample_count;
daily TTFT p50/p90/avg now weighted by ttft_sample_count.
- dashboard hourly-row merge and cross-segment combine weight TTFT by
the streaming sample count; queryUsageLatency returns it for raw
head/tail fragments.
duration stays weighted by success_count (recorded for every request);
p95/p99/max keep the conservative MAX merge (weight-independent).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The OpenAI/Codex 5h "used %" inversion that caused fresh accounts to show
~96-99% used (PR #2918, commit b65dde63) was already reverted in #2993, so the
stored value is now the correct "used %" again. This commit hardens that fix:
1. Regression test locking in direct "used %" semantics. The semantics have
flip-flopped twice (#2918 -> #2993) with no value-level guard — a fresh
account (secondary_used_percent=1, 5h window) must store
codex_5h_used_percent=1, not 99.
2. Stale-bounded self-heal in resolveOpenAIQuotaUtilization (the single
auto-pause chokepoint). An account poisoned with an inflated used% gets
excluded from scheduling, and a paused account never receives traffic to
refresh its snapshot — so it stayed stuck until the window's reset_at passed
(up to 5h/7d). When codex_usage_updated_at is older than 2h, the account is
no longer auto-paused on that snapshot; it gets one request whose response
headers refresh the snapshot and self-heal it. A missing timestamp is treated
as fresh (stays paused), and an actively-served exhausted account refreshes
the timestamp every response so it never crosses the bound — it cannot escape
auto-pause.
No change to Normalize(); no 100-x reintroduced; no new dependency wiring.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
When an OpenAI Chat Completions client targets an Anthropic-platform group,
ForwardAsChatCompletions converts the request CC → Responses → Anthropic
(ChatCompletionsToResponses → ResponsesToAnthropicRequest) before forwarding it
upstream. The Responses→Anthropic converter emits each function_call as its own
assistant message and each function_call_output as its own user message and
relies solely on mergeConsecutiveMessages to alternate roles. That is not enough
to satisfy Anthropic's tool-pairing invariants, so a trimmed or partial tool
history produces an upstream 400, e.g.:
tool_use_id found in tool_result blocks: call_00_...
Each tool_result block must have a corresponding tool_use block in the
previous message.
The failures this leaves unrepaired:
- orphan tool_result — a client that does sliding-window context management
keeps a recent tool result but drops the assistant tool_calls message that
announced it, so the tool_result has no matching tool_use;
- unanswered/dangling tool_use — a parallel call whose sibling result never
came back, or a call left dangling, which Anthropic also rejects.
Add normalizeAnthropicToolPairing, run between two merge passes: the first merge
groups parallel calls and their results; the pairing pass indexes every
tool_result by its tool_use id, keeps only answered tool_use blocks (dropping
unanswered/dangling calls, and the assistant message entirely when nothing else
remains) and re-emits the matching tool_result blocks as the immediately
following user message; standalone/orphan tool_results are dropped from their
original position; the second merge restores alternation. This mirrors
normalizeChatMessages on the Responses→Chat path.
Tested two ways: responses_to_anthropic_tool_pairing_test.go covers the repair
on direct Responses input (developer message between call and output, parallel
both-answered kept grouped, parallel one-unanswered dropped, orphan tool_result,
dangling call, single-call baseline); responses_to_anthropic_cc_chain_test.go
drives the real ChatCompletionsToResponses → ResponsesToAnthropicRequest chain
and reproduces the production 400 (orphan and unanswered-parallel) — both fail
without the repair and pass with it. The full apicompat suite stays green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The /v1/images/generations and /v1/images/edits paths routed every
non-failover upstream error through the generic handleErrorResponse,
whose final switch collapses anything that isn't 401/402/403/429 into a
hardcoded 502 "Upstream request failed" — discarding the actual upstream
status code, type, code, message, and param. So a gpt-image-2 400
(invalid_request_error, moderation, unsupported parameter, ...) reached
the client as an opaque 502.
The sibling Chat Completions and Messages compat paths already avoid this
via handleCompatErrorResponse, which preserves the real status and
message. Add the equivalent for images: handleOpenAIImagesErrorResponse
keeps all existing side-effects (ops logging, error-passthrough rules,
ShouldHandleErrorCode, account-disable/secondary-failover) but surfaces
the real upstream status + type/code/message/param by reusing the
existing OpenAIImagesUpstreamError machinery. Both forward paths now call
it instead of handleErrorResponse for non-failover errors.
Failover behavior (5xx / 401 / 403 / 429 / 529) is unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Call redis.replicate_commands() at the top of every Lua script that
reads redis.call('TIME'), so the resulting writes (ZADD, ZREMRANGEBYSCORE,
SET, DEL, EXPIRE) replicate correctly on Redis 3.2-4.x.
This is a no-op on Redis 5.0+, where effects replication is the default.
The function remains a public API on Redis 7.0+ (deprecated but supported)
and is still the documented way to opt into effects replication on older
versions, so calling it unconditionally is safe and forward-compatible.
Affected scripts (8 in total, all already using redis.call('TIME') by
design to avoid client-side clock skew across multiple app instances):
- concurrency_cache.go: acquireScript, getCountScript,
cleanupExpiredSlotsScript
- session_limit_cache.go: registerSessionScript, refreshSessionScript,
getActiveSessionCountScript, isSessionActiveScript
- user_msg_queue_cache.go: releaseLockScript
Scripts that do not invoke non-deterministic commands are intentionally
left untouched.
GET /api/v1/keys/:id previously returned distinct HTTP status
codes for 'key not found' (404) vs 'key exists but belongs to
another user' (403). This oracle allowed attackers to enumerate
valid API key IDs by observing response differences.
Now returns 404 in both cases so the response is identical
regardless of whether a key exists.
Fixes: CWE-204 (Information Disclosure via ID Oracle)
HTML-encode user-supplied key names in both Create and Update
endpoints. Previously, names were stored verbatim — an attacker
could inject script tags that would execute in admin panels or
any view using innerHTML/v-html rendering.
Fixes: CWE-79 (Stored Cross-Site Scripting)
Include codex-auto-review in the OpenAI fallback models list so /v1/models exposes it when no account mapping is configured. Keep the entry aligned with the existing default model catalog.