normalizeOpenAIParallelToolCallsWithoutTools only looked at the top-level
"tools" array, but normalizeOpenAIResponsesLiteTools moves namespace tools
into an input item of type "additional_tools" and deletes the top-level key.
A Responses Lite request that carries tools therefore looks like it has none,
and the parallel_tool_calls:false that ensureOpenAIResponsesLiteParallelToolCalls
had just pinned gets deleted on the way out.
OpenAI then applies its default of true and rejects the request:
400 unsupported_value: "X-OpenAI-Internal-Codex-Responses-Lite requires
`parallel_tool_calls` to be false."
Note the field cannot simply be pinned to false unconditionally: without tools
OpenAI rejects it with "'parallel_tool_calls' is only allowed when 'tools' are
specified", so the two constraints have to be honoured together.
Reuse the same tool-detection standard the Lite path already uses by adding
openAIRequestBodyHasTools, the []byte counterpart of openAIResponsesLiteHasTools,
so both sides of the repo agree on what "has tools" means.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The upstream names one offending index per response, but a replayed
conversation routinely carries dozens of items of the same type, each with a
status the upstream schema does not accept. Clearing a single index per round
trip needs one retry per item, so a conversation with more than
maxOpenAIResponsesRejectedFieldRetries such items exhausts the bounded budget
and the 400 reaches the client. Reported against tool_search_output items,
where the rejection surfaced as "Unknown parameter: 'input[60].status'".
Clear the status of every input item sharing the rejected item's type in the
same pass. Items of other types keep theirs: the rejection only proves that
the rejected item's type has no status field. When the rejected item carries
no type to match on, fall back to clearing the named index alone.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Add gateway.models_list_read_max_bytes with the existing 8 MiB behavior as its default, and apply it consistently to generic, Codex, and Antigravity model-list reads.
Read one sentinel byte for Codex manifests so oversized responses return an explicit bounded upstream error instead of malformed JSON.
- cc_pipeline: observe Chat Completions chunks/bodies as untyped payloads
(empty event type) so the upstream-echoed service_tier is trusted, matching
the upstream constraint that only terminal events and untyped bodies report
the actual processing tier.
- tests: add model field to terminal SSE frames (observation only triggers on
model-bearing frames) and assert response.created tier echo is ignored.
- openai_ws_forwarder_ingress: resolve billing tier from the local
upstreamResponseModelObserver (upstream echo first) instead of the raw
request payload, matching the HTTP->WS bridge and WS v2 forwarder.
- openai_ws_http_bridge_test: add fast-alias + upstream default case
proving the local observer's echoed tier wins.
- handler tests: keep only invalid service_tier -> 400 (short-circuits at
validation); valid/omitted semantics covered by the pure service-level
validation tests, avoiding real account selection in tests.
- Accept fast|priority (canonical priority), flex|auto|default|scale on
/v1/responses and /v1/chat/completions; reject unknown/empty/non-string
with HTTP 400; omitted and null stay compatible.
- Propagate service_tier through JSON/SSE, Responses<->Chat conversions,
fallback paths and HTTP->upstream WebSocket bridge.
- Billing prefers the upstream terminal tier; the outbound (policy-
transformed) tier is used only when upstream omits the field.
Explicit upstream default bills Standard even when Fast was requested.
- Pricing: Fast premium 2x Standard for gpt-5.6-sol/terra/luna and
gpt-5.4; 2.5x for gpt-5.5; channel FastMultiplier stays authoritative.
- Live verification (official Codex 0.149.0 + gateway, HTTP & WS):
upstream ChatGPT backend may return terminal default even when the
account catalog advertises priority; billing follows the actual tier.
DashScope/DeepSeek later tool_call deltas send empty id and
function.name. Clients that merge with !== undefined overwrite
the first delta's identity and dispatch unknown tool "". Drop
those empty fields on the raw Chat Completions SSE path.
A terminal event that arrives with an empty output was rebuilt from delta
accumulation. BufferedResponseAccumulator models only one reasoning item, one
message, and N function calls, and records no item id, status, or phase, so a
turn carrying several items collapsed into a single fabricated message: the
reasoning item disappeared, the real message id was replaced, and phase was
lost.
reconstructResponseOutputFromSSE already prefers the raw output_item.done
items over accumulation for buffered responses. The streaming path had no
equivalent because it never sees the whole body at once. Collect the raw item
of each output_item.done keyed by output_index and rebuild from those, falling
back to accumulation only when the stream reported no done item at all.
Items are stored as raw JSON, so vendor extensions and item types this gateway
does not model survive the rebuild verbatim.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>