When converting a Chat Completions stream into Responses events, the first
tool_call delta chunk was copied wholesale into stream state (including
function.arguments), then the same chunk's arguments were accumulated again by
the shared `+=` block. For OpenAI this is harmless because its first tool_call
chunk carries empty arguments, but upstreams that pack id+name+arguments into a
single chunk (e.g. GLM/Zhipu) end up with doubled arguments such as
{"cmd":"ls"}{"cmd":"ls"}. Codex then fails to parse the tool call with
"trailing characters", breaking every tool invocation.
Reset the copied arguments so the shared accumulator counts them exactly once,
keeping the emitted delta and the final done/arguments consistent.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The concurrency / switch-rate / throughput cards on the ops dashboard
sit in grid cells that only set `min-h-[360px]` (no definite height).
Their inner card uses `h-full`, which resolves to `auto` when the parent
height is `auto`. Combined with the Chart.js `responsive` +
`maintainAspectRatio: false` charts, this forms a height feedback loop:
the canvas reads the parent height to size itself, the content then grows,
the next ResizeObserver tick reads an even larger height, and the cards
stretch downward without bound.
On wide screens (`lg:grid-cols-4`) a sibling card usually fixes the row
height via `align-items: stretch`, masking the issue. It surfaces when no
sibling bounds the row height — e.g. the single-column (`grid-cols-1`)
stacked layout on narrow viewports, or when the concurrency card collapses
to little content. `min-h` only sets a floor, not a ceiling.
Fix: give the two Chart.js canvas cells a definite height (`h-[360px]`) so
the responsive resize has a fixed reference and the loop cannot run. The
concurrency card is not a responsive canvas, so it keeps `min-h-[360px]`
to avoid clipping its content.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
When gpt-image upstream returns response.completed with no image but a text
refusal (content moderation, e.g. the model replies "this request was judged
unsafe to generate"), the soft-failure path treated it as a probabilistic
upstream failure and returned a retryable UpstreamFailoverError (502). Retrying
or switching accounts is futile for a content-policy block — it just burns
other accounts' quota and still surfaces an opaque 502 to the client.
Distinguish the two no-image cases:
(A) model text refusal -> 400 content_policy_violation, no retry, refusal
reason passed through to the client.
(B) truly empty response -> unchanged: retryable UpstreamFailoverError (502).
Adds extractOpenAIImagesModelRefusal (extracts the refusal text from
output_text.delta / message output_text, capped at 600 chars) and two unit
tests; the existing empty-response retry test is unaffected.