- cc_pipeline: observe Chat Completions chunks/bodies as untyped payloads
(empty event type) so the upstream-echoed service_tier is trusted, matching
the upstream constraint that only terminal events and untyped bodies report
the actual processing tier.
- tests: add model field to terminal SSE frames (observation only triggers on
model-bearing frames) and assert response.created tier echo is ignored.
- openai_ws_forwarder_ingress: resolve billing tier from the local
upstreamResponseModelObserver (upstream echo first) instead of the raw
request payload, matching the HTTP->WS bridge and WS v2 forwarder.
- openai_ws_http_bridge_test: add fast-alias + upstream default case
proving the local observer's echoed tier wins.
- handler tests: keep only invalid service_tier -> 400 (short-circuits at
validation); valid/omitted semantics covered by the pure service-level
validation tests, avoiding real account selection in tests.
- Accept fast|priority (canonical priority), flex|auto|default|scale on
/v1/responses and /v1/chat/completions; reject unknown/empty/non-string
with HTTP 400; omitted and null stay compatible.
- Propagate service_tier through JSON/SSE, Responses<->Chat conversions,
fallback paths and HTTP->upstream WebSocket bridge.
- Billing prefers the upstream terminal tier; the outbound (policy-
transformed) tier is used only when upstream omits the field.
Explicit upstream default bills Standard even when Fast was requested.
- Pricing: Fast premium 2x Standard for gpt-5.6-sol/terra/luna and
gpt-5.4; 2.5x for gpt-5.5; channel FastMultiplier stays authoritative.
- Live verification (official Codex 0.149.0 + gateway, HTTP & WS):
upstream ChatGPT backend may return terminal default even when the
account catalog advertises priority; billing follows the actual tier.
DashScope/DeepSeek later tool_call deltas send empty id and
function.name. Clients that merge with !== undefined overwrite
the first delta's identity and dispatch unknown tool "". Drop
those empty fields on the raw Chat Completions SSE path.
DOMPurify <=3.3.1 (and the mermaid-transitive 3.3.3) carry ~18 disclosed
sanitizer-bypass/XSS advisories, including GHSA-cj63-jhhr-wcxv
(CVE-2026-65913): with USE_PROFILES enabled, ALLOWED_ATTR is rebuilt as a
plain array and looked up via ALLOWED_ATTR[lcName], so a polluted
Array.prototype property (e.g. onclick) is treated as an allow-listed
attribute and survives sanitization -- this app calls
DOMPurify.sanitize(svg, { USE_PROFILES: { svg: true, svgFilters: true } })
in src/utils/sanitize.ts, whose output is rendered via v-html in
ImageUpload.vue's SVG upload preview.
Bumped to 3.4.14 (latest, OSV-clean) and pinned via pnpm.overrides so the
mermaid-transitive copy dedupes to the same patched version instead of
staying pinned at 3.3.3. Lockfile-only regen via pnpm 9, no other package
changes.
A terminal event that arrives with an empty output was rebuilt from delta
accumulation. BufferedResponseAccumulator models only one reasoning item, one
message, and N function calls, and records no item id, status, or phase, so a
turn carrying several items collapsed into a single fabricated message: the
reasoning item disappeared, the real message id was replaced, and phase was
lost.
reconstructResponseOutputFromSSE already prefers the raw output_item.done
items over accumulation for buffered responses. The streaming path had no
equivalent because it never sees the whole body at once. Collect the raw item
of each output_item.done keyed by output_index and rebuild from those, falling
back to accumulation only when the stream reported no done item at all.
Items are stored as raw JSON, so vendor extensions and item types this gateway
does not model survive the rebuild verbatim.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>