Fixes#5148: with aggregator upstreams (e.g. newapi) that end SSE
streams without a proper terminal event, every such request was
silently missing from usage logs and billing.
Root cause (tracked via the nested audit issue): the low-level
Anthropic SSE readers already return the partially collected usage
together with the stream error (missing terminal event, read error,
interval timeout), but Forward converted every such result to
(nil, err) and the handler returned before submitting RecordUsage.
Changes:
- Add partialStreamUsageResult: on stream errors, wrap observed usage
into a ForwardResult and return it alongside the error, for both the
regular Anthropic path and the API-key passthrough path. Invariants:
UpstreamFailoverError always keeps result=nil (failover retries are
billed as the successful attempt, never twice), and zero observed
usage returns no partial result (no phantom zero-usage records).
- Messages handler: hoist the usage submission block into a closure
shared by the success path and the new partial-result error path.
- Usage record worker pool: distinguish pool-stopped drops
(dropped_stopped) from operator-configured drop/sample overflow
drops; billing tasks now fall back to inline synchronous execution
only during the shutdown window, while explicit drop/sample overflow
semantics are preserved. Image usage keeps its mandatory fallback
for both drop kinds via the new mode.Dropped() helper.
Tests: Forward-level regressions for missing-terminal / read-error /
no-usage / failover-invariant on both paths, plus handler-level
stopped-pool sync fallback and drop-policy preservation tests.
Real claude-cli/2.1.220 auto-mode classifier requests carry two system
entries: the security-monitor prompt plus an appended session-context
block. The previous len(systemEntries) != 1 guard rejected them before
any content check ran, so claude_code_only groups kept refusing the
classifier (#5152, follow-up to #5041/#5048).
Scan every entry for the monitor prompt instead of requiring exactly
one. Discrimination is unchanged: the matching entry still needs the
10k-char minimum, the fixed prefix, and all eight markers.
OpenAI Responses may return code=rate_limit_exceeded in a response.failed SSE event while the HTTP status remains 200. Classify these failures as 429 so configured pool-mode retries and account failover are applied. No account-level rate-limit state is written on this path: the 200-stream response headers carry normal quota snapshots, and retry semantics stay owned by the failover engine.
Ignore compact keepalive bytes when determining whether semantic output has started, preserving safe retries before real output.
Implements #2646: the risk-control content audit can now send OpenAI
Moderations requests through a proxy from IP Management - Proxy Servers.
Backend:
- ContentModerationConfig gains proxy_id (nil = direct, unchanged default)
- update semantics: null keeps, 0 clears, >0 selects (validated to exist)
- moderation calls build the client via the shared httpclient pool; proxy
resolution failure surfaces as a moderation error and never silently
falls back to direct connection
- proxy_id -> URL resolution cached 60s (single-entry, invalidated on
config save) so the pre-block hot path does not hit the DB per request
- test-key endpoint accepts proxy_id too (null = saved config's proxy,
0 = force direct), so input-key/saved-key tests exercise the same path
- proxy usage/inactivity logged (content_moderation.proxy_enabled /
proxy_not_active) without leaking credentials
Frontend:
- ProxySelector in the risk-control basic settings tab, proxy list loaded
non-blockingly; save and test payloads carry proxy_id; zh/en i18n
- UseTLS now tries implicit TLS first (port 465 semantics) and, when the
server answers in plaintext (tls.RecordHeaderError, e.g. port 587
submission), automatically retries with mandatory STARTTLS; encryption
is never silently downgraded (fixes#1470, supersedes #1488)
- TestSMTPConnectionWithConfig now shares connectSMTP with the send path,
adding the opportunistic STARTTLS upgrade the send path gained in
b402c367d; this removes the 'test connection fails but test email
sends' mismatch reported in #1488
- test-connection now also honors dial/IO timeouts and ignores
non-standard QUIT responses, matching the send path
Reason:
- Responses retries can carry account-bound encrypted compaction items that OpenAI rejects with invalid_encrypted_content.
Changes:
- Drop encrypted compaction and compaction_summary items only during the existing recovery retry.
- Preserve unencrypted compaction items and cover HTTP and WebSocket recovery paths.
glm-5.2 has no entry in the fallback table, so getFallbackPricing falls
through to `strings.Contains(modelLower, "glm-5")` and prices it at
GLM-5 rates ($1.00 in / $3.20 out per MTok) instead of the official
z.ai rates ($1.40 / $4.40) — roughly 27% under.
LiteLLM carries no bare `glm-5.2` key either (only provider-prefixed
`cloudflare/@cf/zai-org/glm-5.2` and `fireworks_ai/.../glm-5p2`, which
the lookup candidates never match), so the request always lands on the
fallback path and the discrepancy shows up directly in usage logs.
Add the glm-5.2 entry (same price as glm-5.1 per docs.z.ai) and match it
before the bare `glm-5` branch, with a note that dotted variants must
precede it. The existing regression test asserting the old glm-5 price
is updated accordingly.
Source: https://docs.z.ai/guides/overview/pricing
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The count_tokens path now strips max_tokens; this sibling test also
exercises ForwardCountTokens but still asserted max_tokens preservation.
Flip its assertion to expect the field is filtered, matching
CountTokensFiltersGenerationFields.
- Move pool mode check before the status switch so 401/402/403/5xx
all skip tempUnschedule consistently
- Skip rateLimitGrok in updateGrokUsageSnapshot for pool mode
- Expand test coverage to all affected status codes
The proxy stream circuit introduced in v0.1.164 (#4749) removes every
account behind a quarantined proxy from scheduling. When all schedulable
accounts share one proxy (a common deployment), two mid-stream
disconnects within a minute zeroed out capacity for 10 minutes and every
request failed with 502. One HTTP/2 connection loss also killed all
multiplexed streams at once, tripping the threshold from a single event.
- Quarantine now degrades to a preference: when the only reason no
account is available is proxy quarantine, selection retries once with
the quarantine bypassed, so capacity can never reach zero.
- Disconnects within 3s per proxy collapse into one failure event.
- Add gateway.openai_proxy_stream_circuit.disabled escape hatch.
- A completed stream still clears the quarantine immediately; TTL,
thresholds and recording guards are unchanged.