Successful upstream billing probes already persist the upstream-declared
rate as a display-only snapshot. Add a per-account opt-in that writes
that declared rate back to the account's rate_multiplier, so the account
cost basis follows upstream repricing automatically instead of drifting
until an operator notices.
- new per-account flag upstream_billing_rate_sync_enabled stored next to
the probe flag in account extra: enabling sync force-enables the
probe, disabling the probe cascades sync off, and eligibility follows
IsUpstreamBillingProbeIdentity (tightened from any non-empty platform
to an explicit whitelist of the five supported API-key platforms so
future platforms do not silently inherit probe/sync semantics)
- only a successful probe whose declared rate survives validation
(finite, within bounds, not rounded to zero at the rate_multiplier
decimal(10,4) scale) writes back; failed/unsupported/invalid probes
leave rate_multiplier unchanged
- the writeback rides the existing snapshot CAS transaction:
UpdateUpstreamBillingProbeSnapshot takes an optional rateMultiplier
and applies it atomically with the snapshot under the same
identity/snapshot compare-and-swap, so a probe result observed on a
stale account cannot clobber a concurrent admin edit
- admin edit goes through UpdateWithAccountBillingSettings, which
applies the form without overwriting a rate that a probe synchronized
after the edit form was loaded (nil rateMultiplier = not edited);
once sync is enabled the edit form shows the rate as managed
- bulk update rejects a manual rate_multiplier change when any target
account has rate sync enabled (whole batch fails with a dedicated
error so partial writes cannot bypass the sync ownership)
- frontend: sync toggle with hints in the edit modal (probe/sync
enable/disable coupling enforced in the form), synced-rate tooltip on
the rate cell, bulk edit modal warns and blocks rate edits that hit
sync-enabled accounts; en/zh copy updated
- tests: service unit tests for sync gating/validation/cascade, sqlmock
repo tests for the extended CAS, real-PostgreSQL integration tests
(rate written only for successful+enabled accounts, manual rate
protected after sync disabled, admin edit preserved across concurrent
probe sync), handler/API contract updates, frontend specs for modal
coupling, bulk rejection and rate cell
coder/websocket arms a context.AfterFunc that hard-closes the connection
when a write context is canceled, and AfterFunc stop does not wait for a
callback that already started. An external cancellation (e.g. ingress
lease loss) landing inside the disarm window of an already-successful
downstream write could therefore kill the TCP connection before the
retry close frame (1013) was written, leaving the client with a bare
EOF. Mirror the read side: bound downstream writes with the write
timeout only, and rely on the explicit Close/CloseNow performed by every
relay exit path for teardown.
Fixes the flaky TestPassthroughLifecycle_LeaseLossSendsRetryClose.
/v1/sub2api/billing is a key-scoped sub2api convention: any API-key
account whose base_url points at a sub2api-compatible upstream answers
it regardless of the account platform. Widen probe eligibility from
platform=openai to every API-key account (OAuth/Bedrock stay excluded:
no static key to present).
- central predicate exported as IsUpstreamBillingProbeIdentity; runner,
manual probe, SetAccountEnabled, admin create/update/bulk validation
and CRS reconcile all follow it
- probe target resolution reads credentials.api_key/base_url directly.
OpenAI keeps its official-default base URL and openai transport
profile. Other platforms whose base_url is empty or points at an
official provider API domain (anthropic.com, googleapis.com, x.ai,
grok.com, openai.com - matched on the normalized hostname, port and
trailing dot stripped, as the exact host or any subdomain) persist
"unsupported" without sending a request: the create form fills empty
base_url with official defaults and offers official regional presets
(e.g. us-east-1.api.x.ai), and official APIs cannot answer
/v1/sub2api/billing, so probing would only send the account key to a
nonexistent official path
- due-scan SQL, BulkUpdate probe WHERE, UpdateCredentials stale-snapshot
CASE and proxy-change invalidation drop their platform filters
(type='apikey' retained)
- frontend: rate cell, edit/create/bulk modals gate on type==='apikey';
the antigravity upstream create flow (its own helper and form section)
shows the auto-probe toggle and passes upstream_billing_probe_enabled;
bulk WS-mode section stays OpenAI-only; settings copy de-scoped
- tests: multiplatform service unit tests (relay success, official/empty
base_url unsupported without request, normalized official-host matrix,
OpenAI defaults preserved), real-PostgreSQL due-scan coverage for
openai/anthropic/grok plus oauth/disabled exclusion, updated
sqlmock/CRS/frontend specs; probe stays opt-in per account
Fixes#5148: with aggregator upstreams (e.g. newapi) that end SSE
streams without a proper terminal event, every such request was
silently missing from usage logs and billing.
Root cause (tracked via the nested audit issue): the low-level
Anthropic SSE readers already return the partially collected usage
together with the stream error (missing terminal event, read error,
interval timeout), but Forward converted every such result to
(nil, err) and the handler returned before submitting RecordUsage.
Changes:
- Add partialStreamUsageResult: on stream errors, wrap observed usage
into a ForwardResult and return it alongside the error, for both the
regular Anthropic path and the API-key passthrough path. Invariants:
UpstreamFailoverError always keeps result=nil (failover retries are
billed as the successful attempt, never twice), and zero observed
usage returns no partial result (no phantom zero-usage records).
- Messages handler: hoist the usage submission block into a closure
shared by the success path and the new partial-result error path.
- Usage record worker pool: distinguish pool-stopped drops
(dropped_stopped) from operator-configured drop/sample overflow
drops; billing tasks now fall back to inline synchronous execution
only during the shutdown window, while explicit drop/sample overflow
semantics are preserved. Image usage keeps its mandatory fallback
for both drop kinds via the new mode.Dropped() helper.
Tests: Forward-level regressions for missing-terminal / read-error /
no-usage / failover-invariant on both paths, plus handler-level
stopped-pool sync fallback and drop-policy preservation tests.
Real claude-cli/2.1.220 auto-mode classifier requests carry two system
entries: the security-monitor prompt plus an appended session-context
block. The previous len(systemEntries) != 1 guard rejected them before
any content check ran, so claude_code_only groups kept refusing the
classifier (#5152, follow-up to #5041/#5048).
Scan every entry for the monitor prompt instead of requiring exactly
one. Discrimination is unchanged: the matching entry still needs the
10k-char minimum, the fixed prefix, and all eight markers.
OpenAI Responses may return code=rate_limit_exceeded in a response.failed SSE event while the HTTP status remains 200. Classify these failures as 429 so configured pool-mode retries and account failover are applied. No account-level rate-limit state is written on this path: the 200-stream response headers carry normal quota snapshots, and retry semantics stay owned by the failover engine.
Ignore compact keepalive bytes when determining whether semantic output has started, preserving safe retries before real output.
Implements #2646: the risk-control content audit can now send OpenAI
Moderations requests through a proxy from IP Management - Proxy Servers.
Backend:
- ContentModerationConfig gains proxy_id (nil = direct, unchanged default)
- update semantics: null keeps, 0 clears, >0 selects (validated to exist)
- moderation calls build the client via the shared httpclient pool; proxy
resolution failure surfaces as a moderation error and never silently
falls back to direct connection
- proxy_id -> URL resolution cached 60s (single-entry, invalidated on
config save) so the pre-block hot path does not hit the DB per request
- test-key endpoint accepts proxy_id too (null = saved config's proxy,
0 = force direct), so input-key/saved-key tests exercise the same path
- proxy usage/inactivity logged (content_moderation.proxy_enabled /
proxy_not_active) without leaking credentials
Frontend:
- ProxySelector in the risk-control basic settings tab, proxy list loaded
non-blockingly; save and test payloads carry proxy_id; zh/en i18n
- UseTLS now tries implicit TLS first (port 465 semantics) and, when the
server answers in plaintext (tls.RecordHeaderError, e.g. port 587
submission), automatically retries with mandatory STARTTLS; encryption
is never silently downgraded (fixes#1470, supersedes #1488)
- TestSMTPConnectionWithConfig now shares connectSMTP with the send path,
adding the opportunistic STARTTLS upgrade the send path gained in
b402c367d; this removes the 'test connection fails but test email
sends' mismatch reported in #1488
- test-connection now also honors dial/IO timeouts and ignores
non-standard QUIT responses, matching the send path
Reason:
- Responses retries can carry account-bound encrypted compaction items that OpenAI rejects with invalid_encrypted_content.
Changes:
- Drop encrypted compaction and compaction_summary items only during the existing recovery retry.
- Preserve unencrypted compaction items and cover HTTP and WebSocket recovery paths.