When an upstream API gateway (e.g. new-api) relays real Claude Code
requests, the User-Agent becomes Go-http-client while the body retains
the full Claude Code fingerprint (billing attribution block +
metadata.user_id + cache_control breakpoints).
Previously, the OAuth mimicry path relied solely on UA matching to
detect Claude Code clients. Without a matching UA, the gateway would
rewrite the system prompt — replacing the client's carefully structured
system blocks and cache_control breakpoints with its own injection.
This breaks Anthropic's prefix-based prompt cache: since the cache key
evaluates tools → system → messages in order, a changed system
invalidates all downstream message caching.
Symptoms observed:
- cache_read permanently locked at ~25K (only system prompt cached)
- cache_creation growing monotonically every turn (full messages rewrite)
- Single-request costs $17-27 instead of normal $1-2
Fix: when UA does not match but the body contains a valid billing
attribution block (x-anthropic-billing-header with cc_entrypoint=),
treat the request as proxied Claude Code traffic and skip mimicry.
This preserves the client's original system structure and cache_control
breakpoints, allowing Anthropic's prompt cache to function correctly.
Manual Grok connection tests previously surfaced upstream HTTP 402 errors without changing account availability. Persist the same 30-minute payment-required cooldown used by the live forwarding path so refreshed account lists no longer present the account as schedulable.
Constraint: Keep manual-test HTTP 402 handling aligned with existing Grok forwarding semantics.
Rejected: Mark the account permanently error | payment state can recover and the forwarding path intentionally uses a bounded cooldown.
Confidence: high
Scope-risk: narrow
Directive: Keep the manual-test cooldown reason and duration aligned with handleGrokAccountUpstreamError.
Tested: go test -tags=unit ./internal/service -count=1; go vet -tags=unit ./internal/service; production package compile check
Not-tested: Live xAI account with an exhausted subscription
Related: #4794
Go persists upstream_billing_probe.next_probe_at via RFC3339Nano, but
jsonpath datetime() parses at most 6 fractional digits, so every stored
timestamp failed to parse and was treated as malformed: fail-open due,
ordered into the invalid bucket by id ASC. With more enabled accounts
than the per-cycle limit, the same lowest IDs monopolized every cycle
and higher IDs were never probed again. Trim the fraction to
microseconds before datetime(), mirroring ListDueOllamaCloudUsageAccounts,
and pin the behavior with integration regressions: nanosecond parsing,
due-time ordering beyond the limit, preserved fail-open for truly
invalid dates.
All four pass on a quiet box and fail on a loaded one, each for its own
reason. None of them is testing the clock, so none of them should be
failing on it.
- ollama_cloud_usage_test.go: the second caller was released with
`close(release)` right after its goroutine was started, not after it
had reached the singleflight group. When the first refresh won that
race, the second became a new singleflight execution, re-read the
account, saw the LastAttemptAt the first one had just written, and
came back with the 30-second manual-refresh 429 at line 685. It now
counts account loads and waits for the second caller's own load,
which happens right before it joins the group. Adds a counting
GetByID to the test repo.
- gateway_hotpath_optimization_test.go: a 20ms sleep was meant to let
all 12 callers reach the cache before the loader was released. A
caller that arrived after the load had finished got a hit, not a
miss, so the miss count came out 11 of 12. It now waits on the miss
counter itself, which is the value the test asserts on.
- token_refresh_pool_health_test.go: the floor was `configuredSpacing`
minus 10ms, i.e. 40ms out of 50ms. Each start timestamp is taken
after the rate gate releases the goroutine, so scheduler delay can
compress one observed gap with the gate behaving correctly — seen at
37ms and again at 13ms. The floor is now a tenth of the configured
spacing. Measured with providerQPS=20, 8 attempts, concurrency 2:
gate at 50ms gives a minimum gap of 49.97ms, gate at 0 gives 22µs.
So an unpaced gate sits three orders of magnitude under the 5ms floor
and is still caught, while jitter has room to move. A comment warns
against replacing this with an assertion on the total span of the
starts: the span is set by how long each attempt takes under the
concurrency limit, not by the gate — 471ms paced against 241ms
unpaced — so a span check passes with the gate disabled.
- prompt_guard_test.go: the bound only has to show the failover shared
the first endpoint's 70ms deadline and did not take the second
endpoint's own 500ms one. An unshared deadline lands near 535ms, so
350ms still fails loudly (seen: 224ms against a 180ms bound).
Tests only; no product code is touched. Each fix was checked in both
directions: it passes with the behaviour intact, and it still fails when
the behaviour is broken on purpose (for the QPS one, by swapping the
shared rate gate for a zero-interval one).
Registration duplicate checks compared the email verbatim (lowercase +
trim only), so a single inbox could spawn unlimited accounts via provider
alias features: plus addressing (user+tag@gmail.com) and the Gmail dot
trick (u.s.e.r@gmail.com) both deliver to the same mailbox but were seen
as distinct, verifiable addresses. This lets abusers bulk-register to farm
signup grants while the domain whitelist and email verification pass.
Add NormalizeEmailForAliasDedup to collapse these variants to a single
"inbox identity" and use it (via existsByEmailOrAlias) on the three local
email-registration paths: register, send-verify-code, and its async
variant. The exact ExistsByEmail check runs first; only on a miss do we
scan the candidate domains (gmail.com/googlemail.com are mutual aliases)
and compare normalized forms. Lookup errors fail closed, matching the
existing check, so the path cannot be bypassed by inducing errors.
Scoped to registration only — email storage, display, login, and delivery
are unchanged. OAuth-bound emails (already provider-verified) are out of
scope.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Avoid recording the generic account-model transient cooldown for pool-mode statuses explicitly configured for same-account retry. This lets the bounded request-local retry budget complete while preserving cooldowns for other statuses and non-pool accounts.