When an upstream API gateway (e.g. new-api) relays real Claude Code
requests, the User-Agent becomes Go-http-client while the body retains
the full Claude Code fingerprint (billing attribution block +
metadata.user_id + cache_control breakpoints).
Previously, the OAuth mimicry path relied solely on UA matching to
detect Claude Code clients. Without a matching UA, the gateway would
rewrite the system prompt — replacing the client's carefully structured
system blocks and cache_control breakpoints with its own injection.
This breaks Anthropic's prefix-based prompt cache: since the cache key
evaluates tools → system → messages in order, a changed system
invalidates all downstream message caching.
Symptoms observed:
- cache_read permanently locked at ~25K (only system prompt cached)
- cache_creation growing monotonically every turn (full messages rewrite)
- Single-request costs $17-27 instead of normal $1-2
Fix: when UA does not match but the body contains a valid billing
attribution block (x-anthropic-billing-header with cc_entrypoint=),
treat the request as proxied Claude Code traffic and skip mimicry.
This preserves the client's original system structure and cache_control
breakpoints, allowing Anthropic's prompt cache to function correctly.
Manual Grok connection tests previously surfaced upstream HTTP 402 errors without changing account availability. Persist the same 30-minute payment-required cooldown used by the live forwarding path so refreshed account lists no longer present the account as schedulable.
Constraint: Keep manual-test HTTP 402 handling aligned with existing Grok forwarding semantics.
Rejected: Mark the account permanently error | payment state can recover and the forwarding path intentionally uses a bounded cooldown.
Confidence: high
Scope-risk: narrow
Directive: Keep the manual-test cooldown reason and duration aligned with handleGrokAccountUpstreamError.
Tested: go test -tags=unit ./internal/service -count=1; go vet -tags=unit ./internal/service; production package compile check
Not-tested: Live xAI account with an exhausted subscription
Related: #4794
Go persists upstream_billing_probe.next_probe_at via RFC3339Nano, but
jsonpath datetime() parses at most 6 fractional digits, so every stored
timestamp failed to parse and was treated as malformed: fail-open due,
ordered into the invalid bucket by id ASC. With more enabled accounts
than the per-cycle limit, the same lowest IDs monopolized every cycle
and higher IDs were never probed again. Trim the fraction to
microseconds before datetime(), mirroring ListDueOllamaCloudUsageAccounts,
and pin the behavior with integration regressions: nanosecond parsing,
due-time ordering beyond the limit, preserved fail-open for truly
invalid dates.
All four pass on a quiet box and fail on a loaded one, each for its own
reason. None of them is testing the clock, so none of them should be
failing on it.
- ollama_cloud_usage_test.go: the second caller was released with
`close(release)` right after its goroutine was started, not after it
had reached the singleflight group. When the first refresh won that
race, the second became a new singleflight execution, re-read the
account, saw the LastAttemptAt the first one had just written, and
came back with the 30-second manual-refresh 429 at line 685. It now
counts account loads and waits for the second caller's own load,
which happens right before it joins the group. Adds a counting
GetByID to the test repo.
- gateway_hotpath_optimization_test.go: a 20ms sleep was meant to let
all 12 callers reach the cache before the loader was released. A
caller that arrived after the load had finished got a hit, not a
miss, so the miss count came out 11 of 12. It now waits on the miss
counter itself, which is the value the test asserts on.
- token_refresh_pool_health_test.go: the floor was `configuredSpacing`
minus 10ms, i.e. 40ms out of 50ms. Each start timestamp is taken
after the rate gate releases the goroutine, so scheduler delay can
compress one observed gap with the gate behaving correctly — seen at
37ms and again at 13ms. The floor is now a tenth of the configured
spacing. Measured with providerQPS=20, 8 attempts, concurrency 2:
gate at 50ms gives a minimum gap of 49.97ms, gate at 0 gives 22µs.
So an unpaced gate sits three orders of magnitude under the 5ms floor
and is still caught, while jitter has room to move. A comment warns
against replacing this with an assertion on the total span of the
starts: the span is set by how long each attempt takes under the
concurrency limit, not by the gate — 471ms paced against 241ms
unpaced — so a span check passes with the gate disabled.
- prompt_guard_test.go: the bound only has to show the failover shared
the first endpoint's 70ms deadline and did not take the second
endpoint's own 500ms one. An unshared deadline lands near 535ms, so
350ms still fails loudly (seen: 224ms against a 180ms bound).
Tests only; no product code is touched. Each fix was checked in both
directions: it passes with the behaviour intact, and it still fails when
the behaviour is broken on purpose (for the QPS one, by swapping the
shared rate gate for a zero-interval one).
PUT /api/v1/admin/settings is a whole-document write. The admin UI always sends
the complete document, so saving from the settings page is unaffected both
before and after this change. The bug is only reachable when an API client calls
the endpoint directly and sends just the fields it wants to change, which is the
natural assumption for a PUT on a settings resource.
Value-typed fields of UpdateSettingsRequest bind to their zero value when the
payload omits them, and buildSystemSettingsUpdates writes every key
unconditionally, so such a caller has no way to say "leave this one alone".
Omitting a field and explicitly clearing it are indistinguishable on the wire.
A caller that sends only the field it wants to change, e.g.
{"risk_control_enabled": true}
sets that flag and clears every other unguarded field in the same request.
Measured against a fully configured store, one such call empties site_name,
site_subtitle, api_base_url, contact_info and doc_url, and turns
registration_enabled, email_verify_enabled, invitation_code_enabled and
turnstile_enabled off. turnstile_enabled alone gates the captcha on login,
register, forgot-password and both verify-code endpoints, and
email_verify_enabled is a precondition of IsPasswordResetEnabled.
The damage is easy to miss. site_name has a built-in fallback, so
getStringOrDefault renders the cleared value as the default product name and the
login page visibly changes, while the toggles just go quiet. Reopening the
settings page reads the already-cleared state back into the form, so correcting
the one visible field and saving persists the rest of the damage.
Fields that grew their own guard already survive this: the SMTP block falls back
to the previous values when smtp_host arrives empty, secret fields are written
only when non-empty, and 132 request fields are pointers whose handler merges an
omitted field with the stored value. This generalizes that pattern rather than
adding a fourth ad-hoc guard.
The handler now decodes the payload a second time as a raw field map, resolves
the setting key each absent field would have written, and hands that set to the
service, which drops those keys before SetMultiple, so the stored value is never
touched. Fields the payload does carry are written as before, giving the caller
the partial-update semantics it was already assuming. The mapping is reflected off
the request's json tags so new fields are covered without maintaining a list;
smtp_from_email is the only field whose json name differs from its setting key
and is aliased explicitly.
Only value-typed fields are filtered. Pointer fields keep whole-document
behaviour on purpose: forwarded_client_ip_headers and
api_key_acl_trust_forwarded_ip depend on being rewritten on every save to
re-normalize fail-closed state, which the malformed forwarded-client-IP header
test pins down.
An explicitly sent empty value is still a deliberate clear; only absent fields
are preserved. A partial write refreshes the in-process caches from storage
instead of from the request struct, which holds zero values for whatever the
caller omitted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Registration duplicate checks compared the email verbatim (lowercase +
trim only), so a single inbox could spawn unlimited accounts via provider
alias features: plus addressing (user+tag@gmail.com) and the Gmail dot
trick (u.s.e.r@gmail.com) both deliver to the same mailbox but were seen
as distinct, verifiable addresses. This lets abusers bulk-register to farm
signup grants while the domain whitelist and email verification pass.
Add NormalizeEmailForAliasDedup to collapse these variants to a single
"inbox identity" and use it (via existsByEmailOrAlias) on the three local
email-registration paths: register, send-verify-code, and its async
variant. The exact ExistsByEmail check runs first; only on a miss do we
scan the candidate domains (gmail.com/googlemail.com are mutual aliases)
and compare normalized forms. Lookup errors fail closed, matching the
existing check, so the path cannot be bypassed by inducing errors.
Scoped to registration only — email storage, display, login, and delivery
are unchanged. OAuth-bound emails (already provider-verified) are out of
scope.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>