Review follow-ups on the reset-credit caching flow:
- Recover account state BEFORE (and independently of) the reset-credit
display cache. A failed cache refresh could previously abort the run and
leave the account rate-limited — the very reason the credit was spent
(#3672 / #3740). The recovered account row is now returned even when the
cache refresh fails.
- Run the post-reset bookkeeping on a detached, time-boxed context and give
the panel reset call a larger timeout. A client abort no longer strands a
consumed (non-refundable) credit with an unrecovered account, and the
chained upstream calls can no longer exceed the client timeout and invite a
retry that spends a second credit.
- Persist the reset-credit snapshot through POST /accounts/:id/quota/refresh
instead of a side-effecting GET flag, so the write is covered by the audit
middleware. A rejected snapshot write now degrades to cache_persisted=false
instead of turning a successful upstream read into a 502 that left the card
without a credit count and the reset button permanently disabled.
- Reject snapshots whose positive count carries no expiration entries, and
drop expired credits (clamping the count) when rehydrating, so a stale
cache can no longer light up the reset button.
- Keep nil quota / rate-limit services nil in the handler's interface fields;
storing a nil *Service made the "not enabled" guards non-nil.
- Time-box the usage-refresh suppression and reuse handleAccountUpdated so the
patched row also enters the auto-refresh silent window.
AdminRefundDialog already implements a `requireForce` prop that renders the
force checkbox and blocks submit until it is checked, but AdminOrdersView
never consumed `require_force` from the refund response nor bound the prop,
so the checkbox could never render. Admins hitting a require-force case (for
example a user who spent their balance after requesting the refund) only saw
an error toast and had no way to complete the refund from the UI.
Keep the dialog open when the backend answers `require_force`, surface the
checkbox and the backend warning, and reset the flag whenever the dialog is
opened or closed.
Verified with `npm run typecheck` (vue-tsc --noEmit, clean).
Group pricing (rate multiplier, peak windows, per-user overrides) and
account cost (accounts.rate_multiplier) already live side by side, but
nothing stops the scheduler from handing a request to an account whose
cost multiplier exceeds what the group's pricing can profitably serve.
Add an opt-in per-group profit gate that filters scheduling candidates
by a margin rule, while ordering, scoring, stickiness and breakers keep
working unchanged among qualified accounts.
Admission rule: an account qualifies iff U <= D * (1 - min_margin -
safety_buffer) within a small relative epsilon, where U is
accounts.rate_multiplier (0 is legal; missing/negative/NaN/Inf are
conservatively rejected as invalid) and D is the requester's effective
downstream multiplier (user-group override ?? group default, times the
group peak factor) frozen at the request's pricing instant.
- groups gain profit_control_enabled / profit_min_margin /
profit_safety_buffer (migration 191); the durable auth-cache
invalidation trigger additionally watches the profit and pricing
columns (migration 192) so out-of-band group edits cannot leave
stale auth snapshots; GetByKeyForAuth explicitly projects the new
columns and the API-key auth snapshot version is bumped to force a
refresh of pre-existing snapshots
- request-level pricing instant: token entry points install pricingAt
into ctx; the profit threshold D and the RecordUsage peak factor
read the same instant, so one request never changes price mid-flight
across waits/retries/failover (media and unwired paths keep the
existing record-time semantics)
- the gate covers token requests on openai, anthropic, gemini, grok
and antigravity groups: OpenAI-family handlers via
WithOpenAIRequestPricingContext (responses incl. WS bridge, chat
completions, messages, embeddings, alpha search), the shared gateway
via WithGatewayTokenRequestPricing (messages, chat completions,
responses, gemini model actions); composite groups cannot enable it
directly; image/video/models/usage/count_tokens stay ungated and an
explicit image-generation intent suppresses the gate end to end
- post-slot recheck: after a slot is acquired the account is re-read
via SchedulerSnapshotService.GetAccount (scheduler cache, then DB;
only when both fail the check fails open with WARN + metric); a
vetoed account releases its slot and joins the request's exclusion
set for reselection; sticky bindings are written only after the
final check passes, and an over-threshold sticky account is skipped,
not deleted, so it comes back once its rate recovers
- sticky-session cache contract: GatewayCache.GetSessionAccountID now
returns ErrStickySessionNotFound on a miss (mapped from redis.Nil in
the repository implementation, mirroring ErrRefreshTokenNotFound) so
the profit sticky path can distinguish "no binding yet" from a real
read failure without importing the cache driver in service code
- cross-group re-entry (composite parent -> member group) resolves the
gate against the member group and clears a stale parent gate instead
of letting a foreign threshold veto accounts
- per-platform/group activity counters (installs, threshold vetoes,
invalid-rate vetoes, refresh failures) for observability
- admin UI: profit-control section on the five platforms' group forms
with percent input, validation and platform-switch reset; group
create/update/duplicate normalize and validate the config at a
single choke point
- cmd/profit-preview: offline what-if tool that replays the production
admission semantics over an exported config/account/override/model
dump, reports per-model admitted-account counts under the default
and the worst-case (lowest user override) D, and surfaces probe-sync
staleness as warnings without affecting admission
Tests: service unit coverage for gate resolution/veto/threshold
epsilon/pricing instant/suppress marker/scheduler filtering and
post-slot recheck (incl. -race on the profit surface), unit-tagged
handler slot-recheck and capability-mapping regressions, sqlmock and
real-PostgreSQL integration regressions for the GetByKeyForAuth
projection and the migration-192 trigger watch list, API contract
update, and frontend specs for the five-platform form helpers.
Successful upstream billing probes already persist the upstream-declared
rate as a display-only snapshot. Add a per-account opt-in that writes
that declared rate back to the account's rate_multiplier, so the account
cost basis follows upstream repricing automatically instead of drifting
until an operator notices.
- new per-account flag upstream_billing_rate_sync_enabled stored next to
the probe flag in account extra: enabling sync force-enables the
probe, disabling the probe cascades sync off, and eligibility follows
IsUpstreamBillingProbeIdentity (tightened from any non-empty platform
to an explicit whitelist of the five supported API-key platforms so
future platforms do not silently inherit probe/sync semantics)
- only a successful probe whose declared rate survives validation
(finite, within bounds, not rounded to zero at the rate_multiplier
decimal(10,4) scale) writes back; failed/unsupported/invalid probes
leave rate_multiplier unchanged
- the writeback rides the existing snapshot CAS transaction:
UpdateUpstreamBillingProbeSnapshot takes an optional rateMultiplier
and applies it atomically with the snapshot under the same
identity/snapshot compare-and-swap, so a probe result observed on a
stale account cannot clobber a concurrent admin edit
- admin edit goes through UpdateWithAccountBillingSettings, which
applies the form without overwriting a rate that a probe synchronized
after the edit form was loaded (nil rateMultiplier = not edited);
once sync is enabled the edit form shows the rate as managed
- bulk update rejects a manual rate_multiplier change when any target
account has rate sync enabled (whole batch fails with a dedicated
error so partial writes cannot bypass the sync ownership)
- frontend: sync toggle with hints in the edit modal (probe/sync
enable/disable coupling enforced in the form), synced-rate tooltip on
the rate cell, bulk edit modal warns and blocks rate edits that hit
sync-enabled accounts; en/zh copy updated
- tests: service unit tests for sync gating/validation/cascade, sqlmock
repo tests for the extended CAS, real-PostgreSQL integration tests
(rate written only for successful+enabled accounts, manual rate
protected after sync disabled, admin edit preserved across concurrent
probe sync), handler/API contract updates, frontend specs for modal
coupling, bulk rejection and rate cell
/v1/sub2api/billing is a key-scoped sub2api convention: any API-key
account whose base_url points at a sub2api-compatible upstream answers
it regardless of the account platform. Widen probe eligibility from
platform=openai to every API-key account (OAuth/Bedrock stay excluded:
no static key to present).
- central predicate exported as IsUpstreamBillingProbeIdentity; runner,
manual probe, SetAccountEnabled, admin create/update/bulk validation
and CRS reconcile all follow it
- probe target resolution reads credentials.api_key/base_url directly.
OpenAI keeps its official-default base URL and openai transport
profile. Other platforms whose base_url is empty or points at an
official provider API domain (anthropic.com, googleapis.com, x.ai,
grok.com, openai.com - matched on the normalized hostname, port and
trailing dot stripped, as the exact host or any subdomain) persist
"unsupported" without sending a request: the create form fills empty
base_url with official defaults and offers official regional presets
(e.g. us-east-1.api.x.ai), and official APIs cannot answer
/v1/sub2api/billing, so probing would only send the account key to a
nonexistent official path
- due-scan SQL, BulkUpdate probe WHERE, UpdateCredentials stale-snapshot
CASE and proxy-change invalidation drop their platform filters
(type='apikey' retained)
- frontend: rate cell, edit/create/bulk modals gate on type==='apikey';
the antigravity upstream create flow (its own helper and form section)
shows the auto-probe toggle and passes upstream_billing_probe_enabled;
bulk WS-mode section stays OpenAI-only; settings copy de-scoped
- tests: multiplatform service unit tests (relay success, official/empty
base_url unsupported without request, normalized official-host matrix,
OpenAI defaults preserved), real-PostgreSQL due-scan coverage for
openai/anthropic/grok plus oauth/disabled exclusion, updated
sqlmock/CRS/frontend specs; probe stays opt-in per account
Implements #2646: the risk-control content audit can now send OpenAI
Moderations requests through a proxy from IP Management - Proxy Servers.
Backend:
- ContentModerationConfig gains proxy_id (nil = direct, unchanged default)
- update semantics: null keeps, 0 clears, >0 selects (validated to exist)
- moderation calls build the client via the shared httpclient pool; proxy
resolution failure surfaces as a moderation error and never silently
falls back to direct connection
- proxy_id -> URL resolution cached 60s (single-entry, invalidated on
config save) so the pre-block hot path does not hit the DB per request
- test-key endpoint accepts proxy_id too (null = saved config's proxy,
0 = force direct), so input-key/saved-key tests exercise the same path
- proxy usage/inactivity logged (content_moderation.proxy_enabled /
proxy_not_active) without leaking credentials
Frontend:
- ProxySelector in the risk-control basic settings tab, proxy list loaded
non-blockingly; save and test payloads carry proxy_id; zh/en i18n
- Add compact_home_enabled setting to provide a minimal landing page
- Preserves custom home_content priority over compact mode
- Renders only site identity, navigation, and login/dashboard link
- Avoids marketing copy that triggers anti-fraud systems
- Includes focused backend/frontend tests and i18n support
Fixes#5065
The PASSKEY_DISABLED silence guard compared the string error code
against error.code, but the api client puts the numeric envelope code
there and the string code in error.reason, so the guard never matched
and every /profile visit on deployments without WebAuthn configured
showed a spurious "failed to load passkeys" toast.
Read error.reason instead, and skip the credentials request entirely
when the feature is disabled so the card no longer issues a request
that is guaranteed to fail with 403.
A hijacked session must not be able to silently add a passkey as a
persistent backdoor or remove the victim's credentials. Registration
(begin) and deletion now verify the account password server-side,
reusing the existing PASSWORD_REQUIRED / PASSWORD_INCORRECT errors.
The password is used instead of TOTP step-up so the guard also protects
deployments that never configured a TOTP encryption key. The password
key in both request bodies is covered by the audit middleware's
key-substring redaction, so no credential material reaches audit_logs.
Frontend: the add-passkey form gains a current-password field, and the
delete confirmation is now a dialog with a password input (replacing
window.confirm), mirroring the TOTP disable dialog. Backend error
messages (e.g. wrong password) are surfaced instead of the generic
failure toast. Rename remains password-free as it is cosmetic.
Administrators often need exact model identifiers while editing account whitelists. Add a dedicated copy action without changing model selection or upstream sync behavior.
Constraint: Keep the contribution frontend-only and avoid model routing or persistence changes
Confidence: high
Scope-risk: narrow
Reversibility: clean
Directive: Keep copy and selection as separate actions
Tested: focused Vitest, account component regression tests, frontend typecheck, lint, and production build
Not-tested: Authenticated browser screenshot
Related: Wei-Shaw/sub2api#2151
Flattening Codex `type:"namespace"` tool declarations into `namespace__name`
renames the tool but cannot rewrite the contract the client already handed the
model: Codex tells it to call `to=functions.collaboration.spawn_agent`. Under
`tool_mode: code_mode_only` (gpt-5.6-*) collaboration tools are the model's only
direct channel, so the rename left it unable to address them — it fell back to
`functions.exec` and looped on meaningless shell commands until the session
collapsed, with no upstream error anywhere. Fixes#4978.
Preserving is what the upstream actually wants. OAuth egress is unconditionally
chatgpt.com/backend-api/codex/responses (base_url is read only for API-key
accounts), i.e. the party that defines the extension; codex-rs builds one
request for both WS and HTTP and falls back WS->HTTP mid-session, so these
declarations already reach that host over HTTP — as they do today via the
responses-lite carrier and the WS/HTTP bridge, neither of which ever flattened.
Namespace names cannot be allowlisted: `features.multi_agent_v2.tool_namespace`
is user-configurable (openai/codex#31864 recommends renaming it to `agents`) and
MCP/connector namespaces are generated at runtime (mcp__codex_apps__gmail).
So the default is inverted rather than extended with more names.
- Preserve namespace declarations by default for OpenAI OAuth; the new
per-account `extra.openai_responses_flatten_namespaces` restores the previous
behavior for deployments routing OAuth traffic to a relay that rejects them.
- Keep `input[].namespace` on tool-call items for OAuth non-compact requests.
The upstream requires the round-trip ("Missing namespace for function_call
'...'. Round-trip the model's function_call item with its namespace field
included."), so preserving declarations while stripping the calls would 400 on
the second turn. The item-type allowlist mirrors the existing reactive strip.
- Keep compact on its current behavior: that endpoint rejects the field outright
("Unknown parameter: 'input[894].namespace'", #4761) and there is no evidence
either way about namespace declarations there, so this change does not widen
its surface. API-key accounts are untouched — their upstream is a standard
Responses API and the reactive retry only clears one index per round trip.
- Clear the flatten mapping at the start of each Forward attempt so an account
reached through failover cannot restore responses with the previous account's
mapping.
Admin UI gets an OAuth-only toggle (create/edit/bulk) plus zh/en copy noting the
compact exception.