The passthrough streaming loop buffers leading events in `pendingLines` until
the first visible output, so a client receives **zero bytes** — not even the
HTTP response headers, since gin's ResponseWriter only commits them on the
first write — for as long as the upstream takes to produce its first visible
event.
The Forward path already recognises this exact hazard and solves it with a
downstream keepalive:
// Track downstream writes separately from upstream reads: pre-output
// failover can buffer response.created / response.in_progress, so
// keepalive must be based on downstream idle time.
The passthrough path has the same buffering but never starts a keepalive.
Reasoning models routinely think for several minutes before emitting a visible
event, so intermediate proxies close the connection on idle timeout.
Production evidence from one deployment (12h window, gpt-5.6 family):
* 44 `/v1/responses` requests produced their first visible output only after
600–900 s, each carrying 30k–50k output tokens — the upstream had finished
the work.
* Every one of them was killed by an intermediate nginx `proxy_read_timeout`
(600 s) and returned 504. The user received nothing.
* Cutting early does not even save resources: the code deliberately keeps
draining the upstream for usage after the client disconnects.
Fix: start the existing keepalive helper for the passthrough stream as well.
* `StartOpenAICompactSSEKeepalive` keeps its compact-marker gate; the gate-free
body is extracted as `startOpenAISSEKeepalive` for callers that already know
they are in an SSE context.
* The passthrough loop starts it when `gateway.stream_keepalive_interval > 0`
and stops it immediately before its first real write, so the ResponseWriter
is never shared.
Pre-output failover is unaffected: keepalive bytes are already excluded by
`OpenAICompactKeepaliveAdjustedWrittenSize`, which `openAIStreamClientOutputStarted`
on this very path already consults (#3887). No pendingLines and no
account-specific headers are ever emitted by a keepalive.
Behaviour is unchanged when `gateway.stream_keepalive_interval` is 0.
Public groups have always been bindable by every user: CanBindGroup returned
true for any non-exclusive group, and user_allowed_groups only ever carried the
exclusive groups an admin had granted. Admins had no way to hand a single user
a subset of the public groups short of converting a group to exclusive, which
changes it for everyone already using it.
A user now carries restrict_public_groups. While it is false, which is the
default and what every existing row migrates to, nothing changes: every public
group stays bindable. Once an admin turns it on for a user, that user's public
groups are narrowed to the ones listed in user_allowed_groups, the same table
that already gates exclusive groups.
The flag is an administrative control, so it rides on the admin user DTO only
and leaves the shape of the end-user endpoints alone.
The model plaza filter honours the flag as well, so a restricted user is not
shown groups they would be refused when binding a key. Anonymous visitors have
no user record and keep the previous view.
Enforcement rides on the existing CanBindGroup choke point, so it covers key
creation, key updates, and per-request authorization together. An API key bound
to a group that is later withdrawn stops working at request time rather than
lingering as a key that can be listed but not used.
The admin dialog gains a toggle over the public group list. Turning it off
re-checks every public group, so an admin cannot save a list that reads as
restrictive while the restriction itself is disabled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
When running a custom build or pre-release version (e.g. v0.1.183-custom),
parseVersion failed to parse the patch segment because of the non-numeric suffix,
causing compareVersions to incorrectly report an available update.
This commit strips any hyphenated suffix before splitting SemVer components.
Signed-off-by: WUJI-Labs <zoujincool@gmail.com>
Merge origin/main into contrib/routed-codex-model-catalog.
Single conflict in openai_gateway_forward.go: main added
ClearActualOpenAIUpstreamEndpoint + SetActualOpenAIUpstreamEndpoint
(commit 4795650d2) at the same insertion point where the PR added
filterOpenAIResponsesNoneReasoningEffortForAccount. Both changes
are independent — endpoint observability first, then body filtering.
Copy cached API-key manifests before group-specific mutation, use DeepSeek
model IDs for Codex fallbacks, omit unsupported config.toml effort, and
drop wildcard mapping keys from generated catalogs.
Co-authored-by: Cursor <cursoragent@cursor.com>