IanShaw027
5315896b30
fix(grok): align free soft-gate tests with async fail-open cache
...
Treat cacheTTL=0 as non-expiring known entries, always store negative
refresh markers, and update sticky getSchedulableAccount tests to expect
first-hit fail-open then block after background stats warm.
2026-08-08 23:38:50 +08:00
IanShaw027
a54a4b674b
fix(grok): clear golangci errcheck and gofmt on free-quota path
...
Check the sync.Map type assertion in free-quota refresh coalescing, and
gofmt migration checksum rules plus prompt-audit route map alignment.
2026-08-08 22:52:31 +08:00
IanShaw027
7eb1310701
fix(grok): close free-by-default billing and related review blockers
...
H1/H2: bill search and voice with code defaults when group prices are
nil (explicit 0 remains free); bump API key auth snapshot to v19 and
refresh incomplete media/search/audio projections.
M1–M6: free-quota soft gate fails open on cache miss with background
refresh and 60s default TTL; correct password_auth config docs; default
cross-client model map to true (→ grok-4.5); audit /tts and /web_search;
exclude composite from migration 220 video-price clears; never let a
search surcharge mask token pricing failures.
2026-08-08 14:39:22 +08:00
IanShaw027
cec922d335
fix(grok): clear golangci-lint findings on complete-integration branch
...
Check Close/CloseNow errors, drop unused helpers and dead constants,
lowercase ST1005 error strings, and stop discarding unwrap status as an
unused assignment so CI golangci-lint passes.
2026-08-08 12:56:40 +08:00
IanShaw027
1f58e25ab3
Merge upstream/main into feat/grok-complete-integration
...
冲突集中在 chat completions / messages 两条 Responses 转发路径:
upstream 给 OpenAIForwardResult 增加了 UpstreamResponseModel 与
UpstreamResponseModelConflict(配套 beginUpstreamResponseModelObservation
观测器),本分支在同样位置把返回值改成了具名变量以便挂 Grok 原生搜索计数。
两侧不互斥,合并结果同时保留上游的响应模型观测字段与 Grok SearchCount 逻辑。
frontend/pnpm-lock.yaml 取 upstream 版本:package.json 与 upstream 完全一致,
本地差异只是 pnpm install 的重解析噪音。
2026-08-08 11:12:56 +08:00
IanShaw027
07b46e93e4
fix(grok): 修正 voice 路由推导、视频价归一化与门禁缓存驻留
...
- custom-voices endpoint 改由匹配到的路由模板推导(c.FullPath())。
原实现按请求 URL 字面后缀判 /audio,voice_id 恰为 "audio" 时
GET /custom-voices/audio 会被改写成 custom-voices/audio/audio,
把档案查询变成音频下载。补 voice_id="audio" 的路由推导测试。
- NormalizeVideoModelPrices 不再把无法识别的分辨率静默折算成 480p:
新增 LookupVideoBillingResolution 报告未知档位,配置解析路径丢弃并告警,
运行时计费仍走 OrDefault 兜底。model/tier 两层遍历改为排序遍历,
多个别名收敛到同一 family 时结果不再随 Go map 顺序漂移,冲突单价告警。
- grok free quota 软门禁缓存新增过期淘汰。条目按 account_id 键控且只写不删,
账号下线后会驻留至进程结束;淘汰只挂在已受 TTL 约束的查询路径上,
不影响缓存命中热路径。
- 迁移 220 清空非 Grok 分组视频价前先落快照表 groups_video_price_backup_220,
原 UPDATE 不可回滚。
- 简化 grok_search_count 两处等价冗余的事件类型分支。
2026-08-08 11:02:22 +08:00
Wesley Liddick
cc67b1aca1
Merge pull request #5406 from bestony/fix/openai-oauth-routing-hints
...
fix(openai): forward OAuth routing hints
2026-08-08 10:58:53 +08:00
IanShaw027
f3bac4619e
feat(keys): expand Grok client samples and tune free soft-gate default
...
Ship Use Key templates that match Grok Build / Codex best practice: env
vars + multi-model config.toml with api_backend=responses, env_key over
hardcoded secrets, and clearer shell/path guidance for Claude/Codex/OpenCode.
Also set free_quota_token_limit default to 500k (24h soft-gate), clean up
personal-dev-only comments, and keep billing test fixtures aligned.
2026-08-08 10:26:16 +08:00
IanShaw027
e01ce90d47
fix(grok): harden voice request ids, video pending, and search pricing alerts
...
Mint durable grok_audio/grok_realtime usage ids, avoid CLI headers on api.x.ai
voice, retry video pending store and fail-closed when snapshot is missing without
status duration, align pure-video ImageCount tests, and escalate unset search
price_per_1k to error-level logs.
2026-08-08 09:45:12 +08:00
IanShaw027
12db0f906a
fix(grok): drop account-test ZDR path and align media CLI headers
...
Remove optional upload_url / fake connectivity-only success from admin video
tests. Stamp Grok CLI headers only on the CLI proxy so OAuth media against
api.x.ai can complete and preview video like the gateway path.
2026-08-08 09:45:12 +08:00
IanShaw027
8399a30417
fix(grok): treat free-usage and billing exhaustion as recoverable
...
Classify free-usage bodies during account tests without quarantining content
policy, and mark billing/spending-limit refresh failures as transient so
accounts stay probe-eligible.
2026-08-08 08:48:38 +08:00
IanShaw027
35faaa6d21
feat(grok): register custom-voices CRUD and audio download gateway routes
...
Forward list/get/patch/delete and reference-audio paths with safe path segment
encoding, method passthrough, and empty-body GET/DELETE handling.
2026-08-08 08:48:38 +08:00
IanShaw027
85b65284ec
fix(grok): set async video duration_ms from create accept to done discovery
...
Store CreatedAt on pending billing at video create and use wall-clock E2E
latency when status/content first observes official done+video.url, so usage
logs no longer record only the single poll hop.
2026-08-08 08:48:38 +08:00
IanShaw027
165b072908
fix(grok): gateway media/voice routing, models, and status UI polish
...
Align gateway Grok media/voice paths and model lists, harden upstream failure
and quota handling, clear non-Grok video generation config migration, and polish
temp-unsched/status indicators with model whitelist updates.
2026-08-08 01:07:27 +08:00
IanShaw027
2526a04226
fix(admin): empty web-search config on missing setting and reset dialog scroll
...
Return a disabled empty web-search-emulation config when the setting key is
absent, and reset BaseDialog body scroll on open for long modals.
2026-08-08 01:07:27 +08:00
IanShaw027
68faeac837
fix(grok): restore base URL resolution and operator settings wiring
...
Honor account GetGrokBaseURLOr policy for official vs custom endpoints,
and wire settings resolution used by responses/chat URL builders.
2026-08-08 01:07:19 +08:00
IanShaw027
6d632eec45
fix(grok): tighten OAuth SSO flow and hide password login
...
Require oauth state/redirect consistency, fail closed on missing proxy,
and remove password login from create/reauth UI (admin-only password path stays off by default).
2026-08-08 01:07:19 +08:00
IanShaw027
d0767eab9d
feat(grok): admin account test modes with real media preview
...
Add mode-first connectivity probes for text/image/video/search/tts/stt/realtime,
standalone voice and web-search paths, media upload options, and in-browser
image/audio/video preview (including ZDR-safe b64 images and edit validation).
2026-08-08 01:07:08 +08:00
Wesley Liddick
155c494964
Merge pull request #5399 from fengshao1227/fix/responses-anthropic-invalid-content-blocks
...
fix(apicompat): Responses→Anthropic 转换不再发出上游会拒收的 content block
2026-08-07 23:21:59 +08:00
Wesley Liddick
8991574873
Merge pull request #5345 from puppywang/fix/oauth-pending-account-takeover
...
fix(security): block OAuth account takeover via pending exchange
2026-08-07 23:20:22 +08:00
Wesley Liddick
8f7b0a314d
Merge pull request #5398 from Wei-Shaw/fix/openai-capacity-shed-stream-recovery
...
fix(gateway): 流内降载错误恢复 pre-output failover 并对客户端改写为可重试错误码
2026-08-07 23:11:20 +08:00
li
64090de664
fix(apicompat): Responses→Anthropic 转换不再发出上游会拒收的 content block
...
convertResponsesInputToAnthropic 的 default 分支把未知 item 的 content 逐字
透传成 Anthropic user 消息,Responses 专有的分片类型会原样进入上游请求体。
最典型的是工具执行后回放的 reasoning item:带 content 数组时,reasoning_text
块直接发给 Anthropic,上游回 400 Request body format invalid,而该 item 会一直
留在会话历史里,导致此后每一轮都继续失败。
同时修正两处会产出空内容消息的路径——Anthropic 拒收空内容消息与空白 text 块:
分片全部不可识别时,user 消息退化成 content:""、assistant 消息退化成单个空
text 块。
改动:
- type=reasoning 显式跳过。Anthropic 无法摄入 OpenAI reasoning:encrypted_content
不透明,thinking 重放需要上游签发的 signature。Codex 常见形态(只带 summary +
encrypted_content)本来就会被丢弃,这里让带 content 的形态行为一致。
- default 分支改走 convertResponsesUserToAnthropicContent 白名单转换,保留其中
可识别的文本/图片,丢弃其余分片。
- user / assistant 分支在转换结果为空内容或纯空白文本时跳过该消息。
Fixes #5329
2026-08-07 23:09:03 +08:00
shaw
14a27f1960
test(gateway): 校准 error 帧边界 flush 期望至 pre-output failover 新契约
...
可重试类 error 帧按设计不再算客户端输出(保留在 attempt 缓冲中,
为随后的 response.failed 保住 pre-output failover),因此不在自身
边界单独 flush,而是与终止帧一起出站(1 次 flush)。原用例按旧行为
断言 2 次 flush 导致 CI 失败。
同时补充不可重试类(invalid_request)error 帧仍在边界 flush 的对照
用例,锁住两类错误帧的行为区分。
2026-08-07 22:55:42 +08:00
shaw
c33c3208e3
fix(gateway): 流内降载错误恢复 pre-output failover 并对客户端改写为可重试错误码
...
上游容量降载的真实序列是「event: error → event: response.failed」。此前
{"type":"error"} 帧被当作首个客户端输出立即 flush,clientOutputStarted 被固化,
随后的 response.failed 永远进不了 pre-output failover 分支,降载错误被原样转发;
Codex CLI 对 server_is_overloaded/slow_down 按闭集判致命,直接终止会话并提示
"Selected model is at capacity. Please try a different model."。
修复:
- 可重试类 error 帧不再算客户端输出,恢复既有的同账号重试+切号链路;
不可重试类(content_policy/invalid_request 等)维持原样转发,保留上游错误细节
- 必须转发给客户端时(流中途已有输出 / WS 桥接),把 server_is_overloaded、
slow_down 改写为致命集之外的 server_error,触发 Codex 内置退避重试;
错误消息原样保留,rate_limit_exceeded 等其他错误码一律不动
- 监控、计费与账号状态判定均基于改写前的原始事件
2026-08-07 22:42:33 +08:00
白宦成
de349187d9
fix(openai): harden priority routing hints
2026-08-07 21:47:54 +08:00
白宦成
815035fcc9
fix(openai): send OAuth routing hints
2026-08-07 21:47:53 +08:00
白宦成
915cc7e7bd
fix(openai): stop injecting legacy beta on OAuth responses
2026-08-07 21:47:53 +08:00
Brisbanehuang
db0bff82c7
feat(usage): audit upstream response models
...
(cherry picked from commit 839036224f795c8ee5dc6718a2a14372a45eea44)
2026-08-07 09:40:11 -04:00
Wesley Liddick
e88fc52ce6
Merge pull request #5383 from fengshao1227/fix/responses-tool-parameters-null-type
...
fix(openai): 修正 Responses 工具 Schema 中显式为 null 的 parameters.type
2026-08-07 21:06:10 +08:00
Wesley Liddick
045b620d0d
Merge pull request #5391 from fengshao1227/fix/codex-plan-gated-image-model-no-cooldown
...
fix(ratelimit): 图片模型被 Codex 文本端点拒绝时不再写模型冷却
2026-08-07 21:05:58 +08:00
li
02fbcbe3ad
fix(ratelimit): 守卫按端点来源门控,并与冷却键对齐模型口径
...
上一版守卫只看模型类型,不区分请求从哪个端点进来。OAuth 账号的 /v1/images/*
上游同样是 Codex Responses(openai_images_responses.go → handleOpenAIImagesErrorResponse
→ handleOpenAIAccountUpstreamError → HandleUpstreamModelNotFound),所以专用生图
端点也会命中 plan-gated 分支。账号确实不具备生图能力时跳过冷却,会让调度层失去
唯一的刹车:每个请求都完整走一遍号池,对上游形成无上界的 400 放大。
改动:
- 新增 ctxkey.OpenAIImagesEndpoint 与 WithOpenAIImagesEndpoint /
OpenAIImagesEndpointFromContext,在 handler/openai_images.go 入口置位;
与 OpenAIImageGenerationIntent 区分——后者在 /v1/responses 带图片模型时也会置位。
- 守卫下移到 modelKey 计算之后,抽成 shouldSkipCodexPlanGatedImageModelCooldown,
仅在 plan-gated 分支、且非 /v1/images/* 入站时生效。
- 同时判断 requestedModel 与最终 modelKey:冷却键走 account.GetMappedModel,
账号可以把文本别名映射到 gpt-image-*,只判请求模型会漏掉这种形态。
2026-08-07 20:49:39 +08:00
li
b5d9fd21b0
fix(ratelimit): 图片模型被 Codex 文本端点拒绝时不再写模型冷却
...
gpt-image-* 被误发到 /v1/chat/completions 或 /v1/responses 时,Codex 会稳定
返回 400 codex_plan_gated_model。HandleUpstreamModelNotFound 把它当成账号能力
缺失,给每个账号写 30 分钟 model_rate_limits 冷却。failover 走完整个号池后,
池内每个账号都冷却了该图片模型,于是后续走对端点的 /v1/images/generations
请求也一并 503 no available accounts。
在 codex plan-gated 分支里对图片模型提前返回:仍然 failover 掉本次尝试,
但跳过 SetModelRateLimit 写入,不把号池对正确的生图请求整体下线。
Fixes #4828
2026-08-07 20:41:34 +08:00
li
146b8b6684
revert(grok): 撤回 405 账号级踢号,只保留 failover 分类
...
tempUnscheduleGrok 是账号级、全端点、全模型的下线,而 handleGrokAccountUpstreamError
有 8 个调用点(raw chat / Grok 媒体 / CC 管线 / WS HTTP 桥接 / responses),
/v1/responses 的 405 只证明该端点缺失,账号在其他端点上仍然健康。
仓库里已有相反的既定不变式:shouldApplyOpenAIAlphaSearchAccountErrorSideEffects
把 404/405 显式排除在账号状态写入之外,注释写明「端点不存在只说明该上游不支持
独立搜索,账号本身健康」。
failover 分类器加 405 单独就能解开 sticky 锁死(调度器选号时会即时改绑 sticky),
账号级冷却不是必需的。「405 不踢号但记住端点能力」按端点级负缓存另行处理。
2026-08-07 20:41:23 +08:00
li
a071b27b41
fix(grok): 405 纳入 failover 与临时踢号,修复 Grok 会话粘性锁死
...
Grok 分组里若有账号的自定义 base_url 只实现 /v1/chat/completions,对
/v1/responses 固定返回 405,会话 sticky 到该账号后就永久锁死:
1. shouldFailoverUpstreamError 不含 405,请求不进换号循环
2. handleGrokAccountUpstreamError 的 switch 不含 405,账号不进临时不可调度
3. 账号仍可调度,sticky 不清理,客户端自动重连继续绑同一个坏号
改动两处:
- shouldFailoverUpstreamError 增加 405,让请求可换号,sticky 随之更新
- handleGrokAccountUpstreamError 增加 405 → 30 分钟冷却,与 403 同级
Fixes #4668
2026-08-07 20:37:56 +08:00
li
379db19141
perf(openai): 工具 Schema 净化改为单次拼接,去掉逐路径全量重写
...
sjson.SetBytes 在 opts 为 nil 时 optimistic=false、inplace=false,每次调用都要
重扫整个文档并新建缓冲区拷贝,命中 N 处就是 N 次全量拷贝。/v1/responses 的 body
上限是 gateway.max_body_size(默认 256MB),构造请求能塞进百万级命中,会被放大
成 TB 级 memcpy。
改为收集 gjson 给出的绝对字节偏移后一次性拼接:全程一次分配、一次拷贝。
偏移不可用(Index 为 0)或与 body 对不上时跳过该处,而不是猜位置拼出损坏的 JSON。
实测(2000 处命中):旧写法 17740 次分配,新写法 18 次。
2026-08-07 20:37:44 +08:00
li
f3c94d2099
fix(openai): 修正 Responses 工具 Schema 中显式为 null 的 parameters.type
...
Codex Desktop 内置的 automation_update 工具会带 parameters.type = null,
OpenAI 直接回 400 invalid_function_parameters,而网关把该状态归一成可重试的
502 upstream_error。该工具定义还会沉进多轮会话历史,之后每一轮都继续失败,
同一份坏 Schema 在账号池里被反复重放。
在 OpenAIGatewayService.Forward 分流到 passthrough / Codex transform /
原生 ChatCompletions 之前,统一把显式为 null 的 parameters.type 补成
"object",覆盖顶层 tools[] 与多轮历史里 input[].tools[] 的嵌套工具定义。
只修正显式 null:缺失 type 的 Schema 本身合法,补写会收窄客户端语义,
因此保持原样。
Fixes #5364
2026-08-07 20:30:44 +08:00
Wesley Liddick
b8c0c00091
Merge pull request #5384 from fengshao1227/fix/ops-system-log-sink-flush-backoff
...
fix(ops): 系统日志落库失败后退避重试,避免拖垮数据库连接池
2026-08-07 20:20:20 +08:00
IanShaw027
72a56f862c
fix(grok): free 500k + 对齐 personal-dev tier/时间窗
...
- soft-gate 默认额度 500k tokens / 滚动 24h / 95%(门禁 475k)
- soft-gate 仅显式 free OAuth(subscription_tier/plan_type == free)
- isKnownGrokFreeAccount 按 personal-dev(usage% 为 paid 证据、仅 credentials tier)
- 调度阈值仅 grok_sched header quota 窗,去掉 billing 7d/30d 候选
2026-08-07 18:52:53 +08:00
li
e687ca3e9d
fix(ops): 系统日志落库失败后退避重试,避免拖垮数据库连接池
...
OpsSystemLogSink 的 flush 没有任何失败退避:ticker 每 1 秒触发一次,每次都
带 5 秒超时打一批 COPY,失败后立刻在下一个 tick 再来一次。远程 PostgreSQL 上
COPY 被 context 取消会导致协议失步、连接被服务端 FATAL 终止,于是这条循环
变成持续的「占用连接 → 超时取消 → 销毁重建」;在 max_open 较小的连接池里,
日志通道会长期占满连接,业务侧最终报 Billing 503。
改为连续失败后指数退避(2s 起,翻倍,封顶 60s):
- 退避窗口内直接丢弃当前批次并计入 dropped_count,不再触达数据库;
- 任意一次成功立即清空失败计数与抑制窗口;
- 失败日志随退避天然收敛为每个窗口至多一条,并附带 failures/backoff。
Fixes #5265
2026-08-07 18:45:05 +08:00
IanShaw027
8f5657b91c
fix(grok): 收窄 free 推断 — 仅月度探测,勿把周度 OK 当 free
...
- isKnownGrokFreeAccount 不再因 weekly StatusOK 推断 free
- 修复周度付费快照被 media soft-block 的回归
- 外部 OpenAI token 对比测试在 401 时 skip
2026-08-07 18:34:55 +08:00
IanShaw027
3a415e6d04
fix(grok): 五轮评审 — messages SearchCount 与 OpenAI free-gate
...
- /v1/messages Grok:SSE/buffered 路径累计 SearchCount
- OpenAI 遗留 sticky/list 应用 free soft-gate(advanced 关闭时)
- 单测:OpenAI getSchedulableAccount free-gate + messages 计数器
2026-08-07 18:23:28 +08:00
IanShaw027
3ae94df72b
fix(grok): 四轮评审 — SearchCount 接线与 sticky free-gate
...
- forwardGrokResponses 传播 stream/JSON 的 SearchCount 与 ImageCount
- Chat 桥接 SSE/buffered 对 Grok 累计 SearchCount(附加费)
- getSchedulableAccount sticky 应用 free soft-gate
- 无 call_id 时用合成 key 去重,避免 SSE ~2× 超扣
- 单测:JSON/SSE 接线、sticky free-gate、no-id dedup
2026-08-07 18:15:08 +08:00
IanShaw027
d7c9e7167b
fix(grok): 三轮评审 — 流式 Search 去重与调度/计费加固
...
- P0: 直播 SSE SearchCount 跨事件 call_id 去重,避免 ~2× 附加费
- SearchCount/Audio/WebSearchCalls 走 mandatory usage task
- web_search:uuid 等 forced request_id 优先于 client/local
- Sanitize 始终剥离 cookie;ApplyOAuth 清 grok_needs_reauth_at
- free 判定:paid 证据压过陈旧 free 凭据
- Gateway 列表应用 free soft-gate;token/body-read 可 failover
- web_search 重试支持 WaitPlan 获取;周 PeriodEnd 不再回填月 end
2026-08-07 17:58:23 +08:00
shaw
99b357083e
fix(subscription): restore midnight daily quota reset
...
日额度窗口恢复日历日语义(配置时区每天 0 点刷新),修复 v0.1.170 引入的回归:
- 手动重置配额后日窗口锚点漂移到重置时刻,此后 0 点不再刷新
- 新建/续费订阅日窗口按购买或首用时刻滚动,而非 0 点刷新
日窗口三处锚点(激活/手动重置/续期)统一写入当天 0 点,自动重置改为跨日历日
边界触发;存量非 0 点锚点在下一个 0 点自愈,无需数据库迁移。周/月窗口保持
期限对齐滚动语义(含到期约束)不变,不回归 issue #5051 的月额度翻倍问题。
日卡一次性额度豁免逻辑保留。
2026-08-07 17:42:51 +08:00
IanShaw027
245d069602
fix(grok): 二轮评审残留 — Search 叠加计费与 fail-closed 安全
...
- SearchCost 叠加 token(openai/gateway),未定价 warn
- Token URL 校验失败回落 DefaultTokenURL,禁止 Effective 旁路
- free 判定收窄 paidSignal(仅 plan/月额度),usage% 不否决 free
- web_search: mandatory 计费、uuid request_id、上游 failover 重选账号
- 调度阈值:7d/30d 不跨期 until + 48h stale 可选跳过
- SanitizeStoredCredentials 接入 create/update/bulk/SSO/ApplyOAuth
- ApplyOAuth 成功清除 grok_needs_reauth;SSO 允许 header_override_enabled
- VideoModelPrices 视为媒体定价完整;realtime 正常关闭仍计费
2026-08-07 17:40:21 +08:00
IanShaw027
ccf7ba3ba7
fix(grok): P2 分层清理、主对话 SearchCount 与调度 7d/30d
...
- DoGrokNativeResponsesJSON 去 gin.Context,UA 固定 CLI 身份
- GrokOAuthService 去掉 redis 直依赖,session store 由 wire 注入
- 主对话/流式响应统计 web_search_call/x_search/tool_search 写入 SearchCount
- Grok 调度阈值候选并入官方 weekly/monthly billing %
2026-08-07 17:21:01 +08:00
IanShaw027
7a81468282
fix(grok): 按 review 优先级修复计费漏扣、auth 投影与 free 门禁
...
P0: content 与 status 共用 claim 计费;稳定 grok-video request_id;
auth 热路径投影 video_model_prices。
P1: claim 失败释放可重试;视频绑定 TTL≥24h;SSO 凭据白名单与脱敏;
Token URL 校验;free soft-gate 默认 1M 并统一 free 判定;
Voice 预检余额;STT 抗低报;Realtime 失败不计费;search 未定价告警。
2026-08-07 17:15:43 +08:00
IanShaw027
93a04567d4
style(grok): 清理搜索处理器末尾空行
2026-08-07 17:03:48 +08:00
IanShaw027
be5e4226d4
test(grok): 补齐模型映射与调度设置契约
2026-08-07 17:01:20 +08:00
IanShaw027
5bb206ee99
test(grok): 补齐音频与搜索计价接口契约
2026-08-07 16:56:53 +08:00