fix(backend): 启动时跑一次真推理预热 OCR 模型,避免首请求超时

## 现象
首次启动后端立刻抢购时,第一个验证码识别会失败:浏览器 console 报
"Permission was denied for ... loopback",后端日志报 WinError 10053
(客户端中止已建立连接)。

## 根因
PaddleOCR 首次推理需要 7-12 秒做 JIT 编译(CPU 模式实测 OCR 单步
7.8s + YOLO 等其他几秒)。原 preload_worker 只启动子进程不跑真推理,
"识别子进程就绪" 日志骗了用户——进程就绪 ≠ 模型预热完成。

如果首个验证码请求撞上 JIT 编译,客户端(Tampermonkey GM_xmlhttpRequest
或浏览器 fetch)等不了 12 秒就断开连接,后端推理出结果时 socket 已经
reset,点击错过最佳时机。

后续请求虽然 56ms 就能返回,但腾讯 captcha SDK 可能已经触发自我重置,
导致看上去 "成功几次就不灵"。

## 修法
在 preload_worker 启动子进程后,立即跑一次 dummy 真推理触发完整的
yolo+ocr JIT 编译。优先用 dataset/debug_captcha_direct/ 下已有的真
验证码样本(同尺寸、同分布、JIT 路径最贴近实战),找不到时兜底用全白
480x672 图。

## 验证
- 启动后端,GUI 日志会多出 "模型预热完成 (xxxxms)" 一行
- 立即触发第一个验证码:端到端 50-100ms 完成(vs 修复前 12000ms)
- WinError 10053 不再出现
This commit is contained in:
ygjj
2026-06-25 05:45:58 +08:00
committed by OLmatter
parent 1b1c6b5c0c
commit 2531183237
+25 -1
View File
@@ -654,7 +654,31 @@ def run_gui():
try:
log_to_gui("预热识别子进程...")
_get_worker()
log_to_gui("识别子进程就绪!")
log_to_gui("识别子进程就绪,开始模型预热...")
# 子进程启动 ≠ 模型预热完成。PaddleOCR 第一次真推理会做 JIT 编译,
# 实测耗时 7-12s。如果首个真验证码撞上 JIT,客户端必断(WinError 10053
# 或浏览器 fetch timeout),结果识别出来了也来不及点。
# 这里跑一次 dummy 推理把整条 yolo+ocr 链路打热。
try:
debug_dir = ROOT / "dataset" / "debug_captcha_direct"
warmup_path = None
if debug_dir.exists():
for p in debug_dir.iterdir():
if p.suffix.lower() == ".png":
warmup_path = p
break
if warmup_path is None:
# 兜底:生成一张全白 480x672 图,跟真验证码同尺寸
from PIL import Image as _Img
warmup_path = ROOT / "dataset" / "_warmup.png"
warmup_path.parent.mkdir(parents=True, exist_ok=True)
if not warmup_path.exists():
_Img.new("RGB", (672, 480), "white").save(warmup_path)
_wt0 = time.perf_counter()
_ = recognize_captcha(str(warmup_path), list("测试图"), crop_rect=None)
log_to_gui(f"模型预热完成 ({(time.perf_counter()-_wt0)*1000:.0f}ms)")
except Exception as we:
log_to_gui(f"预热失败(不影响功能,仅首次会慢): {we}")
except Exception as e:
log_to_gui(f"子进程启动失败: {e}")