6.3 KiB
Backend Configuration
The backend can choose the fastest available runtime automatically and can be overridden from the command line or environment variables.
Recommended Startup
First-time setup on Windows when Python may not be installed:
powershell -ExecutionPolicy Bypass -File scripts\bootstrap_windows.ps1 -Target auto
bootstrap_windows.ps1 will:
- find an existing Python 3.12 installation if available
- install Python 3.12 with
wingetwhen possible - fall back to downloading the official Python installer
- create the CPU/GPU backend virtual environment
- install backend dependencies
- check the detector weight path
First-time setup when Python is already installed:
powershell -ExecutionPolicy Bypass -File scripts\setup_backend.ps1 -Target auto
CPU-only setup:
powershell -ExecutionPolicy Bypass -File scripts\setup_backend.ps1 -Target cpu
GPU setup:
powershell -ExecutionPolicy Bypass -File scripts\setup_backend.ps1 -Target gpu
The setup script creates local virtual environments:
.venv_paddlefor CPU inference.venv_paddle_gpufor GPU inference
The GPU requirements include CUDA runtime wheels used by Paddle, so users do not need to manually install the CUDA toolkit for the default Windows setup. They still need an NVIDIA driver that supports the GPU runtime.
GUI mode:
python scripts\tools\start_backend.py --mode auto
Windows wrapper:
powershell -ExecutionPolicy Bypass -File scripts\start_backend.ps1 -Mode auto
Headless mode:
python scripts\tools\start_backend.py --headless --mode auto
Force GPU:
python scripts\tools\start_backend.py --mode gpu
Force CPU with explicit worker count:
python scripts\tools\start_backend.py --mode cpu --cpu-workers 3
Automatic Defaults
--mode autoprobes.venv_paddle_gpuand selects GPU when Paddle can see a CUDA device.- If GPU is unavailable, the backend falls back to
cpu_parallel. - CPU workers default to:
1worker on 1-2 CPU cores2workers on 3-6 CPU cores3workers on 7+ CPU cores
- Prompt-constrained OCR decoding is enabled by default on both CPU and GPU.
Setup Script
Python entry:
python scripts\setup_backend.py --target auto
Targets:
| Target | Behavior |
|---|---|
auto |
Uses nvidia-smi to choose GPU when an NVIDIA GPU is visible, otherwise CPU |
cpu |
Creates .venv_paddle and installs CPU dependencies |
gpu |
Creates .venv_paddle_gpu and installs GPU dependencies |
both |
Creates both environments |
Useful options:
python scripts\setup_backend.py --target cpu --recreate
python scripts\setup_backend.py --target gpu --no-smoke-test
python scripts\setup_backend.py --target cpu --pip-arg -i --pip-arg https://pypi.tuna.tsinghua.edu.cn/simple
The setup script also checks for the YOLO detector weight:
models/weights/yolo-captcha-detector.pt
Command-Line Options
| Option | Environment variable | Default |
|---|---|---|
--host |
CNCAPTCHA_HOST |
0.0.0.0 |
--port |
CNCAPTCHA_PORT |
8888 |
--mode auto/gpu/cpu |
CNCAPTCHA_OCR_MODE |
auto |
--cpu-workers N |
CNCAPTCHA_CPU_OCR_WORKERS |
auto |
--yolo-device DEVICE |
CNCAPTCHA_YOLO_DEVICE |
0 on GPU, cpu on CPU |
--yolo-imgsz N |
CNCAPTCHA_YOLO_IMGSZ |
448 |
--cpu-model MODEL |
CNCAPTCHA_CPU_OCR_MODEL |
hybrid |
--gpu-model MODEL |
CNCAPTCHA_GPU_OCR_MODEL |
PP-OCRv5_server_rec |
--no-constrained |
CNCAPTCHA_OCR_CONSTRAINED=0 |
constrained enabled |
Additional advanced variables:
$env:CNCAPTCHA_CPU_OCR_FAST_MODEL='PP-OCRv5_mobile_rec'
$env:CNCAPTCHA_CPU_OCR_FALLBACK_MODEL='PP-OCRv5_server_rec'
$env:CNCAPTCHA_GPU_OCR_DEVICE='gpu:0'
$env:CNCAPTCHA_SKIP_GPU_DETECT='1'
Health Check
The server exposes the resolved configuration:
Invoke-RestMethod http://127.0.0.1:8888/health
The response includes backend.ocr_mode, backend.cpu_workers,
backend.gpu_available, and the selected YOLO/OCR settings.
Browser Capture Tips
The backend captures the visible browser window when the userscript sends a
captcha request. Chrome is not required. Current browser matching covers
Chrome, Edge, Firefox, Brave, Opera, bigmodel, Z.ai, GLM, and Chinese
Zhipu page titles.
The released userscript does not need a browser-specific change for this
backend update. Keep the backend GUI dropdown on Auto for normal use. If more
than one browser window is open and the backend captures the wrong one, choose
the exact window from the Browser window dropdown and click Refresh after
opening or switching browser windows.
If the backend says it cannot find the browser window:
- Keep the backend GUI dropdown on
Autounless there are multiple similar browser windows. - If needed, use the backend GUI
Browser windowdropdown to select the exact browser window. ClickRefreshafter opening or switching browser windows. - Keep the GLM Coding page visible and not minimized.
- Use a normal browser window instead of a PWA/app window.
- Make sure the title contains GLM, Z.ai, bigmodel, or Zhipu. Opening the
normal
bigmodel.cn/glm-codingpage is preferred. - Run this diagnostic command:
python scripts\monitor\window_helper.py
If window_helper.py can capture the browser but the backend cannot, update to
the latest script/backend files and restart the backend.
If the backend says it did not get a qualified captcha image:
- Put the browser and captcha dialog on the same monitor as the backend session.
- Avoid hiding the captcha dialog behind the backend GUI or another window.
- Try a medium browser size first, for example 1280x800 to 1920x1080.
- If Windows display scaling is very high, try 100%, 125%, or 150%.
- Keep browser zoom at 50% or above. 67%, 75%, 90%, and 100% are safer choices.
- If the monitor is very large or high-DPI, avoid a tiny browser window; make the captcha image at least about 180 px wide and 150 px tall on screen.
- Keep the whole captcha dialog visible, including the prompt line, image, and confirm button.
The GUI checkboxes that previously appeared in the backend window were removed because they did not control the released userscript flow.