mirror of
https://github.com/OLmatter/glm-coding-helper.git
synced 2026-10-07 14:58:39 +08:00
202 lines
6.3 KiB
Markdown
202 lines
6.3 KiB
Markdown
# Backend Configuration
|
|
|
|
The backend can choose the fastest available runtime automatically and can be
|
|
overridden from the command line or environment variables.
|
|
|
|
## Recommended Startup
|
|
|
|
First-time setup on Windows when Python may not be installed:
|
|
|
|
```powershell
|
|
powershell -ExecutionPolicy Bypass -File scripts\bootstrap_windows.ps1 -Target auto
|
|
```
|
|
|
|
`bootstrap_windows.ps1` will:
|
|
|
|
- find an existing Python 3.12 installation if available
|
|
- install Python 3.12 with `winget` when possible
|
|
- fall back to downloading the official Python installer
|
|
- create the CPU/GPU backend virtual environment
|
|
- install backend dependencies
|
|
- check the detector weight path
|
|
|
|
First-time setup when Python is already installed:
|
|
|
|
```powershell
|
|
powershell -ExecutionPolicy Bypass -File scripts\setup_backend.ps1 -Target auto
|
|
```
|
|
|
|
CPU-only setup:
|
|
|
|
```powershell
|
|
powershell -ExecutionPolicy Bypass -File scripts\setup_backend.ps1 -Target cpu
|
|
```
|
|
|
|
GPU setup:
|
|
|
|
```powershell
|
|
powershell -ExecutionPolicy Bypass -File scripts\setup_backend.ps1 -Target gpu
|
|
```
|
|
|
|
The setup script creates local virtual environments:
|
|
|
|
- `.venv_paddle` for CPU inference
|
|
- `.venv_paddle_gpu` for GPU inference
|
|
|
|
The GPU requirements include CUDA runtime wheels used by Paddle, so users do
|
|
not need to manually install the CUDA toolkit for the default Windows setup.
|
|
They still need an NVIDIA driver that supports the GPU runtime.
|
|
|
|
GUI mode:
|
|
|
|
```powershell
|
|
python scripts\tools\start_backend.py --mode auto
|
|
```
|
|
|
|
Windows wrapper:
|
|
|
|
```powershell
|
|
powershell -ExecutionPolicy Bypass -File scripts\start_backend.ps1 -Mode auto
|
|
```
|
|
|
|
Headless mode:
|
|
|
|
```powershell
|
|
python scripts\tools\start_backend.py --headless --mode auto
|
|
```
|
|
|
|
Force GPU:
|
|
|
|
```powershell
|
|
python scripts\tools\start_backend.py --mode gpu
|
|
```
|
|
|
|
Force CPU with explicit worker count:
|
|
|
|
```powershell
|
|
python scripts\tools\start_backend.py --mode cpu --cpu-workers 3
|
|
```
|
|
|
|
## Automatic Defaults
|
|
|
|
- `--mode auto` probes `.venv_paddle_gpu` and selects GPU when Paddle can see a
|
|
CUDA device.
|
|
- If GPU is unavailable, the backend falls back to `cpu_parallel`.
|
|
- CPU workers default to:
|
|
- `1` worker on 1-2 CPU cores
|
|
- `2` workers on 3-6 CPU cores
|
|
- `3` workers on 7+ CPU cores
|
|
- Prompt-constrained OCR decoding is enabled by default on both CPU and GPU.
|
|
|
|
## Setup Script
|
|
|
|
Python entry:
|
|
|
|
```powershell
|
|
python scripts\setup_backend.py --target auto
|
|
```
|
|
|
|
Targets:
|
|
|
|
| Target | Behavior |
|
|
| --- | --- |
|
|
| `auto` | Uses `nvidia-smi` to choose GPU when an NVIDIA GPU is visible, otherwise CPU |
|
|
| `cpu` | Creates `.venv_paddle` and installs CPU dependencies |
|
|
| `gpu` | Creates `.venv_paddle_gpu` and installs GPU dependencies |
|
|
| `both` | Creates both environments |
|
|
|
|
Useful options:
|
|
|
|
```powershell
|
|
python scripts\setup_backend.py --target cpu --recreate
|
|
python scripts\setup_backend.py --target gpu --no-smoke-test
|
|
python scripts\setup_backend.py --target cpu --pip-arg -i --pip-arg https://pypi.tuna.tsinghua.edu.cn/simple
|
|
```
|
|
|
|
The setup script also checks for the YOLO detector weight:
|
|
|
|
```text
|
|
models/weights/yolo-captcha-detector.pt
|
|
```
|
|
|
|
## Command-Line Options
|
|
|
|
| Option | Environment variable | Default |
|
|
| --- | --- | --- |
|
|
| `--host` | `CNCAPTCHA_HOST` | `0.0.0.0` |
|
|
| `--port` | `CNCAPTCHA_PORT` | `8888` |
|
|
| `--mode auto/gpu/cpu` | `CNCAPTCHA_OCR_MODE` | `auto` |
|
|
| `--cpu-workers N` | `CNCAPTCHA_CPU_OCR_WORKERS` | auto |
|
|
| `--yolo-device DEVICE` | `CNCAPTCHA_YOLO_DEVICE` | `0` on GPU, `cpu` on CPU |
|
|
| `--yolo-imgsz N` | `CNCAPTCHA_YOLO_IMGSZ` | `448` |
|
|
| `--cpu-model MODEL` | `CNCAPTCHA_CPU_OCR_MODEL` | `hybrid` |
|
|
| `--gpu-model MODEL` | `CNCAPTCHA_GPU_OCR_MODEL` | `PP-OCRv5_server_rec` |
|
|
| `--no-constrained` | `CNCAPTCHA_OCR_CONSTRAINED=0` | constrained enabled |
|
|
|
|
Additional advanced variables:
|
|
|
|
```powershell
|
|
$env:CNCAPTCHA_CPU_OCR_FAST_MODEL='PP-OCRv5_mobile_rec'
|
|
$env:CNCAPTCHA_CPU_OCR_FALLBACK_MODEL='PP-OCRv5_server_rec'
|
|
$env:CNCAPTCHA_GPU_OCR_DEVICE='gpu:0'
|
|
$env:CNCAPTCHA_SKIP_GPU_DETECT='1'
|
|
```
|
|
|
|
## Health Check
|
|
|
|
The server exposes the resolved configuration:
|
|
|
|
```powershell
|
|
Invoke-RestMethod http://127.0.0.1:8888/health
|
|
```
|
|
|
|
The response includes `backend.ocr_mode`, `backend.cpu_workers`,
|
|
`backend.gpu_available`, and the selected YOLO/OCR settings.
|
|
|
|
## Browser Capture Tips
|
|
|
|
The backend captures the visible browser window when the userscript sends a
|
|
captcha request. Chrome is not required. Current browser matching covers
|
|
Chrome, Edge, Firefox, Brave, Opera, `bigmodel`, `Z.ai`, `GLM`, and Chinese
|
|
Zhipu page titles.
|
|
|
|
The released userscript does not need a browser-specific change for this
|
|
backend update. Keep the backend GUI dropdown on `Auto` for normal use. If more
|
|
than one browser window is open and the backend captures the wrong one, choose
|
|
the exact window from the `Browser window` dropdown and click `Refresh` after
|
|
opening or switching browser windows.
|
|
|
|
If the backend says it cannot find the browser window:
|
|
|
|
- Keep the backend GUI dropdown on `Auto` unless there are multiple similar
|
|
browser windows.
|
|
- If needed, use the backend GUI `Browser window` dropdown to select the exact
|
|
browser window. Click `Refresh` after opening or switching browser windows.
|
|
- Keep the GLM Coding page visible and not minimized.
|
|
- Use a normal browser window instead of a PWA/app window.
|
|
- Make sure the title contains GLM, Z.ai, bigmodel, or Zhipu. Opening the
|
|
normal `bigmodel.cn/glm-coding` page is preferred.
|
|
- Run this diagnostic command:
|
|
|
|
```powershell
|
|
python scripts\monitor\window_helper.py
|
|
```
|
|
|
|
If `window_helper.py` can capture the browser but the backend cannot, update to
|
|
the latest script/backend files and restart the backend.
|
|
|
|
If the backend says it did not get a qualified captcha image:
|
|
|
|
- Put the browser and captcha dialog on the same monitor as the backend session.
|
|
- Avoid hiding the captcha dialog behind the backend GUI or another window.
|
|
- Try a medium browser size first, for example 1280x800 to 1920x1080.
|
|
- If Windows display scaling is very high, try 100%, 125%, or 150%.
|
|
- Keep browser zoom at 50% or above. 67%, 75%, 90%, and 100% are safer choices.
|
|
- If the monitor is very large or high-DPI, avoid a tiny browser window; make
|
|
the captcha image at least about 180 px wide and 150 px tall on screen.
|
|
- Keep the whole captcha dialog visible, including the prompt line, image, and
|
|
confirm button.
|
|
|
|
The GUI checkboxes that previously appeared in the backend window were removed
|
|
because they did not control the released userscript flow.
|