When OCR misreads a character (e.g. 钡→部 lookalike) or multi-window requests
cross, prompt chars don't match box chars exactly -> indices_for_prompt raised
RuntimeError -> worker crashed -> captcha failed. Now map_prompt_to_boxes
catches the mismatch and falls back to candidate_scores: each prompt char
picks the box with highest candidate score for that char. Only warns to stderr,
never crashes. Worst case the click order is wrong but the worker stays alive
for the next captcha.
User reported captcha visible but takes very long to reach backend (OCR itself
only 24ms). The main-page captcha IIFE polled every 1200ms, adding up to 1.2s
delay before processing starts. Reduce to 200ms so detection is near-instant.
The iframe IIFE already runs at 50ms; this brings the main-page path in line.
paddle GPU available != torch GPU available. User may have GPU paddlepaddle
but CPU-only torch in .venv_paddle_gpu. YOLO/Ultralytics needs torch CUDA,
so device=0 crashes with 'Invalid CUDA device=0 requested, torch.cuda False'.
Now backend_config probes torch.cuda.is_available() separately; if torch has
no CUDA, YOLO silently falls back to CPU (OCR still uses GPU) + prints install
hint. Prevents the hard crash reported by users.
When all packages show '抢购人数过多,请刷新再试', the script used to
auto-refresh via location.replace - too aggressive, triggers heavier risk
control. Now it stops and tells the user to manually refresh (F5) until they
see the '特惠订阅' button. State set to DONE so the script pauses cleanly;
page reload on manual refresh restarts the scan fresh.
- RTX 50 (Blackwell) needs CUDA 12.8+ but project uses cu11; use CPU mode
- GPU debugging checklist: paddlepaddle-gpu vs CPU, torch GPU build, driver
- CPU slow? check for old release (v5 server 1189ms vs v6 tiny 110ms)
#32: canBuy/isSoldOut regex was only catching 售罄/补货/暂时; buttons showing
'抢购人数过多,请刷新再试' still passed canBuy -> wasted clicks. Add those +
'请稍后' to the block regex so the script skips and waits instead of clicking
a button that won't go through.
#33: macos-setup.md already had python-tk but framed as optional; Tk is actually
required (one-click-start.command launches a Tk GUI). Reword to make it clear
Tk is required, mention python.org installer bundles tkinter, list the error
messages users see when it's missing.
Also: CRLF fix from earlier, Linux support (PR #34), version bump 23.5 -> 23.6.
Root cause (per Dixon): repo had no .gitattributes, so on Windows git checkout
converted shell scripts to CRLF; release zips built on Windows then carried
CRLF -> macOS/Linux users hit /bin/bash^M: bad interpreter.
Fix (two layers):
- Add .gitattributes: *.sh/*.command = eol=lf (force LF on checkout/commit),
*.ps1/*.cmd = eol=crlf (PS/cmd friendly), JS/py/md = lf, binaries = binary.
- build_release_zips.ps1 / build_portable.ps1: when writing .sh/.command into
the zip, replace CRLF with LF (defense-in-depth even if autocrlf slipped).
Also normalize scripts/setup_backend_macos.sh (was CRLF in working tree).
Verified: new zips have ALL .sh/.command as LF.
The file table still listed start-backend-pipeline-gui.cmd but we removed it
from Windows packages to avoid confusion with one-click-start.cmd. Repository
keeps the file for macOS/manual use, but it shouldn't be in the user-facing
file table. Linux entry from PR #34 already present.
The Linux one-click-start.sh and setup_backend_linux.sh scripts landed in
the previous setup commits but were not surfaced in user-facing docs or
in the release zip layout, so Linux users had no on-ramp. This commit
adds the missing user-facing documentation and ships the Linux entry
point in both release zip builders.
- README.md: extend 快速开始 / 启动后端 to cover Linux alongside Windows
and macOS, update the zip table to mark online-installer as the
macOS / Linux recommendation, add Linux quick-start snippet, and
refresh 常用文件 / 常用启动方式 to list one-click-start.sh and the
new docs/linux-setup.md. Wording tightened so Windows users are not
told to follow Linux commands and vice versa.
- docs/linux-setup.md: new Chinese-language install guide covering scope,
prerequisites (Python 3.12 / uv, NVIDIA GPU optional), one-click and
manual setup paths, virtualenv layout, known limitations, port
troubleshooting, post-install verification, and a Windows/macOS/Linux
comparison table.
- one-click-start.sh: refresh the top-of-file comment so it describes
the actual entry point (start_backend.py --headless ->
captcha_server_headless) instead of the old "pipeline backend" copy
from when the Linux path was a stub.
- scripts/release/build_portable.ps1: include one-click-start.sh in the
portable zip and add a Linux section to the embedded portable README
pointing users at online-installer for Linux with a fallback chmod +
run snippet.
- scripts/release/build_release_zips.ps1: include one-click-start.sh in
the common zip items and rewrite ONLINE_INSTALLER_README.txt to give
per-platform launch instructions, document auto PyPI mirror detection,
the .venv_paddle / .venv_paddle_gpu layout, and link to
docs/linux-setup.md.
No code logic change; only docs + packaging so Linux is a first-class
release target alongside Windows and macOS.
The Linux one-click-start and setup_backend_linux scripts were both
assuming interactive stdin:
- one-click-start.sh called `read -r -p ...` on every fatal exit path
(non-Linux host, missing release files, install failure, missing
weight). When the script is launched from a desktop entry, systemd
user unit, CI, or any other context where stdin is not a TTY, `read`
would block waiting for input that will never come, hanging the
process instead of exiting promptly. Replace each site with a small
`pause_if_tty` helper that only prompts when `[ -t 0 ]`, so the
script still pauses for a user at a real terminal but returns
immediately when run unattended.
- scripts/setup_backend_linux.sh printed a single, hard-coded launch
block that only mentioned `.venv_paddle` (CPU). After the previous
--target auto/cpu/gpu/both refactor, the user could have built any
combination of CPU and GPU venvs but the hints always pointed at the
CPU one. Wrap the post-install message in `print_completion_hints`,
iterate over the actually-selected modes to print the correct
`<venv>/bin/python` for each (CPU -> .venv_paddle, GPU ->
.venv_paddle_gpu), and add an auto-mode command line when both
venvs exist or a hint to install the CPU fallback when only GPU
was built.
No behaviour change for the install flow itself; only the prompt
handling and the post-install guidance.
Previously invoke_setup splatted the script path with its arguments directly
(""${args[@]}""), which relies on the script having a valid shebang and
the executable bit set when invoked from a different process state or via
some shell wrappers. The Windows Invoke-Bootstrap counterpart launches
its bootstrap via a stable interpreter entry (powershell -File), and the
Linux path should be at least as robust.
- Resolve the setup script to an absolute path under $SCRIPT_DIR and run
it explicitly with 'bash "$setup_script" "${setup_args[@]}"', so the
shebang/exec bit is not on the critical path.
- Rename the local array to setup_args / setup_script to make intent
clearer at the call sites.
This keeps the recreate-on-existing-python behaviour from the previous
fix and matches the Windows pipeline's "stable interpreter invocation"
convention.
The Linux one-click-start invoke_setup helper only recreated the venv when
NEEDS_RECREATE was set by the portability/import check. If the target
python (e.g. <venv>/bin/python) already existed from a previous foreign
install or partial state, the setup script would try to reuse it and fail
in confusing ways. This mirrors the Windows Invoke-Bootstrap behaviour.
- invoke_setup now takes the selected venv python as $2, with the recreate
flag moved to $3 (call sites in the auto/cpu fallback branches updated).
- Recreate is forced when force_recreate=1 OR the target venv python path
already exists on disk, so any pre-existing python triggers --recreate.
- Comments added in Chinese to match the rest of the Linux script and
reference the Windows counterpart.
Fixes the case where running one-click-start.sh on a host that already
has a broken/foreign <venv>/bin/python leaves the user with a half-set-up
environment instead of a clean rebuild.
Linux one-click-start.sh and setup_backend_linux.sh previously installed
dependencies from the default PyPI source, which is often slow or blocked
in mainland China. Mirror the Windows pipeline's PyPI mirror behaviour by
probing well-known mirrors and falling back to the official source.
- New scripts/pypi_mirror.sh: ordered list of Tsinghua / Aliyun / USTC /
Tencent / pypi.org; pypi_mirror_probe checks each with curl or wget
(3s timeout); pypi_mirror_select returns the first reachable URL;
ensure_pypi_mirror_pip_args sets PIP_ARGS=("-i" <url>) only when the
caller has not already supplied --pip-arg. Reusable via 'source'.
- one-click-start.sh sources pypi_mirror.sh and calls
ensure_pypi_mirror_pip_args before invoking the backend installer, so a
bare one-click-start.sh on a CN host picks the fastest mirror.
- scripts/setup_backend_linux.sh does the same for direct invocations
(e.g. ./scripts/setup_backend_linux.sh), with the same skip-if-pip-arg-
set semantics. Header comment updated to document the behaviour.
Users who pass --pip-arg keep full control; everyone else gets a working
mirror automatically and the install no longer wedges on pypi.org.
Mirror the Windows one-click-start.cmd / setup_backend.ps1 flow for
Linux users:
- scripts/setup_backend_linux.sh: detect NVIDIA GPU, prefer uv for
Python 3.12 / venv / pip install of requirements-backend-{cpu,gpu}.txt,
run core-import smoke tests, verify yolo-captcha-detector.pt weight.
Supports --target auto/cpu/gpu/both, --recreate, --skip-install,
--no-smoke-test, --pip-arg.
- one-click-start.sh: Linux entry point that picks auto/cpu/gpu venv,
rebuilds foreign or broken venvs, falls back from GPU -> CPU on auto
failure, ensures a CPU fallback venv exists alongside GPU, and execs
scripts/tools/start_backend.py --headless with CNCAPTCHA_PORT and
CPU/GPU python overrides.
Together with one-click-start.cmd/.command this gives every platform a
single double-click path to a working captcha backend.
Greasy Fork script page description (docs/greasyfork-description.html) is what
users see on the script listing. Add the click-block risk control as the first
item in 重要提醒: blocked -> wait 10s or manually click 特惠订阅, OCR still auto.
One-liner: 订阅入口点不动就手动点特惠订阅,验证码自动打.
The standalone userscript README had no mention of the 2026-06-23 click-block
risk control or the config panel options. Add:
- Risk control section: blocked click -> wait 10s or manually click 特惠订阅,
OCR still auto. One-liner: 入口点不动就手动点特惠订阅,验证码自动打.
- Config panel guide: click delay, first-click delay, auto-close, auto-click.
- F8/F9 hotkeys that were missing.
The @description shows on Greasy Fork/Tampermonkey list and update notices -
first thing users see. Add: (1) PP-OCRv6 in parens, (2) the risk-control
workaround one-liner '订阅入口被风控拦截时手动点特惠订阅即可,验证码自动打'.
Zhipu upgraded risk control this morning: auto-click subscribe sometimes
blocked/rejected (manual click too), shows block_message. Advisory:
- wait ~10s and retry when blocked
- if auto-click keeps failing, manually click the '特惠订阅' entry
- OCR/captcha solving still fully automatic regardless of how the entry was
triggered (auto-click and OCR are decoupled)
One-liner: 入口点不动就手动点特惠订阅,验证码自动打。
OCR is now so fast (v6 tens of ms) that the first character click can land
before the captcha DOM finishes rendering/animating -> first click misses.
Previously the inter-click delay only applied between characters; the first
click had zero wait. Add a configurable pre-first-click random delay (default
150-300ms): new CAPTCHA_FIRST_CLICK_DELAY_MIN_MS/MAX_MS config, exposed in the
config panel under '验证码点字延时策略' as '首击前延时', passed through
getCaptchaDirectDelayConfig + normalize, slept at the start of
clickCaptchaDirectCoords. Set to 0 to disable.
Previous fix (run bootstrap in-process to avoid pipe buffering) broke
parameter binding: '& script.ps1 @argsList' splatting treated '-Target'
as the value of a positional param, not the param name -> ValidateSet
rejected it. Revert to external '& powershell -File bootstrap.ps1 @argsList'
(param binding reliable), drop the | Tee pipe (was the buffer root cause),
let child stdout inherit console directly for real-time pip progress.
Windows users seeing two .cmd files (one-click-start.cmd and
start-backend-pipeline-gui.cmd) get confused which to click. pipeline-gui
runs the optional backend/ FastAPI pipeline, not the main path. Remove its
.cmd/.ps1 from both portable and online package Include lists; keep the
.command for macOS. one-click-start.cmd is now the only Windows entry.
The quick-start step 4 and the rush steps both told users to double-click
start-backend-pipeline-gui.cmd, but that runs the optional backend/ FastAPI
pipeline, not what one-click-start runs (captcha_server.py). Windows users
should double-click one-click-start.cmd for both first-time install and daily
launch. pipeline-gui.cmd now only mentioned as optional alternative.
The '验证码识别说明' section described backend/ FastAPI pipeline as the main
backend and recommended start-backend-pipeline-gui.cmd, but one-click-start
actually runs scripts/tools/captcha_server.py (different Tk GUI, different
model display). Rewrote to:
- Describe captcha_server.py as the main backend one-click-start runs
- Document the Tk GUI fields (status/model/prompt/result/log with timing)
- Clarify backend/ is an optional alternative pipeline, not the default
- Update file table: one-click-start.cmd is the main entry, mention
captcha_server.py explicitly
- Mention v6 tiny+medium and the config/env override for ocr_model
The & powershell -File | Tee-Object pipeline buffered bootstrap output, so
during the ~1min pip install users saw nothing between 'Installing cpu
environment...' and 'Starting backend on 8888' - looked frozen. Run bootstrap
in-process (no child powershell, no pipe) so pip progress streams to console;
use Start-Transcript for the log file instead of Tee pipe (no buffering).
- Small hidden set: add row 10 PP-OCRv6 tiny + constrained (~30ms/img)
- Stress test 379: add PP-OCRv6 tiny row (100%, ~110ms/img, ~11x faster than
v5 server), keep v5 server as comparison, note test conditions
- Strict click radius table: add v6 tiny column (100% at all radii)
- Update conclusion: v6 tiny now best on accuracy/speed/stability
- Update model journey line to mention v6
Repro: auto-mode -> F8 pause -> manually close invalid popup -> F8 resume
-> stuck in WAITING for full timeout. Root cause: tick returns immediately
while paused, so taskClickTime stays frozen at the pre-pause click; on resume
elapsed includes pause duration. Worse, the manually-closed popup is gone,
so findRLModal/isPayDialog return null and WAITING has no branch to detect
'the popup I was waiting for is gone' -> falls through to the timeout tail.
Fix in setRuntimePaused(resume): if currently WAITING, reset taskClickTime to
now (re-count elapsed from resume); if no popup present anymore (closed during
pause), drop straight to IDLE to retry instead of waiting out the timeout.
Explains why not-closing-the-popup was fine (popup still detectable on resume).
The 30s wait (rare but should never happen) came from layered captcha-state
guessing (captchaSeen/GM counter/PS.inProgress branches) that misjudged some
edge cases as 'captcha in progress, wait patiently'. Strip all of it: after
click, just wait for a result (preview response / dialog), hard cap 10s. Normal
recognition (even on slow CPU) finishes in 5-8s; over 10s = stuck = retry.
The [direct] 识别成功 log line (captcha_server.py line 492) was the one users
actually saw - it had no timing. My earlier [captcha] print landed on the
[CAPTURE] path (auto-capture), a DIFFERENT code path. Replace the [direct]
success/fail logs with the same [captcha] summary including end-to-end + yolo
+ ocr timing. Now both paths print timing consistently.
The [CAPTURE] got result line truncated JSON at 200 chars, cutting off
elapsed_ms - user saw no timing. Replace with a concise [captcha] summary:
[captcha] 城匙称 -> 称匙城 | conf=1.00 end-to-end=312ms (worker total=285ms yolo=120ms ocr=165ms) | engine=yolo+hybrid_cpu_parallel
This is the code path one-click-start actually runs (scripts/tools/), not
backend/server.py where I added a similar line earlier.
Add a '识别模型' row below status showing the resolved model: for hybrid mode
'hybrid | 主 PP-OCRv6_tiny_rec → 兜底 PP-OCRv6_medium_rec', for GPU the gpu
model, otherwise the cpu model. Updated every tick so it reflects config even
if resolved after GUI start.
User log showed PP-OCRv5_mobile_rec still loading despite v6 'upgrade'. Root
cause: one-click-start -> start_backend.ps1 -> scripts/tools/start_backend.py
-> captcha_server.py (the EARLY pipeline), NOT backend/server.py (PR#10). The
v6 default I set in backend/ was never reached.
Set v6 defaults in the code path users actually run:
- backend_config.py: cpu_fast_model v5_mobile -> v6_tiny, cpu_fallback_model
v5_server -> v6_medium, gpu_model v5_server -> v6_tiny (hybrid fast path).
- ppocr_cpu_pool_worker.py: MODEL_NAME v5_server -> v6_tiny (non-hybrid path).
- backend_config.py line 134 gpu probe default also v6_tiny.
The pipeline refactor dropped the per-recognition log line, so the GUI log
box (which tails backend stdout) showed nothing useful - no prompt, no result,
no timing. Add a concise [captcha] line after each result is assembled:
[captcha] 畅倍标 -> 畅倍标 | conf=0.99 total=162ms yolo=103ms ocr=59ms | req=1
Also bump _http_get_json timeout 1s -> 3s for safety under load.
Instead of hardcoding a single mirror (if it's down, install fails), probe a
list of mirrors at install time and pick the first reachable one:
1. Tsinghua 2. Aliyun 3. USTC 4. Tencent 5. pypi.org (official fallback)
HEAD probe with 3s timeout per mirror (~seconds total). Users can still override
via -PipArg. Verified: probe correctly selects Tsinghua when reachable.
Real root cause of 100% failure (finally visible via tee log):
'bootstrap_windows.ps1 : Missing an argument for parameter PipArg'
one_click_start Invoke-Bootstrap passed '-PipArg -i -PipArg https://...'
but PowerShell splatting treats the dash-prefixed '-i' as the NEXT parameter
name, not the PipArg value -> bootstrap aborts before pip even runs.
Same class of bug as the argparse one, one layer up. Fix:
- Invoke-Bootstrap: join PipArg into a single ;-delimited string, pass one
-PipArg value (no dash ambiguity at the PS param layer).
- bootstrap_windows.ps1: PipArg is now [string]; split on ';' to recover the
array, then emit '--pip-arg=VALUE' to setup_backend as before.
- Verified: bootstrap_windows.ps1 -PipArg '-i;https://...' runs cleanly,
setup_backend receives --pip-arg=-i --pip-arg=https://... correctly.
Users hitting [FAIL] with no visible pip error - the nested powershell +
setup_backend errors were getting buried. Tee the full bootstrap output to
logs/backend-install.log and tell the user to share that file when reporting.
Root cause of 100% install failure: setup_backend.py --pip-arg used action=append,
so '--pip-arg -i' made argparse treat '-i' (a dash-prefixed value) as the next
option and abort with 'error: argument --pip-arg: expected one argument'.
The mirror set by one_click_start.ps1 never reached pip -> mainland users hit
PyPI timeout -> 'Backend environment repair failed'.
Fix:
- bootstrap_windows.ps1: emit '--pip-arg=VALUE' (= form) so dash-prefixed values
like -i / --index-url are not misparsed as options.
- setup_backend.py: also pass mirror to the 'pip install --upgrade pip setuptools
wheel' step (line 137 previously had none, would time out before requirements).
- Verified end-to-end: fresh venv, full CPU install via Tsinghua mirror,
paddleocr-3.7.0 + paddlepaddle-3.3.1 installed, smoke test passed.
PS 5.1 on Chinese Windows reads non-BOM files as the system ANSI codepage
(GBK), so UTF-8 Chinese bytes corrupt parsing and surface as
'Missing expression after comma' at param blocks. Adding EF BB BF BOM forces
PS 5.1 to interpret as UTF-8. Affects one_click_start.ps1 (the .cmd entry),
bootstrap_windows.ps1, start_backend.ps1, setup_backend.ps1,
start-backend-pipeline-gui.ps1, and the release scripts.
Users in mainland China hitting ReadTimeoutError / RemoteDisconnected when
one-click-start.cmd installs backend deps from files.pythonhosted.org.
one_click_start.ps1 -PipArg now defaults to -i https://pypi.tuna.tsinghua.edu.cn/simple;
users can override by passing -PipArg explicitly. Chain already supports it:
one_click_start.ps1 -> bootstrap_windows.ps1 -> setup_backend.py --pip-arg.
- build_release_zips.ps1 / build_portable.ps1: Windows bsdtar -a misidentifies
.zip extension and produces corrupt archives; Compress-Archive fails on long/
non-ASCII paths. Fall back to python zipfile (no MAX_PATH limit, standard zip).
- build_portable.ps1: add requirements-backend-gpu.txt to Include list. The
one_click_start.ps1 Assert-RequiredFiles check requires BOTH requirements
files in every package; portable-cpu was missing gpu one, causing
'[FAIL] Release package is incomplete' on first run.
Backend:
- Default OCR model PP-OCRv5_server_rec -> PP-OCRv6_tiny_rec (~14x faster,
83ms/img vs 1189ms on 379 real captchas, accuracy still 100%)
- Configurable via config.json ocr_model or env CNCAPTCHA_CPU_OCR_MODEL/GLM_OCR_MODEL
- GUI OCR model choices updated to v6 (tiny/medium) + v5 server fallback
- worker.py: split 3 crops into independent OCR tasks for parallel recognition
- ppocr_worker.py: batch forward (predict_batch_*) + single-crop path
- server.py: partial results aggregation + OCR_MODEL env passthrough
- requirements: paddleocr>=3.7.0 / paddlex>=3.7.0 (v6 requires)
Userscript v23.3:
- Fix: auto-click subscribe then wait 15s when click didn't trigger captcha.
Synthetic click sometimes gets swallowed by Zhipu frontend (button DOM ready
but component state machine not ready), no captcha pops up, main loop stuck in
WAITING until MODAL_WAIT=15000 timeout.
- Fix: use 'iframe got new prompt+bg image' as the signal. iframe increments GM
counter glm_captcha_seen_seq each time it sees a new captcha (prompt or bg
changed); main loop records baseline at click time, compares in WAITING -
counter increased = captcha popped, wait patiently; no increase after 1.5s =
click didn't trigger captcha, immediately retry subscribe. Counter is
monotonic, no residue, no timestamp race.