When OCR misreads a character (e.g. 钡→部 lookalike) or multi-window requests
cross, prompt chars don't match box chars exactly -> indices_for_prompt raised
RuntimeError -> worker crashed -> captcha failed. Now map_prompt_to_boxes
catches the mismatch and falls back to candidate_scores: each prompt char
picks the box with highest candidate score for that char. Only warns to stderr,
never crashes. Worst case the click order is wrong but the worker stays alive
for the next captcha.
User reported captcha visible but takes very long to reach backend (OCR itself
only 24ms). The main-page captcha IIFE polled every 1200ms, adding up to 1.2s
delay before processing starts. Reduce to 200ms so detection is near-instant.
The iframe IIFE already runs at 50ms; this brings the main-page path in line.
paddle GPU available != torch GPU available. User may have GPU paddlepaddle
but CPU-only torch in .venv_paddle_gpu. YOLO/Ultralytics needs torch CUDA,
so device=0 crashes with 'Invalid CUDA device=0 requested, torch.cuda False'.
Now backend_config probes torch.cuda.is_available() separately; if torch has
no CUDA, YOLO silently falls back to CPU (OCR still uses GPU) + prints install
hint. Prevents the hard crash reported by users.
When all packages show '抢购人数过多,请刷新再试', the script used to
auto-refresh via location.replace - too aggressive, triggers heavier risk
control. Now it stops and tells the user to manually refresh (F5) until they
see the '特惠订阅' button. State set to DONE so the script pauses cleanly;
page reload on manual refresh restarts the scan fresh.
- RTX 50 (Blackwell) needs CUDA 12.8+ but project uses cu11; use CPU mode
- GPU debugging checklist: paddlepaddle-gpu vs CPU, torch GPU build, driver
- CPU slow? check for old release (v5 server 1189ms vs v6 tiny 110ms)
#32: canBuy/isSoldOut regex was only catching 售罄/补货/暂时; buttons showing
'抢购人数过多,请刷新再试' still passed canBuy -> wasted clicks. Add those +
'请稍后' to the block regex so the script skips and waits instead of clicking
a button that won't go through.
#33: macos-setup.md already had python-tk but framed as optional; Tk is actually
required (one-click-start.command launches a Tk GUI). Reword to make it clear
Tk is required, mention python.org installer bundles tkinter, list the error
messages users see when it's missing.
Also: CRLF fix from earlier, Linux support (PR #34), version bump 23.5 -> 23.6.
Root cause (per Dixon): repo had no .gitattributes, so on Windows git checkout
converted shell scripts to CRLF; release zips built on Windows then carried
CRLF -> macOS/Linux users hit /bin/bash^M: bad interpreter.
Fix (two layers):
- Add .gitattributes: *.sh/*.command = eol=lf (force LF on checkout/commit),
*.ps1/*.cmd = eol=crlf (PS/cmd friendly), JS/py/md = lf, binaries = binary.
- build_release_zips.ps1 / build_portable.ps1: when writing .sh/.command into
the zip, replace CRLF with LF (defense-in-depth even if autocrlf slipped).
Also normalize scripts/setup_backend_macos.sh (was CRLF in working tree).
Verified: new zips have ALL .sh/.command as LF.
The file table still listed start-backend-pipeline-gui.cmd but we removed it
from Windows packages to avoid confusion with one-click-start.cmd. Repository
keeps the file for macOS/manual use, but it shouldn't be in the user-facing
file table. Linux entry from PR #34 already present.
The Linux one-click-start.sh and setup_backend_linux.sh scripts landed in
the previous setup commits but were not surfaced in user-facing docs or
in the release zip layout, so Linux users had no on-ramp. This commit
adds the missing user-facing documentation and ships the Linux entry
point in both release zip builders.
- README.md: extend 快速开始 / 启动后端 to cover Linux alongside Windows
and macOS, update the zip table to mark online-installer as the
macOS / Linux recommendation, add Linux quick-start snippet, and
refresh 常用文件 / 常用启动方式 to list one-click-start.sh and the
new docs/linux-setup.md. Wording tightened so Windows users are not
told to follow Linux commands and vice versa.
- docs/linux-setup.md: new Chinese-language install guide covering scope,
prerequisites (Python 3.12 / uv, NVIDIA GPU optional), one-click and
manual setup paths, virtualenv layout, known limitations, port
troubleshooting, post-install verification, and a Windows/macOS/Linux
comparison table.
- one-click-start.sh: refresh the top-of-file comment so it describes
the actual entry point (start_backend.py --headless ->
captcha_server_headless) instead of the old "pipeline backend" copy
from when the Linux path was a stub.
- scripts/release/build_portable.ps1: include one-click-start.sh in the
portable zip and add a Linux section to the embedded portable README
pointing users at online-installer for Linux with a fallback chmod +
run snippet.
- scripts/release/build_release_zips.ps1: include one-click-start.sh in
the common zip items and rewrite ONLINE_INSTALLER_README.txt to give
per-platform launch instructions, document auto PyPI mirror detection,
the .venv_paddle / .venv_paddle_gpu layout, and link to
docs/linux-setup.md.
No code logic change; only docs + packaging so Linux is a first-class
release target alongside Windows and macOS.
The Linux one-click-start and setup_backend_linux scripts were both
assuming interactive stdin:
- one-click-start.sh called `read -r -p ...` on every fatal exit path
(non-Linux host, missing release files, install failure, missing
weight). When the script is launched from a desktop entry, systemd
user unit, CI, or any other context where stdin is not a TTY, `read`
would block waiting for input that will never come, hanging the
process instead of exiting promptly. Replace each site with a small
`pause_if_tty` helper that only prompts when `[ -t 0 ]`, so the
script still pauses for a user at a real terminal but returns
immediately when run unattended.
- scripts/setup_backend_linux.sh printed a single, hard-coded launch
block that only mentioned `.venv_paddle` (CPU). After the previous
--target auto/cpu/gpu/both refactor, the user could have built any
combination of CPU and GPU venvs but the hints always pointed at the
CPU one. Wrap the post-install message in `print_completion_hints`,
iterate over the actually-selected modes to print the correct
`<venv>/bin/python` for each (CPU -> .venv_paddle, GPU ->
.venv_paddle_gpu), and add an auto-mode command line when both
venvs exist or a hint to install the CPU fallback when only GPU
was built.
No behaviour change for the install flow itself; only the prompt
handling and the post-install guidance.
Previously invoke_setup splatted the script path with its arguments directly
(""${args[@]}""), which relies on the script having a valid shebang and
the executable bit set when invoked from a different process state or via
some shell wrappers. The Windows Invoke-Bootstrap counterpart launches
its bootstrap via a stable interpreter entry (powershell -File), and the
Linux path should be at least as robust.
- Resolve the setup script to an absolute path under $SCRIPT_DIR and run
it explicitly with 'bash "$setup_script" "${setup_args[@]}"', so the
shebang/exec bit is not on the critical path.
- Rename the local array to setup_args / setup_script to make intent
clearer at the call sites.
This keeps the recreate-on-existing-python behaviour from the previous
fix and matches the Windows pipeline's "stable interpreter invocation"
convention.
The Linux one-click-start invoke_setup helper only recreated the venv when
NEEDS_RECREATE was set by the portability/import check. If the target
python (e.g. <venv>/bin/python) already existed from a previous foreign
install or partial state, the setup script would try to reuse it and fail
in confusing ways. This mirrors the Windows Invoke-Bootstrap behaviour.
- invoke_setup now takes the selected venv python as $2, with the recreate
flag moved to $3 (call sites in the auto/cpu fallback branches updated).
- Recreate is forced when force_recreate=1 OR the target venv python path
already exists on disk, so any pre-existing python triggers --recreate.
- Comments added in Chinese to match the rest of the Linux script and
reference the Windows counterpart.
Fixes the case where running one-click-start.sh on a host that already
has a broken/foreign <venv>/bin/python leaves the user with a half-set-up
environment instead of a clean rebuild.
Linux one-click-start.sh and setup_backend_linux.sh previously installed
dependencies from the default PyPI source, which is often slow or blocked
in mainland China. Mirror the Windows pipeline's PyPI mirror behaviour by
probing well-known mirrors and falling back to the official source.
- New scripts/pypi_mirror.sh: ordered list of Tsinghua / Aliyun / USTC /
Tencent / pypi.org; pypi_mirror_probe checks each with curl or wget
(3s timeout); pypi_mirror_select returns the first reachable URL;
ensure_pypi_mirror_pip_args sets PIP_ARGS=("-i" <url>) only when the
caller has not already supplied --pip-arg. Reusable via 'source'.
- one-click-start.sh sources pypi_mirror.sh and calls
ensure_pypi_mirror_pip_args before invoking the backend installer, so a
bare one-click-start.sh on a CN host picks the fastest mirror.
- scripts/setup_backend_linux.sh does the same for direct invocations
(e.g. ./scripts/setup_backend_linux.sh), with the same skip-if-pip-arg-
set semantics. Header comment updated to document the behaviour.
Users who pass --pip-arg keep full control; everyone else gets a working
mirror automatically and the install no longer wedges on pypi.org.
Mirror the Windows one-click-start.cmd / setup_backend.ps1 flow for
Linux users:
- scripts/setup_backend_linux.sh: detect NVIDIA GPU, prefer uv for
Python 3.12 / venv / pip install of requirements-backend-{cpu,gpu}.txt,
run core-import smoke tests, verify yolo-captcha-detector.pt weight.
Supports --target auto/cpu/gpu/both, --recreate, --skip-install,
--no-smoke-test, --pip-arg.
- one-click-start.sh: Linux entry point that picks auto/cpu/gpu venv,
rebuilds foreign or broken venvs, falls back from GPU -> CPU on auto
failure, ensures a CPU fallback venv exists alongside GPU, and execs
scripts/tools/start_backend.py --headless with CNCAPTCHA_PORT and
CPU/GPU python overrides.
Together with one-click-start.cmd/.command this gives every platform a
single double-click path to a working captcha backend.
Greasy Fork script page description (docs/greasyfork-description.html) is what
users see on the script listing. Add the click-block risk control as the first
item in 重要提醒: blocked -> wait 10s or manually click 特惠订阅, OCR still auto.
One-liner: 订阅入口点不动就手动点特惠订阅,验证码自动打.
The standalone userscript README had no mention of the 2026-06-23 click-block
risk control or the config panel options. Add:
- Risk control section: blocked click -> wait 10s or manually click 特惠订阅,
OCR still auto. One-liner: 入口点不动就手动点特惠订阅,验证码自动打.
- Config panel guide: click delay, first-click delay, auto-close, auto-click.
- F8/F9 hotkeys that were missing.
The @description shows on Greasy Fork/Tampermonkey list and update notices -
first thing users see. Add: (1) PP-OCRv6 in parens, (2) the risk-control
workaround one-liner '订阅入口被风控拦截时手动点特惠订阅即可,验证码自动打'.
Zhipu upgraded risk control this morning: auto-click subscribe sometimes
blocked/rejected (manual click too), shows block_message. Advisory:
- wait ~10s and retry when blocked
- if auto-click keeps failing, manually click the '特惠订阅' entry
- OCR/captcha solving still fully automatic regardless of how the entry was
triggered (auto-click and OCR are decoupled)
One-liner: 入口点不动就手动点特惠订阅,验证码自动打。
OCR is now so fast (v6 tens of ms) that the first character click can land
before the captcha DOM finishes rendering/animating -> first click misses.
Previously the inter-click delay only applied between characters; the first
click had zero wait. Add a configurable pre-first-click random delay (default
150-300ms): new CAPTCHA_FIRST_CLICK_DELAY_MIN_MS/MAX_MS config, exposed in the
config panel under '验证码点字延时策略' as '首击前延时', passed through
getCaptchaDirectDelayConfig + normalize, slept at the start of
clickCaptchaDirectCoords. Set to 0 to disable.
Previous fix (run bootstrap in-process to avoid pipe buffering) broke
parameter binding: '& script.ps1 @argsList' splatting treated '-Target'
as the value of a positional param, not the param name -> ValidateSet
rejected it. Revert to external '& powershell -File bootstrap.ps1 @argsList'
(param binding reliable), drop the | Tee pipe (was the buffer root cause),
let child stdout inherit console directly for real-time pip progress.