mirror of
https://github.com/OpenHands/OpenHands.git
synced 2026-10-07 12:58:49 +08:00
* fix: bind static-server dual-stack to fix Docker E2E ECONNREFUSED on ::1 The Docker mock-LLM E2E tests were flaky because static-server.mjs bound to 0.0.0.0 (IPv4 only), but Playwright and Chromium often resolve `localhost` to ::1 (IPv6) on Ubuntu CI runners, causing intermittent ECONNREFUSED. The non-Docker tests didn't have this problem because they go through ingress.mjs, which calls server.listen(port) without a host argument — Node.js defaults to :: (dual-stack: IPv4 + IPv6). Changes: - static-server.mjs: default host from "0.0.0.0" to null; when null, call server.listen(port) without host so Node binds to :: - docker/entrypoint.sh: drop --host 0.0.0.0 from both static-server invocations so they use the dual-stack default - playwright.mock-llm.config.ts: drop --host 0.0.0.0 from the public-mode static server (tests hit it directly via localhost) Callers behind ingress.mjs (dev-with-automation, dev-static) still pass --host 0.0.0.0 explicitly, which is fine since the ingress itself is already dual-stack. Co-authored-by: openhands <openhands@all-hands.dev> * fix: use 127.0.0.1 in Docker E2E config to avoid IPv6 ECONNREFUSED The Docker mock-LLM E2E tests fail with ECONNREFUSED ::1:18300 because Playwright/Chromium resolve `localhost` to ::1 (IPv6) on Ubuntu CI runners, but the Docker container's static-server binds to 0.0.0.0 (IPv4 only). The non-Docker tests don't have this issue because they go through ingress.mjs which binds to :: (dual-stack: IPv4 + IPv6). Rather than changing the server binding (which could have side-effects inside the Docker container), this fix changes the Docker Playwright config to use 127.0.0.1 directly for all URLs: INGRESS_URL, MOCK_LLM_BACKEND_URL, MOCK_LLM_PUBLIC_MODE_URL, and the webServer health-check probe. This bypasses DNS resolution entirely and connects via IPv4. Co-authored-by: openhands <openhands@all-hands.dev> * fix: keep Docker container alive when a backend service exits The Docker entrypoint used `wait -n` which exits the entire container when ANY child process exits. After heavy automation tests (which spawn multiple conversations), the agent-server or automation backend could exit, taking down the static-server with it — causing ECONNREFUSED for subsequent tests. In the non-Docker path, each service is an independent host process, so one crashing doesn't affect the others. The ingress proxy returns 502 for the dead backend but stays up. Change the entrypoint to monitor children in a loop: log crashes but keep the container running as long as any service is still alive. Only exit when ALL tracked children are dead. The SIGTERM trap still handles clean shutdown. Co-authored-by: openhands <openhands@all-hands.dev> * fix: simplify entrypoint keep-alive to avoid wait -n interaction issues Replace the complex PID monitoring loop with a simple sleep loop. The previous wait -n based loop regressed automation tests (5/14 vs 8/14 on the simpler IPv4-only commit), likely due to bash wait -n signal handling interacting poorly with child processes. The sleep loop keeps the container alive indefinitely. The existing SIGTERM/SIGINT trap handles clean shutdown. Co-authored-by: openhands <openhands@all-hands.dev> * fix: address review feedback on entrypoint and docs - Wait on STATIC_PID only (not all children): the static-server is the critical ingress process. If it dies the container exits with a meaningful exit code. Backend crashes (agent-server, automation) are tolerated — the proxy returns 502. (Copilot feedback) - Add `exit 0` to cleanup(): ensures the script terminates after a SIGTERM-triggered trap return instead of falling through. The `wait` builtin is signal-interruptible so SIGTERM is processed immediately — no stale sleep blocking delivery. (all-hands-bot) - Update AGENTS.md to match the actual behavior (wait on static-server PID, not a monitoring loop or infinite sleep). (Copilot feedback) Co-authored-by: openhands <openhands@all-hands.dev> * fix: use signal-safe sleep-wait loop with static-server liveness check The bare `wait "$STATIC_PID"` approach failed the same way as the original `wait -n` (8/14 — conversation/model-switch tests get ECONNREFUSED after automation). The `while true; do sleep 86400; done` pattern from the previous commit was the only one that passed 14/14, but had two issues flagged in review: 1. Bare `sleep` as foreground blocks SIGTERM delivery (all-hands-bot) 2. Container stays alive forever even if static-server dies (Copilot) This commit addresses both: - `sleep 10 & wait $!` — `wait` (builtin) is the foreground op, so SIGTERM interrupts it immediately and the trap fires. - `while kill -0 "$STATIC_PID"` — loop exits when the critical ingress process dies; container exits with a meaningful code. - `cleanup()` keeps `exit 0` so the script terminates after a signal-triggered trap return. Co-authored-by: openhands <openhands@all-hands.dev> --------- Co-authored-by: openhands <openhands@all-hands.dev>