Files
OpenHands/docker
Rohit Malhotraandopenhands e408deea9e fix: bind static-server dual-stack to fix Docker E2E ECONNREFUSED on ::1 (#1032)
* fix: bind static-server dual-stack to fix Docker E2E ECONNREFUSED on ::1

The Docker mock-LLM E2E tests were flaky because static-server.mjs
bound to 0.0.0.0 (IPv4 only), but Playwright and Chromium often
resolve `localhost` to ::1 (IPv6) on Ubuntu CI runners, causing
intermittent ECONNREFUSED.

The non-Docker tests didn't have this problem because they go
through ingress.mjs, which calls server.listen(port) without a
host argument — Node.js defaults to :: (dual-stack: IPv4 + IPv6).

Changes:
- static-server.mjs: default host from "0.0.0.0" to null; when
  null, call server.listen(port) without host so Node binds to ::
- docker/entrypoint.sh: drop --host 0.0.0.0 from both
  static-server invocations so they use the dual-stack default
- playwright.mock-llm.config.ts: drop --host 0.0.0.0 from the
  public-mode static server (tests hit it directly via localhost)

Callers behind ingress.mjs (dev-with-automation, dev-static) still
pass --host 0.0.0.0 explicitly, which is fine since the ingress
itself is already dual-stack.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use 127.0.0.1 in Docker E2E config to avoid IPv6 ECONNREFUSED

The Docker mock-LLM E2E tests fail with ECONNREFUSED ::1:18300
because Playwright/Chromium resolve `localhost` to ::1 (IPv6) on
Ubuntu CI runners, but the Docker container's static-server binds to
0.0.0.0 (IPv4 only).

The non-Docker tests don't have this issue because they go through
ingress.mjs which binds to :: (dual-stack: IPv4 + IPv6).

Rather than changing the server binding (which could have
side-effects inside the Docker container), this fix changes the
Docker Playwright config to use 127.0.0.1 directly for all URLs:
INGRESS_URL, MOCK_LLM_BACKEND_URL, MOCK_LLM_PUBLIC_MODE_URL, and
the webServer health-check probe. This bypasses DNS resolution
entirely and connects via IPv4.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: keep Docker container alive when a backend service exits

The Docker entrypoint used `wait -n` which exits the entire
container when ANY child process exits. After heavy automation
tests (which spawn multiple conversations), the agent-server or
automation backend could exit, taking down the static-server with
it — causing ECONNREFUSED for subsequent tests.

In the non-Docker path, each service is an independent host process,
so one crashing doesn't affect the others. The ingress proxy
returns 502 for the dead backend but stays up.

Change the entrypoint to monitor children in a loop: log crashes
but keep the container running as long as any service is still
alive. Only exit when ALL tracked children are dead. The SIGTERM
trap still handles clean shutdown.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: simplify entrypoint keep-alive to avoid wait -n interaction issues

Replace the complex PID monitoring loop with a simple sleep loop.
The previous wait -n based loop regressed automation tests (5/14
vs 8/14 on the simpler IPv4-only commit), likely due to bash
wait -n signal handling interacting poorly with child processes.

The sleep loop keeps the container alive indefinitely. The existing
SIGTERM/SIGINT trap handles clean shutdown.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: address review feedback on entrypoint and docs

- Wait on STATIC_PID only (not all children): the static-server is
  the critical ingress process. If it dies the container exits with
  a meaningful exit code. Backend crashes (agent-server, automation)
  are tolerated — the proxy returns 502.  (Copilot feedback)

- Add `exit 0` to cleanup(): ensures the script terminates after a
  SIGTERM-triggered trap return instead of falling through. The
  `wait` builtin is signal-interruptible so SIGTERM is processed
  immediately — no stale sleep blocking delivery.  (all-hands-bot)

- Update AGENTS.md to match the actual behavior (wait on static-server
  PID, not a monitoring loop or infinite sleep).  (Copilot feedback)

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use signal-safe sleep-wait loop with static-server liveness check

The bare `wait "$STATIC_PID"` approach failed the same way as
the original `wait -n` (8/14 — conversation/model-switch tests
get ECONNREFUSED after automation).  The `while true; do sleep
86400; done` pattern from the previous commit was the only one
that passed 14/14, but had two issues flagged in review:

1. Bare `sleep` as foreground blocks SIGTERM delivery (all-hands-bot)
2. Container stays alive forever even if static-server dies (Copilot)

This commit addresses both:
- `sleep 10 & wait $!` — `wait` (builtin) is the foreground op,
  so SIGTERM interrupts it immediately and the trap fires.
- `while kill -0 "$STATIC_PID"` — loop exits when the critical
  ingress process dies; container exits with a meaningful code.
- `cleanup()` keeps `exit 0` so the script terminates after a
  signal-triggered trap return.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-02 18:33:39 +00:00
..