Rohit Malhotraandopenhands 86f3fcfa8e feat: add mock-LLM e2e test for automation creation and run dispatch (#905)
* feat: add mock-LLM e2e test for automation creation and run dispatch

Add a new e2e test spec (mock-llm-automation.spec.ts) that exercises the
full automation lifecycle through the UI and mock LLM:

1. Navigates to the Automations page
2. Clicks 'Add Automation' → 'Create Automation' to launch a conversation
3. Mock LLM returns scripted terminal tool calls that:
   - curl POST to create a cron automation (echo hello world at 9am daily)
   - curl POST to dispatch a run using the created automation ID
4. Verifies the automation was created correctly (name, schedule, enabled)
5. Verifies the dispatched run reaches COMPLETED status
6. Verifies the automation appears on the automations page

Supporting changes:

- Mock automation server (mock-automation-server.py): Lightweight Python
  HTTP server implementing automation API endpoints in-memory with
  auto-completing runs (~0.5s PENDING → RUNNING → COMPLETED)

- Mock LLM server admin API: Added /admin/reset, /admin/trajectory/register,
  and /admin/trajectory/activate endpoints so tests can inject custom
  trajectories per test scenario

- Playwright config: Added mock automation server as additional webServer
  (port 18299)

- Test helpers: Added routeAutomationApiToMock() (Playwright page.route
  proxy), ensureMockLLMProfile() (API-based profile setup), trajectory
  registration/activation helpers, and automation verification helpers

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: scope modal button click and reset LLM trajectory between test suites

- Scope 'Create Automation' click to the modal element to avoid strict
  mode violation (empty-state page also renders a button with the same
  testId)
- Reset mock LLM to default trajectory at the start of conversation
  test step 3, preventing failures when the automation test's custom
  trajectory wasn't fully consumed due to earlier test failures

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use home chat launcher instead of modal auto-submit for automation test

The Automations modal's 'Create Automation' button uses
useLaunchSkillInChat which navigates to /conversations and sets
messageToSend in the Zustand store. However, the home page's
ChatInputLogic explicitly nullifies messageToSend when there is no
active conversationId, so the message never auto-submits.

Rewrite step 2 to type the prompt directly into the home chat launcher
and click submit — the same path real users take. This bypasses the
broken modal→store→auto-submit chain and reliably creates a
conversation.

Co-authored-by: openhands <openhands@all-hands.dev>

* refactor: use real automation backend instead of mock server

Remove the mock automation server and Playwright route interception.
The automation test now hits the real automation backend running inside
the bin/agent-canvas.mjs stack (through the ingress proxy).

Key changes:
- Terminal curl commands use $OPENHANDS_AUTOMATION_API_KEY for auth
  and hit the ingress URL for the automation API
- Verification queries go through the real /api/automation/v1 endpoints
  with X-Session-API-Key auth
- Removed mock-automation-server.py and all mock automation helpers
- Removed MOCK_AUTOMATION_PORT from Playwright config
- Test is fully e2e: mock LLM → real agent-server → real automation
  backend

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: add retry logic for automation backend startup delay

The Playwright webServer health check passes once the ingress serves
the static frontend, but the automation backend (started via uvx) may
still be initializing. This caused 502 errors when the test tried to
list automations immediately.

Changes:
- listAutomations retries on 502/503 up to 30 times (60s)
- listAutomationRuns returns empty on non-OK responses (waitForRunStatus
  retries anyway)
- Step 1 waits for the automation backend to be ready before registering
  the trajectory
- Increased timeouts for steps 1 (120s) and 2 (180s)

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: wait for automation backend before starting tests

Change the Playwright webServer URL check from the ingress root (/) to
the automation API endpoint (/api/automation/v1). This ensures Playwright
waits for the full stack — agent-server + automation backend + ingress —
to be ready before running tests.

The automation backend is the last service to start (installed via uvx
from PyPI) and can take 30-60s in CI. The previous root URL check only
verified the ingress served static files, allowing tests to begin while
the automation backend was still initializing (returning 502).

Playwright accepts 2xx, 3xx, and 400-403 as 'ready' responses, so the
401 from an unauthenticated GET to the automation endpoint correctly
signals readiness.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: pre-warm automation backend in CI and fix 0-test reporter

Two fixes for the mock-LLM E2E automation test:

1. CI workflow: pre-install openhands-automation into the uvx cache
   before the Playwright run starts. Without this, the automation
   backend's first-time uvx install (~60-90s for boto3, google-cloud,
   etc.) exceeded the 180s Playwright webServer timeout.

2. Done-marker reporter: treat 0 completed tests as a failure. The
   previous logic defaulted allPassed=true, so a webServer timeout
   (0 tests ran) wrote .all-passed and the CI wrapper reported success.

3. Playwright webServer URL probes the automation endpoint through the
   ingress (/api/automation/v1) instead of the root (/). This ensures
   ALL services are up before tests start — the automation backend is
   the last to start.

Co-authored-by: openhands <openhands@all-hands.dev>

* debug: capture agent-canvas startup logs for CI diagnostics

Tee the agent-canvas binary output to a log file that gets uploaded
as a CI artifact. This exposes the full startup sequence including
agent-server, automation backend, and ingress startup messages that
are currently invisible due to Playwright's single-line ANSI rendering.

Also bumped webServer timeout to 300s and global timeout to 900s.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use absolute path for state dir to fix automation DB crash

Root cause: the automation backend's SQLite DB URL was derived from a
RELATIVE state dir path (.tmp/mock-llm-state). Since the automation
child process also has its cwd set to the state dir, the DB path
resolved to the nested path:
  .tmp/mock-llm-state/.tmp/mock-llm-state/automations.db
which doesn't exist → sqlite3.OperationalError → backend never starts.

Fix: resolve() the state dir to an absolute path so the DB URL works
regardless of the child process cwd.

Also reverts the diagnostic tee/timeout changes from the previous commit
since the root cause is now identified.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use X-Session-API-Key for automation API auth in curl commands

The automation backend authenticates via X-Session-API-Key header (set by
AUTOMATION_LOCAL_API_KEY env var), matching the frontend's automation
service. The previous Authorization: Bearer header was not accepted.

The env var $OPENHANDS_AUTOMATION_API_KEY carries the session API key
value (injected by buildAgentServerAutomationEnv in dev-with-automation.mjs)
and is inherited by the terminal tool's bash subprocess.

Co-authored-by: openhands <openhands@all-hands.dev>

* debug: add verbose curl output to diagnose automation API failures

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: hardcode session API key in curl commands

The agent-server terminal tool may not inherit all parent process env
vars (the SDK sandboxes the execution environment). Hardcoding the
session API key directly in the curl commands ensures authentication
works regardless of the terminal's env setup.

Co-authored-by: openhands <openhands@all-hands.dev>

* debug: dump conversation events to diagnose terminal command output

Add a diagnostic step that fetches all conversation events via the
API and logs terminal command outputs, so we can see what the curl
command actually returned inside the agent-server terminal.

Co-authored-by: openhands <openhands@all-hands.dev>

* debug: dump raw event JSON for diagnostics

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: add padding response for internal pre-agent LLM call

The agent-server makes an internal LLM call (likely condenser or
skill analysis) before the agent's main loop starts, which consumes
one scripted response. Add a throwaway empty text response at the
start of the automation trajectory so the agent's first real turn
gets the correct create command.

Also improves diagnostic event dump to show raw JSON.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: accept any run status for dispatched automation verification

The dispatched automation run won't reach COMPLETED in mock LLM mode
because the automation's conversation needs additional LLM responses
that would exhaust the mock server. Verify that a run was dispatched
(exists with any valid status) instead of waiting for COMPLETED.

Also adds waitForAnyRun helper and removes diagnostic event dump.

Key fixes in this commit series:
- Padding response for internal pre-agent LLM call (condenser)
- Hardcoded session API key in curl commands
- X-Session-API-Key auth header (matching frontend convention)
- waitForAnyRun instead of waitForRunStatus COMPLETED

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: move mock LLM reset to afterEach; fix automations page wait

- Move resetMockLLM() from a cleanup test into afterEach so it always
  runs even when preceding serial tests fail. This prevents the
  conversation test suite from getting exhausted mock LLM responses.

- Replace instant page.textContent check with Playwright's auto-retry
  expect(getByText).toBeVisible for the automations page load.

- Remove now-redundant cleanup test.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: defer automation cleanup to step 3 so list page verification works

afterEach was deleting the automation after step 2, so step 3's
navigation to /automations found an empty list. Move automation
deletion into step 3's final cleanup sub-step. Conversations and
mock LLM reset still happen in afterEach.

Co-authored-by: openhands <openhands@all-hands.dev>

* docs: update AGENTS.md with automation e2e learnings

Document the real automation backend approach, padding response
requirement for skill-activated conversations, and afterEach mock
LLM reset pattern.

Co-authored-by: openhands <openhands@all-hands.dev>

* feat: verify run COMPLETED + conversation link click-through

Strengthen the automation e2e test to verify the full lifecycle:

- Add extra mock LLM responses (indices 4-6) for the automation run's
  spawned conversation so it can complete and fire the callback
- Wait for run status COMPLETED (not just any status)
- Assert the completed run has a conversation_id
- Navigate to automation detail page and verify the COMPLETED badge
- Click the run's conversation link and verify it navigates to
  /conversations/{id}

This exercises:
  1. Automation creation via terminal curl → real backend
  2. Run dispatch → real backend starts a new conversation
  3. Run completion → automation script fires callback → COMPLETED
  4. UI list page → detail page → run conversation link click-through

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use data-testid for run status badge, not translated text

RunStatusBadge renders i18n 'Successful' (not 'Completed') for
completed runs. Use the stable data-testid='run-status-icon-completed'
instead of fragile text matching.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: handle multiple conversation link elements in run detail

The automation detail page may render multiple <a> elements with the
same conversation href (e.g. header + activity row). Use .first() to
select and click the first matching link.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: address review comments

1. mock-llm-server.py: add explicit 404 for unknown POST paths so
   typos in admin endpoints produce a clear error instead of silently
   falling through to the LLM completion handler.

2. mock-llm-automation.spec.ts: expand the padding response comment
   to document the specific agent-server feature (skill-activation
   pipeline) that triggers it, why the conversation test doesn't need
   it, and how misalignment manifests as a fast timeout failure.

3. mock-llm-automation.spec.ts: add test.afterAll safety net to
   delete leftover automations even if step 3 fails before its
   cleanup sub-step runs.

4. mock-llm-helpers.ts: include active_profile in the PATCH payload
   so ensureMockLLMProfile actually guarantees the named profile is
   activated, not just that settings are applied to whatever profile
   happens to be active.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: address second round of review comments

1. Fix stale JSDoc on listAutomations — now correctly notes the
   health check probes /api/automation/v1 and retries are a safety net.

2. listAutomationRuns throws immediately on non-retriable HTTP errors
   (401, 500, etc.) instead of silently returning empty results that
   cause confusing timeouts. Only 502/503 are treated as retriable.

3. Warn on malformed trajectory turns — _parse_trajectory_turns now
   prints a stderr warning when a turn has neither 'tool_call' nor
   'text', making typos like 'tool_calls' visible at registration
   time instead of producing a confusing 'Mock LLM exhausted' error.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: revert active_profile in PATCH — breaks conversation test profile flow

The agent-server's PATCH /api/settings with active_profile created a
profile via API that conflicted with the conversation test's UI-based
profile creation flow. The conversation test's step 2 ('Set as active')
failed because the profile state was inconsistent.

Reverted to the original approach: ensureMockLLMProfile configures LLM
settings on whatever profile is currently active, without creating or
switching profiles. Updated the early-return check to compare model and
base_url instead of profile name, and expanded the JSDoc to clarify
this is NOT a profile management function.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: address third round of review comments

1. Confirm probe URL returns 200 without auth — added comment in
   playwright.mock-llm.config.ts explaining the automation list
   endpoint serves 200 unauthenticated (confirmed in CI).

2. _read_body JSON error handling — wrap json.loads in try/except
   and return 400 invalid_json on malformed payloads. Replace
   __import__('threading') with a normal top-level import.

3. Remove unasserted token constants — AUTOMATION_CREATE_TOKEN and
   AUTOMATION_DISPATCH_TOKEN were never asserted; inline the printf
   breadcrumbs and add a comment clarifying AUTOMATION_REPLY_TOKEN
   is the only asserted token.

4. Capture remaining_responses inside the lock — consistent with
   the threaded-safety pattern even though HTTPServer is single-
   threaded. Applied to both /admin/reset and /admin/trajectory/activate.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: address fourth round of review comments

1. Add explicit assertion that runConversationId is set before the
   click-through verification in step 3 — prevents silent pass when
   step 2 fails before populating the ID.

2. Fix _read_body double-response on malformed JSON: return None
   instead of {} on parse failure, add guards in both callers
   (/admin/trajectory/register and /admin/trajectory/activate) to
   return early when body is None (error already sent).

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: address fifth round of review suggestions

1. /admin/reset now clears _named_trajectories so the server
   returns to full initial state.

2. Malformed trajectory turns raise ValueError immediately at
   registration time instead of silently skipping — caught in
   the register handler and returned as 400 bad_request.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: move resetMockLLM from afterEach to afterAll

The /admin/reset endpoint now clears _named_trajectories (previous
fix), which broke the serial step flow: step 1 registers the
trajectory, afterEach cleared it, step 2 tried to activate it → 404.

Moving the reset to afterAll preserves named trajectories across
the serial steps while still cleaning up after the full suite.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: address PR review feedback (#905)

- Document cumulative wait budget for step 2 timeout (180s) and note
  fallback to 240s if CI proves flaky
- Add clarifying comment for defensive re-activation of trajectory
  at the start of step 2 (belt-and-suspenders pattern)

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-29 16:06:33 +00:00
2026-04-24 17:33:22 -04:00

agent-canvas

Warning

This project is in alpha phase. It may be vibecoded, untested, or out of date. Learn more.

OpenHands is a platform for orchestrating coding agents across different environments. You can:

  • ⌨️ prompt agents manually
  • 🕐 run agents on a schedule
  • ⚡ trigger agents automatically — e.g. from Slack, GitHub, or Datadog.

Agents can run anywhere:

  • 🧑‍💻 on your laptop
  • 🖥️ on a remote virtual machine
  • ☁️ in our hosted cloud
  • 🏢 or inside your company’s infrastructure

The same Agent Canvas frontend can swap between each of these environments, so you can see everything in one place.

OpenHands works with any agent harness (e.g. Claude Code, Codex) or connect directly to an LLM (e.g. Anthropic, OpenAI, Gemini, Mistral, Minimax, Kimi).

If you have questions or feedback, please open a GitHub issue or join the #proj-agent-canvas channel in Slack

Screenshot 2026-05-11 at 10 13 19 AM

Quickstart

You can install OpenHands to run agents on any machine: on your laptop, on a dedicated computer like a Mac Mini, or on a server in the cloud.

The most powerful way to run OpenHands is on a server in the cloud. This allows your agents to continue running even when your laptop is shut, and makes it easier to trigger your agents through third-party services like Slack, GitHub, and Datadog. See SELF_HOSTING.md for details, especially with respect to security hardening.

Notably, you can run the backend in multiple different environments, and switch between them from the same Agent Canvas frontend. E.g. you can share an Agent Server with your team for agents doing code review and dependency updates, then have your personal agents running on your laptop.

Option 1: Without a Sandbox

Warning

This runs the agent-server directly on the machine you're installing on — the agent will have full access to your filesystem!

Prerequisites: Node.js 22.12.x or later, uv

npm install -g @openhands/agent-canvas
agent-canvas

Option 2: With a Docker Sandbox

docker pull ghcr.io/openhands/agent-canvas:1.0.0-alpha.8

export PROJECTS_PATH=~/projects  # directory containing your project folders

docker run -it --rm \
  -p 8000:8000 \
  -v ~/.openhands:/home/openhands/.openhands \
  -v ${PROJECTS_PATH}:/projects \
  ghcr.io/openhands/agent-canvas:1.0.0-alpha.8

The agent will be able to access any project under PROJECTS_PATH.

Option 3: From Source

Prerequisites: Node.js 22.12.x or later, npm, uv (for running the agent server via uvx)

git clone https://github.com/OpenHands/agent-canvas.git
cd agent-canvas
npm install
npm run dev

Access the UI at http://localhost:8000. You can add additional backends directly from the UI.

Architecture

Agent Canvas is powered by the OpenHands Agent Server, a REST API for running multiple agents on a single machine. Each Agent Server runs on a single host/port; the Agent Canvas can connect to multiple Agent Servers and easily flip between them.

You can run an Agent Server anywhere:

  • Directly on your laptop (be careful!)
  • On a dedicated machine like a Mac Mini
  • On a virtual machine in the cloud
  • Inside OpenHands Cloud (our commercial offering)

The Agent Server is often paired with an Automation Server, which lets you set up agents that run on a schedule or in response to events.

image

More documentation

For contributor and developer workflows, including frontend-only mode, mock mode, environment variables, and build/test commands, see DEVELOPMENT.md.

S
Languages
TypeScript 93.7%
JavaScript 4.7%
Python 0.9%
Shell 0.3%
CSS 0.2%
Other 0.1%