Rohit Malhotraandopenhands 92b25e0af3 ci: selective E2E test execution based on changed files (#1286)
* ci: add paths filters to E2E workflows to skip irrelevant PRs

Add paths: filters to the pull_request triggers of the three E2E
workflows so they are skipped when a PR only touches files that
cannot affect the test suite (docs, specs, .agents/, unrelated
test directories, etc.).

- mock-llm-e2e.yml: triggers on src/, public/, scripts/, bin/,
  config/, tests/e2e/mock-llm/, tests/e2e/support/, package.json,
  package-lock.json, build/TS/styling configs, and its own workflow
  file.

- snapshot-tests.yml: triggers on src/, public/,
  tests/e2e/snapshots/, tests/e2e/support/, package.json,
  package-lock.json, build/TS/styling/playwright configs, and its
  own workflow file. Both pull_request and push-to-main triggers
  are filtered with the same path set.

- mock-llm-docker-e2e.yml: same paths as mock-llm-e2e.yml plus
  docker/** and playwright.mock-llm-docker.config.ts. The
  workflow_run trigger (post-Docker-build on main) is unaffected
  by path filters and always runs.

workflow_dispatch is unaffected by paths: filters in all three
workflows, so a manual run always executes the full suite.

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: organize mock-LLM E2E tests into feature subdirectories with selective execution

Reorganize the 15 mock-LLM spec files from a flat directory into feature
subdirectories that mirror the source code structure:

  tests/e2e/mock-llm/
    settings/    — LLM profiles, ACP agent, model switching
    conversations/ — core conversation flow, image upload
    automations/ — automation lifecycle, preset cards
    onboarding/  — first-run onboarding flow
    backends/    — auth modes, cross-connect, partial stack
    home/        — workspace selection, folder browser
    skills/      — skill loading and activation
    regressions/ — CSS isolation, event pagination, etc.

Add a test-mapping config (test-mapping.json) and resolver script
(scripts/resolve-affected-tests.mjs) that maps changed source files
to the affected test subdirectories. The resolver has three modes:

  1. Feature-isolated changes (e.g. src/components/features/settings/**)
     → run only the mapped subdirs + regressions
  2. Cross-cutting changes (src/api/**, package.json, shared helpers,
     or any unmapped src/ file) → run the full suite (__ALL__)
  3. Non-relevant changes (docs, specs) → nothing (workflow paths
     filter already skipped)

The mock-llm-e2e.yml workflow now has a 'Resolve affected test
directories' step that queries PR changed files via the GitHub API,
runs the resolver, and passes the result to Playwright. workflow_dispatch
always runs the full suite.

All relative imports in moved spec files are updated. The Playwright
config discovers specs recursively so no config change is needed.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: update PROJECT_ROOT paths in backend specs moved to subdirectory

The partial-stack and cross-connect specs resolve PROJECT_ROOT from
import.meta.url using ../../.. (3 levels). After moving them from
tests/e2e/mock-llm/ to tests/e2e/mock-llm/backends/, they need
../../../.. (4 levels) to reach the repo root. Without this fix,
bin/agent-canvas.mjs resolves to a nonexistent path and the
backend-only test fails with MODULE_NOT_FOUND.

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: include new mock-LLM specs in selective runs

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: fail closed for mock e2e selection

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: avoid pending skipped e2e checks

Co-authored-by: openhands <openhands@all-hands.dev>

* test: align folder workspace e2e with auto-selection

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-13 00:30:09 -04:00
2026-04-24 17:33:22 -04:00

agent-canvas

Warning

This project is in the Beta phase. It may be vibecoded, untested, or out of date. OpenHands takes no responsibility for the code or its support. Learn more.

Project Status: Beta

OpenHands is a platform for orchestrating coding agents across different environments. You can:

  • ⌨️ prompt agents manually
  • 🕐 run agents on a schedule
  • ⚡ trigger agents automatically — e.g. from Slack, GitHub, or Datadog.

Agents can run anywhere:

  • 🧑‍💻 on your laptop
  • 🖥️ on a remote virtual machine
  • ☁️ in our hosted cloud
  • 🏢 or inside your company’s infrastructure

The same Agent Canvas frontend can swap between each of these environments, so you can see everything in one place.

OpenHands works with any agent harness (e.g. Claude Code, Codex) or connect directly to an LLM (e.g. Anthropic, OpenAI, Gemini, Mistral, Minimax, Kimi).

If you have questions or feedback, please open a GitHub issue or join the #proj-agent-canvas channel in Slack.

Screenshot 2026-05-11 at 10 13 19 AM

Project ownership and support

  • Current status: Beta.
  • Support channel: #proj-agent-canvas.
  • Support level: Best effort while the project remains in Beta.

Quickstart

You can install OpenHands to run agents on any machine: on your laptop, on a dedicated computer like a Mac Mini, or on a server in the cloud.

The most powerful way to run OpenHands is on a server in the cloud. This allows your agents to continue running even when your laptop is shut, and makes it easier to trigger your agents through third-party services like Slack, GitHub, and Datadog. See SELF_HOSTING.md for details, especially with respect to security hardening.

Notably, you can run the backend in multiple different environments, and switch between them from the same Agent Canvas frontend. E.g. you can share an Agent Server with your team for agents doing code review and dependency updates, then have your personal agents running on your laptop.

Option 1: Without a Sandbox

Warning

This runs the agent-server directly on the machine you're installing on — the agent will have full access to your filesystem!

Prerequisites: Node.js 22.12.x or later, uv

npm install -g @openhands/agent-canvas
agent-canvas

The agent-canvas command starts the full local stack by default. You can also split it when you want to run pieces separately:

agent-canvas --frontend-only  # static frontend + ingress only
agent-canvas --backend-only   # agent server + automation backend + ingress only

Option 2: With a Docker Sandbox

Prerequisites:

  • Docker: Docker Desktop on macOS/Windows, or Docker Engine/Docker Desktop on Linux.
  • A host directory for PROJECTS_PATH containing the project folders you want the agent to access. Create it before starting the container.

macOS / Linux:

export PROJECTS_PATH="$HOME/projects"  # directory containing your project folders
mkdir -p "$PROJECTS_PATH" "$HOME/.openhands"

docker run -it --rm \
  -p 8000:8000 \
  -v "$HOME/.openhands:/home/openhands/.openhands" \
  -v "${PROJECTS_PATH}:/projects" \
  ghcr.io/openhands/agent-canvas:1.0.0-rc.6

Windows (PowerShell / Windows Terminal): See README.windows.md for the equivalent commands.

The agent will be able to access any project under PROJECTS_PATH.

Option 3: From Source

Warning

This runs the agent-server directly on the machine you're installing on — the agent will have full access to your filesystem!

Prerequisites: Node.js 22.12.x or later, npm, uv (for running the agent server via uvx)

git clone https://github.com/OpenHands/agent-canvas.git
cd agent-canvas
npm install
npm run dev

Access the UI at http://localhost:8000. You can add additional backends directly from the UI.

Architecture

Agent Canvas is powered by the OpenHands Agent Server, a REST API for running multiple agents on a single machine. Each Agent Server runs on a single host/port; the Agent Canvas can connect to multiple Agent Servers and easily flip between them.

You can run an Agent Server anywhere:

  • Directly on your laptop (be careful!)
  • On a dedicated machine like a Mac Mini
  • On a virtual machine in the cloud
  • Inside OpenHands Cloud (our commercial offering)

The Agent Server is often paired with an Automation Server, which lets you set up agents that run on a schedule or in response to events.

image

More documentation

S
Languages
TypeScript 93.7%
JavaScript 4.7%
Python 0.9%
Shell 0.3%
CSS 0.2%
Other 0.1%