chuckbutkusandopenhands 2ed2c571e6 fix(upload): resolve relative working dirs against /api/file/home (#1106)
* fix(upload): resolve relative working dirs against /api/file/home

The agent-server's /api/file/upload endpoint requires an absolute path
and `mkdir -p`'s the parent. The frontend's `toAbsoluteWorkspacePath`
was naively prepending `/` to the default relative `workspace/project`
working dir, producing `/workspace/project/<hex>/...`. On macOS and
fresh Docker images that path lives under a read-only filesystem root,
so uploads failed with `OSError: [Errno 30] Read-only file system:
'/workspace'`.

Conversations themselves kept working because the agent-server resolves
relative `workspace.working_dir` against its own process CWD (which is
writable in dev), so the worktree landed elsewhere and only uploads
mistargeted the read-only root.

Fix: introduce `getAgentServerHomeDir` (cached per backend, backed by
`FileClient.getHome` → `GET /api/file/home`) and a
`resolveAbsoluteAgentServerPath` helper. Both the conversation-start
payload and the file-upload destination now go through this resolver,
so they always agree on a single absolute path anchored at the
agent-server's home directory (e.g. `~/workspace/project/<hex>`).
Absolute paths pass through unchanged, so explicit workspace selections
and `VITE_WORKING_DIR` overrides are unaffected.

Spec: WUP-001 in specs/workspace-upload-path.md.

Co-authored-by: openhands <openhands@all-hands.dev>

* docs: clarify resolveConversationUploadWorkingDir returns raw (possibly-relative) working dir

Add JSDoc explaining that callers must funnel the result through
buildWorkspaceUploadPath (which calls resolveAbsoluteWorkspacePath) to
get an absolute path for the upload endpoint.  Also document that the
UNC path case is already covered by the existing isAbsolutePath regex.

Addresses review comment on PR #1106.

Co-authored-by: openhands <openhands@all-hands.dev>

* test(e2e): add mock-llm image-upload test

Adds an end-to-end mock-LLM test that exercises the full image-attachment
pipeline:

1. Attaches a minimal 1×1 PNG to the home-page chat via the hidden file
   input (data-testid="upload-image-input") using Playwright's
   setInputFiles.
2. Submits "What is in this image?" — creating a conversation and sending
   the message via sendMessageWithAttachments.
3. Verifies the agent replies with IMAGE_REPLY_TOKEN in the chat UI.
4. Verifies the user MessageEvent in the conversation events API has
   image_urls populated (base64 data: URL).
5. Verifies at least one /v1/chat/completions call to the mock server
   contained an image_url content block — confirming the image was
   forwarded to the LLM as expected.

Supporting changes:
- mock-llm-server.py: add GET /admin/requests endpoint that exposes all
  captured completion request bodies since the last reset; reset also
  clears the history.
- mock-llm-helpers.ts: add getMockLLMRequests(), IMAGE_REPLY_TOKEN,
  and MINIMAL_PNG_BASE64 exports.
- AGENTS.md: document the new endpoint and test spec.

Co-authored-by: openhands <openhands@all-hands.dev>

* test(e2e): add padding response to image-upload trajectory

The agent-server makes one internal LLM call for skill-analysis before
the main agent loop starts. The original 1-response trajectory was
consumed by that internal call, leaving the agent with a 500 error and
retry storm.

Add 1 padding response (turn 0: empty text) + 1 safety buffer (turn 2)
following the same pattern as mock-llm-automation.spec.ts.

Co-authored-by: openhands <openhands@all-hands.dev>

* test(e2e): use gpt-4o model name for vision-capable LLM requests

litellm strips image_url content blocks for unknown model names like
'openai/mock-test-model'. Switch to 'openai/gpt-4o' so litellm knows
the model accepts vision content and includes base64 image_url blocks
in the completion request body.

The base_url still points at the local mock server; the model name is
only a hint to litellm's request formatter.

Also improve assertion diagnostics to print all captured LLM requests
(not just the first) when the assertion fails.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-03 17:47:45 -04:00
2026-04-24 17:33:22 -04:00

agent-canvas

Warning

This project is in the Beta phase. It may be vibecoded, untested, or out of date. OpenHands takes no responsibility for the code or its support. Learn more.

Project Status: Beta

OpenHands is a platform for orchestrating coding agents across different environments. You can:

  • ⌨️ prompt agents manually
  • 🕐 run agents on a schedule
  • ⚡ trigger agents automatically — e.g. from Slack, GitHub, or Datadog.

Agents can run anywhere:

  • 🧑‍💻 on your laptop
  • 🖥️ on a remote virtual machine
  • ☁️ in our hosted cloud
  • 🏢 or inside your company’s infrastructure

The same Agent Canvas frontend can swap between each of these environments, so you can see everything in one place.

OpenHands works with any agent harness (e.g. Claude Code, Codex) or connect directly to an LLM (e.g. Anthropic, OpenAI, Gemini, Mistral, Minimax, Kimi).

If you have questions or feedback, please open a GitHub issue or join the #proj-agent-canvas channel in Slack.

Screenshot 2026-05-11 at 10 13 19 AM

Project ownership and support

  • Current status: Beta.
  • Support channel: #proj-agent-canvas.
  • Support level: Best effort while the project remains in Beta.

Quickstart

You can install OpenHands to run agents on any machine: on your laptop, on a dedicated computer like a Mac Mini, or on a server in the cloud.

The most powerful way to run OpenHands is on a server in the cloud. This allows your agents to continue running even when your laptop is shut, and makes it easier to trigger your agents through third-party services like Slack, GitHub, and Datadog. See SELF_HOSTING.md for details, especially with respect to security hardening.

Notably, you can run the backend in multiple different environments, and switch between them from the same Agent Canvas frontend. E.g. you can share an Agent Server with your team for agents doing code review and dependency updates, then have your personal agents running on your laptop.

Option 1: Without a Sandbox

Warning

This runs the agent-server directly on the machine you're installing on — the agent will have full access to your filesystem!

Prerequisites: Node.js 22.12.x or later, uv

npm install -g @openhands/agent-canvas
agent-canvas

The agent-canvas command starts the full local stack by default. You can also split it when you want to run pieces separately:

agent-canvas --frontend-only  # static frontend + ingress only
agent-canvas --backend-only   # agent server + automation backend + ingress only

Option 2: With a Docker Sandbox

Prerequisites:

  • Docker: Docker Desktop on macOS/Windows, or Docker Engine/Docker Desktop on Linux.
  • A host directory for PROJECTS_PATH containing the project folders you want the agent to access. Create it before starting the container.

macOS / Linux:

export PROJECTS_PATH="$HOME/projects"  # directory containing your project folders
mkdir -p "$PROJECTS_PATH" "$HOME/.openhands"

docker run -it --rm \
  -p 8000:8000 \
  -v "$HOME/.openhands:/home/openhands/.openhands" \
  -v "${PROJECTS_PATH}:/projects" \
  ghcr.io/openhands/agent-canvas:1.0.0-beta.6

Windows (PowerShell / Windows Terminal): See README.windows.md for the equivalent commands.

The agent will be able to access any project under PROJECTS_PATH.

Option 3: From Source

Warning

This runs the agent-server directly on the machine you're installing on — the agent will have full access to your filesystem!

Prerequisites: Node.js 22.12.x or later, npm, uv (for running the agent server via uvx)

git clone https://github.com/OpenHands/agent-canvas.git
cd agent-canvas
npm install
npm run dev

Access the UI at http://localhost:8000. You can add additional backends directly from the UI.

Architecture

Agent Canvas is powered by the OpenHands Agent Server, a REST API for running multiple agents on a single machine. Each Agent Server runs on a single host/port; the Agent Canvas can connect to multiple Agent Servers and easily flip between them.

You can run an Agent Server anywhere:

  • Directly on your laptop (be careful!)
  • On a dedicated machine like a Mac Mini
  • On a virtual machine in the cloud
  • Inside OpenHands Cloud (our commercial offering)

The Agent Server is often paired with an Automation Server, which lets you set up agents that run on a schedule or in response to events.

image

More documentation

S
Languages
TypeScript 93.7%
JavaScript 4.7%
Python 0.9%
Shell 0.3%
CSS 0.2%
Other 0.1%