* fix(upload): resolve relative working dirs against /api/file/home The agent-server's /api/file/upload endpoint requires an absolute path and `mkdir -p`'s the parent. The frontend's `toAbsoluteWorkspacePath` was naively prepending `/` to the default relative `workspace/project` working dir, producing `/workspace/project/<hex>/...`. On macOS and fresh Docker images that path lives under a read-only filesystem root, so uploads failed with `OSError: [Errno 30] Read-only file system: '/workspace'`. Conversations themselves kept working because the agent-server resolves relative `workspace.working_dir` against its own process CWD (which is writable in dev), so the worktree landed elsewhere and only uploads mistargeted the read-only root. Fix: introduce `getAgentServerHomeDir` (cached per backend, backed by `FileClient.getHome` → `GET /api/file/home`) and a `resolveAbsoluteAgentServerPath` helper. Both the conversation-start payload and the file-upload destination now go through this resolver, so they always agree on a single absolute path anchored at the agent-server's home directory (e.g. `~/workspace/project/<hex>`). Absolute paths pass through unchanged, so explicit workspace selections and `VITE_WORKING_DIR` overrides are unaffected. Spec: WUP-001 in specs/workspace-upload-path.md. Co-authored-by: openhands <openhands@all-hands.dev> * docs: clarify resolveConversationUploadWorkingDir returns raw (possibly-relative) working dir Add JSDoc explaining that callers must funnel the result through buildWorkspaceUploadPath (which calls resolveAbsoluteWorkspacePath) to get an absolute path for the upload endpoint. Also document that the UNC path case is already covered by the existing isAbsolutePath regex. Addresses review comment on PR #1106. Co-authored-by: openhands <openhands@all-hands.dev> * test(e2e): add mock-llm image-upload test Adds an end-to-end mock-LLM test that exercises the full image-attachment pipeline: 1. Attaches a minimal 1×1 PNG to the home-page chat via the hidden file input (data-testid="upload-image-input") using Playwright's setInputFiles. 2. Submits "What is in this image?" — creating a conversation and sending the message via sendMessageWithAttachments. 3. Verifies the agent replies with IMAGE_REPLY_TOKEN in the chat UI. 4. Verifies the user MessageEvent in the conversation events API has image_urls populated (base64 data: URL). 5. Verifies at least one /v1/chat/completions call to the mock server contained an image_url content block — confirming the image was forwarded to the LLM as expected. Supporting changes: - mock-llm-server.py: add GET /admin/requests endpoint that exposes all captured completion request bodies since the last reset; reset also clears the history. - mock-llm-helpers.ts: add getMockLLMRequests(), IMAGE_REPLY_TOKEN, and MINIMAL_PNG_BASE64 exports. - AGENTS.md: document the new endpoint and test spec. Co-authored-by: openhands <openhands@all-hands.dev> * test(e2e): add padding response to image-upload trajectory The agent-server makes one internal LLM call for skill-analysis before the main agent loop starts. The original 1-response trajectory was consumed by that internal call, leaving the agent with a 500 error and retry storm. Add 1 padding response (turn 0: empty text) + 1 safety buffer (turn 2) following the same pattern as mock-llm-automation.spec.ts. Co-authored-by: openhands <openhands@all-hands.dev> * test(e2e): use gpt-4o model name for vision-capable LLM requests litellm strips image_url content blocks for unknown model names like 'openai/mock-test-model'. Switch to 'openai/gpt-4o' so litellm knows the model accepts vision content and includes base64 image_url blocks in the completion request body. The base_url still points at the local mock server; the model name is only a hint to litellm's request formatter. Also improve assertion diagnostics to print all captured LLM requests (not just the first) when the assertion fails. Co-authored-by: openhands <openhands@all-hands.dev> --------- Co-authored-by: openhands <openhands@all-hands.dev>
103 KiB
Repository Notes
General
- This repository is a near-direct port of the OpenHands frontend, adapted to talk straight to
software-agent-sdk/agent_serverwithout the usual OpenHands app backend. - Frontend API adaptation lives mainly in
src/api/:option-servicefabricates an OSS web-client config and reads models/providers through@openhands/typescript-clientLLM endpoints.settings-serviceuses@openhands/typescript-clientsettings APIs for persistence; reads schemas from/api/settings/agent-schemaand/api/settings/conversation-schema, fetches settings with optionalX-Expose-Secrets: encryptedheader for conversation start payloads, and saves settings via PATCH with diffs.agent-server-conversation-service,event-service,agent-server-git-service, andskills-serviceroute local agent-server access through@openhands/typescript-clientrather than direct HTTP calls.
- Supported env vars for deployment:
VITE_BACKEND_BASE_URLfor the agent server base URL.VITE_SESSION_API_KEYfor optional session auth.VITE_WORKING_DIRfor the default workspace path sent when starting conversations.VITE_WORKER_URLSas a comma-separated list of browser worker URLs if you want the Browser tab to probe exposed app hosts.VITE_ENABLE_BROWSER_TOOLS=falseto omitBrowserToolSetfrom new conversation payloads.VITE_LOAD_PUBLIC_SKILLS=falseto disable loading public skills from the OpenHands extensions marketplace (https://github.com/OpenHands/extensions). Defaults to true (opt-out).
- Default working-dir fallback is now the relative path
workspace/project(exported asDEFAULT_WORKING_DIRfromsrc/api/agent-server-config.ts); git-path heuristics and the default PLAN preview path should reuse that constant instead of hardcoding/workspace/project. - The UI keeps most OpenHands routes/layout intact, but hosted-only behavior (org, account management, integrations) has been removed via the fabricated OSS config because there is no separate app backend.
- Verification command:
npm run typecheck && npm run build. - GitHub automation now includes
.github/workflows/ci.ymlfornpm ci,npm test, andnpm run build, plus.github/dependabot.ymlwith weekly npm/github-actions updates gated by a 7-day cooldown.
Tracking / Analytics Architecture
Two distinct PostHog systems exist. Never mix them at a call site.
System 1 — telemetry.ts (library-level, anonymous)
- Purpose: anonymous npm-consumer telemetry (
canvas_install,canvas_new_session) - Keys: hardcoded staging/prod keys in
telemetry.ts; routed throughhttps://z.openhands.dev - Consent:
localStorage["openhands-telemetry-consent"]viauseTelemetry/TelemetryConsentBanner canvas_installfires once, pre-consent, per installation- Exports:
trackEvent,useTelemetry,TelemetryConsentBanner, etc. fromsrc/lib/index.ts— these are the public library API for npm consumers - Rule: Do NOT import
trackEventfrom#/services/telemetryin app routes or components
System 2 — useTracking hook (app-level, identified)
- Purpose: product analytics for app behaviour events
- Key:
VITE_POSTHOG_CLIENT_KEYenv var →OptionService.getConfig()→PostHogWrapper→PostHogProvider; routed tohttps://us.i.posthog.com - Consent:
user_consents_to_analytics(backend setting) +useSyncPostHogConsentinroot-layout;AnalyticsConsentFormModalalso callssetTelemetryConsentto keep both systems in sync - All events are typed, named functions in
src/hooks/use-tracking.ts— add a new function there for every new event; never callposthog.capture()raw from a component commonProperties(current_url,user_email) are attached automatically by the hook- Rule: Do NOT use raw
usePostHog()+posthog.capture()in components — always go throughuseTracking
Adding a new event
- Add a typed function to
useTrackinginsrc/hooks/use-tracking.ts - Add the function to the hook's
returnobject - Destructure and call it from the component:
const { trackFoo } = useTracking()
Env var
VITE_POSTHOG_CLIENT_KEY — see .env.sample. Without it, PostHogProvider never mounts and all useTracking calls are silently dropped (safe default for local dev).
Runtime Services in Dev Stacks
- When the agent-canvas dev launchers (
npm run dev/dev:minimal/ the publishedagent-canvasbinary) start a stack, they set aVITE_RUNTIME_SERVICES_INFOenv var on the frontend describing which services are running and how the agent should reach them. The frontend forwards this verbatim asAgentContext.system_message_suffixon everyPOST /api/conversations, so conversations land with a<RUNTIME_SERVICES>block appended to the system prompt. - The block lists URLs from the agent's point of view:
- The Agent Server is always reachable as
http://localhost:<port>from inside the sandbox — but that is you, not the automation backend. - Host-side services (ingress, Vite, automation) are reachable as
http://localhost:<port>.
- The Agent Server is always reachable as
- Agents should treat the
<RUNTIME_SERVICES>block as authoritative: don't hardcodelocalhost:8000for "the automation server", and don't probe random ports trying to discover services. If the block says automation is not running, skip/api/automationcalls; otherwise use the listedurl_from_agent+api_prefix(default/api/automation) and theX-Session-API-Key: $OPENHANDS_AUTOMATION_API_KEYheader. - The launcher → frontend → suffix plumbing is:
scripts/dev-safe.mjs::buildRuntimeServicesInfo()— pure helper that constructs the info object.scripts/dev-with-automation.mjs::buildAutomationRuntimeServicesInfo()— wraps it with automation details; called from both Vite spawn (startVite) and the static build (static-build.mjs).src/api/agent-server-adapter.ts::buildRuntimeServicesSystemSuffix()readsVITE_RUNTIME_SERVICES_INFOand renders the<RUNTIME_SERVICES>markdown block;createAgentFromSettings()attaches it toagent_context.system_message_suffixwhen present.
VITE_RUNTIME_SERVICES_INFO shape
The env var is a JSON string of:
{
"mode": "dev:automation",
"services": {
"agent_server": {
"description": "The OpenHands Agent Server this agent is running inside. ...",
"url_from_agent": "http://localhost:18000"
},
"ingress": {
"description": "Unified entry point. Routes /api/automation/* ...",
"url_from_agent": "http://localhost:8000"
},
"frontend": {
"kind": "vite",
"description": "Vite dev server hosting the agent-canvas frontend.",
"url_from_agent": "http://localhost:3001"
},
"automation": {
"description": "OpenHands Automations service. All routes are mounted under '/api/automation'. Authenticate with header 'X-Session-API-Key: $OPENHANDS_AUTOMATION_API_KEY'.",
"url_from_agent": "http://localhost:18001",
"api_prefix": "/api/automation",
"docs_url": "http://localhost:18001/api/automation/docs",
"openapi_url": "http://localhost:18001/api/automation/openapi.json",
"auth_env_var": "OPENHANDS_AUTOMATION_API_KEY"
}
}
}
All keys under services are optional and omitted when the corresponding service isn't running. frontend.kind is "vite" for dev launchers running the Vite dev server and "static" for stacks serving a pre-built build/ directory (dev:static, the published agent-canvas binary). services.vite is accepted as a legacy alias for services.frontend by the renderer.
Example <RUNTIME_SERVICES> block (dev with automation)
<RUNTIME_SERVICES>
You are running inside an agent-canvas dev stack started in 'dev:automation' mode.
The following services are reachable from your sandbox. URLs are written
from your point of view (i.e., as you should curl/fetch them).
* Agent Server (you): http://localhost:18000
The OpenHands Agent Server this agent is running inside. Tool calls (terminal, file_editor, browser, etc.) execute here.
* Ingress: http://localhost:8000
Unified entry point. Routes /api/automation/* to the automation backend, /api/* and /sockets to the agent-server, and /* to the frontend.
* Frontend: http://localhost:3001
Vite dev server hosting the agent-canvas frontend.
* Automation backend: http://localhost:18001
OpenHands Automations service. All routes are mounted under '/api/automation'. Authenticate with header 'X-Session-API-Key: $OPENHANDS_AUTOMATION_API_KEY'.
Docs: http://localhost:18001/api/automation/docs
OpenAPI: http://localhost:18001/api/automation/openapi.json
Auth: header 'X-Session-API-Key: $OPENHANDS_AUTOMATION_API_KEY'
Trust this block over guessing: do not assume any other URLs are running.
In particular, http://localhost:18000 inside your sandbox is the Agent Server
you are running inside of — NOT the automation backend.
</RUNTIME_SERVICES>
Visual Snapshot Testing
- Snapshot tests live in
tests/e2e/snapshots/and compare screenshots against baselines stored as GitHub Actions artifacts (NOT in git). - Baseline storage: Baselines are stored as the
snapshot-baselinesartifact (90-day retention), uploaded on every push tomain. They are never committed to the repository —tests/e2e/__snapshots__/is gitignored. The artifact is found by querying the artifacts API by name (not by workflow run status) so only runs that actually uploaded the artifact are matched. - Run locally with
npm run test:e2e:snapshots; generate/update snapshots withnpm run test:e2e:snapshots:update. - CI workflow (
snapshot-tests.yml):- On
main: Runstest:e2e:snapshots:update, uploadstests/e2e/__snapshots__/as thesnapshot-baselinesartifact (90 days). - On PRs: Downloads the latest
snapshot-baselinesartifact, runstest:e2e:snapshotsagainst it, generates current snapshots viatest:e2e:snapshots:update, then posts a fresh PR comment (old comment is deleted first so image URLs always point to the current run). Changed snapshots are shown in a side-by-side expected/actual/diff table; new snapshots show the full screenshot. Images are force-pushed to a dedicatedsnapshot-artifacts/pr-<N>orphan branch (NOT the PR branch) so required CI checks are never invalidated. URLs areraw.githubusercontent.com/<owner>/<repo>/<sha>/changed/...or.../new/.... Triggers:opened,synchronize,reopened,labeled,unlabeled. - Force-regenerate baselines: Trigger the
Snapshot Testsworkflow manually withforce_update=true.
- On
- PR comment:
tests/e2e/snapshots/scripts/post-snapshot-comment.mjsposts a fresh comment (<!-- snapshot-test-report -->marker) with collapsed<details>sections — 🔴 Changed (side-by-side expected/actual/diff), 🆕 New (full screenshot), ✅ Unchanged (list of names). - Critical ordering note: Playwright clears
test-results/at the start of every new run. The workflow runs two Playwright passes: (1)test:e2e:snapshots(comparison, writes*-diff.png), then (2)test:e2e:snapshots:update(regenerates baselines, clearstest-results/, no diffs written). The "Save comparison test-results" step copiestest-results/to/tmp/comparison-resultsbetween these two passes and passes it asCOMPARISON_RESULTS_DIRto the comment script, which readsTEST_RESULTS_DIRfrom that env var. Without this step all snapshots appear "unchanged" because the diff files are gone. - Acknowledging intentional changes: If snapshots changed on purpose (UI redesign, etc.), add the
update-snapshotslabel to the PR. This causes: (1) the CI failure step to be skipped so the check passes, (2) the comment status to flip to ✅ with a note that changes are acknowledged, (3) thelabeledtrigger fires a fresh CI run automatically so no manual re-run is needed. Removing the label re-enables the failure. When the PR merges, the main-branch run uploads the new screenshots as the updated baseline — no separate "regenerate on main" step required. - Viewing diffs: On failure, Playwright generates
*-actual.png,*-expected.png, and*-diff.pngintest-results/. Runnpx playwright show-reportto view the HTML report. The PR comment also embeds these images directly. - Bootstrap: When no
snapshot-baselinesartifact exists on main yet, all snapshots are classified as "🆕 New" and CI passes. The first main-branch run after this state uploads the initial artifact. - Image branch cleanup: the
snapshot-artifacts/pr-<N>branch is deleted automatically by thecleanup-snapshot-artifactsjob inpr-artifacts.ymlwhen the PR is closed (merged or abandoned). No manual cleanup needed. - Key patterns for writing snapshot tests:
- Use
setupMocks(page, showConsentModal)helper to configure API mocks consistently. - Use
dismissConsentModal(page)after navigation to dismiss the analytics modal. - Use
animations: "disabled"andmaxDiffPixelRatio: 0.01intoHaveScreenshot()to reduce flakiness. - Target specific elements via
page.getByTestId()rather than full-page screenshots when possible. - Hidden checkbox pattern:
SettingsSwitchrenders<input hidden data-testid="...">. BothtoBeVisible()andclick({ force: true })fail on hidden inputs (no layout dimensions). Use the enclosing<label>:page.locator('label:has([data-testid="my-toggle"])').click()— the browser's label→control activation firesonChangeon the hidden input. - HeroUI Autocomplete testId forwarding:
Autocompletedoes NOT forwarddata-testidto any DOM element. UsegetByRole("combobox", { name: /label text/ })to locate/assert on dropdown fields generated bySchemaField. - SdkSectionPage early return:
SdkSectionPagereturns a plain<p>(nodata-testidwrapper) whenfilteredSchema.sections.length === 0. Always ensure the mock schema includes the section the page requests. The condenser section is inagent_settings_schema(default source), notconversation_settings_schema.
- Use
- Composing tests: Use
test.step()for iterative snapshots within a single test, or extract helper functions (likenavigateToSettings(page)) to share setup across tests. For heavier reuse, use Playwright fixtures viatest.extend(). - Snapshots are organized by
{snapshotDir}/{testFilePath}/{projectName}/{arg}.png(configured inplaywright.config.ts). - Conversation page snapshot tests: The dev server uses MSW service workers for API mocking. For conversation-page tests, rely on MSW's pre-defined mock conversations (IDs "1", "2", "3" in
src/mocks/conversation-handlers.ts) rather than fighting Playwright route interception. MSW's service worker intercepts requests before Playwrightpage.route()can; Playwright route interceptors only see requests that escape the service worker. Stub WebSocket viapage.addInitScript()and inject events into the Zustand store via the exposedwindow.__OH_EVENT_STORE__API. Usetest.describe.configure({ mode: "serial" })for conversation tests since the WebSocket stub + heavier page setup can cause intermittent failures in parallel mode. - Baseline generation for CI: Baselines generated locally will NOT match CI (different OS, fonts, rendering). Baselines are regenerated automatically on every push to
main. After adding new snapshot tests, open a PR — the new snapshots will be shown as "🆕 New" in the PR comment and become the baseline when the PR merges. To force-refresh baselines from main without waiting for a code push, trigger the "Snapshot Tests" workflow manually withforce_update=true. - Mock conversation timestamps must ALL be
now-relative:src/mocks/conversation-handlers.tsdefines mock conversations sorted byupdated_atdescending in the sidebar. Every conversation'screated_at/updated_atmust usenow - X * days(relative to module-load time), never a fixed absolute date likePAGINATION_BASE_TIME. Mixing the two strategies causes a sort-order crossover as real time passes — the fixed-date conversation ages past a relative one and they swap positions, breaking every snapshot that includes the sidebar.PAGINATION_BASE_TIMEis kept only for internal event timestamps used by pagination tests; conversation listing timestamps are decoupled from it. - MSW handler state is PAGE-level JS, not service-worker state: In MSW 2.x browser mode the request handlers (including mutable Maps like
automations) are compiled into the client bundle and run in the main thread.page.reload()re-initialises all module-level state (e.g.const automations = new Map(...)runs fresh on every page load). Tests that need to show an "empty list" state must NOT callpage.reload()after deleting items. Instead: make the DELETE fetches frompage.evaluate()(which DO go through MSW), then callwindow.__TEST_INVALIDATE_QUERIES__()(exposed in mock mode byentry.client.tsx) to trigger React Query's cache invalidation in-place without a reload. Example: the automations empty-state snapshot test intests/e2e/snapshots/automations.snapshot.spec.ts.
Live End-to-End Test Framework
- The live QA path is intentionally separate from ordinary mocked Playwright coverage. If ordinary browser tests are added, keep them outside
tests/e2e/live/soplaywright.config.tscan run them while ignoring**/live/**; live LLM-backed tests must never run as part ofnpm run test:e2e. - Live tests live under
tests/e2e/live/and are run only throughnpm run test:e2e:live, which usesplaywright.live.config.ts. Keep the spec names descriptive; the primary conversation smoke test istests/e2e/live/real-agent-server-conversation.spec.ts. npm run test:e2e:liveloads.envthrough Node's--env-file-if-existsflag and invokestests/e2e/live/scripts/run-live-e2e.mjs. The runner validates the required local environment, explains missing credentials/prerequisites, and then runsplaywright test --config=playwright.live.config.ts. Usenpm run test:e2e:live -- --checkto validate local setup without running the test, and pass Playwright flags after--(for examplenpm run test:e2e:live -- --headed).- Local live E2E requires one LLM credential:
LIVE_E2E_LLM_API_KEY,OPENAI_API_KEY,ANTHROPIC_API_KEY, orLLM_API_KEY. Optional overrides areLIVE_E2E_LLM_BASE_URL,LIVE_E2E_LLM_MODEL,LIVE_E2E_SESSION_API_KEY,LIVE_E2E_BACKEND_URL, andLIVE_E2E_FRONTEND_PORT. The local runner prints which variables are missing without printing secret values. - Live-test-only helpers belong under
tests/e2e/live/utils/. The current helper module istests/e2e/live/utils/agent-server-conversation.ts; do not put live-only helpers in the sharedtests/e2e/support/directory. playwright.live.config.tsstarts the real local Agent Server/UI stack vianpm run dev:minimal, not MSW mocks. It usesLIVE_E2E_SESSION_API_KEYwhen set, otherwise generates a per-run random session key and passes it throughSESSION_API_KEY,OH_SESSION_API_KEYS_0, andVITE_SESSION_API_KEY; specs that need direct backend requests must injectX-Session-API-Keyonly for the configured backend origin throughrouteBackendSessionApiKey(page), never through global PlaywrightextraHTTPHeaders. Live tests default to frontend port3101and Agent Serverhttp://127.0.0.1:18100so they do not accidentally reuse a normal local dev stack.tests/e2e/live/utils/agent-server-conversation.tsconfigures the running Agent Server before each live conversation by PATCHing${LIVE_E2E_BACKEND_URL ?? "http://127.0.0.1:18100"}/api/settingswith LLM settings and low-risk conversation settings. LLM credentials are read fromLIVE_E2E_LLM_API_KEY,OPENAI_API_KEY,ANTHROPIC_API_KEY, orLLM_API_KEY; CI defaults useLIVE_E2E_LLM_BASE_URL(defaulthttps://llm-proxy.app.all-hands.dev) andLIVE_E2E_LLM_MODEL(defaultopenhands/claude-haiku-4-5-20251001).- The live conversation test should stay cheap and as deterministic as possible while still exercising one real tool call: it asks the model to run the exact
EXPECTED_BASH_COMMAND, waits for the bash output token to appear outside the user's message in the UI, confirms a successfulExecuteBashObservation/TerminalObservationthrough the real Agent Server events API, and then waits for the finalEXPECTED_REPLY_TOKEN. This exercises the real UI, Agent Server settings API, conversation creation, websocket/event path, terminal tool execution, and LLM response path. Because LLM behavior is not perfectly deterministic even at temperature 0, CI keeps one retry for live E2E; future live tests should document any expected variance and avoid prompts that require unnecessary formatting obedience. - Live E2E must not pollute analytics.
playwright.live.config.tsstarts the app withVITE_DO_NOT_TRACK=1; the live helper seeds local storage with telemetry/analytics opt-out values before app code runs; and each live spec should installguardAgainstPostHogRequests(page)before navigation so any attempted request to*.posthog.comorz.openhands.devis blocked locally and fails the test. - Live Playwright videos are intentionally recorded for PR QA debugging when
LIVE_E2E_RECORD_VIDEO=onis set by CI; local default video mode isretain-on-failure. Do not add live tests that render API keys, tokens, secret values, or credential-bearing error messages in the browser. Screenshots should target a safe app/chat region such asdata-testid="chat-interface"instead ofpage.screenshot({ fullPage: true }), and should applygetLiveArtifactMask(page)for text/field redaction; if a future live test must exercise sensitive UI, change that test/media path to redact the sensitive output or retain video only on failure. .github/workflows/ci.ymlruns live E2E only at PR level: either manually viaworkflow_dispatchwith a requiredpr_number, or from a same-repository PR carrying thelive-e2elabel. The live job must skip fork PRs before checking out PR code so LLM credentials and artifact-push tokens are never exposed to untrusted code. Do not broaden it to run on every PR push unless cost and credential policy are explicitly revisited.- Keep live E2E secrets out of job-level
env. The workflow should check whether credentials exist before checkout, but inject the LLM key only into the trusted step that actually runs the live test. - The live job uploads the Playwright HTML report plus screenshot/video output as a GitHub Actions artifact, and also extracts the primary screenshot/video attachments. It converts the WebM recording to a GIF preview with
ffmpegso GitHub PR comments can inline the preview. Keep Playwright trace capture disabled for live tests because the setup flow sends LLM credentials to the Agent Server settings API, and traces can record request bodies. Failure messages around live Agent Server settings must not print response bodies from credential-bearing requests. - Inline PR-comment media is stored as PR-only files under
.pr/live-e2e/<github_run_id>/on the PR branch, not on a long-lived orphan media branch. The comment usesraw.githubusercontent.com/<repo>/<artifact_commit>/.pr/live-e2e/...URLs for the GIF and PNG so GitHub can render them inline. The WebM is linked as the full recording because GitHub comments do not reliably inline WebM. .github/workflows/pr-artifacts.ymlowns two cleanup responsibilities: (1).pr/live-e2e/— comments when.pr/artifacts exist and removes them after PR approval for same-repo PRs (fork PRs require manual cleanup); (2)snapshot-artifacts/pr-<N>— deletes the snapshot image orphan branch when the PR is closed (merged or abandoned), via thecleanup-snapshot-artifactsjob triggered onpull_request: [closed].- The live reporting scripts live beside the live tests under
tests/e2e/live/scripts/:run-live-e2e.mjs,extract-live-e2e-media.mjs,render-live-e2e-report.mjs, andupsert-pr-comment.mjs. Keep report/comment/local-runner logic there rather than in top-levelscripts/, because these scripts are part of the live E2E framework. - When changing any part of this framework — live workflow triggers, artifact publishing,
.prcleanup, live Playwright config, live test file layout, helper locations, local runner behavior, or report/comment scripts — update thisAGENTS.mdsection in the same PR so future agents have the current operating model.
Mock-LLM E2E Test Framework
- Mock-LLM tests live under
tests/e2e/mock-llm/and exercise the complete stack — from the browser through the real agent-server to a scripted mock LLM server — without any real LLM credentials. Run locally withnpm run test:e2e:mock-llm. - Production-fidelity launch: The Playwright config (
playwright.mock-llm.config.ts) starts the fullagent-canvasstack viabin/agent-canvas.mjs— the same binary thatnpx @openhands/agent-canvasexecutes when users install the npm package. This means mock-LLM tests exercise the actual production path: pre-built static frontend + static-server.mjs + agent-server via uvx + automation backend via uvx + ingress proxy, all behind a single port. - A pre-built
build/directory is required. The Playwright webServer command runsnpm run build:appwhenbuild/index.htmlis absent, but CI should run the build step explicitly for caching (npm run build:appin.github/workflows/mock-llm-e2e.yml). - Single ingress URL: Tests use one URL for both the browser (
baseURL) and backend API assertions (BACKEND_URL). The ingress proxy routes/api/*to the agent-server,/api/automation/*to the automation backend, and/*to the static frontend. Default ingress port for tests is18300(override viaMOCK_LLM_INGRESS_PORTenv var). - State isolation:
OH_CANVAS_SAFE_STATE_DIR=.tmp/mock-llm-stateisolates test state from the user's real~/.openhands/agent-canvas/directory. BothSTATE_DIR(.tmp/mock-llm-state) and the automation DB dir (.tmp/automation/) are cleaned before each test run — the automation DB now lives outside STATE_DIR atdirname(STATE_DIR)/automation/automations.db, mirroring Docker's~/.openhands/automation/automations.db. - Session API key: A random key is generated per test run and passed to the stack via
SESSION_API_KEY/OH_SESSION_API_KEYS_0/VITE_SESSION_API_KEY. The static server injects it intoindex.htmlat serve time so the frontend authenticates automatically. - Mock LLM server (
tests/e2e/mock-llm/scripts/mock-llm-server.py): Python HTTP server using openhands-sdk'sTestLLMto return scripted tool-call + text trajectories. Supports admin API endpoints for dynamic trajectory management:POST /admin/reset— reset to the default trajectory (terminal printf + text reply); also clears the stored completion-request historyPOST /admin/trajectory/register— register a named trajectory (JSON body:{name, turns}where each turn is{tool_call: {name, arguments}}or{text: "..."})POST /admin/trajectory/activate— activate a previously registered trajectoryGET /admin/requests— return the list of all/v1/chat/completionsrequest bodies captured since the last reset (used by the image-upload test to verify the image was forwarded to the LLM)
- Real automation backend: The automation test uses the production automation backend (started by
bin/agent-canvas.mjs), NOT a mock server. Terminalcurlcommands from the agent hit the automation API through the ingress proxy at the test'sBACKEND_URL(defaulthttp://localhost:18300). Auth uses theX-Session-API-Keyheader matching the stack's session key. Themock-automation-server.pyfile still exists as a reference but is not used by current tests. - Test helpers (
tests/e2e/mock-llm/utils/mock-llm-helpers.ts): ExportsregisterTrajectory(),activateTrajectory(),resetMockLLM(),ensureMockLLMProfile(),getMockLLMRequests()(fetches captured completion bodies fromGET /admin/requests),IMAGE_REPLY_TOKEN+MINIMAL_PNG_BASE64(constants for the image-upload spec), and more. - Padding response for internal LLM call: The agent-server makes an internal LLM call (condenser/skill-analysis) before the agent's main loop starts when skills are activated. This consumes one trajectory response. Automation tests prepend a throwaway
{ text: "" }response as padding. The conversation test does NOT need this because its user message doesn't trigger skill activation. - Test specs:
mock-llm-conversation.spec.ts— Creates LLM profile via UI, runs a conversation with a terminal tool call, verifies bash execution and agent reply.mock-llm-image-upload.spec.ts— Attaches a 1×1 PNG via the hidden file input, sends "What is in this image?", verifies the agent replies, that the user message event stores image_urls, and that the mock LLM received an image_url content block with a base64 data: URL in at least one completion call.mock-llm-automation.spec.ts— Full automation lifecycle: registers a trajectory (7 responses total — 4 for the main conversation + 3 for the automation run's spawned conversation) where the LLM creates a cron automation and dispatches a run via terminalcurlcommands to the real automation backend. Verifies: automation created with correct schedule, run reaches COMPLETED status with a conversation_id, automation appears on the/automationslist page, detail page shows COMPLETED badge (data-testid="run-status-icon-completed"), and clicking the run's conversation link navigates to the correct/conversations/{id}page.mock-llm-partial-stack.spec.ts— Partial stack mode tests. Unlike other specs, these spawn their ownbin/agent-canvas.mjschild processes instead of relying on the config's webServer entries. Three describe blocks: (1)--frontend-onlyverifies static frontend is served (200 on/), backend routes return 503 (/server_info,/api/settings,/api/automation/v1), and the browser shows the manage-backends modal; (2)--backend-onlyverifies/server_inforeturns 200,/api/settingsis reachable, automation endpoint works, and root/asset requests return 503; (3) port conflict verifies the process exits non-zero with a clear error message when the ingress port is occupied, then starts successfully on a free port. Each test uses isolated state dirs and high port numbers (18310+ range) to avoid collisions with the main full-stack instance.
- Tests run serially (
workers: 1,mode: "serial"per describe block). Files are discovered alphabetically so automation tests run before conversation tests; each spec is self-contained (automation test configures its own LLM profile via the settings API). TheafterEachhook resets the mock LLM to its default trajectory so subsequent specs start fresh even when a preceding test fails. - CI workflow:
.github/workflows/mock-llm-e2e.ymlruns on PRs with thee2e-testslabel or on manual dispatch. It builds the frontend, starts the mock LLM server, runs the tests, and posts a PR comment with results. - The custom
DoneMarkerReporterwrites.mock-llm-markers/.tests-doneafter all tests complete (before webServer teardown) so the CI wrapper can detect completion and kill the lingering teardown process.
Docker Image Testing (Shared Specs)
- The same test specs and helpers are reused to validate the Docker image via
playwright.mock-llm-docker.config.ts. Run locally withnpm run test:e2e:mock-llm:docker(requires Docker daemon and a built image). - Architecture: The Docker config replaces the npm path's
bin/agent-canvas.mjswebServer with adocker run --network hostcommand. The mock LLM server still runs on the host. On Linux (including CI),--network hostlets the container share the host's network stack so all127.0.0.1URLs work identically. On macOS/Windows Docker Desktop (bridge networking), setMOCK_LLM_AGENT_URL=http://host.docker.internal:<port>so the agent-server inside Docker can reach the host-side mock LLM server. - Dual-stack binding: Both
scripts/static-server.mjsandscripts/ingress.mjsdefault to::(dual-stack, accepting IPv4 and IPv6 connections). The Docker entrypoint passes--host ::explicitly. This meanslocalhostis safe in both the Docker and npm Playwright configs — whether it resolves to127.0.0.1(IPv4) or::1(IPv6), the server accepts the connection. The mock LLM server URL (MOCK_LLM_URL) still uses127.0.0.1because the Python mock server is a separate process whose bind behavior we don't control. - Entrypoint crash resilience:
docker/entrypoint.shuses awhile kill -0 "$STATIC_PID"; do sleep 10 & wait $!; doneloop instead ofwait -n "${PIDS[@]}"(any child). If the agent-server or automation backend exits mid-test, the static-server proxy stays up and returns 502s for backend routes — the container doesn't disappear withECONNREFUSED. The container exits only when the static-server (ingress) dies or on SIGTERM/SIGINT. Thesleep & wait $!pattern ensureswait(a bash builtin) is the foreground op, so trapped signals fire immediately.cleanup()includesexit 0so the script terminates after a signal-triggered trap return. - URL split:
mock-llm-helpers.tsexports two mock LLM URL constants:MOCK_LLM_BASE_URL— alwayshttp://127.0.0.1:<port>, used by tests for the mock LLM admin API (register/activate/reset trajectories).MOCK_LLM_AGENT_URL— defaults toMOCK_LLM_BASE_URL, overridable viaMOCK_LLM_AGENT_URLenv var. Used when configuring the LLM profile (base_urlfield) — this is the URL the agent-server uses for inference calls. The npm path and Docker-with---network hostpath use the same value; Docker on macOS needs the override.
- Docker image: Set
MOCK_LLM_DOCKER_IMAGEto the image tag (default:ghcr.io/openhands/agent-canvas:latest). The container is started with--rm --network hostand a unique--namefor cleanup. - State isolation: The Docker container uses its internal state directory (no host mount needed for tests). Each test run starts a fresh container.
- CI workflow:
.github/workflows/mock-llm-docker-e2e.ymlhas three triggers — all pull the already-built image from GHCR (no rebuild): (1)workflow_runfires automatically after theDockerworkflow completes on main; (2)pull_requestwith thee2e-testslabel polls the Docker workflow until it finishes for the PR's head SHA, then pulls the image (needed becauseworkflow_runonly fires for workflow files already on the default branch); (3)workflow_dispatchaccepts a customdocker_imageinput. The image tag is derived from the commit SHA (ghcr.io/openhands/agent-canvas:sha-<short>-amd64). Fork PRs are skipped (no GHCR push). Report artifacts go totest-results-mock-llm-docker/andplaywright-report-mock-llm-docker/.
Additional Notes
-
Published binary auth fix: When users install the npm package globally (
npm install -g @openhands/agent-canvas) and runagent-canvas, the pre-built static frontend has NOVITE_SESSION_API_KEYbaked in (npm publish runsnpm run buildwith no such env var). The runtime session key is generated when the CLI launches and reaches the frontend viascripts/static-server.mjs --session-api-key <key>, which injects a<head>script that does two things: (a) setswindow.__AGENT_CANVAS_SESSION_API_KEY__ = <key>— read bygetBakedSessionApiKey()insrc/api/agent-server-config.tsas a fallback when the env var is empty, symmetric with__AGENT_CANVAS_AUTH_REQUIRED__/isAuthRequired(); (b) writes the same key intolocalStorage['openhands-agent-server-config'].sessionApiKey, always overwriting when the value differs, so any code path that still reads the legacy storage key (e.g. e2e fixtures) sees the live key. The window-global path is the load-bearing one — without it,makeDefaultLocalBackend()returns null on a fresh install, the backend registry seeds empty, androot.tsxtraps the user behind the Manage Backends modal instead of onboarding.scripts/dev-with-automation.mjsandscripts/dev-static.mjsboth pass--session-api-key ${config.sessionApiKey}when starting the static server. -
Direct
dependenciesanddevDependenciesinpackage.jsonare exact-pinned (no caret ranges); reproducible installs should use the committedpackage-lock.jsonplusnpm ci, and targeted transitive fixes still belong inoverrides. -
package-lock.jsonmust also retain the optional peer entry fornode_modules/vite-tsconfig-paths/node_modules/typescript@5.9.3; without that nested lock entry, cleannpm ciinstalls on CI fail withMissing: typescript@5.9.3 from lock file. -
npm testnow runsnpm run make-i18nfirst so clean environments generatesrc/i18n/declaration.tsbefore Vitest loads aliased imports. -
__tests__/vite-config.test.tsshould importvite.configdirectly under// @vitest-environment node; spawning plainnode -e 'import ./vite.config.ts'is not portable across Node patch releases in CI. -
vitest.setup.tsmust guard DOM-specific globals (HTMLCanvasElement,HTMLElement,window) because some suites run in the Node environment instead of jsdom. -
__tests__/components/providers/posthog-wrapper.test.tsxmust wrapPostHogWrapperin aQueryClientProvider; the wrapper now reads its client from React Query context instead of importing the global singleton. -
WebSocket hook regression note:
__tests__/hooks/use-websocket.test.ts'sonClosecallback assertion was flaky against the shared MSW websocket server in CI; keep that single test on a deterministic stubbedWebSocketclose path instead of relying on MSW close timing. -
Library i18n regression note:
__tests__/i18n/library-namespace.test.tsimports../../src/index, which can take >5s under the full Vitest suite aftervi.resetModules(). Keep an explicit per-test timeout (currently 15s) so the suite doesn't fail on slow workers. -
src/components/shared/buttons/styled-tooltip.tsxshould keep HeroUI tooltip animations disabled in Vitest (disableAnimationwhenimport.meta.env.MODE === "test"); otherwise full-suite runs can end with unhandledwindow is not definedrejections fromframer-motionafter jsdom teardown (seen viarecent-conversationtests in CI). -
__tests__/i18n/library-namespace.test.tsimports the full library entry and can exceed Vitest's default 5s timeout under full-suite load; keep an explicit higher timeout on that case unless the test is substantially narrowed. -
@openhands/typescript-clientshould be pinned to a released git tag/version rather than an unreleased commit SHA; when agent-canvas needs new client API, release/tag the client first and then update the dependency to that tag. Released versions should include the typed clients, agent-server version compatibility helpers,WorkspacesClient,ConversationClient.switchLLM, and subpath exports forclient/http-client,events/remote-events-list, andworkspace/remote-workspaceneeded by the agent-canvas agent-server integration.RemoteWorkspace.gitChanges/gitDiffaccept an optional{ ref }option; agent-canvas passes'HEAD'so the changes panel reflects working-tree + index versus the latest commit (i.e. staged + unstaged) instead of a diff against the upstream/default branch. -
The
@openhands/typescript-clientgit dep must be expressed as agit+https://github.com/...URL in bothpackage.jsonand the top-level dep entry ofpackage-lock.json; thegithub:OpenHands/...shorthand normalizes togit+ssh://inside the lockfile, and Vercel's build environment has no GitHub SSH key, so an ssh-pinned lockfile makes Vercel fall back to a stale cached tarball and the bundler then fails with[MISSING_EXPORT] ConversationClient/FileClient/SharedClient is not exported by .../dist/clients.js.scripts/vercel-install.sh(wired up viavercel.json'sinstallCommand) defensively rewrites any leftovergit+ssh://git@github.com/resolved URLs togit+https://github.com/and adds matchinggit config --global url..insteadOfaliases before invokingnpm ci, so a future regression that re-introduces an ssh-pinned lockfile entry still builds on Vercel. See GitHub issue #384 for the original failure and PR #382 for the prior single-shot lockfile fix that this generalizes.
API Access Rules
Two strict conventions govern every REST call in the frontend. Violations break CI
via src/api/no-direct-agent-server-calls.test.ts.
Rule 1 -- Agent-server calls must use @openhands/typescript-client
All calls that target the local agent-server (/api/*, /server_info, /sockets)
must go through typed client classes from @openhands/typescript-client, never
raw axios, fetch, or the legacy shared openHands axios instance.
Available clients and their subpath imports:
ConversationClient--@openhands/typescript-client/clientsFileClient--@openhands/typescript-client/clientsVSCodeClient--@openhands/typescript-client/clientsServerClient--@openhands/typescript-client/clientsHttpClient--@openhands/typescript-client/client/http-clientRemoteWorkspace--@openhands/typescript-client/workspace/remote-workspaceRemoteEventsList--@openhands/typescript-client/events/remote-events-list
Client options are always assembled via helpers in src/api/agent-server-client-options.ts:
getAgentServerClientOptions(overrides?)-- for SDK client constructorsgetAgentServerHttpClientOptions(overrides?)-- forHttpClient-based callers
These helpers read host, session API key, and working directory from the active backend registry and env config, so callers never hardcode URLs or auth tokens.
// CORRECT
const data = await new ConversationClient(
getAgentServerClientOptions(),
).getConversation(id);
const file = await new FileClient(
getAgentServerClientOptions(),
).downloadTextFile(path);
// WRONG -- raw axios/fetch calls fail the no-direct-agent-server-calls.test.ts guard
const data = await axios.get(`${host}/api/conversations/${id}`);
const data = await fetch(`/api/conversations/${id}`);
Allowed exceptions (files that may use axios directly for infrastructure reasons):
src/api/automation-service/automation-service.api.tssrc/api/cloud/proxy.ts-- the proxy envelope POST itself
Rule 2 -- Cloud backend routes must go through callCloudProxy
Any call from the browser to the cloud backend (app.all-hands.dev) or a cloud
runtime sandbox (*.prod-runtime.all-hands.dev) must go through callCloudProxy()
in src/api/cloud/proxy.ts. These origins do not permit CORS from localhost;
callCloudProxy POSTs the request envelope to /api/cloud-proxy on the local
agent-server, which forwards it server-side.
import { callCloudProxy } from "../cloud/proxy";
// CORRECT -- cloud endpoint
const result = await callCloudProxy<ResponseType>({
backend,
method: "GET",
path: `/api/v1/app-conversations/search?${params}`,
});
// CORRECT -- cloud runtime sandbox, auth via session key
const result = await callCloudProxy<ResponseType>({
backend,
method: "GET",
hostOverride: buildHttpBaseUrl(conversationUrl),
path: `/api/git/changes?path=${path}`,
authMode: "session-api-key",
sessionApiKey,
});
// WRONG -- direct fetch/axios to a cloud host is blocked by CORS in the browser
const result = await axios.get(`${backend.host}/api/v1/app-conversations`);
callCloudProxy key options:
backend-- the cloudBackendobject (provides host and bearer token)hostOverride-- override for runtime-sandbox calls; replacesbackend.hostauthMode--"bearer"(default, cloud) |"session-api-key"(runtime sandbox) |"none"sessionApiKey-- required whenauthMode === "session-api-key"
Standard cloud/local branch pattern used throughout the service layer:
if (getActiveBackend().backend.kind === "cloud") {
return callCloudProxy({ backend: active, ... });
}
return new ConversationClient(getAgentServerClientOptions()).someMethod(...);
No Magic Strings
Inline string literals that carry meaning (user-facing copy, identifiers, keys, route paths, storage keys, event types, query keys, env-var names, etc.) must not appear at call sites. Magic strings drift across files, defeat search/refactor, bypass tsc's spell-checking, and ship untranslated UI text. The i18next/no-literal-string rule is set to "error" in eslint.config.js and CI fails on violations — do not silence it with eslint-disable unless the string is genuinely non-localizable (e.g. the ⌘↩ keyboard glyph in plan-preview.tsx / conversation-tabs.tsx).
Rule 1 — User-facing strings go through i18n
Every visible string (button labels, headings, validation messages, aria-label, title, alt, toast copy, placeholders) must be routed through react-i18next's t() keyed by an I18nKey enum member. Keys are declared once in src/i18n/translation.json with values for all 15 supported languages (see src/i18n/index.ts::AvailableLanguages), and npm run make-i18n regenerates src/i18n/declaration.ts + public/locales/<lang>/openhands.json. npm run check-translation-completeness fails CI if any key is missing a language.
// CORRECT
import { useTranslation } from "react-i18next";
import { I18nKey } from "#/i18n/declaration";
const { t } = useTranslation("openhands");
return (
<button aria-label={t(I18nKey.CHAT$DISMISS_LABEL)}>
{t(I18nKey.CHAT$DISMISS)}
</button>
);
// WRONG -- ships English to every locale; flagged by i18next/no-literal-string
return <button aria-label="Dismiss">Dismiss</button>;
Key naming follows the existing CATEGORY$IDENTIFIER convention (see src/i18n/translation.json — common prefixes: CHAT_INTERFACE$, SETTINGS$, COMMON$, BUTTON$, HOME$, MICROAGENT$, etc.). Reuse an existing prefix; only introduce a new one when no sensible bucket exists.
Caveat: eslint-plugin-i18next's recommended config catches JSX text children but NOT string-literal prop values like aria-label="…" or string ternaries passed to props. Even when the rule does not flag them, treat them as user-facing strings and route them through t(). If you find a hardcoded prop string, fix it; do not assume the linter's silence is approval.
Rule 2 — Non-UI identifiers live in named constants, not inline literals
For strings the user never sees but the program reads (storage keys, event names, query keys, route paths, env-var names, header names, hardcoded paths, feature-flag identifiers), declare a single named constant in the closest module that owns the concept and import it everywhere else. Co-locate related constants in a tiny dedicated file (*-keys.ts, *-constants.ts) when more than two callers need them.
// CORRECT
const ONBOARDING_COMPLETED_KEY = "openhands-onboarded";
localStorage.setItem(ONBOARDING_COMPLETED_KEY, "true");
// CORRECT -- query keys go through SETTINGS_QUERY_KEYS / SECRETS_QUERY_KEYS / …
// in src/hooks/query/query-keys.ts (enforced by no-restricted-syntax)
queryClient.invalidateQueries({ queryKey: SETTINGS_QUERY_KEYS.all });
// WRONG -- duplicated literal across files, no compile-time link, silent typo risk
localStorage.setItem("openhands-onboarded", "true");
queryClient.invalidateQueries({ queryKey: ["settings"] });
Already-named constants in this repo include DEFAULT_WORKING_DIR (src/api/agent-server-config.ts), OPENHANDS_I18N_NAMESPACE (src/i18n/index.ts), BUNDLED_BACKEND_ID (backend registry), and the *_QUERY_KEYS helpers in src/hooks/query/query-keys.ts. Reuse these instead of re-inlining the literal.
Rule 3 — Discriminated-union tags use string-literal types, not bare strings
When a string is part of a discriminated union or enum-like set (event kinds, backend kinds, tab IDs, agent statuses, observation result statuses), the type itself should constrain the literal. Pass values typed against that union, not raw string, so callers get autocomplete and the compiler catches typos.
// CORRECT
type BackendKind = "local" | "cloud";
if (backend.kind === "cloud") { … }
// WRONG -- `backend.kind` typed as `string`; "clould" compiles fine
if (backend.kind === "clould") { … }
Allowed exceptions
- Test fixtures (
__tests__/,tests/e2e/) may use inline literals for setup data — tests are the boundary where strings stop being magic. - Non-localizable display glyphs (keyboard shortcuts like
⌘↩, currency symbols, etc.) may stay inline behind aneslint-disable-next-line i18next/no-literal-stringcomment. Keep the disable on the single offending line; never widen it to a file-level disable for a single glyph. - Generated files (
src/i18n/declaration.ts,public/locales/<lang>/openhands.json) are produced bynpm run make-i18n; do not hand-edit, do not lint-target.
When adding code that needs a new string, decide up front which rule it falls under: if a user reads it → Rule 1; if the program reads it → Rule 2; if it tags a union → Rule 3. Do not commit code that fails any of these rules just because the linter happens not to catch it.
-
Use
@openhands/typescript-clientclasses directly for agent-server-backed REST/workspace/event/VS Code calls. Centralize host/session API key/working-directory option assembly throughsrc/api/agent-server-client-options.ts; the backend fallback policy itself lives insrc/api/backend-registry/active-store.ts. -
Local verification/build gotchas:
npm run typecheckassumes generated translation types exist; runnpm run make-i18nfirst ifsrc/i18n/declaration.tsis missing.
-
Merge note:
mainremoved the old project-management integration subcomponents/hooks and their related feature-flag/i18n surface. If a feature branch still keeps the top-level/integrationsgit-token page, retainsrc/routes/git-settings.tsxplus the git-provider token inputs/hooks, but do not blindly restoresrc/components/features/settings/project-management/*or the old integration mutation/query hooks unless the corresponding option types and i18n keys are also reintroduced. -
The OSS cleanup removed hosted-only auth, org, account, onboarding, and invitation codepaths, routes, and tests. Keep
integrations,git-settings,secrets, MCP settings, and other local/self-hosted flows intact when simplifying OSS behavior. -
When merging main into this branch, keep the new agent-server compatibility bootstrap in
src/root.tsx, but do not reintroduce hosted-only invitation cleanup or marketing CTA chrome in the OSS user menu; the OSS account menu should just render settings links plus Docs. -
During the OSS cleanup audit, the runtime removals held up, but route-level regression coverage for still-active OSS settings pages had been deleted too aggressively. Keep focused tests for local/self-hosted screens like
app-settings,llm-settings,git-settings,mcp-settings, andsecrets-settingseven when stripping hosted-only code. -
npm run dev:mockneeds MSW handlers for the direct agent-server routes used by the adapted frontend, not the original OpenHands mock paths. Key routes that must stay covered are:- bootstrap/model loading:
/server_info,/api/llm/models/verified,/api/llm/providers - settings schemas:
/api/settings/agent-schema,/api/settings/conversation-schema - settings CRUD:
GET /api/settings,PATCH /api/settings - secrets CRUD:
GET /api/settings/secrets(list),GET /api/settings/secrets/:name(value),PUT /api/settings/secrets(upsert),DELETE /api/settings/secrets/:name - conversation browsing/loading:
/api/conversations/search,/api/conversations?ids=...,/api/conversations/:id,/api/conversations/:id/events/* - runtime git panels:
/api/git/changes,/api/git/diff
- bootstrap/model loading:
-
Static mock verification needs a build created with
VITE_MOCK_API=true(usenpm run build:mock); the client must start MSW whenever that flag is enabled, even in production/static builds, otherwise routes like/settingsand the conversations pane fall through to the static server and crash on undefined.filter/.mapassumptions. -
No frontend version guard:
OptionService.getConfig()callsloadAgentServerInfo()which fetches/server_infoonly to (a) detect an unreachable agent server (renders the onboarding screen viaAgentServerUnavailableError) and (b) cacheusable_toolsfor tool gating. All advertised versions are accepted. The "manage backends" modal displays each local backend's/server_info.versionin light text via theBACKEND$VERSION_LABELtranslation key. -
Backend registry: there is no longer a separate "bundled" backend. On first read of the
openhands-backendslocalStorage key (raw === null),readStoredBackends()seeds the registry with one default local backend (makeDefaultLocalBackend(), idBUNDLED_BACKEND_ID = "default-local", host/api-key fromagent-server-config). After that the seed is just an ordinary registered backend — users can rename or remove it like any other.getEffectiveLocalBackend()returns the first registered local, falling back to a synthesized default if the registry has no locals (used by API clients that need a baselinelocaltarget). The "Manage backends" modal and the BackendSelector dropdown both read from the single registered list, so the seeded default appears in both without any special-casing. -
Shared
Dropdownopen behavior: when the menu opens, it clears the input/search text so callers can show the current selection viaplaceholderwhile still rendering the full option list. Generic dropdown tests should not expect the selected label to remain in the input after reopening unless the parent explicitly controls that display. -
useLoadOlderEventsneeds ref-basedisLoading/hasMoreguards in addition to React state becauseChatInterfacecan trigger pagination fromonScroll,onWheel, and the no-overflow effect in the same tick; closure-based state alone allows duplicate page requests. -
ChatInterfacecontinuity tests should assert that conversation messages render without the fullchat-messages-skeleton, not thatdata-testid="loading-spinner"is absent: the lazy older-events indicator reuses the sharedLoadingSpinnercomponent and legitimately renders that inner test id while history backfill is running. -
useConversationHistorynow mirrors the older-events pagination fallback when the first page is exactlyINITIAL_HISTORY_PAGE_SIZE: treatnext_page_idor a full page ashasMore, so older agent-server variants that omitnext_page_idstill allow one more backfill request. The hook anduseLoadOlderEventsalso defensively reject mocked/malformedpage.itemsresponses before reversing them. -
/server_infotool capability metadata fromsoftware-agent-sdkPR #3028 ended up shipping asusable_tools(notavailable_tools). Frontend browser-tool gating should key offusable_tools, and still default to allowing tools when the server does not advertise tool metadata. -
Useful regression tests for mock mode live in
__tests__/api/option-service.test.ts,__tests__/api/mock-conversation-handlers.test.ts, and__tests__/api/mock-settings-handlers.test.ts. -
Browser-verified mock-mode tour artifact was generated at
artifacts/frontend-tour.gif. -
Live
agent_servercompatibility quirks discovered during browser verification:- Latest
openhands-agent-serverlive-mode notes (verified against 1.18.1):/api/settings/agent-schemaand/api/settings/conversation-schemaexist on recent servers, but they return401when the server was started withSESSION_API_KEYorOH_SESSION_API_KEYS_0; the frontend must send the same value asVITE_SESSION_API_KEY/X-Session-API-Key.- The provider/model picker should use
/api/llm/providers,/api/llm/models, and/api/llm/models/verified;/api/v1/config/providers/searchand/api/v1/config/models/search404 on current live agent-server releases. - When the browser is accessing the frontend through a remote host (for example an All Hands work URL) but
VITE_BACKEND_BASE_URLpoints at127.0.0.1/localhost, browser-side REST calls must fall back to the frontend origin so Vite can proxy/apiand/socketsto the local backend.
GET /api/conversationsexpects repeatedidsparams (?ids=a&ids=b), not Axios's default bracket form (ids[]=a), so the shared Axios client needs a custom params serializer.- Runtime git panels should prefer the conversation's reported
workspace.working_dirwhen present; falling back to/workspace/projectcan produce 500s likeNot a git repositoryfor direct local workspaces such as/workspace/project/agent-canvas. - For development,
npm run devnow usesuvxto run a temporary agent-server installation, so no permanentuv tool installis required. For standalone installations,uv tool install -U --with openhands-tools --with openhands-workspace openhands-agent-serverwould expose the executable asagent-server(notopenhands-agent-server), possibly requiring~/.local/binonPATH. - Current SDK / agent-server conversation start payloads must use SDK-registered snake_case tool names, not the old class-style names. Working names against SDK v1.18.1 were:
terminalfile_editortask_trackerbrowser_tool_setUsingTerminalTool/FileEditorTool/TaskTrackerTool/BrowserToolSetcaused live/api/conversations/{id}/eventsruns to fail withToolDefinition '<name>' is not registered.
- The root compatibility bootstrap now treats
/server_infonetwork/timeout failures as a first-classAgentServerUnavailableError, uses a short 5s timeout for that probe, and disables React Query retries/toasts for the initial config fetch so missing backends fail fast with an explicit full-screen notice. - For local verification in this repo, setting
VITE_WORKING_DIR=/workspace/project/agent-canvasavoids initial Changes-tab 500s from pointing conversations at the non-repo parent/workspace/project. - A successful end-to-end live run in this environment required a real LLM config (
LLM_MODEL+LLM_API_KEY). The defaultlitellm_proxy/...model with nollm_api_keyfailed at runtime with alitellm.AuthenticationError.
- Latest
-
Agent-server recovery UX gotchas:
- Keep
/settings/agent-serverin the intermediate-page bypass path (use-is-on-intermediate-page) souseConfig()-driven layout/sidebar queries do not block the recovery screen behind a global spinner. PostHogWrappershould treat config-fetch failures as silent/optional (no user-facing toast), otherwise onboarding/recovery screens show a duplicate incompatible-server toast on top of the friendly guidance.- Keep the settings route on the compact
AgentServerConnectionFormvariant withshowSectionHeader={false}and no checklist; the blocked root onboarding should stay similarly minimal, with only the status card plus a single sentence that links to the repo setup instructions. - For local screenshot/GIF capture of SPA routes, serve
build/with an SPA fallback (for examplesirv build --single) and restart the static server after each rebuild so hashed asset URLs stay in sync.
- Keep
-
Git provider token persistence note:
src/api/secrets-service.tsstores git provider tokens in TWO places:- Agent-server secrets API (
PUT /api/settings/secrets) with naming conventionGIT_PROVIDER_{PROVIDER}_TOKEN- for agent runtime use - localStorage (
openhands-agent-server-git-provider-tokens) - for frontend git API calls (repo search, branches, etc.) TheaddGitProvidermethod stores to server FIRST (must succeed), then updates localStorage. This ensures server-side persistence is the source of truth.
- Agent-server secrets API (
-
Agent server connection settings now live at
Settings > Agent Server(/settings/agent-server). The page reads deployment defaults fromVITE_BACKEND_BASE_URL/VITE_SESSION_API_KEY, saves user overrides in theopenhands-agent-server-configlocalStorage key, and must stay reachable even when the backend compatibility probe fails so users can recover from missing or wrong backend configuration. -
Auth modes for
agent-canvas(dev and production):- Local mode (default, no
--publicflag): A session API key is auto-generated and persisted to~/.openhands/agent-canvas/session-api-key.txt. The key is baked into the Vite dev server viaVITE_SESSION_API_KEYor injected into static builds viastatic-server.mjs --session-api-key. Users never need to paste a key. - Public mode (
--publicflag): RequiresLOCAL_BACKEND_API_KEYenv var. The key is used as the agent-server session key (OH_SESSION_API_KEYS_0) but is NOT baked into the frontend (noVITE_SESSION_API_KEY, no--session-api-keyto static-server). The frontend detects a 401 from/server_infoviaisAgentServerAuthError()and showsApiKeyEntryScreen(src/components/features/backends/api-key-entry-screen.tsx). The screen reusesBackendFormwith the host pre-filled (read-only) and prompts for the API key. On submit, the key is persisted tolocalStorage['openhands-agent-server-config']and the page reloads. - Dev usage:
LOCAL_BACKEND_API_KEY=my-secret npm run dev -- --public - Production usage:
LOCAL_BACKEND_API_KEY=my-secret npx @openhands/agent-canvas --public - The
--publicflag is supported by bothscripts/dev-with-automation.mjs(parsed inparseArgs(), propagated viaconfig.isPublic) andbin/agent-canvas.mjs(passed asisPublictomain()). - The 401 detection lives in
src/api/agent-server-compatibility.ts(isAgentServerAuthError()), and the gate is insrc/root.tsx'sAppcomponent, between theAgentServerUnavailableErrorcheck and the<Outlet />render. - Key rotation resilience (non-public): Stale registry entries are reconciled in two places. (1)
syncLauncherDefaultLocalBackend()insrc/api/backend-registry/storage.tsre-runs at module init: for any stored backend whose id isdefault-localand whose host matches (or is loopback-equivalent to) the launcher's default, itsapiKeyis overwritten with the currentmakeDefaultLocalBackend().apiKey(sourced fromVITE_SESSION_API_KEYor, in the published-binary path,window.__AGENT_CANVAS_SESSION_API_KEY__). (2)scripts/static-server.mjs's injection script always overwriteslocalStorage['openhands-agent-server-config'].sessionApiKeywhen it differs from the runtime key, keeping any code that still reads that legacy storage in sync. E2E coverage:tests/e2e/mock-llm/mock-llm-auth-modes.spec.ts(fresh-install, key-rotation, and public-mode scenarios).
- Local mode (default, no
-
Backend/footer actions that launch modals from inside a dropdown or popover should intercept
onMouseDownto keep the menu mounted, then perform the actual open ononClick. Current examples:Add backend/Manage backendsinsrc/components/features/backends/backend-selector.tsx, plus the mirrored workspace-footer buttons insrc/components/features/conversation-panel/new-conversation-button.tsx. -
BackendSelector's cloud-org switch paths should never rethrow from the dropdownonChangehandler: unexpected non-Axios failures need a generic error toast instead of an unhandled promise rejection, and the malformed(cloud backend, null org)self-heal path should fall back to the bundled backend if/switchfails. -
NewConversationButtonshould support keyboard dismissal (Escape) for its inline popover, while still keeping the popover open when its modal children (FolderBrowserModal,ManageWorkspacesModal) are active. -
README expectation: keep the first section as a concrete, chronological from-scratch quickstart for running this frontend against a real
openhands-agent-server(clone, install prerequisites, optional.env, runnpm run dev). -
Keep README user-focused and move contributor/developer-specific workflows (
dev:safe, mock mode, detailed env vars/build-test notes) intoDEVELOPMENT.md. -
Windows-specific command syntax (PowerShell) lives in
README.windows.md. When changing install / Docker sandbox instructions inREADME.md, updateREADME.windows.mdin the same PR to keep them in sync. -
scripts/dev-safe.mjsusesuvxfor temporary agent-server installation — no permanentuv tool installneeded. Environment variables (highest precedence first):OH_AGENT_SERVER_LOCAL_PATH— absolute path to a localsoftware-agent-sdkcheckout. Runs the local checkout viauvxwith--with-editableforopenhands-sdk/openhands-tools/openhands-workspaceand--reinstallforopenhands-agent-server, so SDK edits are picked up on restart. Highest precedence.OH_AGENT_SERVER_GIT_REF— git commit SHA or branch name (takes precedence over version)OH_AGENT_SERVER_VERSION— specific PyPI version (e.g., "1.24.0")OH_SECRET_KEY— secret key for settings encryption; auto-generated and persisted to~/.openhands/agent-canvas/secret-key.txton first run (same file Docker uses), ensuring dev mode and Docker share the same key when both mount the same~/.openhandsdirectory. Override with the env var to pin a specific key.SESSION_API_KEY/OH_SESSION_API_KEYS_0/VITE_SESSION_API_KEY— session API key for agent-server authentication; auto-generated usingcrypto.randomBytes(32)if not set, passed to both agent-server (OH_SESSION_API_KEYS_0) and frontend (VITE_SESSION_API_KEY)- Default: released PyPI version
1.24.0for agent-server SDK libraries
-
Security:
scripts/dev-safe.mjsandscripts/dev-with-automation.mjsauto-generate random API keys when needed and persist the defaults so static builds, localStorage, and restarted services stay in sync:SESSION_API_KEY— 64-character hex (256-bit) for agent-server API authentication; persisted at~/.openhands/agent-canvas/session-api-key.txtunless overridden via env varAUTOMATION_LOCAL_API_KEY— 64-character hex for automation backend auth; persisted at~/.openhands/agent-canvas/automation-api-key.txtunless overriddenOH_SECRET_KEY— 64-character hex (256-bit) for settings encryption; persisted at~/.openhands/agent-canvas/secret-key.txtunless overridden via env var. Same file used bydocker/entrypoint.sh, so dev mode and Docker share the same key automatically.
-
scripts/dev-safe.mjsshould fail fast ifuvxcannot be spawned (for example missing PATH entries). -
npm run devruns the full local stack viauvx(agent-server + automation backend + Vite dev server + ingress proxy) with no Docker dependency.npm run dev:staticdoes the same but serves a production build of the frontend instead of the Vite dev server. -
scripts/dev-with-automation.mjsruns the full stack: agent-server, automation backend (both via uvx), frontend server, and ingress proxy. It defaults to Vite when run directly, supports--staticfor an existing build, and supports--dynamicso wrappers that default static can opt back into Vite. Uses a standalone ingress proxy (scripts/ingress.mjs) to route traffic:/api/automation/*→ automation backend (:18001)/api/*,/sockets, etc. → agent server (:18000)/*(default) → frontend server (:3001), either Vite or static depending on launcher mode- Environment variables:
PORT(ingress port, default: 8000),OH_AUTOMATION_GIT_REF(git ref, overrides default version),OH_AUTOMATION_VERSION(default:1.0.0a3),AUTOMATION_LOCAL_API_KEY(optional, use a fixed key; default: persisted generated key),OH_AUTOMATION_API_KEY_PATH(override the persisted default key path) scripts/check-sdk-version-sync.mjschecks the releasedopenhands-automationpackage againstversions.automationSdkinconfig/defaults.json; that value may intentionally lagversions.agentServerwhile automation has not yet published a matching release.- Access points:
http://localhost:8000/(main UI),http://localhost:8000/api/automation/docs(API docs) - Security:
AUTOMATION_LOCAL_API_KEYdefaults to a generated key persisted across restarts because static frontend builds bake it intoVITE_AUTOMATION_API_KEY. Set the env var explicitly to rotate or pin it. The cipher key (OH_SECRET_KEY) is persisted at~/.openhands/agent-canvas/secret-key.txt(same file used bydocker/entrypoint.sh); both modes share the same key automatically when using the same~/.openhandsdirectory.
-
scripts/ingress.mjsis a standalone HTTP reverse proxy that can be used independently to route traffic to multiple backends based on URL path prefix. -
scripts/dev-safe.mjs(nownpm run dev:minimal) runs just agent-server + Vite without automation. -
Vite dev mode can black-screen on first load with
504 Outdated Optimize Depif core client-entry deps are not prebundled; keepreact,react/jsx-runtime,react-dom/client, andreact-router/dominoptimizeDeps.include. -
Bundle/dev-graph hygiene (Tier 1 cleanup landed):
src/i18n/translation.json(~1 MB) is imported only bysrc/i18n/resources.ts, whichsrc/i18n/index.tsre-exports astranslationResourcesfor the@openhands/agent-canvas/i18nsubpath. The re-export is aexport … fromplus/* @__PURE__ */annotation, so rollup drops the JSON from the app build (the prodcustom-toast-handlerschunk went from ~909 KB to ~74 KB). Do not move the JSON import back intosrc/i18n/index.ts— that immediately re-bundles all translations into every chunk that importsi18n.- The environment-switch overlay is split: lightweight store/triggers live in
components/features/backends/environment-switch-store.ts; the React component lives inenvironment-switch-overlay.tsx(re-exports the store API for back-compat). Eagerly-mounted callers (e.g.backend-selector.tsx) MUST import trigger helpers from the store, not the overlay file. The overlay isReact.lazy'd fromroutes/root-layout.tsx. - Other always-conditional UI is
React.lazy'd to keep the root layout's eager graph small:AnalyticsConsentFormModal,AlertBanner(root-layout),SettingsModal(sidebar),AgentServerConnectionForm(root.tsx). Tests that assert on these mounted nodes needawait screen.findByTestId(...)instead ofgetByTestId(...). - The terminal tab (
components/features/terminal/terminal.tsx) isReact.lazy'd inconversation-tab-content.tsxalongside the other tabs, so xterm + addon-fit + xterm.css don't enter the conversation route's eager graph. - Avoid importing app code through
#/components/conversation-events/chator itsevent-message-components/index.tsbarrel — they exist forlib/index.ts(npm subpath) consumers only. Internal callers use deep paths (./messages,./event-message-components/<name>,./event-content-helpers/should-render-event) so Vite dev doesn't fan out the barrel.
-
Vercel deployment note: React Router builds for this repo must keep
build/clientintact on actual Vercel builds and includepresets: [vercelPreset()]from@vercel/react-router/vite; flatteningbuild/clientduring a Vercel build produces deployments with empty outputs (routes: null, no static files) and a production 404. -
The repo should include a root
LICENSEfile to satisfy the incubator-program requirements. -
OpenHands repo bootstrap files live under
.openhands/:.openhands/setup.shinstallsuv(viacurl -LsSf https://astral.sh/uv/install.sh | sh) if not present, installs frontend dependencies withnpm ciwhen needed, creates.envfrom.env.sampleif missing, appendsVITE_WORKING_DIRfor this repo when unset, and generatessrc/i18n/declaration.tsvianpm run make-i18n..openhands/hooks.jsonregisters.openhands/hooks/on_stop.shas a Stop hook so OpenHands runs the local quality gate (npm run lintandnpm test) before finishing.
-
GitHub PR-review automation should stay aligned with the current OpenHands repo conventions: keep the review workflow at
.github/workflows/pr-review-by-openhands.yml, keep the companion.github/workflows/pr-review-evaluation.yml, auto-run on newly opened non-draft PRs andready_for_reviewevents from established contributors, still support thereview-thislabel /openhands-agent/all-hands-botreviewer triggers, use the OpenHands app LLM proxy defaults, and use the dual-trigger pattern (pull_requestfor same-repo PRs,pull_request_targetfor forks) so workflow changes can self-verify without widening fork secret exposure. -
The repo now includes
.agents/skills/custom-codereview-guide.md, adapted fromOpenHands/software-agent-sdk, to force PR reviews to always leave either an APPROVE or COMMENT review instead of silently finishing with no review object. -
HeroUI rollback / migration notes:
- The attempted HeroUI v3 upgrade changed global theme wiring and homepage design tokens enough that the repo currently prefers
@heroui/react@2.8.10until a broader visual validation pass is done. - Keep the v2 Tailwind integration active via
@plugin '../hero.ts'insrc/tailwind.cssand source HeroUI classes fromnode_modules/@heroui/theme/dist/**/*. - The settings UI currently relies on the v2
Autocomplete+AutocompleteItem/AutocompleteSectionAPIs insettings-dropdown-input.tsxandmodel-selector.tsx; a future v3 retry will need to replace those controls again.
- The attempted HeroUI v3 upgrade changed global theme wiring and homepage design tokens enough that the repo currently prefers
-
Library i18n is now namespace-scoped under
openhands:src/i18n/index.tsexportsOPENHANDS_I18N_NAMESPACE,translationResources, andwaitForI18n(),scripts/make-i18n-translations.cjsemitspublic/locales/<lang>/openhands.json, standalonesrc/entry.client.tsxexplicitly awaits i18n init, and host apps can register bundles via the@openhands/agent-canvas/i18nsubpath export. -
Route decoupling note:
src/components/should stay free of directreact-routerimports. Route state now flows throughsrc/context/navigation-context.tsx, the standalone app bridges router state withsrc/routes/react-router-navigation-provider.tsx, and link-like UI should usesrc/components/shared/navigation-link.tsx. -
Test helper note:
test-utils.tsxnow wraps renders with a defaultNavigationProvider(currentPath: "/",conversationId: "test-conversation-id"). Navigation-sensitive tests can override that viarenderWithProviders(..., { navigation: { ... } }). -
CSS isolation for embeddable/hosted use now relies on a scoped wrapper attribute: all bundled CSS is prefixed under
[data-agent-server-ui]viapostcss-prefix-selectorinvite.config.ts, with selector exceptions handled bytransformAgentServerUISelector()insrc/styles/agent-server-ui-style-scope.ts. That transform must remap global selectors like:root,html, andbodydirectly onto the scoped shell instead of emitting impossible descendants such as[data-agent-server-ui] :root. -
Public embedding entry points should use
AgentServerUIProviders(scoped root on by default) orAgentServerUIRootfor manual control. The standalone app already renders its own scoped root insrc/root.tsx, sosrc/entry.client.tsxmust passwithStyleRoot={false}to avoid nesting duplicate shells. KeepAgentServerUIRootand the scoping constants re-exported fromsrc/lib/index.tsso library consumers can customize the host wrapper without reaching into private paths. -
AgentServerUIRoot's themed inner wrapper must set a defaultcolor: var(--foreground)in addition to thedark/data-thememarkers; otherwise inherited text andcurrentColorSVG icons fall back to dark browser defaults after CSS scoping, causing dark-on-dark regressions on pages like the home screen. -
Theme/customization tokens for the embedded shell are exposed as
--oh-*CSS variables. Override them throughstyleOverrides,style, or host CSS targeting[data-agent-server-ui]; Tailwind theme tokens insrc/tailwind.cssshould continue to reference those variables with@theme inlineso host apps can restyle the UI without reworking component class names. -
Regression coverage for the CSS isolation work lives in
__tests__/agent-server-ui-providers.test.tsx,__tests__/agent-server-ui-style-scope.test.ts, and the browser-leveltests/e2e/regressions/css-isolation.spec.tsPlaywright test. -
Conversation history is loaded lazily, REST-first then WebSocket:
-
useConversationHistory(insrc/hooks/query/use-conversation-history.ts) fetches only the most recentINITIAL_HISTORY_PAGE_SIZE(default 50) events usingsort_order='TIMESTAMP_DESC', then reverses to chronological order. Older pages are paginated in viauseLoadOlderEventswhen the user scrolls near the top of the chat. -
EventService.searchEvents(conversationId, conversationUrl, sessionApiKey, options)returns the rawEventSearchPage({ items, next_page_id }); options supportlimit,pageId,sortOrder,timestampGte,timestampLt. Both the local and cloud-proxy code paths forward the new params. -
The main
ConversationWebSocketProviderwaits for the REST query to settle before opening its socket, then connects withresend_mode='since'andafter_timestamp=<latest preloaded event ts>(falling back to'all'when the REST result is empty or errored). The legacyresend_all=trueflag is removed for the main connection; the planning-agent sub-conversation still usesresend_alluntil it is migrated to the same REST-then-WS pattern. -
The event store gained a bulk
addEvents(events)action (used for the initial REST seed and for "scroll-up" pagination) that re-sorts by timestamp once at the end so older pages can be merged in cheaply. Per-event dedup still works via the existingeventIdsset. -
ChatInterfacewiresuseLoadOlderEventsinto its scroll handler (threshold 80px from the top), shows adata-testid="loading-older-events"spinner during pagination, and preserves the visible scroll offset by storing the previousscrollHeightand adding the height delta after the older page renders. -
useLoadOlderEventsintentionally distinguishes between "no anchor yet" (empty store before the initial REST seed, soloadOlder()should no-op) and "malformed oldest event" (store has an oldest event with no timestamp, so the hook throws, flipshasMorefalse, andChatInterfacesurfaces the failure via the shared error banner instead of failing silently).
-
-
Action grouping in the chat stream:
src/components/conversation-events/chat/group-events.tsfolds runs of consecutive groupable events (regularActionEvent/ObservationEventcards, but notFinishAction,ThinkAction,PlanningFileEditorObservation,TaskTrackerObservation, hooks, errors, or message events) into singleRenderedItemgroups. The threshold lives inEVENT_GROUP_MIN_SIZE(currently 2, so even pairs of back-to-back actions get folded).EventGroup(src/components/conversation-events/chat/event-message-components/event-group.tsx) is the collapsible header that wraps each run. Default state is collapsed; the header showsEVENT_GROUP$ACTIONS_COMPLETED(with a success check) when the group is done, orEVENT_GROUP$ACTIONS_PROGRESSplus the currently-running action's title (fromgetEventContent) while a memberActionEventhas not yet been replaced by its observation in the UI events array. Expanding renders the originalEventMessages verbatim so each card still expands the way it did before.- Agent thoughts attached to an
ActionEvent(event.thought) are hoisted out of groups:groupEventsemits a thirdRenderedItemkind"thought"whenever a groupable event carries (or, for an observation, originates from) a non-empty thought, flushing the current run and starting a new one.messages.tsxrenders that item viaThoughtEventMessageand passessuppressThoughttoEventMessageso the inline thought isn't duplicated inside the group's expanded content.ThinkActionis excluded from this hoisting because the thought IS its action body and is rendered through its own codepath. groupEventsnow de-duplicates hoisted thoughts by action ID so mixed UI arrays that temporarily contain both an action and its replacement observation do not emit the same thought twice;minSizeis treated as a validated internal invariant (>= 1).EventGroupshould returnnullfor an emptyeventsarray and wire the toggle button to the expanded body witharia-controls/role="region"/aria-labelledby.src/components/conversation-events/chat/messages.tsxis the only consumer; the grouping is transparent to upstream code. Coverage lives in__tests__/components/conversation-events/chat/group-events.test.ts(pure logic, including thought hoisting) and__tests__/components/conversation-events/chat/event-message-components/event-group.test.tsx(rendering/interaction).
-
Home page workspace UX (agent-server backend):
RepoConnectorno longer renders a tabbed launcher; it just rendersWorkspaceSelectionFormbecause this build only ever talks to an agent-server backend (no cloud backend is wired up). The oldLaunchTabscomponent was removed; if a cloud backend is ever supported again, branch on backend mode inRepoConnectorand renderRepositorySelectionFormfor that path.FolderBrowserModal's "Use this folder" button adds only the currently navigated directory as a single workspace (named by its basename). It no longer iteratessubdirsand adds each child as a separate workspace.- The
WorkspaceDropdownsticky footer now exposes both "+ Add Workspace" (opens the folder browser) and "Manage Workspaces" (opensManageWorkspacesModal, which lets users remove individual workspaces viauseWorkspacesStore.removeWorkspace). The Manage button is hidden when there are no workspaces yet. - The sidebar "+ New Conversation" trigger (
NewConversationButtoninsrc/components/features/conversation-panel/) opens a popover that is a flat list, not the home-screen combobox: a leading "No workspace" entry plus one entry per stored workspace, each clicking through touseCreateConversationimmediately (no separate Launch button). It mirrors the dropdown footer actions/pattern (+ Add Workspace,Manage Workspaces) locally rather than embeddingWorkspaceDropdownitself. useResolvedWorkspaces()now returnsisLoading/isErrorfor parent-directory scans;WorkspaceSelectionFormshould surface that state (status text and disabling the empty dropdown while parent results are still loading) instead of assuming the merged list is immediately ready.ManageWorkspacesModalshould require a confirmation step before removing either a saved workspace or a workspace parent; parent removals should mention the child-workspace impact, and tests should assert both the confirmation flow and that removing the selected workspace clears the launch selection.- In
useWorkspacesStore, keepclearWorkspaces()scoped to literal workspaces only; use explicit helpers likeclearWorkspaceParents()/clearAll()for broader resets so future callers do not accidentally wipe parent registrations.
-
Default LLM model —
DEFAULT_SETTINGS.llm_model("openhands/minimax-m2.7", defined insrc/services/settings.ts) is the canonical frontend default.buildConfiguredOpenHandsAgentSettingsinsrc/api/agent-server-adapter.tsalways sends this value explicitly when the resolvedllm.modelis absent, empty, or whitespace-only — the frontend never relies on the agent-server SDK's own default (gpt-5.5). If you change the default model, updateDEFAULT_SETTINGS.llm_modelinsrc/services/settings.tsand the checklist inspecs/llm-defaults.md. Spec:@spec LLD-001. -
Custom secrets are NOT auto-attached by the agent-server.
POST /api/conversationsonly persists what the client sends inrequest.secrets; the persisted secrets store (/api/settings/secrets) is never read at conversation-start.buildStartConversationRequestWithEncryptedSettingsenumeratesSecretsService.getSecrets()and turns each entry into aLookupSecretwhoseurlpoints back at/api/settings/secrets/{name}and whoseheaderscarryX-Session-API-Keyfor auth. Pre-1.21.x agent-server SDKs would silently drop that header during validation whensecrets_encrypted=true(the cipher in the validation context tried tocipher.decrypt(plaintext_session_key), failed, and the validator removed the header — the conversation runtime then got 401s for every saved secret). The SDK fix preserves plaintext header values when decryption fails; if you still see saved secrets unavailable inside a conversation, verify the running agent-server bundles aLookupSecret._validate_secretsthat falls back to plaintext on decrypt failure. -
MCP page layout: MCP is a top-level nav entry at
/mcp(rendered bysrc/routes/mcp.tsx), shown right below "Skills" insrc/components/features/sidebar/sidebar.tsx. The legacy/settings/mcproute still works as a redirect viasrc/routes/mcp-settings-redirect.tsx, andsrc/routes/mcp-settings.tsxre-exports the new page so the publishedMCPSettingslibrary symbol (insrc/components/settings/index.ts) keeps the same shape. Marketplace catalog data and MCP logo mappings live in the MCP-capable entries from@openhands/extensions/integrations; the Slack API catalog option should point athttps://github.com/zencoderai/slack-mcp-serverand use@zencoderai/slack-mcp-server. Deprecated marketplace entries removed upstream (for example GitLab / Google Maps / Postgres / Puppeteer / SQLite) should disappear from the marketplace grid. The Installed section still needs to render and search arbitrary non-catalog custom servers via the raw servername/commandfallback insrc/utils/mcp-marketplace-utils.ts+InstalledServerCard. Tavily is a regular stdio MCP entry (tavily-mcp+TAVILY_API_KEY), not a special built-in sentinel anymore. Components are colocated undersrc/components/features/mcp-page/and reuse the existingMCPServerFormfor the "Add custom server" / edit flow. -
Library packaging notes:
- Public npm entrypoints now come from
src/index.ts→src/lib/index.ts, with domain barrels undersrc/components/{conversation,terminal,browser,files,settings,sidebar}/index.ts. npm run buildremains the standalone app build (react-router build), whilenpm run build:librunsvite buildin library mode plustsc -p tsconfig.lib.jsonto emit.d.tsfiles intodist/.- The library build relies on
vite.config.tswithBUILD_LIB=true, preserved modules indist/, and packageexportsentries that map root/subpaths todist/**/*.jsplus matching declaration files. - Declaration emit needs
src/library-env.d.tsand the narrowedtsconfig.lib.json; broadsrc/**/*.tsxdeclaration builds pulled in route-only files and missed?react/window globals.
- Public npm entrypoints now come from
-
Bundle/dev-graph hygiene (Tier 1 cleanup landed):
src/i18n/translation.json(~1 MB) is imported only bysrc/i18n/resources.ts, whichsrc/i18n/index.tsre-exports astranslationResourcesfor the@openhands/agent-canvas/i18nsubpath. The re-export is aexport … fromplus/* @__PURE__ */annotation, so rollup drops the JSON from the app build (prodcustom-toast-handlerschunk: 909 KB -> 74 KB;conversationchunk: 728 KB -> 392 KB). Do not move the JSON import back intosrc/i18n/index.ts— that immediately re-bundles all translations into every chunk that importsi18n.- The environment-switch overlay is split: lightweight store/triggers live in
components/features/backends/environment-switch-store.ts; the React component lives inenvironment-switch-overlay.tsx(re-exports the store API for back-compat). Eagerly-mounted callers (e.g.backend-selector.tsx) MUST import trigger helpers from the store, not the overlay file. The overlay isReact.lazy'd fromroutes/root-layout.tsx. - Other always-conditional UI is
React.lazy'd to keep the root layout's eager graph small:AnalyticsConsentFormModal,AlertBanner(root-layout),SettingsModal(sidebar), and the unreachable-backend modal path inroot.tsx(ManageBackendsModal). Tests that assert on these mounted nodes may needawait screen.findByTestId(...)/waitFor(...)instead of synchronousgetByTestId(...). - The terminal tab (
components/features/terminal/terminal.tsx) isReact.lazy'd inconversation-tab-content.tsxalongside the other tabs, so xterm + addon-fit + xterm.css don't enter the conversation route's eager graph (they ship as a separateterminal-*.jschunk now). - Avoid importing app code through
#/components/conversation-events/chator itsevent-message-components/index.tsbarrel — they exist forlib/index.ts(npm subpath) consumers only. Internal callers use deep paths (./messages,./event-message-components/<name>,./event-content-helpers/should-render-event) so Vite dev doesn't fan out the barrel.
-
Backend dropdown connectivity indicator:
useBackendsHealth(src/hooks/query/use-backends-health.ts) polls each registered backend every 10s. Local agent-server backends are probed viaServerClient.getServerInfo()(/server_info); cloud backends are probed viagetCurrentCloudApiKey()(/api/keys/currentthrough the bundled/api/cloud-proxy). Verdicts are surfaced as a colored dot rendered throughDropdownOption.prefix(added tosrc/ui/dropdown/types.ts); the trigger reads its prefix from the liveoptionsarray (not downshift's frozenselectedItem) so the indicator updates without remounting. The same dot is also rendered in each row ofManageBackendsModal, which now opts into a one-shot re-probe for previously disabled backends so opening the modal can clear stale persisted error state when a server has recovered. Tests live in__tests__/hooks/query/use-backends-health.test.tsx, theconnection indicatorblock of__tests__/components/backends/backend-selector.test.tsx, and__tests__/components/backends/manage-backends-modal.test.tsx. -
Manage Backends modal:
src/components/features/backends/manage-backends-modal.tsxlets users edit (host/name/api-key/kind) and remove existing backends, plus add new ones inline via a "+ Add Backend" footer button that opens aBackendFormModal. Both the dropdown footer's "Add backend" and the manage modal's "+ Add Backend" reuseBackendFormModal(seebackend-form-modal.tsx), withmode="add"ormode="edit";AddBackendModalis now a thin compatibility wrapper forBackendFormModal mode="add". The modal is also auto-rendered (with a no-oponClose) bysrc/root.tsxwhen the active backend is unreachable, replacing the old full-screenMissingAgentServerNoticeonboarding screen. -
Conversation right-panel regression note:
ConversationTabsnow owns the moved refresh/build buttons, so__tests__/components/features/conversation/conversation-tabs.test.tsxshould cover that behavior directly. The drawer's open/closed state (isRightPanelShown/hasRightPanelToggled) is intentionally session-only: it always starts closed on app load (or on opening a fresh/existing conversation after a restart), but it survives in-app navigation because the ZustanduseConversationStorestays alive across React Router transitions. TheConversationStatelocalStorage blob (conversation-state-{id}) deliberately does not carry arightPanelShownfield —useConversationLocalStorageStatedoes not expose asetRightPanelShownsetter,sanitizeStoredStatestrips the legacyrightPanelShownkey from older persisted blobs on read, andRightPanelToggle/useSelectConversationTabonly mutate the in-memory store. In tests, seed the Zustand store directly forselectedTab/isRightPanelShown/hasRightPanelToggled(the component sync effect currently restores onlyselectedTabfrom localStorage, so localStorage alone will not make a tab read as active or the drawer read as open). -
Changes tab /
FileDiffViewerdeleted-file note: the agent-server's/api/git/diffendpoint callspath.exists()first (seeopenhands-sdk/openhands/sdk/git/git_diff.py→get_git_diff), so requesting a diff for aD(deleted) file returnsGitPathError→ HTTP 400 and trips the global QueryCache error toast.useUnifiedGitDiffdisables the query whentype === "D"andFileDiffViewerrenders a localized "file deleted" placeholder (DIFF_VIEWER$FILE_DELETED,data-testid="file-deleted-message") instead of the view-mode toolbar / Monaco editor for that case. -
Onboarding modal:
src/components/features/onboarding/onboarding-modal.tsxis a 4-step welcome flow rendered by<OnboardingHost />(mounted on the home route) and gated by theopenhands-onboardedlocalStorage flag (use-onboarding-completion.ts). The four steps live understeps/: choose-agent (Step 0 – OpenHands selectable, Claude Code & Codex disabled with a "coming soon" note), check-backend (embeds the newBackendFormextracted frombackend-form-modal.tsxplus a colored connection banner driven byuseBackendsHealth), setup-llm (renders<LlmSettingsScreen onSaveSuccess={onNext} />so the existing settings UI keeps owning validation), and say-hello (text input pre-filled fromONBOARDING$HELLO_DEFAULT_MESSAGE, launches a no-workspace conversation viauseCreateConversationand closes the modal). Animation: all four panels are mounted as siblings inside a horizontal rail; advancing/retreating just setscurrentStep, which translates the rail by-(step * 100)%for the slide effect. Progress is rendered byOnboardingProgressBarwithdata-stateper segment (completed|current|upcoming). When extending, refactorBackendFormModalcarefully — the innerBackendFormis the public surface used both by the modal and byCheckBackendStep; the modal version still owns dirty/save tracking so it keeps "Save"/"Cancel" footer behavior. -
Worktree policy (this conversation): commits are made on the worktree branch and the user expects the worktree to stay attached to that branch. Do NOT run
git switch --detachin the worktree and reattach the branch to the main workspace after each commit — only do that when the user explicitly asks. See~/.openhands/skills/worktree-switch/SKILL.mdfor the manual procedure the user invokes. -
Files tab diff-view default logic: keyed off
useHasAttachedSource()(src/hooks/use-has-attached-source.ts), which is true when the user explicitly attached either a repo (conversation.selected_repository) or a local workspace (getStoredConversationMetadata(id).selected_workspace, persisted bycreateConversationwhenworkingDirOverrideis supplied). The agent-server pre-initialises every conversation workspace as a git worktree for its own change tracking, so do NOT use a filesystem probe (git status/useUnifiedGetGitChanges) as the attachment signal — that was tried in earlier iterations and made every fresh no-attachment conversation incorrectly default to diff view. The companionuseHasGitCommitsprobe (src/hooks/query/use-has-git-commits.ts) then suppresses diff view for attached-but-empty cases (unborn HEAD, non-git workspace). -
Files tab diff-view default logic: keyed off
useHasAttachedSource()(src/hooks/use-has-attached-source.ts), which is true when the user explicitly attached either a repo (conversation.selected_repository) or a local workspace (getStoredConversationMetadata(id).selected_workspace, persisted bycreateConversationwhenworkingDirOverrideis supplied). The agent-server pre-initialises every conversation workspace as a git worktree for its own change tracking, so do NOT use a filesystem probe (git status/useUnifiedGetGitChanges) as the attachment signal — that was tried in earlier iterations and made every fresh no-attachment conversation incorrectly default to diff view. The companionuseHasGitCommitsprobe (src/hooks/query/use-has-git-commits.ts) then suppresses diff view for attached-but-empty cases (unborn HEAD, non-git workspace). -
Collapsible thinking:
ThinkActionevents and LLM extended reasoning (reasoning_content/thinking_blocksonActionEvent) are rendered as collapsible sections viaCollapsibleThinking(src/components/conversation-events/chat/event-message-components/collapsible-thinking.tsx). Collapsed by default to keep the chat compact — the thinking is often in English regardless of the user's conversation language. ThegetReasoningContent()helper inevent-thought-helpers.tsextracts the content, preferringreasoning_content(plain string) and falling back to Anthropicthinking_blocks. i18n keys:THINKING$TITLE,THINKING$EXPAND,THINKING$COLLAPSE. Tests:__tests__/components/conversation-events/chat/event-message-think-action.test.tsx. -
Agent delegation settings: the
Settings > Agentpage (src/routes/agent-settings.tsx) is intentionally NOT aSdkSectionPagewrapper. It mirrors upstream OpenHands#14418 — it flatMaps every section ofagent_settings_schemaand finds theenable_sub_agentsfield by key, so it works regardless of which section the real backend exposes the field in. Don't refactor it back toSdkSectionPageunless you also know the real backend's section name and add a fallback for the live "SDK schema unavailable" path. The toggle persists viaagent_settings_diff. Nav item lives inOSS_NAV_ITEMS(settings-nav.tsx) with the robot icon (SETTINGS$NAV_AGENT). The mock schema insettings-handlers.tsputs the field in ageneralsection. Client-side gate:getAgentTools()inagent-server-adapter.tsonly attachestask_tool_setto new conversations whenagent_settings.enable_sub_agents === true. Without that gate the agent server would still receive the tool whenever it advertised it in/api/server_info, so the toggle had no effect on running conversations. -
Settings naming is backend-aware today: local
/settingsis profile-oriented (use-settings-nav-items.tsrenames the first settings item/title/subtitle toLLM Profilesandchat-input-model.tsx/chat-input-actions.tsxlink there asLLM Profiles), while cloud keeps the genericLLM Settingscopy because cloud still edits raw settings rather than saved profiles. The local profile editor (llm-settings-local-view.tsx) should keep explicit create/edit profile headings plus helper text so users know they are saving a profile, not mutating the current conversation directly. -
ESLint config (flat, ESLint 9): the project uses
eslint.config.js(not.eslintrc) and runs oneslint@9.x, not 10. The constraint pinning us below 10 iseslint-plugin-react@7.37.x, which still callscontext.getFilename()at rule-load time — that API was removed in ESLint 10 and@eslint/compat'sfixupPluginRulesdoes NOT shim it. Don't try to bump eslint past 9 until eslint-plugin-react ships a v10-compatible release. Import rules come fromeslint-plugin-import-x(the maintained fork ofeslint-plugin-import) but are registered under bothimport-x/andimport/prefixes viaplugins: { import: importXPlugin, ... }so existing// eslint-disable-next-line import/...directives keep working.linterOptions.reportUnusedDisableDirectivesis set to"warn"(not "off") so stale airbnb-era disable comments still surface in lint output without failing CI. The TS-overrides block has anignores: ["src/hooks/query/query-keys.ts"]so theno-restricted-syntaxrule banning raw["settings", ...]query keys doesn't fire on the file that defines the helpers themselves. No.npmrc/legacy-peer-depsflag is needed — all our plugins declare ESLint 9 peer compatibility. -
Centralized config:
config/defaults.jsonis the single source of truth for version pins (agent-server, automation, automation SDK), port defaults, persistence paths, and package names. All consumers read from this file:- JS scripts (
dev-safe.mjs,dev-with-automation.mjs,check-sdk-version-sync.mjs) read it viaJSON.parse(readFileSync(...)). - Docker: a
config-genbuild stage converts the JSON to/opt/agent-canvas/defaults.env(shell-sourceable);entrypoint.shsources it at startup. - CI workflow: a
Read defaults from config/defaults.jsonstep usesnode -pto extract values into$GITHUB_OUTPUT. - Dockerfile ARG defaults are kept as fallbacks for local
docker buildwithout the CI workflow; CI always passes--build-argoverrides from the JSON. - To bump a version, edit
config/defaults.jsononly — the JS scripts, Docker build, and CI workflow all derive their values from it.
- JS scripts (
-
Docker all-in-one image:
.github/workflows/docker.ymlbuilds and publishesghcr.io/openhands/agent-canvas— a combined image that bundles the agent-server (fromghcr.io/openhands/agent-server), the automation server (openhands-automationvia pip), and the agent-canvas frontend (static build). The Dockerfile lives atdocker/Dockerfile, the entrypoint atdocker/entrypoint.sh. The workflow structure mirrors the SDK repo'sserver.yml: abuild-and-push-imagematrix job (2 × arch: amd64 onubuntu-24.04, arm64 onubuntu-24.04-arm) pushes arch-suffixed tags, thenmerge-manifestscreates multi-arch manifests viadocker buildx imagetools create, thenconsolidate-build-infoaggregates artifacts, andupdate-pr-descriptionupdates the PR body (using<!-- AGENT_CANVAS_DOCKER_START -->/<!-- AGENT_CANVAS_DOCKER_END -->markers). The workflow triggers on push to main,v*tags (releases), PRs, andworkflow_dispatch. On release tags it also pushes semver tags (e.g.1.2.3,1.2,1,latest). Fork PRs are skipped (no GHCR auth). The image exposes port 8000 as a unified entry point:/api/automation/*→ automation (:18001),/api/*→ agent-server (:18000),/*→ static frontend. The Dockerfile accepts aVITE_APP_ENVbuild arg (default empty → staging PostHog key); the CI workflow passesVITE_APP_ENV=productiononly for tagged releases (refs/tags/v*), so PR and main-branch images use the staging key while release images use the production key, matching thebuild:libnpm path. The entrypoint auto-generates both the session API key andOH_SECRET_KEY(persisted to~/.openhands/agent-canvas/session-api-key.txtandsecret-key.txtrespectively) when none is provided, so the image runs secure by default. Users can override either via env var (OH_SECRET_KEY,SESSION_API_KEY/OH_SESSION_API_KEYS_0).scripts/dev-safe.mjsuses the samesecret-key.txtfile, so dev mode and Docker share the same key when both use the same~/.openhandsdirectory. -
Spec files live under
specs/. Spec IDs are stable — never renumber. Mark deprecated specs withstrikethrough. Tag implementation code and tests with// @spec BM-002 — Short titlecomments so specs are grep-able across the codebase (grep -rn '@spec BM-' src/ __tests__/). Place the comment on the line immediately above the relevant code block or test. When multiple tests cover the same spec, useit.eachif the test structure is identical. -
Release automation: releases use a long-lived release branch model — a
rel-X.Y.Zbranch is created frommain, QA/fixes land there, and publishing is triggered by pushing av*tag directly to that branch (the branch is never merged back to main). The dist-tag is resolved dynamically at publish time by querying npm: if no stable version (no pre-release suffix) has ever been published, all releases use--tag latest; once a stable version exists, pre-release versions get their own dist-tag (alpha/beta/rc) and only stable versions keeplatest. The transition is automatic — no workflow change is needed when the first stable release ships. Three workflows fire in parallel on everyv*tag push:create-release.yml(creates the GitHub Release object with auto-generated notes and marks pre-release for hyphenated versions),npm-publish.yml(builds and publishes to npm with the correct dist-tag), anddocker.yml(builds multi-arch Docker images). The release skill (.agents/skills/release.md, keyword trigger:release) guides agents through the full process. -
Cloud conversation resume gating: when a cloud conversation is closed from the UI (
pauseCloudSandboxis called), the conversation'sconversation_urlis NOT cleared -- it still points to the old sandbox host.WebSocketProviderWrappermust suppress the URL (passnulltoConversationWebSocketProvider) whilesandbox_status === "PAUSED", otherwise the WebSocket immediately tries the stale URL before the sandbox wakes. Symmetrically,useActiveConversation's refetch interval must fast-poll (3 s) on both!conversation_urlANDsandbox_status === "PAUSED"-- checking only the missing URL would leave the hook on the 30 s interval while the sandbox is resuming. The resume sequence: navigate -> sandbox PAUSED detected ->resumeCloudSandboxcalled (inconversation.tsx) -> fast-poll detects RUNNING ->conversationUrlunblocked -> WebSocket connects.