* fix(streaming): dedupe streamed reasoning against the action that renders it
The intermediate-action reconciliation only strips a streamed delta once the
streamed text is matched against the action's `thought`. A model that streams
reasoning without a matching thought — a reasoning-only delta, or an action
whose thought is empty — never reaches that check, so the delta's reasoning and
the action's own reasoning both render: two identical "Thinking" bubbles.
Reconcile reasoning independently of the text match. When the arriving action
renders its own reasoning, clear the reasoning from the current step's trailing
delta run, keeping the delta only for streamed text the action does not render.
A delta that is the sole reasoning carrier is left untouched.
* fix(streaming): scope streamed-reasoning dedup to the same sender
`getTrailingReasoningDeltas` collected the whole trailing delta run
regardless of which agent produced it. The main and planning sockets share
one event store, so an arriving main-agent `ActionEvent` stripped the
planning agent's still-live reasoning along with its own — the planning
"Thinking" disappeared from the chat entirely.
Scope the selector with `isSameStreamingSender`, the same rule
`getCurrentTurnContentDeltas` already applies on the finalize path (#1656).
The typescript-client (pinned 1.34.0) now exports available_models and
default_model directly on ACPProviderInfo, so the local intersection
workaround that claimed those fields were missing from the client type is
obsolete. Use getAcpProvider()'s return type directly and add a test
locking in that the registry data fields (display_name, default_command,
available_models, default_model) are sourced from the SDK, not a local
mirror.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix(streaming): render final event authoritatively and dedupe reconnect side-effects
Fixes the reconciliation and reconnect correctness bugs audited in #1656.
- Reconciliation: the final MessageEvent/FinishAction is authoritative. Drop the
provisional streamed deltas and render the canonical final event, so an
assistant reply renders exactly once — never holey after a mid-stream
disconnect, never duplicated when the SDK reshapes the finalized text (e.g.
strips leading whitespace) — and its metadata (critic_result,
activated_microagents) renders too. Stream-only reasoning is preserved.
Adopts the design from #1695.
- Reconnect: a WebSocket reconnect replays the backlog from a stale anchor. The
event store dedups by id, but the side-effects did not. Guard the terminal
append, error banner/analytics, cache invalidation and one-shot UI actions
behind isDuplicateEvent in both the main and planning handlers.
- Merge guard: only merge consecutive streaming deltas that share a sender, so a
planning-agent delta cannot concatenate onto a main-agent one and be
misattributed.
* fix(streaming): scope final-event reconciliation to the same sender
finalizeStreamingDeltasInPlace collected every content delta after the
last user message without comparing isFromPlanningAgent. Because the main
and planning sockets share the event store, a final event from one agent
could strip the other agent's still-live streamed deltas before pushing
its own canonical event. Scope the candidate deltas with
isSameStreamingSender so finalization only supersedes its own stream, and
cover both interleaving directions with regression tests (#1656).
* Suppress telemetry consent modal in Cloud Canvas
Treat same-origin locked Cloud cookie deployments as already consented for Canvas library telemetry so the modal does not flicker before the main app login flow.
Co-authored-by: openhands <openhands@all-hands.dev>
* Stabilize mock LLM settings tests after ACP
Reset the default agent profile back to OpenHands through the agent-profile API before LLM-profile setup paths that need /settings/llm. This prevents an ACP profile left by the previous serial spec from redirecting later settings tests to /settings/agents.
Co-authored-by: openhands <openhands@all-hands.dev>
* Stabilize files tab mock E2E git setup
Ensure the attached-workspace conversation has both an origin remote and a real HEAD commit before asserting that the Files tab defaults to diff view. The diff default now intentionally depends on both attached source metadata and an available commit base.
Co-authored-by: openhands <openhands@all-hands.dev>
* Fix desktop right panel toggle visibility
Update the desktop right-panel toggle to set both the user-toggled flag and the visible state. The missing visibility update left the panel visually closed in mock E2E while off-screen tab controls remained mounted.
Co-authored-by: openhands <openhands@all-hands.dev>
* Make files tab mock E2E open the panel explicitly
Wait for the desktop panel toggle to report an open state before interacting with Files tab controls, then verify the Diff segment can be selected for the attached-workspace conversation.
Co-authored-by: openhands <openhands@all-hands.dev>
---------
Co-authored-by: openhands <openhands@all-hands.dev>
Dependabot's npm config used commit-message prefix `deps`, which is not in
the conventional type list the pr-title check accepts, so every npm PR failed
the title lint. Switch it to `chore`.
The PR description check ran on all non-draft PRs, but bots write their own
bodies and cannot follow the HUMAN/AGENT template, so it failed on every
dependabot and release-please PR. Skip it when the author is a Bot.
Co-authored-by: aivong-openhands <ai.vong@openhands.dev>
On Windows, spawnService ran uvx through cmd.exe (shell: true), so the `<`
in the `agent-client-protocol<0.11` version constraint was parsed as input
redirection and agent-server exited immediately with "The system cannot find
the file specified."
Resolve the command to its absolute path with where.exe and spawn it directly,
with no shell, so argument metacharacters stay literal. npm is unaffected: it is
already wrapped in cmd.exe by buildNpmScriptCommand before it reaches
spawnService.
The Kimi Code membership (Moonshot) is reachable via a Console API key
against https://api.kimi.com/coding/v1 (model kimi-for-coding), distinct
from the OpenPlatform at api.moonshot.ai/v1. Two small canvas changes so
the membership endpoint behaves like a first-class provider rather than a
custom one:
- map-provider: label the moonshot provider as Moonshot (was the raw id)
- llm-settings: treat api.kimi.com/coding/v1 as a known provider default
base_url so the LLM form opens in basic view instead of the advanced view
Pairs with a software-agent-sdk change that verifies kimi-for-coding in
VERIFIED_MOONSHOT_MODELS, earning it the Verified badge in the model list.
* feat: support serving canvas under subpath
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: redirect root app routes to canvas base path
Co-authored-by: openhands <openhands@all-hands.dev>
* Support cookie auth for locked Cloud Canvas
Co-authored-by: openhands <openhands@all-hands.dev>
* Use main app login for cookie Cloud Canvas
Co-authored-by: openhands <openhands@all-hands.dev>
* Allow main app auth probe exception
Co-authored-by: openhands <openhands@all-hands.dev>
* Use OpenHands returnTo login parameter
Co-authored-by: openhands <openhands@all-hands.dev>
* Prefer Cloud current org in backend selector
Co-authored-by: openhands <openhands@all-hands.dev>
* Render Canvas home at root path
Co-authored-by: openhands <openhands@all-hands.dev>
* Fix locked Cloud auth during domain transition
Treat openhands.dev and all-hands.dev Cloud hosts as equivalent for locked cookie auth, seed cookie backends on the current origin, and let main-app login redirects run before Canvas onboarding.
Co-authored-by: openhands <openhands@all-hands.dev>
* Fix locked backend seeding in mocked config tests
Preserve the existing early exit when no Cloud lock is configured so tests and local flows with partial agent-server-config mocks still seed the default local backend.
Co-authored-by: openhands <openhands@all-hands.dev>
---------
Co-authored-by: openhands <openhands@all-hands.dev>
* refactor: define canvas UI as an SDK client tool
Send a JSON-defined canvas_ui_client tool on new, profile-based, and resumed conversation requests while retaining the legacy Python registration for persisted conversations. Normalize the new SDK event kinds to the existing Canvas UI rendering.
Co-authored-by: smolpaws <engel@enyst.org>
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: omit canvas client tool from ACP launches
* refactor: rename canvas client tool
Use the semantic canvas_ui_control name and contain the SDK-generated action discriminator behind exported constants.
Co-authored-by: Engel Nyst <engel.nyst@gmail.com>
---------
Co-authored-by: Engel Nyst <engel.nyst@gmail.com>
Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: Debug Agent <157206163+simonrosenberg@users.noreply.github.com>
* feat(helm): add Helm chart for StatefulSet deployment
Adds helm/agent-canvas/ chart that deploys the all-in-one agent-canvas
image as a StatefulSet with:
- PVC via volumeClaimTemplates mounted at $HOME/.openhands so
settings, encrypted secrets, conversation history, automation
SQLite DB, workspaces, and auto-generated keys all persist across
pod restarts and image upgrades. Supports BYO PVC via
persistence.existingClaim.
- Ingress template with the usual knobs (className, annotations,
hosts[].host + paths[].path/pathType, tls[].hosts/.secretName).
- ClusterIP + headless Services (headless required by StatefulSet).
- ServiceAccount always created for stable pod identity; RBAC
bindings gated by rbac.enabled, with:
* rbac.namespaces: list of namespaces to bind the SA into
via one RoleBinding per namespace against the built-in
`admin` ClusterRole (full namespace-scoped access).
* rbac.clusterAdmin: bool, off by default, adds a
ClusterRoleBinding to `cluster-admin` when true.
- Probes on /alive, podSecurityContext matching UID 1000 with
fsGroup 1000 so the PVC is writable, extraEnv passthrough for
LLM keys, optional external Postgres via AUTOMATION_DB_URL, and
optional existingSecret refs for OH_SECRET_KEY / session key.
helm lint passes; helm template verified against defaults,
full RBAC + ingress + TLS, existingClaim, and persistence.enabled=false.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix(helm): pin appVersion to published agent-canvas tag 1.1.0
The initial commit accidentally pulled 1.23.0 from config/defaults.json,
which is the agent-server SDK version — not the ghcr.io/openhands/agent-canvas
image tag. Point appVersion at versions.agentCanvas (1.1.0) instead so the
default 'helm install' pulls a tag that actually exists on GHCR.
* feat(helm): mount entire $HOME on the PVC instead of just ~/.openhands
The previous default mounted the PVC at /home/openhands/.openhands, so
only the .openhands subtree survived pod restarts. Anything the agent
wrote under $HOME/workspace (cloned repos, generated files, worktrees)
was lost the next time the pod was rescheduled — which is the common
case since agents default to $HOME/workspace as the CWD.
Move the mount up to /home/openhands so the whole HOME persists:
* ~/.openhands — as before
* ~/workspace — the agent's actual working tree
* dotfiles the user creates (~/.gitconfig, ~/.cache, ~/.local, ~/.ssh)
Safe because the upstream image doesn't ship any dotfiles or venvs
under /home/openhands — Python packages are --system-installed under
/opt/agent-canvas — so an empty PVC on first boot doesn't shadow
anything, and the entrypoint recreates the .openhands subtree.
Migration note for existing installs: this changes the default
persistence.mountPath. Existing releases already using this chart
have a PVC whose ROOT is the .openhands directory; if they upgrade
without pinning the old mountPath, they'll get a fresh empty HOME
mounted over their existing .openhands contents and their state will
appear to reset. To preserve existing state, either:
* set persistence.mountPath: /home/openhands/.openhands to keep the
old behaviour, or
* delete the STS + PVC and start fresh.
* fix(helm): correct UID to 10001, mount PVC at ~/.openhands + ~/workspace via subPath
Two related bugs.
1. Wrong UID/GID/fsGroup
podSecurityContext pinned runAsUser/runAsGroup/fsGroup to 1000 based
on a stale assumption. The upstream agent-server image actually
runs as `openhands` at UID/GID **10001** (see
software-agent-sdk openhands-agent-server/openhands/agent_server/
docker/Dockerfile: `ARG USERNAME=openhands / UID=10001 / GID=10001`).
On Debian trixie, UID 1000 happens to be occupied by an unrelated
`pn` user carried in from a base layer, so `kubectl exec` shows
the wrong username and the process can't write to a PVC chowned to
gid 1000. Pin everything to 10001.
2. Wrong mount point
The previous commit moved persistence.mountPath from
/home/openhands/.openhands up to /home/openhands. That was wrong:
the image ships .bashrc, .profile, .bash_logout in that directory,
so mounting an empty PVC over the whole HOME shadows them and gives
users a bare-bones shell on `kubectl exec`.
Instead, mount the SAME PVC at multiple well-known subdirectories
via subPath. The default 'mounts' list covers the two paths that
actually need to persist:
- /home/openhands/.openhands (subPath: openhands)
- /home/openhands/workspace (subPath: workspace)
The pristine HOME (including dotfiles) stays intact from the image,
and both trees survive pod restarts and image upgrades on the same
disk. Users can extend the list with e.g. ~/.cache or ~/.config
without allocating additional PVCs.
Breaking change for existing installs relative to the previous
helm-chart-upstream tip commit — the values shape changed from a
single `persistence.mountPath` to a list `persistence.mounts`, and
the default UID is now 10001 instead of 1000.
* chore(helm): bump appVersion to 1.2.0
Point the chart's default image tag at the newly published
ghcr.io/openhands/agent-canvas:1.2.0. Chart version itself is
unchanged.
* fix(helm): keep version labels out of volumeClaimTemplates so upgrades work
Symptom: 'helm upgrade' failing with
StatefulSet.apps 'agent-canvas' is invalid: spec: Forbidden: updates
to statefulset spec for fields other than 'replicas', 'ordinals',
'template', 'updateStrategy', 'revisionHistoryLimit',
'persistentVolumeClaimRetentionPolicy' and 'minReadySeconds' are
forbidden
Cause: templates/statefulset.yaml stamps the full 'agent-canvas.labels'
set into 'volumeClaimTemplates[].metadata.labels'. That block lives
under STS.spec (NOT spec.template), which the apiserver treats as
immutable. The full label set includes 'app.kubernetes.io/version'
(from Chart.appVersion) and 'helm.sh/chart' (from Chart.version), both
of which change when either is bumped — turning routine version bumps
into forbidden-diff failures.
Fix: introduce a new helper 'agent-canvas.immutableLabels' that emits
only the labels that never change for a given release (name, instance,
managed-by), and use it in place of the full label set inside
volumeClaimTemplates. Object-level metadata.labels on the STS itself,
Services, ServiceAccount, etc. remain unchanged since those metadata
subtrees ARE mutable.
Existing installs one-shot recovery (before their next 'helm upgrade'):
# STS gone, pod + PVC preserved; helm upgrade will recreate STS
# from the fixed template and re-adopt both.
kubectl -n <ns> delete sts <release>-agent-canvas --cascade=orphan
helm upgrade <release> ./helm/agent-canvas -n <ns> -f values.yaml
* Apply suggestions from code review
Co-authored-by: Robert Brennan <accounts@rbren.io>
* Apply suggestion from @rbren
* Remove duplicate warning from README.md
Removed duplicate warning about the experimental nature of the Helm chart.
* docs(helm): clarify when to use this chart and relationship to OpenHands Enterprise
Address Graham's review feedback on PR #1641: the README didn't explain
why someone would want to use this chart or how it relates to OpenHands
Enterprise.
- Add a 'When to use this' section describing the self-hosted persistent
backend and internal vibecoding-platform use cases (pulled from the
companion docs PR OpenHands/docs#614).
- Add a 'Relationship to OpenHands Enterprise' section making the
unauthenticated / single-tenant / all-agents-comingled nature of this
chart explicit, and contrasting it with OHE's authentication,
role-based access control, multi-tenancy, and isolated agent
sandboxes.
- Add a 'Security' section covering authenticated ingress, LoadBalancer
exposure risk, and cluster-admin guidance, cross-referencing OHE.
Co-authored-by: openhands <openhands@all-hands.dev>
---------
Co-authored-by: openhands <openhands@all-hands.dev>
* fix(onboarding): cloud-only setup polish
- Update BACKEND$CLOUD_DESCRIPTION to 'Connect instantly to OpenHands Cloud to try out Agent Canvas'
- Show Add Backend modal instead of onboarding flow for locked Cloud mode
- Removes confusing 4 progress bars for Cloud-only setup
- Uses BackendFormModal with hideCloseButton=true for cleaner UX
- Hides 'Add Backend' header in locked mode, shows only Cloud login UI
- Increase CloudLoginColumn bottom padding pb-7 -> pb-8 for better spacing
Fixes#1536
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: prettier formatting
* test(e2e): deflake mock-LLM profile activation helper
The activateProfileViaUI helper was flaky under CI load: the menu-click
gesture (open menu → click "Set active") sometimes didn't register,
leaving the profile inactive and the subsequent badge poll to time out
after 30s. This caused intermittent mock-llm-e2e failures across many
PRs (including #1607) with different test subsets failing each run.
Drive activation through the agent-server API directly
(POST /api/profiles/:name/activate) instead of the menu UI. The
activation is enforced synchronously server-side, so by the time the
response returns the profile is active. The helper still polls the
Settings UI for the "Active" badge, preserving the verification that
the UI reflects the activated state.
Co-authored-by: openhands <openhands@all-hands.dev>
* test(e2e): retry ACP conversation on transient session_id race
The openhands-sdk's ACP agent has an intermittent race in the npm/uvx
CI environment where the first prompt is sent before the ACP session is
fully established, producing "validation errors for PromptRequest
session_id". The Docker E2E path doesn't hit this; the npm path does
only sometimes, flaking mock-llm-e2e for unrelated PRs.
Wrap the ACP conversation creation in a retry loop (up to 3 attempts).
On each failed attempt, delete the failed conversation and start a
fresh one from the home page. The reply-token wait is reduced to 45s
per attempt so the total stays within the test timeout.
Increase the Playwright globalTimeout from 600s to 720s and the
workflow deadline from 660s to 780s to accommodate the retry budget
without hitting the hard cap.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix(acp): pin agent-client-protocol<0.11 for SDK 1.29.x compatibility
agent-client-protocol 0.11.0 swapped the positional args of
ClientSideConnection.prompt() from (prompt, session_id) to
(session_id, prompt). openhands-sdk through at least 1.29.x calls
conn.prompt(prompt_blocks, session_id) positionally, so acp>=0.11
causes "2 validation errors for PromptRequest" — session_id receives
the prompt list and prompt receives the session_id string. This broke
the mock-LLM ACP E2E test (and any ACP agent conversation) for all PRs,
not just #1607.
Root cause confirmed via a minimal reproduction: spawning mock-acp-server.py
as a subprocess and calling conn.prompt(prompt_blocks, session_id) with
acp 0.11.0 produces the exact ValidationError; with acp 0.10.1 the call
succeeds and the mock server replies correctly.
Fix by pinning agent-client-protocol>=0.10.1,<0.11 in three places:
1. scripts/dev-safe.mjs — adds --with agent-client-protocol<0.11 to
all PyPI-based agent-server uvx invocations (both version-pinned
and default)
2. .github/workflows/mock-llm-e2e.yml — pins acp in .mock-llm-venv
(used by the mock ACP server subprocess in the npm/uvx path)
3. .github/workflows/mock-llm-docker-e2e.yml — same pin for the
Docker E2E path's mock-llm-venv
Also reverts the ACP test retry loop and timeout increases from
afefeb2 — the retry masked the real bug and made the suite slower.
The diagnostic events-API dump on failure is kept (with a larger
truncate limit) since it was useful for root-causing this issue.
Co-authored-by: openhands <openhands@all-hands.dev>
* test: update buildAgentServerCommand expectations for acp pin
Update the two PyPI-based test cases in dev-safe.test.ts to expect the
new --with agent-client-protocol>=0.10.1,<0.11 arg added by the
previous commit. Also reflow the AGENT_SERVER_EXTRA_WITH constant to
satisfy prettier.
Co-authored-by: openhands <openhands@all-hands.dev>
* test: cover locked Cloud first-run modal
* docs: add issue 1536 blocked QA evidence
Co-authored-by: OpenHands <openhands@all-hands.dev>
* chore: Remove PR-only artifacts
* chore: remove unrelated PR changes
---------
Co-authored-by: neubig <398875+neubig@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: Graham Neubig <gneubig@users.noreply.github.com>
Co-authored-by: hieptl <hieptl.developer@gmail.com>
Co-authored-by: allhands-bot <allhands-bot@users.noreply.github.com>
Co-authored-by: neubig <neubig@users.noreply.github.com>