Files
OpenHands/tests/e2e/live-acp
ec4616c1c7 feat(acp): containerized + cloud ACP — onboarding, secrets, and recycled-sandbox resume (#1013/#1014/#988) (#1102)
* feat(acp): containerized ACP — credential onboarding + inline secrets (#1013/#1014)

Wire the canvas halves of agent-canvas#1014 (Docker) and #1013 (credential
onboarding) so a user can run an ACP agent (Codex / Claude Code / Gemini)
against a containerized agent-server through Canvas, with credentials supplied
in the UI.

Credential onboarding UX (#1013):
- Extend the ACP secrets step beyond the API key to the per-provider reserved
  credentials a fresh container needs: Codex CODEX_AUTH_JSON, Claude
  CLAUDE_CODE_OAUTH_TOKEN, Gemini GOOGLE_APPLICATION_CREDENTIALS_JSON +
  GOOGLE_CLOUD_PROJECT/LOCATION + GOOGLE_GENAI_USE_VERTEXAI. File-content blobs
  render as multiline fields.
- Make the step capability-driven: required on a backend with no host login
  (cloud, or a logged-out local/Docker backend per the auth probe), optional
  when a login is detected or the probe can't classify (native dev).
- Fix the orphaned-secret bug: warn instead of toasting "Saved" when the active
  backend can't consume the credential (cloud can't yet read file secrets).

Send secrets + model (start request):
- buildStartConversationRequest emits reserved ACP credentials inline as
  StaticSecrets (overriding any same-named LookupSecret) and mirrors them onto
  agent_context.secrets, so the SDK's acp_file_secrets defaults materialise the
  *_JSON blobs before the CLI spawns. The orchestrator reads back the saved
  reserved values for the active provider (local backends only).
- Preselect a Vertex-safe acp_model for Gemini (gemini-2.5-flash) so a fresh
  container doesn't hit gemini-cli's preview default that 404s on Vertex.
- Never auto-promote *_BASE_URL to an inline secret (an inherited base URL
  breaks the Claude OAuth token's bearer auth).

Docker setup + docs:
- examples/acp-docker/ docker-compose (persistent volume + canvas_ui tool mount
  + credential notes); .env.sample + docs point VITE_BACKEND_BASE_URL at it.
- docs/ACP_AGENTS.md gains a "Running ACP agents in a Docker container" section.

Per-conversation isolation (acp_isolate_data_dir) left as a documented TODO —
the field isn't exposed on ACPAgentSettings in the released typescript-client.

Tests + e2e:
- Unit tests for the StaticSecret emission, reserved-credential sets, Vertex
  model default, getSecretValues read-back, and the required-credentials matrix.
- tests/e2e/live-acp/: a vite-node harness that builds each provider's request
  via buildStartConversationRequest and POSTs it to a real container. Validated
  with REAL API calls against agent-server c950fdb-python: Codex ✅, Claude ✅,
  Gemini ✅ (materialise ADC -> vertex-ai -> real reply). Gemini's default-config
  init is blocked by an SDK/gemini-cli set_session_mode("yolo") issue (documented
  caveat, not a credential problem).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(acp): make containerized credentials survive the real conversation-start path

Validating end-to-end through the application's own orchestrator
(buildStartConversationRequestWithEncryptedSettings) against a live container —
rather than the request builder in isolation — surfaced two real bugs that would
have broken the feature in the product:

1. secrets_encrypted mangled the plaintext reserved StaticSecrets. The app always
   fetches settings in encrypted mode, so the start request carried
   secrets_encrypted=true. The agent-server then runs every secret value through
   cipher.decrypt() during validation — including our reserved ACP creds, which
   are read back as PLAINTEXT. Result: the credential was silently dropped
   (decrypt fails → None) on a cipher backend, or a hard 500 ("cipher not
   configured") on a fresh container with no OH_SECRET_KEY. Fix: don't set
   secrets_encrypted for ACP conversations — an ACP agent has no encrypted agent
   secret (no LLM api_key), and its provider creds ride as plaintext StaticSecrets.

2. A different provider's leftover file-content secret broke the active provider.
   A CODEX_AUTH_JSON saved while onboarding Codex leaks into a later Claude
   conversation via the global-secrets → LookupSecret path. The SDK materialises
   file secrets eagerly at spawn by resolving the secret source, and a LookupSecret
   resolution stalls → ReadTimeout → "Failed to start ACP server: timed out". Fix:
   reserved file-content blobs (the multiline *_JSON creds) never travel as
   LookupSecrets — the active provider's is sent inline as a StaticSecret, any
   other provider's is dropped (getAllReservedAcpFileSecretNames).

Re-validated through the app orchestrator against agent-server c950fdb-python
(onboarding createSecret → buildAcpAgentSettingsDiff PATCH → orchestrator
read-back → real reply): Codex ✅, Claude ✅ (leftover CODEX_AUTH_JSON correctly
dropped). Gemini's app path is correct (StaticSecrets emitted, vertex-ai auth
reached); this run hit the documented invalid_rapt stale-ADC caveat (host ADC
expired since the prior fresh-ADC pass) — an environment issue, not code.

Adds regression tests (secrets_encrypted suppressed for ACP / kept for non-ACP;
leftover file blob dropped not LookupSecret'd; getAllReservedAcpFileSecretNames)
and the app-path e2e harness (tests/e2e/live-acp/acp-docker-app-e2e.mts). Notes
OH_SECRET_KEY as optional (secret persistence) in the compose example.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: address PR review feedback (#1102)

- retag acp_isolate_data_dir TODO #1014 (this PR) -> #1019 (the
  per-conversation isolation follow-up the knob serves)
- note the Gemini Vertex scalars (PROJECT/LOCATION/USE_VERTEXAI) are
  plain config / a routing flag, not secrets

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(acp): order subscription credential before API key in onboarding

Show each provider's reserved subscription/Vertex credential first
(Claude CLAUDE_CODE_OAUTH_TOKEN, Codex CODEX_AUTH_JSON, Gemini Vertex SA),
then the API key, then the base URL — the subscription token is the
primary auth path for ACP providers, with the API key as the fallback.
Display order only; getAcpProviderSecrets consumers are order-independent.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(i18n): disable i18next value escaping so React handles it

i18next's default escapeValue double-escapes interpolated values on top
of React's own escaping, rendering paths like ~/.codex/auth.json as
~&#x2F;.codex&#x2F;auth.json. Set interpolation.escapeValue=false (the
standard react-i18next config); React still escapes at render time.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(acp): unify secret wire-delivery; keep "reserved" as onboarding-only

Drop the reserved-vs-custom split in how secrets reach the agent-server.
Previously, provider credentials ("reserved") rode inline as StaticSecrets
while user secrets rode as loopback LookupSecrets — a fork introduced only
to dodge a deadlock: the SDK resolved an ACP agent's secrets synchronously
on its event loop at CLI spawn, so a loopback LookupSecret self-deadlocked.

That deadlock is fixed at the source in software-agent-sdk#3510 (ACP
cold-start runs off the event loop), so the workaround is no longer needed.
Now every secret — env-var credential, file-content blob, or user secret —
ships uniformly as a LookupSecret, for ACP and non-ACP alike. The SDK
resolves and (for file blobs) materialises them off the loop, so the
loopback fetch is safe.

"Reserved" survives only as an onboarding/validation concept (which fields
to prompt for per provider, capability-driven required steps) — it no
longer affects the wire.

Removed: StaticSecret type, acpStaticSecrets option + the inline path, the
file-blob lookupSkip, SecretsService.getSecretValues, and the reserved-name
value read-back. Kept: secrets_encrypted suppression for ACP (an ACP
request carries no encrypted payload, and a fresh ACP container may have no
OH_SECRET_KEY cipher).

Note: getReservedAcpSecretNames / getAllReservedAcpFileSecretNames in
constants/acp-providers.ts are now unused by the wire; the former is still
useful for validation, the latter can be pruned.

Depends on software-agent-sdk#3510.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(acp): prune now-dead reserved-secret wire helpers

Follow-up to the wire-delivery unification: getReservedAcpSecretNames and
getAllReservedAcpFileSecretNames were only ever consumed by the inline
StaticSecret / file-blob-skip path, which is gone. They have no remaining
production callers, so remove them (and their tests). The reserved-credential
field definitions (ACP_RESERVED_CREDENTIALS, getAcpProviderSecrets) and the
``reserved`` / ``multiline`` flags stay — onboarding still reads them.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(acp): re-point containerized ACP at SDK 1.25.0 (#3510) + fix e2e harnesses

The unified LookupSecret delivery (e076e9bb) depends on software-agent-sdk#3510
(ACP cold-start off the event loop), which first ships in v1.25.0. The example
compose/docs/e2e all still defaulted to agent-server:c950fdb-python, which
predates #3510 and deadlocks the first ACP turn ("Failed to start ACP server:
timed out"). Bump every default to 1.25.0-python and document it as the minimum.

Also realign the live-acp e2e harnesses, which still encoded the removed
StaticSecret API (the PR's headline evidence predated the unification):
- acp-docker-e2e.mts: store each credential via SecretsService.createSecret,
  send name-only customSecrets, assert every emitted secret is a LookupSecret.
- acp-docker-app-e2e.mts: flip the assertion StaticSecret -> LookupSecret; drop
  the stale getSecretValues reference.
- Both: fix a polling bug where "idle" (the transient pre-run state) was treated
  as terminal, so the loop bailed before the agent ran and read an empty reply.
  Terminal is now {finished, error, stuck, stopped}.

Correct the stale StaticSecret doc comments in constants/acp-providers.ts
(reserved is now an onboarding/validation marker, not a wire distinction).

Re-validated in-container against agent-server:1.25.0-python: Codex and Claude
pass end-to-end on both harnesses (LookupSecret resolves off-loop, no deadlock,
even with leftover cross-provider file-secrets present). Gemini's credential
path is proven (vertex-ai auth reached) but the turn is blocked by gemini-cli
0.45.x ignoring the requested acp_model and running gemini-3-flash — an SDK
model-selection concern tracked in software-agent-sdk#3532, not a Canvas bug;
the docs/e2e notes are corrected accordingly.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(acp): improve credential hint text with fetch commands

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(acp): show provider credentials in Settings → Agent

Adds a Credentials section to /settings/agent when an ACP provider is
selected, so users can set or rotate tokens/keys after onboarding without
hunting through Settings → Secrets. Mirrors the onboarding fields exactly
(same hints, same already-saved placeholders, Optional tag on multiline
fields) with its own Save button that writes directly to the secret store.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(acp): drop the agent_context.secrets mirror — request.secrets is the sole channel

The mirror's justification ("ACPAgent's spawn-time env loop reads from
agent_context.secrets, not the registry") predates the pinned minimum
agent-server: 1.25.0 already injects the ACP spawn env from
secret_registry, seeded from request.secrets (sdk#3299/#3464), and
sdk#3528 removes the agent_context drain entirely. Keeping the mirror
preserved a second, dead credential channel — the exact coupling
agent-canvas#1039 is eliminating.

Canvas now sends every credential in top-level request.secrets only.
Tests inverted to pin the single-channel contract; adapter/type
comments updated to match.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(acp): non-flash Gemini default, shared credential form, review cleanups

- ACP_VERTEX_SAFE_MODEL → gemini-2.5-pro: gemini-cli 0.45.x re-resolves any
  *-flash id at generation time to its current default flash (sdk#3532), so a
  flash pin is never honored; docs + e2e defaults updated to match
- extract AcpSecretField + useSaveAcpSecrets and move AcpCredentialsSection
  to components/ — onboarding and Settings → Agent share one field renderer
  and one save flow (incl. the orphaned-file-credential warning on cloud)
- a required credentials step is only satisfied by an actual credential (a
  masked `secret` field) — a base URL or GCP scalar alone no longer unblocks
- warn inline when CLAUDE_CODE_OAUTH_TOKEN and ANTHROPIC_BASE_URL are both
  set (typed or saved) — the pair silently breaks bearer auth
- drop the near-dead `reserved` field flag; collapse the leftover two-block
  secrets scaffolding in buildStartConversationRequest
- sync 14 stale locales on the OAuth/file-blob hints; fix issue refs
  (TODO #1019→#1014 — #1019 is closed; OpenHands#1016→agent-canvas#1016)
- tests: settings credentials-section coverage, non-flash pin, conflict
  matrix, tightened-gate cases

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(acp): unify default-model surfaces + dedupe credential forms and e2e harness

Review-pass cleanups:

- Route ALL three default-model surfaces (onboarding diff builder,
  Settings -> Agent seeding, start-request null fallback, + chat-input
  display) through getAcpPreferredDefaultModel, so the Vertex-safe
  Gemini override can't diverge between surfaces. New regression tests
  pin the diff-builder and start-request fallbacks to it.
- Extract useAcpCredentialForm + AcpConflictWarnings: the onboarding
  step and the Settings credentials section now share the values state,
  existing-secret lookups, conflict pairs, and save flow.
- Extract tests/e2e/live-acp/harness.mts: provider plans, host
  credential collectors, and HTTP/poll helpers shared by both live
  scripts (a model default can no longer drift between them).
- Restore the TODO(#1019) retag (accidentally reverted to the
  self-referencing #1014 in the last cleanup commit); same fix in
  docs/ACP_AGENTS.md.
- Drop the tautological ACP_VERTEX_SAFE_MODEL literal assertion, fix a
  dead key-ternary in getAcpProviderSecrets, TODO(#1016) on the
  cloud file-credential capability check, and document that baked .env
  creds don't satisfy the onboarding login probe.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: address PR review feedback (#1102)

- Restore package-lock.json to main — the npm-install churn (29 dropped
  "dev": true flags) was never meant to ship with this PR
- Note why global escapeValue:false is safe (React escapes at render;
  no translated string hits dangerouslySetInnerHTML)
- Note the non-macOS skip path in the e2e claudeOAuthToken collector

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(acp): tighten the credential gate + clarify base-URL docs (#1102 review)

- A file blob no longer satisfies the required credential step on a
  backend that can't materialise it (cloud, #1016) — the save flow
  already warned it was orphaned, so it can't be what opens the gate.
  consumesFileCredentials moves into useAcpCredentialForm so the gate
  and the save warning share one capability check.
- Next stays disabled while the login probe is still classifying a
  local backend, so a fast click can't slip past a gate about to come
  up "unauthenticated". A probe that completes as "unknown" stays
  permissive.
- Docs: a saved *_BASE_URL secret does ride along on every start
  request like any other saved secret; Canvas only never derives one
  from LLM settings. Reword the two claims that suggested otherwise.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(e2e): record 2026-06-07 re-validation — all three providers pass

Fresh 1.25.0-python container + fresh volume at the branch tip: Codex and
Claude pass both scripts; Gemini's full turn now passes too (fresh ADC +
gemini-2.5-pro + session-mode override), upgrading the previous
"blocked on model selection" row.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(acp): resume a recycled cloud ACP conversation via bootstrap prompt (#988)

A cloud ACP conversation whose sandbox was recycled (STOPPED/MISSING, e.g. the
runtime idle-stopped or hit its TTL) was a read-only dead end: the chat input
was replaced by the archived banner, and cloud createConversation never
re-provisions an existing conversation_id. The backend already supports
resuming such a conversation — re-issuing the start with the same
conversation_id rebuilds it and, for ACP, replays the durable event store as a
bootstrap prompt (OpenHands#14640) — but nothing in canvas triggered it.

Surface it:
- AppConversationStartRequest.conversation_id so the cloud start path can target
  an existing conversation.
- wakeRecycledCloudConversation(id, repoSelection): re-POST /api/v1/app-conversations
  with the conversation_id (and repo selection, so the rebuilt working dir
  matches the original cwd an ACP resume keys off).
- useWakeConversation mutation: wakes + invalidates the conversation queries so
  the active-conversation poll reconnects once the fresh sandbox is RUNNING.
- A Resume button in the archived banner for an ACP conversation whose sandbox
  is MISSING (ERROR stays read-only).

Validated e2e against a local SaaS-equivalent stack (OpenHands main app_server +
a main-built agent-server image, Docker sandboxes): create an ACP conversation,
docker rm -f the sandbox, wake → fresh sandbox + bootstrap-prompt resume, the
agent recalls prior context (codeword) and the <<RESUMED CONVERSATION>> marker
is present.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(acp): consume file-content credentials on cloud too (#988)

Cloud now materialises reserved file-content credentials (Codex auth.json,
Gemini Vertex SA) from the per-user encrypted secret store via
agent_context.secrets at conversation start (the cloud backend pins an SDK that
materialises reserved file secrets), so a pasted blob is consumable on every
supported backend — not just local. Drop the local-only gate on
consumesFileCredentials: a Codex/Gemini file blob now satisfies the onboarding
credential gate on cloud and saving it toasts success instead of the
orphaned-credential warning.

Folds the remaining cloud-enablement piece in from the native-resume canvas
branch (the wake/bootstrap-resume path landed separately); native session/load
is a backend-only concern (SDK + OpenHands), so canvas needs nothing further.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Debug Agent <debug@example.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 12:14:14 +02:00
..

Live ACP-in-Docker e2e

Proves the containerized ACP credential path through Canvas's own code — each credential is saved to the agent-server's secret store (as onboarding does) and buildStartConversationRequest references it as a LookupSecret the server resolves at spawn time — against a real agent-server container, with real provider API calls. It's the "it actually works" companion to the unit tests in __tests__/api/agent-server-adapter.test.ts — those assert the request shape; this asserts a real agent reply.

Requires agent-server:1.25.0-python or newer (software-agent-sdk#3510): the ACP credentials ride as loopback LookupSecrets, and only #3510 resolves them off the event loop. An older image deadlocks the first turn.

It is not part of npm test (it lives under tests/, which Vitest excludes, and needs a running container + real host credentials).

Run it

# 1. Agent-server container with the canvas_ui tool mounted (as the dev stack does).
#    Minimum 1.25.0-python (software-agent-sdk#3510); override for a newer build.
docker run -d --name oh-acp -p 8010:8000 \
  -v oh-acp-data:/workspace \
  -v "$(pwd)/tools:/canvas-tools:ro" -e OH_EXTRA_PYTHON_PATH=/canvas-tools \
  ghcr.io/openhands/agent-server:1.25.0-python

# 2. Run the e2e (all providers, or a subset).
npx vite-node -c tests/e2e/live-acp/vite-node.config.mts \
  tests/e2e/live-acp/acp-docker-e2e.mts -- codex claude gemini

# 3. Tear down (holds real creds).
docker rm -f oh-acp

Credentials are read from the host and never printed: Codex ~/.codex/auth.json, the Claude Code OAuth token from the macOS keychain, and the gcloud ADC for Gemini Vertex (gcloud auth application-default login first). A provider whose creds aren't present is skipped.

The provider plans (models, credential collectors) and the HTTP/poll helpers are shared between both scripts via harness.mts — change a model default or credential knob there, not per script.

Last validated result (agent-server 1.25.0-python, unified LookupSecret path)

Re-validated 2026-06-07 against ghcr.io/openhands/agent-server:1.25.0-python (the first release with software-agent-sdk#3510), on a fresh volume — every credential seeded from the secret store, no leftover state. Each credential rides as a loopback LookupSecret; the logs confirm the agent-server resolved it (GET /api/settings/secrets/<name> 200) during ACP cold-start — no deadlock, no "Failed to start ACP server: timed out", which is exactly what #3510 fixes. Codex and Claude also passed the app-orchestrator script (acp-docker-app-e2e.mts).

Provider Result Evidence (agent-server logs)
Codex ✅ real reply ACPOK-CODEX (both scripts) Materialised ACP file-secret 'CODEX_AUTH_JSON' -> …/acp/codex/auth.json; codex-acp 0.15.0; Authenticating with ACP method: chatgpt
Claude Code ✅ real reply ACPOK-CLAUDE (both scripts) claude-agent-acp 0.30.0; CLAUDE_CODE_OAUTH_TOKEN env path (no ANTHROPIC_BASE_URL)
Gemini CLI ✅ real reply ACPOK-GEMINI¹ Materialised ACP file-secret 'GOOGLE_APPLICATION_CREDENTIALS_JSON' -> …/acp/gemini-cli/gcloud-credentials.json; gemini-cli 0.45.1; Authenticating with ACP method: vertex-ai → real Vertex inference on gemini-2.5-pro

¹ Gemini prerequisites. The full turn passes with: a fresh host ADC (gcloud auth application-default login — a stale one fails as invalid_rapt, a credential problem, not a Canvas one), the non-flash gemini-2.5-pro model (gemini-cli 0.45.x re-resolves any *-flash id at generation time to its current default flash, which 404s on projects that don't serve it — software-agent-sdk#3532; this is why Canvas preselects gemini-2.5-pro), and ACP_E2E_GEMINI_SESSION_MODE=default to clear the separate gemini-cli ≥0.43 set_session_mode("yolo") headless-init blocker.

Knobs

  • ACP_E2E_BASE_URL (default http://localhost:8010)
  • ACP_E2E_CODEX_MODEL / ACP_E2E_CLAUDE_MODEL / ACP_E2E_GEMINI_MODEL
  • ACP_E2E_GEMINI_SESSION_MODE (set default to bypass the SDK yolo blocker)
  • GOOGLE_CLOUD_PROJECT / GOOGLE_CLOUD_LOCATION (else read from gcloud / us-central1)