Commit Graph
116 Commits
Author SHA1 Message Date
dc99e98615 fix: preserve backend scope in conversation links (#16091)
Co-authored-by: neubig <398875+neubig@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: neubig <neubig@users.noreply.github.com>
Co-authored-by: allhands-bot <allhands-bot@users.noreply.github.com>
2026-08-14 20:20:19 -04:00
Juan Pedro Michelini Jorgeandopenhands fa21e01a6c fix(onboarding): preselect OpenHands LLM provider after picking OpenHands agent (#16531)
Co-authored-by: openhands <openhands@all-hands.dev>
2026-08-14 11:05:27 -03:00
83ba34cce6 chore: bump SDK deps (software-agent-sdk 1.42.1, automation 1.7.1, typescript-client 1.38.0) (#16554)
Co-authored-by: neubig <neubig@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>
2026-08-12 21:16:34 -04:00
bf2e37dcad fix: preserve MCP credentials during Canvas mutations (#16144)
Co-authored-by: neubig <neubig@users.noreply.github.com>
Co-authored-by: Rohit Malhotra <rohitvinodmalhotra@gmail.com>
Co-authored-by: openhands <openhands@all-hands.dev>
2026-08-04 22:48:42 -04:00
Hiep Le 565c10daa2 chore: bump @openhands/extensions to 0.16.0 (#16321) 2026-08-05 02:32:41 +07:00
Rishav Naskarandhieptl 947d9a0358 fix(backends): pin backend identity on sidebar conversation links (#16243)
Co-authored-by: hieptl <hieptl.developer@gmail.com>
2026-08-04 13:15:32 +07:00
Hiep Le 5d5a0648db feat: mock the setup contract from the published fixtures (#16221) 2026-08-01 00:16:05 +07:00
932edbf812 feat(backends): compact Cloud vs Agent-server add-backend chooser (#16211)
Co-authored-by: Devin <devinvinson@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Graham Neubig <neubig@gmail.com>
Co-authored-by: Graham Neubig <neubig@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>
2026-07-31 02:35:47 -04:00
Rohit Malhotraandopenhands 8dfa1d510c fix: Revive local telemetry consent banner (#16183)
Co-authored-by: openhands <openhands@all-hands.dev>
2026-07-29 17:58:08 -04:00
simonrosenbergandopenhands f4612aaafb fix: remove ACP settings access gating (#1908)
* Remove ACP settings access gating

* Test ACP settings access in mobile hub

* test: stabilize mock-LLM skills navigation

Co-authored-by: openhands <openhands@all-hands.dev>

* test: retry mock-LLM skills navigation

Co-authored-by: openhands <openhands@all-hands.dev>

* test: avoid redundant model switch reload

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-07-25 08:39:56 +00:00
Hiep Le 7b1e07d6f8 feat: add windows desktop installer build, docs, and win32 fixes (#1897)
* feat: add Windows desktop installer build, docs, and win32 fixes

* fix: failing tests

* refactor: update the code based on feedback
2026-07-24 06:25:07 +00:00
Rohit Malhotraandopenhands efd20f7d56 fix: suppress telemetry consent prompt in Cloud Canvas (#1848)
* Suppress telemetry consent modal in Cloud Canvas

Treat same-origin locked Cloud cookie deployments as already consented for Canvas library telemetry so the modal does not flicker before the main app login flow.

Co-authored-by: openhands <openhands@all-hands.dev>

* Stabilize mock LLM settings tests after ACP

Reset the default agent profile back to OpenHands through the agent-profile API before LLM-profile setup paths that need /settings/llm. This prevents an ACP profile left by the previous serial spec from redirecting later settings tests to /settings/agents.

Co-authored-by: openhands <openhands@all-hands.dev>

* Stabilize files tab mock E2E git setup

Ensure the attached-workspace conversation has both an origin remote and a real HEAD commit before asserting that the Files tab defaults to diff view. The diff default now intentionally depends on both attached source metadata and an available commit base.

Co-authored-by: openhands <openhands@all-hands.dev>

* Fix desktop right panel toggle visibility

Update the desktop right-panel toggle to set both the user-toggled flag and the visible state. The missing visibility update left the panel visually closed in mock E2E while off-screen tab controls remained mounted.

Co-authored-by: openhands <openhands@all-hands.dev>

* Make files tab mock E2E open the panel explicitly

Wait for the desktop panel toggle to report an open state before interacting with Files tab controls, then verify the Diff segment can be selected for the attached-workspace conversation.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-07-20 15:33:24 +00:00
Hiep LeandGraham Neubig a8edaa7cf4 feat: show actionable connection health on installed server cards (#1833)
* feat: show actionable connection health on installed server cards

* fix: failing tests

---------

Co-authored-by: Graham Neubig <neubig@gmail.com>
2026-07-20 16:00:02 +02:00
2b7ceea667 refactor: define canvas UI as an SDK client tool (#1797)
* refactor: define canvas UI as an SDK client tool

Send a JSON-defined canvas_ui_client tool on new, profile-based, and resumed conversation requests while retaining the legacy Python registration for persisted conversations. Normalize the new SDK event kinds to the existing Canvas UI rendering.

Co-authored-by: smolpaws <engel@enyst.org>

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: omit canvas client tool from ACP launches

* refactor: rename canvas client tool

Use the semantic canvas_ui_control name and contain the SDK-generated action discriminator behind exported constants.

Co-authored-by: Engel Nyst <engel.nyst@gmail.com>

---------

Co-authored-by: Engel Nyst <engel.nyst@gmail.com>
Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: Debug Agent <157206163+simonrosenberg@users.noreply.github.com>
2026-07-15 09:40:57 +00:00
Hiep Le 133d8d9494 test: deflake mock-LLM Docker CSS-scope regression test (#1643) 2026-07-09 22:32:02 +07:00
simonrosenbergandClaude Opus 4.8 54d718ad4e feat(agent-profiles): Agent Profiles — Settings → Agent as the profile library (local + cloud) (#1571)
* feat(agent-profiles): minimal local Agent Profiles library reusing the Agent settings form

Adds a Settings → Agent profiles library (local backends only) that mirrors
the LLM-profiles UX: a list of named profiles with a create/edit view that
reuses the existing Agent settings form as the editor — you just add a name
(and, for OpenHands agents, pick an LLM profile).

Deliberately minimal vs the full Phase-4 UX: no chat-input picker, no live
switch, no Settings information-architecture rework. Condenser / verification /
MCP stay global, exactly as on main.

- Data layer: AgentProfilesService + list/save/delete/rename/activate hooks
  wrapping the ts-client AgentProfilesClient (endpoints shipped in
  agent-server v1.29.0).
- Editor: AgentSettingsScreen gains an opt-in `embedded` mode (hides its
  header + global Save, seeds from an override, and reports state via a save
  control) — mirroring how LlmSettingsScreen is embedded in the LLM-profiles
  view. The global Agent settings page is unchanged.
- Library: AgentProfilesLocalView (list/create/edit) + manager/body/row/menu +
  delete modal, at the additive route /settings/agents, gated to local
  backends (cloud has no /api/agent-profiles surface yet, epic #3730).
- Maps the form to AgentProfileSaveInput: OpenHands requires an llm_profile_ref
  (via a picker); ACP stores acp_server/acp_model and the command as a shell
  string. Validated end-to-end against a real agent-server.

Part of OpenHands/software-agent-sdk#3713 (Phase 4). An alternative to the
larger #1550.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(agent-profiles): add chat-input agent-profile picker + live in-conversation switch

Adds the full chat integration for Agent Profiles (epic #3713, #3727), keeping
the simplified library/editor from the previous commit:

- New-conversation picker (home): an agent-profile toggle replaces the LLM-
  profile toggle. Selecting activates the profile so the next conversation
  launches from it; conversations start via `agent_profile_id` (resolved
  server-side) instead of an inline agent_settings dump.
- Mid-conversation switch, capability-gated by the running agent:
  - OpenHands conversation → live LLM-profile switch (`/switch_profile`).
  - ACP conversation → live model switch (`set_session_model`, existing
    ChatInputModel).
  - Home / cloud fall back to the agent-profile picker / model picker.
- Threads `agent_profile_id` through the conversation-start path
  (buildStartConversationRequest: agent_profile_id XOR agent_settings; skip the
  ACP tag / encrypted-settings / subscription check on the profile path) and
  reads the server's `launched_agent_profile` provenance to mark the current
  profile without settings-matching.
- Replaces the old SwitchProfileButton/context-menu with the new pickers.

Validated end-to-end against a real agent-server (SDK main): starting a
conversation with `agent_profile_id` returns 201 and stamps
`launched_agent_profile { agent_profile_id, revision }`.

Ported from #1550's chat implementation. Gates green: typecheck, eslint,
prettier, i18n (15 langs), vitest (3496 passed).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(agent-profiles): extract + unit-test buildAgentProfileFields mapping

Addresses the code-review feedback that the profile-fields builder — the ACP
"built-in default command → null vs verbatim shell string" branch plus the
schema-driven tool_concurrency_limit coercion — was the most novel logic in the
PR yet had no automated coverage (every test mocked the embedded form away).

- Extracts the closure into a pure exported `buildAgentProfileFields()` in
  agent-settings.tsx; the embedded control now just snapshots state into it.
- Adds 8 unit tests locking the round-trip: ACP built-in-default → null, custom
  command → shell string, custom preset, blank-model → null, OpenHands
  enable_sub_agents passthrough, concurrency coercion (valid / empty / throws).
- Clarifies the service header (client ships in ts-client 1.28.0; the server
  endpoints it targets shipped in agent-server v1.29.0) per the version-doc nit.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(agent-profiles): address review feedback + fix e2e regression

- Fix mock-LLM E2E regression: the profile-identity spec still targeted the
  removed `switch-profile-button`; point it at the new `chat-input-llm-profile`
  picker (mirrors #1550's e2e update).
- Use the `useRenameAgentProfile` hook in the editor instead of calling the
  service directly (the hook was otherwise dead code; now the rename gets list
  invalidation for free).
- Drop the unreachable in-conversation branch from the home AgentProfile picker:
  the picker only renders on home (a running conversation shows the LLM/model
  picker), so `useChatInputProfileState` is now home-only (activate as launch
  default), and the "start new with profile" hint + its
  CHAT$START_NEW_WITH_PROFILE_HINT key (15 langs) are removed.
- Update the two affected tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(agent-profiles): correctness fixes from #1571 code review

- switch-llm-profile: run the inline "Switched to" message, #1082 metadata
  persist, and error reporting in mutation-level callbacks so they survive the
  switcher menu unmounting on select
- agent-server-adapter: derive acp_server from agent.acp_server when the
  acpserver tag is absent, so a profile-launched ACP conversation keeps its
  model picker and provider chip
- use-create-conversation: await the LLM-profile list before the
  dangling-llm_profile_ref launch guard so a mid-load send can't launch blind
- chat-input pickers: read switch/activate pending state via useIsMutating so
  the pill button actually disables during an in-flight switch
- use-activate-agent-profile: surface activation errors (drop disableToast) and
  optimistically flip active_agent_profile_id with rollback
- chat-input-actions: fall back to the LLM picker on the home page when the
  backend has no /api/agent-profiles surface

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(agent-profiles): pass embedded props to reused Agent settings form

The profile editor reused AgentSettingsScreen via the route module's default
export. React Router's Vite plugin wraps a route default with
withComponentProps, which invokes it with route props and drops any props a
parent passes — so `embedded`/`onSaveControlChange` never reached it,
`saveControl` stayed null, and the Save button was permanently disabled
(couldn't create or edit a profile at all).

Split the route into a named `AgentSettingsScreen` export (the reusable
component embedded consumers import) plus a thin default `AgentSettingsRoute`
wrapper, mirroring `LlmSettingsRoute`. The local-view now imports the named
export. Updated the unit-test mock to provide the named export (the old mock
only stubbed `default`, which is exactly what masked this at unit level).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(agent-profiles): un-gate Agent Profiles on cloud backends

The cloud enterprise app-server now exposes the same /api/agent-profiles
contract as the local agent-server (OpenHands #15060, epic #3730), so lift
the local-only gating and route cloud calls through the cloud proxy.

Transport:
- cloud/agent-profiles-service.api.ts: CRUD via callCloudProxy (bearer +
  X-Org-Id) against the identical /api/agent-profiles paths; org resolved
  server-side from the session, so no {org_id} segment.
- cloud/org-profiles-service.api.ts: list org LLM profiles at
  /api/organizations/{org_id}/profiles so the editor's llm_profile_ref
  picker works on cloud. Only listing is cloud-routed.
- AgentProfilesService + ProfilesService.listProfiles branch to the cloud
  transport when the active backend is cloud (mirrors SettingsService).

Surfaces un-gated:
- Settings → Agent profiles nav item + route (no more redirect to /settings/agent).
- Home chat-input agent-profile picker (fetch + pickerKind) on cloud.
- Launch-from-profile: cloud AppConversationStartRequest now carries
  agent_profile_id (added to the type + the cloud create request), which the
  backend resolves and stamps as launched_agent_profile.

In-conversation live switch on cloud is intentionally left on the model
picker for now: the cloud backend has no per-conversation profile-switch
endpoint yet and org LLM-profile detail masks the api_key, so a client-side
switch isn't possible — tracked as a follow-up for full parity.

Tests updated for the new nav behavior (agent-profiles shown on both).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(agent-profiles): update useAgentProfiles docstring for cloud support

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(agent-profiles): clarify cloud in-conversation switch is intentionally local-only

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(agent-profiles): preserve acp_server through the wire normalizer

An ACP conversation launched from an agent profile (agent_profile_id) showed
a generic chip and an empty in-conversation model picker: the provider
identity never reached the UI.

Root cause: #1571 taught the conversation adapter to source acp_server from
`agent.acp_server` (SDK #3692) when the `acpserver` tag is absent — which is
exactly the profile-launch case, since that path doesn't stamp the tag. But
`normalizeAgent` (the wire parser feeding the adapter) projected only
`{kind, acp_model, llm}` and dropped `acp_server`, so the adapter's fallback
always saw undefined → acp_server null → no ACP provider → generic chip + no
model list.

Add `acp_server` to the normalizeAgent projection (the type already declared
it). Regression test covers the no-tag / agent-sourced path.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(agent-profiles): make Settings → Agent the profile library

Collapse the two Settings sections ("Agent" global form + "Agent profiles"
library) into a single "Agent" entry that IS the Agent Profile library: it
lists the user's profiles and its create/edit view is the reused Agent
settings form plus a name (the embedded AgentSettingsScreen). The active
profile is the current agent.

- settings-nav: one "Agent" item → /settings/agents (the library).
- /settings/agent redirects to /settings/agents; default settings path +
  ACP route-guard target updated accordingly.

Also derive the ACP-enabled state from the ACTIVE AGENT PROFILE rather than
settings.agent_settings.agent_kind. Activate is pointer-only and never writes
agent_settings, so the global settings are stale when an ACP profile is
active; the nav-disable, home ACP context, useLlmConfigured, and the ACP
route guard now read the active profile (new useActiveAgentProfile hook) and
fall back to settings only while the profile list is loading.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(agent-profiles): gate the LLM-setup banner on the active agent profile's LLM

useLlmConfigured decided "is the LLM ready" from the standalone active LLM
profile, but conversations now launch from the active AGENT profile. For an
OpenHands profile the relevant LLM is the one it references via
llm_profile_ref — not whichever LLM profile happens to be "active". So the
"Your LLM isn't set up" banner could be wrong in both directions (e.g. the
active LLM profile has a key but the agent profile references a keyless one).

Resolve the LLM profile to check from the active agent profile's
llm_profile_ref (openhands), falling back to the active LLM profile only when
there's no ref yet. ACP agent profiles stay always-configured (subprocess
owns its LLM). New unit test covers the discriminating case.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(agent-profiles): relabel LLM profile "Active" → "Default"

The active LLM profile no longer drives new conversations (the active AGENT
profile does) — it's just the default llm_profile_ref seeded into new agent
profiles. Relabel the LLM-profile badge "Active" → "Default" and the row
action "Set as active" → "Set as default" to stop implying it launches
conversations. New i18n keys (SETTINGS$PROFILE_DEFAULT / _SET_DEFAULT, 15
langs). The agent-profile "Active" badge is unchanged — that one IS active.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(onboarding): land the user's choice on the active agent profile

Onboarding configured global agent_settings + an LLM profile (OpenHands) or
ACP secrets, but never touched an AGENT profile — so the active agent profile
stayed the seeded `default` (openhands → ref `default`), disconnected from what
onboarding set up. Result: an OpenHands user who entered a key still hit "LLM
isn't set up" (the active agent profile referenced a keyless profile), and ACP
users never got an ACP agent profile at all.

Add useApplyOnboardingAgentProfile: upsert + activate the well-known `default`
agent profile from the onboarding choice. The OpenHands LLM step now points it
at the LLM profile it just created; the ACP secrets step makes it an ACP
profile for the chosen provider (opus[1m]/valid default, no LLM key needed).

Verified e2e: OpenHands onboarding → default agent profile refs the configured
LLM + banner clears; Claude Code onboarding → default agent profile is
acp/claude-code, active, no LLM required.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: drop unused eslint-disable in onboarding agent-profile hook

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(agent-profiles): gate mutate controls for cloud view-only members

Reuse #1532's org-permission gating for the Agent Profiles UI. Agent
profiles are org-scoped on cloud (bearer + X-Org-Id, edit_org_settings),
so a cloud member previously saw Add/Edit/Delete/Set-active controls that
would 403 server-side — the same flash-then-403 problem #1532 fixed for
LLM profiles.

- Generalize useCanManageLlmProfiles -> useCanManageOrgProfiles (it reads
  the generic edit_org_settings permission; local users always true).
- Thread canManage through AgentProfilesManager -> Body -> Row, mirroring
  LlmProfilesManager: hide the Add button and the row actions menu for
  view-only members.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: fix stale comments surfaced by PR review

Comment-only. No behavior change.

- chat-input-actions.tsx: the pickerKind summary claimed "cloud → model
  picker (cloud has no profile surface)", contradicting the code, which
  uses the AgentProfile picker on cloud home too (#15060). Rewrite to
  match the actual cases; trim the duplicated render-site recap.
- acp-route-guard.ts / settings-nav.tsx / settings.tsx: the ACP redirect
  target moved to /settings/agents (plural) in this PR, but three
  docstrings still said /settings/agent. Update them.

* test(mock-llm-e2e): wire the active agent profile to the mock LLM

Fixes the mock-LLM e2e regression where the home composer stayed blocked
(submit disabled / launcher never ready) so conversation-launching specs
timed out. Conversations now launch from the active AGENT profile (#1571),
and `useLlmConfigured` follows that profile's `llm_profile_ref` — not the
active LLM profile. The specs seed `openhands-onboarded` and configure an
LLM profile the old way, so the seeded "default" agent profile still
pointed at a keyless LLM and the composer never unblocked.

Mirror what onboarding does for a real user: after activating the mock LLM
profile, upsert + activate the "default" agent profile referencing it. Add
a shared `ensureMockLLMAgentProfile` helper (called from ensureMockLLMProfile
and from the conversation spec, which sets up inline).

Verified locally: the full mock-llm-conversation spec passes 4/4 (real
conversation runs against the mock LLM) with this change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(agent-profiles): address PR #1571 review feedback

Human review (VascoSch92):
- useLlmConfigured: fall back to the active LLM profile when the active agent
  profile's llm_profile_ref is stale/absent, mirroring the launch-time fallback
  in useCreateConversation. Without this the two contradicted each other: launch
  succeeded via the fallback but the hook reported unconfigured and spuriously
  disabled the composer + banner (even inside a running conversation). Adds a
  regression test for the stale-ref scenario.
- Drop the dead launched_profile plumbing (wire parse + types + adapter map):
  it had zero readers (the home picker keys off active_agent_profile_id and the
  in-conversation picker is LLM/model by design), so the "Consumed by the
  picker" comments were misleading.
- Point the remaining /settings/agent links at /settings/agents (ACP model
  context, chat-input model state, chat error re-auth, command menu) so the
  route rename doesn't cost an extra redirect hop.

/codereview-roasted:
- Extract the triple-nested pickerKind ternary into a pure, unit-tested
  resolvePickerKind() helper.
- Document why cloud OpenHands onboarding intentionally does not repoint the
  active agent profile (persistAsProfile is local-only; cloud resolves the
  agent-profile/LLM wiring server-side).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(agent-profiles): scope the profile-launch enrichment gap (#1571 review)

- Document at buildStartConversationRequest that the profile path relies on the
  server/SDK to restore exec tools + public skills (software-agent-sdk#3967),
  and that canvas_ui + the RUNTIME_SERVICES suffix are intentionally canvas-only.
- Point the createConversation positional-args TODO at the tracked issue (#1587).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(agent-profiles): preserve unmodeled fields on edit-save + await profiles at launch

Two fixes from the #1571 review:

Edit-save wiped every profile field the minimal editor doesn't model
(condenser, verification, system_message_suffix, skill/MCP refs, embedded
skills, ACP session mode/timeout): the save endpoint is a whole-profile
overwrite, and the editor posted only its own fields. The save payload now
spreads the stored profile under the edited fields via a pure, kind-aware
mergeAgentProfileSaveInput — a kind switch stays a clean variant replacement
(the server's extra="forbid" union rejects mongrel payloads), and
server-managed identity (id/name/revision) is stripped. The edit fetch now
uses X-Expose-Secrets: encrypted so any skills[].mcp_tools values round-trip
as Fernet tokens instead of persisting the mask literally (same pattern as
the LLM-profile editor).

Launch raced the agent-profiles query: useCreateConversation read the hook's
maybe-unresolved data, so a send fired before the list loaded fell through to
the stale global agent_settings path — which activation (pointer-only) never
updates — and silently launched the wrong agent. The launch now awaits the
list via queryClient.ensureQueryData on the shared query key (mirroring the
LLM-profile ref validation below it), with retry: false so backends without
the surface degrade to the legacy launch immediately. The dangling-llm-ref
downgrade also logs a console.warn so the silent fallback is diagnosable.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(acp): drive ACP spec through the Agent Profile editor, not the retired route

Settings → Agent is now the Agent Profile library (#1571): the standalone
/settings/agent form redirects to /settings/agents, whose editor reuses the
same embedded agent-settings-screen form. The ACP mock-llm spec and the
resetToOpenHandsAgentViaUI cleanup helper still navigated the old route and
waited on the retired agent-save-button, so they timed out — and the cleanup
helper's failure (swallowed by afterAll's try/catch) left the "default" agent
profile stuck in ACP mode, poisoning downstream specs that share the backend.

- Add openAgentProfileEditor(page, name): navigate /settings/agents, open the
  named profile's editor via its row action menu (row located by the
  profile-name span[title], mirroring activateProfileViaUI).
- Rewrite resetToOpenHandsAgentViaUI to drive the new editor (switch kind →
  OpenHands, pick an LLM profile, save via save-agent-profile-btn).
- Point ACP spec steps 1 & 2 at the editor; swap agent-save-button →
  save-agent-profile-btn.
- Verify step 1 against GET /api/agent-profiles/default (the new source of
  truth) instead of legacy /api/settings; acp_command is a shell string there
  (ts-client AgentProfile.acp_command: string | null), not a token array.

* fix(build): keep the styling core in one chunk to avoid a tv() init-order crash

This PR's new imports grew/shifted the auto-split `vendor` chunk enough that
Rolldown's size-based splitter (`maxSize`) sliced the styling core apart —
separating a HeroUI component's top-level `tv()` recipe from tailwind-variants'
core within the emitted init order. The recipe then evaluated before
tailwind-variants initialized, throwing `TypeError: s is not a function` at
module load. React Router reported "Error loading route module root-layout,
reloading page", looped, and rendered a blank page — deterministically crashing
the whole app and failing 13 mock-llm-e2e specs (npm + docker) that load the
shell.

Give the styling core (@heroui/react + tailwind-variants + tailwind-merge +
clsx) its own group that is never size-split, so it initializes as a coherent
unit before any consumer's top-level `tv()` call. Verified locally: the home
route renders (was a blank page) with zero console errors.

* test(mock-llm): make LLM-profile setup idempotent and fix stale ACP launch assertion

With the crash fixed, the app renders and a second class of failure surfaced:
specs that call `ensureMockLLMProfile` after the first one deadlocked on a stuck
"Delete Profile" modal, and the ACP spec's payload assertion checked the old
launch shape.

- ensureMockLLMProfile: create the mock LLM profile only when absent instead of
  delete-then-recreate. Once the active agent profile references it (wired right
  after, via ensureMockLLMAgentProfile — #1571), the LLMProfile FK guard rejects
  deletion; the delete-confirm modal then silently stays open and its backdrop
  blocks every later click (`add-llm-profile` timed out across files, home,
  automations, mcp, model-switch, preset-automation). The mock config is
  deterministic, so reusing an existing same-named profile is correct.
  deleteProfileIfExists is unchanged — it still works for the non-referenced
  profiles that other specs delete.
- mock-llm-acp-agent step 3: conversations now launch from the active
  AgentProfile (#1571), so the POST /api/conversations payload carries
  `agent_profile_id` and omits `agent_settings` (mutually exclusive, per
  agent-server-adapter). Assert that shape instead of the retired
  `agent_settings.agent_kind`; the ACP reply-token check still proves the ACP
  agent ran.

* ci: degrade gracefully when the linked SDK reference isn't a PR

"Resolve linked SDK PR" (mock-llm-docker-e2e.yml) greps the PR description
for OpenHands/software-agent-sdk#NNNN or .../pull/NNNN and tries to build
against that PR's branch. GitHub's "#NNNN" shorthand looks identical for
issues and PRs, so a description that links a tracking issue (e.g. #3713)
matches the same regex — and /pulls/{number} 404s for an issue number,
failing the whole job under `bash -e` instead of falling back to the
released SDK version like the "no match" branch already does.

Treat a failed PR lookup the same as "no linked PR found": log and exit 0,
leaving git_ref unset so the job falls through to the released version.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(test): make ensureMockLLMProfile/AgentProfile converge, not skip

Two real e2e failures traced to test-helper bugs surfaced only once #3968
(SDK 1.32.0) let profile-launched conversations actually run:

- ensureMockLLMProfile: the earlier idempotent-reuse fix (deadlock guard
  against the LLMProfile FK constraint) skipped writing the profile's
  config entirely whenever a same-named profile already existed —
  correct for repeat calls with the SAME config, but silently ignored a
  DIFFERENT one. mock-llm-image-upload requests a vision-capable model
  ("openai/gpt-4o") to get past the mock LLM's default; when an earlier
  spec in the same CI run had already created "mock-llm" with the
  default model, the override never applied and the agent replied "the
  currently selected model does not support image understanding" —
  confirmed via the CI screenshot. Fixed by editing the existing profile
  in place (via the LLM settings UI's Edit flow, never deleting it) so
  every call converges on the requested model/apiKey/baseUrl regardless
  of what an earlier test left behind.

- ensureMockLLMAgentProfile: OpenHandsAgentProfile.skill_refs defaults to
  `[]` (none discovered) when omitted from the save payload. Workspace-
  scoped project skills are discovered independently of this and keep
  working, but a profile-launched conversation's agent never sees any
  public/preset skill (e.g. an installed automation's bundled skill)
  without an explicit skill_refs. Set it to `null` (all discovered),
  matching what a real onboarding-seeded profile effectively gets.

mock-llm-model-switch step 2's post-switch reply timeout is left
unaddressed: its trajectory hard-codes one padding turn for "the
agent-server's internal condenser/skill-analysis call before the main
loop" (a documented, historically-fragile assumption per the test's own
comment) — plausibly now off by one now that #3968 lets the agent make
additional real tool-use calls around a /model switch. Fixing this
requires an empirical trajectory-turn count from a real 1.32.0
conversation trace, which needs a CI cycle to observe correctly rather
than guessing blind.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* revert(test): drop skill_refs=null from ensureMockLLMAgentProfile

CI showed this regressed mock-llm-skills.spec.ts (project skill in
workspace/.agents/skills/), which passed before this change: fixed the
narrow preset-automation slash-command skill-activation case at the
cost of breaking a more fundamental, previously-solid #3968 validation
— a net-negative trade, not a clean win.

The shared "default" agent profile backs every spec in the suite;
widening its skill_refs to "all discovered" has global blast radius
across unrelated tests, evidently including some interaction with
project-skill discovery/activation tracking that isn't understood yet.
A fix for preset-automation's specific skill needs to be scoped to that
one profile/test, not applied to the profile every other spec shares.

Keeps the ensureMockLLMProfile edit-in-place fix (proven, isolated,
fixes mock-llm-image-upload with no observed side effects).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(agent-profiles): default new profiles' skill_refs to "all discovered"

OpenHandsAgentProfile.skill_refs defaults to `[]` (none) server-side when
omitted from a save payload. Neither onboarding's profile seed nor the
Settings "Add Agent Profile" editor exposes a skill_refs control, so every
newly-created profile silently gets zero public/user/project skills — a
profile-launched conversation's agent can't activate any of them (#1571
launches conversations from the active agent profile). This is exactly
the mock-llm-preset-automation regression: a slash-command-triggered
skill never activates because the "default" test profile has no
skill_refs, matching what a real user's fresh profile would also hit.

useSaveAgentProfile is the single choke point for every profile save
(onboarding seed + Settings create/edit), so default skill_refs to `null`
("all discovered") there whenever the caller hasn't set it explicitly —
matches what users actually expect (a new agent has access to their
skills unless deliberately scoped down) and requires no SDK change. The
pinned typescript-client doesn't type skill_refs on AgentProfileSaveInput
yet (SDK/wire drift), so this reaches it via an untyped merge; `in`
checks the runtime object since mergeAgentProfileSaveInput's edit-preserve
spread can carry it at runtime despite the missing type.

Re-applies the equivalent default to ensureMockLLMAgentProfile (the e2e
test helper bypasses this hook via a raw fetch) so the test suite mirrors
real behavior.

Verified: full unit suite green (3618 passed), typecheck clean,
agent-profiles-local-view.test.tsx passes unaffected (it mocks
useSaveAgentProfile at the hook boundary, so this change is invisible to
it). Locally reproduced the fix: mock-llm-preset-automation's slash-
command skill-activation test now passes; mock-llm-image-upload
(previously fixed) still passes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(agent-profiles): self-heal LLM profile stream=true for profile-launched conversations

A profile-launched conversation (agent_profile_id) never sends
agent_settings, so PR #1474's `llm.stream = true` (buildConfiguredOpenHands
AgentSettings, agent-server-adapter.ts) never reaches it — the referenced
LLM profile's stored `stream` field (SDK default: false) is used as-is by
resolve_agent_profile/_build_openhands_settings, with no override, unlike
the legacy path.

The agent-server decides once, at conversation construction, whether to
wire the `on_token` streaming callback — based on whether any of the
agent's LLMs has stream=True at that moment — and never re-evaluates it
afterward (confirmed by reading LocalConversation.switch_llm: it swaps the
LLM but never touches _on_token). So a profile-launched conversation whose
LLM profile was never saved with stream=true gets on_token=None for its
entire lifetime. switchProfile's switch_llm call (unconditionally sending
stream: true, unchanged by this fix) then crashes the next completion with
"Streaming requires an on_token callback", since on_token can never be
(re-)wired post-construction. Confirmed via real agent-server tracebacks in
both mock-llm-e2e and mock-llm-docker-e2e CI runs.

Streaming is a pre-existing, independently-shipped feature (PR #1474) that
must not regress for legacy-launched conversations — ruling out simply
dropping switch_llm's stream:true (would silently disable streaming after
a switch for the one case that works today). And since existing users'
LLM profiles predate this fix, defaulting stream:true only at future
profile-save time (mirroring the skill_refs fix) would still crash on
their first profile-launched conversation post-deploy.

ensureLlmProfileStreams is a migration shim: at the one call site
guaranteed to run for every profile-launched conversation (already
fetching the LLM-profiles list to validate llm_profile_ref exists), check
the referenced LLM profile's full config and, if stream isn't already
true, save it with stream:true — self-healing both new and existing
profiles on first use, memoized per profile name for the session so it's
a no-op read on every subsequent launch. Mirrors the profile-duplicate
flow's exact pattern for round-tripping the encrypted secret
(getProfile(name, "encrypted") + saveProfile(..., include_secrets: true))
so the stored api_key is never clobbered. Touches neither the legacy
agent_settings path nor switch_llm — both keep working exactly as before.

Safe to delete once virtually all users are migrated, or once
resolve_agent_profile forces stream=true for OpenHands profiles upstream
(same category of fix as the skill_refs default — likely the same #3967
umbrella), whichever comes first.

Verified: full unit suite green (3620 passed, +2 new tests exercising
this exact self-heal/no-op branching), typecheck clean. Locally
reproduced the fix: mock-llm-model-switch's on_token crash no longer
occurs; preset-automation and image-upload remain passing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(agent-profiles): link the skill_refs/streaming shims to their tracking issue

References OpenHands/agent-canvas#1619 (the cleanup-tracking issue for both
workarounds) and the specific upstream SDK issues, so the removal criteria
is discoverable from the code itself, not just the PR description.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(agent-profiles): drop skill_refs/streaming migration shims (SDK #4017 landed)

software-agent-sdk#4017 (PR #4018) fixes both gaps these shims worked
around: OpenHandsAgentProfile.skill_refs now defaults to null (all
discovered) server-side, and the agent-server forces llm.stream=true
for profile-launched conversations. Both shims are now dead code.

Validated end-to-end against the SDK branch (OH_AGENT_SERVER_LOCAL_PATH)
before removing: real HTTP round-trips confirmed skill_refs defaults to
null and the launched agent's LLM streams even though the underlying LLM
profile is stored with stream=false; the full mock-llm-skills.spec.ts and
mock-llm-profile-management.spec.ts suites pass unchanged.

Removes:
- withDefaultSkillRefs (src/hooks/mutation/use-save-agent-profile.ts)
- ensureLlmProfileStreams + its two dedicated tests
  (src/hooks/mutation/use-create-conversation.ts)

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(agent-profiles): correct comments for the disabled_skills deny-list

SDK #4017 replaced the profile's skill_refs allow-list (and embedded skills)
with a disabled_skills deny-list. Canvas is already deny-list-native — the
user-level disabled_skills UI exists and the generic profile merge carries the
field automatically — so only two stale comments referencing embedded skills /
skill refs needed correcting. No functional change; the per-profile skill
picker stays out of scope for the minimal editor.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(agent-profiles): drop stale skill_refs from fixtures for the deny-list

SDK #4017 replaced the profile's skill_refs allow-list (and embedded skills)
with a disabled_skills deny-list. Update the fixtures/comments that still
referenced the removed fields (they ride untyped through `as unknown` casts /
raw POST bodies, so the generic merge round-trips them regardless):
- merge-agent-profile-save-input.test.ts + agent-profiles-local-view.test.tsx:
  skill_refs -> disabled_skills, drop embedded `skills`, schema_version 3,
  ACP fixtures drop the skill field (ACP has none). Correct the stale
  exposeSecrets/mcp_tools comment (profiles are secret-free now).
- mock-llm-helpers.ts: the omitted-field comment now describes the deny-list.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(agent-profiles): profile fixtures use the v1 baseline schema_version

SDK #4017 collapsed the pre-ship AgentProfile schema history to a clean v1
baseline (no v2/v3, no migrations). Update the two profile fixtures to
schema_version: 1 to match the shipped model.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(agent-profiles): launch the `default` profile via agent_settings; stamp the launched LLM ref

The seeded `default` agent profile is the enriched baseline that mirrors global
agent_settings, not a deliberate profile pick. Launching it via `agent_profile_id`
made the server rebuild the agent purely from the profile, dropping the canvas-only
enrichments the profile-resolution path can't carry — the `<RUNTIME_SERVICES>`
system-message suffix, the `canvas_ui` tool, and project-skill loading. Route the
well-known `default` profile through the agent_settings launch instead; named
profiles are deliberate custom configs and keep the profile path. Fixes the
mock-llm-docker-e2e automation RUNTIME_SERVICES failure.

Also from #1571 review:
- Stamp the launched OpenHands profile's `llm_profile_ref` into conversation
  metadata (not the standalone active LLM profile) so the switcher pill names the
  exact profile the conversation runs when the two differ (#1082).
- Add `retry: false` to the LLM-ref validation fetch, matching the sibling
  agent-profiles fetch, so a slow/erroring /api/profiles falls back promptly.

Hoist the well-known name to `WELL_KNOWN_DEFAULT_AGENT_PROFILE_NAME` (shared by the
launch path and onboarding). Re-onboarding intentionally overwrites `default`.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(agent-profiles): gate the LLM-setup banner on the active agent-profile load

`useLlmConfigured` derives `isAcpAgent` and the referenced LLM from the active
agent profile but omitted that query's loading state from `isLoading`. On a cold
cache an ACP agent (which needs no key) briefly read as an unconfigured OpenHands
agent, flashing the "LLM not set up" banner until the profiles query resolved.
Thread the `useActiveAgentProfile` loading signal into the indeterminate state so
consumers render nothing until the active agent profile is known (#1571 review).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(agent-profiles): scope the default→agent_settings launch to OpenHands profiles

The `default`→agent_settings shortcut (which preserves <RUNTIME_SERVICES>/canvas_ui)
must not apply to an ACP `default` profile: activation is pointer-only, so global
agent_settings is stale (still OpenHands) when an ACP profile is active — routing it
via agent_settings launched the wrong agent (mock-llm-acp-agent.spec.ts step 3
expected agent_profile_id, got OpenHands agent_settings). ACP also carries no
<RUNTIME_SERVICES>/canvas_ui enrichment, so there's nothing to preserve. Gate the
shortcut on agent_kind === "openhands"; ACP defaults keep the profile path.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(agent-profiles): address PR #1571 review findings (VascoSch92)

- Gate the default-profile agent_settings downgrade to local backends
  only; cloud always launches from the resolved agent_profile_id.
- Emit an explicit schema-default (not an omitted key) when
  tool_concurrency_limit is cleared, so edit-save actually resets it.
- Restore the tailored "Switched to {name} failed" toast via
  meta.disableToast + a dedicated onError.
- Share one AGENT_PROFILES_RETRY_OPTIONS constant across the launch
  path, redirectIfAcpActive, and useAgentProfiles so retry policy
  can't drift between call sites.
- Fix a stale comment on optimisticActiveProfile's write path.
- Self-heal a dangling llm_profile_ref in the agent-profile editor by
  validating it against the live LLM-profiles list on load.

* fix(agent-profiles): restore cloud in-conversation LLM-profile switching

resolvePickerKind hard-coded cloud conversations to the read-only
model picker, on the premise that cloud has no per-conversation
switch endpoint. That's not true: POST
/api/v1/app-conversations/{id}/switch_profile has existed since
OpenHands#14288 (2026-05-05), predating this PR, and the frontend
plumbing to call it (AgentServerConversationService.switchProfile's
cloud branch) was already implemented and just unreachable.

Cloud OpenHands conversations now resolve to the LLM-profile picker,
same as local, matching how ACP already behaves identically on both
backends. main's old SwitchProfileButton had no cloud gate either, so
this restores previously-working behavior rather than adding new
scope.

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-08 17:22:55 +00:00
Vasco Schiavo d8f7097112 test(e2e): deflake mock-LLM profile activation helper (#1580) 2026-07-08 16:17:00 +02:00
Graham Neubigandopenhands c552545926 feat(mcp): add OAuth support to MCP install flow
Squash merge PR #1583.

This merge commit was created by an AI agent (OpenHands) on behalf of Graham Neubig.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-07-07 12:03:36 +02:00
Hiep Le b6727865b7 feat: instrument PostHog analytics for the onboarding funnel (#1547)
* feat: instrument PostHog analytics for the onboarding funnel

* test: de-flake onboarding layout probe and profile activation
2026-07-02 22:26:39 +07:00
Graham NeubigandCodex 09d56a330b [codex] Always show first-run onboarding (#1528)
* Always show first-run onboarding

* Fix first-run onboarding CI regressions

* Fix root onboarding launch navigation

---------

Co-authored-by: Codex <codex@openai.com>
2026-06-29 10:01:05 +00:00
Vasco Schiavo 17e3328812 test(e2e): de-flake git control bar workspace pill assertion (#1475)
The mock-LLM files-and-git step 3 asserted the control bar shows the
workspace folder basename ("my-app"). But once the local
`git remote get-url origin` probe detects the remote that step 2 adds,
the bar shows the repo slug instead — the folder pill is only the
pre-detection fallback (GitControlBarRepoButton: selectedRepository ||
workspaceName). The assertion raced that probe and failed intermittently.

Accept either the folder basename or the configured remote slug, and
share the slug via a single EXPECTED_REPO_SLUG constant used by both the
trajectory's git remote add and the assertion.
2026-06-24 16:06:28 +02:00
Hiep Le 8d8ae5f8a0 feat: add deployment-choice modal for GitHub and Slack responders (#1379)
* feat: add deployment-choice modal for GitHub and Slack responders

* fix: failing e2e tests
2026-06-23 18:42:27 +00:00
Engel Nystandopenhands 01d141d0cd ci: collapse mock e2e PR comment tables (#1465)
Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-23 15:56:16 +02:00
a1c68313b2 Add lock-to-cloud backend setup mode (#1389)
* Show onboarding before public backend auth gate

Co-authored-by: openhands <openhands@all-hands.dev>

* Make backend setup the first onboarding step

Co-authored-by: openhands <openhands@all-hands.dev>

* Restore Cloud backend option in onboarding

Co-authored-by: openhands <openhands@all-hands.dev>

* Make first-run backend onboarding calmer

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: update public onboarding e2e expectation

* fix: cover onboarding-first public auth e2e

* test: keep ProgressEvent polyfill through teardown

* chore: refresh PR checks after QA

* Add lock-to-cloud backend setup mode

Co-authored-by: openhands <openhands@all-hands.dev>

* Hide skip on locked Cloud backend onboarding

Co-authored-by: openhands <openhands@all-hands.dev>

* Remove add-backend onboarding subtitle

Co-authored-by: openhands <openhands@all-hands.dev>

* Skip healthy backend onboarding step

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: support skipped backend step in onboarding e2e

* chore: Remove PR-only artifacts

* fix: address onboarding review nits

* fix: show onboarding for locked cloud first run

* ci: support stacked mock llm runs

* test: assert scoped shell background

* fix: resolve merge conflicts with main (fix-public-onboarding stacking)

- Remove duplicate handleConnected/actionRowClassName/titleKey declarations
  in check-backend-step.tsx that resulted from merging the parent PR's
  changes on top of our lock-to-cloud additions
- Remove erroneous waitFor(onboarding-backend-connected) steps from the
  'shows a connection error' test which uses a no-backend context where
  the connection banner is never shown

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: remove unused isLockedToCloud export

All callsites use getLockedCloudHost() !== null directly.
Remove the redundant helper to keep the public API intentional.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: show onboarding first in locked-cloud mode when a session key is present

On PR #1389 Hiep reported that `static-server.mjs --lock-to-cloud ...`
landed on the Manage Backends recovery modal ("Add Backend") instead of
first-run onboarding after a fresh `~/.openhands`.

Root cause: when the build had a baked-in `VITE_SESSION_API_KEY` (or one
was injected via `--session-api-key`), `makeDefaultLocalBackend()` seeded
a Local backend even in locked-to-Cloud mode. That made `isNoBackend()`
false, so `lockedNoBackend` was false and first-run onboarding was
skipped; the subsequent `/server_info` probe failed and `root.tsx`
rendered `MissingAgentServerScreen` (Manage Backends recovery modal).

Fix:
- `makeDefaultLocalBackend()` returns null when `getLockedCloudHost()` is
  set, so locked mode never auto-seeds a Local backend.
- `root.tsx` broadens the gate to `lockedNeedsOnboarding`: locked + (no
  backend OR active backend is not Cloud) triggers onboarding, covering a
  stale persisted Local backend from a previous non-locked session too.

Verified by building with a baked `VITE_SESSION_API_KEY` and serving with
`--lock-to-cloud`: the app now shows the first-run onboarding Cloud-login
screen instead of the recovery modal, and no Local backend is seeded.
Non-locked mode still seeds the Local backend as before.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: locked-cloud onboarding layout + restore CI test mock

CI fix:
- `use-create-conversation-metadata.test.ts` mocks the whole
  `agent-server-config` module but was missing `getLockedCloudHost`, which
  `makeDefaultLocalBackend()` now imports. Add it (returning null) so the
  default local backend seeds and the create-conversation mutation
  succeeds again.

Onboarding layout (locked-to-Cloud first-run step):
- Drop the `max-w-sm` cap on the locked CloudLoginColumn so the "Skip the
  setup — connect instantly with your OpenHands Cloud account." text fills
  the modal content width instead of wrapping in a narrow centered column.
- Add `pb-7` to the onboarding scroll area so the "Login with OpenHands
  Cloud" button is no longer flush with / cut off by the modal bottom.
  Widening the text (fewer lines) plus the bottom padding together give
  the button breathing room.

Co-authored-by: openhands <openhands@all-hands.dev>

* test: add getLockedCloudHost to agent-server-config test mocks

`makeDefaultLocalBackend()` now imports `getLockedCloudHost` from
`agent-server-config`. Two tests that fully mock that module were missing
the export, so the default local backend never seeded and every create-/
read-conversation path threw `NoBackendAvailableError`:

- `agent-server-conversation-service.test.ts` (23 failures on ubuntu CI)
- `use-create-conversation-metadata.test.ts` (already fixed in prev commit)

Add `getLockedCloudHost: vi.fn(() => null)` to both mocks so the non-locked
default-backend seeding path works again.

Co-authored-by: openhands <openhands@all-hands.dev>

* Enhance conversation sidebar with pinned section and grouped organization (#1144)

* Add pinned conversations and reorderable workspace folders to the sidebar.

Persist pins per backend with a capped pinned section, pin-on-hover cards that keep the icon aligned with hover actions via an invisible ellipsis spacer, and drag-and-drop folder ordering stored in panel preferences.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Simplify grouped folder rows for drag and expand.

Drop the grip and chevron controls, remove selection highlight and layout animation, and drag or click the folder label directly while keeping row hover feedback.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Polish folder drag-and-drop and pinned section visuals.

Drag the whole folder (and contents) as the drag image, show an accent drop
line between folders with position-aware reordering, and animate sibling
folders into place only around a reorder. Swap the folder icon to its open or
closed counterpart on hover, add a chronological-view divider plus an outline
pin icon to the pinned section header, and render that header in normal weight.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add hover metadata popover for sidebar conversations.

Show a modal-styled popover on conversation hover with the full title, status
dot, and repo/branch-or-directory, model, and created-date rows. Reserve the
action overlay width so titles truncate instead of colliding with the pin,
drop the small status tooltip, and gate the popover behind a new "Hover
metadata" toggle in the filter dropdown (persisted, on by default).

Co-authored-by: Cursor <cursoragent@cursor.com>

* Improve folder drag preview and placeholder.

Show a rounded, surfaced drag image anchored to the grab point and blank the
original row (preserving its height) via opacity so Chrome does not cancel the
native drag.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Harden sidebar "Load more" pagination

Dedupe loaded conversations by id and keep fetching pages until the
visible list actually grows, so a single "Load more" click reliably
surfaces new rows despite the 10s background refetch dropping in-flight
fetchNextPage calls or pages yielding zero visible rows. Show the
skeleton throughout. Also drop the native title tooltip on card titles
and record the still-intermittent double-click symptom as a KNOWN ISSUE.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Keep pinned conversations exclusive to the pinned section.

Filter pinned threads out of grouped/chronological lists to prevent duplicates, add regression coverage for both list modes, and add the missing upgrade-button translation key with typed i18n usage.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: failing tests

* fix: lint

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: hieptl <hieptl.developer@gmail.com>

* Fix locked cloud onboarding follow-ups

Co-authored-by: openhands <openhands@all-hands.dev>

* Slow down onboarding follow-up GIFs

Co-authored-by: openhands <openhands@all-hands.dev>

* Skip onboarding when active backend already has a configured LLM

Detect returning users via flat `llm_api_key_set` + `agent_settings.llm.model` (or subscription auth), regardless of backend kind. Locked-Cloud-not-logged-in and stale local backend still fall through to the modal so the existing recovery paths kick in.

Co-authored-by: openhands <openhands@all-hands.dev>

* Scope onboarding skip rule to Cloud backends only

Local agent-servers can be started with an env-injected `LLM_API_KEY`, which makes `llm_api_key_set` an unreliable returning-user signal — Mock-LLM E2E fresh-install tests were tripping on the SDK default model + env key combo. For Local backends the skip stays driven by the existing `openhands-onboarded` localStorage flag; Cloud backends continue to use the settings-based rule.

Co-authored-by: openhands <openhands@all-hands.dev>

* Trigger CI re-run (empty commit)

Workflows didn't fire on 80ea575a — pushing empty commit to nudge the webhook.

Co-authored-by: openhands <openhands@all-hands.dev>

* Always pre-fill onboarding LLM step with OpenAI GPT-5.5 default

The returning-Cloud-user case is now handled at the host level (OnboardingHost skips the whole modal). Users who actually reach the LLM step are first-time installs who want the default pre-filled — restoring the pre-PR-1389 behavior that the onboarding-regressions E2E asserts. Also drops the now-empty unit test that mirrored the old step-level preservation.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: Update PR QA artifacts

* Generalize onboarding-skip to Local backends with configured LLMs

Hiep flagged that the onboarding modal still walks users through Set Up
your LLM after they connect to a pre-configured backend. Investigation:

  * On Cloud, the fast-path keyed off settings.llm_api_key_set + a
    non-empty llm.model. That worked.
  * On Local, the fast-path bailed early on backend.kind !== 'cloud'.
    But the local agent-server reports the exact same readiness signal
    via llm_api_key_is_set (and the local settings-service mapper
    already remaps that to llm_api_key_set on the way through). The
    only reason the skip didn't fire was the explicit kind gate.

Drop the gate, accept either field name, and rename the predicate to
reflect what it actually checks (isBackendLlmReady). A truly fresh
agent-server reports both flags as false, so the modal still shows for
genuine first-run setup.

Tests:
  * Updated 'does not skip onboarding for a Local backend' to its
    inverse: 'skips for a Local backend with an LLM already configured'.
  * Added 'still shows the modal for a fresh Local agent-server with no
    API key set' to lock in the fresh-install case.
  * All 3328 vitest tests pass; typecheck clean.

Refs Hiep's review comment on PR #1389.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(onboarding): address Hiep's review on PR #1389 (#1389)

Resolves the three issues Hiep reported on PR #1389:

1. **Choose Agent step gets skipped after Cloud login.** When the
   backend slide finished via Cloud login and `skipBackendStep` flipped
   true, the slide indices renumbered (agent: 1→0, setup: 2→1). The
   user's numeric `currentStep` of 1 — pointing at Choose Agent before
   the flip — now pointed at Set Up LLM, and the corrective effect that
   decremented it ran a render too late. Track the user's *phase*
   ("backend" | "agent" | "setup" | "hello") instead of a numeric
   step. The visible slide index is derived from phase + slideOrder, so
   renumbering can never move the user onto a different logical step.
   The previous `wasSkippingBackendStep` ref + decrement effect is
   replaced by a single effect that snaps phase forward only when the
   current phase is no longer in slideOrder (e.g. "backend" right
   after the slide collapsed).

2. **Existing Cloud LLM settings not shown to returning users.** The
   skip-onboarding fix from commit 78254e1b already routes returning
   users with a configured LLM around the onboarding modal entirely,
   so they never hit the Set Up LLM step in the first place. The new
   phase-based flow preserves that behavior; no further change needed.

3. **Redundant 'Or' divider** between manual and Cloud columns in
   BackendConnectionOptions. Both columns have prominent titles
   ("OpenHands Cloud" with logo on the right) and a generous gap
   already; the explicit divider added visual noise without
   information. Remove the divider markup.

Also gitignores local static-server runtime artifacts (workspace/,
build-fresh/) that were getting picked up by 'git add -A'.

Two regression tests cover the standard (non-locked-cloud) flow: one
verifies the user stays on Choose Agent after completing Cloud login
from the side-by-side picker, and one verifies the 'Or' divider is
gone. All 3,241 unit tests pass; lint and typecheck are clean.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: suppress Add Backend modal in locked-to-cloud mode

Resolves hieptl's review feedback on PR #1389: when the static server is
launched with --lock-to-cloud, navigating to the app showed the Manage
Backends recovery modal ("Add Backend") instead of going straight to
Cloud onboarding/login.

Root cause: the `openhands-onboarded` localStorage flag is origin-scoped
and persists across deployments. A user who previously completed
onboarding in a non-locked session on the same origin carries that flag
into a locked-to-Cloud session. The stale flag suppressed
first-run onboarding (`shouldShowFirstRunOnboarding` was gated on
`!onboardingCompleted`), so the app fell through to the
`/server_info` probe. With no usable local backend in locked mode the
probe throws `AgentServerUnavailableError`, and root.tsx renders the
`MissingAgentServerScreen` / `ManageBackendsModal` recovery modal.

Fix: when `lockedNeedsOnboarding` is true, ignore the completion flag
and force first-run onboarding (which owns the Cloud login). The
non-locked path is unchanged — `onboardingCompleted` still suppresses
the modal for returning users with a configured backend.

Also confirms the minor cleanup from the bot review: `isLockedToCloud()`
was already removed in commit addda40e; no remaining references.

Adds a regression test reproducing hieptl's exact scenario (stale
`openhands-onboarded` flag + locked-to-Cloud + no backend) and asserting
the onboarding modal renders instead of the Manage Backends modal.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(onboarding): don't skip onboarding modal for launcher-seeded backend

PR #1389 generalized the OnboardingHost "returning user with a
configured LLM" skip from Cloud-only to all backends (commit 78254e1b).
That broke the mock-LLM E2E fresh-install / onboarding-happy-path /
onboarding-regressions specs:

  tests/e2e/mock-llm/backends/mock-llm-auth-modes.spec.ts:57
    "auth mode: fresh install with runtime-injected key ›
     reaches the onboarding modal without pre-seeded localStorage"

The mock-LLM E2E stack runs every spec serially against a single
shared agent-server. Earlier specs configure an LLM profile that
persists in the server's settings, so by the time the fresh-install
spec runs (with a clean browser context, no `openhands-onboarded`
flag, and a launcher-seeded default-local backend), the server
reports `llm_api_key_is_set: true` + a non-empty model.
`OnboardingHost.isBackendLlmReady` then returned true, so the host
marked onboarding complete and returned null — the first-run modal
never mounted and the test timed out waiting for
`onboarding-step-choose-agent`. Main is green on the same test
because main's skip was Cloud-only.

The settings-based LLM-ready signal is unreliable for the
launcher-seeded default-local backend: the agent-server can be
started with an env-injected LLM key, and shared-server deployments
retain configured LLMs across browser sessions. Keying first-run
onboarding off the server's LLM state would suppress the modal for
a genuinely fresh browser install.

Fix: keep the skip for Cloud backends and for Local backends the
user explicitly added via "Add Backend" (which carry a non-default
id), but suppress it for the launcher-seeded default-local backend
(`SEEDED_DEFAULT_BACKEND_ID`). First-run detection for that backend
stays driven by the `openhands-onboarded` localStorage flag, matching
main's behavior and restoring the E2E fresh-install contract. The
PR's core intent (suppress the Add Backend recovery modal in
locked-to-Cloud mode, commit 47619f11) is unchanged.

Tests:
  * Updated "skips the modal for a Local backend..." to seed a
    user-added Local backend (non-default id) so the skip still
    fires for the Add-Backend scenario.
  * Added "still shows the modal for a launcher-seeded default-local
    backend even when the agent-server reports a configured LLM" to
    lock in the fresh-install regression.
  * All 3332 vitest tests pass; typecheck + lint + build clean.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(onboarding): don't auto-complete onboarding for launcher-seeded backend

Commit 9029e036 fixed OnboardingHost so the first-run onboarding modal
shows for the launcher-seeded default-local backend even when the
shared mock-LLM agent-server reports a configured LLM. But the same
over-suppression existed in src/root.tsx: a separate
`isBackendLlmReady` check (no default-local exclusion) fed a
`markCompleted()` effect that persisted `openhands-onboarded=1`
whenever the active backend reported a ready LLM — including the
launcher-seeded default-local backend.

That root-level effect was the remaining cause of the
mock-llm-onboarding-regressions.spec.ts:16 failure
("keeps the modal open on backdrop click and Escape"):

  * The OnboardingModal already renders with no `onClose` on its
    ModalBackdrop, so backdrop clicks and Escape are no-ops — the
    modal itself was never closeable that way.
  * The test failure was actually the `expect.poll` asserting
    `openhands-onboarded` stays null: root.tsx's `markCompleted`
    effect fired (agent-server had a configured LLM from earlier
    serial specs) and persisted completion, even though the modal
    stayed mounted.

Fix: apply the same `SEEDED_DEFAULT_BACKEND_ID` exclusion to
root.tsx's `isBackendLlmReady` that OnboardingHost already uses.
The settings-based LLM-ready signal is unreliable for the
launcher-seeded default backend (env-injected keys, shared-server
LLM persistence across browser sessions), so first-run detection
there stays driven by the `openhands-onboarded` localStorage flag.
The skip still fires for Cloud backends and for Local backends the
user explicitly added via "Add Backend" (non-default id).

The OnboardingModal's non-dismissible backdrop/Escape behavior is
unchanged and already correct (ModalBackdrop receives no `onClose`,
so `closeOnEscape`/`closeOnBackdropClick` default-true handlers
call `onClose?.()` which is a no-op).

Tests:
  * Added root.test.tsx case "does not mark onboarding complete for
    the launcher-seeded default-local backend even when the
    agent-server reports a configured LLM" — verified it fails
    without the root.tsx fix and passes with it.
  * All 3333 vitest tests pass; typecheck + lint + build clean.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: force Cloud replacement for stale Local backend in locked mode

Critical fixes for the locked-to-Cloud flow (PR #1389 review):

1. root.tsx: the ready-backend fast-path in locked mode now requires the
   active backend to match the locked Cloud host (normalized via the new
   isSameCloudHost helper), not just . A reachable stale
   Local backend (or a Cloud backend on a different host) that reports a
   configured LLM no longer bypasses the Cloud login/replacement flow.
   The markCompleted effect is also guarded so it only persists completion
   for the legitimate locked Cloud host.

2. onboarding-modal.tsx: in locked mode, CheckBackendStep is only skipped
   when the active backend IS the locked Cloud host. A reachable stale
   Local backend keeps the backend slide visible so Cloud login can
   replace it.

Also addresses minor review suggestions:
- LOCK_TO_CLOUD_WINDOW_KEY is now module-private (only getLockedCloudHost
  reads it; static-server.mjs/tests use the literal string).
- Extract shared isBackendLlmReady helper into its own module
  (is-backend-llm-ready.ts) so root.tsx and OnboardingHost stay in sync
  without duplicating the rule and without pulling the onboarding modal
  graph into root's eager bundle.
- Inline the no-op initialValueOverrides intermediate in setup-llm-step.

Adds regression tests for the stale-Local-backend and other-Cloud-host
scenarios in both root.test.tsx and onboarding-modal.test.tsx.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(onboarding): close stale-backend lock-to-Cloud bypass in CheckBackendStep (#1389)

PR-review bot pointed out (HEAD 55d382be) that keeping the backend
slide visible for a non-matching backend in locked mode is insufficient:
CheckBackendStep itself still hits its connected-backend shortcut for
a reachable stale Local backend, hiding the Cloud login UI and showing
a Next button that lets the user continue as Local.

Apply the same host-match guard inside CheckBackendStep. A new local
`treatAsNoBackend` (= noBackendSelected || lockedCloudHostMismatch)
drives:
  - title: ONBOARDING$LOGIN_TO_CLOUD_TITLE (not BACKEND_TITLE)
  - render: BackendConnectionOptions (Cloud login UI), no ConnectionBanner
  - no "Show configuration" toggle and no Next-shortcut action row

`noBackendSelected` still governs whether handleConnected calls
`addBackend` or `updateBackend`, so the stale backend is replaced
rather than duplicated.

Strengthen the regression test the bot flagged: it now asserts the
Cloud login title and login button are visible, and that the
`onboarding-backend-show-configuration` toggle, `onboarding-backend-next`
button, and the (misleading) Connected subtitle are all absent.

All 3,249 unit tests pass; lint and typecheck are clean.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(onboarding): clear stale active org_id when replacing a Cloud backend host (#1389)

PR-review bot raised one remaining state carry-over: replacing a
mismatched Cloud backend updates its host/apiKey via `updateBackend`,
but the persisted `active.orgId` (X-Org-Id) is keyed to the OLD
host's org list. The newly-locked Cloud backend would keep sending an
invalid `X-Org-Id` until the user manually re-picked an org.

Fix in CheckBackendStep.handleConnected: when the submitted payload's
host differs from the previously-active backend's host, call
`setActive(backend.id, null)` to drop the now-invalid org selection.
The user re-picks an org on the new host via the usual org switcher.

Local-only edits are unaffected because Local backends always carry
`active.orgId === null`, so the conditional is a no-op there.

New regression test seeds a Cloud backend at other-cloud.example.com
with `orgId="stale-org-from-other-host"`, drives the Cloud login
button, and asserts `getActiveSelection().orgId === null` while the
backend row is updated in place (same id).

All 3,250 unit tests pass; lint and typecheck are clean.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(onboarding): dismiss modal immediately after Cloud login in locked mode (#1389)

Resolves the flicker hieptl reported on PR #1389: after logging into
OpenHands Cloud in locked-to-Cloud mode, the onboarding modal advanced
to the Choose Agent slide (the "next window"), then got torn down by
the root first-run gate, then briefly remounted via OnboardingHost —
appearing to flash in and out.

Cloud login IS the onboarding completion in locked mode, so:
- CheckBackendStep now calls onClose (dismiss) instead of onNext when
  a Cloud login succeeds in locked-to-Cloud mode, so the next slide
  never shows. Standard (non-locked) mode still walks the user through
  agent/LLM setup via onNext.
- root.tsx's locked-mode first-run gate now treats onboardingCompleted
  as authoritative once the active backend IS the locked Cloud host,
  so the first-run screen hides immediately on login (without waiting
  for the Cloud settings probe to confirm a configured LLM). The flag
  is still ignored when the active backend is not the locked Cloud
  host, preserving the stale-flag bypass protection.

Added failing tests (now passing) reproducing both halves of the flicker:
- onboarding-modal: Cloud login in locked mode calls onClose, not onNext.
- root: the first-run screen hides immediately after Cloud login
  completes (post-login state with no configured LLM), instead of
  reopening via OnboardingHost.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: Remove PR-only artifacts

* ci: revert docker.yml pull_request branch filter change

Reverts the removal of `branches: [main]` from the `pull_request`
trigger in .github/workflows/docker.yml (introduced in 5bb8049f). That
change is unrelated to the locked-to-Cloud onboarding work on this PR
and is out of scope. Restores the file to match main exactly so the
Docker workflow again only runs on PRs targeting `main`.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: Graham Neubig <gneubig@users.noreply.github.com>
Co-authored-by: neubig <398875+neubig@users.noreply.github.com>
Co-authored-by: allhands-bot <allhands-bot@users.noreply.github.com>
Co-authored-by: hieptl <hieptl.developer@gmail.com>
Co-authored-by: FraterCCCLXIII <panentheum@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-22 16:06:12 +00:00
5728144593 Show onboarding before public backend auth gate (#1385)
* Show onboarding before public backend auth gate

Co-authored-by: openhands <openhands@all-hands.dev>

* Make backend setup the first onboarding step

Co-authored-by: openhands <openhands@all-hands.dev>

* Restore Cloud backend option in onboarding

Co-authored-by: openhands <openhands@all-hands.dev>

* Make first-run backend onboarding calmer

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: update public onboarding e2e expectation

* fix: cover onboarding-first public auth e2e

* test: keep ProgressEvent polyfill through teardown

* chore: refresh PR checks after QA

* chore: Remove PR-only artifacts

* fix: address onboarding review nits

* test: assert scoped shell background

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: Graham Neubig <gneubig@users.noreply.github.com>
Co-authored-by: neubig <398875+neubig@users.noreply.github.com>
Co-authored-by: allhands-bot <allhands-bot@users.noreply.github.com>
Co-authored-by: hieptl <hieptl.developer@gmail.com>
2026-06-18 17:07:35 -04:00
944679d9fd [AgentProfile][canvas] Give SSE/shttp MCP servers referenceable names (custom editor + marketplace) (#1386)
* feat(mcp): give SSE/shttp MCP servers a user-given name

SSE/shttp servers were serialized under auto-generated dict keys
("sse", "shttp", "shttp_1"), making them unreferenceable by name in
AgentProfile mcp_server_refs. Surface an optional "Server name" field
so they get a stable, user-meaningful key — consistent with stdio.

- Add name? to MCPSSEServer / MCPSHTTPServer
- toSdkMcpConfig: reserve(entry.name || "sse"/"shttp") so the name
  becomes the dict key; reserve() still de-dups collisions
- parseMcpConfig: round-trip user-given names, but treat auto-generated
  keys as nameless so existing configs re-serialize unchanged
- Plumb name through the add/update mutations and flattenMcpConfig
- Add optional "Server name" input to the SSE/shttp editor form

Part of #3726 (AgentProfile epic #3713).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(mcp): name marketplace remote installs after the catalog slug

The marketplace install path builds its own payload, separate from the
custom editor. Stdio entries already carry serverName, but remote
(sse/shttp) installs set no name — so GitHub, Linear, etc. landed under
the auto-generated "sse"/"shttp" key and were unreferenceable in
mcp_server_refs, the same gap the custom-editor fix addressed.

Set name: entry.id on the remote-install payload so e.g. GitHub keys as
"github". reserve() still de-dups repeat installs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(e2e): expect GitHub MCP install under the "github" key

The marketplace remote-install path now names servers after the catalog
slug, so a GitHub install persists under mcpServers.github (referenceable
in mcp_server_refs) rather than the auto-generated "shttp" fallback.
Update the install/delete e2e assertions accordingly.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(mcp): validate the optional SSE/shttp server name as a safe key

The Server name we added for SSE/shttp becomes the mcp_config dict key
(and the mcp_server_refs reference), but the form only validated stdio
names. An sse/shttp name with spaces/special chars would produce a
malformed key. Apply the same ^[a-zA-Z0-9_-]+$ rule stdio uses, while
keeping the field optional.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Debug Agent <simon@openhands.dev>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 21:39:21 +02:00
Hiep Le 57d078a271 refactor: extract TreeNode out of file-tree-view (#1358)
* refactor: extract TreeNode out of file-tree-view

* fix: failing tests
2026-06-16 18:53:17 +00:00
Hiep Le e011407118 fix: detect invalid credentials in MCP Test Connection (#1175) 2026-06-15 16:03:39 +00:00
a289be7593 [codex] Fix ACP MCP auth header forwarding (#1164)
* Fix ACP MCP auth header forwarding

* Fix MCP auth header review feedback

* Expect redacted MCP auth headers

---------

Co-authored-by: neubig <398875+neubig@users.noreply.github.com>
Co-authored-by: Graham Neubig <gneubig@users.noreply.github.com>
2026-06-13 23:18:07 -04:00
Rohit Malhotraandopenhands 92b25e0af3 ci: selective E2E test execution based on changed files (#1286)
* ci: add paths filters to E2E workflows to skip irrelevant PRs

Add paths: filters to the pull_request triggers of the three E2E
workflows so they are skipped when a PR only touches files that
cannot affect the test suite (docs, specs, .agents/, unrelated
test directories, etc.).

- mock-llm-e2e.yml: triggers on src/, public/, scripts/, bin/,
  config/, tests/e2e/mock-llm/, tests/e2e/support/, package.json,
  package-lock.json, build/TS/styling configs, and its own workflow
  file.

- snapshot-tests.yml: triggers on src/, public/,
  tests/e2e/snapshots/, tests/e2e/support/, package.json,
  package-lock.json, build/TS/styling/playwright configs, and its
  own workflow file. Both pull_request and push-to-main triggers
  are filtered with the same path set.

- mock-llm-docker-e2e.yml: same paths as mock-llm-e2e.yml plus
  docker/** and playwright.mock-llm-docker.config.ts. The
  workflow_run trigger (post-Docker-build on main) is unaffected
  by path filters and always runs.

workflow_dispatch is unaffected by paths: filters in all three
workflows, so a manual run always executes the full suite.

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: organize mock-LLM E2E tests into feature subdirectories with selective execution

Reorganize the 15 mock-LLM spec files from a flat directory into feature
subdirectories that mirror the source code structure:

  tests/e2e/mock-llm/
    settings/    — LLM profiles, ACP agent, model switching
    conversations/ — core conversation flow, image upload
    automations/ — automation lifecycle, preset cards
    onboarding/  — first-run onboarding flow
    backends/    — auth modes, cross-connect, partial stack
    home/        — workspace selection, folder browser
    skills/      — skill loading and activation
    regressions/ — CSS isolation, event pagination, etc.

Add a test-mapping config (test-mapping.json) and resolver script
(scripts/resolve-affected-tests.mjs) that maps changed source files
to the affected test subdirectories. The resolver has three modes:

  1. Feature-isolated changes (e.g. src/components/features/settings/**)
     → run only the mapped subdirs + regressions
  2. Cross-cutting changes (src/api/**, package.json, shared helpers,
     or any unmapped src/ file) → run the full suite (__ALL__)
  3. Non-relevant changes (docs, specs) → nothing (workflow paths
     filter already skipped)

The mock-llm-e2e.yml workflow now has a 'Resolve affected test
directories' step that queries PR changed files via the GitHub API,
runs the resolver, and passes the result to Playwright. workflow_dispatch
always runs the full suite.

All relative imports in moved spec files are updated. The Playwright
config discovers specs recursively so no config change is needed.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: update PROJECT_ROOT paths in backend specs moved to subdirectory

The partial-stack and cross-connect specs resolve PROJECT_ROOT from
import.meta.url using ../../.. (3 levels). After moving them from
tests/e2e/mock-llm/ to tests/e2e/mock-llm/backends/, they need
../../../.. (4 levels) to reach the repo root. Without this fix,
bin/agent-canvas.mjs resolves to a nonexistent path and the
backend-only test fails with MODULE_NOT_FOUND.

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: include new mock-LLM specs in selective runs

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: fail closed for mock e2e selection

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: avoid pending skipped e2e checks

Co-authored-by: openhands <openhands@all-hands.dev>

* test: align folder workspace e2e with auto-selection

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-13 00:30:09 -04:00
Engel Nystandopenhands 2888a5ca4e fix: preserve hidden LLM base URL on basic saves (#1347)
Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-13 05:05:25 +02:00
Engel Nystandopenhands af9c653f3c settings: trust public OpenHands provider models (#1280)
* settings: trust public OpenHands provider models
* test: update OpenHands profile E2E expectation
* test: allow transitional OpenHands model shape in E2E
* test: update OpenHands provider profile E2E
* ci: pin mock e2e to SDK provider PR

Temporarily run mock-LLM E2E against OpenHands/software-agent-sdk#3548 so the provider settings refactor is validated with the matching SDK end state.

* test: stabilize automation trajectory for SDK pin

Use non-empty safety replies for the automation-run conversation so the mock trajectory completes whether the pinned SDK does or does not consume an internal LLM turn first.

* ci: isolate SDK PR pin to smoke test

Keep the full mock-LLM and Docker E2E suites on the released agent-server/automation stack, and add a targeted profile-management E2E job that runs against the SDK PR git ref.

* ci: remove temporary SDK PR pin

Revert the temporary OpenHands/software-agent-sdk#3548 git ref before switching the PR to the release-branch SDK pin.

* ci: pin SDK smoke to release PR

Point the temporary SDK smoke-test pin at OpenHands/software-agent-sdk#3638 (Release v1.28.0) so PR #1280 validates against the release branch hash.

Co-authored-by: openhands <openhands@all-hands.dev>

Remove the temporary SDK release-branch pin now that software-agent-sdk v1.28.0 is published, and run the mock E2E workflow against the released agent-server package.

* test: align agent server version references

Update docs and drift-detection expectations for the 1.28.0 agent-server pin.

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-13 01:56:11 +00:00
Graham Neubigandneubig df93ca36a7 Add Windows portability guards for workspace flows (#1311)
* Add Windows portability guards for workspace flows

* Increase snapshot workflow timeout

* Trigger CI after timeout update

* Fix windows portability PR after main merge

---------

Co-authored-by: neubig <398875+neubig@users.noreply.github.com>
2026-06-12 19:14:03 -04:00
OpenHands Bot 82ea2b609a feat: save hosted MCP credentials as secrets (#1331) 2026-06-12 20:03:46 +00:00
Rohit Malhotraandopenhands c7c8862c11 Remove visual snapshot tests and bump Docker E2E timeout (#1332)
* Remove visual snapshot tests

Co-authored-by: openhands <openhands@all-hands.dev>

* Bump Docker E2E timeout

Co-authored-by: openhands <openhands@all-hands.dev>

* Align Docker E2E timeout caps

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-12 13:03:58 -04:00
270ef6a876 Remove obsolete OpenHands proxy base URL handling (#1321)
Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: Engel Nyst <engel.nyst@gmail.com>
2026-06-12 00:40:21 +00:00
Tim O'Farrellandopenhands 910b19ae76 chore: bump agent-server → 1.28.1, automation → 1.0.0a9, extensions → 0.4.1 (#1319)
* chore: bump agent-server → 1.28.1, automation → 1.0.0a9, extensions → 0.4.1

Co-authored-by: openhands <openhands@all-hands.dev>

* Test fixes

* fix: inject proxy base_url for litellm_proxy/* when server omits it (agent-server ≥1.28)

Agent-server ≥1.28 may return base_url:null when fetching a litellm_proxy/*
profile config, even when the profile was saved with the All-Hands proxy URL.
This caused the Basic-tab re-save flow in LlmSettingsLocalView.handleSave to
call isOpenHandsProxyModel(model, null) → false, hitting the else-branch that
deletes base_url and stranding the profile (issue #1146).

Fix: add a secondary check — litellm_proxy/* with a missing base_url is treated
the same as litellm_proxy/* with the proxy URL already set, and
OPENHANDS_LLM_PROXY_BASE_URL is injected before the save request is sent.

Also updates the mock-LLM E2E test to accept both storage representations:
- litellm_proxy/* + proxyBaseUrl  (pre-1.28, guards issue #1146 regression)
- openhands/*     + null          (1.28+, server-managed routing)

And adds a unit test exercising the base_url:null path.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-11 19:16:03 -04:00
Tim O'Farrellandopenhands 8071edf72a Revert "chore: bump agent-server → 1.28.1, automation → 1.0.0a9, extensions → 0.4.1 (#1315)" (#1318)
This reverts commit 1917b5d39fbf09dc51213b4b484698fe394314c7.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-11 21:28:04 +00:00
Tim O'Farrellandopenhands 15a52fea75 chore: bump agent-server → 1.28.1, automation → 1.0.0a9, extensions → 0.4.1 (#1315)
* chore: bump agent-server → 1.28.1, automation → 1.0.0a9, extensions → 0.4.1

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: update doc examples to reference agent-server 1.28.1

Update version references in AGENTS.md, scripts/dev-safe.mjs, and
scripts/check-sdk-version-sync.mjs from 1.27.0 → 1.28.1 to stay
in sync with the agentServer pin in config/defaults.json.

Fixes: docs-version-sync.test.ts failures

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: inject proxy base_url for litellm_proxy/* when server omits it (agent-server ≥1.28)

Agent-server ≥1.28 may return base_url:null when fetching a litellm_proxy/*
profile config, even when the profile was saved with the All-Hands proxy URL.
This caused the Basic-tab re-save flow in LlmSettingsLocalView.handleSave to
call isOpenHandsProxyModel(model, '') → false, hitting the else-branch that
deletes base_url and stranding the profile (issue #1146).

Fix: add a secondary check for litellm_proxy/* models with a missing base_url
(null/undefined/empty), treating them the same as a stored proxy URL and
injecting OPENHANDS_LLM_PROXY_BASE_URL before the save request is sent.

Also adds a unit test exercising the base_url:null path.

Co-authored-by: openhands <openhands@all-hands.dev>

* test(e2e): accept agent-server 1.28 model rewrite in proxy profile test

Agent-server 1.28 normalises litellm_proxy/* → openhands/* on storage
and manages the proxy URL internally (returning base_url:null). The old
assertions hard-coded the pre-1.28 storage format (litellm_proxy/* +
explicit proxy URL), causing the test to fail on every 1.28 run.

Extract assertProxyProfileConfig() helper that accepts both storage
representations:
- litellm_proxy/* + proxyBaseUrl   (pre-1.28, guards issue #1146 regression)
- openhands/*     + null           (1.28+, server-managed routing)

The issue #1146 guard is preserved: a litellm_proxy/* profile without a
proxy URL is still flagged as a stranded profile.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-11 21:00:05 +00:00
Engel Nystandopenhands 39c816d38d fix: stabilize snapshot consent handling (#1304)
* fix: stabilize snapshot consent handling

Seed the legacy analytics-consent key in snapshot tests so the server-backed analytics dialog closes before page interactions, and make snapshot reports distinguish visual diffs/new snapshots from actual workflow failures.

* fix: only seed analytics consent for snapshots

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-11 19:55:51 +02:00
Rohit Malhotraandopenhands b969162027 test(mock-llm): add E2E coverage for Files tab, Git control bar, and Browser tab (#1029)
* test(mock-llm): add E2E coverage for Files tab, Git control bar, and Browser tab

Add mock-LLM E2E tests exercising conversation panel tabs and git
integration against the real agent-server:

- Files tab defaults to diff view when a workspace is attached
  (selected_workspace seeded in conversation metadata localStorage)
- Files tab defaults to file-tree view when NO workspace is attached
- Git control bar shows workspace-name pill for folder-attached
  conversations
- Browser tab renders empty state when no page has been browsed

All tests run serial in a single describe block, sharing one
conversation for the workspace-attached cases (steps 3-5) and
creating a fresh conversation for the no-attachment case (step 6).

Issue #511

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(test): reset mock LLM trajectory before each conversation creation

The mock-LLM E2E test failed because the default 2-turn trajectory
was exhausted by preceding test suites (automation, conversation).
After exhaustion every /chat/completions returns 500, so the agent
never produces REPLY_TOKEN and waitForNonUserMessageText times out.

Fix: call resetMockLLM(request) at the top of step 2 and step 6
(before each conversation creation), matching the pattern used by
mock-llm-conversation.spec.ts step 3.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(ci): report timeout instead of '0/0 passed' when test suite is killed

When the CI wrapper kills Playwright after the 5-minute deadline
(exit code 124), no results.json or marker files exist. Previously
the PR comment showed '0/0 passed' with an empty table, which was
misleading.

Now the render script accepts --exit-code from the workflow. When
exit code is 124 and no results exist, it renders a clear timeout
entry: '⏱️ (test suite timed out before completing)' with a note
pointing to workflow logs.

Both mock-llm-e2e.yml and mock-llm-docker-e2e.yml pass the exit
code through.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(test): re-seed workspace metadata in each test step

Each Playwright test() gets a fresh browser context, so localStorage
from step 2 is gone when steps 3-5 run. Extract seedWorkspaceMetadata()
helper and call it in steps 3 and 4 (which assert on workspace-dependent
UI: git control bar name pill and files tab diff-view default).

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(test): assert git control bar buttons instead of workspace name

The agent-server creates conversation worktrees inside the agent-canvas
repo, so git detection always finds the real repo ('OpenHands/agent-canvas')
and the workspace-name fallback ('my-app') never renders. Assert that
Pull/Push buttons are visible instead — these only appear when the git
control bar has successfully detected a repository.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(test): retry ensureMockLLMProfile on transient socket failures

The automation spec's step 1 intermittently fails with 'socket hang up'
on GET /api/settings because the agent-server briefly drops connections
between test suites (while processing cleanup from the previous spec's
afterAll).

Add retryOnTransient() helper that retries up to 5 times (1s delay) on
socket hang up, ECONNRESET, ECONNREFUSED, 502, and 503. Apply it to
both the GET and PATCH calls in ensureMockLLMProfile.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(static-server): handle WebSocket proxy socket errors

The static-server's proxyWebSocket function was missing error handlers
on the piped client/backend sockets. When a WebSocket connection tears
down abruptly during test cleanup (ECONNRESET, EPIPE), the unhandled
'error' event crashes the Node.js process, killing the Docker container
and causing ECONNREFUSED for all subsequent tests.

Add .on('error') handlers to both proxySocket and socket, matching the
pattern already used in ingress.mjs (lines 273-278).

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(test): make git control bar assertion work in npm and Docker

In the npm path the agent-server creates worktrees inside the host repo
so git detection finds 'OpenHands/agent-canvas' and shows Pull/Push
buttons. In the Docker path there's no git repo inside the container,
so the git control bar only shows the workspace name pill.

Use Playwright's locator.or() to assert on whichever indicator appears:
Pull button (npm) or workspace basename text (Docker). Re-add
seedWorkspaceMetadata so the Docker path has a workspace name to show.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(test): use git-init trajectory for cross-environment git detection

Instead of making the test assertion fuzzy, ensure the conversation
workspace is always a proper git repo. Register a custom trajectory
that runs 'git init && git commit' when no repo exists (Docker path)
and skips init when already inside a git worktree (npm path).

This lets the git control bar consistently show Pull/Push buttons in
both environments, making the assertion deterministic.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(test): add git remote in trajectory for Pull/Push button detection

The git control bar shows Pull/Push only when it can parse a
provider+repository from 'git remote get-url origin'. A bare git init
without a remote means the buttons never appear.

Update the trajectory to add a fake GitHub remote when bootstrapping
a new repo (Docker path). Skip when the workspace already has an
origin remote (npm path — inherits the host repo).

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(test): fix shell syntax in git bootstrap trajectory

The if/then/else joined with spaces produced invalid bash: 'then true
else' (missing semicolons). Rewrite using || operator which avoids
the issue entirely. Also increase Pull button timeout to 25s since
useLocalGitInfo polls every 10s.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(test): assert workspace pill as primary gate, soft-check Pull/Push

The useLocalGitInfo probe requires a connected bash WebSocket that may
not be available in Docker after agent completion. The workspace pill
('my-app') is the primary user-facing behavior for folder-attached
conversations and renders reliably from localStorage.

Make the workspace pill the hard assertion (primary gate). Treat
Pull/Push buttons as a soft check that logs a message instead of
failing when the git probe hasn't completed in time.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(test): increase diff toggle assertion timeout for Docker API latency

The toHaveAttribute('aria-checked', 'true') assertion had only a 5s
timeout. useHasAttachedSource depends on useActiveConversation fetching
the conversation API first — in Docker the round-trip can be slower.
Increase to 15s so the React Query response has time to arrive and
trigger the re-render that flips the toggle.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(test): configure git user in Docker trajectory for commit to work

git commit --allow-empty fails in Docker containers without user.name
and user.email configured. Add git config commands to the bootstrap
trajectory so the initial commit actually creates a HEAD ref.

Without a valid commit, useHasGitCommits returns false and the diff
toggle defaults to off — matching the design ('no commits means no
diff base') but not the test expectation.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(test): make diff toggle test environment-agnostic

In Docker, useHasGitCommits may not fire (workspace.working_dir may
be absent or the bash probe may not execute for finished conversations).
This causes the diff toggle to default to 'off' instead of 'on'.

Rather than asserting a specific default, verify:
1. Both toggle options render (diff on / diff off)
2. Clicking 'on' switches the toggle to checked state

This still exercises the full Files tab rendering pipeline and toggle
interactivity without being fragile to the git probe's environment
dependencies.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(test): add animation waits before panel/tab interactions

The right panel uses a 300ms CSS transition. Clicking the diff toggle
immediately after opening the panel causes click interception by the
animation overlay in Docker. Add explicit waits after panel open and
tab switch clicks.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(test): robust panel/tab/toggle waits + force click in step 4

- Wait for tab bar visibility (proves panel animation completed)
- Wait for diff toggle itself (not the files-tab container which may
  be 'hidden' during CSS transition)
- Use force click to bypass residual animation overlay
- Simplify into a single test.step

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(e2e): use parent toggle container instead of .or() to avoid strict mode violation

The SegmentedToggle renders both option buttons simultaneously as a
radio group. Using .or() on two always-visible elements triggers
Playwright's strict mode ('resolved to 2 elements'). Wait for the
parent radiogroup container (files-tab-diff-toggle) instead.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(e2e): wait for diff toggle instead of files-tab container in step 6

The files-tab main container reports 'hidden' during the right-panel
drawer animation. Wait for the inner diff toggle radio group (same
approach as step 4) which is visible once the tab content renders.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: address review feedback — trim verbose comments, fix dead code

- Trim seedWorkspaceMetadata JSDoc to keep only the addInitScript timing note
- Remove self-evident 're-seed' comments in steps 3 and 4
- Trim step 1 trajectory block comment to two lines
- Remove step 2 seed rationale comment (function name is sufficient)
- Remove box-header section dividers added in this PR
- Fix unreachable throw in retryOnTransient via lastError pattern
- Tighten retryOnTransient JSDoc to just list the retried conditions

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: increase Docker E2E timeout from 15 to 25 minutes

The 15-minute job timeout is too tight for PR-triggered runs that must
first wait for the Docker workflow to complete (up to ~5 min) and then
pull the image (up to ~12 min with a cold runner cache), leaving no
room for setup and test execution.

Successful PR runs already take 12-13 minutes typically. With an
unlucky cold Docker cache (observed on the 04:11 UTC run for PR 1029),
the image pull alone took 11+ minutes, causing the job to hit the
15-minute timeout before tests even started.

Increasing to 25 minutes provides sufficient headroom for:
- Docker workflow wait: ~3-5 min typical
- Docker image pull (cold cache): up to ~12 min
- Test infrastructure setup: ~2 min
- Playwright test execution: ~6-7 min

Co-authored-by: openhands <openhands@all-hands.dev>

* Apply suggestion from @malhotra5

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-09 15:39:15 -04:00
chuckbutkusandopenhands f93cb3c9ee settings: persist app preferences and disabled_skills on the agent-server (#1191)
* settings: persist app preferences and disabled_skills on the agent-server

The local agent-server now exposes app_preferences on the persisted
settings (OpenHands/software-agent-sdk#3539): language, sound
notifications, analytics consent, git identity, and disabled_skills are
returned on GET /api/settings under app_preferences and updated via a
new app_preferences_diff field on PATCH /api/settings.

This brings the local agent-server to parity with the cloud, which has
always accepted the same keys at the top level. Drops the localStorage
workaround that mirrored these fields in two keys
(openhands-agent-server-app-preferences and
openhands-agent-server-disabled-skills), along with the
app-preferences-store.ts module and the DISABLED_SKILLS_STORAGE_KEY
helpers it depended on.

- SettingsService.transformApiResponse reads app_preferences from the
  server response and hoists each field onto the flat Settings shape so
  consumers (settings.language, settings.disabled_skills, …) keep
  working unchanged.
- SettingsService.saveSettings routes the same set of fields through
  the new app_preferences_diff for local backends and through the
  existing app_preferences flat-spread path for cloud backends.
- New legacy-app-preferences-migration.ts runs once on first
  getSettings() after upgrade: when the server reports an
  app_preferences block AND legacy localStorage values are still
  present, it pushes them up via app_preferences_diff and clears the
  legacy keys. Pre-1.27 servers (which omit app_preferences entirely)
  cause the migration to no-op so existing data isn't dropped before
  the server can accept it.
- Updated MSW handlers to round-trip app_preferences and
  app_preferences_diff so the mock backend matches production.
- Test coverage: 5 new tests in __tests__/api/settings-service.test.ts
  for the local round-trip, the mixed diff routing, the legacy
  migration, and the pre-1.27 skip path.

Closes the localStorage workaround called out in the recent audit of
agent-canvas localStorage usage (items 3 and 4: disabled_skills and
app-preferences fields).

Depends on agent-server 1.27 / SDK PR #3539.

Co-authored-by: openhands <openhands@all-hands.dev>

* settings: read/write app preferences via misc_settings container

Follow-up to the localStorage cleanup in this PR + SDK refactor in
openhands/software-agent-sdk#3543. The agent-server now exposes
frontend-owned settings under a generic misc_settings container instead
of a top-level app_preferences field.

Wire shape changes:

  Before:                                 After:
  GET /api/settings                       GET /api/settings
    -> { app_preferences: {...} }           -> { misc_settings: { app_preferences: {...} } }

  PATCH /api/settings                     PATCH /api/settings
    body.app_preferences_diff (shallow      body.misc_settings_diff (deep-merged,
    overlay, replaces named fields)         same semantics as agent_settings_diff)

Why the rename to misc_settings: the previous name pinned the API to a
single 'frontend-owned' namespace. Adding a future category like
ui_preferences (sidebar layout / view modes) would have required either
yet another top-level field or shoehorning unrelated UI state into
AppPreferences. With misc_settings as a container, new categories drop
in as nested fields without churning the top-level shape.

Changes:

- settings-service.api.ts
  * SettingsApiResponse.app_preferences -> .misc_settings (typed)
  * SettingsUpdateRequest.app_preferences_diff -> .misc_settings_diff
  * Add MiscSettings interface
  * transformApiResponse reads response.misc_settings?.app_preferences
  * saveSettings emits { misc_settings_diff: { app_preferences } }
  * Local 'has any diffs' check tracks misc_settings_diff
  * Doc comments updated; semantics noted as deep-merge
- legacy-app-preferences-migration.ts
  * Gate on serverResponse.misc_settings, not .app_preferences
  * pushDiff callback now wraps the diff in { app_preferences: ... }
- src/mocks/settings-handlers.ts
  * GET handler returns misc_settings.app_preferences
  * PATCH handler accepts misc_settings_diff; deep-merges nested
    app_preferences into the persisted block
  * Internal mock state stores under misc_settings to match wire shape
- __tests__/api/settings-service.test.ts
  * Four tests updated to assert the new wire shape (local PATCH body,
    GET round-trip, mixed-diff routing, legacy localStorage migration)
  * Pre-1.27 detection test now keys off missing misc_settings
- AGENTS.md
  * App-preferences note rewritten for the misc_settings container,
    explains deep-merge semantics, and documents the in-flight rename
    (flat shape introduced in #3539 never shipped to users)

Cloud path is unchanged: cloud /api/v1/settings still accepts the
fields as flat top-level keys, mirrored by saveCloudSettings.

Verification:

  $ npm run typecheck
  exit 0

  $ npm test -- __tests__/api/settings-service.test.ts \
                __tests__/api/mock-settings-handlers.test.ts
  23 tests passed

  $ npm test
  3009 passed | 12 skipped | 9 todo

  $ npm run lint
  All matched files use Prettier code style!

  $ npm run build
  built in 1.50s

Co-authored-by: openhands <openhands@all-hands.dev>

* Bump agent-server default to 1.27.0

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-09 18:53:13 +00:00
Rohit Malhotraandopenhands 8c2cc3997d fix: GitHub MCP server works in Docker without Docker-in-Docker (#1282)
* fix: GitHub MCP server works in Docker without Docker-in-Docker

The GitHub MCP catalog entry uses `docker run` as its transport command,
which fails inside the agent-canvas Docker container because Docker is not
available (no daemon, no CLI). This is the only MCP integration affected —
all others use `npx` or `uvx`.

Fix:
- Pre-install the `github-mcp-server` Go binary in the Docker image via a
  new multi-arch download stage (supports amd64/arm64)
- Export `getDeploymentMode()` from agent-server-adapter to expose the
  runtime services info mode ("docker", "dev:automation", etc.)
- Add `patchGitHubEntry()` in mcp-marketplace-utils.ts that rewrites the
  catalog entry from `docker run … ghcr.io/github/github-mcp-server` to
  `github-mcp-server stdio` when deployment mode is "docker"
- The patch follows the existing `patchLinearEntry` pattern: immutable
  spread, conditional on entry id, wired into `getMcpMarketplaceCatalog()`

Closes #1190

* docs: document GitHub MCP catalog patching in AGENTS.md

Co-authored-by: openhands <openhands@all-hands.dev>

* test: add E2E test for GitHub MCP install flow via marketplace UI

Exercises the full MCP page UI flow:
- Navigate to /mcp, verify GitHub marketplace card is visible
- Open install modal, verify fields (command, PAT input)
- Validate empty PAT shows error
- Fill PAT, submit with mocked /api/mcp/test success, verify installed
- Delete installed server via toggle + confirmation modal

Intercepts POST /api/mcp/test to return mock success since the real
github-mcp-server binary is not available in the test environment.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: track github-mcp-server version in config/defaults.json

Move the hardcoded GITHUB_MCP_SERVER_VERSION=1.2.0 from the Dockerfile
default into config/defaults.json (versions.githubMcpServer) alongside
the other external dependency pins.

- Dockerfile: ARG no longer has a default; CI and local builds must
  pass it explicitly
- docker.yml: reads the version from config and passes it as a build-arg
- docker-build.mjs: reads the version from config and passes it too

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: correct GitHub MCP binary download URL and remove flaky validation test

- Fix Dockerfile: release assets use github-mcp-server_Linux_{arch}.tar.gz
  (no version in the filename), not github-mcp-server_{version}_Linux_{arch}.tar.gz
- Remove step 3 (empty PAT validation test) which relied on CSS class
  selector that doesn't work reliably in Playwright with compiled Tailwind
- Renumber remaining steps (4→3, 5→4)

Co-authored-by: openhands <openhands@all-hands.dev>

* docs: address review comments — document docker command assumption and arch fallback

- mcp-marketplace-utils.ts: explain why we match on command === 'docker'
  and what happens if upstream changes the catalog entry
- Dockerfile: document the *) arch fallback and when to update it

Co-authored-by: openhands <openhands@all-hands.dev>

* test: assert Docker-specific command patching in GitHub MCP E2E test

The test now asserts the command field value based on the deployment mode:
- Docker E2E: expects 'github-mcp-server stdio' (native binary)
- npm E2E: expects 'docker' (original catalog transport)

Uses MOCK_LLM_DOCKER_IMAGE env var presence (set only by the Docker
Playwright config) to determine which assertion to make. This ensures
the patchGitHubEntry runtime rewrite is exercised in Docker E2E.

Co-authored-by: openhands <openhands@all-hands.dev>

* test: add unit tests for patchGitHubEntry Docker command rewrite

Addresses review feedback to add unit test coverage for the runtime
catalog patching. Three new tests via getMcpMarketplaceCatalog:
- Non-Docker mode: GitHub entry keeps original 'docker run' command
- Docker mode: command rewritten to 'github-mcp-server stdio'
- Docker mode: other entries (Tavily) unaffected

Uses vi.mock to control getDeploymentMode return value.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-09 14:08:40 -04:00
976413534d test(e2e): add folder browser → workspace → conversation E2E test (#1264)
* test(e2e): add folder browser → workspace → conversation E2E test

Adds a mock-LLM E2E test covering the workspace selection flow from
issue #511:
  - Browse local folders via the folder browser UI
  - Add a directory as a workspace
  - Select it in the dropdown and launch a conversation
  - Verify POST /api/conversations receives the correct working_dir
  - Verify selected_workspace is persisted in localStorage metadata

Docker compat: volume-mounts the test directory into the container so
the agent-server's folder browser can list it. Uses a host/container
path split (MOCK_LLM_FOLDER_WORKSPACE_HOST_DIR env var) following the
same pattern as the skill test mounts.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(e2e): address review feedback on folder-workspace test

- Export MOCK_LLM_FOLDER_WORKSPACE_CONTAINER_DIR from Docker config for
  consistency with MOCK_LLM_SKILL_REPOS_CONTAINER_DIR pattern
- Use os.tmpdir() for host-side fallback instead of hardcoding /tmp
- Read container-side path from MOCK_LLM_FOLDER_WORKSPACE_CONTAINER_DIR
  env var (set by Docker config) with os.tmpdir() fallback for npm mode
- Navigate folder browser path segments dynamically instead of
  hardcoding individual directory names

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(e2e): use path.posix for container paths, assert root after nav-up

- Use path.posix.join for CONTAINER_DIR_BASE and TEST_DIR since the
  agent-server filesystem is always POSIX (Linux container or Linux host)
- Replace hardcoded 10-iteration nav-up loop with while(!disabled) plus
  an explicit assertion that we reached '/' before navigating down

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: Remove PR-only artifacts

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: allhands-bot <allhands-bot@users.noreply.github.com>
2026-06-08 17:33:40 -04:00
a248bf0e08 feat: render critic results in conversation events (#485)
* feat: add critic result types, component, and event rendering

Migrate critic visualization from OpenHands PR #14133.

Types (src/types/agent-server/core/base/critic.ts):
- CriticResult — score (0-1), message, metadata
- CriticFeature — name, display_name, probability
- CriticCategorizedFeatures — agent_behavioral_issues, user_followup_patterns,
  infrastructure_issues, other
- CriticMetadata — wraps categorized features and event IDs
- Added optional critic_result field to ActionEvent and MessageEvent

Component (critic-result-display.tsx):
- Star rating (0-5) with color coding: green ≥60%, yellow ≥40%, red <40%
- Percentage display
- Expandable categorized feature breakdown with per-feature probabilities
- Iterative refinement hint when disabled in settings
- Full i18n support (15 languages)

Integration:
- FinishEventMessage renders CriticResultDisplay below the finish message
  when critic_result is present
- UserAssistantEventMessage renders CriticResultDisplay for agent messages
  when critic_result is present

Tests:
- 14 unit tests covering score rendering, star ratings, color coding, label
  rendering, expand/collapse, iterative refinement hint, and multiple
  feature categories

Co-authored-by: openhands <openhands@all-hands.dev>

* refactor: SdkSectionPage multi-source support + verification settings

Refactor SdkSectionPage to accept a `settingsSources` array instead of
single `settingsSource`/`sectionKeys` props, enabling a page to render
fields from multiple schema sources (e.g. both agent_settings and
conversation_settings).

Key changes:
- SdkSectionPage: new `settingsSources: SettingsSourceConfig[]` prop
  replaces `settingsSource`/`sectionKeys`; tracks values/dirty state
  per source; emits combined save payload with per-source diff keys
  (agent_settings_diff, conversation_settings_diff)
- verification-settings: simplified to declarative multi-source config
  pulling critic fields from agent_settings and confirmation/security
  fields from conversation_settings
- condenser-settings, llm-settings: updated to new `settingsSources` API
- Mock handlers: merged critic fields into verification section; updated
  defaults to match upstream schema structure
- All tests updated for new API shape; full suite passes (2313 tests)

Co-authored-by: openhands <openhands@all-hands.dev>

* feat(verification): require user-supplied API key when enabling the critic

Instead of silently reusing the OpenHands provider's LLM API key,
expose a dedicated verification.critic_api_key schema field that
appears (required) once the critic is enabled. The hint underneath
reuses the existing OpenHands Cloud copy from the LLM provider screen
so users know any LLM API key — easiest, their OpenHands Cloud key —
will power the critic.

- Add field to mock agent_settings schema with secret/required/critical
  flags and depends_on: [verification.critic_enabled].
- Extend FIELD_HELP_LINKS with an optional suffixKey so the schema-
  driven help row can render OpenHands Cloud copy without forking it.
- Add SCHEMA$VERIFICATION$CRITIC_API_KEY$LABEL and $DESCRIPTION across
  all 15 locales.
- Cover both enabled (field + help link visible, password, required)
  and disabled (field hidden) states in
  __tests__/routes/verification-settings.test.tsx.

Co-authored-by: openhands <openhands@all-hands.dev>

* ui(verification): drop critic API key into a full-width row below the toggles

The two-column settings grid was placing the critic API key beside
Enable Critic, leaving a tall stretch of whitespace under the toggle
and squeezing the help link copy. Instead:

- Reorder the mock schema so the Critic API Key field comes after
  Enable Iterative Refinement, freeing the right column for the second
  toggle on the first row.
- Introduce FIELD_FULL_WIDTH_KEYS in schema-field.tsx (small UI-only
  set, mirrors the FIELD_HELP_LINKS pattern) and have
  sdk-section-page apply xl:col-span-2 to those fields. The critic
  API key is the only entry for now.

Result in Basic view with the critic enabled:
- Row 1: Enable Critic  |  Enable Iterative Refinement
- Row 2: Critic API Key (full width, with help link)

Co-authored-by: openhands <openhands@all-hands.dev>

* test(snapshots): update verification helper for schema-driven page

The hand-written 'Enable Confirmation Mode' header was removed when
verification-settings.tsx switched to a pure SdkSectionPage, and
confirmation_mode is a prominence: 'major' schema field — so it only
appears in Advanced/All views. The snapshot test helper was still
waiting for the old text and using the old confirmation-mode-toggle
testId, which made all three verification snapshots time out.

- waitForVerificationPage now waits for 'Enable Critic' (the first
  critical-prominence field, always visible), then clicks the
  sdk-section-all-toggle to switch to the 'All' view, then waits for
  the rendered 'Confirmation Mode' label (i18n SCHEMA$…$LABEL gives
  it a capital M, not the schema's raw 'Confirmation mode').
- The on/off tests use the new sdk-settings-confirmation_mode testId
  emitted by SchemaField's SettingsSwitch wrapper. The label-click
  pattern is preserved (the <input type=checkbox> is hidden by
  SettingsSwitch).
- Security-analyzer locator is now case-insensitive (/security
  analyzer/i) since the i18n label is 'Security Analyzer'.

Baselines will needBaselines will needBaselines will needBaselines will needBaselull-width row, and switching to the 'All' view all change
the rendered pixels. Apply the 'update-snapshots' label after this
commit lands.

Co-authored-by: openhands <openhands@all-hands.dev>

* Clarify critic API key guidance

* test: snapshot verification critic settings

* fix: improve critic score rendering accessibility

* test: cover multi-source settings save

* test: cover verification settings dedupe

* style: use strict null checks in critic helper

* test: cover critic result e2e rendering

* test: make live critic e2e start conversation directly

* test: fix live e2e llm profile setup

* chore: Update PR QA artifacts

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: Rohit Malhotra <rohitvinodmalhotra@gmail.com>
Co-authored-by: allhands-bot <allhands-bot@users.noreply.github.com>
2026-06-08 12:37:36 -04:00
Rohit Malhotraandopenhands e6e61b0fa6 refactor(e2e): drive mock-LLM test interactions through the UI (#1222)
Replace direct API calls for seeding/configuring state in mock-LLM E2E
tests with UI-driven interactions wherever possible, ensuring downstream
API calls are covered by the test.

Changes:

- mock-llm-profile-management.spec.ts: All three scenarios (active
  profile deletion, same-model identity, litellm_proxy base_url
  preservation) now create and activate profiles through the Settings →
  LLM Profiles UI instead of raw POST/activate API calls. Cleanup in
  afterAll uses the UI delete flow via exported deleteProfileIfExists.

- mock-llm-model-switch.spec.ts: The switch-target profile B is now
  created through the Settings UI (createProfileViaUI) instead of a
  raw POST to /api/profiles. Cleanup uses UI-driven deletion.

- mock-llm-skills.spec.ts: Replaced the inlined configureMockLLM()
  helper (which did a raw PATCH to /api/settings) with the UI-driven
  ensureMockLLMProfile(page) that creates and activates the profile
  through the Settings screen.

- mock-llm-acp-agent.spec.ts: afterAll cleanup now resets agent type
  back to OpenHands via the Settings → Agent UI (resetToOpenHandsAgentViaUI)
  instead of a raw PATCH to /api/settings. The local selectDropdownOption
  is removed in favor of the shared export from mock-llm-helpers.

- mock-llm-conversation.spec.ts: Removed the redundant API pre-check
  in step 3 that verified the profile was active via GET /api/profiles.
  Steps 1+2 already verified this through the UI (Active badge check).

- mock-llm-helpers.ts:
  - Extracted createProfileViaUI() from ensureMockLLMProfile() as a
    standalone exported helper for tests that need to create profiles
    without activating them.
  - Exported deleteProfileIfExists() and activateProfileViaUI() so
    tests can compose profile lifecycle operations through the UI.
  - Added selectDropdownOption() (consolidated from ACP spec's local copy).
  - Added resetToOpenHandsAgentViaUI() for UI-driven agent type reset.
  - Marked the API-based resetToOpenHandsAgent() as @deprecated.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-08 15:50:50 +00:00
ec4616c1c7 feat(acp): containerized + cloud ACP — onboarding, secrets, and recycled-sandbox resume (#1013/#1014/#988) (#1102)
* feat(acp): containerized ACP — credential onboarding + inline secrets (#1013/#1014)

Wire the canvas halves of agent-canvas#1014 (Docker) and #1013 (credential
onboarding) so a user can run an ACP agent (Codex / Claude Code / Gemini)
against a containerized agent-server through Canvas, with credentials supplied
in the UI.

Credential onboarding UX (#1013):
- Extend the ACP secrets step beyond the API key to the per-provider reserved
  credentials a fresh container needs: Codex CODEX_AUTH_JSON, Claude
  CLAUDE_CODE_OAUTH_TOKEN, Gemini GOOGLE_APPLICATION_CREDENTIALS_JSON +
  GOOGLE_CLOUD_PROJECT/LOCATION + GOOGLE_GENAI_USE_VERTEXAI. File-content blobs
  render as multiline fields.
- Make the step capability-driven: required on a backend with no host login
  (cloud, or a logged-out local/Docker backend per the auth probe), optional
  when a login is detected or the probe can't classify (native dev).
- Fix the orphaned-secret bug: warn instead of toasting "Saved" when the active
  backend can't consume the credential (cloud can't yet read file secrets).

Send secrets + model (start request):
- buildStartConversationRequest emits reserved ACP credentials inline as
  StaticSecrets (overriding any same-named LookupSecret) and mirrors them onto
  agent_context.secrets, so the SDK's acp_file_secrets defaults materialise the
  *_JSON blobs before the CLI spawns. The orchestrator reads back the saved
  reserved values for the active provider (local backends only).
- Preselect a Vertex-safe acp_model for Gemini (gemini-2.5-flash) so a fresh
  container doesn't hit gemini-cli's preview default that 404s on Vertex.
- Never auto-promote *_BASE_URL to an inline secret (an inherited base URL
  breaks the Claude OAuth token's bearer auth).

Docker setup + docs:
- examples/acp-docker/ docker-compose (persistent volume + canvas_ui tool mount
  + credential notes); .env.sample + docs point VITE_BACKEND_BASE_URL at it.
- docs/ACP_AGENTS.md gains a "Running ACP agents in a Docker container" section.

Per-conversation isolation (acp_isolate_data_dir) left as a documented TODO —
the field isn't exposed on ACPAgentSettings in the released typescript-client.

Tests + e2e:
- Unit tests for the StaticSecret emission, reserved-credential sets, Vertex
  model default, getSecretValues read-back, and the required-credentials matrix.
- tests/e2e/live-acp/: a vite-node harness that builds each provider's request
  via buildStartConversationRequest and POSTs it to a real container. Validated
  with REAL API calls against agent-server c950fdb-python: Codex ✅, Claude ✅,
  Gemini ✅ (materialise ADC -> vertex-ai -> real reply). Gemini's default-config
  init is blocked by an SDK/gemini-cli set_session_mode("yolo") issue (documented
  caveat, not a credential problem).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(acp): make containerized credentials survive the real conversation-start path

Validating end-to-end through the application's own orchestrator
(buildStartConversationRequestWithEncryptedSettings) against a live container —
rather than the request builder in isolation — surfaced two real bugs that would
have broken the feature in the product:

1. secrets_encrypted mangled the plaintext reserved StaticSecrets. The app always
   fetches settings in encrypted mode, so the start request carried
   secrets_encrypted=true. The agent-server then runs every secret value through
   cipher.decrypt() during validation — including our reserved ACP creds, which
   are read back as PLAINTEXT. Result: the credential was silently dropped
   (decrypt fails → None) on a cipher backend, or a hard 500 ("cipher not
   configured") on a fresh container with no OH_SECRET_KEY. Fix: don't set
   secrets_encrypted for ACP conversations — an ACP agent has no encrypted agent
   secret (no LLM api_key), and its provider creds ride as plaintext StaticSecrets.

2. A different provider's leftover file-content secret broke the active provider.
   A CODEX_AUTH_JSON saved while onboarding Codex leaks into a later Claude
   conversation via the global-secrets → LookupSecret path. The SDK materialises
   file secrets eagerly at spawn by resolving the secret source, and a LookupSecret
   resolution stalls → ReadTimeout → "Failed to start ACP server: timed out". Fix:
   reserved file-content blobs (the multiline *_JSON creds) never travel as
   LookupSecrets — the active provider's is sent inline as a StaticSecret, any
   other provider's is dropped (getAllReservedAcpFileSecretNames).

Re-validated through the app orchestrator against agent-server c950fdb-python
(onboarding createSecret → buildAcpAgentSettingsDiff PATCH → orchestrator
read-back → real reply): Codex ✅, Claude ✅ (leftover CODEX_AUTH_JSON correctly
dropped). Gemini's app path is correct (StaticSecrets emitted, vertex-ai auth
reached); this run hit the documented invalid_rapt stale-ADC caveat (host ADC
expired since the prior fresh-ADC pass) — an environment issue, not code.

Adds regression tests (secrets_encrypted suppressed for ACP / kept for non-ACP;
leftover file blob dropped not LookupSecret'd; getAllReservedAcpFileSecretNames)
and the app-path e2e harness (tests/e2e/live-acp/acp-docker-app-e2e.mts). Notes
OH_SECRET_KEY as optional (secret persistence) in the compose example.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: address PR review feedback (#1102)

- retag acp_isolate_data_dir TODO #1014 (this PR) -> #1019 (the
  per-conversation isolation follow-up the knob serves)
- note the Gemini Vertex scalars (PROJECT/LOCATION/USE_VERTEXAI) are
  plain config / a routing flag, not secrets

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(acp): order subscription credential before API key in onboarding

Show each provider's reserved subscription/Vertex credential first
(Claude CLAUDE_CODE_OAUTH_TOKEN, Codex CODEX_AUTH_JSON, Gemini Vertex SA),
then the API key, then the base URL — the subscription token is the
primary auth path for ACP providers, with the API key as the fallback.
Display order only; getAcpProviderSecrets consumers are order-independent.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(i18n): disable i18next value escaping so React handles it

i18next's default escapeValue double-escapes interpolated values on top
of React's own escaping, rendering paths like ~/.codex/auth.json as
~&#x2F;.codex&#x2F;auth.json. Set interpolation.escapeValue=false (the
standard react-i18next config); React still escapes at render time.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(acp): unify secret wire-delivery; keep "reserved" as onboarding-only

Drop the reserved-vs-custom split in how secrets reach the agent-server.
Previously, provider credentials ("reserved") rode inline as StaticSecrets
while user secrets rode as loopback LookupSecrets — a fork introduced only
to dodge a deadlock: the SDK resolved an ACP agent's secrets synchronously
on its event loop at CLI spawn, so a loopback LookupSecret self-deadlocked.

That deadlock is fixed at the source in software-agent-sdk#3510 (ACP
cold-start runs off the event loop), so the workaround is no longer needed.
Now every secret — env-var credential, file-content blob, or user secret —
ships uniformly as a LookupSecret, for ACP and non-ACP alike. The SDK
resolves and (for file blobs) materialises them off the loop, so the
loopback fetch is safe.

"Reserved" survives only as an onboarding/validation concept (which fields
to prompt for per provider, capability-driven required steps) — it no
longer affects the wire.

Removed: StaticSecret type, acpStaticSecrets option + the inline path, the
file-blob lookupSkip, SecretsService.getSecretValues, and the reserved-name
value read-back. Kept: secrets_encrypted suppression for ACP (an ACP
request carries no encrypted payload, and a fresh ACP container may have no
OH_SECRET_KEY cipher).

Note: getReservedAcpSecretNames / getAllReservedAcpFileSecretNames in
constants/acp-providers.ts are now unused by the wire; the former is still
useful for validation, the latter can be pruned.

Depends on software-agent-sdk#3510.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(acp): prune now-dead reserved-secret wire helpers

Follow-up to the wire-delivery unification: getReservedAcpSecretNames and
getAllReservedAcpFileSecretNames were only ever consumed by the inline
StaticSecret / file-blob-skip path, which is gone. They have no remaining
production callers, so remove them (and their tests). The reserved-credential
field definitions (ACP_RESERVED_CREDENTIALS, getAcpProviderSecrets) and the
``reserved`` / ``multiline`` flags stay — onboarding still reads them.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(acp): re-point containerized ACP at SDK 1.25.0 (#3510) + fix e2e harnesses

The unified LookupSecret delivery (e076e9bb) depends on software-agent-sdk#3510
(ACP cold-start off the event loop), which first ships in v1.25.0. The example
compose/docs/e2e all still defaulted to agent-server:c950fdb-python, which
predates #3510 and deadlocks the first ACP turn ("Failed to start ACP server:
timed out"). Bump every default to 1.25.0-python and document it as the minimum.

Also realign the live-acp e2e harnesses, which still encoded the removed
StaticSecret API (the PR's headline evidence predated the unification):
- acp-docker-e2e.mts: store each credential via SecretsService.createSecret,
  send name-only customSecrets, assert every emitted secret is a LookupSecret.
- acp-docker-app-e2e.mts: flip the assertion StaticSecret -> LookupSecret; drop
  the stale getSecretValues reference.
- Both: fix a polling bug where "idle" (the transient pre-run state) was treated
  as terminal, so the loop bailed before the agent ran and read an empty reply.
  Terminal is now {finished, error, stuck, stopped}.

Correct the stale StaticSecret doc comments in constants/acp-providers.ts
(reserved is now an onboarding/validation marker, not a wire distinction).

Re-validated in-container against agent-server:1.25.0-python: Codex and Claude
pass end-to-end on both harnesses (LookupSecret resolves off-loop, no deadlock,
even with leftover cross-provider file-secrets present). Gemini's credential
path is proven (vertex-ai auth reached) but the turn is blocked by gemini-cli
0.45.x ignoring the requested acp_model and running gemini-3-flash — an SDK
model-selection concern tracked in software-agent-sdk#3532, not a Canvas bug;
the docs/e2e notes are corrected accordingly.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(acp): improve credential hint text with fetch commands

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(acp): show provider credentials in Settings → Agent

Adds a Credentials section to /settings/agent when an ACP provider is
selected, so users can set or rotate tokens/keys after onboarding without
hunting through Settings → Secrets. Mirrors the onboarding fields exactly
(same hints, same already-saved placeholders, Optional tag on multiline
fields) with its own Save button that writes directly to the secret store.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(acp): drop the agent_context.secrets mirror — request.secrets is the sole channel

The mirror's justification ("ACPAgent's spawn-time env loop reads from
agent_context.secrets, not the registry") predates the pinned minimum
agent-server: 1.25.0 already injects the ACP spawn env from
secret_registry, seeded from request.secrets (sdk#3299/#3464), and
sdk#3528 removes the agent_context drain entirely. Keeping the mirror
preserved a second, dead credential channel — the exact coupling
agent-canvas#1039 is eliminating.

Canvas now sends every credential in top-level request.secrets only.
Tests inverted to pin the single-channel contract; adapter/type
comments updated to match.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(acp): non-flash Gemini default, shared credential form, review cleanups

- ACP_VERTEX_SAFE_MODEL → gemini-2.5-pro: gemini-cli 0.45.x re-resolves any
  *-flash id at generation time to its current default flash (sdk#3532), so a
  flash pin is never honored; docs + e2e defaults updated to match
- extract AcpSecretField + useSaveAcpSecrets and move AcpCredentialsSection
  to components/ — onboarding and Settings → Agent share one field renderer
  and one save flow (incl. the orphaned-file-credential warning on cloud)
- a required credentials step is only satisfied by an actual credential (a
  masked `secret` field) — a base URL or GCP scalar alone no longer unblocks
- warn inline when CLAUDE_CODE_OAUTH_TOKEN and ANTHROPIC_BASE_URL are both
  set (typed or saved) — the pair silently breaks bearer auth
- drop the near-dead `reserved` field flag; collapse the leftover two-block
  secrets scaffolding in buildStartConversationRequest
- sync 14 stale locales on the OAuth/file-blob hints; fix issue refs
  (TODO #1019→#1014 — #1019 is closed; OpenHands#1016→agent-canvas#1016)
- tests: settings credentials-section coverage, non-flash pin, conflict
  matrix, tightened-gate cases

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(acp): unify default-model surfaces + dedupe credential forms and e2e harness

Review-pass cleanups:

- Route ALL three default-model surfaces (onboarding diff builder,
  Settings -> Agent seeding, start-request null fallback, + chat-input
  display) through getAcpPreferredDefaultModel, so the Vertex-safe
  Gemini override can't diverge between surfaces. New regression tests
  pin the diff-builder and start-request fallbacks to it.
- Extract useAcpCredentialForm + AcpConflictWarnings: the onboarding
  step and the Settings credentials section now share the values state,
  existing-secret lookups, conflict pairs, and save flow.
- Extract tests/e2e/live-acp/harness.mts: provider plans, host
  credential collectors, and HTTP/poll helpers shared by both live
  scripts (a model default can no longer drift between them).
- Restore the TODO(#1019) retag (accidentally reverted to the
  self-referencing #1014 in the last cleanup commit); same fix in
  docs/ACP_AGENTS.md.
- Drop the tautological ACP_VERTEX_SAFE_MODEL literal assertion, fix a
  dead key-ternary in getAcpProviderSecrets, TODO(#1016) on the
  cloud file-credential capability check, and document that baked .env
  creds don't satisfy the onboarding login probe.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: address PR review feedback (#1102)

- Restore package-lock.json to main — the npm-install churn (29 dropped
  "dev": true flags) was never meant to ship with this PR
- Note why global escapeValue:false is safe (React escapes at render;
  no translated string hits dangerouslySetInnerHTML)
- Note the non-macOS skip path in the e2e claudeOAuthToken collector

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(acp): tighten the credential gate + clarify base-URL docs (#1102 review)

- A file blob no longer satisfies the required credential step on a
  backend that can't materialise it (cloud, #1016) — the save flow
  already warned it was orphaned, so it can't be what opens the gate.
  consumesFileCredentials moves into useAcpCredentialForm so the gate
  and the save warning share one capability check.
- Next stays disabled while the login probe is still classifying a
  local backend, so a fast click can't slip past a gate about to come
  up "unauthenticated". A probe that completes as "unknown" stays
  permissive.
- Docs: a saved *_BASE_URL secret does ride along on every start
  request like any other saved secret; Canvas only never derives one
  from LLM settings. Reword the two claims that suggested otherwise.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(e2e): record 2026-06-07 re-validation — all three providers pass

Fresh 1.25.0-python container + fresh volume at the branch tip: Codex and
Claude pass both scripts; Gemini's full turn now passes too (fresh ADC +
gemini-2.5-pro + session-mode override), upgrading the previous
"blocked on model selection" row.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(acp): resume a recycled cloud ACP conversation via bootstrap prompt (#988)

A cloud ACP conversation whose sandbox was recycled (STOPPED/MISSING, e.g. the
runtime idle-stopped or hit its TTL) was a read-only dead end: the chat input
was replaced by the archived banner, and cloud createConversation never
re-provisions an existing conversation_id. The backend already supports
resuming such a conversation — re-issuing the start with the same
conversation_id rebuilds it and, for ACP, replays the durable event store as a
bootstrap prompt (OpenHands#14640) — but nothing in canvas triggered it.

Surface it:
- AppConversationStartRequest.conversation_id so the cloud start path can target
  an existing conversation.
- wakeRecycledCloudConversation(id, repoSelection): re-POST /api/v1/app-conversations
  with the conversation_id (and repo selection, so the rebuilt working dir
  matches the original cwd an ACP resume keys off).
- useWakeConversation mutation: wakes + invalidates the conversation queries so
  the active-conversation poll reconnects once the fresh sandbox is RUNNING.
- A Resume button in the archived banner for an ACP conversation whose sandbox
  is MISSING (ERROR stays read-only).

Validated e2e against a local SaaS-equivalent stack (OpenHands main app_server +
a main-built agent-server image, Docker sandboxes): create an ACP conversation,
docker rm -f the sandbox, wake → fresh sandbox + bootstrap-prompt resume, the
agent recalls prior context (codeword) and the <<RESUMED CONVERSATION>> marker
is present.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(acp): consume file-content credentials on cloud too (#988)

Cloud now materialises reserved file-content credentials (Codex auth.json,
Gemini Vertex SA) from the per-user encrypted secret store via
agent_context.secrets at conversation start (the cloud backend pins an SDK that
materialises reserved file secrets), so a pasted blob is consumable on every
supported backend — not just local. Drop the local-only gate on
consumesFileCredentials: a Codex/Gemini file blob now satisfies the onboarding
credential gate on cloud and saving it toasts success instead of the
orphaned-credential warning.

Folds the remaining cloud-enablement piece in from the native-resume canvas
branch (the wake/bootstrap-resume path landed separately); native session/load
is a backend-only concern (SDK + OpenHands), so canvas needs nothing further.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Debug Agent <debug@example.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 12:14:14 +02:00
Rohit Malhotraandopenhands e32608b609 test(e2e): add mock-LLM coverage for litellm_proxy base_url preservation (#1183)
Add a third regression test to mock-llm-profile-management.spec.ts that
exercises the fix from PR #1148 (issue #1146) end-to-end:

1. Creates a profile via the API with a litellm_proxy/* model paired with
   the All-Hands proxy base_url — the exact state the SDK persists after
   rewriting an openhands/* model selection during onboarding.
2. Opens the profile in the UI in edit mode.
3. Ensures the Basic tab is active and clicks Save.
4. Reads the profile back via the API and asserts the proxy base_url was
   preserved (not stripped by the Basic-tab save logic).
5. Reloads the page and re-verifies persistence.

Also adds a getProfileConfig() helper for reading profile details via the
API with encrypted secrets, and extends saveProfile() with an optional
baseUrl parameter so callers can set a non-default base URL.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-07 18:15:52 -04:00
c39354f2b3 feat: load public skills from @openhands/extensions npm package (#1199)
* build(deps): move @openhands/extensions to npm 0.2.0

* feat: load public skills from @openhands/extensions npm package

Public skills are now loaded from the @openhands/extensions npm package
via a standard JS module import instead of fetching them through the
agent-server (which cloned the extensions GitHub repo at runtime).

  import { SKILLS_CATALOG } from '@openhands/extensions/skills';

SkillsService maps each SkillCatalogEntry to a SkillInfo and merges the
bundled public catalog with user/project skills fetched from the
agent-server (load_public: false). If the agent-server is unreachable,
the bundled catalog is returned alone.

Changes:
- SkillsService: imports SKILLS_CATALOG from @openhands/extensions/skills,
  maps entries to SkillInfo, merges with user/project skills from
  agent-server (load_public: false).
- agent-server-adapter: hardcodes load_public_skills: false in
  buildAgentContext().
- agent-server-config: removes shouldLoadPublicSkills() and its
  VITE_LOAD_PUBLIC_SKILLS env var.
- dev-safe.mjs: removes getExtensionsRef() / DEFAULT_EXTENSIONS_REF
  and EXTENSIONS_REF injection in buildAgentServerEnv().
- Docker: removes CONFIG_EXTENSIONS_REF from config-gen stage and
  EXTENSIONS_REF from entrypoint.sh.
- .env.sample: removes VITE_LOAD_PUBLIC_SKILLS comment.
- Tests updated to match new architecture.

Depends on OpenHands/extensions#310 which adds the SKILLS_CATALOG export.

Co-authored-by: openhands <openhands@all-hands.dev>

* test: remove activated_skills assertion from preset-automation E2E

With load_public_skills: false the agent-server no longer loads public
skills at runtime, so activated_skills is always empty. The conversation
itself works (slash command sent, agent replies) — only the server-side
skill activation metadata is gone.

Co-authored-by: openhands <openhands@all-hands.dev>

* feat: pass bundled public skills via agent_context.skills for SDK-side activation

Instead of doing frontend-side trigger matching, pass the bundled
SKILLS_CATALOG entries directly in agent_context.skills at conversation
start. The SDK performs trigger matching, sets activated_skills on user
events, and injects skill content into the system prompt — the exact
same behavior as when load_public_skills was true, but without cloning
the extensions repo at runtime.

buildBundledSkills() converts each catalog entry into the SDK Skill JSON
shape with KeywordTrigger ({ type: 'keyword', keywords: [...] }) for
skills with triggers, or null for always-active skills.

Restores the activated_skills E2E assertion in the preset-automation
test since the SDK now handles activation.

Co-authored-by: openhands <openhands@all-hands.dev>

* test: add E2E tests for project/user skill loading and deletion

Add mock-llm-skills.spec.ts with three tests:
1. Project skill in workspace/.agents/skills/ triggers on matching keyword
2. User skill in ~/.openhands/skills/ triggers on matching keyword
3. Deleting a user skill removes it from subsequent conversations

Tests create ephemeral SKILL.md files with unique trigger keywords,
send messages through the real agent-server stack, and verify
activated_skills in the conversation events API.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use explicit APIRequestContext type import for CI TS6 compatibility

Replace inline `import('@playwright/test').APIRequestContext` type
references with a proper top-level type import. Also align afterEach
fixture destructuring with other specs' pattern.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: remove node: prefix from imports to fix CI TS resolution

TypeScript 6 on CI (Node 24) has a type resolution conflict when
`node:` prefixed imports (node:path, node:fs, node:os) coexist with
`@playwright/test` types in the same file. This caused
`APIRequestContext` to be incorrectly resolved as `Page`. Use
unprefixed imports (path, fs, os) which work identically in Node.js.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: split fs helpers into separate file to fix CI TS6 type resolution

Move node built-in imports (path, fs, os) and filesystem helpers to
`utils/skill-test-helpers.ts`. The spec file now only imports from
`@playwright/test` and the two helper modules, avoiding the type
resolution conflict between node builtins and Playwright fixture types
that caused `APIRequestContext` to be incorrectly inferred as `Page`
on CI (TypeScript 6 / Node 24 / Ubuntu).

API assertion logic is now inline within each test step, using the
`request` fixture directly instead of standalone functions with
explicit `APIRequestContext` type annotations.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use namespace imports to avoid TS6 type inference issue

Switch from named imports to namespace imports (`import * as helpers`)
with subsequent destructuring. This changes how TypeScript resolves the
imported function signatures, avoiding a Node 24 / TS6 type inference
bug where `ensureMockLLMProfile` was incorrectly resolved as expecting
`Page` instead of `APIRequestContext`.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: add typed wrapper for ensureMockLLMProfile to fix CI TS2345

Add a local `configureMockLLM` wrapper with an explicit
`APIRequestContext` type annotation. This works around a CI-specific
TypeScript 6 type inference issue where the imported
`ensureMockLLMProfile` signature is incorrectly resolved as expecting
`Page` instead of `APIRequestContext` when called from a Playwright
test body that also imports from `skill-test-helpers` (a module with
node built-in imports). The wrapper's explicit type annotation forces
correct type checking at the call site.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: inline ensureMockLLMProfile logic to fix CI TS2345

Instead of importing ensureMockLLMProfile from mock-llm-helpers (which
triggers a CI-specific TS6 type inference bug when combined with
skill-test-helpers imports), inline the same logic as a local function
with explicit APIRequestContext typing. This avoids the cross-module
type resolution issue entirely.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: resolve WORKSPACE_DIR relative to agent-server CWD, not STATE_DIR

The agent-server resolves the relative working_dir ("workspace/project")
from its own CWD (the project root), not from STATE_DIR/workspaces.
The test was writing skill files to the wrong directory so the SDK
never found them, causing activated_skills to be empty.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: create standalone git repo for project skill E2E test

The agent-server creates a git worktree for each conversation, and only
committed files appear in worktrees. The previous approach wrote skill
files to the filesystem without committing them, so the worktree never
contained them and load_project_skills found nothing.

Now the test:
1. Creates a standalone git repo (.tmp/mock-llm-skill-repos/) with the
   skill file committed
2. Creates the conversation via API with that repo as working_dir
3. The agent-server worktree includes the committed skill
4. load_project_skills discovers it in the worktree

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: add secrets_encrypted flag to skill test conversation creation

The GET /api/settings with X-Expose-Secrets: encrypted returns cipher-
encrypted secret values. The POST /api/conversations needs
secrets_encrypted: true to tell the server to decrypt them, otherwise
the request fails with HTTP 422.

Co-authored-by: openhands <openhands@all-hands.dev>

* refactor: use UI workspace selection for project skill E2E test

Instead of creating conversations via API (bypassing the frontend code),
the test now exercises the full UI flow:

1. Creates a standalone git repo with the skill committed
2. Registers the repo as a workspace via POST /api/workspaces
3. Opens the 'Open workspace' dialog in the UI
4. Selects the workspace from the dropdown
5. Types the message and submits via the chat input

This exercises the actual frontend code paths (workspace dropdown,
workspace selection form, createConversation with workingDirOverride)
that real users go through.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: add padding response for skill-analysis in deletion test

The agent-server makes a skill-analysis LLM call even when no user/project
skills are loaded, because public skills from the npm package are still
present. The deletion test only had 1 trajectory response, causing the
agent to hang waiting for the 2nd response (the actual reply).

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: simplify deletion test to not depend on specific event type

The deletion test was failing because it waited for an event with
source='agent' and event_type='message' in the events API, but the
mock LLM text reply may produce a different event type. Since
waitForNonUserMessageText already confirms the agent replied in the
UI, we just need to verify no activated_skills contains the deleted
skill name.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: mount skill test dirs into Docker container for e2e tests

The Docker E2E skills test was failing because the agent-server inside
the Docker container couldn't access skill repos and user skill files
created on the host filesystem.

Fix by:
- Adding volume mounts for skill repos (.tmp/mock-llm-skill-repos/ →
  /tmp/mock-llm-skill-repos/) and user skills (.tmp/mock-llm-user-skills/
  → /home/openhands/.openhands/skills/) to the Docker run command
- Setting env vars (MOCK_LLM_SKILL_REPOS_CONTAINER_DIR,
  MOCK_LLM_USER_SKILLS_HOST_DIR) so skill-test-helpers.ts can
  distinguish host-side vs agent-side paths
- Updating createProjectSkillRepo to return both hostDir and agentDir
  so the test registers the container-side path with the agent-server

In npm mode (no env vars set), all paths fall back to the existing
host-side values — no behavior change for the npm test path.

Co-authored-by: openhands <openhands@all-hands.dev>

* docs: document Docker skill test volume mounts in AGENTS.md

Co-authored-by: openhands <openhands@all-hands.dev>

* feat: mark newly added mock-LLM E2E tests with 🆕 badge in PR comments

The render-mock-llm-report.mjs script now accepts a --new-files flag
with a comma-separated list of spec file paths added in the PR. Tests
from those files get a 🆕 badge in the results table, and the summary
line shows the count (e.g. '🆕 2 new').

Both CI workflows (mock-llm-e2e.yml and mock-llm-docker-e2e.yml) add
a 'Detect newly added spec files' step that queries the GitHub API
for files with status=='added' matching the mock-LLM spec pattern,
avoiding shallow-clone issues with git diff.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: match Playwright basename file paths against repo-relative --new-files

Playwright's JSON reporter emits file paths relative to testDir
(e.g. 'mock-llm-skills.spec.ts') while the GitHub API returns
repo-relative paths (e.g. 'tests/e2e/mock-llm/mock-llm-skills.spec.ts').
The isNewTest() matcher now compares basenames in addition to exact/suffix
matching, so 🆕 badges render correctly.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: stabilize pagination loading-indicator test + improve new-test callout

1. Flaky test fix: the 'loads older events when scrolling up' test
   asserts that the loading-older-events indicator appears, but the
   instant mock response lets React batch isLoading true→false in one
   commit — the DOM element never materialises. Add a 300ms delay to
   older-events mock responses so the indicator renders reliably.

2. Better new-test visibility: replace the subtle inline 🆕 emoji with
   a prominent green blockquote callout above the results table that
   lists each new test with its status icon and spec file.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: address PR review — type safety, docs, test assertions

1. Define BundledSkill interface for buildBundledSkills() return type
   instead of the opaque SettingsRecord[] (review thread #1).

2. Document PUBLIC_SKILLS as an immutable build-time snapshot that is
   baked into the bundle and requires a dependency bump to update
   (review thread #2).

3. Add migration note to buildAgentContext() explaining that the former
   VITE_LOAD_PUBLIC_SKILLS env var was removed because bundled skills
   have no clone latency. load_public_skills: false is still passed to
   tell the SDK to skip its own clone (review thread #3).

4. Add structural assertions for individual skill entries in the adapter
   test: name, content, source, is_agentskills_format, and trigger
   shape (review testing gap).

5. Update stale VITE_LOAD_PUBLIC_SKILLS comments in E2E test files.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: Joe Laverty <joe.laverty@openhands.dev>
Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-07 21:52:59 +00:00
John-Mason P. Shackelfordandopenhands 9df17f3022 fix(tests): remove racy home-route wait in collapsible-thinking snapshot spec (#1201)
fix(tests): stop waiting for home-route text in collapsible-thinking snapshot (#1201)

Collapsible-thinking snapshot tests intermittently timed out at 20 s on
`getByText("Let's start building!")` and never reached the screenshot
assertion. The wait was racy and unnecessary: `navigateToConversation`
goes directly to `/conversations/<CONVERSATION_ID>`, so the home-route
copy only renders for the brief moment before the conversation route
hydrates — and on a slow CI worker that flash can be skipped entirely.

Replace it with a route-stable readiness signal. After
`page.goto(/conversations/<id>)`, wait on
`window.__OH_EVENT_STORE__.getState().loadedConversationId === <id>`
before calling `injectEvents`. That flag flips from null → CONVERSATION_ID
inside `ConversationWebSocketContext`'s `useLayoutEffect`, which is the
same effect that calls `clearEventsForConversation(<id>)`. Observing it
means the clear-and-set has already happened, so any subsequent
`addEvents` will survive the merge (dedup + sort, not replace) instead
of being wiped by a late-arriving clear.

This restores the original intent of the test (inject events into a
mounted conversation route and screenshot the result) without depending
on transient home-route rendering.

Note on `Visual Snapshot Tests`: the 1 remaining baseline diff
(`think-action-collapsed.png`) is not a regression from this change —
it's pre-existing UI drift. The `snapshot-baselines` artifact on `main`
still dates from commit `bbac3b53`, the last green snapshot run before
#1128 introduced the `LlmNotConfiguredBanner` and `SwitchProfileButton`.
Three of the four collapsible-thinking snapshots are only counted as
"unchanged" because `mode: "serial"` skipped them after test 1 failed.
The `update-snapshots` label is applied so the post-merge run on `main`
uploads a fresh baseline reflecting the current UI; future PRs will
compare cleanly against it.

Fixes #1200

Co-authored-by: openhands <openhands@all-hands.dev>
2026-06-06 12:28:45 -04:00