Commit Graph
90 Commits
Author SHA1 Message Date
Tim O'Farrellandopenhands 0206d0e9ee fix(skills): save disabled_skills when server omits the field (#837)
* fix(skills): save disabled_skills when settings has no prior disabled_skills field

The hydration effect in SkillsSettingsScreen only set hasHydratedInitialSettings=true
when settings?.disabled_skills was truthy. When no skills had ever been disabled,
the server omits the field entirely (undefined), so the condition was always false:

  if (settings?.disabled_skills) { ... }   // skipped when field is absent

Because settings?.disabled_skills is undefined both before *and* after settings
load, the dependency array [settings?.disabled_skills] never changed value on
load either, so the effect never ran a second time. hasHydratedInitialSettings
stayed false, and the save effect's early-return guard blocked every toggle.

Fix by gating on settingsLoading instead and defaulting the missing field with
?? []. Also add a localStorage stub for Node.js 25+ which ships a built-in
localStorage that is non-functional without --localstorage-file, breaking the
zustand persist middleware in the test environment.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(tests): use mockResolvedValue(true) for saveSettings spy

saveSettings returns Promise<boolean>, not Promise<void>, so
mockResolvedValue(undefined) caused TS2345 type errors in CI.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(skills): persist disabled_skills to localStorage for local backend

The local agent-server has no concept of disabled_skills and its
PATCH /api/settings strips the field before sending. As a result,
toggling a skill off appeared to work in the UI but was never
persisted -- the value was silently discarded on every save and
page refresh reverted the toggle.

Fix by mirroring the existing app-preferences pattern:

- saveSettings (local path): call writeStoredDisabledSkills before
  stripping the field from the PATCH payload.
- transformApiResponse: call readStoredDisabledSkills and overlay it
  onto the partial Settings returned from the API, so every getSettings
  call (cached or fresh) surfaces the stored value.

Export DISABLED_SKILLS_STORAGE_KEY so tests can reference the key
without duplicating the string literal.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-27 17:49:25 -06:00
Tim O'Farrellandopenhands 463deff24a feat: validate MCP server connectivity before saving (#785) (#816)
* feat: pre-flight MCP test before save (#785)

Validate MCP server connectivity via POST /api/mcp/test before saving
in both the marketplace install modal and the custom-server editor.

- Add McpService.testServer() using MCPClient from @openhands/typescript-client
- Add useTestMcpServer() useMutation hook
- InstallServerModal: test before addMcpServer; show inline error on failure,
  keep modal open (onClose not called); button shows Verifying… then Saving…
- CustomServerEditor: same pre-flight pattern before add/update; also exposes
  a standalone Test Connection button via MCPServerForm's new onTest prop
- MCPServerForm: add onTest/isTestPending/testMessage props; extract buildConfig
  helper; add handleTestClick using formRef; render Test Connection button and
  testMessage inline above action buttons
- Add i18n keys: MCP$TEST_BUTTON, MCP$VERIFYING, MCP$TEST_SUCCESS,
  MCP$TEST_ERROR_TIMEOUT, MCP$TES  MCP$TEST_ERROR_TIMEOUT, MCP$TES  MCP$TEST_ERROR_TIMEOUT, MCP$TES  MCP$TEST_ERROR_TIMEOUT, MCP$TES  MCP$TEST_ERROR_TIMEOUT, MCP$TES  + keeps modal open, success path, Verifying label)

Closes #785

Co-authored-by: openhands <openhands@all-hands.dev>

* test: stub McpService.testServer in tests that save through the pre-flight

Three test suites click a submit button whose handler now runs the
pre-flight connectivity test before calling saveSettings.  None of
them mocked McpService.testServer, so the mutation errored out before
reaching the save step.

Add vi.spyOn(McpService, 'testServer').mockResolvedValue({ ok: true, tools: [] })
to the beforeEach of each affected suite so the test-then-save flow
resolves as expected and the existing assertions remain valid.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: decode HTML entities in MCP test-server error messages

The Python backend passes error strings through html.escape(), so
apostrophes arrive as &#39; and slashes as &#x2F;.  Add a one-liner
decodeHtmlEntities() helper in McpService that uses a temporary
<textarea> to let the browser's own HTML parser decode the string
before it reaches any UI component.

Decoding happens once, at the API boundary, so every consumer
(InstallServerModal, CustomServerEditor, etc.) automatically gets
clean text without needing its own unescaping logic.

Add a focused McpService unit test (4 cases) that mocks MCPClient
via vi.hoisted + vi.mock to exercise the real decoding path:
- success responses pass through unchanged
- &#39; / &#x2F; entities decoded in the error field
- numeric (&lt;) and hex (&#x73;) entities decoded
- stdio config mapped to the correct MCPServerSpec shape

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: render MCP error strings as plain text with newlines

Two bugs in the error message display path:

1. HTML-entity encoding: react-i18next escapes interpolated values by
   default (\' -> &#39;, / -> &#x2F;).  Fix by using the no-escape
   prefix {{-error}} in the MCP$TEST_ERROR_UNKNOWN translation string
   so the raw error text is interpolated verbatim.

2. Newlines ignored: the \n in multi-line Python tracebacks was
   swallowed by the browser.  Fix by adding whitespace-pre-wrap to the
   <p> elements in install-server-modal (globalError) and
   mcp-server-form (testMessage) so \n renders as a visual line break.

Also revert the previous decodeHtmlEntities approach from the service
layer — the backend does not HTML-escape thlayer — the backend does nwas introduced by i18next, not the server.  Update McpService telayer — the backend does not HTML-escape thlayer — the backend dngs in the
translation layer instead.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-27 12:56:25 -06:00
Rohit Malhotraandopenhands 5f5311272f fix: use backend registry apiKey for automation auth instead of build-time env var (#830)
The localAutomationAxios interceptor was reading the session API key from
import.meta.env.VITE_SESSION_API_KEY (baked in at build/publish time),
causing 401 errors for users of the published npm package because their
runtime session key (injected into localStorage by static-server.mjs) was
never included in requests to the automation backend.

The fix reads the API key from getEffectiveLocalBackend().apiKey on every
request, which dynamically resolves the current backend registry entry —
the same source the host/baseURL was already using. This ensures:
- Published npm package users get the runtime-injected key from localStorage
- Users who edit their backend via the Manage Backends UI get their updated key
- No more falling back to Keycloak cookie auth (which always returns 401)

Fixes #829

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-27 14:31:31 -04:00
3d724dd8fd feat(acp): versioned ACP model picker + per-conversation agent chip (#730)
* feat(acp): save concrete model defaults in canvas

* test(acp): verify provider models from subprocesses

* feat(acp): per-conversation agent chip + versioned Claude model picker

User-visible changes:

- Conversation cards now show a single inline chip `[brand mark] {model}`
  on every conversation. ACP cards always render (identity info); OpenHands
  native cards render whenever `agent.llm.model` is present. New
  `AgentBrandIcon` covers Claude / Codex / Gemini brand marks with a
  terminal-glyph fallback; the OpenHands logo is recolored via
  `[&_path]:fill-current` so it inherits the muted-grey chip color
  (the shipped SVG hardcodes `fill="white"`).

- Settings → Agent dropdown for Claude Code lists 10 versioned options
  (Opus 4.7, 4.6, 4.6/1M, 4.5, 4.1; Sonnet 4.6, 4.6/1M, 4.5; Haiku 4.5;
  opusplan). Default switched from `sonnet` to `claude-opus-4-7`.
  Canonical IDs were verified against the bundled `claude` CLI binary's
  model registry — the static SDK `.mjs` shims don't carry the full
  registry. `[1m]` aliases used for 1M-context variants (no canonical
  `claude-*-1m` IDs ship in the SDK).

- `ConversationCardFooter` chip resolves the ACP model string from
  `current_model_name → current_model_id → agent.acp_model → agent.llm.model`
  (the `"acp-managed"` sentinel is filtered out), so the chip works
  whether or not the agent-server populates the SDK runtime fields.

- Removed dead `showLlmProfiles` plumbing on `ConversationCard` /
  `CompactConversationRow` / `ConversationPanel` pass-throughs that
  used to gate the OpenHands model line behind a metadata-menu toggle.
  Chip is now always shown when a model exists.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(acp): adapter cleanup + drop unauthenticated model-check script

- Extract `resolveAcpDisplayModel(info)` from the triple-nested ternary
  in `toAppConversation`. Same precedence chain (runtime name → runtime
  id → configured model → llm.model minus sentinel), just expressed once
  with a loop and named conditions.

- Name the `"acp-managed"` literal as `ACP_MANAGED_SENTINEL` (exported)
  and the `"default" / "default (recommended)"` strings as
  `ACP_DEFAULT_PLACEHOLDERS`. The sentinel was referenced from 7 places
  across source and tests.

- Trim `DirectConversationInfo.agent.kind` JSDoc to 3 lines (was 7).

- Delete `scripts/check-acp-provider-models.mjs` and revert the CI step +
  triggers in `.github/workflows/acp-providers-sync.yml`. The script
  spawned each ACP wrapper unauthenticated and validated Canvas's lists
  against the wrapper's fallback model set — which is only 3 models for
  Claude Code (sonnet / sonnet[1m] / haiku) regardless of what the
  wrapper actually accepts. The check flagged every legitimately-added
  Opus entry as drift. Removing it; followup tracked in #740 for moving
  the lists into `@openhands/typescript-client`.

- Remove `test:acp-models` npm script (only invoked the deleted file).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): plug review-flagged gaps in model resolution + chip path

Five behavioral fixes + two refactors from review feedback on PR #730:

- F1: ``normalizeAgent`` and ``requireDirectConversationInfo`` were
  dropping ``agent.acp_model``, ``current_model_id``, and
  ``current_model_name`` on the way in from the wire — so the chip
  resolver only worked in unit tests (which build ``DirectConversationInfo``
  in-process) and silently fell back to provider-name labels in production.
  Normalizer now preserves all three.

- F2: ChatInputActions gated rendering ChatInputModel on
  ``isCloud || conversation?.agent_kind === "acp"``, so on a local home
  screen with no active conversation it always picked SwitchProfileButton.
  SwitchProfileButton then hid itself for ACP — net result: the home-ACP
  model label never appeared. Added ``isHomeAcp`` derivation from
  ``settings.agent_settings.agent_kind`` so the chat input picks
  ChatInputModel in that case too.

- F3: Switching the preset dropdown to Custom set ``isCustomAcpModel``
  but left ``acpModel`` untouched, so a user moving Claude Code → Custom
  + typing a custom command could save ``acp_model: "claude-opus-4-7"``
  on an unrelated wrapper. Now clears ``acpModel`` on Custom selection.

- F4: ``buildConfiguredAcpAgentSettings`` was stripping null / empty
  ``acp_model`` and not falling back to ``provider.default_model``, so
  existing users with ``acp_model: null`` saw the new registered default
  in Settings → Agent but their next conversation still started with the
  agent-server's own default (UI/runtime mismatch). Conversation creation
  now substitutes the provider default for empty values, matching what
  the form displays.

- F5: ``isAcpDefaultPlaceholder`` was applied only to runtime fields, not
  to the ``configured`` (``agent.acp_model``) or ``sdkLlm``
  (``agent.llm.model``) fallback rungs of the precedence chain. Older
  settings that persisted the literal ``"Default (recommended)"`` could
  surface it on chips. Filter now applies uniformly.

- R1: All five surfaces (Settings form, conversation creation,
  ChatInputModel, conversation adapter, chip) now route through one
  helper, ``resolveEffectiveAcpModel({ runtimeName, runtimeId,
  configured, sdkLlm, providerDefault })`` in ``acp-providers.ts``.
  Placeholder + sentinel filtering live in one place. ``providerDefault``
  is opt-in — chip omits it (don't lie about what's running), settings
  / creation / chat-input pass it (silently substitute the registry
  default).

- R2: ``ACPProviderIcon`` no longer includes ``"openhands"`` — it's
  ACP-only again. Reintroduced ``AgentBrandIconKind = "openhands" |
  ACPProviderIcon`` for surfaces that can render either harness's mark.

Tests: ``acp-server-conversation-service.test.ts`` now exercises the wire
path with the new fields. ``agent-server-adapter.test.ts`` adds two cases
covering placeholder filtering on configured + sdkLlm rungs.
``agent-settings.test.tsx`` adds a Custom-preset-clears-default
regression. ``chat-input-model.test.tsx`` swaps the
"home + ACP renders nothing" assertion for "home + ACP shows the provider
default" — that's the new correct behavior.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): translate new agent-settings keys + harmonize Gemini labels

Addresses the latest review pass:

- Translates ``SETTINGS$AGENT_CUSTOM_MODEL`` and ``SETTINGS$AGENT_MODEL_HINT``
  into all 14 non-English locales (ja, zh-CN, zh-TW, ko-KR, no, it, pt, es,
  ar, fr, tr, de, uk, ca). Both keys previously fell back to English text
  on every non-English client, blocking the model dropdown's hint copy and
  Custom-model label from being legible.

- Harmonizes the Gemini model labels with the Claude / Codex pattern —
  raw IDs like ``gemini-3.1-pro-preview`` become ``Gemini 3.1 Pro
  (preview)`` so the three providers read consistently in the dropdown.

- Adds provenance notes to ``CODEX_MODELS`` and ``GEMINI_MODELS`` mirroring
  the Claude one — naming the upstream source the list was extracted from
  and pointing at agent-canvas#740 (the long-term "move ACP model lists to
  ``@openhands/typescript-client``" plan).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): address review — chip labels, dead toggle, icon de-dup, shared hook

- Chip shows the provider's picker label (e.g. "Claude Opus 4.7") instead of
  the raw acp_model ID via new labelForAcpModel(); falls back to the raw ID
  for custom overrides and the provider name when no model. (#1)
- Drop the stale version from the 1M labels (opus[1m]/sonnet[1m] →
  "Claude Opus (1M)" / "Claude Sonnet (1M)") so the version-agnostic alias
  and its label can't disagree. (#2)
- Remove the now no-op "Show LLM profiles" filter-menu row + its wiring
  (chip is unconditional now; the preference no longer affects cards). (#4)
- Onboarding AgentOptionIcon delegates to AgentBrandIcon; delete the
  duplicated brand-mark path constants and per-kind SVG markup. (#5)
- Name the OpenHands logo aspect ratio (3:2) so the chip and onboarding
  tile render identically. (#7)
- Extract useAcpModelContext() shared by chat-input-actions and
  chat-input-model to kill the duplicated isHomeAcp / destination-path /
  label logic. (#6)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): gate the agent chip behind the restored "LLM model" toggle

The chip was unconditional, which (a) silently flipped OpenHands cards from
main's model-hidden-by-default to always-shown and (b) left the panel's
"LLM model" toggle a no-op (then deleted). Restore one toggle that gates the
unified chip for BOTH ACP and OpenHands cards (default OFF), matching the
sibling "Repo and branch" metadata toggle.

- Footer: chip (ACP brand mark + model, OpenHands logo + model) now renders
  only when showAgentChip is set.
- Restore showLlmProfiles plumbing: filter-menu row + ConversationPanel/
  ConversationCard/CompactConversationRow pass-throughs (panel + filter-menu
  are now byte-identical to main).
- ConversationCard.shouldRenderFooter gates ACP under the toggle too.

Net effect vs main for OpenHands: identical by default (hidden); when the
toggle is on, the only change is the OpenHands logo added next to the model
name. ACP identity chip becomes opt-in via the same toggle.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): reconcile model on command edit, unify chip labels, dedupe provider lookup

Review follow-ups for the versioned ACP model picker + agent chip:

- agent-settings: editing the command textarea into a different provider
  (or a custom command) now reconciles the model selector, so Save can no
  longer silently persist e.g. claude-opus-4-7 against a Codex/custom
  wrapper. The preset dropdown already did this; the textarea is the other
  way a user switches providers. Also always show the custom-model input
  when the dropdown is on "Custom" so a mismatched value is visible/editable
  rather than hidden.
- chat input: surface the provider's human label (e.g. "Claude Opus 4.7")
  for ACP conversations, matching the conversation-list chip, instead of the
  raw acp_model id.
- relabel the conversation-panel "LLM model" toggle to "Agent / model" — it
  now governs the ACP brand chip too.
- extract getAcpProvider() and replace the repeated ACP_PROVIDERS.find()
  lookups across the adapter, settings, and the constants module.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): stop filling the OpenHands logo's transparent hand paths

AgentBrandIcon recolored the OpenHands logo with a blanket
``[&_path]:fill-current``, which also overrode the two ``fill="transparent"``
hand shapes — turning the mark into a solid white blob on the onboarding tile
(and a filled blob on the conversation chip). The logo asset has 5
``fill="white"`` wordmark paths and 2 ``fill="transparent"`` hands; only the
former should inherit ``currentColor``.

Scope the override to ``[&_path:not([fill=transparent])]:fill-current`` so the
hands stay transparent (negative space), restoring the original two-tone logo.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Debug Agent <debug@example.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 17:54:54 +02:00
e49f6d721e feat(chat): attachment UX, upload-as-file, and home/cloud submit fixes (#712)
* feat(chat): merge tools menu into plus button with file upload footer

Combine the chat tools dropdown with the + control and add an
Add Files and Images action with a paperclip icon at the bottom.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(chat): support clipboard image paste on home and conversation inputs

Read pasted screenshots from clipboard items, wire the home launcher
through the shared attachment upload flow, and send first messages with
attachments after creating a conversation.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(chat): per-pasted-image upload-as-file control on thumbnails

Replace the global checkbox with a circular upload button on clipboard
pasted images, track paste source in the store, and fix overlay stacking.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): polish pasted-image upload toggle styling and tooltips

Match the remove button size and corner inset, show a checkmark when
active with hover feedback, and swap the tooltip to "Do not upload as file".

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): add bottom padding below attachment thumbnails row

Match the chat input container's top inset so pasted images and files
are spaced evenly above and below the thumbnail strip.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): use surface grey for upload-as-file toggle and nudge position

Replace invalid primary tokens with oh-surface/oh-muted styling for both
states and raise the button slightly from the thumbnail corner.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(home): defer attachment sends until cloud start task is ready

Cloud conversation creation can return a provisional `task-{uuid}` URL
while the sandbox provisions. Sending messages or uploading files against
that id caused 422 UUID parsing errors on the home `/conversations` input.

- Queue attachments in memory keyed by start-task id when provisioning
- Flush uploads and the user message once `useTaskPolling` sees READY
- Always include typed text in the start request, even with attachments
- Add store and flush helper tests

Paste/drag on the home input were already wired; this fixes starting a
conversation with attachments (or text + attachments) on cloud backends.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Revert "fix(home): defer attachment sends until cloud start task is ready"

This reverts commit 9a443088ca5b14613c72b60459e26f1d8fb0f545.

* Reapply "fix(home): defer attachment sends until cloud start task is ready"

This reverts commit 7ba6a267cedaa67dbf08948ddcd00b9c5d6e269c.

* fix(chat): upload attachments into the conversation workspace

File uploads targeted read-only /workspace in Docker dev stacks and
cloud task flush used raw i18next before app init. Resolve the
conversation working_dir for upload paths and use the initialized
openhands i18n instance when flushing deferred attachments.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cloud): route home attachments through the runtime sandbox

Cloud file uploads and the first message were hitting the bundled
local agent-server and 404ing. Defer home attachments until the start
task is ready, upload via the provisioned runtime URL, and send events
through the cloud proxy.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): send home attachments in a single first message

Skip initial_message when starting with attachments and avoid enqueueing
optimistic duplicates after the attachment send already persisted the message.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cloud): show home submit as first message during task provisioning

Enqueue optimistic pending messages on cloud start-task routes so the chat
shows the user's message while the sandbox provisions, hide empty-state
suggestions during provisioning, and reassign pending bubbles to the real
conversation id when the task is ready.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): suppress spurious older-messages errors on new conversations

Skip auto-pagination when the initial history page is complete, on cloud
start-task routes, or while provisioning, and stop surfacing an error
banner when older events cannot be anchored.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): hide history skeleton when pending message is visible

Treat optimistic home-submit bubbles as loaded content so the feed
skeleton does not flash over the user's first message on new conversations.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cloud): stop empty-state suggestions flashing after home submit

Keep pending bubbles linked across task-to-conversation redirect, reassign
them before paint on the real route, and tighten suggestion gating so the
Let's Start Building overlay does not flash during provisioning.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): show upload-as-file toggle on all attached images

Mark every attached image for the per-image upload control, not just
clipboard pastes, so file picker and drag-and-drop previews match pasted
screenshot behavior.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): use file-plus icon for upload-as-file toggle

Replace the generic upload glyph with Lucide FilePlus so the per-image
control reads more clearly as "add this image as a workspace file."

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor: remove unrelated files

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: hieptl <hieptl.developer@gmail.com>
2026-05-24 17:13:10 +07:00
9626fc92ab Customize extensions UI and polish automations experience (#739)
* fix(ui): move skills type filters into search row dropdown

Put skill type filters in a dropdown beside full-width search and keep
the result count below. Add SkillsTypeFilterDropdown and update tests
to open the menu before selecting a filter option.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): full-width MCP toolbar with section filter

Replace the half-width MCP search with a toolbar matching skills: flex search plus Installed/Library section filter. Remove the skills list count row for a cleaner toolbar layout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(ui): unify skills and MCP extension cards with detail modal and add toggles

Polish the Skills and MCP pages with shared card chrome, circle plus/check
install toggles, and richer skill inspection. Skills open a detail modal with
metadata pills and a traditional enable switch; MCP library and installed cards
use the same toggle and white-outline hover treatment, with installed servers
opening the editor on card click.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): drop hover ring glow from extension module cards

Remove the white ring on hover/focus so skills and MCP tiles keep the border
highlight without the outer glow.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): align installed MCP cards with marketplace module layout

Match the library tile structure: transport under the title, command or URL
on a bottom grey detail line, and the same min-height and spacing.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(ui): add enable/remove tooltips and hover X on circle toggles

Show Enable on the plus state and swap the checkmark to an X on hover when
selected, with a Remove tooltip, across skills and MCP add toggles.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): show catalog description on installed MCP server cards

Restore the marketplace description text above the command detail line so
installed tiles match the library card content layout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): clamp installed MCP descriptions and grey the command line

Limit catalog descriptions to two lines and style the command/URL metadata
row with tertiary-alt so only that detail line reads as muted grey.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Thicken remove X and add destructive hover on toggle.

Use stroke-based x-mark at 14px to match checkmark weight, with a soft red background when hovering to remove.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Use cool grey styling for Auto-discovery skill type badge.

Replace yellow/gold agentskills pill colors with neutral cool-grey tones aligned to the design system.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add top-right close buttons to skill and MCP modals.

Introduce a shared ModalCloseButton and wire it into the skill detail, MCP install, and custom server editor panels.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Match MCP custom editor modal width to other modals.

Use 520px panel width so the edit form aligns with the skill detail and install modals.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Disable resizing on skill detail content textarea.

Use resize-none on the readonly content field so the modal layout stays fixed.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Move skill detail enable toggle below header row.

Place the switch in a full-width settings bar so it no longer collides with the top-right close button.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add delete with confirmation to MCP custom editor.

Show a destructive footer action in edit mode and reuse the shared confirmation modal before removing the server.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Style MCP editor delete as secondary button with trash icon.

Match the bordered secondary treatment used by Cancel while keeping a clear delete affordance via Trash2.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Keep marketplace MCP cards in add-only toggle state.

Library tiles always show the plus affordance since users may install multiple instances of the same server template.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Use solid border on MCP installed empty states.

Replace dashed outlines on the installed-servers empty and no-results panels with a standard border.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Use shared Dropdown for MCP and skills filter controls.

Replace bespoke filter menus with EnumFilterDropdown so toolbar filters match the site-wide combobox styling and behavior.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Use tertiary-alt grey for MCP and skills secondary copy.

Align card descriptions, modal hints, and empty-state text with the muted tone used for paths and subtitles elsewhere.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Use tertiary-light for MCP and skills descriptive copy.

Match card descriptions, modal body text, and empty states to the page header subtitle tone.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Align MCP empty-state hint with tertiary-light copy tone.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Make filter dropdown triggers hug their label width.

Add fitContent mode to the shared Dropdown and use it for MCP/skills filters instead of a fixed 9rem trigger.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Match filter dropdown trigger styling to toolbar search fields.

Use rounded-lg, base-secondary background, and shared border/focus treatment; keep the active-filter highlight when not on All.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Replace filter combobox with a simple right-aligned menu.

Use a button-only toolbar filter with a menu anchored to the trigger's right edge and no type-to-search input.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add shared constants for the skills installation docs link.

Centralize the OpenHands adding-skills guide URL and example /add-skill command so the skills UI can reference them without inline literals.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add an Add skill button and instructions modal on the skills page.

Opens a fixed-header modal that explains using /add-skill in chat, highlights key paths with inline code chips, provides a copyable example command, and links to the OpenHands adding-skills documentation.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Keep disabled skill cards at full opacity on the skills page.

Remove the card-wide opacity-70 dimming so only the enable toggle reflects disabled state, matching clearer card readability when a skill is turned off.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add Use skill actions that open a new chat with a prefilled command.

Share launch logic across the add-skill and skill detail modals, disable Use skill when a skill is off, and use the message-square-share icon for the primary action.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Match recommended automation card hover styling to skills and MCP tiles.

Reuse the shared extension module card surface and interactive hover classes so automation cards get the same border and background lift on hover.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Use the same monochrome styling for all skill type badges.

Trigger-based and Always active now share the cool grey Auto-discovery treatment instead of blue and green accents.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Use container queries so Skills and MCP card grids collapse in narrow columns.

Switch the shared grid from viewport md breakpoints to a 600px column-width container query so cards stay single-column when the extensions content area is tight, even on desktop layouts with the sidebar visible.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Keep Skills extension pills on a single nowrap row with overflow.

Harden pill chip styling and reuse SkillCardPillRow in the detail modal so badges never wrap to a second line in narrow cards or modals.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Hide skill copy actions when the source is a scope label, not a path.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Centralize extension module pill styling for reuse across Skills UI.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add MCP logo stack badge for multi-server automation card icons.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Align recommended automation cards with Skills layout and shared pill row.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Match automation list section headings to MCP Installed typography.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Restore visible hover feedback on saved automation cards.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix extension module card hover using valid tokens and a visible ring.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Align saved automation cards with shared extension module surface chrome.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add a decorative plus badge to recommended automation cards.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Limit extension card hover to a border highlight without glow or fill shift.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Match saved automation card typography and pills to recommended cards.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Make the automations search field span the full content width.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add an Add Automation header button that opens create instructions in a modal.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add hover and tooltip to recommended automation plus badges.

Match the unselected CirclePlusCheckToggle affordance so decorative plus icons highlight on card hover and show a launch tooltip.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add a shared bordered empty-state panel class for extension pages.

Centralize the rounded border/padding chrome used by MCP Installed and Automations list empty states.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Redesign the automations empty state to match MCP installed styling.

Use the shared bordered panel, drop the database icon, and add a secondary hint pointing users to the recommendations below.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Align MCP installed empty state with shared panel and white title text.

Reuse the shared empty-state container and render the primary line in white to match the automations empty state.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Use the shared modal close button in Add Automation modal.

Switch to ModalBackdrop and ModalCloseButton so the header and dismiss control match Skills and MCP modals.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add dispatchAutomation API for manual automation runs.

Expose POST /api/automation/v1/{id}/dispatch for local and cloud backends with service-layer tests.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Mock automation dispatch endpoint for manual runs.

Return a pending run from POST /api/automation/v1/:id/dispatch in MSW with handler coverage.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add useDispatchAutomation mutation hook.

Invalidate automation list, detail, and run history queries after a successful manual dispatch.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Wire manual Run now dispatch through the automations list page.

Pass dispatch handlers and pending state from the list route into automation groups and cards.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add i18n keys for automation card Run now and actions menu.

Introduce AUTOMATIONS$RUN_NOW and AUTOMATIONS$ACTIONS_MENU for the saved automation card controls.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Align automation kebab menu with shared context menu styling.

Use ContextMenu rows, muted-to-foreground icon hover tokens, portal rendering to avoid overflow clipping, and drop the red delete styling.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Replace saved automation card toggle with play Run now and menu actions.

Add a play trigger, move Run now and View into the kebab menu with Lucide FileText, and keep turn on/off and delete under manage permissions.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Show visible Run now label on saved automation cards.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Tighten spacing between saved automation card title and description.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Limit recommended automation plus badge hover to direct hover only.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Set recommended automation plus badge tooltip to Add Automation.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add grid and list view toggle to the automations page.

Let users switch saved automations between cards and table rows from a toggle beside search, with the choice persisted locally.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Polish automations list view and replace view toggle with icon dropdown.

Use secrets-style table surfaces and row hover, place pills beside the title, and swap the segmented control for a 36px icon menu without a caret.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Show a checkmark on the selected automations view mode in the dropdown.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Polish automations list layout, toolbar, and table row standards.

Use two-column card grids, align search and view controls with skills/MCP toolbars, add shared compact table row sizing with lighter hover, and square icon actions with a Run now tooltip.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Remove resting borders from extension module cards.

Skills, MCP, and automation tiles now stay borderless until hover or keyboard focus, with shared CSS handling click focus without a stuck highlight.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Match automation card Run now button height to the kebab menu.

Use a shared h-8 text button class so grid card actions align with the 32px three-dot control.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Tighten automation detail section header padding.

Use symmetric py-3 on Configuration and Activity Log headers to match other modal section bars.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Use theme tokens for skill type badge pill colors.

Hardcoded cool-grey rgba values did not follow the active color theme; switch skill type badges to semantic text-secondary opacity tokens only.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix skill toggle hover state and selected check styling.

Use pointer hover instead of focus for the remove icon swap, show a white outline check when enabled, and let skill cards use a Disable tooltip.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Show MCP stdio transport labels as STDIO in card headers.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Remove Use skill action from the add skill modal.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Redesign add-automation instructions and place chat caret after prefilled prompts.

Replace the modal's plugin/conversation cards with inline guidance and a launch-in-chat action, and focus prefilled messages at the end for automation and skill flows.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor: remove unrelated file

* fix: failing tests

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: hieptl <hieptl.developer@gmail.com>
2026-05-24 01:06:30 +07:00
Graham Neubigandneubig 56fcdc6b2e Remove direct HttpClient usage from agent-server APIs (#743)
* Show workspace version errors in canvas

* Use current agent server version in mocks

* Pin merged typescript client dependency

* Use typescript client v1.23 release

* Clarify typescript client release pinning guidance

* Make snapshot baseline wait non-fatal on API errors

* Remove direct HttpClient usage for agent server APIs

* Validate LLM profile config before switching

---------

Co-authored-by: neubig <398875+neubig@users.noreply.github.com>
2026-05-23 13:07:42 +00:00
Graham Neubigandneubig 8cab54e0b1 Show agent-server version errors for workspaces (#742)
* Show workspace version errors in canvas

* Use current agent server version in mocks

* Pin merged typescript client dependency

* Use typescript client v1.23 release

* Clarify typescript client release pinning guidance

---------

Co-authored-by: neubig <398875+neubig@users.noreply.github.com>
2026-05-23 12:53:53 +00:00
Graham Neubigandopenhands 8df546950c Use SDK agent settings and model switch UI (#457)
* Use SDK agent settings for local conversations

* chore: address PR review feedback (#457)

Co-authored-by: openhands <openhands@all-hands.dev>

* Show switch LLM tool results in model UI

* chore: fix model switch formatting

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-22 22:23:39 -04:00
2e3f8b64dd Add automation LLM profile controls (#536)
* Add automation LLM profile controls

Co-authored-by: openhands <openhands@all-hands.dev>

* Add automation LLM profile QA capture

Co-authored-by: openhands <openhands@all-hands.dev>

* Make automation LLM profile read-only

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: Remove PR-only artifacts

* test: cover automation manual dispatch

Co-authored-by: openhands <openhands@all-hands.dev>

* Display automation LLM profile names

Co-authored-by: openhands <openhands@all-hands.dev>

* Treat automation model as profile name

Co-authored-by: openhands <openhands@all-hands.dev>

* Use model field for automation profiles

Co-authored-by: openhands <openhands@all-hands.dev>

* Use model profile wording in automation test

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: allhands-bot <allhands-bot@users.noreply.github.com>
2026-05-23 01:45:33 +00:00
Vasco Schiavoandopenhands 1c5be9f906 chore: change default model to MiniMax-M2.7 (#697)
* chore: change default model to MiniMax-M2.7

- Update DEFAULT_SETTINGS.llm_model to openhands/minimax-m2.7
- Update agent_settings.llm.model to openhands/minimax-m2.7

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: update test assertion and mock for new default model

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-21 12:50:01 +00:00
Hiep Le b3d5a61fe4 feat(frontend): persist local workspaces on the agent-server (#653)
* feat: persist local workspaces on the agent-server

* refactor: update the code based on feedback

* fix: lint
2026-05-21 13:29:14 +07:00
Rohit Malhotra eee7fe1b7a feat: use async conversations and interrupt endpoint for local mode (#670) 2026-05-20 22:16:20 -04:00
2aa262acdc feat(acp): show ACP agent name on conversation cards (#658)
The ``acpserver`` conversation tag has been stamped at create time since
the ACP integration landed, but the read path silently dropped it — so
the sidebar had no way to tell a Claude-Code conversation from a Codex
one. Plumb the tag through ``DirectConversationInfo`` →
``AppConversation`` → ``ConversationCard`` / ``CompactConversationRow``
and render a small pill above the LLM-model line. Closes the
"conversation list with agent tags" half of agent-canvas#405.

Co-authored-by: Debug Agent <debug@example.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 15:41:27 +00:00
a4534c15f4 feat: add agent delegation settings page (#483)
* feat: add agent delegation settings page

Migrate sub-agent delegation UI from OpenHands/OpenHands#14418.

- Add Settings > Agent page with an Enable sub-agents toggle backed by
  agent_settings_diff
- Add /settings/agent route and Agent nav item with robot icon
- Add enable_sub_agents field to mock agent settings schema (general
  section) and default settings
- Add SCHEMA$ENABLE_SUB_AGENTS$DESCRIPTION and
  SCHEMA$ENABLE_SUB_AGENTS$LABEL i18n translations for all supported
  languages
- Add enable_sub_agents: false to DEFAULT_SETTINGS agent_settings

Co-authored-by: openhands <openhands@all-hands.dev>

* docs: update AGENTS.md with agent delegation settings context

Co-authored-by: openhands <openhands@all-hands.dev>

* refactor: align agent delegation settings page with repo conventions

* fix(agent-tools): gate task_tool_set on enable_sub_agents setting

* chore: change icon

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: Vasco Schiavo <115561717+VascoSch92@users.noreply.github.com>
Co-authored-by: VascoSch92 <vasco.schiavo@protonmail.com>
2026-05-19 21:39:38 +00:00
Rohit Malhotraandopenhands edb998220a Remove Docker dependency from dev workflow (#635)
- Delete scripts/dev-docker.mjs and its test
- Simplify package.json: 'npm run dev' now runs local uvx stack directly
  (agent-server + automation + Vite + ingress), no Docker needed
- Remove dev:docker, dev:docker:dynamic, dev:dangerously-dockerless scripts
- Add dev:static for production-build frontend variant
- Update bin/agent-canvas.mjs CLI to use uvx-based stack
- Rename Docker-specific variables: DOCKER_PROJECTS_PATH → PROJECTS_PATH,
  shouldDefaultToDockerProjects → shouldDefaultToProjectsPath
- Update i18n: HOST_HOME_NOT_MOUNTED_HINT no longer references Docker
- Update all docs (README, DEVELOPMENT, SELF_HOSTING, AGENTS.md, CHANGELOG)
- Rename e2e snapshot: docker-workspace-browser → projects-workspace-browser
- Fix all tests to reflect new script names and remove Docker references

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-19 15:27:25 -04:00
f0c36bac8f feat(acp): Settings → Agent + onboarding + chat-UI gating for ACP-driven conversations (#416)
* feat(acp): add minimal ACP agent UI (parity with OpenHands#14401)

Adds a Settings → Agent page so users can switch the conversation
between the built-in OpenHands agent and an external ACP (Agent Client
Protocol) subprocess (Claude Code, Codex, Gemini CLI, or custom command)
without hand-editing settings.

Discriminates in agent-server-adapter: when `agent_settings.agent_kind
=== "acp"`, build an `ACPAgent` payload (kind, acp_command, acp_model)
instead of the LLM-shaped Agent, and skip the LLM defaults that would
otherwise be rejected as extras. Stamps the provider key onto
`tags.acpserver` so the chip can resolve a brand name from a single
source.

Tag-key constant note: the conventional `acp_server` form is invalid —
agent-server validates tag keys against `^[a-z0-9]+$` and returns 422.
The flattened `acpserver` form survives validation; the named constant
`ACP_SERVER_TAG_KEY` keeps the regex and the key colocated.

Gates the LLM and Condenser nav items behind a `disabledByAcp` flag,
greys them out with a tooltip, and redirects to `/settings/agent` in
the settings loader (not a per-route useEffect, so there is no one-
frame flash of the LLM page before bouncing).

E2E validated against `ghcr.io/openhands/agent-server:fa29ae2-python`:
- PATCH /api/settings with `agent_kind: "acp"` round-trips
- POST /api/conversations with the adapter's ACP payload returns 201,
  `agent.kind=ACPAgent`, `acp_command` preserved, tags stamped.

Closes #412

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(acp): wire onboarding ChooseAgent step into ACP settings

Drops the "Support for other agents coming soon!" banner now that
the support exists. Enables the Claude Code / Codex tiles and adds a
Gemini CLI tile so the four options here match ``ACP_PROVIDERS`` from
the Settings → Agent page.

Selecting an ACP option and clicking Next persists ``agent_kind:"acp"``
plus the registry provider key (``acp_server``) via ``useSaveSettings``,
mirroring the diff the Settings page emits. The advance only happens
on save success — a failed PATCH stays on the step and surfaces a toast.

Skips the embedded LLM-setup step (index 2) on both forward and back
navigation when an ACP agent is active: the subprocess owns its own
LLM and authenticates through Secrets, so the form has nothing to
configure. OpenHands path is untouched.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(acp): seamless Claude Code + Codex CLI auth via dev-docker

Live-validated against `ghcr.io/openhands/agent-server:1.22.1-python`
(the canvas's default pin, which now ships ACPAgent natively — no
SHA override needed). Two changes surfaced by the run:

1. **Mount `~/.claude.json` in dev:docker.**  Recent Claude Code CLI
   versions persist auth + workspace state in `~/.claude.json` next
   to (not inside) `~/.claude/`.  Without this single-file mount,
   `@agentclientprotocol/claude-agent-acp` can't see the user's
   existing login and prompts to re-auth inside the sandbox.

2. **Use the new ACP package name in `ACP_PROVIDERS`.**  Upstream
   renamed `@zed-industries/claude-code-acp` → `@agentclientprotocol/
   claude-agent-acp`.  The old name still works but emits an npm
   deprecation warning; the agent-server's own OpenAPI example uses
   the new name.  Test fixtures pinning the legacy name updated to
   match.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(acp): import ACP_PROVIDERS from typescript-client

Canvas was carrying its own copy of the ACP provider registry, which
drifted out of sync with the canonical Python SDK source and ended up
encoding an invalid Codex invocation (``@openai/codex acp`` — codex
CLI has no ``acp`` subcommand, so the spawn deadlocked silently with
``Error: stdin is not a terminal`` and no log line).

This change deletes ``src/constants/acp-providers.ts`` and imports the
registry from ``@openhands/typescript-client`` instead, which now
mirrors the Python SDK (see OpenHands/typescript-client#167). The TS
SDK pin in ``package.json`` is bumped to the PR-branch SHA
(``45a803c``) for now; once #167 merges and a new tagged release is
cut, the pin can flip to the tag in a follow-up commit.

Shape changes consumers needed to absorb:
- ``ACPProviderConfig[]`` → ``Record<string, ACPProviderInfo>``
  (lookup by key replaces ``.find``; ``Object.values`` where an array
  is needed)
- ``display_name`` → ``displayName`` (camelCase matches TS conventions)
- ``default_command`` → ``defaultCommand`` (and now ``readonly string[]``;
  components spread into a fresh array before passing to consumers that
  expect mutability)

``ACP_CUSTOM_PRESET_KEY`` is the only ACP-related constant that stays
canvas-local — it's a synthetic sentinel for the "Custom" dropdown
option, not a real provider, so it has no SDK counterpart. Moved to
``src/constants/acp-presets.ts``.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* revert(acp): keep ACP_PROVIDERS local to canvas

Reverts the brief detour through `@openhands/typescript-client` for
the ACP provider registry.  Splitting the registry across two repos
adds publish-coordination friction and doesn't actually eliminate
the drift problem — it just moves it from
"canvas vs. python-sdk" to "ts-sdk vs. python-sdk", with extra steps.

Now:

- `src/constants/acp-providers.ts` is the canvas-local copy again,
  with the corrected `codex` command (`@zed-industries/codex-acp`,
  the real ACP-protocol stdio server — not `@openai/codex acp`,
  which is the codex CLI's interactive mode and deadlocks the agent
  handshake when spawned without a TTY).
- The package.json pin reverts to `v0.6.0` (the typescript-client
  release that does not include the unmerged `ACP_PROVIDERS` export
  from #167, which is now closed).
- The split `acp-presets.ts` file is folded back in.

Drift risk between this file and the Python SDK source is tracked
in #587, with a longer-term plan to address it (TS-SDK mirror,
code-gen, or runtime endpoint).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): resolve empty acp_command from registry in adapter

PR #416 ships the Settings → Agent page (and onboarding) with a
"default preset" shortcut that stores ``acp_command: []`` and trusts
the agent-server to resolve it from ``acp_server``.  It doesn't.

The agent-server's ``ACPAgent`` model has no ``acp_server`` field and
no registry resolution — it just hands ``acp_command`` straight to a
subprocess spawn.  Empty list trips ``acp_agent.py:1013`` with
``IndexError: list index out of range``, the agent loop dies silently
inside the agent-server's run thread, and the conversation hangs in
``idle`` with the user's message persisted but never answered.  No
error reaches the UI; from the user's perspective they sent a message
and nothing happened.

Caught while exercising the live ``dev:safe`` stack: a fresh
conversation seeded from the onboarding "Claude Code" tile produced
``Failed to start ACP server: list / IndexError: list index out of
range`` in the agent-server log.

The fix is purely client-side — expand ``acp_command`` against
``ACP_PROVIDERS`` (canvas's local mirror of the Python SDK registry,
see #587) before the payload leaves the adapter, when the user picked
a built-in preset.  ``acp_server: "custom"`` and any unknown key are
left untouched — those genuinely depend on the user's explicit
command, and silently inventing one would mask a real config bug.

Three new adapter tests cover:
- ``acp_command: []`` + ``acp_server: "claude-code"`` → command
  resolved to ``["npx","-y","@agentclientprotocol/claude-agent-acp"]``
- ``acp_command`` omitted entirely + ``acp_server: "codex"`` → same
  resolution path
- ``acp_command: []`` + ``acp_server: "custom"`` → left untouched

2273 tests pass, lint + typecheck clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): pass disabled state into SettingsDesktopSidebar

When ACP is the active agent, ``useSettingsNavItems`` correctly tags
the LLM and Condenser entries with ``disabled: true``.  The mobile
drawer (rendered via ``SettingsNavLink``) already respected that.
The desktop sidebar (rendered via ``SidebarNavLink``, came in with
the recent sidebar refactor) was constructing the link without
forwarding the flag, so both items stayed fully clickable / styled
as enabled while the conversation was running on an ACP subprocess.

Two tiny changes:

1. ``SettingsDesktopSidebar`` passes ``renderedItem.disabled`` through
   to ``SidebarNavLink``.  That alone gives the right visual state
   (``opacity-50``, ``pointer-events-none``) and keyboard behaviour
   (``tabIndex=-1`` + ``onClick preventDefault``) — both already
   implemented by ``SidebarNavLink``.
2. ``SidebarNavLink`` additionally sets ``aria-disabled="true"`` when
   ``disabled``, closing a screen-reader gap that existed independently
   of this regression (the link sounded actionable to assistive tech
   even though it wasn't).

The ``clientLoader`` redirect in ``routes/settings.tsx`` continues to
handle direct URL navigation to a disabled-by-ACP page, so even if
someone bookmarks ``/settings/condenser`` and lands there while ACP
is active, they get bounced to ``/settings/agent``.

Two new tests in ``settings-navigation.test.tsx``:
- Disabled-by-ACP items in the desktop sidebar carry ``aria-disabled``.
- Enabled items don't.

2275 tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): address PR #416 review feedback

Addresses both human and all-hands-bot review comments on #416:

**Critical bugs fixed**

- ``acp_args`` duplication on load (bot critical #1): the textarea
  is the single source of truth for the launch tokens, but save only
  wrote ``acp_command`` — any API-set ``acp_args`` survived and
  concatenated at spawn time. Save now always writes ``acp_args: []``.

- ``tokenizeCommand`` corrupted quoted Custom commands (human bug #2):
  ``bash -c "echo hello"`` got split into
  ``["bash","-c","\"echo","hello\""]`` and silently misbehaved. New
  ``src/utils/acp-command.ts`` wraps ``shell-quote`` with selective
  re-quoting (so ``npx -y @org/pkg`` renders verbatim, not ``\@org/pkg``)
  and filters non-string entries (redirects, env-var refs) out of the
  parsed argv. Round-trip tests pin the contract.

- Loader/component settings cache mismatch (bot critical #4): loader
  used ``SETTINGS_QUERY_KEYS.byScope("personal")``; ``useSettings``
  used ``[...byScope("personal"), backend.id, orgId]``. They didn't
  share cache. Aligned + set ``staleTime: 0`` on the loader read so
  cross-tab kind flips are picked up immediately (the in-render hook
  keeps its 5-minute stale window).

- ``getFirstAvailablePath`` ignored the new agent route (human bug #3):
  ``/settings/agent`` now precedes the others in the fallback list,
  so first-time / hide_llm_settings users land on the agent picker
  rather than ``/settings/app``.

- ``ACP_SETTINGS_KEYS`` documentation (human #4): pre-empts the
  "why not trim this list to UI-visible fields" question by spelling
  out that it serves as both the ACP allow-list and the OpenHands
  deny-list — trimming would silently leak API-set ``acp_*`` state.

**Refactor (human #2 + #3)**

- ``description_key`` moves into ``ACP_PROVIDERS``; the onboarding
  ``AGENT_OPTIONS`` is now derived from the registry so adding a new
  provider only needs one edit.
- One ``buildAcpAgentSettingsDiff`` helper replaces the two near-copies
  in ``choose-agent-step.tsx`` and ``agent-settings.tsx``; both call
  sites are now under a single contract for the agent_settings_diff
  shape.

**UX (bot)**

- Onboarding progress bar shows the actual visited-step count when
  the LLM step is skipped (3 segments for ACP, 4 for OpenHands).
  Previously segment 2 popped "completed" on a slide the user never
  visited.

**Test coverage (bot)**

- New ``__tests__/utils/acp-command.test.ts`` covers parseCommand /
  formatCommand round-trips, quoted args, embedded escapes, shell-
  operator filtering, package-style tokens.
- Adapter: empty ``acp_model: ""``, unknown ``acp_server`` key,
  ACP→OH→ACP round trip (no field leakage either direction).
- agent-settings: cleared input keeps Save disabled, whitespace-only
  same, full Custom command with quoted args round-trips through
  shell-quote.
- choose-agent-step: provider switching (claude-code → codex) rebuilds
  the diff cleanly, no leak from the prior selection.

**Acknowledged (no action)**

- Bot critical #2 (supply chain drift) — same problem as the existing
  agent-canvas#587, already tracked.
- Bot critical #3 (desktop sidebar disabled) — fixed in 27a3e79 a few
  commits before this review was written; review snapshot was stale.
- Translation duplication (human #1) — matches the existing
  ``translation.json`` convention (every key has all 15 locales).
- Option-bag → split functions (human #5) — cosmetic; defer.

2292 tests pass, lint + typecheck clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): pre-bundle shell-quote so the Vite dev server can load it

``shell-quote`` is a CommonJS module that does ``module.exports = {
parse, quote }``. The previous commit wired it into ``src/utils/
acp-command.ts`` with a named ESM import, which the dev server
rejected on the first ``agent-settings.tsx`` load:

    SyntaxError: The requested module '/node_modules/shell-quote/
    index.js?v=...' does not provide an export named 'parse'

Switching to a namespace import (``import * as shellQuote from
"shell-quote"; const { parse, quote } = shellQuote;``) makes the
named-export check pass, but Vite then served the raw CJS file
to the browser unchanged and the next request died with:

    ReferenceError: exports is not defined

This second failure is because ``vite.config.ts`` sets
``optimizeDeps.noDiscovery: true`` — new dependencies must be listed
in ``optimizeDeps.include`` or Vite won't run them through its
CJS-to-ESM prebundler. Adding ``"shell-quote"`` there fixes it; the
existing entry has a comment block explaining the same constraint
for other deps. The Rollup-based prod build was unaffected.

Verified: dev server boots clean, ``GET /settings/agent`` returns
200, no ``exports is not defined`` in the Vite client log, 9 unit
tests in ``__tests__/utils/acp-command.test.ts`` pass on the Node
test runner (vitest) where the CJS interop already worked.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): address second-pass review on PR #416

Addresses the second all-hands-bot review's critical + improvements:

**Critical: load path mis-merged acp_command + acp_args**

Settings stored with the registry-default shortcut (``acp_command:
[]``, ``acp_server: "claude-code"``) plus a non-empty ``acp_args``
showed only the args in the textarea — no registry prefix. Saving
then sent ``acp_command: ["--extra-arg"]`` and flipped the preset
to ``custom``, silently losing the ``npx -y @agentclientprotocol/
claude-agent-acp`` prefix. The fix expands the registry default
*before* concatenating with args, so the textarea always shows
the full launch command and round-trips cleanly.

**Improvement: formatCommand drops empty-string args**

``formatCommand(["bash", "-c", ""])`` rendered as ``"bash -c "``
which parsed back to ``["bash", "-c"]``, silently losing the empty
slot. Now quotes empty tokens explicitly so they survive.

**Improvement: desktop sidebar disabled tooltip parity**

Mobile drawer's ``SettingsNavLink`` already showed "Disabled while
{agentName} is active" on greyed-out items; the desktop
``SidebarNavLink`` had no explanation. Added a ``disabledReason``
prop (i18n-agnostic; the caller formats the string) and wrap with
``StyledTooltip`` when disabled-with-reason. ``SettingsDesktopSidebar``
now forwards the formatted message — same UX on both surfaces.

**Test coverage gaps the bot flagged**

- ``agent-settings``: new regression guard for ``acp_command:[]`` +
  non-empty ``acp_args`` load (would have caught the critical bug
  above).
- ``acp-command``: empty-string round-trip case + explicit assertion;
  five more shell-operator filters (pipe, ``;``, ``&&``, ``||``,
  ``>>``).
- ``settings-navigation``: desktop sidebar wraps disabled items in
  StyledTooltip when ``disabledReason`` is supplied; not when omitted.

Plus a clean merge from ``origin/main`` (one-line import conflict
in ``agent-server-adapter.ts``).

2350 tests pass, lint + typecheck clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): address third-pass review on PR #416

- parseCommand: try/catch around shell-quote.parse so a malformed
  command in the textarea can't crash Settings → Agent mid-render
- Rewrite shell-metasyntax tests to pin the *actual* shell-quote
  behaviour (it's a parser, not a security filter) — operators,
  globs, and comments are dropped; backticks / $VAR / $(...) survive
  as literal tokens but are NOT expanded at parse time
- Add npm URLs + verification date (2026-05-19) to each ACP_PROVIDERS
  entry so future maintainers can re-check upstream packages
- Document the silent preset-switch behaviour on detectPreset (the
  dropdown follows the textarea; the textarea is the source of truth)
- Use the exported ACP_SERVER_TAG_KEY constant in the adapter test
  so a rename surfaces as a compile error rather than a runtime
  schema mismatch
- Restore the canonical typescript-client lock entry (drop the
  git+ssh:// + SHA bump that crept in from a local npm install)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): don't expose LLM-switch UI on ACP conversations

The SDK's ACPAgent carries a sentinel ``llm`` (``acp-managed``) for
cost-attribution only — the real model lives on the ACP subprocess
via ``acp_model`` and isn't visible on ``agent.llm.model``. Without
this fix, ``toAppConversation`` surfaced the sentinel as the
conversation's ``llm_model``, and the chat header's
SwitchProfileButton happily let users "change the model" while the
running Claude-Code / Codex / Gemini subprocess kept its own. A
confusing silent no-op.

Two layers of defence so no future consumer has to re-derive the rule:

  1. Boundary normalisation: ``toAppConversation`` reads the
     pydantic discriminator (``info.agent.kind === "ACPAgent"``),
     surfaces it as ``agent_kind: "acp" | "openhands"`` on
     AppConversation, and nulls ``llm_model`` for ACP. Mirrors
     OpenHands PR #14401.

  2. UI gate: SwitchProfileButton returns null when
     ``conversation.agent_kind === "acp"``. The right control for ACP
     model switching is the ``acp_model`` field on Settings → Agent,
     not this picker.

Tests cover both: a new ``toAppConversation`` case asserts
``agent_kind === "acp"`` + ``llm_model === null`` for an
``{kind: "ACPAgent"}`` payload, and a new SwitchProfileButton case
asserts the button hides for an ``agent_kind: "acp"`` conversation
even when profiles are present.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): bridge Settings → Secrets into the ACP subprocess env

The bare ``payload.secrets`` channel lands in the agent-server's
``secret_registry`` server-side, which the OpenHands ``Agent`` reads
directly — but ``ACPAgent._start_acp_server`` builds its subprocess
env from ``agent_context.secrets``, not from the registry. Without a
bridge, a Settings → Secrets entry like ``ANTHROPIC_API_KEY`` is
silently invisible to the ACP CLI (Claude Code, Codex, Gemini), so
users hit "authentication failed" with no on-screen hint that their
configured secret never reached the subprocess.

Mirror the same LookupSecret map onto
``payload.agent.agent_context.secrets`` when ``acpMode === true``,
so the agent-server's existing env-injection loop picks them up.
The bare ``payload.secrets`` channel is also kept (it serves other
consumers + remains the canonical "conversation secrets" wire). The
mirroring fires only when there's something to bridge; non-ACP
payloads are unchanged.

This is a shim. Once canvas pins to an agent-server build that
includes software-agent-sdk PR #3299 (which teaches ACPAgent to
also read from ``state.secret_registry``), the ``if (acpMode)``
branch can be deleted with no behaviour change.

Tests:
- New: ACP payload mirrors customSecrets onto agent_context.secrets
- New: empty customSecrets does NOT synthesize an empty bridge map
- New: non-ACP payload does NOT get an agent_context.secrets bridge
- All 42 adapter tests pass

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): hide MCP nav + cloud LLM-model fallback while ACP is active

Two ACP-leak fixes the review surfaced:

MCP page reachable + editable under ACP
- The SDK's ``ACPAgent`` rejects ``mcp_config`` on init (acp_agent.py:845)
  and the canvas adapter already strips it from start payloads, but the
  /mcp route and the Extensions nav still let users add / edit / delete
  MCP servers — silent no-ops against the running subprocess.
- Add a ``clientLoader`` on /mcp that bounces to /settings/agent when
  ``agent_kind === "acp"``. Grey out the MCP item in
  ExtensionsNavigation with the same explanatory tooltip the LLM /
  Condenser items already use under ACP.
- Extract the redirect into ``utils/acp-route-guard.redirectIfAcpActive``
  so /settings and /mcp share one cache-key + redirect-target
  definition. settings.tsx's clientLoader now calls into it.

Cloud chat ``ChatInputModel`` falls back to ``settings.llm_model`` for ACP
- ``toAppConversation`` writes ``llm_model: null`` on ACP conversations
  (commit 8f0efe62), but ChatInputModel did
  ``conversation?.llm_model ?? settings?.llm_model``, resurrecting the
  user's default OpenHands model on a Claude-Code conversation and
  linking to /settings (which is itself ACP-disabled). Gate on
  ``conversation?.agent_kind === "acp"`` and return null instead.

Tests:
- New: ExtensionsNavigation greys MCP under ACP, leaves Skills + non-ACP
  clickable
- New: /mcp clientLoader redirects under ACP, returns null otherwise +
  on settings-fetch errors (no redirect-loop)
- New: ChatInputModel returns null for ACP even when settings has a model
- All 18 affected tests pass

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): preserve unknown acp_server on no-op saves + reviewer cleanups

Fourth-pass review (PR #416 review comment 4486154133). Triage:

Fixed:
- **Unknown ``acp_server`` demoted to ``"custom"`` on save** (P1, real
  data corruption). A user with ``acp_server`` set out-of-band to a
  provider canvas's registry doesn't carry yet (e.g. a future provider,
  or one removed from the local mirror) would open Settings → Agent
  and lose the original key on the next Save — ``detectPreset`` routes
  every unknown server to ``ACP_CUSTOM_PRESET_KEY``. Now we capture
  the loaded ``acp_server`` + textarea at load time, and on save —
  when both are unchanged and the loaded key is non-empty,
  non-``"custom"``, and absent from ``ACP_PROVIDERS`` — pass it back
  verbatim via a new ``allowUnknownServer`` opt on
  ``buildAcpAgentSettingsDiff``. Editing the command still demotes
  to ``"custom"`` (user is configuring a new thing, so the preset
  name follows the command).
- **Dead ``...existingContext`` spread** in the ACP secret bridge.
  ``createAgentFromSettings`` never populates ``agent_context`` on the
  ACP branch, so the spread always merged into ``{}``. Direct
  assignment — and a comment explaining why a deep-merge would be the
  wrong direction (ACPAgent only treats ``secrets`` as acp_compatible).
- **Misleading ``$VAR`` test comment**. Reworded to lead with the
  no-leak contract (host env values must not end up in the persisted
  ``acp_command``) rather than the implementation-detail tangent.

Documented but not changed:
- **``acp_args: []`` "data loss" concern** — false alarm. Load merges
  ``acp_command + acp_args`` into the textarea before render; save
  persists the merged tokens as ``acp_command`` with ``acp_args: []``.
  Round-trip is correct. Added an inline comment on the load merge so
  the next reviewer doesn't re-flag the reset.
- **Silent preset migration without user feedback** — by design.
  The dropdown re-derives from the textarea so it always reflects
  what will be saved; adding a toast on every keystroke would be
  noise. Already documented as intentional on ``detectPreset``.

Tracked elsewhere:
- Supply-chain drift / npm verification: agent-canvas#587.
- Gemini onboarding icon: agent-canvas#621.

Tests:
- New: ``preserves an unknown loaded acp_server when the user saves
  without editing``
- New: ``demotes an unknown loaded acp_server to 'custom' when the
  user edits the command``
- All 73 affected tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): stop silently corrupting argv, restore AgentContext, fix home-screen gating

Three real review findings, all wired:

1. ``parseCommand`` silently dropped URL tokens with ``?`` query strings
   (and any other shell-glob metacharacter). Reproducer:

     node acp.js --endpoint https://example.com/acp?tenant=abc

   ``shell-quote.parse`` read ``?tenant=abc`` as a glob pattern and
   emitted a non-string AST node; the ``.filter(string)`` then dropped
   the URL entirely, persisting ``["node","acp.js","--endpoint"]``.
   Replaced ``shell-quote.parse`` with a small custom argv tokenizer
   that handles single/double quotes + backslash escapes and treats
   every other character — ``?``, ``*``, ``$``, ``|``, ``>``, ``#``,
   ``&``, ``;``, ``(``, ``)``, backticks — as literal. The agent-server
   passes the argv straight to ``subprocess.create_subprocess_exec``
   anyway (no shell intermediary), so the literal-only model matches
   what actually happens at spawn time. ``shell-quote.quote`` is
   still used by ``formatCommand`` for output.

2. The ACP path skipped the ``agent_context`` block that the OpenHands
   path seeded with ``load_public_skills`` / ``load_user_skills`` /
   optional ``system_message_suffix``. All three are marked
   ``acp_compatible: true`` on the SDK ``AgentContext`` model — the
   ACP CLI renders them via ``ACPAgent._render_suffix`` — so ACP
   conversations were silently shipping a smaller system prompt than
   OpenHands ones. ``createAgentFromSettings`` now seeds the same
   block on both branches. The secret bridge below merges into that
   block (was overwriting it) so ``{ secrets }`` no longer wipes the
   skill flags.

3. ``ChatInputModel`` and ``SwitchProfileButton`` only checked
   ``conversation?.agent_kind``. On the home screen (and during the
   task-startup window) ``conversation`` is undefined, so the
   per-conversation check missed and both surfaces fell back to
   ``settings.llm_model`` / the LLM-profile picker — even when
   ``settings.agent_settings.agent_kind === "acp"`` made it clear
   the next-created conversation would be ACP. Added a settings
   fallback so both controls hide consistently with the rest of the
   ACP nav gating.

Tests:
- parseCommand: new "preserves URLs with query strings" + "preserves
  URLs with multiple query params" + "preserves shell metacharacters as
  literal argv tokens" cases; the old "filters operator" cases flipped
  to "preserves operator as literal". 17 parseCommand cases pass.
- adapter: assertion on the ACP payload's ``agent_context`` updated to
  expect the skill flags instead of ``undefined``.
- chat-input-model + switch-profile-button: new "hides on the home
  page when ACP is the default agent" cases.
- 83 tests pass across the 5 affected files.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): two chat-rendering UX glitches on streaming ACP tool calls

1. Half-formed ACP tool-call cards flashed in the chat before the
   final state arrived. ACP servers stream multiple events per
   ``tool_call_id`` (status flips ``in_progress`` → ``completed`` /
   ``failed``); the intermediate events carry partial
   ``raw_input`` / ``raw_output`` / ``title``. The previous gate
   suppressed only ``in_progress`` and let ``null`` through (a
   "backwards compat" carve-out for older agent-server builds that no
   longer apply at our pinned version). Streaming intermediates often
   arrive without a status set yet, so they leaked.

   Tighten ``shouldRenderEvent`` to require ``status === "completed" ||
   "failed"``. ``handleEventForUI`` already collapses by
   ``tool_call_id`` in place, so the terminal event lands at the
   original position once it arrives — no flash, no double-render.

2. "Reading Read /Users/foo/bar" — Claude Code emits titles like
   ``"Read /Users/foo/bar"`` for a read tool, and our i18n template
   ``"Reading <cmd>{{title}}</cmd>"`` then doubles up the verb.

   Add ``stripRedundantTitlePrefix`` keyed by ``tool_kind``: read →
   strip ``"Read"``, edit → strip ``"Edit"`` / ``"Write"``, execute →
   strip ``"Bash"`` / ``"Run"``, fetch → strip ``"Fetch"`` /
   ``"WebFetch"``. Boundary-checked via trailing whitespace so a token
   like ``"Reads-from"`` is left alone. English-only on purpose: ACP
   servers are anglophone and emit english titles regardless of the
   user's canvas locale; matching translated verbs would go stale the
   moment a new server is added. Titles already lacking a redundant
   prefix (the OpenHands ACP wrapper, future servers) round-trip
   verbatim — the strip is a no-op there.

Tests:
- ``shouldRenderEvent``: ``null`` status now flips to false +
  comment explains why (treated as in-flight, not legacy).
- New ``stripRedundantTitlePrefix`` describe block covers the four
  tool kinds, the no-op case, the word-boundary guard, ``tool_kind:
  null`` (no strip), and empty titles.
- 53 tests pass across the conversation-events helpers.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): drop React Router type import on /mcp clientLoader (CI build)

CI's ``build:lib`` failed with:

  src/routes/mcp.tsx(3,23): error TS6059: File '.../.react-router/types/
  src/routes/+types/mcp.ts' is not under 'rootDir' '/src'.

``tsconfig.lib.json`` sets ``rootDir: "src"`` and pulls in
``src/components/**/*.tsx``. ``src/components/settings/index.ts``
re-exports from ``routes/mcp-settings``, which imports ``routes/mcp``
— so the lib's typecheck graph reaches ``routes/mcp.tsx`` and trips
on the generated ``./+types/mcp`` import that lives under
``.react-router/types/``, outside the lib's rootDir. (``routes/
settings.tsx`` uses the same import pattern but isn't reachable from
the lib graph, which is why local typecheck passed.)

Drop the type import and declare the loader with no parameters —
matches the existing ``index-redirect`` and ``mcp-settings-redirect``
loader pattern. Test calls collapsed to ``clientLoader()`` to match
the new signature.

``npm run build:lib`` now passes locally; 248 tests pass across the
affected suites.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(acp): tighten adapter comments around SDK refs

Two reviewer-flagged comment fixes:

- ``ACP_SETTINGS_KEYS`` docblock no longer claims there's a matching
  ``ACP_SETTINGS_KEYS`` constant in the Python SDK (there isn't;
  the fields are model attributes on ``ACPAgentSettings``). Reworded
  to "Keep aligned with the ``acp_*`` fields on ``ACPAgentSettings``
  in ``openhands-sdk/openhands/sdk/settings/model.py``" with an
  explicit "no matching SDK constant — hand-maintained" note, and
  cross-linked to the existing #587 drift tracker.

- ``createAgentFromSettings`` now spells out where the
  ``acp_compatible`` markers live on each of the three fields we set
  (``system_message_suffix`` L66, ``load_user_skills`` L80,
  ``load_public_skills`` L89 in
  ``openhands-sdk/openhands/sdk/context/agent_context.py``) plus what
  happens when a future SDK bump drops one (422 at conversation start
  → drop the demoted field, don't wrap a workaround). Line refs are
  brittle by design — they're the tripwire that surfaces a regression
  here rather than in production.

No behaviour change; lint + 42 adapter tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Debug Agent <debug@example.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-19 13:04:45 +00:00
Jamie Chicagoandopenhands bec399ca81 Fix stale default-local session key hydration (#607)
Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-18 21:42:23 +02:00
Tim O'Farrellandopenhands 048e390419 feat(automations): add logs modal to activity log items (#600)
* feat(automations): add logs modal to activity log items

Each AutomationRun now surfaces its bash command output via a small
terminal-icon button placed to the left of the run status badge in the
activity log. Clicking the icon opens a modal that fetches the
BashCommand event and all paginated BashOutput events for the run.

- Add bash_command_id to AutomationRun type (and mock data).
- New BashService.getCommandLogs(): cloud-aware reader for bash events
  that routes through callCloudProxy on cloud backends (runtime URL +
  session-api-key auth) and through BashClient directly on local
  backends. Pages through BashOutput events sorted by timestamp.
- New useBashCommandLogs hook: hydrates the run's conversation to
  resolve runtime URL + session API key, then drives the BashService
  query. Surfaces resolution states (loading, conversation missing,
  sandbox gone) so the modal can render meaningful empty states.
- New RunLogsModal: renders interleaved stdout/stderr from the command
  with stderr highlighted, exit code, and explicit messaging when the
  command is missing or the sandbox is no longer alive.
- Activity-log-item gets a logs button that stopPropagation +
  preventDefault on the wrapping conversation link.
- i18n: 8 new AUTOMATIONS$DETAIL$LOGS_* keys across all 15 languages.
- Tests: BashService cloud/local routing + ActivityLogItem button
  behaviour.

Co-authored-by: openhands <openhands@all-hands.dev>

* feat(automations): split logs modal into Output/Error tabs

Address review feedback on the run logs modal:

* Rename the modal title from "Run logs" to "Logs".
* Make the modal actually fire the bash-events search request:
  - In local mode, the agent-server hosts events under a single root,
    so the search query no longer waits for the per-conversation URL
    to resolve before firing. The conversation lookup still runs (so
    session-api-key and per-conversation URL are honoured when
    present), but it no longer gates the fetch.
  - In cloud mode the behaviour is unchanged — runtime endpoints
    require the conversation_url for the cloud-proxy hostOverride.
* Replace the interleaved stdout/stderr pre-block with two tabs:
  Output (stdout, default) and Error (stderr). The body of each tab
  is the chronological concatenation (by timestamp + order) of the
  matching field across every BashOutput event for the command.
* Simplify BashService — drop the extra BashCommand fetch; only the
  BashOutput search (`kind__eq=BashOutput, command_id__eq=<id>`) is
  needed for the rendered view. Note that the agent-server API uses
  `command_id__eq` (not `bash_command_id__eq`) — see the python
  bash_router for the canonical filter name.

i18n: add LOGS_TAB_OUTPUT, LOGS_TAB_ERROR, LOGS_EMPTY across all 15
locales; retranslate LOGS_TITLE to "Logs".

Tests:
* New run-logs-modal.test.tsx covers tab defaults, stdout/stderr
  concatenation (with reverse-order inputs to verify the sort),
  loading state, and Escape-to-close.
* bash-service.test.ts rewritten for the new listOutputs API and a
  cloud-without-conversation-url error path.

Co-authored-by: openhands <openhands@all-hands.dev>

* feat(automations): tolerate paused/missing/unreachable sandboxes in run logs modal

Cloud sandboxes can be in non-RUNNING states (paused, starting,
deleted, errored) and even RUNNING sandboxes can transiently fail at
the network layer. Previously the modal would either show a stuck
'Loading logs...' spinner or dump a raw axios error string. Now each
known-bad state is mapped to a stable `SandboxIssue` code with its
own localized empty-state message.

Behaviour:

* Pre-flight: when `sandbox_status` is MISSING, PAUSED, STARTING, or
  ERROR — or the conversation has no runtime URL at all — the bash
  query is **disabled** (no doomed request is fired). The modal
  renders the matching message instead of a spinner.
* Post-flight: when the request does fire and fails with a 404 or
  5xx response, or a network-level error (no response), the failure
  is classified as `unreachable` and the modal renders the
  'sandbox unreachable' message instead of the raw error. 401/403
  are *not* collapsed — those are auth bugs we want to surface.
* Local backends are unchanged: no sandbox lifecycle, so
  `sandboxIssue` is always null and errors flow through as-is.

API changes:

* `useBashCommandLogs` exposes a `sandboxIssue` discriminated union
  ("missing" | "paused" | "starting" | "errored" | "unreachable")
  instead of the old `hasNoRuntime` boolean. When an issue is set
  the hook clears `error` so the modal doesn't render both an empty
  state AND a raw axios string for the same failure.
* The modal switches over the issue codes via a centralized
  `SANDBOX_ISSUE_I18N` map.

i18n: replace the single `LOGS_SANDBOX_GONE` key with five specific
keys (one per sandbox issue) across all 15 locales.

Tests:

* New `__tests__/hooks/query/use-bash-command-logs.test.tsx`
  exhaustively covers each sandbox_status, the no-runtime-URL case,
  conversationMissing, the happy path, 404/5xx → unreachable
  classification, the network-error case, and the explicit
  401/403-passes-through case.
* `run-logs-modal.test.tsx` extended with a parameterized case
  covering all five issue codes; verifies the empty-state message is
  rendered and the tab body is suppressed.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-18 18:22:26 +00:00
Vasco Schiavo c8bf569d90 feat/fix(profiles): per-conversation /switch_llm in chat (#575)
* feat(profiles): per-conversation /switch_llm in chat, global activate on the home page

* chore(profiles): address feedbacks

* chore(profiles): fix lints
2026-05-18 14:11:11 -04:00
5ea55eaa9e feat(sidebar): conversation list filters, grouping, and loading UX (#530)
* feat(conversation-panel): filters, grouping, and list preferences

Add filter menu for organize/sort/thread scope, metadata toggles, and older
conversations with persisted preferences; optional LLM model labels on cards.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(conversation-panel): sidebar list UX, grouping chrome, and scroll divider

Align grouped rows and cards with main nav inset, add per-workspace/repo thread
launcher and folder-plus picker with shared menu styling, expose group launch
metadata for tests, and tighten card/scroll header borders to match the sidebar.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-card): restore ellipsis trigger and stack context menu above list rows

Bring back the shared vertical EllipsisButton and raise z-index on the open card
and menu so the overflow panel paints above subsequent items in the scroll area.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(sidebar): full-bleed conversations divider and align list rows with nav

Drop list-wrapper overflow clipping so the scrolled header border can span
the aside, keep a stable transparent/colored border, and inset the title row
with pl-4/pr-2. Remove extra card horizontal padding (link px-2 already
applies) and center status dots in an 18px column like SidebarNavLink icons.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-panel): match conversation row hover width to folder rows

Drop link px-2 so cards span the same width as grouped headers, and move
pl-2/pr-1 onto the card and skeleton to mirror folder strip padding.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-panel): portal conversation card menu and align list spacing

Open the ellipsis menu from click only, render it fixed in a body portal with
correct anchor measurement, and ignore outside-click closes on the trigger.
Match grouped-folder vertical rhythm to the main sidebar on md breakpoints.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-panel): polish list skeleton, filter inset, and timestamp nudge

Use one skeleton block that matches conversation card padding and corners; add
light right padding on the filter control; shift relative timestamps slightly
left for alignment.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-panel): restyle conversation list loading skeleton

Show three darker staggered pulse bars with compact list spacing; mount once
from the panel instead of five duplicates.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): unify skeleton styling and staggered pulse

Standardize .skeleton/.skeleton-round on neutral-700 with motion-safe pulse,
add .skeleton-stagger for the three-phase wave, and adopt it across loaders
(automations, chat, skills, tasks, secrets, settings, conversations).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-panel): polish filter menu and grouped list layout

Use clearer section dividers and Lucide sort icons in the conversations
filter menu; tighten the scroll stack and skip empty grouped nav rendering
when there are no workspace groups.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-panel): show older cutoff beside section title

Pair “Older conversations” with a dimmer “Over 1 hour” hint on the right in
the filter menu, with full i18n coverage for the new string.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-panel): drop non-sidebar skeleton stagger usage

Restore automations, chat, tasks, settings, skills, and secrets skeleton
markup to main so list skeleton styling stays scoped to the sidebar.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: group conversations by selected workspace, not per-conversation worktree dir

* fix: enable Delete all for every conversation, regardless of age

* refactor: update the code based on feedback

* refactor: update the code based on feedback

* refactor: update the code based on feedback

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: hieptl <hieptl.developer@gmail.com>
2026-05-18 16:31:54 +07:00
c224d24a9b fix: cloud conversation resume + archived/error sandbox states (#500)
* fix: resume cloud conversations stuck in starting status

Two bugs prevented cloud-backend conversations from resuming properly:

1. useActiveConversation hard-coded a 30 s refetch interval. When a
   cloud sandbox is paused and auto-starts on access, conversation_url
   is null until the sandbox is ready. The WebSocket can't open without
   a URL, so curAgentState stays at LOADING ('starting status') for up
   to 30+ seconds — or forever if the user gave up before the next poll.
   Fix: use the query-state callback form of refetchInterval and drop
   to 3 s whenever conversation_url is null (mirrors the 3 s cadence of
   task polling), falling back to 30 s once the URL is available.

2. updateConversationExecutionStatusInCache called setQueryData with a
   3-element key ["user", "conversation", id] that no longer matches
   the 5-element key stored by useUserConversation
   ["user", "conversation", id, backend.id, orgId] after the
   per-backend cache isolation was added. Optimistic status writes after
   manual pause/resume were silently dropped.
   Fix: switch to setQueriesData with { queryKey: [...] } prefix
   matching so the update hits whichever (backend, org) variant is live.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: auto-resume cloud sandbox when conversation_url is null

Faster polling (prev commit) was not enough. The cloud API returns
conversation_url=null when the sandbox is paused/stopped, and a GET
request alone does not wake it up — you have to POST a new start task
(with sandbox_id to reuse the existing sandbox) and wait for it to
become READY.

Add a useEffect in AppContent that fires once per unique conversation.id
after the initial fetch:
  • skips if not a cloud backend
  • skips if conversation_url is already set (sandbox running)
  • skips if sandbox_id is null (nothing to resume)
  • guards against re-triggering within the same route-mount via a ref

On trigger it calls createConversation(sandbox_id), which POSTs
POST /api/v1/app-conversations with the sandbox_id to the cloud, gets
back a WORKING start task, then navigates to /conversations/task-{id}.
useTaskPolling drives the task to READY and redirects to the real
conversation, now with a conversation_url the WebSocket can connect to.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: navigate back to conversation when resume task fails

When the cloud sandbox fails to start (e.g. 'Sandbox failed to start
within 120s'), the task reaches ERROR status. Previously the user was
left stranded at the task-{id} URL with only a toast to show for it.

Two changes:
1. Pass resumedFromConversationId in React Router navigation state when
   navigating to task-{id} for a cloud resume, so we know where to go
   back if the task fails.
2. In the task-error effect, read that state and navigate back to the
   original conversation (or /conversations if no originator is known).
   The resume effect's ref is still set so it will not re-trigger the
   resume on landing, preventing a retry loop.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use correct sandbox resume endpoint matching OpenHands

Root cause of 'Sandbox failed to start within 120s':
The previous fix called POST /api/v1/app-conversations with sandbox_id,
which is the 'create a new conversation' endpoint. The cloud treats this
as a full sandbox provisioning request with a 120-second cold-start
timeout that can fail on old/stale sandboxes.

The correct endpoint — matching OpenHands' SandboxService.resumeSandbox
and useSandboxRecovery — is POST /api/v1/sandboxes/{id}/resume, which is
a lightweight unpause that simply wakes the existing sandbox without
reprovisioning it.

Three changes:
1. Add SandboxStatus type ('PAUSED'|'RUNNING'|'STARTING'|'MISSING') and
   sandbox_status field to AppConversation, mirroring OpenHands'
   V1SandboxStatus. The cloud API already returns this field; adding the
   type makes it accessible in TypeScript.

2. Add resumeCloudSandbox(sandboxId) to the cloud service, calling
   POST /api/v1/sandboxes/{id}/resume via the cloud proxy — symmetric
   with the existing pauseCloudSandbox.

3. Update the resume effect in conversation.tsx:
   - Detect on sandbox_status === 'PAUSED' (more precise than
     conversation_url === null, which can be null for other reasons).
   - Call resumeCloudSandbox(sandbox_id) instead of createConversation.
   - Stay on the current URL after resume — no task navigation needed.
     The 3-second refetch interval in useActiveConversation polls until
     conversation_url populates, then the WebSocket connects normally.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: add ERROR to SandboxStatus to match OpenHands V1SandboxStatus

Complete enum is MISSING|STARTING|RUNNING|PAUSED|ERROR.

Co-authored-by: openhands <openhands@all-hands.dev>

* feat: archived/error sandbox state — read-only view and sidebar indicators

When a cloud conversation's sandbox_status is MISSING or ERROR it can
never be resumed. These two states now have first-class treatment:

Sidebar / conversation list:
- ConversationStatusDot gains an optional sandboxStatus prop.
  MISSING → gray 'paused' dot with tooltip 'Archived'.
  ERROR   → red 'error' dot with tooltip 'Error'.
  (ExecutionStatus visual is used unchanged for all other states.)
- ConversationCardHeader passes sandboxStatus to the dot and sets
  isConversationArchived on the title, which applies opacity-60.
- ConversationCard renders the existing ConversationStatusBadges pill
  ('Archived' or 'Error' pill badge) for MISSING and ERROR sandboxes.
- CompactConversationRow (collapsed sidebar) passes sandboxStatus to
  both the main dot and the tooltip-preview dot.
- conversation-panel.tsx passes sandbox_status from AppConversation to
  both card variants.

Conversation view (read-only):
- ChatInterface reads sandbox_status via useActiveConversation.
- When MISSING or ERROR, the InteractiveChatBox is replaced by a
  localised banner (title + description) explaining the history is
  read-only. The banner uses data-testid='archived-conversation-banner'
  for testing.
- The auto-resume effect in conversation.tsx already skips MISSING and
  ERROR because it only fires on sandbox_status === 'PAUSED'.

i18n:
  CHAT_INTERFACE$ARCHIVED_SANDBOX_TITLE/DESCRIPTION
  CHAT_INTERFACE$ERROR_SANDBOX_TITLE/DESCRIPTION

Co-authored-by: openhands <openhands@all-hands.dev>

* test: snapshot tests for archived/error sandbox conversation states

Three new Playwright visual snapshots in
tests/e2e/snapshots/archived-conversation.snapshot.spec.ts:

1. conversation-panel-with-archived-badges
   Navigates to /conversations; verifies five conversation cards are
   present; asserts archived-badge and error-badge are both visible;
   captures the conversation panel showing:
   - 'Archived Project'  → gray dot + 'Archived' pill + dimmed title
   - 'Errored Project'   → red dot + 'Error' pill + dimmed title

2. conversation-view-archived
   Navigates to /conversations/4 (sandbox_status: 'MISSING'); stubs
   WebSocket; asserts:
   - archived-conversation-banner is visible (read-only notice)
   - interactive-chat-box is absent (count 0)
   Captures the full chat interface.

3. conversation-view-sandbox-error
   Same as above for /conversations/5 (sandbox_status: 'ERROR');
   captures the 'Sandbox error' banner variant.

Supporting changes:
- src/api/agent-server-adapter.ts
  - Add sandbox_status?: string | null to DirectConversationInfo
  - Import SandboxStatus and map info.sandbox_status → AppConversation
    so the field is no longer silently null for all conversations
- src/mocks/conversation-handlers.ts
  - Add mock conversations 4 (MISSING) and 5 (ERROR) with sandbox_status
  - createConversationResponse now includes sandbox_status in the payload
- src/components/features/chat/chat-interface.tsx
  - Suppress ChatSuggestions ("Let's start building!") for archived
    conversations — showing task suggestions alongside a read-only
    banner is confusing and misleading
- src/components/features/conversation-panel/conversation-card/
  conversation-status-badges.tsx
  - Add data-testid="archived-badge" and data-testid="error-badge"
    so Playwright can assert on their presence without relying on text
- tests/e2e/snapshots/sidebar.snapshot.spec.ts
  - Update toHaveCount(3) → toHaveCount(5) to account for the two new
    mock conversations; existing sidebar snapshot baseline needs
    regeneration on main (intentional diff via update-snapshots label)

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: make sandbox_status optional in AppConversation

sandbox_status is a cloud-only field that local agent-server conversations
never carry. Existing test fixtures built AppConversation objects without
this field, causing TypeScript to error once it became required.

Making it optional (sandbox_status?: SandboxStatus | null) is the
semantically correct choice:
- The field is absent / null for every local conversation
- The adapter still explicitly maps it to null when unset
- ChatInterface reads it with ?? null so undefined is handled safely
- Partial<AppConversation> spreads in test factory functions no longer
  widen to SandboxStatus | null | undefined

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: Prettier formatting — multiline SandboxStatus union and ternary

- SandboxStatus type: expand single-line union to multi-line format
- ConversationCard: wrap sandboxStatus ternary in a multiline JSX block

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: preserve sandbox_status through requireDirectConversationInfo

The validation function that normalises raw API responses into
DirectConversationInfo was not copying sandbox_status, so it was silently
dropped every time a conversation came through the search or batch-get
code paths. This caused the conversation panel to never render the
archived/error badge pills, and the ChatInterface to always treat every
conversation as active (missing read-only banner for MISSING/ERROR sandboxes).

Fix: add sandbox_status: stringOrNull(item.sandbox_status) to the
mapping in requireDirectConversationInfo, mirroring the treatment of
execution_status.

The existing E2E tests for conversations 4 (MISSING) and 5 (ERROR)
were already asserting on archived-badge / error-badge presence and
archived-conversation-banner visibility, so they will now pass.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: don't connect WebSocket while cloud sandbox is PAUSED

When a cloud conversation is closed from the UI (pauseCloudSandbox is
called), the conversation's conversation_url is NOT cleared — it still
points to the old sandbox host. On the next navigation into that
conversation the WebSocket provider saw a non-null URL and immediately
tried to open a connection, which failed because the sandbox had not
yet woken up.

Two-part fix:

1. WebSocketProviderWrapper: suppress conversationUrl (treat it as null)
   while sandbox_status === 'PAUSED', so ConversationWebSocketProvider
   cannot compute a valid wsUrl until the sandbox is actually running.

2. useActiveConversation: add sandbox_status === 'PAUSED' as a
   fast-poll trigger alongside !conversation_url. The old check only
   fast-polled when the URL was absent; for paused sandboxes the URL is
   present but stale, so without this the hook would stay on the slow
   30-second interval while waiting for the sandbox to wake up.

Together these changes let the resume sequence complete correctly:
  navigate → sandbox PAUSED detected → resumeCloudSandbox called →
  fast-poll picks up RUNNING state → conversationUrl unblocked →
  WebSocket connects with a live sandbox host.

Co-authored-by: openhands <openhands@all-hands.dev>

* docs: document cloud PAUSED sandbox WebSocket gating in AGENTS.md

Co-authored-by: openhands <openhands@all-hands.dev>

* test: cover PAUSED sandbox gating and sandbox_status preservation

Three test suites covering the cloud conversation resume bug fixes:

1. agent-server-conversation-service.test.ts — three cases asserting
   that requireDirectConversationInfo preserves sandbox_status through
   batchGetAppConversations (PAUSED, RUNNING, absent → null).

2. websocket-provider-wrapper.test.tsx — five cases asserting that
   WebSocketProviderWrapper passes conversationUrl through when the
   sandbox is RUNNING or null (local backend), suppresses it to null
   when sandbox_status === 'PAUSED', and handles not-yet-fetched data.

3. use-active-conversation.test.ts — five cases asserting that the
   refetchInterval callback returns 3000 when sandbox_status is PAUSED
   (even with a non-null conversation_url) or when conversation_url is
   null, and 30000 in all other ready states.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(test): use null for execution_status fixture field

ExecutionStatus is a string enum — assigning the raw string literal
'idle' triggers TS2322. Null satisfies ExecutionStatus | null and is
irrelevant to what these tests actually exercise.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(lint): prettier format + add missing i18n fallbacks for 7 keys

Two issues from lint-staged pre-commit hook:

1. Prettier: the compound refetchInterval condition in use-active-
   conversation.ts was too long for one line — broke across three lines.

2. Translation completeness: CHAT_INTERFACE$ARCHIVED_SANDBOX_TITLE/
   DESCRIPTION, CHAT_INTERFACE$ERROR_SANDBOX_TITLE/DESCRIPTION, and
   BACKEND$NAME_REQUIRED/HOST_REQUIRED/HOST_INVALID were added in earlier
   commits on this branch but only had English values. Added English
   fallbacks for all 14 other supported locales in translation.json and
   regenerated public/locales/ via make-i18n.

Co-authored-by: openhands <openhands@all-hands.dev>

* i18n: add proper translations for 7 new keys across 14 locales

The previous commit used English as a fallback for all non-English
locales. Replace with proper translations for:

  CHAT_INTERFACE$ARCHIVED_SANDBOX_TITLE
  CHAT_INTERFACE$ARCHIVED_SANDBOX_DESCRIPTION
  CHAT_INTERFACE$ERROR_SANDBOX_TITLE
  CHAT_INTERFACE$ERROR_SANDBOX_DESCRIPTION
  BACKEND$NAME_REQUIRED
  BACKEND$HOST_REQUIRED
  BACKEND$HOST_INVALID

Locales covered: ja, zh-CN, zh-TW, ko-KR, no, ar, de, fr, it, pt,
es, ca, tr, uk.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: also write sandbox_status PAUSED to cache on stop-conversation

Bug: clicking 'Stop conversation' called pauseConversation() and then
only wrote execution_status: PAUSED to the React Query cache via
updateConversationExecutionStatusInCache.  sandbox_status was never
touched, so it remained as whatever the server last returned (null or
'RUNNING').

When the user reopened that conversation:
  • WebSocketProviderWrapper checked sandbox_status === 'PAUSED' → false
    → URL passed through → WebSocket fired at the paused sandbox → failed
  • useActiveConversation saw sandbox_status !== 'PAUSED' AND url !== null
    → 30-second poll interval → 30s before discovering the true state

Fix:
  1. Add patchConversationInCache() to conversation-mutation-utils —
     a generic helper that patches any subset of AppConversation fields
     in both the single-item and paginated-list query caches.
     updateConversationExecutionStatusInCache becomes a thin wrapper.
  2. use-unified-stop-conversation.ts uses patchConversationInCache to
     write BOTH execution_status: PAUSED and sandbox_status: 'PAUSED'
     atomically in onSuccess, so the gate in WebSocketProviderWrapper
     fires immediately on the next render.

Tests: __tests__/hooks/mutation/conversation-mutation-utils.test.ts
  • patchConversationInCache patches single-item cache
  • patchConversationInCache patches paginated list cache
  • patchConversationInCache patches multiple fields atomically
  • patchConversationInCache does not modify unrelated conversations
  • patchConversationInCache is a no-op on empty cache
  • updateConversationExecutionStatusInCache wrapper only touches
    execution_status (sandbox_status is left unchanged)

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(ui): archived/error banner — 'above' copy and readable text colors

Two issues with the archived/error sandbox banner that replaces the
chat input:

1. Copy said 'The history below is read-only' but the banner is
   anchored to the bottom of the chat, so the history is above it.
   Changed to 'above' in all 15 locales (en + ja/zh-CN/zh-TW/ko-KR/
   no/ar/de/fr/it/pt/es/ca/tr/uk).

2. Description text used text-[var(--oh-color-tertiary)] which maps to
   cool-grey-800 — nearly indistinguishable from the cool-grey-925
   surface background. Switched to the palette tokens that the rest of
   the UI uses for readable text on dark surfaces:
   • Title:       --oh-foreground  (cool-grey-100, bold label)
   • Description: --oh-muted       (cool-grey-400, secondary body text)

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): render echo-hello-world trajectory in archived/error views; fix lint

Three changes in one commit:

1. Prettier lint fix (conversation-mutation-utils.ts line 116):
   The one-line arrow body for updateConversationExecutionStatusInCache
   exceeded Prettier's column limit when written inline; split onto its
   own line. This fixes the 'test-and-build (ubuntu) Lint' CI failure.

2. MSW event fixture (src/mocks/conversation-handlers.ts):
   Add ECHO_HELLO_WORLD_TRAJECTORY — three events in TIMESTAMP_DESC
   order (newest-first, as the hook requests) that represent a minimal
   'echo hello world' session:
     archived-evt-1  user MessageEvent    'echo hello world'
     archived-evt-2  agent ExecuteBashAction  echo hello world
     archived-evt-3  env   ExecuteBashObservation  'hello world'
   Wire CONVERSATION_EVENTS map so GET /api/conversations/4/events/search
   and /5/events/search return these events; other conversations still
   get []. useConversationHistory reverses the DESC list back to
   chronological order before storing events.

3. Snapshot test (archived-conversation.snapshot.spec.ts):
   Wait for chatInterface.getByText('echo hello world') to be visible
   before taking the screenshot so the trajectory is guaranteed to have
   rendered above the read-only archived/error banner.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): inject trajectory via Zustand store, not MSW cross-origin fetch

The Service Worker registered at localhost:3001 cannot intercept
cross-origin requests; RemoteEventsList calls GET on the configured
backend host (127.0.0.1:8000), so MSW silently drops the response and
useConversationHistory returns no events.

Fix: pull the injectEvents helper pattern from
collapsible-thinking.snapshot.spec.ts and call it after asserting the
archived/error banner is visible.  The fixture is declared once at the
top of the file alongside a clear comment explaining why it mirrors the
MSW handler rather than importing from it.

Also removes the 10 s timeout from the post-inject getByText check
since injectEvents already polls until the store is populated and then
waits 500 ms for React to flush.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): atomic addEvents+DOM poll in injectEvents, no separate getByText

The previous impl had a two-step race window:
  1. expect.poll passed once store.events.length >= N
  2. 500 ms wait (or DOM waitForFunction) ran afterwards

React Strict-Mode's double clearEvents() invocation could fire between
steps 1 and 2, wiping the store before React flushed the render.

Fix: merge addEvents() and the data-testid="user-message" DOM check into
a single page.waitForFunction() poll.  Playwright polls ~100 ms so on
every tick we both re-seed the store AND verify the DOM element is
present.  addEvents() is idempotent (deduplicates by event ID) so
calling it on every tick is safe.  This eliminates the race entirely.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): wait for archived-banner as settled-state signal in tests 2/3

The previous approach waited for `chat-interface` (h-full flex div) to become
visible, but that container can be present in the DOM with zero computed height
before useActiveConversation resolves — causing intermittent 20 s timeout
failures in CI.

Following the same pattern as collapsible-thinking.snapshot.spec.ts (which
waits for `"Let's start building!"` as its settled-state signal), tests 2/3
now use a dedicated `navigateToArchivedConversation` helper that waits for
`archived-conversation-banner` to be visible (timeout 30 s).

The banner only renders after useActiveConversation returns data with
sandbox_status MISSING or ERROR, so it is a reliable indicator that:
  - the MSW mock responded to GET /api/conversations?ids=<id>
  - React Query received the data and set isFetched = true
  - ChatInterface evaluated isArchivedConversation = true
  - The banner div is both present and has non-zero dimensions

Also removes the now-redundant in-test banner visibility checks (the helper
already asserts them) and the stale `navigateToConversation` helper that was
no longer used.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): MSW ids[] parse bug + inject one event for stable archived-view test

Root cause of snapshot CI failures:

  Axios serializes { ids: ["4"] } as ?ids[]=4 (bracket notation).
  The MSW GET /api/conversations handler read searchParams.getAll("ids"),
  which returns [] when the key is "ids[]". listConversationResponses([])
  then falls back to returning ALL conversations, so results[0] was always
  conversation "1" (first in Map insertion order) regardless of which id
  was requested. Conversation "1" has no sandbox_status, so
  isArchivedConversation was always false and the archived banner never
  rendered — 30 s timeout.

Fix 1 — conversation-handlers.ts:
  Parse both bracket (ids[]) and plain (ids) formats so the mock correctly
  returns only the requested conversation(s).

Fix 2 — archived-conversation.snapshot.spec.ts:
  Rewrite tests 2/3 per user direction:
  • Use seedLocalStorage (same as collapsible-thinking) instead of bespoke
    addInitScript + page.route helpers.
  • Inject ONE minimal ExecuteBashAction event via __OH_EVENT_STORE__ so the
    chat has stable visible content that survives the 3 s polling re-renders
    (conv 4/5 have no conversation_url, so useActiveConversation polls every
    3 s). The injected event stays in the Zustand store across re-renders,
    giving toBeVisible a reliable anchor.
  • Wait for "echo hello" text (event), then wait for the archived banner —
    both are concrete settled-state signals, not the zero-height h-full div.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: trigger CI re-run

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: Update PR QA artifacts

* fix(tests): single injected event for archived-view + useOptionalConversationId mocks

- Remove ECHO_HELLO_WORLD_TRAJECTORY from MSW (was causing 3+1 = 4 events
  in the archived conversation snapshot view). CONVERSATION_EVENTS is now
  empty; the snapshot tests inject exactly one event via __OH_EVENT_STORE__.
- Add useOptionalConversationId to all vi.mock('#/hooks/use-conversation-id')
  calls that were missing it after the main merge refactored that hook.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): wait for banner before injecting events to avoid clearEvents race

The archived-conversation snapshot tests were injecting events via
__OH_EVENT_STORE__ immediately after the store became available on the
window object. However, the conversation route's useEffect (which calls
clearEvents()) fires asynchronously after the first paint — creating a
race where the injected events get wiped.

Fix: wait for the archived-conversation-banner to appear before
injecting events. The banner's presence proves that:
1. The route's clearEvents() effect has already fired
2. useActiveConversation has resolved with the correct sandbox_status
3. The chat interface is ready to accept and display events

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): pre-seed archived conversation events via MSW instead of runtime store injection

The archived-conversation snapshot tests were injecting events into the
Zustand event store at runtime via __OH_EVENT_STORE__. This raced with
the conversation route's useEffect (clearEvents) and React dev-mode
double-mount behavior, making the injected events disappear before the
chat could render them.

Fix: pre-seed CONVERSATION_EVENTS in the MSW mock handlers for
conversations 4 and 5 with one ExecuteBashAction event. The events now
load through the normal REST history path (useConversationHistory →
addEvents) — no runtime Zustand injection, no race condition.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): remove event injection — test banner + hidden input only

The archived-conversation snapshot tests kept crashing because event
injection (both via __OH_EVENT_STORE__ and pre-seeded MSW REST data)
always gets wiped by a React 18 strict mode effect-ordering issue:

In dev mode, strict mode double-fires effects child-before-parent.
ConversationWebSocketProvider (child) calls addEvents() first, then
conversation.tsx (parent) calls clearEvents() second, wiping all events.

Since the feature under test is the read-only banner and hidden chat
input (not event rendering), simplify the tests to verify only those
assertions — no event injection needed.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): add WebSocket stub to archived-conversation tests

The conversation-view snapshots showed a red 'Failed to connect to
server' toast because no agent-server runs at :8000 in CI — the Vite
proxy's ECONNREFUSED propagates to the browser and triggers the error
toast. Other conversation-page snapshot tests already stubbed
WebSocket; this test was missing it.

Extract the duplicated WebSocket stub into a shared helper at
tests/e2e/snapshots/support/stub-websocket.ts and use it in all three
conversation-page snapshot test files (archived-conversation,
collapsible-thinking, changes-tab).

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: update conversation card count from 5 to 6 after main merge

Main added pagination-local conversation fixture, bringing the total
mock conversations to 6. The archived-conversation sidebar test was
still asserting 5.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: update backends-extended snapshot tests for two-column add modal

The add-backend modal was refactored from a single form with radio
buttons (local/cloud kind selection) into a two-column layout:
- Left: manual connection (name, host, API key, Connect)
- Right: cloud OAuth login (device flow)

Kind is now inferred from the host URL, so the old radio button
testids (add-backend-kind-local, add-backend-kind-cloud) no longer
exist. Updated all affected flows:
- Flow 1: removed radio clicks, use URL inference for kind
- Flow 2: replaced radio inference tests with two-column layout test
- Flow 3: replaced OAuth button gating with cloud advanced settings
- Flow 7: removed radio click
- Flow 8: use add-backend-close instead of add-backend-cancel

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: update sidebar snapshot test conversation count from 5 to 6

Same pagination-local fixture issue as archived-conversation test.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: allhands-bot <allhands-bot@users.noreply.github.com>
2026-05-17 13:27:52 -04:00
Hiep Le 5d3e5f7611 chore: remove SaaS references from codebase (#548) 2026-05-17 15:46:17 +07:00
Xingyao Wangandopenhands 2123a2761c feat: enable paginated event loading for cloud mode (#407)
* feat: enable paginated event loading for cloud mode

Enable scroll-up pagination for cloud backends by passing all search
params (sort_order, page_id, timestamp__gte, timestamp__lt) through to
the cloud proxy. The cloud useLoadOlderEvents hook no longer gates on
`isCloud`, so long conversations load the latest 50 events first and
lazily backfill older pages as the user scrolls up.

The loading indicator now shows 'Fetching older messages…' alongside
the spinner so users know what's happening during pagination.

Depends on: OpenHands/OpenHands#14399 (server-side timestamp fix)
Closes #402

Co-authored-by: openhands <openhands@all-hands.dev>

* address review: add fallback for unpatched cloud backends + tests

- Cloud event search now tries full params first, falls back to
  limit-only on error (graceful degradation for servers without
  OpenHands/OpenHands#14399).
- Added JSDoc note about server dependency to useLoadOlderEvents.
- Updated event-service tests: verify all params forwarded, fallback
  on 500, rethrow on limit-only failure, pagination stop on short page.
- Removed obsolete cloud-disabled test and useActiveBackend mock from
  use-load-older-events tests.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: return empty page on fallback to prevent infinite retries

When an unpatched cloud backend rejects timestamp filters, return an
empty page instead of retrying with limit-only params. The limit-only
fallback would return the same most-recent events already in the store,
which get deduped but leave hasMore=true — causing infinite requests.

An empty page makes useLoadOlderEvents set hasMore=false, cleanly
stopping pagination on unpatched backends.

Co-authored-by: openhands <openhands@all-hands.dev>

* docs: fix stale comment about fallback behavior

Co-authored-by: openhands <openhands@all-hands.dev>

* test: add event pagination e2e coverage

Add deterministic mock conversation fixtures for local and cloud event pagination, plus Playwright regression coverage that verifies initial tail loading and scroll-up older-event backfill for both backend modes.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: format event pagination params

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-16 19:53:31 +00:00
Robert Brennanandopenhands 35555bac11 feat: enable sub-agent delegation via task_tool_set (#509)
The agent-server already registers `task_tool_set` (TaskToolSet) at
startup via openhands-agent-server/openhands/agent_server/tool_router.py
and preloads the built-in sub-agents (code-explorer, bash-runner,
web-researcher, general-purpose) through `register_builtins_agents`.
However, agent-canvas never asked for the tool in the agent spec it
sends on POST /api/conversations, so the LLM had no way to delegate.

Add `task_tool_set` to the tools list assembled by `getAgentTools()`,
gated by the existing `isAgentServerToolAvailable` capability probe so
older agent-servers that don't advertise it in /api/server_info's
`usable_tools` are skipped cleanly.

TaskAction / TaskObservation events render through the existing default
event content path, so no new visualizer is required to ship this.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-16 18:00:03 +00:00
Hiep Le 5c33c10b18 feat(frontend): add canvas_ui tool so the agent can drive the UI (#420)
* feat: add canvas_ui tool so the agent can drive the UI

* refactor: update the code based on feedback

* fix: failing tests

* refactor: update the code based on feedback
2026-05-16 16:09:43 +07:00
Rohit Malhotraandopenhands 60e103eec5 fix: route cloud runtime bash/file calls through cloud proxy (#507)
* fix: route cloud runtime bash/file calls through cloud proxy

useWorkspaceFiles, useLocalGitInfo, useHasGitCommits were all building
RemoteWorkspace with getAgentServerClientOptions({ conversationUrl }),
which resolves the host directly to the cloud runtime URL when a cloud
conversation is active (e.g. *.prod-runtime.all-hands.dev).  This caused
CORS errors because the browser made the fetch from localhost.

useWorkspaceFileContent had the same issue for GET /api/file/download.

The fix centralises these operations in a new
AgentServerRuntimeService (src/api/runtime-service/) that mirrors the
pattern already used by agent-server-git-service and event-service:

  if (active.kind === 'cloud' && conversationUrl)
    -> callCloudProxy({ hostOverride: buildHttpBaseUrl(conversationUrl), authMode: 'session-api-key', ... })
  else
    -> SDK typed clients directly

- executeCommand  routes POST /api/bash/execute_bash_command
- downloadFile    routes GET  /api/file/download

use-local-git-info helper functions (probeGitInfoAtDir,
probeNestedRepoInDir) are refactored from taking a RemoteWorkspace
instance to taking a RunCommand callback, keeping the helpers pure.

Co-authored-by: openhands <openhands@all-hands.dev>

* test: cover AgentServerRuntimeService cloud/local routing

Adds 12 unit tests for AgentServerRuntimeService covering:
- executeCommand local path: RemoteWorkspace constructed with resolved
  options; callCloudProxy never called
- executeCommand cloud path: callCloudProxy invoked with POST to
  /api/bash/execute_bash_command, correct hostOverride, body, session-
  api-key auth, and timeoutSeconds; RemoteWorkspace never created; cwd
  omitted when undefined; null stdout/stderr normalised to empty strings;
  null conversationUrl falls back to local
- downloadFile local path: FileClient constructed with resolved options;
  callCloudProxy never called
- downloadFile cloud path: callCloudProxy invoked with GET to
  /api/file/download with URL-encoded path, blob responseType, session-
  api-key auth; FileClient never created; Blob→ArrayBuffer round-trip
  preserves content; null conversationUrl falls back to local

Also adds a cloud-backend integration test to
use-workspace-file-content.test.tsx confirming the hook routes
downloads through callCloudProxy instead of FileClient when a cloud
backend is active, and that callCloudProxy is called with the correct
path/auth/responseType.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-16 04:40:41 +00:00
14d2e9454b test(snapshot): changes tab diff viewer + backend management UI (6 tests) (#450)
* test(snapshot): changes tab diff viewer + backend management UI (6 tests)

Pre-seed MOCK_GIT_CHANGES with M/A/D entries (using AgentServerGitChangeStatus
values: UPDATED/ADDED/DELETED) so changes-tab tests can exercise the file list,
Monaco diff viewer, and deleted-file placeholder without per-test MSW manipulation.

Expose window.__setMockGitChanges__ so the empty-state test can clear the list
after boot and trigger a React Query refetch via __TEST_INVALIDATE_QUERIES__,
avoiding a full page reload that would reinitialise module state.

Backend management tests exercise the selector dropdown, add-backend modal, and
manage-backends modal — all driven by localStorage seeding via addInitScript.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: update baseline snapshots [skip ci]

* ci: trigger re-run against CI-generated baselines

* fix(snapshot-tests): mask Monaco editor for stable CI screenshots; fix unit test

- changes-tab spec: mask data-testid=editor-container so Monaco's sub-pixel
  font hinting (which varies per OS) doesn't cause false pixel-diff failures
- mock-conversation-handlers test: update assertion to match the new pre-seeded
  MOCK_GIT_CHANGES (3 M/A/D entries) instead of the previous empty array

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: update baseline snapshots [skip ci]

* ci: trigger re-run against CI-regenerated baselines (Monaco mask + unit test fix)

* fix(snapshot-tests): normalize RandomTip height via addStyleTag for stable empty-state screenshot

RandomTip renders a randomly-chosen tip whose line-count varies, causing the
flex-1 container above it to have different heights across runs. Fix by injecting
a CSS rule via page.addStyleTag() that pins .text-m.bg-tertiary.p-4 to 80px
(visibility:hidden so the variable text is invisible) — layout is now deterministic.
Switch back to screenshotting the full files-tab panel since dimensions are stable.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: update baseline snapshots [skip ci]

* ci: trigger re-run against baselines (empty-state RandomTip height fix)

* fix(snapshot-tests): use inner content div for empty-state screenshot to avoid left-strip artefact

Screenshot files-tab's last direct div child (the flex-1 content wrapper)
instead of the outer main element.  During CI baseline generation the outer
main's bounding box occasionally captured a ~30px left-panel overlay artefact
that made the baseline permanently diverge from subsequent verification runs.
Targeting the inner wrapper excludes the outer-element overflow while still
showing the full empty-state (icon + 'no changes yet' text + hidden tip area).

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: update baseline snapshots [skip ci]

* ci: validate against fresh inner-div empty-state baseline

* test(snapshot): extended backend UI flows — 12 tests, 19 screenshots

Add backends-extended.snapshot.spec.ts covering 8 behaviour flows
with iterative screenshot captures at each state transition:

Flow 1a  Blank add form — Save disabled until name+host filled
Flow 1b  Local backend — Save enabled with name+host, no API key needed
Flow 1c  Cloud backend — Save disabled without API key, enabled with it
Flow 2a  Host auto-infers Local kind; OAuth section disappears
Flow 2b  Cloud-domain URL keeps Cloud kind; OAuth section shows
Flow 2c  Manual kind selection locks type (touchedKind=true) even when
         a cloud URL is later typed into the Host field
Flow 3   OAuth Login button disabled while host is empty; enabled once filled
Flow 4   Remove backend: shows ConfirmationModal → Cancel keeps row →
         Confirm removes it from the list (4 screenshots)
Flow 5   Edit modal pre-populates name/host/key from stored backend
Flow 6   Switch active backend: environment-switch overlay captured via
         page-level screenshot + animation override so the card is
         opaque at frame-0; after-switch state verified via selector label
Flow 7   Whitespace-only host keeps Save disabled; syntactically invalid
         URL is accepted by the frontend (no URL-format validation)
Flow 8   Cancel add form: dismisses modal, Manage Backends confirms no
         phantom entry was saved

Notable decisions:
- Uses body[data-environment-switching="true"] as the early DOM signal
  before React paints the portal div for the switch overlay
- Adds inline style-tag override before the overlay screenshot because
  .environment-switch-overlay > div has opacity:0 at animation frame 0;
  Playwright's animations:"disabled" pauses there, making the card
  invisible without the override
- Backends seeded via page.addInitScript localStorage injection so
  tests are fully self-contained with no MSW state dependency

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: update baseline snapshots [skip ci]

* ci: validate extended backend snapshot tests against CI baselines

* ci: always post snapshot PR comment even when test generation step fails

The 'Post snapshot report to PR' step was skipped whenever 'Generate
current PR snapshots' exited non-zero (e.g. a test crash like a hidden
element, not just a snapshot diff). GitHub Actions skips steps without
an always() guard when a prior step fails.

Add always() so the comment is posted regardless — showing diffs or
the test failure output — which was the intended behaviour.

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: fix snapshot comment - remove tracked screenshots, add crash reporting

Three fixes:

1. Remove 28 git-tracked snapshot PNGs from this branch.
   These were committed by the old baseline-in-git workflow before #482
   migrated to artifact storage. Because they stayed tracked (gitignore
   doesn't untrack already-indexed files), every CI checkout put them in
   tests/e2e/__snapshots__/ BEFORE the baseline artifact was downloaded.
   The Save step then copied them into /tmp/main-baselines, making the
   new tests appear as 'Unchanged' instead of 'New' in the PR comment.

2. Add 'Clear snapshot directory before downloading baselines' step.
   Wipes tests/e2e/__snapshots__/ before the artifact download so any
   future accidentally-tracked files can never contaminate the baseline.

3. Surface test crashes in the PR comment.
   - Generate step gets continue-on-error + an id so subsequent steps
     can read its outcome.
   - GENERATE_OUTCOME is passed to the comment script.
   - If outcome == 'failure', a GitHub-flavoured WARNING callout is
     prepended to the comment with a direct link to the CI run logs.
   - A dedicated 'Fail if snapshot generation had test crashes' step
     restores the job failure that continue-on-error absorbed.

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: use PR number in snapshot concurrency group for cleaner cancellation

The previous group used github.ref which resolves to refs/pull/{N}/merge
for PR events — correct but opaque. Using github.event.pull_request.number
makes the grouping explicit and human-readable (snapshot-tests-450), and
falls back to github.ref for main pushes and workflow_dispatch.

cancel-in-progress: true was already set, so new commits already cancelled
prior runs. This just makes the intent clearer.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: syntax error in post-snapshot-comment.mjs (] vs ) in lines.push)

lines.push(...) was accidentally closed with ]; instead of ); after
splitting the original lines = [...] array literal into a push call.
Caused a SyntaxError at startup, preventing any comment from being posted.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: snapshot test disabled states, changes-tab crash, and CI false-failures

Three fixes:

1. BrandButton disabled visual styling (brand-button.tsx)
   disabled:opacity-30 pseudo-class was not applying in Vite dev mode
   (Tailwind v4 + postcss-prefix-selector interaction), making disabled
   and enabled buttons visually identical in snapshot screenshots.
   Fix: add isDisabled conditional class directly ('opacity-30
   cursor-not-allowed pointer-events-none') so the disabled appearance
   is applied regardless of whether :disabled pseudo-class works.

2. changes-tab test crash (changes-tab.snapshot.spec.ts)
   Test waited for data-testid='files-tab' but the right panel always
   starts CLOSED (isRightPanelShown = false is session-only Zustand
   state; sanitizeStoredState strips any persisted rightPanelShown key).
   Fix: click data-testid='right-panel-toggle' after navigation to open
   the panel before waiting for files-tab. Also remove the no-op
   rightPanelShown: true from the localStorage seed.

3. CI false-failures for new snapshot tests (snapshot-tests.yml +
   post-snapshot-comment.mjs)
   The 'Fail if comparison found differences' step fired on
   'missing baseline' failures (expected for new tests in a PR) as
   well as actual pixel-diff failures.
   Fix:
   - post-snapshot-comment.mjs outputs has_changes=true/false to
     GITHUB_OUTPUT (true only when changed.length > 0, i.e. real diffs)
   - 'Fail if' step now checks steps.post-comment.outputs.has_changes
     == 'true' instead of compare.outcome == 'failure', so PRs that
     only add new snapshot tests pass CI cleanly.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: reject invalid host URLs in backend form; use http for local addresses

Two related fixes to backend host validation / normalisation:

1. isValidHostUrl() — reject invalid host strings
   canSubmit previously only checked host.trim().length > 0, so
   garbage like 'not://:::a valid url!!!' passed through and enabled
   the Save button. isValidHostUrl() adds two checks before the URL
   constructor: (a) the trimmed value must be non-empty, (b) it must
   contain no whitespace. This catches the test-case input whose spaces
   are the tell-tale sign of a malformed value.

2. normalizeHost() — http:// for local addresses
   Bare hostnames (no explicit scheme) were unconditionally prepended
   with https://, but local servers almost never have TLS certificates.
   The new isLocalAddress() helper detects localhost, 127.x, RFC-1918
   private ranges (10.x, 192.168.x, 172.16-31.x), .local / mDNS names,
   and single-label hostnames — all get http:// instead of https://.
   Hostnames with dots that are not in those ranges (e.g. app.all-hands.dev)
   still default to https://. Explicit http:// or https:// prefixes are
   always preserved as-is.

Test update: the 'backend-add-invalid-url-accepted' snapshot is renamed
to 'backend-add-invalid-url-disabled' and the assertion flips from
not.toBeDisabled() → toBeDisabled(), reflecting the new behaviour.

Co-authored-by: openhands <openhands@all-hands.dev>

* feat: inline error feedback on Name and Host fields in BackendForm

Three parts:

1. SettingsInput gains error / showRequiredTag / onBlur props
   - error?: string — red border on the input plus a small red alert
     paragraph below it (role=alert, data-testid=${testId}-error, linked
     via aria-describedby).
   - showRequiredTag?: boolean — renders a red * after the label to
     signal that the field is mandatory, consistent with OptionalTag.
   - onBlur?: () => void — forwarded directly to the <input>.
   - aria-invalid is set automatically when error is truthy.

2. BackendForm wires touched state → errors → inputs
   - nameTouched / hostTouched (both false on open, set on blur)
   - nameError: 'Name is required' when touched + empty
   - hostError: 'Host is required' when touched + blank/whitespace;
                'Enter a valid URL (e.g. http://localhost:8080)' when
                touched + non-empty but fails isValidHostUrl()
   - Both name and host SettingsInputs get showRequiredTag, the
     computed error, and onBlur={() => setXTouched(true)}.
   Errors are intentionally suppressed until blur so the form does not
   scold the user before they have had a chance to type anything.

3. Three snapshot tests call .blur() after .fill() to reveal errors
   - backend-add-name-only-disabled: focus+blur empty host → 'Host is
     required' appears below the Host field.
   - backend-add-whitespace-host-disabled: blur after fill('   ') →
     same 'Host is required' (whitespace counts as empty).
   - backend-add-invalid-url-disabled: blur after invalid URL fill →
     'Enter a valid URL...' appears below the Host field.
   The backend-add-blank-disabled snapshot is unchanged (neither field
   touched, no errors yet — correct for the fresh-open state).

New i18n keys: BACKEND$NAME_REQUIRED, BACKEND$HOST_REQUIRED,
BACKEND$HOST_INVALID (English only; other locales fall back to en).

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: prettier formatting on nameError / hostError ternaries

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: disable OAuth Login button until name and host are both valid

Previously the 'Login with OpenHands' button was enabled as soon as
a non-empty host was typed, even when the Name field was still blank.
This let users go through the full OAuth device-flow only to find they
still couldn't save because the name was missing.

Gate isDisabled on !name.trim() || !isValidHostUrl(host) so the button
stays disabled until the form is actually ready to save (modulo the
API key that OAuth itself will provide).

Update Flow 3 snapshot test to fill the name before asserting the
button becomes enabled, and update the test description accordingly.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: address PR review feedback (#450)

IPv6 parsing fixes (normalizeHost / isLocalAddress):
- normalizeHost: handle bracket notation [::1]:8080 (extract ::1),
  bare IPv6 addresses with multiple colons (use whole string as
  hostname), and regular host:port as before — prevents split(':')[0]
  from grabbing only the first segment of a multi-colon IPv6 address
- isLocalAddress: strip brackets before comparison; add :: (any-addr),
  ::ffff:127.x.x.x (IPv4-mapped loopback), fe80::/10 (link-local),
  fc00::/7 (unique local); tighten single-label check to exclude
  addresses that contain colons (bare IPv6 non-local addresses)

Mark fields touched on submit attempt:
- handleSubmit sets nameTouched + hostTouched when !canSubmit so
  inline errors appear for keyboard users who press Enter on an
  incomplete form

Snapshot workflow comparison-crash detection:
- Pass COMPARE_OUTCOME=${{ steps.compare.outcome }} to post-comment
- post-snapshot-comment.mjs reads COMPARE_OUTCOME and prepends a
  '[!WARNING]' block when the comparison step itself crashed
  (timeout/OOM) so the comment accurately reflects the run state
  instead of silently showing an incomplete/empty diff table

Remove unnecessary serial mode from backends-extended snapshot suite:
- Each test calls setupPage() with fresh state on its own Playwright
  page; no shared mutable state exists between tests, so serial is
  unnecessary and slows the suite

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-05-15 23:14:30 -04:00
Tim O'Farrellandopenhands b2b71855c6 feat(dev): surface dev-stack runtime services in agent system prompt (#503)
* feat(dev): surface dev-stack runtime services in agent system prompt

Add a structured 'runtime services' info object that the dev launchers
(`dev:safe`, `dev:automation`, `dev:docker`, and the published
`agent-canvas` binary) propagate to the frontend via
`VITE_RUNTIME_SERVICES_INFO`. The frontend renders it into a
`<RUNTIME_SERVICES>` markdown block and attaches it as
`AgentContext.system_message_suffix` on every `POST /api/conversations`.

This means agents start each conversation knowing exactly what services
exist in the current dev stack (ingress URL, automation backend URL +
`/api/automation` prefix, auth header, etc.), instead of having to probe
or — worse — assume `localhost:8000` is the automation server when it is
actually the Agent Server they are running inside of.

URLs are written from the agent's point of view: dockerless modes use
`localhost`, `dev:docker` uses `host.docker.internal`. When automation
isn't running in the current mode (e.g. `dev:safe`), the block says so
explicitly so agents know to skip `/api/automation` calls.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(runtime-services): address review feedback on PR #503

- Validate required `agentServerPort` in `buildRuntimeServicesInfo`;
  previously a missing port baked `http://localhost:undefined` into
  the agent's system prompt.
- Skip the automation entry when the supplied `automation` object has
  no `port` (e.g. a bare `{}` from a misconfigured launcher).
- Rename the JSON service key from `vite` to `frontend` and add a
  `kind: "vite" | "static"` discriminator + mode-aware description,
  so static-build dev stacks (`dev:docker`, the published binary, ...)
  no longer surface a misleading "Vite dev server" line in the agent
  system prompt. The renderer still accepts the legacy `vite` key.
- Anchor the "don't guess" warning to the actual agent-server URL from
  runtime info instead of hardcoded `localhost:8000`, since the
  agent-server uses different ports across dev modes (18000 in
  dev:safe, 8000 in dev:docker, ...).
- Plumb `frontendKind` through `buildAutomationRuntimeServicesInfo`
  and stamp `config.frontendKind` in `dev-with-automation.mjs::main`
  so both Vite spawn and static-build paths describe the frontend
  correctly.
- Expand AGENTS.md with the JSON schema of `VITE_RUNTIME_SERVICES_INFO`
  and a concrete example of the rendered `<RUNTIME_SERVICES>` block.
- Tests: assert the new URL-in-warning behavior, the new `frontend` /
  legacy `vite` rendering, the `agentServerPort`-required guard, the
  `automation: {}` skip, and the legacy `vitePort` alias.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-15 20:53:45 -06:00
2899c7b383 Fix static automation auth and switch LLM tool (#474)
* Fix static automation auth and switch LLM tool

* Allow automation SDK release lag in sync check

* chore: update baseline snapshots [skip ci]

* chore: trigger CI after snapshot update

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-05-15 15:30:29 +00:00
bcd50c9832 chat: conversation UX polish (input controls, drawer state, ellipsis, git controls) (#455)
* fix(conversation): cap chat column width at 800px

Replace responsive max-w-4xl / max-w-6xl with max-w-[800px] so the
middle column stays narrower on large viewports.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(chat): Connect Repo CTA and hide empty branch pill

- Use COMMON$CONNECT_REPO with FolderOpen when no repo/workspace is linked
- Show branch control only when selectedBranch is set (drop No Branch)
- Cap chat interface wrapper at max-w-[800px] without right-panel width coupling

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): refine input controls and local auth fallback

Improve chat input pills and model dropdown interactions while ensuring local agent-server auth uses the configured session key for default-local and cloud-proxy calls to avoid stale-key 401s.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): align attachment and placeholder control styling

Move the file-attach trigger into the chat action controls so it sits before Tools, and restyle it as a grey plus button with a circular hover state to match adjacent controls. Also align the chat input placeholder color with the same neutral control tone for visual consistency.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): simplify agent status labels and tone

Shorten English agent-status messages for the chat pill and align the status text color with the other grey controls for a more consistent compact UI.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): align model popover settings row styling

Add an LLM Settings action to the model popover and normalize its layout, spacing, and divider treatment to match existing dropdown menu patterns while keeping left-aligned positioning.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): restyle status controls and move send action

Make the agent-status control transparent by default with gray-to-white icon hover behavior, and move the submit button to the bottom-right controls area beside agent status.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): tighten spacing above git control bar

Reduce the top margin before the git control bar so it better matches the bottom spacing around the chat action controls.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): gate submit button on input content

Keep the send button inactive until the input has non-whitespace text, and align the revised button sizing/positioning with the bottom action row layout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): streamline overlays and remove legacy event rails

Unify chat control styling and overlay behavior so status/typing/scroll controls float above the thread without adding layout bars, and remove left-rail/checkmark affordances from grouped and generic event cards for a cleaner stream.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): tighten status indicator spacing

Reduce status indicator pill padding and icon size, and add right text padding to balance the compact layout in the chat control overlay.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): prioritize centered scroll control over loader

Keep the scroll-to-bottom control centered and visible whenever the user is away from the bottom, and use solid base/hover fills so it matches the updated chat surface styling.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): soften conversation event header styling

Use the lighter gray chat tone for conversation event header labels/icons and switch those labels to normal weight so grouped event rows match the updated control styling.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): refine markdown spacing and divider styling

Tighten markdown vertical rhythm in chat content, add a shared grey horizontal-rule renderer, and tune heading hierarchy to medium/compact styles for clearer structure without heavy emphasis.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): tighten vertical spacing in action event rows

Reduce stacked margins and paddings across grouped action rows, generic event cards, and collapsible thinking blocks so adjacent conversation entries read as a denser, more consistent stream.

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(design): add app gray palette reference artifacts

Capture the current gray color usage in dedicated SVG references, including both a curated palette and a strict exhaustive inventory for design and UI consistency work.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): align compact input overflow menus with menu conventions

Keep add-file pinned inline, collapse controls only when width truly runs out, and switch overflow entries to standard context-menu row/submenu patterns while preserving the send button layout at tight widths.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-panel): use list filter icon for older filters

Swap the older-conversations summary toggle icon to ListFilter so it matches the intended sidebar filter affordance.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): complete local workspace launch flow in git controls

Switch the local git control CTA from repository connection to workspace launching, including an above-button workspace menu and automatic add-workspace modal when none exist. This also captures the pending chat action/menu styling and test updates in the current working tree.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): relocate desktop vertical padding to input controls

Remove desktop top/bottom padding from the main chat panel and apply equivalent bottom spacing to the chat control area so the open repo/workspace controls and input footer keep consistent breathing room.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation): remove bottom margin from chat pane header

Drop the chat header bottom margin so the conversation title row sits flush with the content below.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): refresh git control bar immediately after Connect Repo

The "Connect Repo" empty-state in the chat input footer kept rendering
even after the Open Repository modal had successfully launched a clone
and the agent had reported the repository as ready. The bar would only
heal after a hard refresh (or never, on cloud backends).

Three independent bugs were stacking:

1. Optimistic update was writing to the wrong React Query cache key.
   `useUpdateConversationRepository.onMutate` called `setQueryData`
   with `["user", "conversation", id]` (3 elements), but
   `useUserConversation` reads from
   `["user", "conversation", id, backendId, orgId]` (5 elements).
   `setQueryData` requires an *exact* key match, so the update landed
   on an orphan cache entry that no observer ever read. Switched to
   `setQueriesData`/`getQueriesData` with the 3-element prefix so the
   optimistic write actually reaches the active query (Tanstack v5
   prefix-matches `setQueriesData` filters). Also normalized
   `branch`/`gitProvider` to `null` to match the shape produced by the
   server-side refetch and prevent identity-flicker between the two
   updates.

2. Cloud `batchGetCloudConversations` / `searchCloudConversations`
   ignored the local repo selection entirely. For local backends
   `toAppConversation` overlays `selected_repository`/`selected_branch`/
   `git_provider` from `localStorage`, but the cloud path returned the
   raw SaaS payload — and the SaaS often returns `null` for those
   fields until its own background hydration finishes. So every
   refetch (mutation invalidation, 30s poll, panel mount) overwrote
   the optimistic value with `null` and the bar snapped back to
   "Connect Repo". Added `overlayStoredRepoSelection` which fills only
   the `null` slots from local storage; populated server values still
   win, so we don't shadow real backend changes.

3. `updateConversationRepository` overwrote the entire metadata blob.
   `setStoredConversationMetadata` is replace-not-merge, so calling it
   with just `{selected_repository, selected_branch, git_provider}`
   silently dropped `selected_workspace` (the local-folder attach
   marker used by the Files tab to default to diff view, see the
   "Files tab diff-view default logic" note in `AGENTS.md`). Now reads
   the existing entry first and spreads it under the new repo fields.

Defense-in-depth changes:

- `useLocalGitInfo` now stays enabled until the conversation reports a
  *complete* repo tuple (`selected_repository` + `git_provider` +
  `selected_branch`), not just `selected_repository`. This lets the
  bar recover from partial-metadata cases (e.g. cloud hydration
  populates only the repo name first, or the user clones into a
  subdirectory of `working_dir`). The probe also gained a nested
  `find . -mindepth 2 -maxdepth 4 -name .git` fallback so a clone
  into `<workingDir>/<repo>/` is still detected after the direct
  `git remote get-url origin` in `<workingDir>` returns "no such
  remote 'origin'" (the agent-server pre-initialises every workspace
  as a worktree, so the parent directory always has a `.git` folder
  with no remote).

- `useUpdateConversationRepository.onSettled` invalidates
  `["local-git-info", conversationId]` so the bar re-probes
  immediately after a connect rather than waiting on the next 10s
  refetch tick.

- `git-control-bar.tsx`'s `hasRepository` predicate now keys off the
  *resolved* `selectedRepository` + `gitProvider` (which include the
  local-git probe's findings), not just the conversation field. This
  lets pull/push/PR buttons light up for local-workspace conversations
  whose repo metadata was inferred from `git remote`, matching what
  the repo + branch chips already showed.

Verification

I traced the failure mode by hitting the live agent-server directly:

  $ curl -s -X POST .../api/bash/execute_bash_command \\
      -H "X-Session-API-Key: \$KEY" \\
      -d '{"command":"git remote get-url origin", "cwd":"<workingDir>"}'

  git remote: error: No such remote 'origin'
  git rev-parse HEAD: ambiguous argument 'HEAD': unknown revision

confirming the worktree-without-remote shape that broke the direct
probe and forced the nested-find fallback.

Tests

  __tests__/hooks/mutation/use-update-conversation-repository.test.tsx
    - optimistically updates the cached conversation under the
      prefix-extended key used by useUserConversation
    - rolls back the prefix-keyed cache entry when the mutation rejects

  __tests__/api/cloud-conversation-service.test.ts (new)
    - overlays locally-stored repo selection onto
      batchGetCloudConversations results when the server returns nulls
    - prefers the cloud server values over locally-stored selections
      when present
    - leaves null entries untouched when the cloud server returns null
      for a missing conversation
    - returns an empty array without calling the proxy when no ids
      are provided
    - overlays repo selection on each item returned from
      searchCloudConversations

Wider sweep:

  npx vitest run __tests__/hooks/mutation \\
                __tests__/api/cloud-conversation-service.test.ts \\
                __tests__/api/conversation-metadata-store.test.ts \\
                __tests__/api/agent-server-adapter.test.ts \\
                __tests__/components/features/chat
  -> 22 files, 152 tests passed.

User-visible behavior after this change:

  1. Clicking Launch in the Connect Repo modal flips the bar to
     repo + branch chips immediately (optimistic update now reaches
     the active query).
  2. The bar stays flipped through the next refetch on cloud
     backends (overlay keeps the local selection visible until the
     SaaS catches up).
  3. Bar picks up nested clones within ~1s on local backends
     (local-git-info invalidation forces a re-probe instead of
     waiting on the 10s poll), and the nested-find fallback handles
     'clone into <workingDir>/<repo>/' flows.
  4. Pull/push/PR buttons now light up for local-workspace
     conversations whose remote was inferred from git remote.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation): make right-panel drawer state session-only

The right-side drawer's open/closed state (`isRightPanelShown` /
`hasRightPanelToggled`) was persisted in localStorage, which made
the panel feel sticky in a way users didn't expect — it would still
be open after reloads or revisits even though they wanted a clean,
focused chat view.

Move drawer state fully into the in-memory Zustand store so it:
- always starts closed on app load (or on opening a conversation
  after a restart),
- survives in-app navigation because Zustand stays alive across
  React Router transitions,
- only persists tab selection (`selectedTab`), which is the part
  users do want to come back to.

The legacy `rightPanelShown` field is silently stripped from older
persisted blobs by `sanitizeStoredState`, so old localStorage data
doesn't churn or leak into the new schema.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(ui): standardize three-dots ellipsis trigger across the app

Different surfaces had drifted to slightly different "more options"
buttons:
- conversation header used a 24x24 icon with a hardcoded fill color,
- conversation cards in the side panel used a separate square
  `ellipsis.svg` glyph,
- the conversation tab bar used a 20x20 icon with bespoke colors,
- LLM profile rows wrapped the icon in a bordered button with
  yet another color.

Promote `EllipsisButton` to be the canonical trigger and route every
inline variant through it so size (w-4 h-4 / 16x16), color
(`text-[#9299AA]`), and hover treatment (`hover:text-white
hover:bg-white/10`) stay consistent everywhere. Layout-only overrides
(e.g. translate, opacity-when-paused) flow through `className`, and a
`testId` escape hatch keeps the existing `profile-menu-trigger`
selector working.

The chat-input overflow button intentionally keeps its pill-shaped
custom variant; a doc comment on `EllipsisButton` calls that out so
future contributors don't replace it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: lint

* fix: failing tests

* refactor: package-lock.json

* refactor: package-lock.json

* refactor: remove artifacts

* refactor: update the code based on feedback

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: hieptl <hieptl.developer@gmail.com>
2026-05-15 14:43:16 +07:00
ed216c38aa feat(chat): add /model slash command for LLM profiles (#418)
* feat: integrate LLM profiles into settings route (PR C)

- Add LlmSettingsLocalView component for integrated profile management
- Extend SdkSectionSaveControl to expose form values for custom save flows
- Update LLM settings route to render profile list with create/edit views
- Add i18n keys for profile create/edit UI (CREATE_PROFILE, EDIT_PROFILE,
  PROFILE_CREATED, PROFILE_UPDATED, MODEL_REQUIRED, STATUS, BUTTON)
- Add test coverage for LlmSettingsLocalView

The integrated view shows:
- Profile list with active badge and action menu
- Add Profile button that opens create form
- Edit button that loads profile config and opens edit form
- Back/Cancel buttons to return to list view

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: address PR review feedback (#393)

- Improve mock typing with properly typed helper functions that provide all
  required React Query fields, eliminating incomplete 'as unknown as' casts
- Add integration test that verifies the save flow (fills in profile name,
  clicks save, verifies UI state transitions)
- Add component documentation noting future refactoring opportunity (extract
  useProfileForm, useProfileSave hooks for better testability)
- Document API key preservation behavior: currently preserves existing encrypted
  key in edit mode with no new key; note about potential 'Clear API Key' UX
  enhancement for future
- Document auto-derive name race condition: client-side uniqueness check uses
  render-time state, so concurrent profile creation by another client would
  result in server conflict error (handled gracefully)
- Document default export change in route file: LlmSettingsLocalView is now
  the default export; named export LlmSettingsScreen remains for embedded use

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: update llm-settings test to use named export

The default export of llm-settings.tsx changed to render LlmSettingsLocalView
(the profiles manager). The test needs to import the named export LlmSettingsScreen
to test the form component directly.

Co-authored-by: openhands <openhands@all-hands.dev>

* Fix LLM profile button/badge sizing and update typescript-client to v0.6.0

- BrandButton: Change padding from p-2 to px-3 py-2 for better text display
- ProfileRow: Increase active badge vertical padding from py-0.5 to py-1
- Update @openhands/typescript-client from commit SHA to v0.6.0 tag

Co-authored-by: openhands <openhands@all-hands.dev>

* Fix LLM settings to show regular form in cloud mode and empty form in create mode

- LlmSettingsRoute: Render LlmSettingsScreen (standard form) for cloud backends
  and LlmSettingsLocalView (profile manager) for local backends only
- LlmSettingsLocalView: Pass empty initial values in create mode to ensure
  fresh form fields, add key prop to force form remount between profiles
- Add unit tests for cloud vs local backend rendering
- Add unit tests for create mode empty form initialization

Co-authored-by: openhands <openhands@all-hands.dev>

* Fix edit mode form initialization to display profile values

- Fix initialValueOverrides logic to properly check for edit mode AND
  existing initialValues before using them
- Add prefix to edit mode key for clearer remount semantics
- Add unit tests verifying edit mode populates profile name correctly
- Add unit tests verifying getProfile is called with encrypted mode

Co-authored-by: openhands <openhands@all-hands.dev>

* Add debug logging to trace edit profile data flow

Co-authored-by: openhands <openhands@all-hands.dev>

* Fix edit profile config parsing - read from config directly not config.llm

The API returns profile config with llm settings at the top level
(config.model, config.api_key, config.base_url), not nested under
config.llm. Fixed the parsing to read directly from detail.config.

Co-authored-by: openhands <openhands@all-hands.dev>

* Handle profile rename during edit and update active profile

When editing a profile and changing its name:
1. Rename the profile first using ProfilesService.renameProfile
2. Then save the profile config to the new name
3. If the renamed profile was the active profile, re-activate it
   after the rename (since rename doesn't update active_profile)

This prevents creating duplicate profiles when just changing the name.

Co-authored-by: openhands <openhands@all-hands.dev>

* Fix package-lock.json to use https protocol for typescript-client

The lock file was using git+ssh:// protocol which causes Vercel build
failures since Vercel doesn't have SSH keys configured. Changed to
git+https:// and removed the integrity hash (git deps don't have one).

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(profiles): Available Profiles heading translation

* fix(profiles): use brand badge for active profile indicator

* feat(profiles): replace form heading with "Back to LLM profiles list"

* fix(profiles): unify profile-name validation and reject any whitespace

* fix(onboarding): persist onboarding LLM choice as an active profile

* refactor(profiles): drop redundant trim/wrapper after validator change

* feat(chat): add /model slash command for LLM profiles

* Add model profile slash completions

* chore: address model command review feedback (#418)

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: address model command follow-up review (#418)

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: Graham Neubig <neubig@gmail.com>
2026-05-15 01:19:55 -04:00
c371afec55 feat: LLM profiles route integration (PR C) (#393)
* feat: integrate LLM profiles into settings route (PR C)

- Add LlmSettingsLocalView component for integrated profile management
- Extend SdkSectionSaveControl to expose form values for custom save flows
- Update LLM settings route to render profile list with create/edit views
- Add i18n keys for profile create/edit UI (CREATE_PROFILE, EDIT_PROFILE,
  PROFILE_CREATED, PROFILE_UPDATED, MODEL_REQUIRED, STATUS, BUTTON)
- Add test coverage for LlmSettingsLocalView

The integrated view shows:
- Profile list with active badge and action menu
- Add Profile button that opens create form
- Edit button that loads profile config and opens edit form
- Back/Cancel buttons to return to list view

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: address PR review feedback (#393)

- Improve mock typing with properly typed helper functions that provide all
  required React Query fields, eliminating incomplete 'as unknown as' casts
- Add integration test that verifies the save flow (fills in profile name,
  clicks save, verifies UI state transitions)
- Add component documentation noting future refactoring opportunity (extract
  useProfileForm, useProfileSave hooks for better testability)
- Document API key preservation behavior: currently preserves existing encrypted
  key in edit mode with no new key; note about potential 'Clear API Key' UX
  enhancement for future
- Document auto-derive name race condition: client-side uniqueness check uses
  render-time state, so concurrent profile creation by another client would
  result in server conflict error (handled gracefully)
- Document default export change in route file: LlmSettingsLocalView is now
  the default export; named export LlmSettingsScreen remains for embedded use

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: update llm-settings test to use named export

The default export of llm-settings.tsx changed to render LlmSettingsLocalView
(the profiles manager). The test needs to import the named export LlmSettingsScreen
to test the form component directly.

Co-authored-by: openhands <openhands@all-hands.dev>

* Fix LLM profile button/badge sizing and update typescript-client to v0.6.0

- BrandButton: Change padding from p-2 to px-3 py-2 for better text display
- ProfileRow: Increase active badge vertical padding from py-0.5 to py-1
- Update @openhands/typescript-client from commit SHA to v0.6.0 tag

Co-authored-by: openhands <openhands@all-hands.dev>

* Fix LLM settings to show regular form in cloud mode and empty form in create mode

- LlmSettingsRoute: Render LlmSettingsScreen (standard form) for cloud backends
  and LlmSettingsLocalView (profile manager) for local backends only
- LlmSettingsLocalView: Pass empty initial values in create mode to ensure
  fresh form fields, add key prop to force form remount between profiles
- Add unit tests for cloud vs local backend rendering
- Add unit tests for create mode empty form initialization

Co-authored-by: openhands <openhands@all-hands.dev>

* Fix edit mode form initialization to display profile values

- Fix initialValueOverrides logic to properly check for edit mode AND
  existing initialValues before using them
- Add prefix to edit mode key for clearer remount semantics
- Add unit tests verifying edit mode populates profile name correctly
- Add unit tests verifying getProfile is called with encrypted mode

Co-authored-by: openhands <openhands@all-hands.dev>

* Add debug logging to trace edit profile data flow

Co-authored-by: openhands <openhands@all-hands.dev>

* Fix edit profile config parsing - read from config directly not config.llm

The API returns profile config with llm settings at the top level
(config.model, config.api_key, config.base_url), not nested under
config.llm. Fixed the parsing to read directly from detail.config.

Co-authored-by: openhands <openhands@all-hands.dev>

* Handle profile rename during edit and update active profile

When editing a profile and changing its name:
1. Rename the profile first using ProfilesService.renameProfile
2. Then save the profile config to the new name
3. If the renamed profile was the active profile, re-activate it
   after the rename (since rename doesn't update active_profile)

This prevents creating duplicate profiles when just changing the name.

Co-authored-by: openhands <openhands@all-hands.dev>

* Fix package-lock.json to use https protocol for typescript-client

The lock file was using git+ssh:// protocol which causes Vercel build
failures since Vercel doesn't have SSH keys configured. Changed to
git+https:// and removed the integrity hash (git deps don't have one).

Co-authored-by: openhands <openhands@all-hands.dev>

* chore(profiles): UI polish, shared validation, onboarding integration (#417)

* fix(profiles): Available Profiles heading translation

* fix(profiles): use brand badge for active profile indicator

* feat(profiles): replace form heading with "Back to LLM profiles list"

* fix(profiles): unify profile-name validation and reject any whitespace

* fix(onboarding): persist onboarding LLM choice as an active profile

* refactor(profiles): drop redundant trim/wrapper after validator change

* Fix LLM profile route mocks and warnings

* chore: update baseline snapshots [skip ci]

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: Vasco Schiavo <115561717+VascoSch92@users.noreply.github.com>
Co-authored-by: Graham Neubig <neubig@gmail.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-05-15 02:46:35 +00:00
Graham Neubigandopenhands a4089e0e4a Default user launchers to static frontend (#434)
Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-14 18:18:36 +00:00
a1befdd89d Fix stale backend health errors after recovery (#424)
Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: hieptl <hieptl.developer@gmail.com>
2026-05-14 11:59:14 +07:00
Hiep Le 05d226a470 fix(frontend): cap backend-health retries and persist failure state (#411)
* fix: cap backend-health retries and persist failure state

* refactor: update the code based on feedback
2026-05-13 14:03:29 +07:00
921ea743ef Fix MCP marketplace: deletes that duplicate, renames on every edit, Slack-twice, and Tavily that silently no-ops (#388)
* Fix MCP server delete causing duplicates and not removing entries

The agent-server's PATCH /api/settings applies agent_settings_diff via
deep-merge (see openhands.agent_server.persistence.models._deep_merge in
the SDK). For scalar fields that's fine, but mcp_config.mcpServers is a
name-keyed map and deep-merge cannot remove keys: a diff that omits a
server leaves the stale key behind, and a diff whose generated names
shift after a deletion produces duplicate entries pointing to the wrong
config.

The frontend's toSdkMcpConfig regenerates server names from a shared
counter on each save, so after deleting any server the resulting key
set shifts and the merge both fails to delete the target and creates
duplicates of the surviving servers.

Compensate inside SettingsService.saveSettings by sending a
{mcp_config: null} PATCH ahead of the real write whenever the diff
sets mcp_config to a non-null value. null is not a dict, so the
deep-merge takes the replace branch; the follow-up PATCH then writes
the new value into a freshly-cleared field. Skip the pre-clear when
the caller is already wiping mcp_config (null) — a single PATCH
already replaces in that case.

Co-authored-by: openhands <openhands@all-hands.dev>

* Stop bumping MCP suffix numbers on unrelated edits

toSdkMcpConfig used a single counter shared across the sse/shttp/stdio
loops, so a stdio server named 'myname' became 'myname_1' when any
sse/shttp entry was persisted ahead of it and got renamed to 'myname_2'
the moment another sse server was added. Every edit that changed the
count of any other server type would shift the suffix on everything
that came after it in the iteration order, which is exactly the
'numbers change every time I edit' behaviour the user hit.

Suffix names only when the same base actually collides, tracked
per-base against the running output dict. Bare 'sse'/'shttp' stay bare
unless there's a real duplicate within their own type, and stdio names
stay verbatim regardless of how many other server types exist.

Co-authored-by: openhands <openhands@all-hands.dev>

* Make Tavily install work, allow multiple instances of the same MCP entry

Two related marketplace bugs:

1. Clicking an already-installed catalog tile (e.g. Slack) opened
   the install modal in *edit* mode, so saving overwrote the
   existing server instead of adding a second one. There was no way
   to install two Slack workspaces, two Postgres connections, etc.

2. Tavily used a fake 'tavily-builtin' template kind that called
   saveSettings({ search_api_key }). That field is not part of
   agent_settings_diff / conversation_settings_diff, isn't in
   APP_PREFERENCE_FIELDS, and isn't forwarded by saveCloudSettings,
   so it was silently dropped on both backends. The SDK has no
   first-class Tavily integration either — the 'wires up the
   Tavily MCP server automatically' comment was aspirational.

Make the marketplace install modal strictly add-only: editing an
existing server already goes through CustomServerEditor from the
installed-server-card's edit button, so the modal's edit path was
redundant and conflicting. Combined with the per-base name
collision suffixing in toSdkMcpConfig, the user can now install a
second Slack and it lands as 'slack_1' alongside the original
'slack' without clobbering it.

Convert the Tavily catalog entry to a regular stdio MCP server
(npx -y tavily-mcp + TAVILY_API_KEY env) so it goes through the
same mcp_config write path as every other catalog entry and works
identically on local and cloud backends.

Strip the now-dead tavily-builtin scaffolding (the MarketplaceTemplate
union variant, the ExistingInstall discriminated union and
isMcpInstall helper, findInstalledMatch's tavily branch, the
InstalledServerCard catalogIdOverride prop, the InstalledServersSection
virtual Tavily card, MCPPage's Tavily-specific Configure/Remove
plumbing, and the marketplace-card transport-label case).

Co-authored-by: openhands <openhands@all-hands.dev>

* Restore mcp_config on rollback when the second PATCH fails

Address review feedback (PR #388) on the two-step PATCH atomicity gap.

The pre-clear (`mcp_config: null`) is destructive at the backend.
If the follow-up write fails after the clear succeeds, the user's
MCP config was previously left silently empty — bad data-loss UX
for what should be an idempotent retry.

Before pre-clearing, snapshot the previous mcp_config in raw SDK
shape (read directly from `fetchCloudSettings` for cloud or
`fetchSettingsFromApi` for local — `getSettings` returns the GUI's
parsed MCPConfig with empty-array defaults, which is not safe to
round-trip back). On second-write failure, attempt a best-effort
rollback PATCH that restores the snapshot, then re-throw the
original error so react-query mutations surface the failure to the
user.

The rollback is intentionally single-shot (no withRetry) — we want
the original error to surface promptly. Rollback errors are
swallowed so they don't shadow the user-actionable failure.

Three new tests cover:
- Successful rollback on cloud-backend second-write failure
- Successful rollback on local-backend second-write failure
- No bogus rollback PATCH when there was nothing to snapshot
  (first-time install where the snapshot has no mcp_config)

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: Remove PR-only artifacts

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: allhands-bot <allhands-bot@users.noreply.github.com>
2026-05-12 13:42:28 -07:00
0569bd77ec feat: add LLM profiles API layer and React Query hooks (#387)
* feat: add LLM profiles API layer and React Query hooks

This PR adds the foundational data layer for the LLM profiles feature:

## API Layer
- ProfilesService: Thin wrapper around SDK ProfilesClient with methods
  for list, get, save, delete, rename, and activate profile operations
- Re-exports SDK types for consumer convenience

## React Query Hooks
- useLlmProfiles: Query hook for listing all profiles
- useSaveLlmProfile: Mutation hook for creating/updating profiles
- useDeleteLlmProfile: Mutation hook for deleting profiles
- useRenameLlmProfile: Mutation hook for renaming profiles
- useActivateLlmProfile: Mutation hook for activating a profile

All mutation hooks properly invalidate both profile list and settings
caches on success, and disable global toasts (consumers handle errors).

## Utilities
- deriveProfileNameFromModel: Derives a clean profile name from model
  strings (e.g., 'openai/gpt-4' -> 'gpt-4')
- PROFILE_NAME_PATTERN: Validation regex for profile names

## Tests
- 47 tests covering all new functionality
- API service method tests
- Hook behavior tests (success, error handling, cache invalidation)
- Utility function tests

Part 1 of LLM Profiles feature (PR A from split plan).
No UI changes - this is purely a data layer addition.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: address PR review feedback

- Remove client.close() calls for consistency with other services
  (SettingsService, SecretsService don't call close())
- Use SETTINGS_QUERY_KEYS.personal() instead of .all for precision
- Add ActiveBackendProvider wrapper in useLlmProfiles tests
- Add test for query key including backend.id and orgId
- Add test for backend-switch cache isolation
- Fix truncation test to actually exercise trailing-dash removal
- Add test for model names that sanitize to empty string

Addresses review feedback from all-hands-bot.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: Remove PR-only artifacts

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: allhands-bot <allhands-bot@users.noreply.github.com>
2026-05-12 16:19:28 -04:00
2281f20f3e feat: Add OAuth 2.0 Device Flow authentication for OpenHands Cloud backends (#381)
* feat: Add OAuth 2.0 Device Flow authentication for OpenHands Cloud backends

Implements one-click login for cloud backends using the OAuth 2.0 Device
Authorization Grant (RFC 8628). When adding a cloud backend with a known
OpenHands Cloud host (*.all-hands.dev, *.openhands.dev), users can click
'Login with OpenHands' to authenticate via browser instead of manually
copying their API key.

Changes:
- Add device-flow-client.ts with startDeviceFlow() and pollForToken()
- Add useDeviceFlow React hook for managing auth state in components
- Add DeviceFlowAuth component with auth UI states (idle, starting,
  awaiting_authorization, success, error)
- Update BackendForm to show device flow auth for cloud backends
- Add i18n translations for all device flow UI strings

Closes #379

* feat: Show device flow login for all cloud backends with clearer UI

- Remove restriction to only known OpenHands Cloud hosts - device flow
  is now available for all cloud backends (including self-hosted)
- Add clear 'OR' divider between login button and manual API key entry
- Add link to API key documentation for manual key generation
- Add new i18n keys: LOGIN_OR, KEY_DOCS_HINT, KEY_DOCS_LINK

* feat: Always show login button for cloud backends, disable when no host

- Login button, OR divider, and manual API key input are now always
  visible when cloud backend type is selected
- Login button is disabled until a valid host URL is entered
- Improves UX by showing the full auth options upfront

* fix: Keep login button visible when typing custom cloud host URL

The kind inference was incorrectly downgrading from 'cloud' to 'local'
when typing a host URL that didn't match known OpenHands Cloud patterns.
Now the inference only upgrades to 'cloud' when a known pattern is
detected, but never downgrades - allowing users to type any custom
cloud host URL while keeping the login button visible.

* fix: Use cloud proxy for device flow to avoid CORS issues

For known OpenHands Cloud hosts (*.all-hands.dev, *.openhands.dev),
device flow requests are now routed through the local agent-server's
cloud-proxy endpoint. This avoids CORS errors when the browser tries
to make direct cross-origin requests to the cloud backend.

Self-hosted instances still use direct requests, assuming they have
CORS properly configured.

* fix: Always use proxy for device flow and fix kind inference regression

1. Device flow now always uses proxy for all hosts (not just known cloud
   hosts). This avoids CORS issues for any custom backend that supports
   device flow.

2. Fix regression where typing a local address (e.g., 127.0.0.1) would
   not switch from cloud to local type. The kind inference now:
   - Auto-infers kind from host in add mode (initial behavior)
   - Only prevents downgrade when user explicitly clicked the cloud
     radio button (not when cloud is just the default)
   - Tracks explicit user selection separately from initial default

* fix: Address PR review feedback for device flow security and RFC compliance

Security fixes:
- Fix isOpenHandsCloudHost() to use URL hostname extraction instead of
  substring matching, preventing attacks like all-hands.dev.evil.com
- Add URL validation in handleStartAuth to check for credential injection
- Sanitize error messages to avoid exposing server error details

RFC 8628 compliance:
- Add required grant_type parameter to token requests
- Make verification_uri_complete optional per RFC Section 3.2
- Build verification_uri_complete if not provided by server

Robustness improvements:
- Validate polling interval to at least 1 second
- Cap slow_down interval to MAX_INTERVAL_MS to prevent DoS
- Open popup on user click to avoid popup blockers

Accessibility:
- Add role='status' and aria-live='polite' to status containers
- Add role='alert' to error container

* fix: Pass abort signal to makeProxiedRequest fetch call

Address review feedback: the abort signal is now properly passed through
makeProxiedRequest to the underlying fetch call, allowing in-flight
proxied requests to be cancelled immediately when the user cancels.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: Address comprehensive review feedback for device flow auth

Security fixes:
- Add URL validation to prevent XSS via javascript: URLs
- Validate verification URLs have https: protocol before use in popup and links
- Add type validation for slow_down interval to prevent NaN tight loops

RFC 8628 compliance:
- Fix slow_down to increment by 5 seconds per Section 3.5 (not double)
- Validate interval is number, finite, and positive before using server value

Robustness:
- Network errors now continue polling instead of failing immediately
- Wrap sleep in try-catch for consistent abort handling
- Add cleanup effect to close popup on unmount

Code cleanup:
- Remove dead userSelectedCloud state (was unreachable)
- Add defensive programming comment for cancellation check
- Fix onSuccess effect to include deviceFlow.reset in deps

Tests:
- Add DoS protection test (caps interval at 30s)
- Add type confusion test (rejects non-numeric interval)
- Add RFC 8628 +5s increment test
- Add network error retry test
- Add unmount cleanup test

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: Remove PR-only artifacts

* fix: Remove noopener from popup to maintain window reference

The 'noopener' option causes window.open() to return null, which means
we lose the reference to the popup and can't update its location when
the verification URL becomes available. This was causing a blank page
to appear instead of the device flow auth page.

Removed 'noopener' from the initial popup open call so we can maintain
the reference and update popupRef.current.location.href when the
verification URL arrives from the device flow.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: allhands-bot <allhands-bot@users.noreply.github.com>
2026-05-12 15:45:52 -04:00
1cd45306df Route agent-server calls through TypeScript client (#278)
* Route agent-server calls through TypeScript client

Co-authored-by: openhands <openhands@all-hands.dev>

* Remove obsolete agent-server API wrappers

Use @openhands/typescript-client directly where practical and remove thin src/api wrappers that no longer add domain-specific behavior. Keep app-specific adapter/config/cloud logic in place and update tests to mock typed clients directly.

Co-authored-by: openhands <openhands@all-hands.dev>

* Inline remaining trivial client wrappers

Remove remaining one-line wrappers around @openhands/typescript-client calls in the local conversation and workspace paths. Keep domain-specific API adapters where they still perform cloud routing, response adaptation, caching, or app-specific defaults.

Co-authored-by: openhands <openhands@all-hands.dev>

* Remove more no-op API facades

Delete additional short methods that only returned constants, no-ops, or delegated directly to another adapter. Keep the remaining services focused on paths that still perform cloud routing, response adaptation, client setup, or app-specific behavior.

Co-authored-by: openhands <openhands@all-hands.dev>

* Address review feedback after latest merge

Co-authored-by: openhands <openhands@all-hands.dev>

* Address follow-up TypeScript client review

Co-authored-by: openhands <openhands@all-hands.dev>

* Fix static dev frontend session auth

Co-authored-by: openhands <openhands@all-hands.dev>

* Handle unexpected conversation response shapes

Co-authored-by: openhands <openhands@all-hands.dev>

* Format conversation response handling

Co-authored-by: openhands <openhands@all-hands.dev>

* Sanitize conversation response fields

Co-authored-by: openhands <openhands@all-hands.dev>

* Show conversation load failures full-screen

Co-authored-by: openhands <openhands@all-hands.dev>

* Harden conversation file path handling

Co-authored-by: openhands <openhands@all-hands.dev>

* Guard SDK settings page against malformed schemas

Co-authored-by: openhands <openhands@all-hands.dev>

* Avoid using Vercel origin as default backend

Co-authored-by: openhands <openhands@all-hands.dev>

* Fix preview API fallback behavior

Co-authored-by: openhands <openhands@all-hands.dev>

* Refresh file previews after workspace edits

Co-authored-by: openhands <openhands@all-hands.dev>

* test: stabilize backend selector cleanup

Co-authored-by: openhands <openhands@all-hands.dev>

* Remove unrelated PR changes

Keep the PR scoped to direct agent-server API removal and TypeScript SDK migration by reverting unrelated backend fallback, static dev auth, conversation error UI, SDK settings hardening, and test-only cleanup changes.\n\nCo-authored-by: openhands <openhands@all-hands.dev>

* Add PR 278 agent run GIF

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: hieptl <hieptl.developer@gmail.com>
2026-05-12 21:32:36 +07:00
Hiep Le 68c9b77147 refactor: remove dead integrations page and third-party Git API layer (#367) 2026-05-12 18:58:19 +07:00
Hiep Le 801c166074 fix(frontend): read cloud VSCode URL from sandbox.exposed_urls (#361)
* fix: read cloud VSCode URL from sandbox.exposed_urls

* fix: lint
2026-05-12 15:57:05 +07:00
Hiep Le 6bb94b8bf0 fix: stop tripping SaaS 500 on cloud conversation history (#360) 2026-05-12 15:30:03 +07:00
Hiep Le 5dc765faf7 fix(frontend): stop mutating cloud current_org_id on backend select (#359)
* fix: stop mutating cloud current_org_id on backend select

* fix: failing tests
2026-05-12 14:11:11 +07:00
Robert Brennanandopenhands 860309f457 chore: remove dead OSS-cleanup leftovers (~560 LoC) (#290)
The OSS cleanup left these files behind, but all of them have zero
runtime references on the current main:

- `src/hooks/mutation/use-accept-tos.ts` (hosted /api/accept_tos flow)
- `src/hooks/mutation/use-logout.ts` (hosted /api/unset-provider-tokens)
- `src/api/invariant-service.ts` (hosted /api/security/* endpoints)
- `src/api/api-keys.ts` (hosted user-API-keys, distinct from the cloud
  org-scoped path that lives in cloud/organization-service.api.ts)
- `src/components/features/settings/new-api-key-modal.tsx` (never
  rendered) + its only consumer's helper `api-key-modal-base.tsx`
- `src/mocks/api-keys-handlers.ts` (mocks an endpoint nothing calls)

Cascading from those removals, `src/api/open-hands-axios.ts` and its
two tests also become dead — every other call site in src/ has been on
`createHttpClient()` (from typescript-client) for some time, and the
custom array-paramsSerializer it provided isn't used by any live caller
(the README example still pointed at the old singleton, also updated).

Verification:
- `grep -rn openHands src/` returns no real-code matches after this commit
- `npm run typecheck` passes (TS would have flagged any stranded import)
- `npm run lint` passes (0 errors, same 10 pre-existing warnings)
- `npx vitest run` — 1859 passed | 12 skipped | 9 todo (6 fewer than
  before, exactly matching the removed openHands-axios tests + the
  two now-stale `vi.mock('#/api/open-hands-axios', ...)` stubs that
  were never asserting anything)

Tiny non-deletion touches:
- `src/api/cloud/proxy.ts` and
  `src/api/automation-service/automation-service.api.ts` had comments
  referring to the gone `openHands` axios — updated to point at
  `createHttpClient()` instead, since the rationale (resolver routes
  to the active backend, this call needs the local one) is unchanged.
- `src/api/README.md` example switched from `openHands.get(...)` to
  `createHttpClient().get(...)` so new services follow the live
  pattern.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-11 19:02:18 -07:00
Robert Brennan bac71ea9a4 Revert "fix: show folder name for local-only repos in git control bar (#322)" (#357)
This reverts commit c26e3ac571eb7ea0be5a3d44e16e875d046a7b6b.
2026-05-11 18:50:07 -07:00
Robert Brennanandopenhands 18c45316a5 feat(backends): remember most recent conversation per backend (#354)
When switching from backend A to backend B while a conversation is
active in A, jump to B's most recently selected conversation instead
of dropping the user on the home page. On any other route (settings,
home, automation list, …), stay on the same path under the new
backend.

Per-backend memory is keyed by (backendId, orgId) and persisted in
localStorage so it survives reloads.

Also fixes a 404 race that surfaced on every conversation-route
backend switch. React Router defers URL transitions while
`setActive` runs at sync priority through `useSyncExternalStore`,
so the conversation route re-rendered once with the new backend id
and the old conversation id — `useUserConversation` then fetched
the old id from the new backend and toasted 'conversation not
available'. Two complementary fixes:

  - BackendSelector now `await`s `navigate(...)` before calling
    `setActive`, so the route transition has committed (and the
    conversation route has unmounted) before backend-scoped query
    keys flip.
  - `routes/conversation.tsx` mirrors the `mountedBackendId` ref
    guard from `routes/automation-detail.tsx`: while a backend
    switch is in flight, the route suppresses the 404 toast and the
    "remember last conversation" write (so we never overwrite B's
    slot with A's stale id), and clears the slot on a confirmed 404
    so the next switch doesn't revisit a deleted conversation.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-11 18:23:11 -07:00
Robert Brennanandopenhands 2cbfa80f32 fix(files-tab): default to Diff view when a workspace or repo is attached (#349)
* fix(files-tab): default to Diff view when a workspace or repo is attached

Previously, the Files tab only defaulted to Diff view when the user
explicitly picked a Git repository on the home page (`selected_repository`
on the conversation). Conversations created via the workspace picker —
which sets a local working directory but not `selected_repository` —
fell through to Files view, even though those conversations also have a
pre-existing working tree the user came in to inspect.

Treat workspace selection the same as repo selection:

- Extend `ConversationMetadata` with an optional `selected_workspace`
  field and persist it client-side when the home-page workspace picker
  supplies a `workingDirOverride` (the agent-server runtime has no
  concept of workspace selection, so this stays client-side, mirroring
  how repo metadata already works).
- Rename `useIsGitRepo` → `useHasAttachedSource` and broaden it to
  return true when *either* `selected_repository` is set on the active
  conversation *or* a `selected_workspace` was stored at creation time.
- Files tab now keys its diff-view default off `hasAttachedSource`
  instead of `isGitRepo`. The existing `useHasGitCommits` probe still
  gates the default off when the attached working tree has no commits
  yet, so non-git workspaces and unborn-HEAD repos correctly fall back
  to Files view.

The user's persisted per-conversation toggle choice continues to win
over the computed default.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore(files-tab): align comments and test names with attached-source semantics

Follow-up cleanup to the diff-default fix. No behaviour change.

- `useHasGitCommits` jsdoc + inline comment now describe all three
  "no diff base" cases (unborn HEAD, non-git workspace, other git
  error) instead of singling out empty repos, and reference
  `useHasAttachedSource` as the gating signal.
- Files-tab test descriptions / inline comment dropped the
  "git repo" framing in favour of "attached source", matching the
  hook the suite is exercising.
- Added an AGENTS.md note capturing the design and warning future
  agents off the filesystem-probe approach that was tried (and reverted)
  in earlier passes at this logic.
- Minor prettier reflow in `use-has-attached-source.ts` picked up by
  `eslint --fix`.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-11 17:30:59 -07:00
Rohit Malhotraandopenhands 05deacbd35 fix(msw): use wildcard URL patterns to match absolute URLs (#344)
* fix(setup): always run npm ci to ensure hooks have dependencies

The on_stop.sh hook runs npm run lint and npm test, which require
node_modules to be installed. Previously, setup.sh only ran npm ci
if node_modules was missing, which could fail if:
- node_modules existed but was incomplete/corrupt
- setup.sh hadn't completed before hooks ran
- package-lock.json was updated but node_modules was stale

Now npm ci always runs during setup, ensuring dependencies are
consistently installed when OpenHands begins working with the repo.

Co-authored-by: openhands <openhands@all-hands.dev>

* Add session_start hook to run setup.sh automatically

The hooks.json was missing a session_start hook, which meant setup.sh
(which installs npm packages via 'npm ci') was never run automatically
when OpenHands began working with this repository.

This caused the stop hook to fail because npm packages weren't installed.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(msw): use wildcard URL patterns to match absolute URLs

MSW handlers using relative paths (e.g., '/api/settings') only intercept
requests made to the same origin as the test runner. When VITE_BACKEND_BASE_URL
is configured (e.g., in a local .env file), the code makes requests to absolute
URLs like 'http://127.0.0.1:8000/api/settings', which MSW treats as a different
origin and lets pass through, causing ECONNREFUSED errors.

This fix updates MSW handlers to use wildcard patterns (e.g., '*/api/settings')
which match both relative paths AND absolute URLs, ensuring tests pass
regardless of whether VITE_BACKEND_BASE_URL is configured.

Changes:
- Update src/mocks/settings-handlers.ts to use '*/' prefix on all routes
- Update src/mocks/secrets-handlers.ts to use '*/' prefix on all routes
- Update test files that use server.use() with test-specific handlers

This resolves the discrepancy between CI (no .env file) and local development
(with .env file containing VITE_BACKEND_BASE_URL).

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(msw): update all remaining handler files with wildcard URL patterns

Additional handler files updated:
- src/mocks/conversation-handlers.ts
- src/mocks/git-repository-handlers.ts
- src/mocks/api-keys-handlers.ts
- src/mocks/auth-handlers.ts
- src/mocks/automation-handlers.ts
- src/mocks/feedback-handlers.ts
- src/mocks/task-suggestions-handlers.ts

Test files updated:
- __tests__/api/option-service.test.ts
- __tests__/components/modals/settings/model-selector-openhands.test.tsx

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-11 18:50:35 -04:00
Rohit Malhotraandopenhands 24da834da3 fix: guard against null provider in useUrlSearch hook (#341)
* fix: guard against null provider in useUrlSearch hook

Prevent unnecessary cloud proxy requests when the provider is null/undefined
(e.g., before providers have loaded from settings).

The useUrlSearch hook was calling GitService.searchGitRepositories()
without validating that the provider was truthy first. When the parent
component passed undefined (from providers[0] when array is empty),
the request would be sent to the cloud proxy with an invalid provider.

Changes:
- Update type signature to accept Provider | null | undefined
- Add early return guard when provider is falsy
- Clear results when provider becomes null

* fix: add defensive guards in GitService for invalid providers

Add a second layer of defense at the GitService level to prevent
cloud proxy requests with invalid providers (null, undefined, empty
string, or stringified 'undefined'/'null').

This fixes the installations search API being called with
'provider=undefined' even when hooks have enabled guards.

Changes:
- Add isInvalidProvider() guard function
- Add guards to all GitService methods that take a provider param
- Return empty results instead of making invalid API requests
- Add comprehensive tests for the guards

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-11 17:14:33 -04:00