Commit Graph
6 Commits
Author SHA1 Message Date
Rohit Malhotraandopenhands 60bc93f0e9 Add backend management specs and @spec annotations (#727)
* Add backend management specs and @spec annotations

- Curate specs/backend-management.md with 4 behavioral specs (BM-001–BM-004)
- Add @spec comments to source and test files for traceability
- Add spec file convention note to AGENTS.md

Co-authored-by: openhands <openhands@all-hands.dev>

* Remove duplicate BM-002 test (non-conversation redirect)

The 'stays on settings' case was covered by two tests; keep the one
with the clearer assertion and drop the redundant copy.

Co-authored-by: openhands <openhands@all-hands.dev>

* Parameterize BM-002 tests with it.each

Collapse 3 identical-structure redirect tests into a single
parameterized it.each covering conversation detail, automation detail,
and non-ID routes.

Co-authored-by: openhands <openhands@all-hands.dev>

* Document @spec tagging convention in AGENTS.md

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-21 18:08:53 -04:00
Rohit Malhotraandopenhands e6d12e3fff fix(BM-001): auto-switch active backend on addBackend (#714)
* spec(BM-001): add spec + failing tests for auto-switch on connect

Adding a backend should automatically switch the active selection to it.
Currently addBackend registers the entry but leaves the user on the
previous backend, forcing a manual switch.

- specs/backend-management.md: BM-001 definition
- Two new test cases (cloud + local) that assert the active backend
  changes after addBackend — both fail against the current implementation.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(BM-001): auto-switch active backend on addBackend

addBackend now calls setActiveSelection after registering the new entry,
so the user lands on the backend they just connected — whether via the
manual host+key form or the cloud OAuth device-flow login.

- src/contexts/active-backend-context.tsx: one-line fix in addBackend
- Updated add-backend-modal test to assert the new active selection
- Updated backend-selector tests that used addBackend purely as setup
  to reset active to the default local backend afterward

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: remove noisy @spec markers from test setup comments

Keep @spec BM-001 only where it marks the implementation or directly
asserts the spec behavior. Setup-only adjustments (resetting active
backend after addBackend) are incidental — plain comments suffice.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: consolidate @spec BM-001 to one test + one implementation site

The spec marker belongs in exactly two places: the line that implements
the behavior and the single test that verifies it. Removed the redundant
local-backend variant (same code path as cloud) and dropped @spec labels
from the modal test (its assertion stays, just without the tag).

Co-authored-by: openhands <openhands@all-hands.dev>

* refactor: remove setActive workarounds from backend-selector tests

Instead of manually resetting active state after addBackend, tests now
work with the auto-switch naturally:

- Local backend tests: click the seeded default 'Local' (which is no
  longer active after auto-switch) to trigger the switch.
- Cloud org tests: add a local backend after the cloud one so the last
  auto-switch lands on local, leaving cloud backends unselected and
  their org rows visible in the dropdown.

No ctx.setActive(DEFAULT_LOCAL_BACKEND_ID) calls remain.

Co-authored-by: openhands <openhands@all-hands.dev>

* refactor: extract shared seed constants in backend-selector tests

SEED_LOCAL_1 and SEED_CLOUD_PRODUCTION replace 12+4 identical inline
config objects. The two remaining inline blocks have different apiKey
values and correctly stay as-is.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-21 18:49:20 +00:00
c224d24a9b fix: cloud conversation resume + archived/error sandbox states (#500)
* fix: resume cloud conversations stuck in starting status

Two bugs prevented cloud-backend conversations from resuming properly:

1. useActiveConversation hard-coded a 30 s refetch interval. When a
   cloud sandbox is paused and auto-starts on access, conversation_url
   is null until the sandbox is ready. The WebSocket can't open without
   a URL, so curAgentState stays at LOADING ('starting status') for up
   to 30+ seconds — or forever if the user gave up before the next poll.
   Fix: use the query-state callback form of refetchInterval and drop
   to 3 s whenever conversation_url is null (mirrors the 3 s cadence of
   task polling), falling back to 30 s once the URL is available.

2. updateConversationExecutionStatusInCache called setQueryData with a
   3-element key ["user", "conversation", id] that no longer matches
   the 5-element key stored by useUserConversation
   ["user", "conversation", id, backend.id, orgId] after the
   per-backend cache isolation was added. Optimistic status writes after
   manual pause/resume were silently dropped.
   Fix: switch to setQueriesData with { queryKey: [...] } prefix
   matching so the update hits whichever (backend, org) variant is live.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: auto-resume cloud sandbox when conversation_url is null

Faster polling (prev commit) was not enough. The cloud API returns
conversation_url=null when the sandbox is paused/stopped, and a GET
request alone does not wake it up — you have to POST a new start task
(with sandbox_id to reuse the existing sandbox) and wait for it to
become READY.

Add a useEffect in AppContent that fires once per unique conversation.id
after the initial fetch:
  • skips if not a cloud backend
  • skips if conversation_url is already set (sandbox running)
  • skips if sandbox_id is null (nothing to resume)
  • guards against re-triggering within the same route-mount via a ref

On trigger it calls createConversation(sandbox_id), which POSTs
POST /api/v1/app-conversations with the sandbox_id to the cloud, gets
back a WORKING start task, then navigates to /conversations/task-{id}.
useTaskPolling drives the task to READY and redirects to the real
conversation, now with a conversation_url the WebSocket can connect to.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: navigate back to conversation when resume task fails

When the cloud sandbox fails to start (e.g. 'Sandbox failed to start
within 120s'), the task reaches ERROR status. Previously the user was
left stranded at the task-{id} URL with only a toast to show for it.

Two changes:
1. Pass resumedFromConversationId in React Router navigation state when
   navigating to task-{id} for a cloud resume, so we know where to go
   back if the task fails.
2. In the task-error effect, read that state and navigate back to the
   original conversation (or /conversations if no originator is known).
   The resume effect's ref is still set so it will not re-trigger the
   resume on landing, preventing a retry loop.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use correct sandbox resume endpoint matching OpenHands

Root cause of 'Sandbox failed to start within 120s':
The previous fix called POST /api/v1/app-conversations with sandbox_id,
which is the 'create a new conversation' endpoint. The cloud treats this
as a full sandbox provisioning request with a 120-second cold-start
timeout that can fail on old/stale sandboxes.

The correct endpoint — matching OpenHands' SandboxService.resumeSandbox
and useSandboxRecovery — is POST /api/v1/sandboxes/{id}/resume, which is
a lightweight unpause that simply wakes the existing sandbox without
reprovisioning it.

Three changes:
1. Add SandboxStatus type ('PAUSED'|'RUNNING'|'STARTING'|'MISSING') and
   sandbox_status field to AppConversation, mirroring OpenHands'
   V1SandboxStatus. The cloud API already returns this field; adding the
   type makes it accessible in TypeScript.

2. Add resumeCloudSandbox(sandboxId) to the cloud service, calling
   POST /api/v1/sandboxes/{id}/resume via the cloud proxy — symmetric
   with the existing pauseCloudSandbox.

3. Update the resume effect in conversation.tsx:
   - Detect on sandbox_status === 'PAUSED' (more precise than
     conversation_url === null, which can be null for other reasons).
   - Call resumeCloudSandbox(sandbox_id) instead of createConversation.
   - Stay on the current URL after resume — no task navigation needed.
     The 3-second refetch interval in useActiveConversation polls until
     conversation_url populates, then the WebSocket connects normally.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: add ERROR to SandboxStatus to match OpenHands V1SandboxStatus

Complete enum is MISSING|STARTING|RUNNING|PAUSED|ERROR.

Co-authored-by: openhands <openhands@all-hands.dev>

* feat: archived/error sandbox state — read-only view and sidebar indicators

When a cloud conversation's sandbox_status is MISSING or ERROR it can
never be resumed. These two states now have first-class treatment:

Sidebar / conversation list:
- ConversationStatusDot gains an optional sandboxStatus prop.
  MISSING → gray 'paused' dot with tooltip 'Archived'.
  ERROR   → red 'error' dot with tooltip 'Error'.
  (ExecutionStatus visual is used unchanged for all other states.)
- ConversationCardHeader passes sandboxStatus to the dot and sets
  isConversationArchived on the title, which applies opacity-60.
- ConversationCard renders the existing ConversationStatusBadges pill
  ('Archived' or 'Error' pill badge) for MISSING and ERROR sandboxes.
- CompactConversationRow (collapsed sidebar) passes sandboxStatus to
  both the main dot and the tooltip-preview dot.
- conversation-panel.tsx passes sandbox_status from AppConversation to
  both card variants.

Conversation view (read-only):
- ChatInterface reads sandbox_status via useActiveConversation.
- When MISSING or ERROR, the InteractiveChatBox is replaced by a
  localised banner (title + description) explaining the history is
  read-only. The banner uses data-testid='archived-conversation-banner'
  for testing.
- The auto-resume effect in conversation.tsx already skips MISSING and
  ERROR because it only fires on sandbox_status === 'PAUSED'.

i18n:
  CHAT_INTERFACE$ARCHIVED_SANDBOX_TITLE/DESCRIPTION
  CHAT_INTERFACE$ERROR_SANDBOX_TITLE/DESCRIPTION

Co-authored-by: openhands <openhands@all-hands.dev>

* test: snapshot tests for archived/error sandbox conversation states

Three new Playwright visual snapshots in
tests/e2e/snapshots/archived-conversation.snapshot.spec.ts:

1. conversation-panel-with-archived-badges
   Navigates to /conversations; verifies five conversation cards are
   present; asserts archived-badge and error-badge are both visible;
   captures the conversation panel showing:
   - 'Archived Project'  → gray dot + 'Archived' pill + dimmed title
   - 'Errored Project'   → red dot + 'Error' pill + dimmed title

2. conversation-view-archived
   Navigates to /conversations/4 (sandbox_status: 'MISSING'); stubs
   WebSocket; asserts:
   - archived-conversation-banner is visible (read-only notice)
   - interactive-chat-box is absent (count 0)
   Captures the full chat interface.

3. conversation-view-sandbox-error
   Same as above for /conversations/5 (sandbox_status: 'ERROR');
   captures the 'Sandbox error' banner variant.

Supporting changes:
- src/api/agent-server-adapter.ts
  - Add sandbox_status?: string | null to DirectConversationInfo
  - Import SandboxStatus and map info.sandbox_status → AppConversation
    so the field is no longer silently null for all conversations
- src/mocks/conversation-handlers.ts
  - Add mock conversations 4 (MISSING) and 5 (ERROR) with sandbox_status
  - createConversationResponse now includes sandbox_status in the payload
- src/components/features/chat/chat-interface.tsx
  - Suppress ChatSuggestions ("Let's start building!") for archived
    conversations — showing task suggestions alongside a read-only
    banner is confusing and misleading
- src/components/features/conversation-panel/conversation-card/
  conversation-status-badges.tsx
  - Add data-testid="archived-badge" and data-testid="error-badge"
    so Playwright can assert on their presence without relying on text
- tests/e2e/snapshots/sidebar.snapshot.spec.ts
  - Update toHaveCount(3) → toHaveCount(5) to account for the two new
    mock conversations; existing sidebar snapshot baseline needs
    regeneration on main (intentional diff via update-snapshots label)

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: make sandbox_status optional in AppConversation

sandbox_status is a cloud-only field that local agent-server conversations
never carry. Existing test fixtures built AppConversation objects without
this field, causing TypeScript to error once it became required.

Making it optional (sandbox_status?: SandboxStatus | null) is the
semantically correct choice:
- The field is absent / null for every local conversation
- The adapter still explicitly maps it to null when unset
- ChatInterface reads it with ?? null so undefined is handled safely
- Partial<AppConversation> spreads in test factory functions no longer
  widen to SandboxStatus | null | undefined

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: Prettier formatting — multiline SandboxStatus union and ternary

- SandboxStatus type: expand single-line union to multi-line format
- ConversationCard: wrap sandboxStatus ternary in a multiline JSX block

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: preserve sandbox_status through requireDirectConversationInfo

The validation function that normalises raw API responses into
DirectConversationInfo was not copying sandbox_status, so it was silently
dropped every time a conversation came through the search or batch-get
code paths. This caused the conversation panel to never render the
archived/error badge pills, and the ChatInterface to always treat every
conversation as active (missing read-only banner for MISSING/ERROR sandboxes).

Fix: add sandbox_status: stringOrNull(item.sandbox_status) to the
mapping in requireDirectConversationInfo, mirroring the treatment of
execution_status.

The existing E2E tests for conversations 4 (MISSING) and 5 (ERROR)
were already asserting on archived-badge / error-badge presence and
archived-conversation-banner visibility, so they will now pass.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: don't connect WebSocket while cloud sandbox is PAUSED

When a cloud conversation is closed from the UI (pauseCloudSandbox is
called), the conversation's conversation_url is NOT cleared — it still
points to the old sandbox host. On the next navigation into that
conversation the WebSocket provider saw a non-null URL and immediately
tried to open a connection, which failed because the sandbox had not
yet woken up.

Two-part fix:

1. WebSocketProviderWrapper: suppress conversationUrl (treat it as null)
   while sandbox_status === 'PAUSED', so ConversationWebSocketProvider
   cannot compute a valid wsUrl until the sandbox is actually running.

2. useActiveConversation: add sandbox_status === 'PAUSED' as a
   fast-poll trigger alongside !conversation_url. The old check only
   fast-polled when the URL was absent; for paused sandboxes the URL is
   present but stale, so without this the hook would stay on the slow
   30-second interval while waiting for the sandbox to wake up.

Together these changes let the resume sequence complete correctly:
  navigate → sandbox PAUSED detected → resumeCloudSandbox called →
  fast-poll picks up RUNNING state → conversationUrl unblocked →
  WebSocket connects with a live sandbox host.

Co-authored-by: openhands <openhands@all-hands.dev>

* docs: document cloud PAUSED sandbox WebSocket gating in AGENTS.md

Co-authored-by: openhands <openhands@all-hands.dev>

* test: cover PAUSED sandbox gating and sandbox_status preservation

Three test suites covering the cloud conversation resume bug fixes:

1. agent-server-conversation-service.test.ts — three cases asserting
   that requireDirectConversationInfo preserves sandbox_status through
   batchGetAppConversations (PAUSED, RUNNING, absent → null).

2. websocket-provider-wrapper.test.tsx — five cases asserting that
   WebSocketProviderWrapper passes conversationUrl through when the
   sandbox is RUNNING or null (local backend), suppresses it to null
   when sandbox_status === 'PAUSED', and handles not-yet-fetched data.

3. use-active-conversation.test.ts — five cases asserting that the
   refetchInterval callback returns 3000 when sandbox_status is PAUSED
   (even with a non-null conversation_url) or when conversation_url is
   null, and 30000 in all other ready states.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(test): use null for execution_status fixture field

ExecutionStatus is a string enum — assigning the raw string literal
'idle' triggers TS2322. Null satisfies ExecutionStatus | null and is
irrelevant to what these tests actually exercise.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(lint): prettier format + add missing i18n fallbacks for 7 keys

Two issues from lint-staged pre-commit hook:

1. Prettier: the compound refetchInterval condition in use-active-
   conversation.ts was too long for one line — broke across three lines.

2. Translation completeness: CHAT_INTERFACE$ARCHIVED_SANDBOX_TITLE/
   DESCRIPTION, CHAT_INTERFACE$ERROR_SANDBOX_TITLE/DESCRIPTION, and
   BACKEND$NAME_REQUIRED/HOST_REQUIRED/HOST_INVALID were added in earlier
   commits on this branch but only had English values. Added English
   fallbacks for all 14 other supported locales in translation.json and
   regenerated public/locales/ via make-i18n.

Co-authored-by: openhands <openhands@all-hands.dev>

* i18n: add proper translations for 7 new keys across 14 locales

The previous commit used English as a fallback for all non-English
locales. Replace with proper translations for:

  CHAT_INTERFACE$ARCHIVED_SANDBOX_TITLE
  CHAT_INTERFACE$ARCHIVED_SANDBOX_DESCRIPTION
  CHAT_INTERFACE$ERROR_SANDBOX_TITLE
  CHAT_INTERFACE$ERROR_SANDBOX_DESCRIPTION
  BACKEND$NAME_REQUIRED
  BACKEND$HOST_REQUIRED
  BACKEND$HOST_INVALID

Locales covered: ja, zh-CN, zh-TW, ko-KR, no, ar, de, fr, it, pt,
es, ca, tr, uk.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: also write sandbox_status PAUSED to cache on stop-conversation

Bug: clicking 'Stop conversation' called pauseConversation() and then
only wrote execution_status: PAUSED to the React Query cache via
updateConversationExecutionStatusInCache.  sandbox_status was never
touched, so it remained as whatever the server last returned (null or
'RUNNING').

When the user reopened that conversation:
  • WebSocketProviderWrapper checked sandbox_status === 'PAUSED' → false
    → URL passed through → WebSocket fired at the paused sandbox → failed
  • useActiveConversation saw sandbox_status !== 'PAUSED' AND url !== null
    → 30-second poll interval → 30s before discovering the true state

Fix:
  1. Add patchConversationInCache() to conversation-mutation-utils —
     a generic helper that patches any subset of AppConversation fields
     in both the single-item and paginated-list query caches.
     updateConversationExecutionStatusInCache becomes a thin wrapper.
  2. use-unified-stop-conversation.ts uses patchConversationInCache to
     write BOTH execution_status: PAUSED and sandbox_status: 'PAUSED'
     atomically in onSuccess, so the gate in WebSocketProviderWrapper
     fires immediately on the next render.

Tests: __tests__/hooks/mutation/conversation-mutation-utils.test.ts
  • patchConversationInCache patches single-item cache
  • patchConversationInCache patches paginated list cache
  • patchConversationInCache patches multiple fields atomically
  • patchConversationInCache does not modify unrelated conversations
  • patchConversationInCache is a no-op on empty cache
  • updateConversationExecutionStatusInCache wrapper only touches
    execution_status (sandbox_status is left unchanged)

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(ui): archived/error banner — 'above' copy and readable text colors

Two issues with the archived/error sandbox banner that replaces the
chat input:

1. Copy said 'The history below is read-only' but the banner is
   anchored to the bottom of the chat, so the history is above it.
   Changed to 'above' in all 15 locales (en + ja/zh-CN/zh-TW/ko-KR/
   no/ar/de/fr/it/pt/es/ca/tr/uk).

2. Description text used text-[var(--oh-color-tertiary)] which maps to
   cool-grey-800 — nearly indistinguishable from the cool-grey-925
   surface background. Switched to the palette tokens that the rest of
   the UI uses for readable text on dark surfaces:
   • Title:       --oh-foreground  (cool-grey-100, bold label)
   • Description: --oh-muted       (cool-grey-400, secondary body text)

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): render echo-hello-world trajectory in archived/error views; fix lint

Three changes in one commit:

1. Prettier lint fix (conversation-mutation-utils.ts line 116):
   The one-line arrow body for updateConversationExecutionStatusInCache
   exceeded Prettier's column limit when written inline; split onto its
   own line. This fixes the 'test-and-build (ubuntu) Lint' CI failure.

2. MSW event fixture (src/mocks/conversation-handlers.ts):
   Add ECHO_HELLO_WORLD_TRAJECTORY — three events in TIMESTAMP_DESC
   order (newest-first, as the hook requests) that represent a minimal
   'echo hello world' session:
     archived-evt-1  user MessageEvent    'echo hello world'
     archived-evt-2  agent ExecuteBashAction  echo hello world
     archived-evt-3  env   ExecuteBashObservation  'hello world'
   Wire CONVERSATION_EVENTS map so GET /api/conversations/4/events/search
   and /5/events/search return these events; other conversations still
   get []. useConversationHistory reverses the DESC list back to
   chronological order before storing events.

3. Snapshot test (archived-conversation.snapshot.spec.ts):
   Wait for chatInterface.getByText('echo hello world') to be visible
   before taking the screenshot so the trajectory is guaranteed to have
   rendered above the read-only archived/error banner.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): inject trajectory via Zustand store, not MSW cross-origin fetch

The Service Worker registered at localhost:3001 cannot intercept
cross-origin requests; RemoteEventsList calls GET on the configured
backend host (127.0.0.1:8000), so MSW silently drops the response and
useConversationHistory returns no events.

Fix: pull the injectEvents helper pattern from
collapsible-thinking.snapshot.spec.ts and call it after asserting the
archived/error banner is visible.  The fixture is declared once at the
top of the file alongside a clear comment explaining why it mirrors the
MSW handler rather than importing from it.

Also removes the 10 s timeout from the post-inject getByText check
since injectEvents already polls until the store is populated and then
waits 500 ms for React to flush.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): atomic addEvents+DOM poll in injectEvents, no separate getByText

The previous impl had a two-step race window:
  1. expect.poll passed once store.events.length >= N
  2. 500 ms wait (or DOM waitForFunction) ran afterwards

React Strict-Mode's double clearEvents() invocation could fire between
steps 1 and 2, wiping the store before React flushed the render.

Fix: merge addEvents() and the data-testid="user-message" DOM check into
a single page.waitForFunction() poll.  Playwright polls ~100 ms so on
every tick we both re-seed the store AND verify the DOM element is
present.  addEvents() is idempotent (deduplicates by event ID) so
calling it on every tick is safe.  This eliminates the race entirely.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): wait for archived-banner as settled-state signal in tests 2/3

The previous approach waited for `chat-interface` (h-full flex div) to become
visible, but that container can be present in the DOM with zero computed height
before useActiveConversation resolves — causing intermittent 20 s timeout
failures in CI.

Following the same pattern as collapsible-thinking.snapshot.spec.ts (which
waits for `"Let's start building!"` as its settled-state signal), tests 2/3
now use a dedicated `navigateToArchivedConversation` helper that waits for
`archived-conversation-banner` to be visible (timeout 30 s).

The banner only renders after useActiveConversation returns data with
sandbox_status MISSING or ERROR, so it is a reliable indicator that:
  - the MSW mock responded to GET /api/conversations?ids=<id>
  - React Query received the data and set isFetched = true
  - ChatInterface evaluated isArchivedConversation = true
  - The banner div is both present and has non-zero dimensions

Also removes the now-redundant in-test banner visibility checks (the helper
already asserts them) and the stale `navigateToConversation` helper that was
no longer used.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): MSW ids[] parse bug + inject one event for stable archived-view test

Root cause of snapshot CI failures:

  Axios serializes { ids: ["4"] } as ?ids[]=4 (bracket notation).
  The MSW GET /api/conversations handler read searchParams.getAll("ids"),
  which returns [] when the key is "ids[]". listConversationResponses([])
  then falls back to returning ALL conversations, so results[0] was always
  conversation "1" (first in Map insertion order) regardless of which id
  was requested. Conversation "1" has no sandbox_status, so
  isArchivedConversation was always false and the archived banner never
  rendered — 30 s timeout.

Fix 1 — conversation-handlers.ts:
  Parse both bracket (ids[]) and plain (ids) formats so the mock correctly
  returns only the requested conversation(s).

Fix 2 — archived-conversation.snapshot.spec.ts:
  Rewrite tests 2/3 per user direction:
  • Use seedLocalStorage (same as collapsible-thinking) instead of bespoke
    addInitScript + page.route helpers.
  • Inject ONE minimal ExecuteBashAction event via __OH_EVENT_STORE__ so the
    chat has stable visible content that survives the 3 s polling re-renders
    (conv 4/5 have no conversation_url, so useActiveConversation polls every
    3 s). The injected event stays in the Zustand store across re-renders,
    giving toBeVisible a reliable anchor.
  • Wait for "echo hello" text (event), then wait for the archived banner —
    both are concrete settled-state signals, not the zero-height h-full div.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: trigger CI re-run

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: Update PR QA artifacts

* fix(tests): single injected event for archived-view + useOptionalConversationId mocks

- Remove ECHO_HELLO_WORLD_TRAJECTORY from MSW (was causing 3+1 = 4 events
  in the archived conversation snapshot view). CONVERSATION_EVENTS is now
  empty; the snapshot tests inject exactly one event via __OH_EVENT_STORE__.
- Add useOptionalConversationId to all vi.mock('#/hooks/use-conversation-id')
  calls that were missing it after the main merge refactored that hook.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): wait for banner before injecting events to avoid clearEvents race

The archived-conversation snapshot tests were injecting events via
__OH_EVENT_STORE__ immediately after the store became available on the
window object. However, the conversation route's useEffect (which calls
clearEvents()) fires asynchronously after the first paint — creating a
race where the injected events get wiped.

Fix: wait for the archived-conversation-banner to appear before
injecting events. The banner's presence proves that:
1. The route's clearEvents() effect has already fired
2. useActiveConversation has resolved with the correct sandbox_status
3. The chat interface is ready to accept and display events

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): pre-seed archived conversation events via MSW instead of runtime store injection

The archived-conversation snapshot tests were injecting events into the
Zustand event store at runtime via __OH_EVENT_STORE__. This raced with
the conversation route's useEffect (clearEvents) and React dev-mode
double-mount behavior, making the injected events disappear before the
chat could render them.

Fix: pre-seed CONVERSATION_EVENTS in the MSW mock handlers for
conversations 4 and 5 with one ExecuteBashAction event. The events now
load through the normal REST history path (useConversationHistory →
addEvents) — no runtime Zustand injection, no race condition.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): remove event injection — test banner + hidden input only

The archived-conversation snapshot tests kept crashing because event
injection (both via __OH_EVENT_STORE__ and pre-seeded MSW REST data)
always gets wiped by a React 18 strict mode effect-ordering issue:

In dev mode, strict mode double-fires effects child-before-parent.
ConversationWebSocketProvider (child) calls addEvents() first, then
conversation.tsx (parent) calls clearEvents() second, wiping all events.

Since the feature under test is the read-only banner and hidden chat
input (not event rendering), simplify the tests to verify only those
assertions — no event injection needed.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): add WebSocket stub to archived-conversation tests

The conversation-view snapshots showed a red 'Failed to connect to
server' toast because no agent-server runs at :8000 in CI — the Vite
proxy's ECONNREFUSED propagates to the browser and triggers the error
toast. Other conversation-page snapshot tests already stubbed
WebSocket; this test was missing it.

Extract the duplicated WebSocket stub into a shared helper at
tests/e2e/snapshots/support/stub-websocket.ts and use it in all three
conversation-page snapshot test files (archived-conversation,
collapsible-thinking, changes-tab).

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: update conversation card count from 5 to 6 after main merge

Main added pagination-local conversation fixture, bringing the total
mock conversations to 6. The archived-conversation sidebar test was
still asserting 5.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: update backends-extended snapshot tests for two-column add modal

The add-backend modal was refactored from a single form with radio
buttons (local/cloud kind selection) into a two-column layout:
- Left: manual connection (name, host, API key, Connect)
- Right: cloud OAuth login (device flow)

Kind is now inferred from the host URL, so the old radio button
testids (add-backend-kind-local, add-backend-kind-cloud) no longer
exist. Updated all affected flows:
- Flow 1: removed radio clicks, use URL inference for kind
- Flow 2: replaced radio inference tests with two-column layout test
- Flow 3: replaced OAuth button gating with cloud advanced settings
- Flow 7: removed radio click
- Flow 8: use add-backend-close instead of add-backend-cancel

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: update sidebar snapshot test conversation count from 5 to 6

Same pagination-local fixture issue as archived-conversation test.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: allhands-bot <allhands-bot@users.noreply.github.com>
2026-05-17 13:27:52 -04:00
Hiep Le 05d226a470 fix(frontend): cap backend-health retries and persist failure state (#411)
* fix: cap backend-health retries and persist failure state

* refactor: update the code based on feedback
2026-05-13 14:03:29 +07:00
Robert Brennanandopenhands 7509bbbe71 Replace bundled backend with seeded default local backend (#235)
* Replace bundled backend with seeded default local backend

Eliminate the special-case 'bundled' backend in favor of a regular registered
backend that is seeded on first read. Users can now rename or remove the
default 'Local' backend like any other.

- Drop `api/backend-registry/bundled.ts` and the `BundledBackend` distinction.
  Add `default-backend.ts` with `makeDefaultLocalBackend()` and
  `getEffectiveLocalBackend()` (used by API clients that need a baseline
  local target when the registry is empty).
- `storage.readStoredBackends()` seeds `[makeDefaultLocalBackend()]` when the
  `openhands-backends` localStorage key has never been written.
- `active-store.getActiveBackend()` falls back to the first local backend in
  the registry (or a synthesized default) instead of the removed bundled.
- `useActiveBackendContext()` no longer exposes `bundledBackend`. The Manage
  Backends modal and BackendSelector dropdown both render directly from the
  registered list — no special-casing — so the seeded default appears
  exactly once in each.
- BackendSelector: `buildOptions()` no longer prepends a synthetic bundled
  entry; the placeholder shows the active backend's name. Drop the
  BACKEND$LOCAL_ROW translation key.
- Dropdown UX: clear the input value on open so the active label does not
  appear both in the trigger and as a highlighted menu row.

Co-authored-by: openhands <openhands@all-hands.dev>

* Fix CI failures for seeded local backend

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-09 12:43:14 -07:00
Hiep Le dd744b13f6 feat: cloud backend support with multi-backend selector and SaaS proxy routing (#145)
* feat: multi-backend support with cloud SaaS proxy routing

* feat: route conversation export through cloud proxy on cloud backends

* fix: route conversation delete through cloud proxy on cloud backends

* fix: forward settings diffs verbatim through cloud proxy save

* fix: surface cloud-aware settings sub-pages and gate local-only routes

* fix: route secrets settings through cloud proxy on cloud backends

* fix: route conversation stop runtime through cloud proxy on cloud backends

* fix: re-expose planning agent UI for cloud backends and route plan file reads through cloud proxy

* fix: route Display Cost runtime fetch through cloud proxy and ungate local metrics without session API key

* fix: handle WAITING_FOR_SANDBOX task status from cloud backends to prevent UI crash

* fix: re-expose Public Share in conversation menu for cloud backends

* fix: redirect to home when switching backends from a conversation page

* fix: hide cloud orgs the API key can't access in backend selector

* feat: support running multiple local agent-servers with shared persistence

* fix: lint

* fix: failing tests
2026-05-08 01:17:40 +07:00