Commit Graph
7418 Commits
Author SHA1 Message Date
Tim O'Farrellandopenhands 048e390419 feat(automations): add logs modal to activity log items (#600)
* feat(automations): add logs modal to activity log items

Each AutomationRun now surfaces its bash command output via a small
terminal-icon button placed to the left of the run status badge in the
activity log. Clicking the icon opens a modal that fetches the
BashCommand event and all paginated BashOutput events for the run.

- Add bash_command_id to AutomationRun type (and mock data).
- New BashService.getCommandLogs(): cloud-aware reader for bash events
  that routes through callCloudProxy on cloud backends (runtime URL +
  session-api-key auth) and through BashClient directly on local
  backends. Pages through BashOutput events sorted by timestamp.
- New useBashCommandLogs hook: hydrates the run's conversation to
  resolve runtime URL + session API key, then drives the BashService
  query. Surfaces resolution states (loading, conversation missing,
  sandbox gone) so the modal can render meaningful empty states.
- New RunLogsModal: renders interleaved stdout/stderr from the command
  with stderr highlighted, exit code, and explicit messaging when the
  command is missing or the sandbox is no longer alive.
- Activity-log-item gets a logs button that stopPropagation +
  preventDefault on the wrapping conversation link.
- i18n: 8 new AUTOMATIONS$DETAIL$LOGS_* keys across all 15 languages.
- Tests: BashService cloud/local routing + ActivityLogItem button
  behaviour.

Co-authored-by: openhands <openhands@all-hands.dev>

* feat(automations): split logs modal into Output/Error tabs

Address review feedback on the run logs modal:

* Rename the modal title from "Run logs" to "Logs".
* Make the modal actually fire the bash-events search request:
  - In local mode, the agent-server hosts events under a single root,
    so the search query no longer waits for the per-conversation URL
    to resolve before firing. The conversation lookup still runs (so
    session-api-key and per-conversation URL are honoured when
    present), but it no longer gates the fetch.
  - In cloud mode the behaviour is unchanged — runtime endpoints
    require the conversation_url for the cloud-proxy hostOverride.
* Replace the interleaved stdout/stderr pre-block with two tabs:
  Output (stdout, default) and Error (stderr). The body of each tab
  is the chronological concatenation (by timestamp + order) of the
  matching field across every BashOutput event for the command.
* Simplify BashService — drop the extra BashCommand fetch; only the
  BashOutput search (`kind__eq=BashOutput, command_id__eq=<id>`) is
  needed for the rendered view. Note that the agent-server API uses
  `command_id__eq` (not `bash_command_id__eq`) — see the python
  bash_router for the canonical filter name.

i18n: add LOGS_TAB_OUTPUT, LOGS_TAB_ERROR, LOGS_EMPTY across all 15
locales; retranslate LOGS_TITLE to "Logs".

Tests:
* New run-logs-modal.test.tsx covers tab defaults, stdout/stderr
  concatenation (with reverse-order inputs to verify the sort),
  loading state, and Escape-to-close.
* bash-service.test.ts rewritten for the new listOutputs API and a
  cloud-without-conversation-url error path.

Co-authored-by: openhands <openhands@all-hands.dev>

* feat(automations): tolerate paused/missing/unreachable sandboxes in run logs modal

Cloud sandboxes can be in non-RUNNING states (paused, starting,
deleted, errored) and even RUNNING sandboxes can transiently fail at
the network layer. Previously the modal would either show a stuck
'Loading logs...' spinner or dump a raw axios error string. Now each
known-bad state is mapped to a stable `SandboxIssue` code with its
own localized empty-state message.

Behaviour:

* Pre-flight: when `sandbox_status` is MISSING, PAUSED, STARTING, or
  ERROR — or the conversation has no runtime URL at all — the bash
  query is **disabled** (no doomed request is fired). The modal
  renders the matching message instead of a spinner.
* Post-flight: when the request does fire and fails with a 404 or
  5xx response, or a network-level error (no response), the failure
  is classified as `unreachable` and the modal renders the
  'sandbox unreachable' message instead of the raw error. 401/403
  are *not* collapsed — those are auth bugs we want to surface.
* Local backends are unchanged: no sandbox lifecycle, so
  `sandboxIssue` is always null and errors flow through as-is.

API changes:

* `useBashCommandLogs` exposes a `sandboxIssue` discriminated union
  ("missing" | "paused" | "starting" | "errored" | "unreachable")
  instead of the old `hasNoRuntime` boolean. When an issue is set
  the hook clears `error` so the modal doesn't render both an empty
  state AND a raw axios string for the same failure.
* The modal switches over the issue codes via a centralized
  `SANDBOX_ISSUE_I18N` map.

i18n: replace the single `LOGS_SANDBOX_GONE` key with five specific
keys (one per sandbox issue) across all 15 locales.

Tests:

* New `__tests__/hooks/query/use-bash-command-logs.test.tsx`
  exhaustively covers each sandbox_status, the no-runtime-URL case,
  conversationMissing, the happy path, 404/5xx → unreachable
  classification, the network-error case, and the explicit
  401/403-passes-through case.
* `run-logs-modal.test.tsx` extended with a parameterized case
  covering all five issue codes; verifies the empty-state message is
  rendered and the tab body is suppressed.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-18 18:22:26 +00:00
Vasco Schiavo c8bf569d90 feat/fix(profiles): per-conversation /switch_llm in chat (#575)
* feat(profiles): per-conversation /switch_llm in chat, global activate on the home page

* chore(profiles): address feedbacks

* chore(profiles): fix lints
2026-05-18 14:11:11 -04:00
Robert Brennan 0444b7df1a Better default install instructions (#585)
* Update README.md

* Update README.md

* Revise OpenHands description in README

Updated project description to specify coding agents.

* Update README.md

* Update project description for clarity
2026-05-18 11:43:53 -04:00
Jamie Chicagoandopenhands 58d5dd5597 docs: add Windows install workaround to README (#581)
* docs: add Windows PowerShell workaround for dev:docker

Co-authored-by: openhands <openhands@all-hands.dev>

* docs: clarify OH_MOUNT_HOST_HOME hint

Co-authored-by: openhands <openhands@all-hands.dev>

* Apply suggestion from @jamiechicago312

* Update README.md

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-18 16:55:20 +02:00
Hiep Le 3e28cb1bbe refactor: rename top-level nav labels for clearer copy (#578) 2026-05-18 21:43:05 +07:00
Robert Brennan f189713e9d Update README.md: elevate dockerless installation and remove extraneous installation instructions (#576)
* Update README.md

* Remove npm package section from README

Removed npm package installation and usage instructions from README.
2026-05-18 10:21:45 -04:00
Hiep Le 42a7a2967a fix: sort full visible list when ordering by Created (#573) 2026-05-18 18:50:21 +07:00
5ea55eaa9e feat(sidebar): conversation list filters, grouping, and loading UX (#530)
* feat(conversation-panel): filters, grouping, and list preferences

Add filter menu for organize/sort/thread scope, metadata toggles, and older
conversations with persisted preferences; optional LLM model labels on cards.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(conversation-panel): sidebar list UX, grouping chrome, and scroll divider

Align grouped rows and cards with main nav inset, add per-workspace/repo thread
launcher and folder-plus picker with shared menu styling, expose group launch
metadata for tests, and tighten card/scroll header borders to match the sidebar.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-card): restore ellipsis trigger and stack context menu above list rows

Bring back the shared vertical EllipsisButton and raise z-index on the open card
and menu so the overflow panel paints above subsequent items in the scroll area.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(sidebar): full-bleed conversations divider and align list rows with nav

Drop list-wrapper overflow clipping so the scrolled header border can span
the aside, keep a stable transparent/colored border, and inset the title row
with pl-4/pr-2. Remove extra card horizontal padding (link px-2 already
applies) and center status dots in an 18px column like SidebarNavLink icons.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-panel): match conversation row hover width to folder rows

Drop link px-2 so cards span the same width as grouped headers, and move
pl-2/pr-1 onto the card and skeleton to mirror folder strip padding.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-panel): portal conversation card menu and align list spacing

Open the ellipsis menu from click only, render it fixed in a body portal with
correct anchor measurement, and ignore outside-click closes on the trigger.
Match grouped-folder vertical rhythm to the main sidebar on md breakpoints.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-panel): polish list skeleton, filter inset, and timestamp nudge

Use one skeleton block that matches conversation card padding and corners; add
light right padding on the filter control; shift relative timestamps slightly
left for alignment.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-panel): restyle conversation list loading skeleton

Show three darker staggered pulse bars with compact list spacing; mount once
from the panel instead of five duplicates.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): unify skeleton styling and staggered pulse

Standardize .skeleton/.skeleton-round on neutral-700 with motion-safe pulse,
add .skeleton-stagger for the three-phase wave, and adopt it across loaders
(automations, chat, skills, tasks, secrets, settings, conversations).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-panel): polish filter menu and grouped list layout

Use clearer section dividers and Lucide sort icons in the conversations
filter menu; tighten the scroll stack and skip empty grouped nav rendering
when there are no workspace groups.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-panel): show older cutoff beside section title

Pair “Older conversations” with a dimmer “Over 1 hour” hint on the right in
the filter menu, with full i18n coverage for the new string.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-panel): drop non-sidebar skeleton stagger usage

Restore automations, chat, tasks, settings, skills, and secrets skeleton
markup to main so list skeleton styling stays scoped to the sidebar.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: group conversations by selected workspace, not per-conversation worktree dir

* fix: enable Delete all for every conversation, regardless of age

* refactor: update the code based on feedback

* refactor: update the code based on feedback

* refactor: update the code based on feedback

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: hieptl <hieptl.developer@gmail.com>
2026-05-18 16:31:54 +07:00
Hiep Le 8e90704a51 chore: remove redundant styling assertions and assert semantic attrs instead (#566) 2026-05-18 12:02:42 +07:00
Hiep Leandallhands-bot 3300768da8 chore: remove unused assets, env var, and obsolete script (#564)
* chore: remove unused assets, env var, and obsolete script

* chore: Remove PR-only artifacts

---------

Co-authored-by: allhands-bot <allhands-bot@users.noreply.github.com>
2026-05-18 11:15:13 +07:00
Hiep Leandallhands-bot 376b3100fc chore: remove unused cloud auth API (#562)
* chore: remove unused cloud auth API

* chore: Remove PR-only artifacts

---------

Co-authored-by: allhands-bot <allhands-bot@users.noreply.github.com>
2026-05-18 11:06:58 +07:00
Hiep Leandallhands-bot 37d5c1dfa0 chore: remove unused utility functions (#560)
* chore: remove unused utility functions

* chore: Remove PR-only artifacts

---------

Co-authored-by: allhands-bot <allhands-bot@users.noreply.github.com>
2026-05-18 10:57:40 +07:00
c224d24a9b fix: cloud conversation resume + archived/error sandbox states (#500)
* fix: resume cloud conversations stuck in starting status

Two bugs prevented cloud-backend conversations from resuming properly:

1. useActiveConversation hard-coded a 30 s refetch interval. When a
   cloud sandbox is paused and auto-starts on access, conversation_url
   is null until the sandbox is ready. The WebSocket can't open without
   a URL, so curAgentState stays at LOADING ('starting status') for up
   to 30+ seconds — or forever if the user gave up before the next poll.
   Fix: use the query-state callback form of refetchInterval and drop
   to 3 s whenever conversation_url is null (mirrors the 3 s cadence of
   task polling), falling back to 30 s once the URL is available.

2. updateConversationExecutionStatusInCache called setQueryData with a
   3-element key ["user", "conversation", id] that no longer matches
   the 5-element key stored by useUserConversation
   ["user", "conversation", id, backend.id, orgId] after the
   per-backend cache isolation was added. Optimistic status writes after
   manual pause/resume were silently dropped.
   Fix: switch to setQueriesData with { queryKey: [...] } prefix
   matching so the update hits whichever (backend, org) variant is live.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: auto-resume cloud sandbox when conversation_url is null

Faster polling (prev commit) was not enough. The cloud API returns
conversation_url=null when the sandbox is paused/stopped, and a GET
request alone does not wake it up — you have to POST a new start task
(with sandbox_id to reuse the existing sandbox) and wait for it to
become READY.

Add a useEffect in AppContent that fires once per unique conversation.id
after the initial fetch:
  • skips if not a cloud backend
  • skips if conversation_url is already set (sandbox running)
  • skips if sandbox_id is null (nothing to resume)
  • guards against re-triggering within the same route-mount via a ref

On trigger it calls createConversation(sandbox_id), which POSTs
POST /api/v1/app-conversations with the sandbox_id to the cloud, gets
back a WORKING start task, then navigates to /conversations/task-{id}.
useTaskPolling drives the task to READY and redirects to the real
conversation, now with a conversation_url the WebSocket can connect to.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: navigate back to conversation when resume task fails

When the cloud sandbox fails to start (e.g. 'Sandbox failed to start
within 120s'), the task reaches ERROR status. Previously the user was
left stranded at the task-{id} URL with only a toast to show for it.

Two changes:
1. Pass resumedFromConversationId in React Router navigation state when
   navigating to task-{id} for a cloud resume, so we know where to go
   back if the task fails.
2. In the task-error effect, read that state and navigate back to the
   original conversation (or /conversations if no originator is known).
   The resume effect's ref is still set so it will not re-trigger the
   resume on landing, preventing a retry loop.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use correct sandbox resume endpoint matching OpenHands

Root cause of 'Sandbox failed to start within 120s':
The previous fix called POST /api/v1/app-conversations with sandbox_id,
which is the 'create a new conversation' endpoint. The cloud treats this
as a full sandbox provisioning request with a 120-second cold-start
timeout that can fail on old/stale sandboxes.

The correct endpoint — matching OpenHands' SandboxService.resumeSandbox
and useSandboxRecovery — is POST /api/v1/sandboxes/{id}/resume, which is
a lightweight unpause that simply wakes the existing sandbox without
reprovisioning it.

Three changes:
1. Add SandboxStatus type ('PAUSED'|'RUNNING'|'STARTING'|'MISSING') and
   sandbox_status field to AppConversation, mirroring OpenHands'
   V1SandboxStatus. The cloud API already returns this field; adding the
   type makes it accessible in TypeScript.

2. Add resumeCloudSandbox(sandboxId) to the cloud service, calling
   POST /api/v1/sandboxes/{id}/resume via the cloud proxy — symmetric
   with the existing pauseCloudSandbox.

3. Update the resume effect in conversation.tsx:
   - Detect on sandbox_status === 'PAUSED' (more precise than
     conversation_url === null, which can be null for other reasons).
   - Call resumeCloudSandbox(sandbox_id) instead of createConversation.
   - Stay on the current URL after resume — no task navigation needed.
     The 3-second refetch interval in useActiveConversation polls until
     conversation_url populates, then the WebSocket connects normally.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: add ERROR to SandboxStatus to match OpenHands V1SandboxStatus

Complete enum is MISSING|STARTING|RUNNING|PAUSED|ERROR.

Co-authored-by: openhands <openhands@all-hands.dev>

* feat: archived/error sandbox state — read-only view and sidebar indicators

When a cloud conversation's sandbox_status is MISSING or ERROR it can
never be resumed. These two states now have first-class treatment:

Sidebar / conversation list:
- ConversationStatusDot gains an optional sandboxStatus prop.
  MISSING → gray 'paused' dot with tooltip 'Archived'.
  ERROR   → red 'error' dot with tooltip 'Error'.
  (ExecutionStatus visual is used unchanged for all other states.)
- ConversationCardHeader passes sandboxStatus to the dot and sets
  isConversationArchived on the title, which applies opacity-60.
- ConversationCard renders the existing ConversationStatusBadges pill
  ('Archived' or 'Error' pill badge) for MISSING and ERROR sandboxes.
- CompactConversationRow (collapsed sidebar) passes sandboxStatus to
  both the main dot and the tooltip-preview dot.
- conversation-panel.tsx passes sandbox_status from AppConversation to
  both card variants.

Conversation view (read-only):
- ChatInterface reads sandbox_status via useActiveConversation.
- When MISSING or ERROR, the InteractiveChatBox is replaced by a
  localised banner (title + description) explaining the history is
  read-only. The banner uses data-testid='archived-conversation-banner'
  for testing.
- The auto-resume effect in conversation.tsx already skips MISSING and
  ERROR because it only fires on sandbox_status === 'PAUSED'.

i18n:
  CHAT_INTERFACE$ARCHIVED_SANDBOX_TITLE/DESCRIPTION
  CHAT_INTERFACE$ERROR_SANDBOX_TITLE/DESCRIPTION

Co-authored-by: openhands <openhands@all-hands.dev>

* test: snapshot tests for archived/error sandbox conversation states

Three new Playwright visual snapshots in
tests/e2e/snapshots/archived-conversation.snapshot.spec.ts:

1. conversation-panel-with-archived-badges
   Navigates to /conversations; verifies five conversation cards are
   present; asserts archived-badge and error-badge are both visible;
   captures the conversation panel showing:
   - 'Archived Project'  → gray dot + 'Archived' pill + dimmed title
   - 'Errored Project'   → red dot + 'Error' pill + dimmed title

2. conversation-view-archived
   Navigates to /conversations/4 (sandbox_status: 'MISSING'); stubs
   WebSocket; asserts:
   - archived-conversation-banner is visible (read-only notice)
   - interactive-chat-box is absent (count 0)
   Captures the full chat interface.

3. conversation-view-sandbox-error
   Same as above for /conversations/5 (sandbox_status: 'ERROR');
   captures the 'Sandbox error' banner variant.

Supporting changes:
- src/api/agent-server-adapter.ts
  - Add sandbox_status?: string | null to DirectConversationInfo
  - Import SandboxStatus and map info.sandbox_status → AppConversation
    so the field is no longer silently null for all conversations
- src/mocks/conversation-handlers.ts
  - Add mock conversations 4 (MISSING) and 5 (ERROR) with sandbox_status
  - createConversationResponse now includes sandbox_status in the payload
- src/components/features/chat/chat-interface.tsx
  - Suppress ChatSuggestions ("Let's start building!") for archived
    conversations — showing task suggestions alongside a read-only
    banner is confusing and misleading
- src/components/features/conversation-panel/conversation-card/
  conversation-status-badges.tsx
  - Add data-testid="archived-badge" and data-testid="error-badge"
    so Playwright can assert on their presence without relying on text
- tests/e2e/snapshots/sidebar.snapshot.spec.ts
  - Update toHaveCount(3) → toHaveCount(5) to account for the two new
    mock conversations; existing sidebar snapshot baseline needs
    regeneration on main (intentional diff via update-snapshots label)

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: make sandbox_status optional in AppConversation

sandbox_status is a cloud-only field that local agent-server conversations
never carry. Existing test fixtures built AppConversation objects without
this field, causing TypeScript to error once it became required.

Making it optional (sandbox_status?: SandboxStatus | null) is the
semantically correct choice:
- The field is absent / null for every local conversation
- The adapter still explicitly maps it to null when unset
- ChatInterface reads it with ?? null so undefined is handled safely
- Partial<AppConversation> spreads in test factory functions no longer
  widen to SandboxStatus | null | undefined

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: Prettier formatting — multiline SandboxStatus union and ternary

- SandboxStatus type: expand single-line union to multi-line format
- ConversationCard: wrap sandboxStatus ternary in a multiline JSX block

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: preserve sandbox_status through requireDirectConversationInfo

The validation function that normalises raw API responses into
DirectConversationInfo was not copying sandbox_status, so it was silently
dropped every time a conversation came through the search or batch-get
code paths. This caused the conversation panel to never render the
archived/error badge pills, and the ChatInterface to always treat every
conversation as active (missing read-only banner for MISSING/ERROR sandboxes).

Fix: add sandbox_status: stringOrNull(item.sandbox_status) to the
mapping in requireDirectConversationInfo, mirroring the treatment of
execution_status.

The existing E2E tests for conversations 4 (MISSING) and 5 (ERROR)
were already asserting on archived-badge / error-badge presence and
archived-conversation-banner visibility, so they will now pass.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: don't connect WebSocket while cloud sandbox is PAUSED

When a cloud conversation is closed from the UI (pauseCloudSandbox is
called), the conversation's conversation_url is NOT cleared — it still
points to the old sandbox host. On the next navigation into that
conversation the WebSocket provider saw a non-null URL and immediately
tried to open a connection, which failed because the sandbox had not
yet woken up.

Two-part fix:

1. WebSocketProviderWrapper: suppress conversationUrl (treat it as null)
   while sandbox_status === 'PAUSED', so ConversationWebSocketProvider
   cannot compute a valid wsUrl until the sandbox is actually running.

2. useActiveConversation: add sandbox_status === 'PAUSED' as a
   fast-poll trigger alongside !conversation_url. The old check only
   fast-polled when the URL was absent; for paused sandboxes the URL is
   present but stale, so without this the hook would stay on the slow
   30-second interval while waiting for the sandbox to wake up.

Together these changes let the resume sequence complete correctly:
  navigate → sandbox PAUSED detected → resumeCloudSandbox called →
  fast-poll picks up RUNNING state → conversationUrl unblocked →
  WebSocket connects with a live sandbox host.

Co-authored-by: openhands <openhands@all-hands.dev>

* docs: document cloud PAUSED sandbox WebSocket gating in AGENTS.md

Co-authored-by: openhands <openhands@all-hands.dev>

* test: cover PAUSED sandbox gating and sandbox_status preservation

Three test suites covering the cloud conversation resume bug fixes:

1. agent-server-conversation-service.test.ts — three cases asserting
   that requireDirectConversationInfo preserves sandbox_status through
   batchGetAppConversations (PAUSED, RUNNING, absent → null).

2. websocket-provider-wrapper.test.tsx — five cases asserting that
   WebSocketProviderWrapper passes conversationUrl through when the
   sandbox is RUNNING or null (local backend), suppresses it to null
   when sandbox_status === 'PAUSED', and handles not-yet-fetched data.

3. use-active-conversation.test.ts — five cases asserting that the
   refetchInterval callback returns 3000 when sandbox_status is PAUSED
   (even with a non-null conversation_url) or when conversation_url is
   null, and 30000 in all other ready states.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(test): use null for execution_status fixture field

ExecutionStatus is a string enum — assigning the raw string literal
'idle' triggers TS2322. Null satisfies ExecutionStatus | null and is
irrelevant to what these tests actually exercise.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(lint): prettier format + add missing i18n fallbacks for 7 keys

Two issues from lint-staged pre-commit hook:

1. Prettier: the compound refetchInterval condition in use-active-
   conversation.ts was too long for one line — broke across three lines.

2. Translation completeness: CHAT_INTERFACE$ARCHIVED_SANDBOX_TITLE/
   DESCRIPTION, CHAT_INTERFACE$ERROR_SANDBOX_TITLE/DESCRIPTION, and
   BACKEND$NAME_REQUIRED/HOST_REQUIRED/HOST_INVALID were added in earlier
   commits on this branch but only had English values. Added English
   fallbacks for all 14 other supported locales in translation.json and
   regenerated public/locales/ via make-i18n.

Co-authored-by: openhands <openhands@all-hands.dev>

* i18n: add proper translations for 7 new keys across 14 locales

The previous commit used English as a fallback for all non-English
locales. Replace with proper translations for:

  CHAT_INTERFACE$ARCHIVED_SANDBOX_TITLE
  CHAT_INTERFACE$ARCHIVED_SANDBOX_DESCRIPTION
  CHAT_INTERFACE$ERROR_SANDBOX_TITLE
  CHAT_INTERFACE$ERROR_SANDBOX_DESCRIPTION
  BACKEND$NAME_REQUIRED
  BACKEND$HOST_REQUIRED
  BACKEND$HOST_INVALID

Locales covered: ja, zh-CN, zh-TW, ko-KR, no, ar, de, fr, it, pt,
es, ca, tr, uk.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: also write sandbox_status PAUSED to cache on stop-conversation

Bug: clicking 'Stop conversation' called pauseConversation() and then
only wrote execution_status: PAUSED to the React Query cache via
updateConversationExecutionStatusInCache.  sandbox_status was never
touched, so it remained as whatever the server last returned (null or
'RUNNING').

When the user reopened that conversation:
  • WebSocketProviderWrapper checked sandbox_status === 'PAUSED' → false
    → URL passed through → WebSocket fired at the paused sandbox → failed
  • useActiveConversation saw sandbox_status !== 'PAUSED' AND url !== null
    → 30-second poll interval → 30s before discovering the true state

Fix:
  1. Add patchConversationInCache() to conversation-mutation-utils —
     a generic helper that patches any subset of AppConversation fields
     in both the single-item and paginated-list query caches.
     updateConversationExecutionStatusInCache becomes a thin wrapper.
  2. use-unified-stop-conversation.ts uses patchConversationInCache to
     write BOTH execution_status: PAUSED and sandbox_status: 'PAUSED'
     atomically in onSuccess, so the gate in WebSocketProviderWrapper
     fires immediately on the next render.

Tests: __tests__/hooks/mutation/conversation-mutation-utils.test.ts
  • patchConversationInCache patches single-item cache
  • patchConversationInCache patches paginated list cache
  • patchConversationInCache patches multiple fields atomically
  • patchConversationInCache does not modify unrelated conversations
  • patchConversationInCache is a no-op on empty cache
  • updateConversationExecutionStatusInCache wrapper only touches
    execution_status (sandbox_status is left unchanged)

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(ui): archived/error banner — 'above' copy and readable text colors

Two issues with the archived/error sandbox banner that replaces the
chat input:

1. Copy said 'The history below is read-only' but the banner is
   anchored to the bottom of the chat, so the history is above it.
   Changed to 'above' in all 15 locales (en + ja/zh-CN/zh-TW/ko-KR/
   no/ar/de/fr/it/pt/es/ca/tr/uk).

2. Description text used text-[var(--oh-color-tertiary)] which maps to
   cool-grey-800 — nearly indistinguishable from the cool-grey-925
   surface background. Switched to the palette tokens that the rest of
   the UI uses for readable text on dark surfaces:
   • Title:       --oh-foreground  (cool-grey-100, bold label)
   • Description: --oh-muted       (cool-grey-400, secondary body text)

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): render echo-hello-world trajectory in archived/error views; fix lint

Three changes in one commit:

1. Prettier lint fix (conversation-mutation-utils.ts line 116):
   The one-line arrow body for updateConversationExecutionStatusInCache
   exceeded Prettier's column limit when written inline; split onto its
   own line. This fixes the 'test-and-build (ubuntu) Lint' CI failure.

2. MSW event fixture (src/mocks/conversation-handlers.ts):
   Add ECHO_HELLO_WORLD_TRAJECTORY — three events in TIMESTAMP_DESC
   order (newest-first, as the hook requests) that represent a minimal
   'echo hello world' session:
     archived-evt-1  user MessageEvent    'echo hello world'
     archived-evt-2  agent ExecuteBashAction  echo hello world
     archived-evt-3  env   ExecuteBashObservation  'hello world'
   Wire CONVERSATION_EVENTS map so GET /api/conversations/4/events/search
   and /5/events/search return these events; other conversations still
   get []. useConversationHistory reverses the DESC list back to
   chronological order before storing events.

3. Snapshot test (archived-conversation.snapshot.spec.ts):
   Wait for chatInterface.getByText('echo hello world') to be visible
   before taking the screenshot so the trajectory is guaranteed to have
   rendered above the read-only archived/error banner.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): inject trajectory via Zustand store, not MSW cross-origin fetch

The Service Worker registered at localhost:3001 cannot intercept
cross-origin requests; RemoteEventsList calls GET on the configured
backend host (127.0.0.1:8000), so MSW silently drops the response and
useConversationHistory returns no events.

Fix: pull the injectEvents helper pattern from
collapsible-thinking.snapshot.spec.ts and call it after asserting the
archived/error banner is visible.  The fixture is declared once at the
top of the file alongside a clear comment explaining why it mirrors the
MSW handler rather than importing from it.

Also removes the 10 s timeout from the post-inject getByText check
since injectEvents already polls until the store is populated and then
waits 500 ms for React to flush.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): atomic addEvents+DOM poll in injectEvents, no separate getByText

The previous impl had a two-step race window:
  1. expect.poll passed once store.events.length >= N
  2. 500 ms wait (or DOM waitForFunction) ran afterwards

React Strict-Mode's double clearEvents() invocation could fire between
steps 1 and 2, wiping the store before React flushed the render.

Fix: merge addEvents() and the data-testid="user-message" DOM check into
a single page.waitForFunction() poll.  Playwright polls ~100 ms so on
every tick we both re-seed the store AND verify the DOM element is
present.  addEvents() is idempotent (deduplicates by event ID) so
calling it on every tick is safe.  This eliminates the race entirely.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): wait for archived-banner as settled-state signal in tests 2/3

The previous approach waited for `chat-interface` (h-full flex div) to become
visible, but that container can be present in the DOM with zero computed height
before useActiveConversation resolves — causing intermittent 20 s timeout
failures in CI.

Following the same pattern as collapsible-thinking.snapshot.spec.ts (which
waits for `"Let's start building!"` as its settled-state signal), tests 2/3
now use a dedicated `navigateToArchivedConversation` helper that waits for
`archived-conversation-banner` to be visible (timeout 30 s).

The banner only renders after useActiveConversation returns data with
sandbox_status MISSING or ERROR, so it is a reliable indicator that:
  - the MSW mock responded to GET /api/conversations?ids=<id>
  - React Query received the data and set isFetched = true
  - ChatInterface evaluated isArchivedConversation = true
  - The banner div is both present and has non-zero dimensions

Also removes the now-redundant in-test banner visibility checks (the helper
already asserts them) and the stale `navigateToConversation` helper that was
no longer used.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): MSW ids[] parse bug + inject one event for stable archived-view test

Root cause of snapshot CI failures:

  Axios serializes { ids: ["4"] } as ?ids[]=4 (bracket notation).
  The MSW GET /api/conversations handler read searchParams.getAll("ids"),
  which returns [] when the key is "ids[]". listConversationResponses([])
  then falls back to returning ALL conversations, so results[0] was always
  conversation "1" (first in Map insertion order) regardless of which id
  was requested. Conversation "1" has no sandbox_status, so
  isArchivedConversation was always false and the archived banner never
  rendered — 30 s timeout.

Fix 1 — conversation-handlers.ts:
  Parse both bracket (ids[]) and plain (ids) formats so the mock correctly
  returns only the requested conversation(s).

Fix 2 — archived-conversation.snapshot.spec.ts:
  Rewrite tests 2/3 per user direction:
  • Use seedLocalStorage (same as collapsible-thinking) instead of bespoke
    addInitScript + page.route helpers.
  • Inject ONE minimal ExecuteBashAction event via __OH_EVENT_STORE__ so the
    chat has stable visible content that survives the 3 s polling re-renders
    (conv 4/5 have no conversation_url, so useActiveConversation polls every
    3 s). The injected event stays in the Zustand store across re-renders,
    giving toBeVisible a reliable anchor.
  • Wait for "echo hello" text (event), then wait for the archived banner —
    both are concrete settled-state signals, not the zero-height h-full div.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: trigger CI re-run

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: Update PR QA artifacts

* fix(tests): single injected event for archived-view + useOptionalConversationId mocks

- Remove ECHO_HELLO_WORLD_TRAJECTORY from MSW (was causing 3+1 = 4 events
  in the archived conversation snapshot view). CONVERSATION_EVENTS is now
  empty; the snapshot tests inject exactly one event via __OH_EVENT_STORE__.
- Add useOptionalConversationId to all vi.mock('#/hooks/use-conversation-id')
  calls that were missing it after the main merge refactored that hook.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): wait for banner before injecting events to avoid clearEvents race

The archived-conversation snapshot tests were injecting events via
__OH_EVENT_STORE__ immediately after the store became available on the
window object. However, the conversation route's useEffect (which calls
clearEvents()) fires asynchronously after the first paint — creating a
race where the injected events get wiped.

Fix: wait for the archived-conversation-banner to appear before
injecting events. The banner's presence proves that:
1. The route's clearEvents() effect has already fired
2. useActiveConversation has resolved with the correct sandbox_status
3. The chat interface is ready to accept and display events

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): pre-seed archived conversation events via MSW instead of runtime store injection

The archived-conversation snapshot tests were injecting events into the
Zustand event store at runtime via __OH_EVENT_STORE__. This raced with
the conversation route's useEffect (clearEvents) and React dev-mode
double-mount behavior, making the injected events disappear before the
chat could render them.

Fix: pre-seed CONVERSATION_EVENTS in the MSW mock handlers for
conversations 4 and 5 with one ExecuteBashAction event. The events now
load through the normal REST history path (useConversationHistory →
addEvents) — no runtime Zustand injection, no race condition.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): remove event injection — test banner + hidden input only

The archived-conversation snapshot tests kept crashing because event
injection (both via __OH_EVENT_STORE__ and pre-seeded MSW REST data)
always gets wiped by a React 18 strict mode effect-ordering issue:

In dev mode, strict mode double-fires effects child-before-parent.
ConversationWebSocketProvider (child) calls addEvents() first, then
conversation.tsx (parent) calls clearEvents() second, wiping all events.

Since the feature under test is the read-only banner and hidden chat
input (not event rendering), simplify the tests to verify only those
assertions — no event injection needed.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshots): add WebSocket stub to archived-conversation tests

The conversation-view snapshots showed a red 'Failed to connect to
server' toast because no agent-server runs at :8000 in CI — the Vite
proxy's ECONNREFUSED propagates to the browser and triggers the error
toast. Other conversation-page snapshot tests already stubbed
WebSocket; this test was missing it.

Extract the duplicated WebSocket stub into a shared helper at
tests/e2e/snapshots/support/stub-websocket.ts and use it in all three
conversation-page snapshot test files (archived-conversation,
collapsible-thinking, changes-tab).

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: update conversation card count from 5 to 6 after main merge

Main added pagination-local conversation fixture, bringing the total
mock conversations to 6. The archived-conversation sidebar test was
still asserting 5.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: update backends-extended snapshot tests for two-column add modal

The add-backend modal was refactored from a single form with radio
buttons (local/cloud kind selection) into a two-column layout:
- Left: manual connection (name, host, API key, Connect)
- Right: cloud OAuth login (device flow)

Kind is now inferred from the host URL, so the old radio button
testids (add-backend-kind-local, add-backend-kind-cloud) no longer
exist. Updated all affected flows:
- Flow 1: removed radio clicks, use URL inference for kind
- Flow 2: replaced radio inference tests with two-column layout test
- Flow 3: replaced OAuth button gating with cloud advanced settings
- Flow 7: removed radio click
- Flow 8: use add-backend-close instead of add-backend-cancel

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: update sidebar snapshot test conversation count from 5 to 6

Same pagination-local fixture issue as archived-conversation test.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: allhands-bot <allhands-bot@users.noreply.github.com>
2026-05-17 13:27:52 -04:00
Hiep Le e3c4cf9d9d chore: remove unused feature components (#557) 2026-05-17 23:56:52 +07:00
Hiep Le 1a1d5556b3 chore: remove unused hooks (#555) 2026-05-17 23:45:34 +07:00
Hiep Le e69217e67e chore: remove unused shared UI components (#552) 2026-05-17 23:19:38 +07:00
Hiep Le 5d3e5f7611 chore: remove SaaS references from codebase (#548) 2026-05-17 15:46:17 +07:00
Hiep Le 82516d4724 fix: suppress Chrome DevTools .well-known probe 404 logs (#546) 2026-05-17 14:56:45 +07:00
Hiep Le 13164d477f refactor: use cn utility for className composition in 21 components (#544) 2026-05-17 14:31:27 +07:00
Hiep Le 7c1cdb858f refactor(frontend): replace inline style props with Tailwind classes and design tokens (#542)
* refactor: replace inline style props with Tailwind classes and design tokens

* fix: failing tests
2026-05-17 14:01:32 +07:00
Hiep Le 48552f5777 fix: use SPA navigation for Start a conversation CTA (#540) 2026-05-17 13:53:27 +07:00
Hiep Le 397b7c06e3 fix: keep backend tray popover open on icon left-click (#538) 2026-05-17 12:19:34 +07:00
Tim O'Farrellandopenhands 0cf58f8657 fix(dev-docker): allow exec on the home tmpfs so stdio MCP servers work (#535)
Docker's `--tmpfs` flag defaults its mount option set to
`rw,noexec,nosuid,nodev`. When dev:docker overlays the agent-server
container's `/home/openhands` with a tmpfs (so the mapped host user can
own it), it was inheriting that `noexec` flag unintentionally.

That breaks any stdio MCP server installed via npx -- e.g.
`npx -y @modelcontextprotocol/server-github` caches its binary under
`~/.npm/_npx/<hash>/node_modules/.bin/mcp-server-github`, npx then tries
to exec it, and the kernel returns EACCES regardless of the 0755 mode
bits on the file. The user sees:

    sh: 1: mcp-server-github: Permission denied
    Failed to connect to MCP server 'github', skipping
    ...
    MCPError: MCP Connection Failure

and the conversation that triggered it aborts during agent init.

Pass `exec` explicitly to override only the `noexec` default. We keep
`nosuid` and `nodev` (the home dir has no business hosting setuid
binaries or device nodes) so we lose no defense-in-depth beyond what's
necessary to make the supported MCP integration actually function.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-16 21:15:42 -06:00
Engel Nystandopenhands 40032f97c8 fix: clarify local llm profiles ui (#534)
Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-17 01:07:22 +02:00
4f46598460 Clarify OpenHands backend login host (#423)
* Add backend modal visual snapshot

Co-authored-by: openhands <openhands@all-hands.dev>

* Clarify OpenHands backend login host

Co-authored-by: openhands <openhands@all-hands.dev>

* Clarify OpenHands Cloud login copy

Co-authored-by: openhands <openhands@all-hands.dev>

* Add missing i18n translations and move host into connection box

- Translate BACKEND$HOST_HELPER, BACKEND$HOST_DOCS_LINK, and
  BACKEND$AUTH_METHOD_HELPER into all 14 non-English locales.
- Move the Host input inside the bordered connection card for cloud-add
  mode so host, login, and API key are visually grouped.
- Keep the host input in a stable DOM position (never unmounted when
  kind flips mid-keystroke) by conditionally styling the wrapper
  rather than rendering two separate input trees.

Co-authored-by: openhands <openhands@all-hands.dev>

* Update add-backend-modal snapshot

Co-authored-by: openhands <openhands@all-hands.dev>

* Reorder: Host + API Key first, then OR Login with OpenHands Cloud

The connection box now shows:
1. Host input + helper text
2. API Key input + docs link
3. ── OR ──
4. Login with OpenHands Cloud button

This groups the manual credentials together and presents the
OAuth device flow as the alternative, which is a clearer mental
model.

Also removes the now-unused BACKEND$AUTH_METHOD_HELPER translation
key.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: update baseline snapshots [skip ci]

* Remove prefilled default host; Login with Cloud always uses app.openhands.dev

- Host field starts empty — no confusing 'use the default' guidance.
- 'Login with OpenHands Cloud' always targets app.openhands.dev
  regardless of what's in the host field, and fills both host and
  API key on success.
- Updated helper text to simply say 'Enter the URL of your OpenHands
  Cloud or self-hosted Agent Server.'
- Removed now-unused BACKEND$AUTH_METHOD_HELPER translation key.
- Updated tests to reflect no-prefill behavior.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: update baseline snapshots [skip ci]

* Remove redundant helper text; form layout is self-explanatory

The Login with OpenHands Cloud button already handles the cloud
case, so the helper text about entering a Cloud URL was confusing.
Removed HOST_HELPER, HOST_DOCS_LINK translation keys and the
OPENHANDS_CLOUD_DOCS_URL constant.

The form now reads cleanly:
  Host  +  API Key  →  OR  →  Login with OpenHands Cloud

Co-authored-by: openhands <openhands@all-hands.dev>

* Add host helper: 'Enter the URL of your agent server or self-hosted OpenHands Cloud'

Links 'self-hosted OpenHands Cloud' to
https://github.com/All-Hands-AI/OpenHands-Cloud

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: update baseline snapshots [skip ci]

* Login button uses host field; prefill cloud default; add OAuth hint

- Host prefilled with https://app.openhands.dev (self-hosted users
  can change it to their own deployment).
- Login button now uses the host from the field (not hardcoded),
  so self-hosted OpenHands Cloud deployments can also use OAuth.
- Added hint below login button: 'Works with OpenHands Cloud or
  your self-hosted OpenHands Cloud deployment.'

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: update baseline snapshots [skip ci]

* Two-column Add Backend modal with i18n translations

Redesign the Add Backend modal as a two-column layout:
- Left column: manual connection (Name, Host, API Key + Connect button)
- Right column: OpenHands Cloud OAuth login with Advanced host override
- Vertical OR divider separating the two approaches

Add 5 new i18n translation keys with all 15 language translations:
- BACKEND$CLOUD_TITLE, BACKEND$CLOUD_DESCRIPTION, BACKEND$CONNECT,
  BACKEND$ADVANCED, BACKEND$NAME_HELPER

Update tests for the new layout (add-backend-modal, backend-selector).

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: lint and prettier formatting in backend-form-modal

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: stabilize add-backend snapshot test with server_info mock and networkidle wait

The e2e test was timing out on the dropdown trigger click because the
backend health check was making real requests to the dev server. Mock
the /server_info endpoint and wait for networkidle before interacting.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: replace networkidle with targeted element wait in snapshot test

The networkidle wait was consuming most of the 60s test timeout due to
periodic health check polls. Wait for the specific dropdown trigger
element instead.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use hover instead of click to open backend dropdown in snapshot test

The BackendSelector uses openOnHover=true. Playwright's click first
hovers (opening the menu), then clicks the toggle (closing it).
Using hover() keeps the menu open for the subsequent menu item click.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: update baseline snapshots [skip ci]

* ci: trigger CI after snapshot baseline update

Co-authored-by: openhands <openhands@all-hands.dev>

* Render backend host helper text

Co-authored-by: openhands <openhands@all-hands.dev>

* Update self-hosted cloud repository link

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use current cloud login host

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-05-16 16:55:44 -04:00
Xingyao Wangandopenhands 2123a2761c feat: enable paginated event loading for cloud mode (#407)
* feat: enable paginated event loading for cloud mode

Enable scroll-up pagination for cloud backends by passing all search
params (sort_order, page_id, timestamp__gte, timestamp__lt) through to
the cloud proxy. The cloud useLoadOlderEvents hook no longer gates on
`isCloud`, so long conversations load the latest 50 events first and
lazily backfill older pages as the user scrolls up.

The loading indicator now shows 'Fetching older messages…' alongside
the spinner so users know what's happening during pagination.

Depends on: OpenHands/OpenHands#14399 (server-side timestamp fix)
Closes #402

Co-authored-by: openhands <openhands@all-hands.dev>

* address review: add fallback for unpatched cloud backends + tests

- Cloud event search now tries full params first, falls back to
  limit-only on error (graceful degradation for servers without
  OpenHands/OpenHands#14399).
- Added JSDoc note about server dependency to useLoadOlderEvents.
- Updated event-service tests: verify all params forwarded, fallback
  on 500, rethrow on limit-only failure, pagination stop on short page.
- Removed obsolete cloud-disabled test and useActiveBackend mock from
  use-load-older-events tests.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: return empty page on fallback to prevent infinite retries

When an unpatched cloud backend rejects timestamp filters, return an
empty page instead of retrying with limit-only params. The limit-only
fallback would return the same most-recent events already in the store,
which get deduped but leave hasMore=true — causing infinite requests.

An empty page makes useLoadOlderEvents set hasMore=false, cleanly
stopping pagination on unpatched backends.

Co-authored-by: openhands <openhands@all-hands.dev>

* docs: fix stale comment about fallback behavior

Co-authored-by: openhands <openhands@all-hands.dev>

* test: add event pagination e2e coverage

Add deterministic mock conversation fixtures for local and cloud event pagination, plus Playwright regression coverage that verifies initial tail loading and scroll-up older-event backfill for both backend modes.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: format event pagination params

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-16 19:53:31 +00:00
Engel Nyst e72b729ce0 Fix project phase to alpha in README (#528)
Updated project phase from sandbox to alpha and improved warning message formatting.
2026-05-16 20:13:02 +02:00
Robert Brennanandopenhands 35555bac11 feat: enable sub-agent delegation via task_tool_set (#509)
The agent-server already registers `task_tool_set` (TaskToolSet) at
startup via openhands-agent-server/openhands/agent_server/tool_router.py
and preloads the built-in sub-agents (code-explorer, bash-runner,
web-researcher, general-purpose) through `register_builtins_agents`.
However, agent-canvas never asked for the tool in the agent spec it
sends on POST /api/conversations, so the LLM had no way to delegate.

Add `task_tool_set` to the tools list assembled by `getAgentTools()`,
gated by the existing `isAgentServerToolAvailable` capability probe so
older agent-servers that don't advertise it in /api/server_info's
`usable_tools` are skipped cleanly.

TaskAction / TaskObservation events render through the existing default
event content path, so no new visualizer is required to ship this.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-16 18:00:03 +00:00
Hiep Le c165ade886 fix: align sidebar title with page title at top of layout (#527) 2026-05-17 00:38:31 +07:00
4713522932 refactor(onboarding): agent step branding and modal layout (#524)
Center Skip under the card, drop the step label from the header, and use white progress segments. Choose-agent options get monochrome logos, per-row Coming soon badges, and refined default/hover/selected card chrome.

Fill BACKEND validation strings across all supported locales so translation completeness checks pass.

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: hieptl <hieptl.developer@gmail.com>
2026-05-16 23:54:19 +07:00
Hiep Le b4d26635b0 fix: prevent LLM profile actions menu from being clipped by scroll container (#523) 2026-05-16 23:32:59 +07:00
Engel Nystandopenhands 5c42bd3492 Rename PROJECT_PATH to PROJECTS_PATH (#521)
* Rename PROJECT_PATH to PROJECTS_PATH

Co-authored-by: openhands <openhands@all-hands.dev>

* Apply suggestion from @enyst

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-16 16:20:44 +00:00
Robert Brennanandopenhands b5bb513c82 feat(conversation-panel): persist filter-menu preferences in localStorage (#510)
Persist the conversation list filter menu's show/hide toggles (older
conversations + repo/branch metadata) across reloads via a new
zustand/persist-backed store, following the same pattern as
home-store and workspaces-store.

The store is intentionally shaped so additional filter-menu options
can be added with just a field, a setter, and a toggle action — no
extra storage plumbing required.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-16 16:09:05 +00:00
Hiep Le 0e94c5a369 fix: right-align action buttons in modals/dialogs (#520) 2026-05-16 22:35:26 +07:00
Hiep Le cd7def89db fix(frontend): auto-focus chat input on mount (#518)
* fix: auto-focus chat input on mount

* refactor: update the code based on feedback
2026-05-16 21:44:05 +07:00
Hiep Le b171515282 test(snapshots): fix docker-workspace-browser for chat-first home (#515)
* ci(snapshots): upload baseline artifact even on partial failure

* test(snapshots): update docker-workspace-browser for chat-first home

* revert: snapshot-tests.yml
2026-05-16 21:14:27 +07:00
f37cd5bd17 Improve UI consistency: themes, tokens, chat chrome, and right-panel tabs (#508)
* fix(conversation): cap chat column width at 800px

Replace responsive max-w-4xl / max-w-6xl with max-w-[800px] so the
middle column stays narrower on large viewports.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(chat): Connect Repo CTA and hide empty branch pill

- Use COMMON$CONNECT_REPO with FolderOpen when no repo/workspace is linked
- Show branch control only when selectedBranch is set (drop No Branch)
- Cap chat interface wrapper at max-w-[800px] without right-panel width coupling

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): refine input controls and local auth fallback

Improve chat input pills and model dropdown interactions while ensuring local agent-server auth uses the configured session key for default-local and cloud-proxy calls to avoid stale-key 401s.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): align attachment and placeholder control styling

Move the file-attach trigger into the chat action controls so it sits before Tools, and restyle it as a grey plus button with a circular hover state to match adjacent controls. Also align the chat input placeholder color with the same neutral control tone for visual consistency.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): simplify agent status labels and tone

Shorten English agent-status messages for the chat pill and align the status text color with the other grey controls for a more consistent compact UI.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): align model popover settings row styling

Add an LLM Settings action to the model popover and normalize its layout, spacing, and divider treatment to match existing dropdown menu patterns while keeping left-aligned positioning.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): restyle status controls and move send action

Make the agent-status control transparent by default with gray-to-white icon hover behavior, and move the submit button to the bottom-right controls area beside agent status.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): tighten spacing above git control bar

Reduce the top margin before the git control bar so it better matches the bottom spacing around the chat action controls.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): gate submit button on input content

Keep the send button inactive until the input has non-whitespace text, and align the revised button sizing/positioning with the bottom action row layout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): streamline overlays and remove legacy event rails

Unify chat control styling and overlay behavior so status/typing/scroll controls float above the thread without adding layout bars, and remove left-rail/checkmark affordances from grouped and generic event cards for a cleaner stream.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): tighten status indicator spacing

Reduce status indicator pill padding and icon size, and add right text padding to balance the compact layout in the chat control overlay.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): prioritize centered scroll control over loader

Keep the scroll-to-bottom control centered and visible whenever the user is away from the bottom, and use solid base/hover fills so it matches the updated chat surface styling.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): soften conversation event header styling

Use the lighter gray chat tone for conversation event header labels/icons and switch those labels to normal weight so grouped event rows match the updated control styling.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): refine markdown spacing and divider styling

Tighten markdown vertical rhythm in chat content, add a shared grey horizontal-rule renderer, and tune heading hierarchy to medium/compact styles for clearer structure without heavy emphasis.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): tighten vertical spacing in action event rows

Reduce stacked margins and paddings across grouped action rows, generic event cards, and collapsible thinking blocks so adjacent conversation entries read as a denser, more consistent stream.

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(design): add app gray palette reference artifacts

Capture the current gray color usage in dedicated SVG references, including both a curated palette and a strict exhaustive inventory for design and UI consistency work.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): align compact input overflow menus with menu conventions

Keep add-file pinned inline, collapse controls only when width truly runs out, and switch overflow entries to standard context-menu row/submenu patterns while preserving the send button layout at tight widths.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation-panel): use list filter icon for older filters

Swap the older-conversations summary toggle icon to ListFilter so it matches the intended sidebar filter affordance.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): complete local workspace launch flow in git controls

Switch the local git control CTA from repository connection to workspace launching, including an above-button workspace menu and automatic add-workspace modal when none exist. This also captures the pending chat action/menu styling and test updates in the current working tree.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): relocate desktop vertical padding to input controls

Remove desktop top/bottom padding from the main chat panel and apply equivalent bottom spacing to the chat control area so the open repo/workspace controls and input footer keep consistent breathing room.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation): remove bottom margin from chat pane header

Drop the chat header bottom margin so the conversation title row sits flush with the content below.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): refresh git control bar immediately after Connect Repo

The "Connect Repo" empty-state in the chat input footer kept rendering
even after the Open Repository modal had successfully launched a clone
and the agent had reported the repository as ready. The bar would only
heal after a hard refresh (or never, on cloud backends).

Three independent bugs were stacking:

1. Optimistic update was writing to the wrong React Query cache key.
   `useUpdateConversationRepository.onMutate` called `setQueryData`
   with `["user", "conversation", id]` (3 elements), but
   `useUserConversation` reads from
   `["user", "conversation", id, backendId, orgId]` (5 elements).
   `setQueryData` requires an *exact* key match, so the update landed
   on an orphan cache entry that no observer ever read. Switched to
   `setQueriesData`/`getQueriesData` with the 3-element prefix so the
   optimistic write actually reaches the active query (Tanstack v5
   prefix-matches `setQueriesData` filters). Also normalized
   `branch`/`gitProvider` to `null` to match the shape produced by the
   server-side refetch and prevent identity-flicker between the two
   updates.

2. Cloud `batchGetCloudConversations` / `searchCloudConversations`
   ignored the local repo selection entirely. For local backends
   `toAppConversation` overlays `selected_repository`/`selected_branch`/
   `git_provider` from `localStorage`, but the cloud path returned the
   raw SaaS payload — and the SaaS often returns `null` for those
   fields until its own background hydration finishes. So every
   refetch (mutation invalidation, 30s poll, panel mount) overwrote
   the optimistic value with `null` and the bar snapped back to
   "Connect Repo". Added `overlayStoredRepoSelection` which fills only
   the `null` slots from local storage; populated server values still
   win, so we don't shadow real backend changes.

3. `updateConversationRepository` overwrote the entire metadata blob.
   `setStoredConversationMetadata` is replace-not-merge, so calling it
   with just `{selected_repository, selected_branch, git_provider}`
   silently dropped `selected_workspace` (the local-folder attach
   marker used by the Files tab to default to diff view, see the
   "Files tab diff-view default logic" note in `AGENTS.md`). Now reads
   the existing entry first and spreads it under the new repo fields.

Defense-in-depth changes:

- `useLocalGitInfo` now stays enabled until the conversation reports a
  *complete* repo tuple (`selected_repository` + `git_provider` +
  `selected_branch`), not just `selected_repository`. This lets the
  bar recover from partial-metadata cases (e.g. cloud hydration
  populates only the repo name first, or the user clones into a
  subdirectory of `working_dir`). The probe also gained a nested
  `find . -mindepth 2 -maxdepth 4 -name .git` fallback so a clone
  into `<workingDir>/<repo>/` is still detected after the direct
  `git remote get-url origin` in `<workingDir>` returns "no such
  remote 'origin'" (the agent-server pre-initialises every workspace
  as a worktree, so the parent directory always has a `.git` folder
  with no remote).

- `useUpdateConversationRepository.onSettled` invalidates
  `["local-git-info", conversationId]` so the bar re-probes
  immediately after a connect rather than waiting on the next 10s
  refetch tick.

- `git-control-bar.tsx`'s `hasRepository` predicate now keys off the
  *resolved* `selectedRepository` + `gitProvider` (which include the
  local-git probe's findings), not just the conversation field. This
  lets pull/push/PR buttons light up for local-workspace conversations
  whose repo metadata was inferred from `git remote`, matching what
  the repo + branch chips already showed.

Verification

I traced the failure mode by hitting the live agent-server directly:

  $ curl -s -X POST .../api/bash/execute_bash_command \\
      -H "X-Session-API-Key: \$KEY" \\
      -d '{"command":"git remote get-url origin", "cwd":"<workingDir>"}'

  git remote: error: No such remote 'origin'
  git rev-parse HEAD: ambiguous argument 'HEAD': unknown revision

confirming the worktree-without-remote shape that broke the direct
probe and forced the nested-find fallback.

Tests

  __tests__/hooks/mutation/use-update-conversation-repository.test.tsx
    - optimistically updates the cached conversation under the
      prefix-extended key used by useUserConversation
    - rolls back the prefix-keyed cache entry when the mutation rejects

  __tests__/api/cloud-conversation-service.test.ts (new)
    - overlays locally-stored repo selection onto
      batchGetCloudConversations results when the server returns nulls
    - prefers the cloud server values over locally-stored selections
      when present
    - leaves null entries untouched when the cloud server returns null
      for a missing conversation
    - returns an empty array without calling the proxy when no ids
      are provided
    - overlays repo selection on each item returned from
      searchCloudConversations

Wider sweep:

  npx vitest run __tests__/hooks/mutation \\
                __tests__/api/cloud-conversation-service.test.ts \\
                __tests__/api/conversation-metadata-store.test.ts \\
                __tests__/api/agent-server-adapter.test.ts \\
                __tests__/components/features/chat
  -> 22 files, 152 tests passed.

User-visible behavior after this change:

  1. Clicking Launch in the Connect Repo modal flips the bar to
     repo + branch chips immediately (optimistic update now reaches
     the active query).
  2. The bar stays flipped through the next refetch on cloud
     backends (overlay keeps the local selection visible until the
     SaaS catches up).
  3. Bar picks up nested clones within ~1s on local backends
     (local-git-info invalidation forces a re-probe instead of
     waiting on the 10s poll), and the nested-find fallback handles
     'clone into <workingDir>/<repo>/' flows.
  4. Pull/push/PR buttons now light up for local-workspace
     conversations whose remote was inferred from git remote.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(conversation): make right-panel drawer state session-only

The right-side drawer's open/closed state (`isRightPanelShown` /
`hasRightPanelToggled`) was persisted in localStorage, which made
the panel feel sticky in a way users didn't expect — it would still
be open after reloads or revisits even though they wanted a clean,
focused chat view.

Move drawer state fully into the in-memory Zustand store so it:
- always starts closed on app load (or on opening a conversation
  after a restart),
- survives in-app navigation because Zustand stays alive across
  React Router transitions,
- only persists tab selection (`selectedTab`), which is the part
  users do want to come back to.

The legacy `rightPanelShown` field is silently stripped from older
persisted blobs by `sanitizeStoredState`, so old localStorage data
doesn't churn or leak into the new schema.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(ui): standardize three-dots ellipsis trigger across the app

Different surfaces had drifted to slightly different "more options"
buttons:
- conversation header used a 24x24 icon with a hardcoded fill color,
- conversation cards in the side panel used a separate square
  `ellipsis.svg` glyph,
- the conversation tab bar used a 20x20 icon with bespoke colors,
- LLM profile rows wrapped the icon in a bordered button with
  yet another color.

Promote `EllipsisButton` to be the canonical trigger and route every
inline variant through it so size (w-4 h-4 / 16x16), color
(`text-[#9299AA]`), and hover treatment (`hover:text-white
hover:bg-white/10`) stay consistent everywhere. Layout-only overrides
(e.g. translate, opacity-when-paused) flow through `className`, and a
`testId` escape hatch keeps the existing `profile-menu-trigger`
selector working.

The chat-input overflow button intentionally keeps its pill-shaped
custom variant; a doc comment on `EllipsisButton` calls that out so
future contributors don't replace it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(design): add 15-stop cool grey palette and complete migration plan

Documents migration of ~98 scattered grey values across agent-canvas to a
unified 15-stop cool blue-grey family (hue ≈ 220–224°), with all existing
hex values, Tailwind utilities, CSS variables, and alpha variants mapped to
the nearest new token by RGB + lightness proximity.

Artifacts:
- cool-grey-palette.svg: visual palette strip + per-shade migration map
- cool-grey-migration.md: CSS/Tailwind definitions, per-file migration
  tables, alpha variant equivalents, and a 6-phase implementation plan

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): remove circular pill and ripple effect from autocomplete caret buttons

Replace the default HeroUI selector button styling (rounded-full, fixed dimensions,
hover fill) with a flat transparent icon and disable the press ripple via
selectorButtonProps={{ disableRipple: true }} on all Autocomplete instances.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(design): enforce single border color token across all UI elements

- Unify --oh-border-input to cool-grey-700 (same as --oh-border), eliminating
  the 3-way border split across inputs, cards, and dividers
- Replace border-neutral-600 in .button-base with border-[var(--oh-border)]
- Replace all border-tertiary usages (27 files) with --oh-border for outer
  borders and --oh-border-subtle for within-panel dividers
- Fix border-tertiary-light on toggle switch OFF state → --oh-border
- Fix border-t-tertiary on app-settings Git section divider → --oh-border-subtle
- Fix divide-tertiary in profiles-body → divide-[var(--oh-border-subtle)]
- Fix secrets table row dividers: --oh-border-subtle → --oh-border
- Fix files-tab toolbar and file-quick-row header lines → --oh-border
- Fix repo-connector and new-conversation card borders → --oh-border
- Remove bg-surface from automations-list and automation-detail routes so
  they inherit bg-base from the root layout, matching all other pages

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(design): migrate automations to shared tokens; fix secondary button, card, and hover strokes

- Replace all legacy tailwind.config.js color tokens in automations/ (27 files):
  bg-surface-card → bg-[var(--oh-surface)], bg-surface-elevated → bg-surface-raised,
  border-border → border-[var(--oh-border)], text-content-muted → text-muted,
  status/toggle/badge tokens → --oh-success/--oh-danger/--oh-muted variants
- Fix active-status-badge and status-badge inactive fills from bg-border → bg-surface-raised
- BrandButton secondary variant: yellow outline+text → border-[var(--oh-border)] text-white
  hover:bg-surface-raised; move hover:opacity-80 off base onto primary/tertiary only
- BrandButton primary: replace text-base (font-size conflict) with text-[var(--oh-color-base)]
  so all variants share the base text-sm font size
- Card primitive default/outlined themes: --oh-border-input → --oh-border (fixes visible
  mismatch between repo-connector and Start from Scratch cards on home screen)
- marketplace-card, skill-card, skills-toolbar: replace hover:border-white/40 with
  hover:border-[var(--cool-grey-500)] and focus:ring-primary/60 with focus:ring-[var(--oh-border)]
- automations routes: remove explicit bg-surface so pages inherit bg-base from root layout

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(design): normalize spinners, borders, and accent colors to design tokens

Replace hardcoded `border-primary`, `border-blue-500`, and `text-primary`
with cool-grey-aligned tokens (`border-white`, `border-white/20`,
`var(--oh-border)`, `var(--oh-muted)`) across loading spinners, modals,
dropdowns, and link styles. Switch dropdown selected-item highlight from
`--oh-interactive-active` to `--oh-interactive-selected`.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(theme): add runtime color theme switcher with OpenHands-Neutral palette

- Add src/themes/color-themes.ts: two themes (OpenHands-DeepSea /
  OpenHands-Neutral) with --cool-grey-* scale overrides and matching
  --heroui-* HSL channel overrides (default, content, background,
  foreground families) so HeroUI components and portalled popovers
  both respond to theme changes.
- Inject overrides via a <style> tag on [data-agent-server-ui] AND
  [data-theme=dark] so portal content rendered to document.body picks
  up the new palette alongside inline components.
- Add ThemeInput (SettingsDropdownInput) to Application Settings;
  applies the theme immediately on selection and persists to
  localStorage under openhands-color-theme.
- Add ColorThemeApplier to root.tsx so the persisted theme is
  re-applied on every page load with no flash.

Additional token fixes found during theme testing:
- Define --color-tertiary-alt → --oh-text-dim in tailwind.css so the
  ~25 placeholder:text-tertiary-alt / text-tertiary-alt usages
  (API key input, helper text, badges) resolve correctly.
- Fix environment-switch-overlay: replace bg-card / border-border /
  text-foreground with --oh-surface / --oh-border / --oh-foreground.
- Fix AutocompleteSection headings in model-selector: add
  classNames={{ heading: "text-[var(--oh-muted)]" }} so Verified /
  Other Models labels are readable against the dropdown background.
- Unify structural panel dividers: sidebar right edge + footer
  separator + files-tab tree divider all changed from
  --oh-border-subtle to --oh-border, matching the right panel.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): improve suggestion card hover to match secondary button style

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(theme): set OpenHands-Neutral as the default color theme

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): use 40px header rows and refine sidebar logo layout

Align chat and tabs headers to h-10, trim sidebar shell padding, and render
the logo with configurable size plus max-w-none so preflight does not shrink
it in the icon column. Add mock stubs for cloud-proxy and folder-browser;
default backend form kind to local; update sidebar tests. Fill SETTINGS$COLOR_THEME
locale strings so translation completeness passes.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): remove slit between chat and right panel resize columns

The flex resize handle used w-1, so percentage widths no longer summed to
100% and left a visible gap. Use zero layout width with an overlaid drag hit
area instead.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): polish right-panel tabs, border, and planner build row

Use a left-only panel border, stack the tabs header in a column with compact
left-aligned tabs, token-based active tab styling, and move the planner Build
control to a dedicated 40px row under a divider.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): measure tab bar overflow and split overflow menu open vs pin

Tabs stay on one row by showing only what fits and moving the rest into the
⋯ menu; opening a tab no longer toggles pin state. Minor polish for loading,
resize-handle hover, and right-panel toggle.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): keep resize grip line highlighted while dragging

Track panel drag state from useResizablePanels plus local hover so the
indicator line stays visible until release and after only when still hovered.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): shared right-drawer empty states and VS Code cross-origin layout

Introduce ConversationTabEmptyState for planner, browser, changes, tasks, and
VS Code embed fallback (muted icon, centered text, secondary BrandButton).
Also refine tab bar overflow grouping, panel toggle idle colors, terminal
transparent canvas, and task-list highlight tests.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): tighten chat input grip drag and commit threshold without param reassign

Align grip styling with the chat/panel resize handle, surface drag state to
the grip component, and require a small movement before treating pointer input
as a resize so post-drag clicks are ignored. Refactor drag tracking to avoid
mutating helper parameters flagged by ESLint.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): keep active unpinned tabs visible and tighten sidebar header padding

Show unpinned right-panel tabs in the strip while selected so the bar matches
the open view; add regression tests. Drop extra right padding on the expanded
sidebar logo row so the collapse control aligns closer to the rail.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor: remove unrelated file

* refactor: remove artifacts

* refactor: remove unrelated file

* refactor: remove unrelated file

* fix: failing tests

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: hieptl <hieptl.developer@gmail.com>
2026-05-16 20:10:42 +07:00
b7abb1e28b fix(settings): layout, SDK save UX, extensions parity, secrets/MCP polish (#512)
* fix(settings): show backend sync notice under left nav

Move BackendSyncedSettingsBadge from the per-section header into the settings
sidebar (desktop and mobile), matching the extensions layout pattern.

Also add missing locale strings for BACKEND$NAME_REQUIRED,
BACKEND$HOST_REQUIRED, and BACKEND$HOST_INVALID so translation completeness
checks pass.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(settings): unify layout scrolling, full-width column, and page subtitles

Let the settings shell own vertical scrolling with a sticky desktop nav and
centered max-w-5xl content. Drop nested 680px field caps so controls match the
column. Add grey subtitles under each section title from nav metadata. Inline
SDK save actions after fields, align profile/app save controls with the start,
and tidy application settings spacing plus onboarding LLM scroll.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(settings): nav layout, 800px column width, SDK view tabs

Split desktop sidebar and mobile drawer so fixed UI stays out of the scroll
flex row; remove extra aside top padding to align with page titles; reserve
badge height. Use 800px max width for settings, skills, and MCP content.
Replace SDK view toggles with white active underline tabs. Wrap settings nav
tests with ActiveBackendProvider for the sync badge.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(settings): match extensions nav scroll; stack SDK fields

Scroll only the settings main column (aside stays pinned like extensions),
mirror extensions horizontal padding on desktop, and drop SDK settings
two-column xl grid so schema fields always use one row per field.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): align extensions layout and polish secrets/MCP lists

Center MCP, Skills, and Plugins content in an 800px column like settings.
Use Lucide edit/delete icons with muted-to-white hover chips, a grey
secrets table panel with tighter rows and scroll cap, plain description
text, and refreshed Settings/Extensions nav labels.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(settings): align nav left inset and Lucide verification/app icons

Match desktop sidebar pl-8 to the main column’s top-8 spacing, and use
Lucide Shield and AppWindow for Verification and Application nav items.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor: translation.json

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: hieptl <hieptl.developer@gmail.com>
2026-05-16 17:45:39 +07:00
Hiep Le 097f743e52 feat: redesign home page as chat-first launcher with optional workspace/repo picker (#514) 2026-05-16 17:08:15 +07:00
Hiep Le 5c33c10b18 feat(frontend): add canvas_ui tool so the agent can drive the UI (#420)
* feat: add canvas_ui tool so the agent can drive the UI

* refactor: update the code based on feedback

* fix: failing tests

* refactor: update the code based on feedback
2026-05-16 16:09:43 +07:00
Rohit Malhotraandopenhands 60e103eec5 fix: route cloud runtime bash/file calls through cloud proxy (#507)
* fix: route cloud runtime bash/file calls through cloud proxy

useWorkspaceFiles, useLocalGitInfo, useHasGitCommits were all building
RemoteWorkspace with getAgentServerClientOptions({ conversationUrl }),
which resolves the host directly to the cloud runtime URL when a cloud
conversation is active (e.g. *.prod-runtime.all-hands.dev).  This caused
CORS errors because the browser made the fetch from localhost.

useWorkspaceFileContent had the same issue for GET /api/file/download.

The fix centralises these operations in a new
AgentServerRuntimeService (src/api/runtime-service/) that mirrors the
pattern already used by agent-server-git-service and event-service:

  if (active.kind === 'cloud' && conversationUrl)
    -> callCloudProxy({ hostOverride: buildHttpBaseUrl(conversationUrl), authMode: 'session-api-key', ... })
  else
    -> SDK typed clients directly

- executeCommand  routes POST /api/bash/execute_bash_command
- downloadFile    routes GET  /api/file/download

use-local-git-info helper functions (probeGitInfoAtDir,
probeNestedRepoInDir) are refactored from taking a RemoteWorkspace
instance to taking a RunCommand callback, keeping the helpers pure.

Co-authored-by: openhands <openhands@all-hands.dev>

* test: cover AgentServerRuntimeService cloud/local routing

Adds 12 unit tests for AgentServerRuntimeService covering:
- executeCommand local path: RemoteWorkspace constructed with resolved
  options; callCloudProxy never called
- executeCommand cloud path: callCloudProxy invoked with POST to
  /api/bash/execute_bash_command, correct hostOverride, body, session-
  api-key auth, and timeoutSeconds; RemoteWorkspace never created; cwd
  omitted when undefined; null stdout/stderr normalised to empty strings;
  null conversationUrl falls back to local
- downloadFile local path: FileClient constructed with resolved options;
  callCloudProxy never called
- downloadFile cloud path: callCloudProxy invoked with GET to
  /api/file/download with URL-encoded path, blob responseType, session-
  api-key auth; FileClient never created; Blob→ArrayBuffer round-trip
  preserves content; null conversationUrl falls back to local

Also adds a cloud-backend integration test to
use-workspace-file-content.test.tsx confirming the hook routes
downloads through callCloudProxy instead of FileClient when a cloud
backend is active, and that callCloudProxy is called with the correct
path/auth/responseType.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-16 04:40:41 +00:00
Rohit Malhotraandopenhands 43e6db6919 docs: add API access rules to AGENTS.md and code review skill (#506)
Document the two mandatory API access conventions that are enforced by
the CI test src/api/no-direct-agent-server-calls.test.ts:

1. All agent-server calls must use typed @openhands/typescript-client
   classes (ConversationClient, FileClient, VSCodeClient, ServerClient,
   RemoteWorkspace, RemoteEventsList) instantiated via
   getAgentServerClientOptions() -- never raw axios/fetch.

2. All cloud SaaS and runtime-sandbox calls must go through
   callCloudProxy() in src/api/cloud/proxy.ts to avoid CORS, using
   hostOverride for runtime-sandbox URLs and authMode='session-api-key'
   for those endpoints.

AGENTS.md gets a full '## API Access Rules' section with client
listings, option helper references, CORRECT/WRONG code examples, and
the allowed-exceptions list.

The custom-codereview-guide.md skill gets a '## Frontend API Access
Conventions' section with DO NOT APPROVE triggers, forbidden pattern
lists, correct examples, and a note about the silent hostOverride bug.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-16 04:30:01 +00:00
Rohit Malhotraandopenhands acb04fd9ba perf(snapshots): skip retries on comparison pass, suppress consent modal (#505)
- Add --retries=0 to test:e2e:snapshots: snapshot pixel-diff failures are
  deterministic — retrying cannot fix them. On a PR that changes 36/60
  snapshots this tripled execution count and added ~2 min to the comparison
  pass.

- Extract seedLocalStorage() helper (tests/e2e/snapshots/support/) that
  seeds openhands-onboarded and openhands-telemetry-consent in a single
  addInitScript call. All 13 snapshot specs now use it instead of
  duplicated inline addInitScript blocks.

- Pre-seeding openhands-telemetry-consent='denied' eliminates the race
  condition where changes-tab timed out at 60 s (x3 retries = 3 min) because
  dismissConsentModal fired before the modal rendered with domcontentloaded.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-16 00:04:29 -04:00
14d2e9454b test(snapshot): changes tab diff viewer + backend management UI (6 tests) (#450)
* test(snapshot): changes tab diff viewer + backend management UI (6 tests)

Pre-seed MOCK_GIT_CHANGES with M/A/D entries (using AgentServerGitChangeStatus
values: UPDATED/ADDED/DELETED) so changes-tab tests can exercise the file list,
Monaco diff viewer, and deleted-file placeholder without per-test MSW manipulation.

Expose window.__setMockGitChanges__ so the empty-state test can clear the list
after boot and trigger a React Query refetch via __TEST_INVALIDATE_QUERIES__,
avoiding a full page reload that would reinitialise module state.

Backend management tests exercise the selector dropdown, add-backend modal, and
manage-backends modal — all driven by localStorage seeding via addInitScript.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: update baseline snapshots [skip ci]

* ci: trigger re-run against CI-generated baselines

* fix(snapshot-tests): mask Monaco editor for stable CI screenshots; fix unit test

- changes-tab spec: mask data-testid=editor-container so Monaco's sub-pixel
  font hinting (which varies per OS) doesn't cause false pixel-diff failures
- mock-conversation-handlers test: update assertion to match the new pre-seeded
  MOCK_GIT_CHANGES (3 M/A/D entries) instead of the previous empty array

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: update baseline snapshots [skip ci]

* ci: trigger re-run against CI-regenerated baselines (Monaco mask + unit test fix)

* fix(snapshot-tests): normalize RandomTip height via addStyleTag for stable empty-state screenshot

RandomTip renders a randomly-chosen tip whose line-count varies, causing the
flex-1 container above it to have different heights across runs. Fix by injecting
a CSS rule via page.addStyleTag() that pins .text-m.bg-tertiary.p-4 to 80px
(visibility:hidden so the variable text is invisible) — layout is now deterministic.
Switch back to screenshotting the full files-tab panel since dimensions are stable.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: update baseline snapshots [skip ci]

* ci: trigger re-run against baselines (empty-state RandomTip height fix)

* fix(snapshot-tests): use inner content div for empty-state screenshot to avoid left-strip artefact

Screenshot files-tab's last direct div child (the flex-1 content wrapper)
instead of the outer main element.  During CI baseline generation the outer
main's bounding box occasionally captured a ~30px left-panel overlay artefact
that made the baseline permanently diverge from subsequent verification runs.
Targeting the inner wrapper excludes the outer-element overflow while still
showing the full empty-state (icon + 'no changes yet' text + hidden tip area).

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: update baseline snapshots [skip ci]

* ci: validate against fresh inner-div empty-state baseline

* test(snapshot): extended backend UI flows — 12 tests, 19 screenshots

Add backends-extended.snapshot.spec.ts covering 8 behaviour flows
with iterative screenshot captures at each state transition:

Flow 1a  Blank add form — Save disabled until name+host filled
Flow 1b  Local backend — Save enabled with name+host, no API key needed
Flow 1c  Cloud backend — Save disabled without API key, enabled with it
Flow 2a  Host auto-infers Local kind; OAuth section disappears
Flow 2b  Cloud-domain URL keeps Cloud kind; OAuth section shows
Flow 2c  Manual kind selection locks type (touchedKind=true) even when
         a cloud URL is later typed into the Host field
Flow 3   OAuth Login button disabled while host is empty; enabled once filled
Flow 4   Remove backend: shows ConfirmationModal → Cancel keeps row →
         Confirm removes it from the list (4 screenshots)
Flow 5   Edit modal pre-populates name/host/key from stored backend
Flow 6   Switch active backend: environment-switch overlay captured via
         page-level screenshot + animation override so the card is
         opaque at frame-0; after-switch state verified via selector label
Flow 7   Whitespace-only host keeps Save disabled; syntactically invalid
         URL is accepted by the frontend (no URL-format validation)
Flow 8   Cancel add form: dismisses modal, Manage Backends confirms no
         phantom entry was saved

Notable decisions:
- Uses body[data-environment-switching="true"] as the early DOM signal
  before React paints the portal div for the switch overlay
- Adds inline style-tag override before the overlay screenshot because
  .environment-switch-overlay > div has opacity:0 at animation frame 0;
  Playwright's animations:"disabled" pauses there, making the card
  invisible without the override
- Backends seeded via page.addInitScript localStorage injection so
  tests are fully self-contained with no MSW state dependency

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: update baseline snapshots [skip ci]

* ci: validate extended backend snapshot tests against CI baselines

* ci: always post snapshot PR comment even when test generation step fails

The 'Post snapshot report to PR' step was skipped whenever 'Generate
current PR snapshots' exited non-zero (e.g. a test crash like a hidden
element, not just a snapshot diff). GitHub Actions skips steps without
an always() guard when a prior step fails.

Add always() so the comment is posted regardless — showing diffs or
the test failure output — which was the intended behaviour.

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: fix snapshot comment - remove tracked screenshots, add crash reporting

Three fixes:

1. Remove 28 git-tracked snapshot PNGs from this branch.
   These were committed by the old baseline-in-git workflow before #482
   migrated to artifact storage. Because they stayed tracked (gitignore
   doesn't untrack already-indexed files), every CI checkout put them in
   tests/e2e/__snapshots__/ BEFORE the baseline artifact was downloaded.
   The Save step then copied them into /tmp/main-baselines, making the
   new tests appear as 'Unchanged' instead of 'New' in the PR comment.

2. Add 'Clear snapshot directory before downloading baselines' step.
   Wipes tests/e2e/__snapshots__/ before the artifact download so any
   future accidentally-tracked files can never contaminate the baseline.

3. Surface test crashes in the PR comment.
   - Generate step gets continue-on-error + an id so subsequent steps
     can read its outcome.
   - GENERATE_OUTCOME is passed to the comment script.
   - If outcome == 'failure', a GitHub-flavoured WARNING callout is
     prepended to the comment with a direct link to the CI run logs.
   - A dedicated 'Fail if snapshot generation had test crashes' step
     restores the job failure that continue-on-error absorbed.

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: use PR number in snapshot concurrency group for cleaner cancellation

The previous group used github.ref which resolves to refs/pull/{N}/merge
for PR events — correct but opaque. Using github.event.pull_request.number
makes the grouping explicit and human-readable (snapshot-tests-450), and
falls back to github.ref for main pushes and workflow_dispatch.

cancel-in-progress: true was already set, so new commits already cancelled
prior runs. This just makes the intent clearer.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: syntax error in post-snapshot-comment.mjs (] vs ) in lines.push)

lines.push(...) was accidentally closed with ]; instead of ); after
splitting the original lines = [...] array literal into a push call.
Caused a SyntaxError at startup, preventing any comment from being posted.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: snapshot test disabled states, changes-tab crash, and CI false-failures

Three fixes:

1. BrandButton disabled visual styling (brand-button.tsx)
   disabled:opacity-30 pseudo-class was not applying in Vite dev mode
   (Tailwind v4 + postcss-prefix-selector interaction), making disabled
   and enabled buttons visually identical in snapshot screenshots.
   Fix: add isDisabled conditional class directly ('opacity-30
   cursor-not-allowed pointer-events-none') so the disabled appearance
   is applied regardless of whether :disabled pseudo-class works.

2. changes-tab test crash (changes-tab.snapshot.spec.ts)
   Test waited for data-testid='files-tab' but the right panel always
   starts CLOSED (isRightPanelShown = false is session-only Zustand
   state; sanitizeStoredState strips any persisted rightPanelShown key).
   Fix: click data-testid='right-panel-toggle' after navigation to open
   the panel before waiting for files-tab. Also remove the no-op
   rightPanelShown: true from the localStorage seed.

3. CI false-failures for new snapshot tests (snapshot-tests.yml +
   post-snapshot-comment.mjs)
   The 'Fail if comparison found differences' step fired on
   'missing baseline' failures (expected for new tests in a PR) as
   well as actual pixel-diff failures.
   Fix:
   - post-snapshot-comment.mjs outputs has_changes=true/false to
     GITHUB_OUTPUT (true only when changed.length > 0, i.e. real diffs)
   - 'Fail if' step now checks steps.post-comment.outputs.has_changes
     == 'true' instead of compare.outcome == 'failure', so PRs that
     only add new snapshot tests pass CI cleanly.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: reject invalid host URLs in backend form; use http for local addresses

Two related fixes to backend host validation / normalisation:

1. isValidHostUrl() — reject invalid host strings
   canSubmit previously only checked host.trim().length > 0, so
   garbage like 'not://:::a valid url!!!' passed through and enabled
   the Save button. isValidHostUrl() adds two checks before the URL
   constructor: (a) the trimmed value must be non-empty, (b) it must
   contain no whitespace. This catches the test-case input whose spaces
   are the tell-tale sign of a malformed value.

2. normalizeHost() — http:// for local addresses
   Bare hostnames (no explicit scheme) were unconditionally prepended
   with https://, but local servers almost never have TLS certificates.
   The new isLocalAddress() helper detects localhost, 127.x, RFC-1918
   private ranges (10.x, 192.168.x, 172.16-31.x), .local / mDNS names,
   and single-label hostnames — all get http:// instead of https://.
   Hostnames with dots that are not in those ranges (e.g. app.all-hands.dev)
   still default to https://. Explicit http:// or https:// prefixes are
   always preserved as-is.

Test update: the 'backend-add-invalid-url-accepted' snapshot is renamed
to 'backend-add-invalid-url-disabled' and the assertion flips from
not.toBeDisabled() → toBeDisabled(), reflecting the new behaviour.

Co-authored-by: openhands <openhands@all-hands.dev>

* feat: inline error feedback on Name and Host fields in BackendForm

Three parts:

1. SettingsInput gains error / showRequiredTag / onBlur props
   - error?: string — red border on the input plus a small red alert
     paragraph below it (role=alert, data-testid=${testId}-error, linked
     via aria-describedby).
   - showRequiredTag?: boolean — renders a red * after the label to
     signal that the field is mandatory, consistent with OptionalTag.
   - onBlur?: () => void — forwarded directly to the <input>.
   - aria-invalid is set automatically when error is truthy.

2. BackendForm wires touched state → errors → inputs
   - nameTouched / hostTouched (both false on open, set on blur)
   - nameError: 'Name is required' when touched + empty
   - hostError: 'Host is required' when touched + blank/whitespace;
                'Enter a valid URL (e.g. http://localhost:8080)' when
                touched + non-empty but fails isValidHostUrl()
   - Both name and host SettingsInputs get showRequiredTag, the
     computed error, and onBlur={() => setXTouched(true)}.
   Errors are intentionally suppressed until blur so the form does not
   scold the user before they have had a chance to type anything.

3. Three snapshot tests call .blur() after .fill() to reveal errors
   - backend-add-name-only-disabled: focus+blur empty host → 'Host is
     required' appears below the Host field.
   - backend-add-whitespace-host-disabled: blur after fill('   ') →
     same 'Host is required' (whitespace counts as empty).
   - backend-add-invalid-url-disabled: blur after invalid URL fill →
     'Enter a valid URL...' appears below the Host field.
   The backend-add-blank-disabled snapshot is unchanged (neither field
   touched, no errors yet — correct for the fresh-open state).

New i18n keys: BACKEND$NAME_REQUIRED, BACKEND$HOST_REQUIRED,
BACKEND$HOST_INVALID (English only; other locales fall back to en).

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: prettier formatting on nameError / hostError ternaries

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: disable OAuth Login button until name and host are both valid

Previously the 'Login with OpenHands' button was enabled as soon as
a non-empty host was typed, even when the Name field was still blank.
This let users go through the full OAuth device-flow only to find they
still couldn't save because the name was missing.

Gate isDisabled on !name.trim() || !isValidHostUrl(host) so the button
stays disabled until the form is actually ready to save (modulo the
API key that OAuth itself will provide).

Update Flow 3 snapshot test to fill the name before asserting the
button becomes enabled, and update the test description accordingly.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: address PR review feedback (#450)

IPv6 parsing fixes (normalizeHost / isLocalAddress):
- normalizeHost: handle bracket notation [::1]:8080 (extract ::1),
  bare IPv6 addresses with multiple colons (use whole string as
  hostname), and regular host:port as before — prevents split(':')[0]
  from grabbing only the first segment of a multi-colon IPv6 address
- isLocalAddress: strip brackets before comparison; add :: (any-addr),
  ::ffff:127.x.x.x (IPv4-mapped loopback), fe80::/10 (link-local),
  fc00::/7 (unique local); tighten single-label check to exclude
  addresses that contain colons (bare IPv6 non-local addresses)

Mark fields touched on submit attempt:
- handleSubmit sets nameTouched + hostTouched when !canSubmit so
  inline errors appear for keyboard users who press Enter on an
  incomplete form

Snapshot workflow comparison-crash detection:
- Pass COMPARE_OUTCOME=${{ steps.compare.outcome }} to post-comment
- post-snapshot-comment.mjs reads COMPARE_OUTCOME and prepends a
  '[!WARNING]' block when the comparison step itself crashed
  (timeout/OOM) so the comment accurately reflects the run state
  instead of silently showing an incomplete/empty diff table

Remove unnecessary serial mode from backends-extended snapshot suite:
- Each test calls setupPage() with fresh state on its own Playwright
  page; no shared mutable state exists between tests, so serial is
  unnecessary and slows the suite

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-05-15 23:14:30 -04:00
Tim O'Farrellandopenhands b2b71855c6 feat(dev): surface dev-stack runtime services in agent system prompt (#503)
* feat(dev): surface dev-stack runtime services in agent system prompt

Add a structured 'runtime services' info object that the dev launchers
(`dev:safe`, `dev:automation`, `dev:docker`, and the published
`agent-canvas` binary) propagate to the frontend via
`VITE_RUNTIME_SERVICES_INFO`. The frontend renders it into a
`<RUNTIME_SERVICES>` markdown block and attaches it as
`AgentContext.system_message_suffix` on every `POST /api/conversations`.

This means agents start each conversation knowing exactly what services
exist in the current dev stack (ingress URL, automation backend URL +
`/api/automation` prefix, auth header, etc.), instead of having to probe
or — worse — assume `localhost:8000` is the automation server when it is
actually the Agent Server they are running inside of.

URLs are written from the agent's point of view: dockerless modes use
`localhost`, `dev:docker` uses `host.docker.internal`. When automation
isn't running in the current mode (e.g. `dev:safe`), the block says so
explicitly so agents know to skip `/api/automation` calls.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(runtime-services): address review feedback on PR #503

- Validate required `agentServerPort` in `buildRuntimeServicesInfo`;
  previously a missing port baked `http://localhost:undefined` into
  the agent's system prompt.
- Skip the automation entry when the supplied `automation` object has
  no `port` (e.g. a bare `{}` from a misconfigured launcher).
- Rename the JSON service key from `vite` to `frontend` and add a
  `kind: "vite" | "static"` discriminator + mode-aware description,
  so static-build dev stacks (`dev:docker`, the published binary, ...)
  no longer surface a misleading "Vite dev server" line in the agent
  system prompt. The renderer still accepts the legacy `vite` key.
- Anchor the "don't guess" warning to the actual agent-server URL from
  runtime info instead of hardcoded `localhost:8000`, since the
  agent-server uses different ports across dev modes (18000 in
  dev:safe, 8000 in dev:docker, ...).
- Plumb `frontendKind` through `buildAutomationRuntimeServicesInfo`
  and stamp `config.frontendKind` in `dev-with-automation.mjs::main`
  so both Vite spawn and static-build paths describe the frontend
  correctly.
- Expand AGENTS.md with the JSON schema of `VITE_RUNTIME_SERVICES_INFO`
  and a concrete example of the rendered `<RUNTIME_SERVICES>` block.
- Tests: assert the new URL-in-warning behavior, the new `frontend` /
  legacy `vite` rendering, the `agentServerPort`-required guard, the
  `automation: {}` skip, and the legacy `vitePort` alias.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-15 20:53:45 -06:00
Tim O'Farrellandopenhands 8db469487b docs(readme): mention OpenAPI docs alongside UI URL (#502)
Both the agent server's FastAPI docs and the automation backend's docs
are reachable through the local ingress proxy on port 8000, so point
users at them right next to the existing UI URL in both the Docker and
non-Docker quickstart sections.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-15 19:16:55 -06:00
Tim O'Farrellandopenhands 4db59b8b94 Proxy agent-server FastAPI docs (/docs, /redoc, /openapi.json) through the ingress (#501)
* feat(ingress): route /docs to the agent server

Add /docs to the list of prefixes proxied to the agent-server in:
  - vite.config.ts (Vite dev server proxy)
  - scripts/dev-with-automation.mjs (ingress + static-server fallback)
  - scripts/dev-static.mjs (ingress + static-server fallback)

Update the explanatory comment in scripts/static-server.mjs to match.

This exposes the agent-server's FastAPI Swagger UI at `/docs` on the
ingress port, alongside the automation backend's existing
`/api/automation/docs`.

Co-authored-by: openhands <openhands@all-hands.dev>

* feat(ingress): also route /redoc and /openapi.json to the agent server

Without /openapi.json, the Swagger UI page served at /docs (added in the
previous commit) renders but fails to load any spec. /redoc is the
FastAPI-served ReDoc alternative and benefits from the same fix.

Routes are added everywhere /docs already is:
  - vite.config.ts (Vite dev server proxy)
  - scripts/dev-with-automation.mjs (ingress + static-server fallback)
  - scripts/dev-static.mjs (ingress + static-server fallback)

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-15 18:46:13 -06:00
Rohit Malhotraandopenhands 9203a72d64 feat(snapshot-ci): group PR comment snapshots by spec file (#499)
Previously every changed/new snapshot got its own ### heading and table,
making it hard to see which snapshots belong to the same feature or flow
(e.g. the four-step MCP Slack install flow, the five-step secrets
lifecycle, or the skills search/filter sequence).

Group all three sections (🔴 Changed, 🆕 New, ✅ Unchanged) by spec file:

- Replaced formatRelPath() with specFromRelPath() + groupBySpec() helpers.
- Changed / New: one ### heading per spec (with count when > 1 snapshot),
  then **bold name** + the 3-column expected|actual|diff table per snapshot.
- Unchanged: compact grouped bullet list — bold spec heading, then one
  bullet per snapshot name, no superfluous intro sentence.

No workflow changes required; the grouping is purely in the comment script.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-15 22:45:14 +00:00
Rohit Malhotraandopenhands 63704f65e5 test: rename sidebar nav label "New" → "Chats" to verify snapshot CI diff comment (#497)
* test: rename sidebar nav label New → Chats to trigger snapshot diff

Intentional one-line change to verify that the snapshot CI workflow
correctly posts a PR comment showing the expected/actual/diff images.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(snapshot-ci): save comparison test-results before update step clears them

Root cause: Playwright wipes its output directory (test-results/) at the
start of each new run.  The workflow runs the tests twice:
  1. npm run test:e2e:snapshots        → comparison, writes *-diff.png files
  2. npm run test:e2e:snapshots:update → regenerates baselines, clears
                                         test-results/ first, no diffs written

By the time post-snapshot-comment.mjs runs, all diff files are gone.
diffBySnapshotName is always empty, so every snapshot is classified as
"unchanged" even when Playwright reported 16 failures.

Fix:
- Add a "Save comparison test-results" step immediately after the
  comparison run that copies test-results/ to /tmp/comparison-results
  before the update pass can delete them.
- Pass COMPARISON_RESULTS_DIR=/tmp/comparison-results to the comment script.
- In post-snapshot-comment.mjs, read TEST_RESULTS_DIR from
  COMPARISON_RESULTS_DIR env var (falls back to "test-results" for local use).

Co-authored-by: openhands <openhands@all-hands.dev>

* docs: document snapshot CI comparison-results ordering in AGENTS.md

Co-authored-by: openhands <openhands@all-hands.dev>

* revert: restore sidebar nav label to New

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-15 18:34:00 -04:00
Rohit Malhotraandopenhands c3de18580d ci: increase Playwright CI workers from 1 to 2 (#495)
ubuntu-24.04 runners have 2 vCPUs. Each Playwright worker runs in its own
browser context so tests are fully isolated (MSW state, localStorage, and
React Query cache are all page-level). Doubling from 1→2 workers should
roughly halve wall-clock test time on CI with no risk of interference.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-15 16:53:27 -04:00