Commit Graph
7538 Commits
Author SHA1 Message Date
Rohit Malhotraandopenhands c76d855149 fix: include tools/ in npm package files (#904)
The tools/ directory containing canvas_ui_tool.py was missing from the
package.json 'files' list, so it was not shipped in the published npm
tarball. Users running the released 'agent-canvas' CLI would get:

  KeyError: "ToolDefinition 'canvas_ui' is not registered"
  Failed to import module 'canvas_ui_tool': No module named 'canvas_ui_tool'

because the agent-server couldn't find the Python module that
dev-safe.mjs exposes via OH_EXTRA_PYTHON_PATH.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-28 19:10:04 +00:00
Rohit Malhotraandopenhands 231367a62b feat(cli): add --info flag to show default stack versions and ports (#902)
agent-canvas --version only shows the package version. The new --info
flag reads config/defaults.json and prints the full picture: agent-canvas
version, default agent-server/automation/SDK versions, default ports,
and the env vars available to override them.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-28 19:02:26 +00:00
Tim O'Farrellandopenhands fc450ef0e9 fix: make load_public_skills default to true (opt-out) (#901)
PR #229 flipped shouldLoadPublicSkills() from opt-out (default true) to
opt-in (default false) for dev-latency reasons. PR #362 patched the
conversation-start path by hardcoding load_public_skills: true, but
that fix was inadvertently lost in the #457 refactor which reintroduced
shouldLoadPublicSkills() inside the new buildAgentContext() helper.

This restores the original opt-out semantics:
- shouldLoadPublicSkills() now returns true unless VITE_LOAD_PUBLIC_SKILLS
  is explicitly set to "false"
- .env.sample updated to document the opt-out form
- Tests updated to match the new default
- AGENTS.md updated (it already said "defaults to true" but the code
  disagreed — now they agree)

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-28 12:52:58 -06:00
Rohit Malhotraandopenhands bd894b0708 feat: add mock-LLM E2E test infrastructure (#833)
* feat: add mock-LLM E2E test infrastructure

Add a new category of E2E tests that exercise the full UI → agent-server →
LLM stack using a scripted mock LLM server instead of real LLM credentials.

The mock server uses openhands-sdk's TestLLM to serve deterministic OpenAI-
compatible responses (tool calls and text replies) over HTTP, so these tests
are fully reproducible and need no API keys.

The Playwright test drives the real UI:
  1. Creates an LLM profile via Settings > LLM Profiles
  2. Sets the profile as active (points at the mock server)
  3. Starts a new conversation from the home page
  4. Sends a user message and verifies the agent responds

Verification is three-layered:
  - Events API: polls for a successful terminal observation
  - Chat UI: asserts the bash output token appears in rendered messages
  - Chat UI: asserts the agent's final reply token appears

New files:
  - tests/e2e/mock-llm/scripts/mock-llm-server.py  (TestLLM HTTP server)
  - tests/e2e/mock-llm/utils/mock-llm-helpers.ts    (shared Playwright helpers)
  - tests/e2e/mock-llm/mock-llm-conversation.spec.ts (the test spec)
  - playwright.mock-llm.config.ts                    (dedicated Playwright config)

Run with: npm run test:e2e:mock-llm

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: make mock-LLM E2E assertions real + add CI workflow

Fixes three broken verification checks in the mock-LLM test:

1. User message no longer contains BASH_TOKEN or REPLY_TOKEN.
   The mock LLM ignores the prompt anyway (TestLLM pops scripted
   responses from a deque), so embedding tokens in the prompt just
   caused the UI assertions to pass vacuously from the user's own
   message text.

2. waitForNonUserMessageText now searches only agent/environment
   output containers (agent-message, environment-message,
   model-messages, event-group) instead of the whole document body.
   This is a positive selector strategy — no risk of false positives
   from sidebar text, nav labels, or user input.

3. Error banner assertion no longer swallows failures. The previous
   .catch(() => {}) meant the step could never fail even when an
   error banner was visible.

Also adds .github/workflows/mock-llm-e2e.yml — triggered on PRs
with the 'e2e-tests' label or manual workflow_dispatch. No secrets
needed (the mock LLM server is self-contained).

Co-authored-by: openhands <openhands@all-hands.dev>

* feat: add PR comment with test results to mock-LLM E2E workflow

The CI workflow now:
1. Captures test exit code without failing the step (so later steps run)
2. Renders a markdown report from Playwright's JSON output showing each
   test name with pass/fail/skip status, duration, and retry count
3. Posts (or updates) a PR comment via the existing upsert-pr-comment.mjs
   script, using a dedicated '<!-- mock-llm-e2e-report -->' marker
4. Expands failure details in a collapsible section with the error message
5. Writes the same report to the GitHub Actions step summary
6. Links to the workflow run and uploaded test artifacts
7. Fails the job at the end if the test exit code was non-zero

Also adds the json reporter to playwright.mock-llm.config.ts.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use venv for openhands-sdk in CI to avoid PEP 668 error

Ubuntu 24.04's system Python is externally managed (PEP 668), so
`uv pip install --system` fails. Fix by creating a dedicated venv
for the mock LLM server and passing the venv's python path via
MOCK_LLM_PYTHON env var.

The Playwright config reads `MOCK_LLM_PYTHON` (default: 'python3')
for the webServer command, so local usage is unchanged.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: retry loop for mock LLM server verification in CI

The litellm import takes ~7 seconds on CI, so the fixed 'sleep 3'
was too short. Replace with a 30-second retry loop that polls the
server every second until it responds, then performs the JSON
validation.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: add GET / health check to mock LLM server

Playwright's webServer readiness probe sends GET / to the configured
URL. The mock server only handled POST, returning 501 for everything
else. Playwright interpreted this as 'not ready' and timed out after
30 seconds.

Add a do_GET handler that returns 200 with a simple JSON status.

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: post fresh PR comment per run + cache Playwright browsers

- Post PR comment: switch from upsert-pr-comment.mjs (which found and
  updated a single marker-tagged comment) to `gh pr comment` so each
  CI trigger leaves its own comment with full test results history.
  Remove the COMMENT_MARKER from render-mock-llm-report.mjs since it
  was only used for the dedup lookup.

- Cache Playwright: add actions/cache for ~/.cache/ms-playwright keyed
  on package-lock.json hash. On cache hit, only install system deps
  (fast apt layer) instead of re-downloading the full Chromium binary.

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: remove Playwright cache (caused extraction hang)

The actions/cache@v4 step for ~/.cache/ms-playwright reproducibly
caused npx playwright install to hang during Chrome zip extraction
(7+ min with no output, vs 24s without caching). The uncached install
completes in ~24s which is fast enough — remove caching for now.

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: move Playwright install before uv/openhands-sdk setup

Playwright's Chrome zip extraction hangs reproducibly when run after
the uv venv + openhands-sdk install steps (7+ min with no output).
The snapshot-tests workflow, which installs Playwright right after
npm ci, completes in ~21s on the same commit at the same time.

Move Playwright install immediately after npm ci — before uv, SDK,
and mock-server verification — to match the working step order.

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: split Playwright install into deps + browser download

Split 'npx playwright install --with-deps chromium' into two steps:
1. install-deps (apt packages only, no browser download)
2. install (browser download + extraction only)

This isolates which phase is hanging: the combined --with-deps flag
runs both in a single process, and the extraction hangs reproducibly
in this workflow despite identical config to snapshot-tests (which
works in 21s). Splitting may avoid whatever interaction causes the
extraction to stall.

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: pin Node 24.15 to fix Playwright install hang

Node 24.16.0 introduced a zip-extraction regression (nodejs/node#63487)
that causes 'playwright install' to hang indefinitely after download
completes for Playwright < 1.60.0. This repo uses Playwright 1.59.1.

The hang was reproduced 4 times on this workflow — download finishes
in ~3s but extraction never completes (7+ minutes of silence).
Meanwhile snapshot-tests (same config) worked because its runner
resolved to Node 24.15.0.

Pin to 24.15.x until the project upgrades to Playwright >= 1.60.0,
which includes a fix for the extract-zip interaction.

Also revert the split install-deps / install experiment back to the
original single 'npx playwright install --with-deps chromium' command.

Ref: microsoft/playwright#41000, microsoft/playwright#40724

Co-authored-by: openhands <openhands@all-hands.dev>

* ci: tighten mock-LLM test timeouts and remove CI retries

Mock LLM responses are instant, so the generous timeouts were causing
CI to hang for 10+ minutes when a test fails:
- retries: 1→0 in CI (mock tests should be deterministic)
- test timeout: 120s→60s (mock responses are instant)
- polling timeouts: 60s→30s for bash observation and chat text checks

Before: 3 tests × 120s timeout × 2 attempts (retry) = up to 12 min
After:  3 tests × 60s timeout × 1 attempt = up to 3 min on failure

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: step 3 conversation creation + add global timeout

Step 3 was clicking the home-chat-launcher container div (a passive
wrapper) instead of using the chat input to create a conversation.
The div click did nothing, and the test timed out waiting for
navigation to /conversations/<id>.

Fix: type into the home-page chat input and click submit — this is how
real users create conversations from the home page.

Also:
- Add globalTimeout (10 min in CI) to cap the entire Playwright run
  so teardown hangs don't waste CI time
- Reduce job timeout-minutes from 20 to 15

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: teardown hang via exec + add diagnostic logging

1. Prefix webServer command with 'exec env' so the shell is replaced
   by the npm process. Without exec, Playwright's SIGTERM kills the
   shell but npm's children (uvx, agent-server, vite) survive as
   orphans, causing the step to hang for 6+ minutes after tests finish.

2. Add diagnostic logging to waitForSuccessfulBashObservation — on
   timeout, the error message now includes the count and kinds of
   events the API actually returned, so we can tell whether the
   conversation never started vs the observation format changed.

3. Add a pre-flight API check in step 3 that verifies the mock-LLM
   profile's base_url is active in server settings before creating a
   conversation. If steps 1+2 didn't persist correctly, this fails
   fast with a clear message instead of timing out on empty events.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: profile check via /api/profiles + timeout wrapper for teardown

1. The pre-flight check was querying /api/settings which doesn't
   contain profile-based LLM config. Fix: query /api/profiles and
   assert active_profile matches the expected profile name.

2. Wrap the Playwright command in 'timeout --kill-after=30 8m' so
   if webServer teardown hangs (orphaned agent-server/vite processes
   ignoring SIGTERM), the entire process tree gets SIGKILL'd after
   8.5 minutes instead of waiting for the 15-min job timeout.

Co-authored-by: openhands <openhands@all-hands.dev>

* debug: dump first observation's raw structure on failure

The events API is returning events (agent-server logs show the bash
command was executed), but isSuccessfulBashObservation can't find a
match. Dump the first observation's full JSON structure so the next
CI run shows exactly what fields the API returns.

Co-authored-by: openhands <openhands@all-hands.dev>

* debug: dump raw event structures to discover API format

Previous diagnostic showed 8 events all with 'unknown' kind —
meaning the events don't have action.kind or observation.kind
properties. Dump the full JSON of the first 3 events to discover
the actual field structure.

Co-authored-by: openhands <openhands@all-hands.dev>

* debug: dump ALL event kinds + first non-stats event structure

Previous dump only showed first 3 events (all stats/state updates).
The observation events are likely in positions [3]-[7]. New diagnostic
shows all event kinds and dumps the first non-stats event so we can
see the actual action/observation format.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use correct APIs for mock-LLM E2E verification

The agent-server's conversation events API returns MessageEvents (not
nested ActionEvent/ObservationEvent), and tool executions live in the
separate bash events API (/api/bash/bash_events/search).

Changes:
- waitForSuccessfulBashObservation: now queries /api/bash/bash_events/search
  with kind__eq=BashOutput, checks stdout/stderr for BASH_TOKEN
- waitForAgentMessageContaining: new helper that checks conversation
  events API for agent MessageEvents containing a given token
- Step 3 verification now:
  1. Bash tool execution via bash events API
  2. Agent reply via conversation events API
  3. Reply token in chat UI (proves full round-trip)
  (Removed BASH_TOKEN UI check — it may not render in chat)

Co-authored-by: openhands <openhands@all-hands.dev>

* perf: cache Playwright browser binaries in CI

Split 'playwright install --with-deps chromium' into two steps:
1. 'playwright install chromium' (only on cache miss) — downloads ~200MB
   of browser binaries, cached via actions/cache keyed on PW version + OS
2. 'playwright install-deps chromium' (always) — installs apt system
   libraries needed by the browser (fast, mostly pre-installed on runner)

This should save ~30-60s on cache-hit runs since the browser download
is the slowest part of the Playwright setup.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: accept BashOutput with null stdout (exit_code=0 proves execution)

The bash events API returns BashOutput events where stdout can be null
even for successful commands (order:0 event with exit_code:0). Accept
null stdout with exit_code 0 as proof of successful execution.
Also dump all bash events (not just first) for CI diagnostics.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: don't fail CI when tests pass but teardown hangs

The Playwright webServer teardown can hang when the agent-server process
doesn't respond to SIGTERM (a known issue with uvicorn child processes).
The timeout wrapper kills the process tree after 8 min, but this was
incorrectly mapped to test failure.

Now when timeout triggers:
1. Check if test-results-mock-llm/results.json exists (Playwright writes
   this before teardown starts)
2. Parse it: if every spec/test has status 'passed', mark as success
3. Only fail if results.json is missing or has actual test failures

Also includes the Playwright cache and bash events API fix.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: background Playwright so shell survives teardown timeout

The previous 'timeout' wrapper killed the entire process group including
our bash shell, so the results.json check never ran. Now:
1. Run Playwright in background (&)
2. Poll every 2s up to 8 min
3. If still running, SIGTERM then SIGKILL the process group
4. Our shell is still alive → check results.json
5. If all tests passed, mark as success despite teardown hang

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: poll for results.json during run, not after kill

Playwright's JSON reporter writes results.json after all tests complete
but before webServer teardown. Poll for the file appearance during the
run (Phase 1), then only kill the hanging process if tests are done
(Phase 2). This way we catch results.json while Playwright is still
alive but stuck in teardown, and can correctly determine pass/fail.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: parse Playwright stdout for pass/fail (not results.json)

Playwright's JSON reporter only writes results.json on process exit,
which is blocked by the webServer teardown hang. The line reporter
prints test results to stdout in real-time BEFORE teardown starts.

New approach:
- Capture stdout with tee to a log file
- Poll the log for 'N passed' summary line
- After killing the hanging process, check the log:
  if 'N passed' exists and no 'N failed' or 'N timed out', mark success

This lets us correctly report passing tests even when the agent-server
process doesn't respond to SIGTERM during webServer teardown.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: write PW output to file directly, add debug logging

Process substitution >(tee ...) is fragile with background kills —
the tee process might be killed alongside npm, leaving an incomplete
log. Write directly to file and tail separately for CI output.

Added explicit debug logging:
- 'Checking PW_LOG for pass/fail...'
- grep output showing what matched
- Different message for failure vs success

Co-authored-by: openhands <openhands@all-hands.dev>

* debug: add verbose logging to post-kill check

Need to see: does PW_LOG exist? What size is it? What does grep find?
Which branch of the if/else is taken? Where exactly does it stop?

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use marker file written by test to detect pass/fail

Playwright's JSON reporter only flushes on clean process exit, and
stdout redirection is unreliable with backgrounded process trees.
Instead, the test itself writes a .all-passed marker file after all
assertions succeed. The CI wrapper polls for this file to detect test
completion, then safely kills the hanging teardown process.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: clean up debug diagnostics from helpers

Remove per-event JSON dumps and verbose diagnostic logging from
waitForSuccessfulBashObservation. Keep concise error messages.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: detect test completion immediately via custom reporter

Playwright's reporter onEnd() fires AFTER all tests complete but
BEFORE webServer teardown starts. A custom DoneMarkerReporter writes:

  .tests-done  — always (content: 'passed' or 'failed')
  .all-passed  — only when all tests pass

The CI wrapper polls for .tests-done, so it detects completion
immediately on both pass AND fail. Previously it only polled for
.all-passed, meaning test failures wasted the full 5-min polling
timeout before the step could finish.

This also moves the marker logic out of the test spec and into the
reporter, which is cleaner — the test code doesn't need to know
about CI infrastructure.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: write marker files outside Playwright's outputDir

Playwright clears its outputDir at the start of each run. Writing
markers to a separate .mock-llm-markers/ directory avoids interference.

Also wrapped onEnd() in try/catch and resolved paths via import.meta.url
to be robust against working directory changes.

Co-authored-by: openhands <openhands@all-hands.dev>

* debug: add console.log to reporter, use process.cwd()

import.meta.url may not work in Playwright's CJS reporter context.
Use process.cwd() instead. Add console.log in onBegin/onEnd to verify
the reporter is loaded and executing.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: write markers in onTestEnd, not onEnd

Playwright's lifecycle: onBegin → tests → onTestEnd → onEnd → cleanup.
WebServer teardown happens during 'cleanup', which hangs indefinitely.
onEnd() fires AFTER cleanup, so it never executes when teardown hangs.

onTestEnd() fires immediately after each test completes, before any
cleanup begins. Track total/completed test counts and write markers
after the last test finishes. This gives both pass and fail signals
before the teardown hang blocks everything.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: report script falls back to marker files when results.json missing

Playwright's JSON reporter only flushes results.json on clean process
exit. When the webServer teardown hangs and the process is killed,
results.json never gets written, so the report showed 0/0 tests.

The render script now checks .mock-llm-markers/.tests-done (written by
DoneMarkerReporter in onTestEnd, before teardown) as a fallback. This
gives correct pass/fail status in the PR comment even without
results.json.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: proper teardown and accurate test durations in PR comment

Two fixes:

1. **Teardown hang resolved**: The webServer command now bypasses npm
   and `exec`s directly into `node scripts/dev-safe.mjs`. Previously
   `exec ... npm run dev:minimal` was used, but npm does NOT forward
   SIGTERM to its child processes. When Playwright sent SIGTERM during
   teardown, npm died but node/uvx/vite survived as orphans, causing
   the hang. Now SIGTERM goes straight to dev-safe.mjs's signal handler
   which kills children via process groups and exits cleanly.

2. **Accurate durations**: DoneMarkerReporter now writes a `.results.json`
   with per-test title, status, duration, and error data (from
   `TestResult.duration` in onTestEnd). The report script reads this
   instead of showing 0ms. Falls back to `.tests-done` (pass/fail only)
   if the JSON is missing.

Co-authored-by: openhands <openhands@all-hands.dev>

* chore: reduce teardown grace period from 10s to 5s

The marker-based detection is immediate (onTestEnd fires before
teardown), so we don't need a long grace period. The remaining hang
is Playwright waiting for the multi-process agent-server tree to
fully exit — expected behavior, not a bug.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-28 14:23:28 -04:00
Rohit Malhotraandopenhands e176cdf9aa ci: cache Playwright browsers and pin Node 24.15 across all workflows (#893)
Two fixes applied to ci.yml and snapshot-tests.yml:

1. **Cache Playwright browsers**: Playwright browser downloads (~150 MB)
   are now cached via actions/cache keyed by OS + Playwright version.
   `npx playwright install chromium` only runs on cache miss; system
   deps (`install-deps`) always run since OS packages aren't cacheable.

2. **Pin Node to 24.15.x**: Node 24.16.0 has a zip-extraction regression
   (nodejs/node#63487) that hangs `playwright install` for Playwright
   < 1.60.0. The ci.yml build job and live E2E job, plus
   snapshot-tests.yml, are all pinned consistently.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-28 18:00:41 +00:00
Tim O'Farrellandopenhands d1813c6cfe refactor: update @openhands/extensions from MCP to integrations (#875)
* refactor: update @openhands/extensions from MCP to integrations

Migrate from @openhands/extensions/mcps to @openhands/extensions/integrations
following the upstream rename in OpenHands/extensions.

Key changes:
- MCP_CATALOG -> INTEGRATION_CATALOG
- McpCatalogEntry -> IntegrationCatalogEntry
- MarketplaceTemplate -> IntegrationTransport
- MCP_LOGOS/MCP_FALLBACK_LOGO -> INTEGRATION_LOGOS/INTEGRATION_FALLBACK_LOGO
- entry.template -> entry.connectionOptions[].transport (via getDefaultTemplate helper)
- automation.requiredMcpIds -> automation.requiredIntegrationIds

The new integration catalog structure supports multiple connection options per
entry (e.g., OAuth + stdio fallback). A getDefaultTemplate() helper extracts
the transport config from the default connection option.

Co-authored-by: openhands <openhands@all-hands.dev>

* Set correct version

* fix: resolve lint errors

- Replace Date.now() with useId() for pure render function compliance
- Use optional chaining in handleStdioSubmit
- Remove unused eslint-disable directive
- Fix prettier formatting issues

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: add getInstallableTemplate for stdio-preferred MCP installations

Integration entries like GitHub and Slack now default to OAuth transport,
but the UI doesn't support OAuth yet. Add getInstallableTemplate() that
prefers stdio (API key-based) connection options over OAuth defaults.

- Add getInstallableTemplate() that finds stdio options first
- Update install-server-modal.tsx to use getInstallableTemplate
- Update recommended-automations-*.tsx to use getInstallableTemplate
- Update findCatalogEntryForServer to check ALL connection options

This ensures the install modal shows the correct input fields (e.g.,
GITHUB_PERSONAL_ACCESS_TOKEN) and correctly detects installed servers.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-28 11:21:35 -06:00
665a258b80 feat(acp): inline live model picker for ACP conversations (#769) (#832)
* feat(acp): inline live model picker for ACP conversations (#769)

Converge ACP model selection onto the native LLM-profile inline picker UX
with live mid-conversation switching, replacing the display-only popover.

- Bump @openhands/typescript-client 1.23.3 -> 1.24.0 (adds switchAcpModel).
- AgentServerConversationService.switchAcpModel(conversationId, model): POST
  /switch_acp_model via ConversationClient, with switchProfile's local-only guard.
- useSwitchAcpModel hook: live switch for a running ACP session; for the
  home/no-session case, persist the choice as the agent-settings default
  (agent_settings_diff { acp_model }) so the next conversation inherits it.
- ChatInputModel popover becomes a picker over the provider's available_models
  (check on the effective model), local backend only; cloud / custom-provider /
  native surfaces keep the display + Settings link.
- New i18n key MODEL$AVAILABLE_MODELS.
- Tests for the hook (live vs settings-default branches) and the picker.

Local backend only (matches native switching); custom/unknown providers and any
app_server route remain out of scope per #769.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): open the model picker on click (don't self-close via click-outside)

The inline picker's trigger button sits outside the popover element, so the
document click-outside handler (useClickOutsideElement) treated the opening
click as an "outside" click and closed the popover in the same interaction —
clicking the chip appeared to do nothing. (A programmatic el.click() worked by
fluke: the popover isn't rendered yet when that click bubbles, so the ref is
null and the close is skipped.)

Pass the trigger button as the hook's ignoreOutsideClickRef so a click on the
chip toggles the popover instead of being treated as an outside click.

Validated end-to-end against a local agent-server 1.24.0: the picker opens and
lists the provider's available_models, and selecting one writes the default via
PATCH /settings (home case), with the chip updating to the new model.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(acp): drop disableToast in useSwitchAcpModel so switch errors surface

useSwitchLlmProfile sets meta.disableToast because it's wrapped by
useSwitchLlmProfileAndLog, which re-surfaces errors via its own onError.
useSwitchAcpModel is called directly (no such wrapper / no onError), so
disableToast was silently swallowing failed switches and settings writes
(e.g. a 409 before the first message, network errors, the cloud guard).

Remove it and let the global mutation error toast report failures — simpler
and gives the user feedback when a switch doesn't take.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(acp): share chat input model picker state

* chore: address PR review feedback (#832)

- Add unit test for useChatInputModelState pinning its branching contract,
  incl. the active-ACP getAcpProvider lookup (was home-only in old component).
- Document why the overflow model submenu uses overflow-y-auto (scroll long
  model lists) rather than overflow-visible — no floating children to clip.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: address PR review feedback (#832)

- Wrap the 'Available models' section label in a presentational <li> so it
  is a valid child of the ContextMenu <ul> (was a bare <div>).
- Drop unnecessary 'as never' casts in use-switch-acp-model tests now that
  the real return types (Promise<void>, Promise<boolean>) are honored.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): bump agent-server pin to 1.24.0 for /switch_acp_model

The inline ACP model picker POSTs to /api/conversations/{id}/switch_acp_model,
which is new in openhands-agent-server 1.24.0. The PR description already
lists agent-server:1.24.0 as a dependency, but config/defaults.json was
left at 1.23.1, so local dev (npm run dev) and Docker installs would 404
on every model switch attempt.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(settings): always land on /settings/agent from /settings

The fallback order in ``getFirstAvailablePath`` put ``/settings/llm``
first whenever ``hide_llm_settings`` was off, so clicking Settings sent
the user to the LLM page. For ACP users that page is disabled and
``redirectIfAcpActive`` only catches them when the *personal* settings
already say ``agent_kind === "acp"`` — being in an ACP conversation
with non-ACP personal settings (the common case during the inline
picker flow) bypassed the guard and dumped them on /settings/llm.

Make ``/settings/agent`` the unconditional first fallback. It is
always available (no feature flag hides it), houses the agent-kind
picker, and the left nav still gets OpenHands users to LLM in one
click — so one extra click for non-ACP users buys a much simpler
routing surface and kills the ACP misroute.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(settings/agent): clear command when switching to Custom preset

Selecting "Custom" in the agent preset dropdown reset ``acpModel`` and
flipped ``isCustomAcpModel`` but left ``commandText`` untouched. On the
next render, ``detectPreset(commandText, ACP_PROVIDERS)`` still matched
the previous provider's ``default_command`` and snapped the dropdown
back off "Custom" — the toggle never stayed on Custom.

Clear ``commandText`` in the Custom branch so ``detectPreset`` falls
through to ``ACP_CUSTOM_PRESET_KEY`` on the next render and the dropdown
stays where the user put it. Empty command also matches the intended
"user supplies their own" semantics of the preset.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(settings): mark Verification page as disabledByAcp

The Verification page writes ``confirmation_mode`` and
``security_analyzer`` into ``conversation_settings_diff``. The ACP
agent loop never reads either: ``openhands/sdk/agent/acp_agent.py``
has zero references to ``confirmation_policy`` or
``security_analyzer``, and the only runtime readers
(``openhands/sdk/agent/agent.py:844,855``) live on the native
``Agent`` class — not on ``ACPAgent``. The backend accepts the values
and stores them on conversation state, but the ACP subprocess never
consults them.

So the page presents real-looking knobs that silently do nothing for
ACP users. Mark it ``disabledByAcp: true`` — same pattern as
``/settings/llm`` and ``/settings/condenser`` — so it greys out in the
nav and the existing route guard at ``src/routes/settings.tsx:47-51``
bounces direct visits to ``/settings/agent``.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(ci): bump doc/script SDK version examples to 1.24.0

The docs-version-sync test enforces that every documented agent-server
version example matches ``config/defaults.json:versions.agentServer``.
The previous commit bumped that pin from 1.23.1 to 1.24.0 for the
``/switch_acp_model`` route, but left the example references in
AGENTS.md, ``scripts/dev-safe.mjs``, and ``scripts/check-sdk-version-sync.mjs``
behind — the drift-detector caught it as ``test-and-build`` failure.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(ci): bump remaining hard-coded 1.23.1 to 1.24.0

``__tests__/scripts/dev-safe.test.ts`` asserts ``buildAgentServerCommand``'s
default ``uvx`` args literally include ``openhands-agent-server==1.23.1`` and
matching ``openhands-{sdk,tools,workspace}==1.23.1``. The CI fix in the prior
commit only updated docs and example references; the central pin bump in
``config/defaults.json`` flowed through to this test's runtime expectation but
the literal expectations were never updated. Bump them.

Also bump the ``MOCK_AGENT_SERVER_VERSION`` placeholder in
``src/mocks/settings-handlers.ts`` for consistency with the central pin —
no test asserts on it, but leaving the mock at 1.23.1 invites future
drift confusion.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Debug Agent <debug@example.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-28 16:42:17 +02:00
Tim O'Farrellandopenhands 082cbc2ddb fix: truncate long conversation title on mobile to keep tab toggle visible (#867)
* fix(skills): save disabled_skills when settings has no prior disabled_skills field

The hydration effect in SkillsSettingsScreen only set hasHydratedInitialSettings=true
when settings?.disabled_skills was truthy. When no skills had ever been disabled,
the server omits the field entirely (undefined), so the condition was always false:

  if (settings?.disabled_skills) { ... }   // skipped when field is absent

Because settings?.disabled_skills is undefined both before *and* after settings
load, the dependency array [settings?.disabled_skills] never changed value on
load either, so the effect never ran a second time. hasHydratedInitialSettings
stayed false, and the save effect's early-return guard blocked every toggle.

Fix by gating on settingsLoading instead and defaulting the missing field with
?? []. Also add a localStorage stub for Node.js 25+ which ships a built-in
localStorage that is non-functional without --localstorage-file, breaking the
zustand persist middleware in the test environment.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(tests): use mockResolvedValue(true) for saveSettings spy

saveSettings returns Promise<boolean>, not Promise<void>, so
mockResolvedValue(undefined) caused TS2345 type errors in CI.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(skills): persist disabled_skills to localStorage for local backend

The local agent-server has no concept of disabled_skills and its
PATCH /api/settings strips the field before sending. As a result,
toggling a skill off appeared to work in the UI but was never
persisted -- the value was silently discarded on every save and
page refresh reverted the toggle.

Fix by mirroring the existing app-preferences pattern:

- saveSettings (local path): call writeStoredDisabledSkills before
  stripping the field from the PATCH payload.
- transformApiResponse: call readStoredDisabledSkills and overlay it
  onto the partial Settings returned from the API, so every getSettings
  call (cached or fresh) surfaces the stored value.

Export DISABLED_SKILLS_STORAGE_KEY so tests can reference the key
without duplicating the string literal.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: truncate long conversation title on mobile to keep tab toggle visible

When a conversation title was long enough it would overflow and hide the
RightPanelToggle button (the tab toggle) in the top-right of the mobile
header. Fixed by:

- Adding `min-w-0` to the left-side flex container in
  ConversationNameWithStatus so it can shrink and the toggle stays in view
- Adding `shrink-0` to the status dot wrapper so it never collapses
- Adding `min-w-0` to ConversationName's root div (flex item) to allow
  it to shrink within the parent
- Removing `w-fit max-w-fit` from the title div (both prevented
  `truncate` from ever activating) so the title now properly truncates
  with an ellipsis when too long
- Adding `shrink-0` to the ellipsis-button wrapper so it is always
  accessible regardless of title length

Fixes #848

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-28 07:41:05 -06:00
Hiep Le a00ebe44f3 fix: hide Planner tab on local backends (#861) 2026-05-28 16:57:59 +07:00
b570b4a2f6 fix: buildNpmScriptCommand always uses cmd.exe on Windows (#734)
* fix: buildNpmScriptCommand always uses cmd.exe on Windows

On Windows, npm sets npm_execpath to a path like
  C:\Program Files\nodejs\node_modules\npm\bin\npm-cli.js
which contains spaces. buildNpmScriptCommand was returning that
path as a spawn argument with the full node.exe path as the command.
spawnService uses shell:true on Windows, so Node.js passes the
unquoted command to cmd.exe:

  cmd.exe /d /s /c C:\Program Files\nodejs\node.exe ...

cmd.exe splits on the space and fails with
  'C:\Program' is not recognized as an internal or external command
causing Vite to exit with code 1 immediately after npm run dev.

Fix: check platform === 'win32' BEFORE checking npm_execpath so
Windows always uses the safe cmd.exe /d /s /c npm run <script> form.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: buildNpmScriptCommand always uses cmd.exe on Windows

On Windows, npm sets npm_execpath to a path like
  C:\Program Files\nodejs\node_modules\npm\bin\npm-cli.js
which contains spaces. buildNpmScriptCommand was returning that
path as a spawn argument with the full node.exe path as the command.
spawnService uses shell:true on Windows, so Node.js passes the
unquoted command to cmd.exe:

  cmd.exe /d /s /c C:\Program Files\nodejs\node.exe ...

cmd.exe splits on the space and fails with
  'C:\Program' is not recognized as an internal or external command
causing Vite to exit with code 1 immediately after npm run dev.

Fix: check platform === 'win32' BEFORE checking npm_execpath so
Windows always uses the safe cmd.exe /d /s /c npm run <script> form.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: AnimatePresence mode=wait expects one child, not two

ChatStatusIndicator had two separately-keyed motion.span children inside
<AnimatePresence mode="wait">. framer-motion's wait mode expects exactly ONE
child to exit before the next enters; two children trigger the repeated warning:
  "attempting to animate multiple children within AnimatePresence,
   but its mode is set to 'wait'"

Fix: wrap both elements in a single motion.span with unified key={status}
and className="contents" (CSS display:contents preserves flex layout).

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: set PYTHONUTF8=1 in agent-server/automation env on Windows

Python on Windows defaults to the system ANSI codepage (cp1252). The agent-server
writes metadata JSON containing emoji (e.g. U+2705 ✅) that cp1252 cannot encode,
producing UnicodeEncodeError → POST /api/conversations 500. Old UTF-8 conversation
files also fail to load at startup (UnicodeDecodeError). PYTHONUTF8=1 enables
Python's UTF-8 mode (PEP 540) for the process, matching Linux/macOS behaviour.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: set PYTHONUTF8=1 in agent-server/automation env on Windows

Python on Windows defaults to the system ANSI codepage (cp1252). The agent-server
writes metadata JSON containing emoji (e.g. U+2705 ✅) that cp1252 cannot encode,
producing UnicodeEncodeError → POST /api/conversations 500. Old UTF-8 conversation
files also fail to load at startup (UnicodeDecodeError). PYTHONUTF8=1 enables
Python's UTF-8 mode (PEP 540) for the process, matching Linux/macOS behaviour.

Co-authored-by: openhands <openhands@all-hands.dev>

* refactor: remove unrelated file

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: hieptl <hieptl.developer@gmail.com>
2026-05-28 16:29:55 +07:00
Hiep Le aac29692a2 fix(frontend): stop fetching conversations for closed run-logs modals (#860)
* fix: stop fetching conversations for closed run-logs modals

* refactor: update the code based on feedback
2026-05-28 14:40:25 +07:00
Hiep Le f9c6c99960 fix: restore disabled Change Agent button on home page chat input (#858) 2026-05-28 13:33:46 +07:00
Hiep Le 996a98f375 fix: add missing translations for MCP connection-test keys (#856) 2026-05-28 13:15:18 +07:00
Hiep Le 26e5f45005 fix: serve favicon.svg as the app favicon via root links export (#854) 2026-05-28 12:48:30 +07:00
Hiep Le f2320accc6 fix: hide Start from a proven workflow on cloud backends (#852) 2026-05-28 12:25:12 +07:00
Hiep Le d7d396df1e fix(frontend): scroll chat to bottom when a new prompt is submitted (#831)
* fix: scroll chat to bottom when a new prompt is submitted

* refactor: update the code based on feedback
2026-05-28 12:14:26 +07:00
Tim O'Farrellandopenhands 0206d0e9ee fix(skills): save disabled_skills when server omits the field (#837)
* fix(skills): save disabled_skills when settings has no prior disabled_skills field

The hydration effect in SkillsSettingsScreen only set hasHydratedInitialSettings=true
when settings?.disabled_skills was truthy. When no skills had ever been disabled,
the server omits the field entirely (undefined), so the condition was always false:

  if (settings?.disabled_skills) { ... }   // skipped when field is absent

Because settings?.disabled_skills is undefined both before *and* after settings
load, the dependency array [settings?.disabled_skills] never changed value on
load either, so the effect never ran a second time. hasHydratedInitialSettings
stayed false, and the save effect's early-return guard blocked every toggle.

Fix by gating on settingsLoading instead and defaulting the missing field with
?? []. Also add a localStorage stub for Node.js 25+ which ships a built-in
localStorage that is non-functional without --localstorage-file, breaking the
zustand persist middleware in the test environment.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(tests): use mockResolvedValue(true) for saveSettings spy

saveSettings returns Promise<boolean>, not Promise<void>, so
mockResolvedValue(undefined) caused TS2345 type errors in CI.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(skills): persist disabled_skills to localStorage for local backend

The local agent-server has no concept of disabled_skills and its
PATCH /api/settings strips the field before sending. As a result,
toggling a skill off appeared to work in the UI but was never
persisted -- the value was silently discarded on every save and
page refresh reverted the toggle.

Fix by mirroring the existing app-preferences pattern:

- saveSettings (local path): call writeStoredDisabledSkills before
  stripping the field from the PATCH payload.
- transformApiResponse: call readStoredDisabledSkills and overlay it
  onto the partial Settings returned from the API, so every getSettings
  call (cached or fresh) surfaces the stored value.

Export DISABLED_SKILLS_STORAGE_KEY so tests can reference the key
without duplicating the string literal.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-27 17:49:25 -06:00
Tim O'Farrellandopenhands 463deff24a feat: validate MCP server connectivity before saving (#785) (#816)
* feat: pre-flight MCP test before save (#785)

Validate MCP server connectivity via POST /api/mcp/test before saving
in both the marketplace install modal and the custom-server editor.

- Add McpService.testServer() using MCPClient from @openhands/typescript-client
- Add useTestMcpServer() useMutation hook
- InstallServerModal: test before addMcpServer; show inline error on failure,
  keep modal open (onClose not called); button shows Verifying… then Saving…
- CustomServerEditor: same pre-flight pattern before add/update; also exposes
  a standalone Test Connection button via MCPServerForm's new onTest prop
- MCPServerForm: add onTest/isTestPending/testMessage props; extract buildConfig
  helper; add handleTestClick using formRef; render Test Connection button and
  testMessage inline above action buttons
- Add i18n keys: MCP$TEST_BUTTON, MCP$VERIFYING, MCP$TEST_SUCCESS,
  MCP$TEST_ERROR_TIMEOUT, MCP$TES  MCP$TEST_ERROR_TIMEOUT, MCP$TES  MCP$TEST_ERROR_TIMEOUT, MCP$TES  MCP$TEST_ERROR_TIMEOUT, MCP$TES  MCP$TEST_ERROR_TIMEOUT, MCP$TES  + keeps modal open, success path, Verifying label)

Closes #785

Co-authored-by: openhands <openhands@all-hands.dev>

* test: stub McpService.testServer in tests that save through the pre-flight

Three test suites click a submit button whose handler now runs the
pre-flight connectivity test before calling saveSettings.  None of
them mocked McpService.testServer, so the mutation errored out before
reaching the save step.

Add vi.spyOn(McpService, 'testServer').mockResolvedValue({ ok: true, tools: [] })
to the beforeEach of each affected suite so the test-then-save flow
resolves as expected and the existing assertions remain valid.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: decode HTML entities in MCP test-server error messages

The Python backend passes error strings through html.escape(), so
apostrophes arrive as &#39; and slashes as &#x2F;.  Add a one-liner
decodeHtmlEntities() helper in McpService that uses a temporary
<textarea> to let the browser's own HTML parser decode the string
before it reaches any UI component.

Decoding happens once, at the API boundary, so every consumer
(InstallServerModal, CustomServerEditor, etc.) automatically gets
clean text without needing its own unescaping logic.

Add a focused McpService unit test (4 cases) that mocks MCPClient
via vi.hoisted + vi.mock to exercise the real decoding path:
- success responses pass through unchanged
- &#39; / &#x2F; entities decoded in the error field
- numeric (&lt;) and hex (&#x73;) entities decoded
- stdio config mapped to the correct MCPServerSpec shape

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: render MCP error strings as plain text with newlines

Two bugs in the error message display path:

1. HTML-entity encoding: react-i18next escapes interpolated values by
   default (\' -> &#39;, / -> &#x2F;).  Fix by using the no-escape
   prefix {{-error}} in the MCP$TEST_ERROR_UNKNOWN translation string
   so the raw error text is interpolated verbatim.

2. Newlines ignored: the \n in multi-line Python tracebacks was
   swallowed by the browser.  Fix by adding whitespace-pre-wrap to the
   <p> elements in install-server-modal (globalError) and
   mcp-server-form (testMessage) so \n renders as a visual line break.

Also revert the previous decodeHtmlEntities approach from the service
layer — the backend does not HTML-escape thlayer — the backend does nwas introduced by i18next, not the server.  Update McpService telayer — the backend does not HTML-escape thlayer — the backend dngs in the
translation layer instead.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-27 12:56:25 -06:00
Rohit Malhotraandopenhands 5f5311272f fix: use backend registry apiKey for automation auth instead of build-time env var (#830)
The localAutomationAxios interceptor was reading the session API key from
import.meta.env.VITE_SESSION_API_KEY (baked in at build/publish time),
causing 401 errors for users of the published npm package because their
runtime session key (injected into localStorage by static-server.mjs) was
never included in requests to the automation backend.

The fix reads the API key from getEffectiveLocalBackend().apiKey on every
request, which dynamically resolves the current backend registry entry —
the same source the host/baseURL was already using. This ensures:
- Published npm package users get the runtime-injected key from localStorage
- Users who edit their backend via the Manage Backends UI get their updated key
- No more falling back to Keycloak cookie auth (which always returns 401)

Fixes #829

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-27 14:31:31 -04:00
Hiep Le a3feade99f fix: keep Activity Log in sync while a run is Pending/Running (#824) 2026-05-28 00:42:00 +07:00
Rohit Malhotraandopenhands b6380f6e55 chore: bump version to 1.0.0-alpha.7 (#822)
Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-27 16:53:16 +00:00
Hiep Le 376569eb01 fix: show local time on Pending run rows instead of Jan 1, 1970 (#812) 2026-05-27 23:26:13 +07:00
Rohit Malhotraandopenhands f2d19c8923 Add GitHub bug report issue template (#813)
* Add GitHub bug report issue template

- Bug report form with install method dropdown (npm, Docker, source, other)
  and version dropdown listing all pre-release versions (alpha.2–alpha.6)
- Includes optional agent-server version, environment, logs/screenshots fields

Co-authored-by: openhands <openhands@all-hands.dev>

* Remove agent server version field from bug report template

Co-authored-by: openhands <openhands@all-hands.dev>

* Remove environment field from bug report template

Co-authored-by: openhands <openhands@all-hands.dev>

* Add OS dropdown to bug report template

Co-authored-by: openhands <openhands@all-hands.dev>

* Address review feedback on bug report template

- Convert version dropdown to free-text input to avoid maintenance burden
- Fix docker inspect description to include image reference
- Add Actual Behavior field between Steps to Reproduce and Expected Behavior
- Split Logs/Screenshots into separate fields so render:shell doesn't break images

Co-authored-by: openhands <openhands@all-hands.dev>

* Add --version flag to CLI and version label to Docker image

- bin/agent-canvas.mjs: add -v/--version flag that reads version from package.json
- docker/Dockerfile: add AGENT_CANVAS_VERSION build arg and org.opencontainers.image.version label
- .github/workflows/docker.yml: extract version from package.json, pass as build arg
- bug_report.yml: update version field description with the actual commands users can run

Co-authored-by: openhands <openhands@all-hands.dev>

* Update version description: use image tag for Docker (label not yet released)

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-27 12:17:35 -04:00
Xingyao Wangandopenhands 4655f7a119 chore: remove PR review GitHub Actions workflow (#814)
Remove `pr-review-by-openhands.yml`, `pr-review-evaluation.yml` — PR review is now handled
via OpenHands Cloud automation.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-27 11:47:46 -04:00
Rohit Malhotraandopenhands 857f3384b8 docs: add Docker and NPM install instructions to README (#806)
* docs: add Docker and NPM install instructions to README

Add three clearly labeled installation options to the Quickstart section:
- Option 1: Docker (pull and run the published image)
- Option 2: NPM (global install of the published package)
- Option 3: From Source (existing clone-and-build workflow)

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use 1.0.0-alpha.6 docker tag instead of latest

Co-authored-by: openhands <openhands@all-hands.dev>

* docs: add explicit export PROJECTS_PATH step before commands

Addresses review feedback: move PROJECTS_PATH setup into an explicit
export step above the docker run / agent-canvas commands so copy-pasting
doesn't produce a broken volume mount.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-27 11:34:27 -04:00
cbeeee002e fix: inject runtime session key into index.html for published binary (#795)
* fix: inject runtime session key into index.html for published binary

The globally installed agent-canvas binary starts the agent-server with a
persisted session API key (~/.openhands/agent-canvas/session-api-key.txt)
as OH_SESSION_API_KEYS_0, making auth required. However, the pre-built
static frontend in the npm package has a different (or empty)
VITE_SESSION_API_KEY baked in at publish time, so every API request gets
401 Unauthorized.

Fix: static-server.mjs now accepts --session-api-key <key> and injects a
tiny bootstrap <script> before </head> in every index.html response. The
script seeds the key into localStorage['openhands-agent-server-config']
only if no key is already stored there, so explicit user overrides (via
Settings > Agent Server) are always preserved.

dev-with-automation.mjs and dev-static.mjs both pass
--session-api-key ${config.sessionApiKey} when spawning the static server,
so the runtime key is always available regardless of what was baked into
the bundle.

Tests: added 8 new cases to __tests__/scripts/static-server.test.ts
covering parseArgs, injection in direct and SPA-fallback index.html
responses, no injection for non-html assets, cache headers, and null key.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(docker): pass runtime session key to static-server so frontend can authenticate

The entrypoint computed EFFECTIVE_SESSION_KEY and forwarded it to the
agent-server (OH_SESSION_API_KEYS_0) and automation backends, but did not
pass it to the static-server. As a result the pre-built index.html served
with no session key injected, so every browser API call received 401.

Wire --session-api-key "$EFFECTIVE_SESSION_KEY" into the static-server
launch command so the runtime key is injected into index.html responses
(via the mechanism added in this branch to static-server.mjs).

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: address review suggestions on session key injection

- static-server.mjs: add comment clarifying replace() targets first
  </head> only; fall back to inserting before </body> when </head> is
  absent (avoids prepending before <!DOCTYPE html>)
- docker/entrypoint.sh: add comment documenting source of
  EFFECTIVE_SESSION_KEY before the static-server invocation
- static-server.test.ts: add test for </head>-absent fallback path
  confirming injection lands before </body>

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: Rohit Malhotra <rohitvinodmalhotra@gmail.com>
2026-05-27 11:12:23 -04:00
Hiep Le c34aa03273 fix: refetch list on revisit so agent-created automations appear (#805) 2026-05-27 21:27:35 +07:00
Tim O'Farrellandopenhands 99108509f6 perf(useLocalGitInfo): consolidate git probe into single bash round-trip (#792)
* perf(useLocalGitInfo): consolidate git probe into single bash round-trip

Replace the former probeGitInfoAtDir + probeNestedRepoInDir pair (which
issued 2-3 serial bash WebSocket round-trips per poll cycle) with a single
GIT_INFO_COMMAND shell script that handles both the root and nested-repo
cases in one run() call.

The script outputs two newline-separated lines: <remote-url>\n<branch>.
The new probeGitInfo function runs the script and splits on the first
newline to recover the same LocalGitInfo struct as before.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: prettier quote style in use-local-git-info.ts

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-27 07:16:13 -06:00
Hiep Le 15356ebbf9 fix(frontend): render Settings icon as a link so it supports Open in new tab (#802)
* fix: render Settings icon as a link so it supports Open in new tab

* fix: failing tests
2026-05-27 18:35:20 +07:00
Hiep Le d2a021e5d3 fix: disable Run Now when the automation is turned off (#800) 2026-05-27 17:10:36 +07:00
Vasco Schiavo b45c4e27cc fix(skills): prefer frontmatter description over body content in slash command menu (#673) 2026-05-27 09:35:35 +00:00
Hiep Le 41e5281509 fix: hide Change Agent button on home page chat input (#798) 2026-05-27 16:23:42 +07:00
Vasco Schiavo 9eaa1f7519 fix: route LLM errors inline and keep conversation/server errors visible in the banner (#787) 2026-05-27 06:55:48 +00:00
Rohit Malhotraandopenhands 907d6bde76 chore: bump version to 1.0.0-alpha.6 (#793)
Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-26 18:30:16 -04:00
Rohit Malhotraandopenhands e2dd1b5f17 fix: unify session and automation API keys into a single credential with consistent header (#681)
* fix: unify session and automation API keys into a single credential

Both the agent-server and automation backend now share the same API key
value. The agent-server validates it via `X-Session-API-Key` and the
automation backend validates it via `Authorization: Bearer …` — different
header formats, same credential.

Changes:
- Frontend: automation axios client reads `VITE_SESSION_API_KEY` instead
  of the now-removed `VITE_AUTOMATION_API_KEY`
- Dev launcher: removed separate `AUTOMATION_LOCAL_API_KEY` generation
  and persistence (`automation-api-key.txt`); `localApiKey` is set to
  `sessionApiKey` so both backends receive the same value
- Static build: stopped baking `VITE_AUTOMATION_API_KEY` (the frontend
  reads from `VITE_SESSION_API_KEY`)
- Docker entrypoint: `OPENHANDS_AUTOMATION_API_KEY`,
  `AUTOMATION_LOCAL_API_KEY`, and `AUTOMATION_AGENT_SERVER_API_KEY` all
  default to the session key when not explicitly overridden
- Tests updated to verify unified key behavior

Fixes the 401 on `/api/automation/v1` when the automation backend is
running but no separate `VITE_AUTOMATION_API_KEY` was configured.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: use X-Session-API-Key header for automation backend auth (consistent with agent-server)

Switch automation backend requests from `Authorization: Bearer …` to
`X-Session-API-Key` header, matching the agent-server's auth pattern.
Both backends now authenticate using the same header and the same key
value (`VITE_SESSION_API_KEY`).

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: address review — remove localApiKey alias, dead constant, add entrypoint guard

- Remove `localApiKey` from config; all call sites now use
  `config.sessionApiKey` directly, making the unified-key intent obvious.
- Delete `DEFAULT_AUTOMATION_API_KEY_PATH` constant and its export
  (no downstream consumers in beta).
- Add fail-fast guard in docker/entrypoint.sh when no session key is
  available, instead of silently exporting empty strings.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: update stale comment on AUTOMATION_LOCAL_API_KEY to reflect unified session key

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-26 22:15:27 +00:00
Rohit Malhotraandopenhands cb831b2860 fix: include config/ in npm package files (#791)
The `config/` directory (containing `defaults.json`) was missing from the
`files` allowlist in package.json, so it was excluded from the published
npm tarball. The CLI entry point imports `scripts/dev-with-automation.mjs`
which reads `config/defaults.json` at startup, causing an ENOENT crash.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-26 22:02:57 +00:00
Rohit Malhotraandopenhands 603ed9d66d chore: bump version to 1.0.0-alpha.5 (#789)
Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-26 17:36:18 -04:00
Tim O'Farrellandopenhands 9c1ae19878 fix: auto-clear backend health on successful conversation list fetch (#647)
* feat: execute git-info probe commands over bash-events WebSocket

Replace per-poll REST calls in useLocalGitInfo with a persistent
WebSocket connection to /sockets/bash-events.

Previously, useQuery's refetchInterval:10_000 caused
AgentServerRuntimeService.executeCommand to fire individual HTTP
requests on every tick (find + 2 git commands × up to 3 candidate
dirs) to probe the workspace for git metadata.

Changes:
- websocket-url.ts: export buildBashWebSocketUrl() using the same
  host/pathPrefix extraction as the conversation-events URL builder
- use-bash-command-runner.ts: new hook that maintains a persistent WS
  connection to /sockets/bash-events and exposes runCommand(command,
  cwd, timeout) → Promise. Commands queued during CONNECTING are
  flushed on open; all in-flight commands are rejected on
  close/error/unmount.
- use-local-git-info.ts: swap AgentServerRuntimeService.executeCommand
  for useBashCommandRunner; keep refetchInterval:10_000 so branch
  changes (e.g. git checkout) are still reflected without a full
  refresh, but each probe now reuses the open socket rather than
  opening new HTTP connections.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix: auto-clear backend health on successful conversation list fetch

When the backend health probe fails 5 times in a row, the disabled flag
is written to localStorage and background polling stops permanently.
This left the status dot red even after the backend recovered, because
no successful API call ever cleared the disabled state.

Fix by wiring a global QueryCache onSuccess handler that calls
recordBackendSuccess(backendId) whenever a query tagged with
meta.backendId succeeds. Tag usePaginatedConversations (which polls
every 10s) with the active backend's id so a recovered backend clears
its stale failure state within one polling cycle automatically.

Any future query that targets a specific backend can opt in by adding
meta: { backendId } to its options.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-26 13:26:59 -06:00
Tim O'Farrellandopenhands eb7c983169 feat: bump @openhands/extensions to include github-repo-monitor and slack-channel-monitor (#781)
Pins @openhands/extensions from 7b33f64 (May 19) to b8c1869 (May 23),
which adds two new automation catalog entries:
- github-repo-monitor: watches GitHub repos for @OpenHands mentions
- slack-channel-monitor: watches Slack channels for @openhands mentions

Updates the popularity-order unit test to reflect the new 7-entry catalog
ranking (github-repo-monitor at rank 98 sits between github-pr-reviewer
at 100 and slack-standup-digest at 94).

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-26 12:57:54 -06:00
Rohit Malhotraandopenhands 45da5606d6 fix: bump agent-server SDK to 1.23.1 (#782)
Fixes two bugs in the upstream agent-server SDK v1.23.0 that caused
500 Internal Server Error on POST /api/conversations:

1. LLM registry duplicate usage_id: The condenser and main agent LLMs
   both used usage_id='default', causing ValueError on conversation
   creation. v1.23.1 checks for existing usage IDs before registering.

2. Validation error handler crash: The _validation_exception_handler
   tried to JSON-serialize raw ValueError objects from Pydantic
   validation contexts, turning 422 errors into 500s. v1.23.1 properly
   sanitizes validation errors before serializing.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-26 18:42:17 +00:00
Rohit Malhotraandopenhands 489070029d fix(ci): publish directly with --tag latest to avoid OIDC dist-tag failure (#778)
* fix(ci): publish directly with --tag latest to avoid OIDC dist-tag failure

The previous workflow published with --tag alpha then ran a separate
npm dist-tag add to set latest. The second call failed with E401
because OIDC trusted publishing tokens don't cover post-publish
registry mutations like dist-tag.

Simplify to a single npm publish --tag latest, which is all we need
until the first stable release (#395).

* docs(ci): note that named prerelease dist-tags are removed under current policy

Add a comment block explaining that alpha/beta/rc dist-tags are intentionally
not published while issue #395's 'everything is latest' policy is active, so
consumers pinning to named prerelease tags are not silently broken without
notice.

Co-authored-by: openhands <openhands@all-hands.dev>

* docs(ci): add OIDC root cause to publish comment per review

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-26 18:16:14 +00:00
Rohit Malhotraandopenhands 5d7f539303 fix(docker): ensure latest tag always points to latest stable release (#780)
Previously, the merge-manifests job unconditionally aliased main→latest
on every push to the main branch. This meant any commit to main after a
release would overwrite the `latest` multi-arch manifest with unreleased
main-branch code instead of the most recent stable tagged version.

The per-arch build already adds `latest-{arch}` tags exclusively for
stable (non-pre-release) version tags (e.g. v1.2.3), and the manifest
merge loop correctly creates the `latest` multi-arch manifest from
those arch-suffixed images. The extra main→latest alias was redundant
for tag pushes and incorrect for plain main pushes.

Remove the main→latest alias so `latest` is only produced by stable
version tag pushes, ensuring `docker pull …:latest` always gets the
highest stable release.

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-26 18:05:00 +00:00
Rohit Malhotraandopenhands 43c10810da fix: switch @openhands/typescript-client from git dep to npm registry (#779)
* fix: switch @openhands/typescript-client from git dep to npm registry

Replace the git+https dependency with the published npm package
(v1.23.3). Git dependencies break `npm install -g` because npm
clones the repo and runs the prepare script, but devDependencies
like rimraf aren't available during global installs.

Co-authored-by: openhands <openhands@all-hands.dev>

* test: add guard against git dependencies in package.json

Git dependencies break `npm install -g` because npm clones the repo
and runs the prepare script without devDependencies. Add a test that
fails if any dependency uses a git URL, with an allowlist for
@openhands/extensions (not yet published to npm).

Co-authored-by: openhands <openhands@all-hands.dev>

* test: also catch bare owner/repo GitHub shorthand in git dep guard

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-26 17:51:37 +00:00
cc63f48a07 refactor(acp): source ACP model lists from typescript-client (closes #740) (#775)
* refactor(acp): source model lists from typescript-client registry

Replace the hand-mined CLAUDE_MODELS / CODEX_MODELS / GEMINI_MODELS lists (and
the duplicated provider metadata) with the @openhands/typescript-client ACP
registry, which mirrors the Python SDK source of truth
(openhands.sdk.settings.acp_providers). acp-providers.ts becomes a thin
adapter: it enriches each upstream record with Canvas-only UI fields (brand
icon + onboarding description) and keeps the helper functions + public export
surface unchanged, so no consumers change.

- Bump the @openhands/typescript-client pin to the #187 merge commit
  (082d4d46), which adds available_models / default_model to the registry.
- Delete the three hardcoded model lists; build ACP_PROVIDERS from
  getAcpProvider() + a small ACP_PROVIDER_UI map.
- Incidentally corrects the Gemini default to auto-gemini-2.5 (the CLI's
  auto-router default), matching the merged SDK/client.

Closes #740.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(acp): pin typescript-client to v1.23.2 tag (was SHA)

Now that typescript-client v1.23.2 is tagged/released (includes #187's ACP
registry, mirroring SDK #3389), pin to the tag instead of the raw #187 merge
SHA. v1.23.2 tracks the SDK's v1.23.2 patch line. Resolves to the same commit
as the prior SHA, so no resolved-content change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci(acp): remove obsolete ACP providers sync check

The acp-providers-sync workflow + scripts/check-acp-providers-sync.mjs existed
to keep Canvas's hand-kept ACP registry mirror in sync with the SDK source
(agent-canvas#587). That mirror is gone — acp-providers.ts now sources its
model data from @openhands/typescript-client, which carries its own
SDK-drift check (check-acp-drift.py). So this canvas-side check is redundant
and was failing on the refactored ACP_PROVIDERS (no longer a literal array).

- Delete .github/workflows/acp-providers-sync.yml + the script.
- Drop the docs-version-sync test case that asserted the script's example.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Debug Agent <debug@example.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 19:25:08 +02:00
Rohit Malhotraandopenhands 3e0d915526 chore: bump version to 1.0.0-alpha.4 (#771)
Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-26 16:45:28 +00:00
Rohit Malhotraandopenhands 89abe93a28 chore: bump automation version to 1.0.0a5 (#774)
Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-26 16:27:21 +00:00
3d724dd8fd feat(acp): versioned ACP model picker + per-conversation agent chip (#730)
* feat(acp): save concrete model defaults in canvas

* test(acp): verify provider models from subprocesses

* feat(acp): per-conversation agent chip + versioned Claude model picker

User-visible changes:

- Conversation cards now show a single inline chip `[brand mark] {model}`
  on every conversation. ACP cards always render (identity info); OpenHands
  native cards render whenever `agent.llm.model` is present. New
  `AgentBrandIcon` covers Claude / Codex / Gemini brand marks with a
  terminal-glyph fallback; the OpenHands logo is recolored via
  `[&_path]:fill-current` so it inherits the muted-grey chip color
  (the shipped SVG hardcodes `fill="white"`).

- Settings → Agent dropdown for Claude Code lists 10 versioned options
  (Opus 4.7, 4.6, 4.6/1M, 4.5, 4.1; Sonnet 4.6, 4.6/1M, 4.5; Haiku 4.5;
  opusplan). Default switched from `sonnet` to `claude-opus-4-7`.
  Canonical IDs were verified against the bundled `claude` CLI binary's
  model registry — the static SDK `.mjs` shims don't carry the full
  registry. `[1m]` aliases used for 1M-context variants (no canonical
  `claude-*-1m` IDs ship in the SDK).

- `ConversationCardFooter` chip resolves the ACP model string from
  `current_model_name → current_model_id → agent.acp_model → agent.llm.model`
  (the `"acp-managed"` sentinel is filtered out), so the chip works
  whether or not the agent-server populates the SDK runtime fields.

- Removed dead `showLlmProfiles` plumbing on `ConversationCard` /
  `CompactConversationRow` / `ConversationPanel` pass-throughs that
  used to gate the OpenHands model line behind a metadata-menu toggle.
  Chip is now always shown when a model exists.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(acp): adapter cleanup + drop unauthenticated model-check script

- Extract `resolveAcpDisplayModel(info)` from the triple-nested ternary
  in `toAppConversation`. Same precedence chain (runtime name → runtime
  id → configured model → llm.model minus sentinel), just expressed once
  with a loop and named conditions.

- Name the `"acp-managed"` literal as `ACP_MANAGED_SENTINEL` (exported)
  and the `"default" / "default (recommended)"` strings as
  `ACP_DEFAULT_PLACEHOLDERS`. The sentinel was referenced from 7 places
  across source and tests.

- Trim `DirectConversationInfo.agent.kind` JSDoc to 3 lines (was 7).

- Delete `scripts/check-acp-provider-models.mjs` and revert the CI step +
  triggers in `.github/workflows/acp-providers-sync.yml`. The script
  spawned each ACP wrapper unauthenticated and validated Canvas's lists
  against the wrapper's fallback model set — which is only 3 models for
  Claude Code (sonnet / sonnet[1m] / haiku) regardless of what the
  wrapper actually accepts. The check flagged every legitimately-added
  Opus entry as drift. Removing it; followup tracked in #740 for moving
  the lists into `@openhands/typescript-client`.

- Remove `test:acp-models` npm script (only invoked the deleted file).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): plug review-flagged gaps in model resolution + chip path

Five behavioral fixes + two refactors from review feedback on PR #730:

- F1: ``normalizeAgent`` and ``requireDirectConversationInfo`` were
  dropping ``agent.acp_model``, ``current_model_id``, and
  ``current_model_name`` on the way in from the wire — so the chip
  resolver only worked in unit tests (which build ``DirectConversationInfo``
  in-process) and silently fell back to provider-name labels in production.
  Normalizer now preserves all three.

- F2: ChatInputActions gated rendering ChatInputModel on
  ``isCloud || conversation?.agent_kind === "acp"``, so on a local home
  screen with no active conversation it always picked SwitchProfileButton.
  SwitchProfileButton then hid itself for ACP — net result: the home-ACP
  model label never appeared. Added ``isHomeAcp`` derivation from
  ``settings.agent_settings.agent_kind`` so the chat input picks
  ChatInputModel in that case too.

- F3: Switching the preset dropdown to Custom set ``isCustomAcpModel``
  but left ``acpModel`` untouched, so a user moving Claude Code → Custom
  + typing a custom command could save ``acp_model: "claude-opus-4-7"``
  on an unrelated wrapper. Now clears ``acpModel`` on Custom selection.

- F4: ``buildConfiguredAcpAgentSettings`` was stripping null / empty
  ``acp_model`` and not falling back to ``provider.default_model``, so
  existing users with ``acp_model: null`` saw the new registered default
  in Settings → Agent but their next conversation still started with the
  agent-server's own default (UI/runtime mismatch). Conversation creation
  now substitutes the provider default for empty values, matching what
  the form displays.

- F5: ``isAcpDefaultPlaceholder`` was applied only to runtime fields, not
  to the ``configured`` (``agent.acp_model``) or ``sdkLlm``
  (``agent.llm.model``) fallback rungs of the precedence chain. Older
  settings that persisted the literal ``"Default (recommended)"`` could
  surface it on chips. Filter now applies uniformly.

- R1: All five surfaces (Settings form, conversation creation,
  ChatInputModel, conversation adapter, chip) now route through one
  helper, ``resolveEffectiveAcpModel({ runtimeName, runtimeId,
  configured, sdkLlm, providerDefault })`` in ``acp-providers.ts``.
  Placeholder + sentinel filtering live in one place. ``providerDefault``
  is opt-in — chip omits it (don't lie about what's running), settings
  / creation / chat-input pass it (silently substitute the registry
  default).

- R2: ``ACPProviderIcon`` no longer includes ``"openhands"`` — it's
  ACP-only again. Reintroduced ``AgentBrandIconKind = "openhands" |
  ACPProviderIcon`` for surfaces that can render either harness's mark.

Tests: ``acp-server-conversation-service.test.ts`` now exercises the wire
path with the new fields. ``agent-server-adapter.test.ts`` adds two cases
covering placeholder filtering on configured + sdkLlm rungs.
``agent-settings.test.tsx`` adds a Custom-preset-clears-default
regression. ``chat-input-model.test.tsx`` swaps the
"home + ACP renders nothing" assertion for "home + ACP shows the provider
default" — that's the new correct behavior.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): translate new agent-settings keys + harmonize Gemini labels

Addresses the latest review pass:

- Translates ``SETTINGS$AGENT_CUSTOM_MODEL`` and ``SETTINGS$AGENT_MODEL_HINT``
  into all 14 non-English locales (ja, zh-CN, zh-TW, ko-KR, no, it, pt, es,
  ar, fr, tr, de, uk, ca). Both keys previously fell back to English text
  on every non-English client, blocking the model dropdown's hint copy and
  Custom-model label from being legible.

- Harmonizes the Gemini model labels with the Claude / Codex pattern —
  raw IDs like ``gemini-3.1-pro-preview`` become ``Gemini 3.1 Pro
  (preview)`` so the three providers read consistently in the dropdown.

- Adds provenance notes to ``CODEX_MODELS`` and ``GEMINI_MODELS`` mirroring
  the Claude one — naming the upstream source the list was extracted from
  and pointing at agent-canvas#740 (the long-term "move ACP model lists to
  ``@openhands/typescript-client``" plan).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): address review — chip labels, dead toggle, icon de-dup, shared hook

- Chip shows the provider's picker label (e.g. "Claude Opus 4.7") instead of
  the raw acp_model ID via new labelForAcpModel(); falls back to the raw ID
  for custom overrides and the provider name when no model. (#1)
- Drop the stale version from the 1M labels (opus[1m]/sonnet[1m] →
  "Claude Opus (1M)" / "Claude Sonnet (1M)") so the version-agnostic alias
  and its label can't disagree. (#2)
- Remove the now no-op "Show LLM profiles" filter-menu row + its wiring
  (chip is unconditional now; the preference no longer affects cards). (#4)
- Onboarding AgentOptionIcon delegates to AgentBrandIcon; delete the
  duplicated brand-mark path constants and per-kind SVG markup. (#5)
- Name the OpenHands logo aspect ratio (3:2) so the chip and onboarding
  tile render identically. (#7)
- Extract useAcpModelContext() shared by chat-input-actions and
  chat-input-model to kill the duplicated isHomeAcp / destination-path /
  label logic. (#6)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): gate the agent chip behind the restored "LLM model" toggle

The chip was unconditional, which (a) silently flipped OpenHands cards from
main's model-hidden-by-default to always-shown and (b) left the panel's
"LLM model" toggle a no-op (then deleted). Restore one toggle that gates the
unified chip for BOTH ACP and OpenHands cards (default OFF), matching the
sibling "Repo and branch" metadata toggle.

- Footer: chip (ACP brand mark + model, OpenHands logo + model) now renders
  only when showAgentChip is set.
- Restore showLlmProfiles plumbing: filter-menu row + ConversationPanel/
  ConversationCard/CompactConversationRow pass-throughs (panel + filter-menu
  are now byte-identical to main).
- ConversationCard.shouldRenderFooter gates ACP under the toggle too.

Net effect vs main for OpenHands: identical by default (hidden); when the
toggle is on, the only change is the OpenHands logo added next to the model
name. ACP identity chip becomes opt-in via the same toggle.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): reconcile model on command edit, unify chip labels, dedupe provider lookup

Review follow-ups for the versioned ACP model picker + agent chip:

- agent-settings: editing the command textarea into a different provider
  (or a custom command) now reconciles the model selector, so Save can no
  longer silently persist e.g. claude-opus-4-7 against a Codex/custom
  wrapper. The preset dropdown already did this; the textarea is the other
  way a user switches providers. Also always show the custom-model input
  when the dropdown is on "Custom" so a mismatched value is visible/editable
  rather than hidden.
- chat input: surface the provider's human label (e.g. "Claude Opus 4.7")
  for ACP conversations, matching the conversation-list chip, instead of the
  raw acp_model id.
- relabel the conversation-panel "LLM model" toggle to "Agent / model" — it
  now governs the ACP brand chip too.
- extract getAcpProvider() and replace the repeated ACP_PROVIDERS.find()
  lookups across the adapter, settings, and the constants module.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(acp): stop filling the OpenHands logo's transparent hand paths

AgentBrandIcon recolored the OpenHands logo with a blanket
``[&_path]:fill-current``, which also overrode the two ``fill="transparent"``
hand shapes — turning the mark into a solid white blob on the onboarding tile
(and a filled blob on the conversation chip). The logo asset has 5
``fill="white"`` wordmark paths and 2 ``fill="transparent"`` hands; only the
former should inherit ``currentColor``.

Scope the override to ``[&_path:not([fill=transparent])]:fill-current`` so the
hands stay transparent (negative space), restoring the original two-tone logo.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Debug Agent <debug@example.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 17:54:54 +02:00
Tim O'Farrellandopenhands c9616c1153 fix: pin openhands-sdk to same version as other SDK packages (#772)
Add openhands-sdk==${VERSION} to the --with list in both PyPI
resolution branches (specific version and default) of
buildAgentServerCommand so all four packages are pinned to the
same versions.agentServer:

  openhands-agent-server, openhands-sdk, openhands-tools, openhands-workspace

Without this pin, openhands-sdk floated to the latest release on
PyPI (a transitive dep with no version bound in the published
openhands-agent-server metadata), causing non-reproducible builds
and version skew between the banner and defaults.json.

The local-path (editable) and git-ref branches are unaffected —
they already source all packages from the same checkout/ref.

Update the two affected unit-test cases to expect the new pinned
openhands-sdk entry in the args array.

Fixes #767

Co-authored-by: openhands <openhands@all-hands.dev>
2026-05-26 15:38:53 +00:00
Neha PrasadandTim O'Farrell f4d45e46fe fix(chat): enable git PR/push/pull on local without provider_tokens_set (#722)
Co-authored-by: Tim O'Farrell <tofarr@gmail.com>
2026-05-26 08:30:38 -06:00
Hiep Le 24a39212a8 fix(frontend): keep composer pinned at the bottom on narrow viewports (#765)
* fix: keep composer pinned at the bottom on narrow viewports

* trigger build
2026-05-26 20:54:15 +07:00