* Update extensions to 0.9
* test: account for jira-issue-to-pr automation in extensions 0.9.0
@openhands/extensions 0.9.0 adds the jira-issue-to-pr recommended
automation (popularityRank 85, beta). Update the recommended-automations
tests to include it in the popularity-ordered list and bump the Beta
section count from 4 to 5.
---------
Co-authored-by: openhands <openhands@all-hands.dev>
* chore: bump typescript-client 1.32.1 + agent-server/openhands-sdk 1.33.0
- @openhands/typescript-client 1.32.0 -> 1.32.1 (package.json + lockfile)
- agent-server SDK libs (openhands-sdk/tools/workspace/agent-server) 1.32.0 -> 1.33.0 via versions.agentServer in config/defaults.json
- sweep doc/test/comment references to 1.33.0 (AGENTS.md, dev-safe.test.ts, check-sdk-version-sync.mjs, dev-safe.mjs, mock-llm-e2e.yml)
- keep the acp <0.11 transitive pin: openhands-sdk 1.33.0 still requires agent-client-protocol>=0.10.1 with no upper bound
Note: check-sdk-version-sync fails until an openhands-automation release pins openhands-sdk 1.33.0. The latest automation (1.1.3) still pins 1.32.0, so versions.automation is left at 1.1.3.
* chore: bump openhands-automation 1.1.3 -> 1.1.4 (pins SDK 1.33.0)
openhands-automation 1.1.4 is now published on PyPI and pins
openhands-sdk/openhands-workspace to 1.33.0, matching versions.agentServer.
This satisfies the check-sdk-version-sync check, which failed while
automation stayed at 1.1.3 (pinned SDK 1.32.0). Consolidates the
SDK+automation bump from #1622 into this PR alongside the ts-client bump.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Debug Agent <157206163+simonrosenberg@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Squash merge PR #1583.
This merge commit was created by an AI agent (OpenHands) on behalf of Graham Neubig.
Co-authored-by: openhands <openhands@all-hands.dev>
* Use libraries for local proxy and static serving
* Fix CI for proxy library refactor
* Fix static server CI failures
Co-authored-by: openhands <openhands@all-hands.dev>
---------
Co-authored-by: Codex <codex@openai.com>
Co-authored-by: openhands <openhands@all-hands.dev>
* fix(examples): inherit acp-docker image from config/defaults.json
examples/acp-docker/docker-compose.yml hardcoded the agent-server image at
`1.25.0-python`. Canvas enforces `compatibility.minimumAgentServer` (1.28.0)
from the repo's single source of truth, so the example default fell below the
floor and rendered "Disconnected — requires 1.28.0 or newer" — a reviewer
following the quickstart as written never reached the feature.
examples/acp-docker was the lone in-repo file hardcoding a version instead of
inheriting from config/defaults.json (14 other files read it; check-sdk-version
-sync only validates the released PyPI package, not in-repo files).
- scripts/gen-acp-docker-env.mjs: read defaults.json, pin AGENT_SERVER_IMAGE to
`${images.agentServer}:${versions.agentServer}-python` in examples/acp-docker
/.env (idempotent upsert; mirrors scripts/docker-build.mjs).
- package.json: `npm run example:acp-docker:env`.
- docker-compose.yml: no-config fallback `1.25.0-python` -> `latest-python`,
always >= the compatibility floor, so zero-config `docker compose up` never
shows "Disconnected"; the generated .env overrides with the pinned SoT
version for the reproducible path.
- .env.example / README.md: document both paths; correct the version narrative
(floor is the defaults.json compatibility pin; #3510 is the deeper functional
floor at/below it).
- __tests__/scripts/acp-docker-env-sync.test.ts: assert the generator's tag
matches defaults.json, the pin satisfies the floor, and the compose fallback
stays `latest-python`. Mirrors docs-version-sync.test.ts — the guard that
makes "can't silently drift" true.
* test(examples): harden acp-docker env-sync per review
Addresses the cli-review-panel findings worth acting on (the rest were
cosmetic or matched the no-validation idiom of scripts/docker-build.mjs):
- gte() in the test guarded with parseSemver — a non-numeric pin (sha /
pre-release) now fails the floor check loudly instead of silently
comparing NaN. The floor check is a CI gate; its one piece of logic
shouldn't mis-compare in silence.
- compose-fallback assertion derives the registry from config.images
.agentServer instead of hardcoding ghcr.io/openhands/... — a registry
change no longer false-fails a test that only cares about the latest-python
tag.
- upsertEnvLine now has unit tests (append / replace-in-place+preserve /
idempotent / commented-template-line / keyless-line guard), making the
"idempotent upsert" claim defensible. It was the one untested piece of real
logic.
- upsertEnvLine guards a keyless line (no "=") with a clear throw, instead of
an empty key matching every line and rewriting the whole file.
* fix(examples): guard acp-docker env-sync entrypoint against undefined argv[1]
The CLI entrypoint guard called pathToFileURL(process.argv[1]) unconditionally.
process.argv[1] is undefined in some ESM contexts (e.g. importing the module for
its exports via `node --input-type=module -e "import(...)"`), so the guard threw
ERR_INVALID_ARG_TYPE at import, before any exported helper was reachable.
Short-circuit on process.argv[1] before pathToFileURL so importing the module is
side-effect-free while the CLI path is unchanged. Add a regression test that
reproduces the bare-import context and asserts a clean exit.
Addresses the review finding on #1434.
* docs(acp-docker): trim verbose comments per review
Address all-hands-bot's review suggestions on #1434:
- test header describes the current invariant, not the prior-state history
(that narration belonged in the PR description)
- docker-compose.yml: condense the image-pin comment to the how-to-override;
the compatibility-floor / #3510 rationale already lives in README §1 + the test
- .env.example: 7-line pin explainer down to 2
Comment-only; env-sync test still 10/10 green, prettier clean.
* Clarify ACP Docker image version guidance
Co-authored-by: openhands <openhands@all-hands.dev>
---------
Co-authored-by: enyst <engel.nyst@gmail.com>
Co-authored-by: openhands <openhands@all-hands.dev>
* chore(deps): bump @openhands/typescript-client to 1.27.0
* test: update ACP provider/model fixtures for typescript-client 1.27.0
1.27.0 refreshed the claude-code/codex ACP registry data: provider command
versions (claude-agent-acp 0.30.0->0.44.0, codex-acp 0.15.0->0.16.0),
claude-code model ids (claude-opus-4-8->opus[1m], claude-sonnet-4-6->sonnet,
claude-haiku-4-5->haiku) plus a new well-labeled "default" option, and the
codex default (gpt-5.5/medium->gpt-5.5).
Canvas sources these lists from the client registry (closes#740), so the
source was already correct -- only the hardcoded test expectations were
stale. Also relaxed the acp-providers placeholder guard to accept the SDK's
intentional "Default (recommended)" entry.
* chore(deps): bump agent-server/openhands-sdk to 1.29.0
Align the spawned agent-server SDK release train (openhands-sdk,
openhands-tools, openhands-workspace, openhands-agent-server) with the
version @openhands/typescript-client 1.27.0 is validated against
(agent-server 1.29.0-python). Bump the coupled openhands-automation pin
to 1.0.0a12, whose SDK deps resolve to 1.29.0, to satisfy the
check-sdk-version-sync gate. minimumAgentServer compat floor unchanged.
Doc/JSDoc/test references updated to keep docs-version-sync green.
* Update typescript-client to v1.25.0 with pinned ACP versions
Upgrades from v1.24.3 to v1.25.0 to get the ACP provider version pins:
- claude-agent-acp@0.30.0
- codex-acp@0.15.0
- gemini-cli@0.38.0
Fixes the 'Method not found' error when starting Codex ACP sessions.
The unpinned npm package was resolving to codex-acp@0.16.0 which
lacks the session/set_model RPC method that the SDK expects.
Closes: agent-canvas#<issue> (codex version incompatibility)
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
* chore: update package-lock.json typescript-client version to 1.25.0
Synchronize lockfile dependency version with package.json to fix npm ci validation.
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
* chore: complete lockfile update for typescript-client@1.25.0
Update both dependency version and package definition with correct integrity hash.
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
* test: update ACP provider command expectations for pinned versions
Update test expectations to match typescript-client@1.25.0 registry which includes
version pins: claude-agent-acp@0.30.0, codex-acp@0.15.0, gemini-cli@0.38.0.
These pins fix the 'Method not found' error on Codex ACP sessions by ensuring
npm resolves to compatible versions.
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
* fix: update ACP test expectations for typescript-client@1.25.0 pinned versions
- Change empty acp_command to fall back to registry defaults (now with pinned versions)
- Update expected Claude Code default model from opus-4-7 to opus-4-8
* fix: update test expectations for typescript-client@1.25.0 registry changes
- Default Claude Code model updated: claude-opus-4-7 → claude-opus-4-8
- Pinned commands now include version: @claude-agent-acp@0.30.0, @codex-acp@0.15.0
- Credential tests load acp_command:[] so preset detection finds claude-code
- Reconciles test types pinned codex command to match registry
---------
Co-authored-by: Debug Agent <simon@openhands.dev>
Co-authored-by: Claude Haiku 4.5 <noreply@anthropic.com>
* Add Windows portability guards for workspace flows
* Increase snapshot workflow timeout
* Trigger CI after timeout update
* Fix windows portability PR after main merge
---------
Co-authored-by: neubig <398875+neubig@users.noreply.github.com>
* chore: bump agent-server → 1.28.1, automation → 1.0.0a9, extensions → 0.4.1
Co-authored-by: openhands <openhands@all-hands.dev>
* Test fixes
* fix: inject proxy base_url for litellm_proxy/* when server omits it (agent-server ≥1.28)
Agent-server ≥1.28 may return base_url:null when fetching a litellm_proxy/*
profile config, even when the profile was saved with the All-Hands proxy URL.
This caused the Basic-tab re-save flow in LlmSettingsLocalView.handleSave to
call isOpenHandsProxyModel(model, null) → false, hitting the else-branch that
deletes base_url and stranding the profile (issue #1146).
Fix: add a secondary check — litellm_proxy/* with a missing base_url is treated
the same as litellm_proxy/* with the proxy URL already set, and
OPENHANDS_LLM_PROXY_BASE_URL is injected before the save request is sent.
Also updates the mock-LLM E2E test to accept both storage representations:
- litellm_proxy/* + proxyBaseUrl (pre-1.28, guards issue #1146 regression)
- openhands/* + null (1.28+, server-managed routing)
And adds a unit test exercising the base_url:null path.
Co-authored-by: openhands <openhands@all-hands.dev>
---------
Co-authored-by: openhands <openhands@all-hands.dev>
* chore: bump agent-server → 1.28.1, automation → 1.0.0a9, extensions → 0.4.1
Co-authored-by: openhands <openhands@all-hands.dev>
* chore: update doc examples to reference agent-server 1.28.1
Update version references in AGENTS.md, scripts/dev-safe.mjs, and
scripts/check-sdk-version-sync.mjs from 1.27.0 → 1.28.1 to stay
in sync with the agentServer pin in config/defaults.json.
Fixes: docs-version-sync.test.ts failures
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: inject proxy base_url for litellm_proxy/* when server omits it (agent-server ≥1.28)
Agent-server ≥1.28 may return base_url:null when fetching a litellm_proxy/*
profile config, even when the profile was saved with the All-Hands proxy URL.
This caused the Basic-tab re-save flow in LlmSettingsLocalView.handleSave to
call isOpenHandsProxyModel(model, '') → false, hitting the else-branch that
deletes base_url and stranding the profile (issue #1146).
Fix: add a secondary check for litellm_proxy/* models with a missing base_url
(null/undefined/empty), treating them the same as a stored proxy URL and
injecting OPENHANDS_LLM_PROXY_BASE_URL before the save request is sent.
Also adds a unit test exercising the base_url:null path.
Co-authored-by: openhands <openhands@all-hands.dev>
* test(e2e): accept agent-server 1.28 model rewrite in proxy profile test
Agent-server 1.28 normalises litellm_proxy/* → openhands/* on storage
and manages the proxy URL internally (returning base_url:null). The old
assertions hard-coded the pre-1.28 storage format (litellm_proxy/* +
explicit proxy URL), causing the test to fail on every 1.28 run.
Extract assertProxyProfileConfig() helper that accepts both storage
representations:
- litellm_proxy/* + proxyBaseUrl (pre-1.28, guards issue #1146 regression)
- openhands/* + null (1.28+, server-managed routing)
The issue #1146 guard is preserved: a litellm_proxy/* profile without a
proxy URL is still flagged as a stranded profile.
Co-authored-by: openhands <openhands@all-hands.dev>
---------
Co-authored-by: openhands <openhands@all-hands.dev>
Keep package metadata in sync with config/defaults.json so npm scripts and mock E2E logs report the current agent-canvas RC version.
Co-authored-by: openhands <openhands@all-hands.dev>
* build(deps): move @openhands/extensions to npm 0.2.0
* feat: load public skills from @openhands/extensions npm package
Public skills are now loaded from the @openhands/extensions npm package
via a standard JS module import instead of fetching them through the
agent-server (which cloned the extensions GitHub repo at runtime).
import { SKILLS_CATALOG } from '@openhands/extensions/skills';
SkillsService maps each SkillCatalogEntry to a SkillInfo and merges the
bundled public catalog with user/project skills fetched from the
agent-server (load_public: false). If the agent-server is unreachable,
the bundled catalog is returned alone.
Changes:
- SkillsService: imports SKILLS_CATALOG from @openhands/extensions/skills,
maps entries to SkillInfo, merges with user/project skills from
agent-server (load_public: false).
- agent-server-adapter: hardcodes load_public_skills: false in
buildAgentContext().
- agent-server-config: removes shouldLoadPublicSkills() and its
VITE_LOAD_PUBLIC_SKILLS env var.
- dev-safe.mjs: removes getExtensionsRef() / DEFAULT_EXTENSIONS_REF
and EXTENSIONS_REF injection in buildAgentServerEnv().
- Docker: removes CONFIG_EXTENSIONS_REF from config-gen stage and
EXTENSIONS_REF from entrypoint.sh.
- .env.sample: removes VITE_LOAD_PUBLIC_SKILLS comment.
- Tests updated to match new architecture.
Depends on OpenHands/extensions#310 which adds the SKILLS_CATALOG export.
Co-authored-by: openhands <openhands@all-hands.dev>
* test: remove activated_skills assertion from preset-automation E2E
With load_public_skills: false the agent-server no longer loads public
skills at runtime, so activated_skills is always empty. The conversation
itself works (slash command sent, agent replies) — only the server-side
skill activation metadata is gone.
Co-authored-by: openhands <openhands@all-hands.dev>
* feat: pass bundled public skills via agent_context.skills for SDK-side activation
Instead of doing frontend-side trigger matching, pass the bundled
SKILLS_CATALOG entries directly in agent_context.skills at conversation
start. The SDK performs trigger matching, sets activated_skills on user
events, and injects skill content into the system prompt — the exact
same behavior as when load_public_skills was true, but without cloning
the extensions repo at runtime.
buildBundledSkills() converts each catalog entry into the SDK Skill JSON
shape with KeywordTrigger ({ type: 'keyword', keywords: [...] }) for
skills with triggers, or null for always-active skills.
Restores the activated_skills E2E assertion in the preset-automation
test since the SDK now handles activation.
Co-authored-by: openhands <openhands@all-hands.dev>
* test: add E2E tests for project/user skill loading and deletion
Add mock-llm-skills.spec.ts with three tests:
1. Project skill in workspace/.agents/skills/ triggers on matching keyword
2. User skill in ~/.openhands/skills/ triggers on matching keyword
3. Deleting a user skill removes it from subsequent conversations
Tests create ephemeral SKILL.md files with unique trigger keywords,
send messages through the real agent-server stack, and verify
activated_skills in the conversation events API.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: use explicit APIRequestContext type import for CI TS6 compatibility
Replace inline `import('@playwright/test').APIRequestContext` type
references with a proper top-level type import. Also align afterEach
fixture destructuring with other specs' pattern.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: remove node: prefix from imports to fix CI TS resolution
TypeScript 6 on CI (Node 24) has a type resolution conflict when
`node:` prefixed imports (node:path, node:fs, node:os) coexist with
`@playwright/test` types in the same file. This caused
`APIRequestContext` to be incorrectly resolved as `Page`. Use
unprefixed imports (path, fs, os) which work identically in Node.js.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: split fs helpers into separate file to fix CI TS6 type resolution
Move node built-in imports (path, fs, os) and filesystem helpers to
`utils/skill-test-helpers.ts`. The spec file now only imports from
`@playwright/test` and the two helper modules, avoiding the type
resolution conflict between node builtins and Playwright fixture types
that caused `APIRequestContext` to be incorrectly inferred as `Page`
on CI (TypeScript 6 / Node 24 / Ubuntu).
API assertion logic is now inline within each test step, using the
`request` fixture directly instead of standalone functions with
explicit `APIRequestContext` type annotations.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: use namespace imports to avoid TS6 type inference issue
Switch from named imports to namespace imports (`import * as helpers`)
with subsequent destructuring. This changes how TypeScript resolves the
imported function signatures, avoiding a Node 24 / TS6 type inference
bug where `ensureMockLLMProfile` was incorrectly resolved as expecting
`Page` instead of `APIRequestContext`.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: add typed wrapper for ensureMockLLMProfile to fix CI TS2345
Add a local `configureMockLLM` wrapper with an explicit
`APIRequestContext` type annotation. This works around a CI-specific
TypeScript 6 type inference issue where the imported
`ensureMockLLMProfile` signature is incorrectly resolved as expecting
`Page` instead of `APIRequestContext` when called from a Playwright
test body that also imports from `skill-test-helpers` (a module with
node built-in imports). The wrapper's explicit type annotation forces
correct type checking at the call site.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: inline ensureMockLLMProfile logic to fix CI TS2345
Instead of importing ensureMockLLMProfile from mock-llm-helpers (which
triggers a CI-specific TS6 type inference bug when combined with
skill-test-helpers imports), inline the same logic as a local function
with explicit APIRequestContext typing. This avoids the cross-module
type resolution issue entirely.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: resolve WORKSPACE_DIR relative to agent-server CWD, not STATE_DIR
The agent-server resolves the relative working_dir ("workspace/project")
from its own CWD (the project root), not from STATE_DIR/workspaces.
The test was writing skill files to the wrong directory so the SDK
never found them, causing activated_skills to be empty.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: create standalone git repo for project skill E2E test
The agent-server creates a git worktree for each conversation, and only
committed files appear in worktrees. The previous approach wrote skill
files to the filesystem without committing them, so the worktree never
contained them and load_project_skills found nothing.
Now the test:
1. Creates a standalone git repo (.tmp/mock-llm-skill-repos/) with the
skill file committed
2. Creates the conversation via API with that repo as working_dir
3. The agent-server worktree includes the committed skill
4. load_project_skills discovers it in the worktree
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: add secrets_encrypted flag to skill test conversation creation
The GET /api/settings with X-Expose-Secrets: encrypted returns cipher-
encrypted secret values. The POST /api/conversations needs
secrets_encrypted: true to tell the server to decrypt them, otherwise
the request fails with HTTP 422.
Co-authored-by: openhands <openhands@all-hands.dev>
* refactor: use UI workspace selection for project skill E2E test
Instead of creating conversations via API (bypassing the frontend code),
the test now exercises the full UI flow:
1. Creates a standalone git repo with the skill committed
2. Registers the repo as a workspace via POST /api/workspaces
3. Opens the 'Open workspace' dialog in the UI
4. Selects the workspace from the dropdown
5. Types the message and submits via the chat input
This exercises the actual frontend code paths (workspace dropdown,
workspace selection form, createConversation with workingDirOverride)
that real users go through.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: add padding response for skill-analysis in deletion test
The agent-server makes a skill-analysis LLM call even when no user/project
skills are loaded, because public skills from the npm package are still
present. The deletion test only had 1 trajectory response, causing the
agent to hang waiting for the 2nd response (the actual reply).
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: simplify deletion test to not depend on specific event type
The deletion test was failing because it waited for an event with
source='agent' and event_type='message' in the events API, but the
mock LLM text reply may produce a different event type. Since
waitForNonUserMessageText already confirms the agent replied in the
UI, we just need to verify no activated_skills contains the deleted
skill name.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: mount skill test dirs into Docker container for e2e tests
The Docker E2E skills test was failing because the agent-server inside
the Docker container couldn't access skill repos and user skill files
created on the host filesystem.
Fix by:
- Adding volume mounts for skill repos (.tmp/mock-llm-skill-repos/ →
/tmp/mock-llm-skill-repos/) and user skills (.tmp/mock-llm-user-skills/
→ /home/openhands/.openhands/skills/) to the Docker run command
- Setting env vars (MOCK_LLM_SKILL_REPOS_CONTAINER_DIR,
MOCK_LLM_USER_SKILLS_HOST_DIR) so skill-test-helpers.ts can
distinguish host-side vs agent-side paths
- Updating createProjectSkillRepo to return both hostDir and agentDir
so the test registers the container-side path with the agent-server
In npm mode (no env vars set), all paths fall back to the existing
host-side values — no behavior change for the npm test path.
Co-authored-by: openhands <openhands@all-hands.dev>
* docs: document Docker skill test volume mounts in AGENTS.md
Co-authored-by: openhands <openhands@all-hands.dev>
* feat: mark newly added mock-LLM E2E tests with 🆕 badge in PR comments
The render-mock-llm-report.mjs script now accepts a --new-files flag
with a comma-separated list of spec file paths added in the PR. Tests
from those files get a 🆕 badge in the results table, and the summary
line shows the count (e.g. '🆕 2 new').
Both CI workflows (mock-llm-e2e.yml and mock-llm-docker-e2e.yml) add
a 'Detect newly added spec files' step that queries the GitHub API
for files with status=='added' matching the mock-LLM spec pattern,
avoiding shallow-clone issues with git diff.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: match Playwright basename file paths against repo-relative --new-files
Playwright's JSON reporter emits file paths relative to testDir
(e.g. 'mock-llm-skills.spec.ts') while the GitHub API returns
repo-relative paths (e.g. 'tests/e2e/mock-llm/mock-llm-skills.spec.ts').
The isNewTest() matcher now compares basenames in addition to exact/suffix
matching, so 🆕 badges render correctly.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: stabilize pagination loading-indicator test + improve new-test callout
1. Flaky test fix: the 'loads older events when scrolling up' test
asserts that the loading-older-events indicator appears, but the
instant mock response lets React batch isLoading true→false in one
commit — the DOM element never materialises. Add a 300ms delay to
older-events mock responses so the indicator renders reliably.
2. Better new-test visibility: replace the subtle inline 🆕 emoji with
a prominent green blockquote callout above the results table that
lists each new test with its status icon and spec file.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: address PR review — type safety, docs, test assertions
1. Define BundledSkill interface for buildBundledSkills() return type
instead of the opaque SettingsRecord[] (review thread #1).
2. Document PUBLIC_SKILLS as an immutable build-time snapshot that is
baked into the bundle and requires a dependency bump to update
(review thread #2).
3. Add migration note to buildAgentContext() explaining that the former
VITE_LOAD_PUBLIC_SKILLS env var was removed because bundled skills
have no clone latency. load_public_skills: false is still passed to
tell the SDK to skip its own clone (review thread #3).
4. Add structural assertions for individual skill entries in the adapter
test: name, content, source, is_agentskills_format, and trigger
shape (review testing gap).
5. Update stale VITE_LOAD_PUBLIC_SKILLS comments in E2E test files.
Co-authored-by: openhands <openhands@all-hands.dev>
---------
Co-authored-by: Joe Laverty <joe.laverty@openhands.dev>
Co-authored-by: openhands <openhands@all-hands.dev>
* non-breaking updates
* fix: update overrides to resolve npm audit vulnerabilities
- ajv: add override to 8.20.0 (fixes ReDoS in $data option, GHSA-2g4f-4pwh-qvx6, affected 7.0.0-alpha.0–8.17.1)
- dompurify: bump override from 3.3.2 to 3.4.7 (fixes XSS bypasses, GHSA-39q2-94rc-95cp and others, affected <=3.3.3)
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: scope ajv override to @vercel/static-config to fix lint crash
The blanket 'ajv: 8.20.0' override forced AJV 8.x globally, breaking
ESLint's @eslint/eslintrc which requires AJV ^6.x (incompatible API).
Scope the override to @vercel/static-config only — the sole vulnerable
consumer (ajv 8.6.3 via @vercel/react-router). ESLint-related packages
now get their own nested ajv@6.15.0 instead of the incompatible 8.x.
npm audit: 0 vulnerabilities npm run lint: ✅
Co-authored-by: openhands <openhands@all-hands.dev>
* Bump react-router to 7.17.0 to fix GHSA-8x6r-g9mw-2r78
Resolves 5 high-severity npm audit findings for the React Router
DoS-via-unbounded-path-expansion advisory affecting react-router
7.0.0 – 7.14.2 and the dependent @react-router/{node,dev,serve,express}
packages.
- Bumped @react-router/node, @react-router/serve, @react-router/dev,
and react-router (incl. peerDep) from 7.14.2 to 7.17.0.
- Regenerated package-lock.json cleanly (the resolver ERESOLVE-looped
when trying to upgrade in place from the existing lockfile).
- Updated AGENTS.md: dropped the obsolete vite-tsconfig-paths
nested-typescript lockfile invariant (no longer a dep), rewrote the
@openhands/typescript-client git-dep note to reflect that it is now
a registry package and the Vercel ssh→https rewrite now protects
@openhands/extensions, and added a tip to regenerate the lockfile
cleanly when bumping pinned versions.
npm audit: 0 vulnerabilities. typecheck + build verified.
Co-authored-by: openhands <openhands@all-hands.dev>
* docs(AGENTS.md): document CVEs addressed by package.json overrides
For each entry in 'overrides' in package.json, record the specific
advisory it patches and (for ajv) why it's scoped to @vercel/static-config
rather than applied globally. Future maintainers can decide when an
override can be dropped (upstream bumps past the fixed version) without
re-deriving the context from git history.
- @vercel/static-config > ajv: 8.20.0 -> GHSA-2g4f-4pwh-qvx6
- dompurify: 3.4.7 -> GHSA-39q2-94rc-95cp
Co-authored-by: openhands <openhands@all-hands.dev>
---------
Co-authored-by: openhands <openhands@all-hands.dev>
* feat: add daily-rotating file logger for dev scripts (issue #815)
Add winston + winston-daily-rotate-file to write all dev-server log
output to logs/agent-canvas.YYYY-MM-DD.log alongside the existing
console output (which is unchanged).
- scripts/logger.mjs — shared module; exports fileLog(level, msg)
and stripAnsi(str). DailyRotateFile transport stores files in
logs/ relative to the project root, rotates at midnight, and
auto-deletes files older than 7 days.
- scripts/dev-with-automation.mjs — logService / logStep /
logSuccess / logError each call fileLog as a side-channel. The
shutdown message, startup title, checkPrerequisites uvx-error, and
printBanner summary are also captured.
- scripts/dev-safe.mjs — spawnProcess errors, main() startup lines,
the unexpected-exit error, and the fatal-error handler all call
fileLog.
- logs/ was already in .gitignore.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: store log files in agent-canvas state dir, not project root
Use OH_CANVAS_SAFE_STATE_DIR (or ~/.openhands/agent-canvas as
the default) to match where all other agent-canvas runtime state
lives, e.g. ~/.openhands/agent-canvas/logs/agent-canvas.YYYY-MM-DD.log
Co-authored-by: openhands <openhands@all-hands.dev>
---------
Co-authored-by: openhands <openhands@all-hands.dev>
* feat: update extensions to slash-command catalog prompts
Update @openhands/extensions to pick up slash-command catalog prompts
(/slack-monitor:poll, /github-monitor:poll).
The buildAutomationPrompt boilerplate still appends API routing info
since the skills don't distinguish local vs cloud backends themselves.
Companion PR: https://github.com/OpenHands/extensions/pull/298
Co-authored-by: openhands <openhands@all-hands.dev>
* chore: update extensions to ed3fcc4 (all catalog entries use slash commands)
Co-authored-by: openhands <openhands@all-hands.dev>
* test: add mock-LLM e2e for preset automation → slash command flow
Exercises the full recommended-automations flow:
1. Configure Slack MCP and mock LLM profile via API
2. Navigate to /automations, click Slack standup digest card
3. Verify conversation opens with /standup-digest:setup
4. Submit the slash command, verify skill activation
5. Mock LLM replies with text and conversation ends
Step 4 (skill activation) will fail until the extensions
PR #298 merges, since the agent-server needs the skill
with the /standup-digest:setup trigger registered.
Co-authored-by: openhands <openhands@all-hands.dev>
* chore: update @openhands/extensions to 8a66900
Includes CI fixes: marketplace entries, README, plugin manifests,
and vendor symlinks for all 5 new automation skills.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: use correct SDK mcp_config format (mcpServers wrapper)
Co-authored-by: openhands <openhands@all-hands.dev>
* fix(test): fix preset automation e2e — verify slash command triggers skill
Root cause: Two issues blocked the activated_skills assertion:
1. MCP configuration with a dummy command (echo/cat) caused the agent-
server to block for 30s in create_mcp_tools() during conversation
startup. The MCP client can't init with a non-MCP process, so the
conversation never processed the user message (0 events, no LLM
calls visible in mock server logs).
2. The events API sort_order parameter was 'TIMESTAMP_ASC' which is
invalid (valid values: TIMESTAMP, TIMESTAMP_DESC), returning 422.
Fix:
- Remove MCP config entirely; send the slash command from the home page
instead of clicking the automation card (which requires a working MCP)
- Install the slash-command skill via POST /api/skills/install so the
agent-server's AgentContext._load_auto_skills picks it up
- Verify the agent actually replies before checking events (proves the
conversation ran end-to-end)
- Poll events API with diagnostic output for debugging
- Clean up installed skill + temp files in afterAll
The activated_skills FE assertion is preserved — the FE renders it.
Co-authored-by: openhands <openhands@all-hands.dev>
---------
Co-authored-by: openhands <openhands@all-hands.dev>
* feat(mcp): render markdown links in helperText; bump extensions to slack field-order PR commit
- Add renderHelperText() to install-server-modal.tsx that converts
[text](url) patterns into <a> elements with target=_blank, so the
Slack workspace-ID helper text (and any future catalog entries) can
embed clickable docs links inline.
- Bump @openhands/extensions to commit 2d43e9c (branch
slack-catalog-field-order-and-helper-links, PR #285) which:
• moves SLACK_TEAM_ID before SLACK_BOT_TOKEN in the install modal
• replaces the plain SLACK_TEAM_ID helper text with linked copy:
'First visit [here](...#find-your-url) to get your Slack URL
and then visit [here](...#find-your-workspace-or-org-id) to
get your workspace ID.'
- Removes stale integrity hash from package-lock.json for the
@openhands/extensions entry; npm install will recompute it.
Co-authored-by: openhands <openhands@all-hands.dev>
* chore: bump @openhands/extensions to d186872 (SLACK_BOT_TOKEN helperText)
Add inline linked helperText for SLACK_BOT_TOKEN in slack.json (PR #285,
commit d186872): 'You'll need to create or update a Slack App as shown
[here](https://github.com/zencoderai/slack-mcp-server#slack-bot-setup).'
Drops the now-redundant helperLink field.
Co-authored-by: openhands <openhands@all-hands.dev>
* chore: bump @openhands/extensions to b45d3a1 (SLACK_TEAM_ID helperText rewrite)
Update SLACK_TEAM_ID helperText to named links:
'First get your [Slack URL](...). Then use that to get your [Workspace ID](...).'
Co-authored-by: openhands <openhands@all-hands.dev>
* chore: bump @openhands/extensions to 84a0a6e (SLACK_BOT_TOKEN named link)
Update SLACK_BOT_TOKEN helperText to:
"You'll need to create or update a [Slack App](...#slack-bot-setup)."
Co-authored-by: openhands <openhands@all-hands.dev>
* chore: bump @openhands/extensions to e07f427 (SLACK_BOT_TOKEN helperText)
Update SLACK_BOT_TOKEN helperText to:
"You'll need to create or update a [Slack App](...) to get a Bot token"
Co-authored-by: openhands <openhands@all-hands.dev>
* chore: bump @openhands/extensions to 5efd1b8
Sync to latest commit on slack-catalog-field-order-and-helper-links (PR #285).
Co-authored-by: openhands <openhands@all-hands.dev>
* chore: bump @openhands/extensions to 952c759
Sync to latest commit on slack-catalog-field-order-and-helper-links (PR #285).
Co-authored-by: openhands <openhands@all-hands.dev>
* chore: bump @openhands/extensions to f30dbfb
Sync to latest commit on slack-catalog-field-order-and-helper-links (PR #285).
Co-authored-by: openhands <openhands@all-hands.dev>
* chore: bump @openhands/extensions to 02715f4
Sync to latest commit on slack-catalog-field-order-and-helper-links (PR #285).
Co-authored-by: openhands <openhands@all-hands.dev>
* chore: bump @openhands/extensions to cb092c8
Sync to latest commit on slack-catalog-field-order-and-helper-links (PR #285).
Co-authored-by: openhands <openhands@all-hands.dev>
* fix(mcp): validate URL scheme in renderHelperText; use matchAll
- Guard href against javascript:/data: XSS via /^https?:\/\//i test
- Replace exec-in-while with matchAll to drop the eslint-disable comment
Addresses review bot feedback on PR #1012.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix(mcp): use double quotes for fallback href to satisfy Prettier
Co-authored-by: openhands <openhands@all-hands.dev>
* chore: update @openhands/extensions to latest main (62594156)
Co-authored-by: openhands <openhands@all-hands.dev>
---------
Co-authored-by: openhands <openhands@all-hands.dev>
* feat: reuse mock-LLM E2E tests for Docker image validation
Add a Docker-specific Playwright config (playwright.mock-llm-docker.config.ts)
that runs the exact same test specs and helpers against the agent-canvas Docker
image instead of the npm build path (bin/agent-canvas.mjs + uvx).
Key changes:
- Split MOCK_LLM_BASE_URL into two constants in mock-llm-helpers.ts:
- MOCK_LLM_BASE_URL: always host-local, used by tests for admin API
- MOCK_LLM_AGENT_URL: env-overridable, used when configuring the LLM
profile (the URL the agent-server uses for inference). Defaults to
MOCK_LLM_BASE_URL for backward compatibility with the npm path.
- New playwright.mock-llm-docker.config.ts:
- Starts the mock LLM server on the host (same as npm path)
- Runs the Docker container with --network host (Linux CI)
- Points to the same testDir (tests/e2e/mock-llm/) and specs
- Separate output dirs to avoid collision with npm path results
- New CI workflow (.github/workflows/mock-llm-docker-e2e.yml):
- Builds the Docker image from current code (or uses a pre-built image)
- Runs the same specs against the container
- Posts PR comment with differentiated report title
- render-mock-llm-report.mjs: accept --title flag for Docker vs npm reports
- npm run test:e2e:mock-llm:docker script added
- .gitignore updated for docker test output dirs
The npm path (test:e2e:mock-llm) is fully backward-compatible — no env var
override needed since MOCK_LLM_AGENT_URL defaults to MOCK_LLM_BASE_URL.
Co-authored-by: openhands <openhands@all-hands.dev>
* refactor: chain Docker E2E off existing Docker CI via workflow_run
Instead of rebuilding the Docker image in the E2E workflow (duplicating
~10-15 min of Docker build time), use workflow_run to trigger automatically
after the existing 'Docker' workflow completes successfully.
The workflow now:
- Triggers on: workflow_run (Docker completed) + workflow_dispatch (manual)
- Derives the image tag from the Docker build's commit SHA
(ghcr.io/openhands/agent-canvas:sha-<short>-amd64)
- Pulls the already-built image from GHCR — no rebuild needed
- Checks out code at the same SHA as the Docker build
- Extracts PR number from workflow_run.pull_requests[] for comments
Removed: Docker build steps, Buildx setup, build-arg resolution.
All image building stays in docker.yml where it belongs.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: replace flaky 1s timeout with polling for Active badge assertion
The 'Active badge' check in step 2 used a hardcoded 1-second
waitForTimeout before reloading. On a loaded CI runner the profile
activation mutation may not persist in time, causing the reload to
show stale state. This is a pre-existing flake (identical test code
passed on the first push and failed on the second).
Replace with expect.poll() that retries the reload+check cycle with
increasing intervals (1s, 2s, 3s) up to 15 seconds total.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: add pull_request trigger for Docker E2E (workflow_run bootstrap)
workflow_run only fires when the workflow file exists on the default
branch (main). Since mock-llm-docker-e2e.yml is new and only on the
PR branch, GitHub doesn't recognize it as a workflow_run listener yet.
Add pull_request trigger (gated by 'e2e-tests' label, skip forks) that
polls the Docker workflow via gh API until it completes for the PR's
head SHA, then pulls the already-built image from GHCR and runs tests.
After merge, workflow_run takes over as the primary automatic trigger.
The pull_request path remains as a fallback for label-gated runs.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: add FILE_STORE, AUTOMATION_BASE_URL, AUTOMATION_WORKSPACE_BASE to Docker entrypoint
The Docker entrypoint was missing several environment variables that the npm
path (dev-with-automation.mjs) sets for the automation backend:
- FILE_STORE=local — without this, the automation backend may fall back to
cloud storage (S3/GCS) which fails without credentials, causing tarball-
based presets (preset/prompt, preset/plugin) to silently error
- LOCAL_STORAGE_PATH — where to store files on the local filesystem
- AUTOMATION_BASE_URL — publicly-reachable base URL for callback URLs
- AUTOMATION_WORKSPACE_BASE — where automation runs unpack tarballs
This explains the Docker E2E failure: the agent's curl to create an automation
via /api/automation/v1/preset/prompt returned an error (likely 500 from missing
storage config), but the mock LLM doesn't care about terminal output and
proceeded to return the scripted final reply. The test then found 0 automations.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix: exclude auth-modes spec from Docker E2E tests
The mock-llm-auth-modes.spec.ts tests npm-binary-specific --auth-required
behaviour (a second static-server instance on port 18301). The Docker image
doesn't provide this second server — it has its own auth handling. Exclude
the spec from the Docker test run via testIgnore.
Co-authored-by: openhands <openhands@all-hands.dev>
* feat: run auth-modes tests inside Docker via PUBLIC_MODE_PORT
Instead of excluding the auth-modes spec from the Docker E2E run or
spinning up a host-side static server with a duplicate build/ directory,
the Docker entrypoint now supports an optional PUBLIC_MODE_PORT env var.
When set, entrypoint.sh starts a second static-server instance from the
same baked-in frontend assets with --auth-required (no session key
injected). This tests the actual Docker image's auth gate behaviour —
not a host-side approximation.
The Playwright Docker config passes -e PUBLIC_MODE_PORT=18301 to the
container and exports MOCK_LLM_PUBLIC_MODE_URL so the auth-modes spec
can reach it. With --network host the port is accessible from the host.
Co-authored-by: openhands <openhands@all-hands.dev>
* address review feedback: drop unlabeled trigger, improve error messages, document env vars
- Drop 'unlabeled' from pull_request trigger types to avoid wasted
workflow runs when any label is removed (the job-level if: condition
would skip immediately anyway)
- Distinguish 'no Docker run found' vs 'didn't complete in time' in
the polling loop's final error message
- Add comment explaining /api/automation/v1 probe returns 200 without
auth so the readiness check won't spin for 180s
- Document FILE_STORE, LOCAL_STORAGE_PATH, AUTOMATION_BASE_URL, and
AUTOMATION_WORKSPACE_BASE in the entrypoint header — these affect
production deployments, not just E2E tests
Co-authored-by: openhands <openhands@all-hands.dev>
---------
Co-authored-by: openhands <openhands@all-hands.dev>
The tools/ directory containing canvas_ui_tool.py was missing from the
package.json 'files' list, so it was not shipped in the published npm
tarball. Users running the released 'agent-canvas' CLI would get:
KeyError: "ToolDefinition 'canvas_ui' is not registered"
Failed to import module 'canvas_ui_tool': No module named 'canvas_ui_tool'
because the agent-server couldn't find the Python module that
dev-safe.mjs exposes via OH_EXTRA_PYTHON_PATH.
Co-authored-by: openhands <openhands@all-hands.dev>