mirror of
https://github.com/OpenHands/OpenHands.git
synced 2026-10-06 14:33:11 +08:00
* feat(helm): add Helm chart for StatefulSet deployment
Adds helm/agent-canvas/ chart that deploys the all-in-one agent-canvas
image as a StatefulSet with:
- PVC via volumeClaimTemplates mounted at $HOME/.openhands so
settings, encrypted secrets, conversation history, automation
SQLite DB, workspaces, and auto-generated keys all persist across
pod restarts and image upgrades. Supports BYO PVC via
persistence.existingClaim.
- Ingress template with the usual knobs (className, annotations,
hosts[].host + paths[].path/pathType, tls[].hosts/.secretName).
- ClusterIP + headless Services (headless required by StatefulSet).
- ServiceAccount always created for stable pod identity; RBAC
bindings gated by rbac.enabled, with:
* rbac.namespaces: list of namespaces to bind the SA into
via one RoleBinding per namespace against the built-in
`admin` ClusterRole (full namespace-scoped access).
* rbac.clusterAdmin: bool, off by default, adds a
ClusterRoleBinding to `cluster-admin` when true.
- Probes on /alive, podSecurityContext matching UID 1000 with
fsGroup 1000 so the PVC is writable, extraEnv passthrough for
LLM keys, optional external Postgres via AUTOMATION_DB_URL, and
optional existingSecret refs for OH_SECRET_KEY / session key.
helm lint passes; helm template verified against defaults,
full RBAC + ingress + TLS, existingClaim, and persistence.enabled=false.
Co-authored-by: openhands <openhands@all-hands.dev>
* fix(helm): pin appVersion to published agent-canvas tag 1.1.0
The initial commit accidentally pulled 1.23.0 from config/defaults.json,
which is the agent-server SDK version — not the ghcr.io/openhands/agent-canvas
image tag. Point appVersion at versions.agentCanvas (1.1.0) instead so the
default 'helm install' pulls a tag that actually exists on GHCR.
* feat(helm): mount entire $HOME on the PVC instead of just ~/.openhands
The previous default mounted the PVC at /home/openhands/.openhands, so
only the .openhands subtree survived pod restarts. Anything the agent
wrote under $HOME/workspace (cloned repos, generated files, worktrees)
was lost the next time the pod was rescheduled — which is the common
case since agents default to $HOME/workspace as the CWD.
Move the mount up to /home/openhands so the whole HOME persists:
* ~/.openhands — as before
* ~/workspace — the agent's actual working tree
* dotfiles the user creates (~/.gitconfig, ~/.cache, ~/.local, ~/.ssh)
Safe because the upstream image doesn't ship any dotfiles or venvs
under /home/openhands — Python packages are --system-installed under
/opt/agent-canvas — so an empty PVC on first boot doesn't shadow
anything, and the entrypoint recreates the .openhands subtree.
Migration note for existing installs: this changes the default
persistence.mountPath. Existing releases already using this chart
have a PVC whose ROOT is the .openhands directory; if they upgrade
without pinning the old mountPath, they'll get a fresh empty HOME
mounted over their existing .openhands contents and their state will
appear to reset. To preserve existing state, either:
* set persistence.mountPath: /home/openhands/.openhands to keep the
old behaviour, or
* delete the STS + PVC and start fresh.
* fix(helm): correct UID to 10001, mount PVC at ~/.openhands + ~/workspace via subPath
Two related bugs.
1. Wrong UID/GID/fsGroup
podSecurityContext pinned runAsUser/runAsGroup/fsGroup to 1000 based
on a stale assumption. The upstream agent-server image actually
runs as `openhands` at UID/GID **10001** (see
software-agent-sdk openhands-agent-server/openhands/agent_server/
docker/Dockerfile: `ARG USERNAME=openhands / UID=10001 / GID=10001`).
On Debian trixie, UID 1000 happens to be occupied by an unrelated
`pn` user carried in from a base layer, so `kubectl exec` shows
the wrong username and the process can't write to a PVC chowned to
gid 1000. Pin everything to 10001.
2. Wrong mount point
The previous commit moved persistence.mountPath from
/home/openhands/.openhands up to /home/openhands. That was wrong:
the image ships .bashrc, .profile, .bash_logout in that directory,
so mounting an empty PVC over the whole HOME shadows them and gives
users a bare-bones shell on `kubectl exec`.
Instead, mount the SAME PVC at multiple well-known subdirectories
via subPath. The default 'mounts' list covers the two paths that
actually need to persist:
- /home/openhands/.openhands (subPath: openhands)
- /home/openhands/workspace (subPath: workspace)
The pristine HOME (including dotfiles) stays intact from the image,
and both trees survive pod restarts and image upgrades on the same
disk. Users can extend the list with e.g. ~/.cache or ~/.config
without allocating additional PVCs.
Breaking change for existing installs relative to the previous
helm-chart-upstream tip commit — the values shape changed from a
single `persistence.mountPath` to a list `persistence.mounts`, and
the default UID is now 10001 instead of 1000.
* chore(helm): bump appVersion to 1.2.0
Point the chart's default image tag at the newly published
ghcr.io/openhands/agent-canvas:1.2.0. Chart version itself is
unchanged.
* fix(helm): keep version labels out of volumeClaimTemplates so upgrades work
Symptom: 'helm upgrade' failing with
StatefulSet.apps 'agent-canvas' is invalid: spec: Forbidden: updates
to statefulset spec for fields other than 'replicas', 'ordinals',
'template', 'updateStrategy', 'revisionHistoryLimit',
'persistentVolumeClaimRetentionPolicy' and 'minReadySeconds' are
forbidden
Cause: templates/statefulset.yaml stamps the full 'agent-canvas.labels'
set into 'volumeClaimTemplates[].metadata.labels'. That block lives
under STS.spec (NOT spec.template), which the apiserver treats as
immutable. The full label set includes 'app.kubernetes.io/version'
(from Chart.appVersion) and 'helm.sh/chart' (from Chart.version), both
of which change when either is bumped — turning routine version bumps
into forbidden-diff failures.
Fix: introduce a new helper 'agent-canvas.immutableLabels' that emits
only the labels that never change for a given release (name, instance,
managed-by), and use it in place of the full label set inside
volumeClaimTemplates. Object-level metadata.labels on the STS itself,
Services, ServiceAccount, etc. remain unchanged since those metadata
subtrees ARE mutable.
Existing installs one-shot recovery (before their next 'helm upgrade'):
# STS gone, pod + PVC preserved; helm upgrade will recreate STS
# from the fixed template and re-adopt both.
kubectl -n <ns> delete sts <release>-agent-canvas --cascade=orphan
helm upgrade <release> ./helm/agent-canvas -n <ns> -f values.yaml
* Apply suggestions from code review
Co-authored-by: Robert Brennan <accounts@rbren.io>
* Apply suggestion from @rbren
* Remove duplicate warning from README.md
Removed duplicate warning about the experimental nature of the Helm chart.
* docs(helm): clarify when to use this chart and relationship to OpenHands Enterprise
Address Graham's review feedback on PR #1641: the README didn't explain
why someone would want to use this chart or how it relates to OpenHands
Enterprise.
- Add a 'When to use this' section describing the self-hosted persistent
backend and internal vibecoding-platform use cases (pulled from the
companion docs PR OpenHands/docs#614).
- Add a 'Relationship to OpenHands Enterprise' section making the
unauthenticated / single-tenant / all-agents-comingled nature of this
chart explicit, and contrasting it with OHE's authentication,
role-based access control, multi-tenancy, and isolated agent
sandboxes.
- Add a 'Security' section covering authenticated ingress, LoadBalancer
exposure risk, and cluster-admin guidance, cross-referencing OHE.
Co-authored-by: openhands <openhands@all-hands.dev>
---------
Co-authored-by: openhands <openhands@all-hands.dev>
19 lines
187 B
Plaintext
19 lines
187 B
Plaintext
# Patterns to ignore when packaging Helm charts.
|
|
.DS_Store
|
|
.git/
|
|
.gitignore
|
|
.bzr/
|
|
.hg/
|
|
.svn/
|
|
*.swp
|
|
*.tmp
|
|
*.orig
|
|
*.bak
|
|
# CI / test
|
|
.github/
|
|
.circleci/
|
|
.travis.yml
|
|
# Editor
|
|
.idea/
|
|
.vscode/
|