Compare commits
43
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
547172de65 | ||
|
|
5f0d690c60 | ||
|
|
de0248c5db | ||
|
|
6643afbcda | ||
|
|
0a612a31c6 | ||
|
|
052319324d | ||
|
|
f885330b34 | ||
|
|
fcddbb0818 | ||
|
|
7691abf76a | ||
|
|
bfe8a87831 | ||
|
|
7ab6bd36d4 | ||
|
|
c4f82e4987 | ||
|
|
4f7697c43b | ||
|
|
0fc0211d26 | ||
|
|
fca30c7e26 | ||
|
|
ce7f45e18d | ||
|
|
070c3d5154 | ||
|
|
e275826236 | ||
|
|
7d40156003 | ||
|
|
bb7a02d8b7 | ||
|
|
c4b69402e1 | ||
|
|
2f5fd90947 | ||
|
|
b43aa104d3 | ||
|
|
d1d2a64d0f | ||
|
|
d5514f5cb3 | ||
|
|
f5915ca9ab | ||
|
|
a93ecee068 | ||
|
|
66daf27910 | ||
|
|
4b787be835 | ||
|
|
f18ff521fc | ||
|
|
5d710413d7 | ||
|
|
bef3da59a7 | ||
|
|
e234dac849 | ||
|
|
26894d835b | ||
|
|
2a5bbbeaae | ||
|
|
252fbabd51 | ||
|
|
85727ca625 | ||
|
|
4bc8622642 | ||
|
|
23bf594a70 | ||
|
|
2f15c1ece1 | ||
|
|
7dae4fcc41 | ||
|
|
d71fd1688b | ||
|
|
7b38b8aae2 |
@@ -0,0 +1,43 @@
|
||||
# GitNexus PR Reviewer Swarm — Claude Code adapter
|
||||
|
||||
This is the **Claude Code** entrypoint for the cross-CLI GitNexus PR reviewer swarm. The
|
||||
review logic itself is CLI-neutral and lives in **[`pr-swarm-review/`](../pr-swarm-review/README.md)**
|
||||
— that README is the canonical guide and covers every CLI (Claude Code, Gemini, Copilot,
|
||||
Cursor, Codex, and any AGENTS.md-aware agent).
|
||||
|
||||
## Invocation (Claude Code)
|
||||
|
||||
```
|
||||
/gitnexus-pr-swarm-review <PR URL or PR number>
|
||||
```
|
||||
|
||||
Runs in **Swarm mode**: the coordinator skill dispatches the seven `gitnexus-*` subagents in
|
||||
parallel (lanes 1–2 first, 3–6 in parallel, lane 7 last as a hard gate).
|
||||
|
||||
## Files in this adapter
|
||||
|
||||
| File | Role |
|
||||
|------|------|
|
||||
| `.claude/skills/gitnexus-pr-swarm-review/SKILL.md` | Coordinator — runs Swarm mode per `pr-swarm-review/orchestration.md` |
|
||||
| `.claude/agents/gitnexus-*.md` | Seven thin subagent wrappers; each reads its canonical persona in `pr-swarm-review/personas/` |
|
||||
|
||||
Each subagent keeps valid Claude Code frontmatter (model, tools, etc.); the mechanical
|
||||
verifier lanes (`test-ci-verifier`, `branch-hygiene-reviewer`) run on Haiku, the analytical
|
||||
lanes on Sonnet.
|
||||
|
||||
## Key properties
|
||||
|
||||
- **Read-only.** Tools limited to Read/Grep/Glob/Bash, and every persona enforces an
|
||||
explicit permitted/prohibited Bash list. No agent edits files, commits, or posts.
|
||||
- **Evidence-grounded**; **missing visibility becomes verification work**; **manually invoked.**
|
||||
|
||||
## Editing
|
||||
|
||||
Edit review behavior in the canonical files under `pr-swarm-review/` (orchestration +
|
||||
personas), **not** in these wrappers. After adding or editing files in `.claude/agents/`,
|
||||
restart Claude Code so it reloads the agent definitions.
|
||||
|
||||
## Relationship to `/gitnexus-pr-review`
|
||||
|
||||
Coexists with the single-agent `/gitnexus-pr-review` skill (a linear checklist using GitNexus
|
||||
MCP tools). This swarm is the multi-persona deep production-readiness review.
|
||||
@@ -0,0 +1,24 @@
|
||||
---
|
||||
name: gitnexus-branch-hygiene-reviewer
|
||||
description: "GitNexus branch hygiene and mergeability reviewer. Use to classify merge state, conflicts, stale branches, merge-from-main commits, unrelated churn, mixed domains, and whether rebase or split is required."
|
||||
tools:
|
||||
- Read
|
||||
- Grep
|
||||
- Glob
|
||||
- Bash
|
||||
model: claude-haiku-4-5-20251001
|
||||
maxTurns: 30
|
||||
---
|
||||
|
||||
# GitNexus Branch Hygiene & Mergeability Reviewer
|
||||
|
||||
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
|
||||
|
||||
**`pr-swarm-review/personas/02-branch-hygiene-reviewer.md`**
|
||||
|
||||
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
|
||||
|
||||
## Rules (always enforced)
|
||||
|
||||
- **Do not edit files.** You are read-only.
|
||||
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
|
||||
@@ -0,0 +1,24 @@
|
||||
---
|
||||
name: gitnexus-docs-dod-reviewer
|
||||
description: "GitNexus docs and Definition-of-Done reviewer. Use to translate repo guidance, linked issues, changed domains, docs requirements, release notes, and acceptance criteria into a PR-specific DoD."
|
||||
tools:
|
||||
- Read
|
||||
- Grep
|
||||
- Glob
|
||||
- Bash
|
||||
model: claude-sonnet-4-6
|
||||
maxTurns: 30
|
||||
---
|
||||
|
||||
# GitNexus Docs & Definition-of-Done Reviewer
|
||||
|
||||
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
|
||||
|
||||
**`pr-swarm-review/personas/06-docs-dod-reviewer.md`**
|
||||
|
||||
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
|
||||
|
||||
## Rules (always enforced)
|
||||
|
||||
- **Do not edit files.** You are read-only.
|
||||
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
|
||||
@@ -0,0 +1,24 @@
|
||||
---
|
||||
name: gitnexus-pr-facts-historian
|
||||
description: "GitNexus PR facts and repository-history investigator. Use to gather PR identity, visible GitHub state, changed files, commits, linked issues, related PRs, historical fixes, regressions, stale follow-ups, and missing visibility."
|
||||
tools:
|
||||
- Read
|
||||
- Grep
|
||||
- Glob
|
||||
- Bash
|
||||
model: claude-sonnet-4-6
|
||||
maxTurns: 40
|
||||
---
|
||||
|
||||
# GitNexus PR Facts & Repository-History Investigator
|
||||
|
||||
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
|
||||
|
||||
**`pr-swarm-review/personas/01-pr-facts-historian.md`**
|
||||
|
||||
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
|
||||
|
||||
## Rules (always enforced)
|
||||
|
||||
- **Do not edit files.** You are read-only.
|
||||
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
|
||||
@@ -0,0 +1,24 @@
|
||||
---
|
||||
name: gitnexus-risk-architect
|
||||
description: "GitNexus production-risk reviewer. Use for risk-model-first review of changed files, runtime behavior, multi-domain changes, user impact, failure modes, compatibility, and merge-blocking risk."
|
||||
tools:
|
||||
- Read
|
||||
- Grep
|
||||
- Glob
|
||||
- Bash
|
||||
model: claude-sonnet-4-6
|
||||
maxTurns: 40
|
||||
---
|
||||
|
||||
# GitNexus Production-Risk Architect
|
||||
|
||||
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
|
||||
|
||||
**`pr-swarm-review/personas/03-risk-architect.md`**
|
||||
|
||||
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
|
||||
|
||||
## Rules (always enforced)
|
||||
|
||||
- **Do not edit files.** You are read-only.
|
||||
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
|
||||
@@ -0,0 +1,24 @@
|
||||
---
|
||||
name: gitnexus-security-boundary-reviewer
|
||||
description: "GitNexus security and trust-boundary reviewer. Use for auth, permissions, secrets, injection, unsafe parsing, external input handling, hidden Unicode, YAML/Docker/workflow risks, and suspicious non-ASCII hygiene."
|
||||
tools:
|
||||
- Read
|
||||
- Grep
|
||||
- Glob
|
||||
- Bash
|
||||
model: claude-sonnet-4-6
|
||||
maxTurns: 35
|
||||
---
|
||||
|
||||
# GitNexus Security & Trust-Boundary Reviewer
|
||||
|
||||
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
|
||||
|
||||
**`pr-swarm-review/personas/05-security-boundary-reviewer.md`**
|
||||
|
||||
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
|
||||
|
||||
## Rules (always enforced)
|
||||
|
||||
- **Do not edit files.** You are read-only.
|
||||
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
|
||||
@@ -0,0 +1,24 @@
|
||||
---
|
||||
name: gitnexus-synthesis-critic
|
||||
description: "GitNexus final review synthesis critic. Use to check whether the final PR review is evidence-grounded, risk-prioritized, GitNexus-specific, non-generic, and follows required verdict rules."
|
||||
tools:
|
||||
- Read
|
||||
- Grep
|
||||
- Glob
|
||||
- Bash
|
||||
model: claude-sonnet-4-6
|
||||
maxTurns: 25
|
||||
---
|
||||
|
||||
# GitNexus Final-Review Synthesis Critic
|
||||
|
||||
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
|
||||
|
||||
**`pr-swarm-review/personas/07-synthesis-critic.md`**
|
||||
|
||||
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
|
||||
|
||||
## Rules (always enforced)
|
||||
|
||||
- **Do not edit files.** You are read-only.
|
||||
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
|
||||
@@ -0,0 +1,24 @@
|
||||
---
|
||||
name: gitnexus-test-ci-verifier
|
||||
description: "GitNexus test and CI reviewer. Use to verify whether changed behavior is covered by targeted tests, whether CI actually runs those tests, and whether workflow changes weaken validation."
|
||||
tools:
|
||||
- Read
|
||||
- Grep
|
||||
- Glob
|
||||
- Bash
|
||||
model: claude-haiku-4-5-20251001
|
||||
maxTurns: 35
|
||||
---
|
||||
|
||||
# GitNexus Test & CI Verifier
|
||||
|
||||
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
|
||||
|
||||
**`pr-swarm-review/personas/04-test-ci-verifier.md`**
|
||||
|
||||
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
|
||||
|
||||
## Rules (always enforced)
|
||||
|
||||
- **Do not edit files.** You are read-only.
|
||||
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
|
||||
@@ -0,0 +1,31 @@
|
||||
---
|
||||
name: gitnexus-pr-swarm-review
|
||||
description: "Run a GitNexus production-readiness pull request review using a coordinated reviewer swarm."
|
||||
---
|
||||
|
||||
# GitNexus PR Swarm Review (Claude Code adapter)
|
||||
|
||||
Use this skill to review a GitNexus pull request and produce a production-readiness review.
|
||||
|
||||
```
|
||||
/gitnexus-pr-swarm-review <PR URL or PR number>
|
||||
```
|
||||
|
||||
You are the **swarm coordinator**. The full review contract — lanes, dependencies,
|
||||
classifications, output structure, finding format, hidden-Unicode checks, and behavior
|
||||
rules — is the canonical, CLI-neutral spec:
|
||||
|
||||
**`pr-swarm-review/orchestration.md`** — read it now and follow it.
|
||||
|
||||
This adapter only pins the Claude Code specifics:
|
||||
|
||||
- **Run in Swarm mode.** Dispatch each lane as its own subagent via the Agent tool. The
|
||||
seven subagents are the project agents named `gitnexus-*` (one per persona); each reads
|
||||
its canonical persona under `pr-swarm-review/personas/`. Run lanes 1–2 first, lanes 3–6
|
||||
in parallel after, and lane 7 last on the draft.
|
||||
- **Lane 7 is a hard gate.** Do not emit the final review while the synthesis critic's
|
||||
"Required corrections before posting" section is non-empty — revise and re-run it.
|
||||
- Stay **read-only**: investigate and report; never edit, commit, or post.
|
||||
|
||||
Do not flatten the review into a generic checklist; delegate to the subagents and
|
||||
synthesize per `orchestration.md`.
|
||||
@@ -5,14 +5,16 @@ description: "Use when the user needs to run GitNexus CLI commands like analyze/
|
||||
|
||||
# GitNexus CLI Commands
|
||||
|
||||
All commands work via `npx` — no global install required.
|
||||
Commands below use `node .gitnexus/run.cjs <command>` — the project-local runner `gitnexus analyze` drops next to the index. It auto-selects an available runner at call time (global `gitnexus`, else `pnpm dlx`, else `npx`), so no package-manager assumption and no global install is required.
|
||||
|
||||
> **Not analyzed yet, or `node .gitnexus/run.cjs` reports `Cannot find module`** (the gitignored runner is absent — e.g. a fresh clone or `git clean`)? (Re)generate it with `npx gitnexus analyze` from the project root. On **npm 11.x**, if `npx` crashes during install (`node.target is null`), install once with `npm i -g gitnexus` (then `gitnexus analyze`) or use `pnpm --allow-build=@ladybugdb/core --allow-build=gitnexus --allow-build=tree-sitter dlx gitnexus@latest analyze`. See [#1939](https://github.com/abhigyanpatwari/GitNexus/issues/1939).
|
||||
|
||||
## Commands
|
||||
|
||||
### analyze — Build or refresh the index
|
||||
|
||||
```bash
|
||||
npx gitnexus analyze
|
||||
node .gitnexus/run.cjs analyze
|
||||
```
|
||||
|
||||
Run from the project root. This parses all source files, builds the knowledge graph, writes it to `.gitnexus/`, and generates CLAUDE.md / AGENTS.md context files.
|
||||
@@ -28,7 +30,7 @@ Run from the project root. This parses all source files, builds the knowledge gr
|
||||
### status — Check index freshness
|
||||
|
||||
```bash
|
||||
npx gitnexus status
|
||||
node .gitnexus/run.cjs status
|
||||
```
|
||||
|
||||
Shows whether the current repo has a GitNexus index, when it was last updated, and symbol/relationship counts. Use this to check if re-indexing is needed.
|
||||
@@ -36,7 +38,7 @@ Shows whether the current repo has a GitNexus index, when it was last updated, a
|
||||
### clean — Delete the index
|
||||
|
||||
```bash
|
||||
npx gitnexus clean
|
||||
node .gitnexus/run.cjs clean
|
||||
```
|
||||
|
||||
Deletes the `.gitnexus/` directory and unregisters the repo from the global registry. Use before re-indexing if the index is corrupt or after removing GitNexus from a project.
|
||||
@@ -49,7 +51,7 @@ Deletes the `.gitnexus/` directory and unregisters the repo from the global regi
|
||||
### wiki — Generate documentation from the graph
|
||||
|
||||
```bash
|
||||
npx gitnexus wiki
|
||||
node .gitnexus/run.cjs wiki
|
||||
```
|
||||
|
||||
Generates repository documentation from the knowledge graph using an LLM. Requires an API key (saved to `~/.gitnexus/config.json` on first use).
|
||||
@@ -66,7 +68,7 @@ Generates repository documentation from the knowledge graph using an LLM. Requir
|
||||
### list — Show all indexed repos
|
||||
|
||||
```bash
|
||||
npx gitnexus list
|
||||
node .gitnexus/run.cjs list
|
||||
```
|
||||
|
||||
Lists all repositories registered in `~/.gitnexus/registry.json`. The MCP `list_repos` tool provides the same information.
|
||||
|
||||
@@ -22,7 +22,7 @@ description: "Use when the user is debugging a bug, tracing an error, or asking
|
||||
4. gitnexus_cypher({query: "MATCH path..."}) → Custom traces if needed
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
||||
## Checklist
|
||||
|
||||
|
||||
@@ -23,7 +23,7 @@ description: "Use when the user asks how code works, wants to understand archite
|
||||
5. READ gitnexus://repo/{name}/process/{name} → Trace full execution flow
|
||||
```
|
||||
|
||||
> If step 2 says "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
> If step 2 says "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
||||
## Checklist
|
||||
|
||||
|
||||
@@ -15,7 +15,7 @@ For any task involving code understanding, debugging, impact analysis, or refact
|
||||
2. **Match your task to a skill below** and **read that skill file**
|
||||
3. **Follow the skill's workflow and checklist**
|
||||
|
||||
> If step 1 warns the index is stale, run `npx gitnexus analyze` in the terminal first.
|
||||
> If step 1 warns the index is stale, run `node .gitnexus/run.cjs analyze` in the terminal first.
|
||||
|
||||
## Skills
|
||||
|
||||
|
||||
@@ -23,7 +23,7 @@ description: "Use when the user wants to know what will break if they change som
|
||||
4. Assess risk and report to user
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
||||
## Checklist
|
||||
|
||||
|
||||
@@ -26,7 +26,7 @@ description: "Use when the user wants to review a pull request, understand what
|
||||
6. Summarize findings with risk assessment
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal before reviewing.
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal before reviewing.
|
||||
|
||||
## Checklist
|
||||
|
||||
|
||||
@@ -22,7 +22,7 @@ description: "Use when the user wants to rename, extract, split, move, or restru
|
||||
4. Plan update order: interfaces → implementations → callers → tests
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
||||
## Checklists
|
||||
|
||||
|
||||
@@ -0,0 +1,17 @@
|
||||
# GitNexus PR Swarm Review
|
||||
|
||||
You are the GitNexus PR review coordinator. Review the pull request named after this command
|
||||
(a PR URL or number for `https://github.com/abhigyanpatwari/GitNexus`). If none was given,
|
||||
ask for one.
|
||||
|
||||
Read `pr-swarm-review/orchestration.md` in this repository and follow it exactly — it is the
|
||||
canonical, CLI-neutral review contract (lanes, classifications, output structure, finding
|
||||
format, hidden-Unicode checks, behavior rules).
|
||||
|
||||
Run in **Solo mode**: you are a single agent, so perform all seven lanes yourself in
|
||||
dependency order, adopting each persona in `pr-swarm-review/personas/0N-*.md` in turn
|
||||
(lanes 1–2 first, then 3–6, then lane 7). Keep every lane's findings in context. Lane 7
|
||||
(synthesis critic) is a hard gate: do not emit the final review until its "Required
|
||||
corrections before posting" section is empty.
|
||||
|
||||
Stay strictly read-only: investigate and report; never edit files, commit, or post to GitHub.
|
||||
@@ -0,0 +1,153 @@
|
||||
# syntax=docker/dockerfile:1
|
||||
|
||||
# Base image: Microsoft's TypeScript+Node devcontainer image. It works on both
|
||||
# linux/amd64 and linux/arm64, gets monthly security patches, and ships the
|
||||
# non-root `node` user (UID 1000, i.e. user ID 1000), zsh + Oh My Zsh, eslint
|
||||
# global, and the `gh` CLI.
|
||||
#
|
||||
# We pin the image by digest, not by tag. That way a silent upstream retag can't
|
||||
# change the build under us. This matches the Dockerfile.cli /
|
||||
# gitnexus/Dockerfile.test convention and issue #1451.
|
||||
#
|
||||
# We pin it as a bare `name@digest` with NO `:tag` prefix on purpose. The
|
||||
# production Dockerfiles use plain `docker build`, but this one is built by
|
||||
# `@devcontainers/cli` / the VS Code Dev Containers resolver. That resolver's
|
||||
# image-name parser rejects the combined `name:tag@sha256:...` form.
|
||||
#
|
||||
# The digest below is for the `1-22-bookworm` tag. It is the multi-arch
|
||||
# manifest-list digest, so it still picks the right platform. To refresh it when
|
||||
# bumping the readable tag, run:
|
||||
# docker buildx imagetools inspect \
|
||||
# mcr.microsoft.com/devcontainers/typescript-node:1-22-bookworm \
|
||||
# --format '{{json .Manifest.Digest}}'
|
||||
FROM mcr.microsoft.com/devcontainers/typescript-node@sha256:7c2e711a4f7b02f32d2da16192d5e05aa7c95279be4ce889cff5df316f251c1d
|
||||
|
||||
# Build args. We deliberately set no version defaults here. devcontainer.json
|
||||
# `build.args` is the single source of truth for versions. A standalone
|
||||
# `docker build .devcontainer/` (for example, a CI smoke test) must pass each
|
||||
# version with --build-arg. Without a default, the build fails loudly instead of
|
||||
# silently drifting from the version pinned in devcontainer.json.
|
||||
ARG CLAUDE_CODE_VERSION
|
||||
ARG CODEX_VERSION
|
||||
# Cursor is pinned by version plus a per-arch tarball sha256 hash. The install
|
||||
# step below verifies that hash. All three values live in devcontainer.json
|
||||
# build.args. They follow the same rule as the others: one source of truth, and
|
||||
# no default so the build fails loudly if a value is missing.
|
||||
ARG CURSOR_VERSION
|
||||
ARG CURSOR_SHA256_X64
|
||||
ARG CURSOR_SHA256_ARM64
|
||||
# Bun is installed via the official remote script (bun.sh/install), pinned by
|
||||
# version. UNLIKE Cursor and the npm packages, this install path runs an
|
||||
# UNVERIFIED remote script — there is no tarball-hash check. Chosen explicitly
|
||||
# at request time over the pin-by-sha256 alternative for install-script
|
||||
# simplicity. To harden later, switch to a pinned tarball + per-arch sha256 in
|
||||
# the Cursor style (release artifacts at github.com/oven-sh/bun/releases).
|
||||
ARG BUN_VERSION
|
||||
ARG TZ=UTC
|
||||
ARG USERNAME=node
|
||||
|
||||
# Copy the build-only ARGs into runtime ENV so shells and lifecycle scripts can
|
||||
# read them. We deliberately do not set CLAUDE_CONFIG_DIR here. Its one true
|
||||
# value lives in devcontainer.json `containerEnv`, and the runtime value wins
|
||||
# anyway.
|
||||
ENV CLAUDE_CODE_VERSION=${CLAUDE_CODE_VERSION} \
|
||||
CODEX_VERSION=${CODEX_VERSION} \
|
||||
CURSOR_VERSION=${CURSOR_VERSION} \
|
||||
BUN_VERSION=${BUN_VERSION} \
|
||||
BUN_INSTALL=/home/${USERNAME}/.bun \
|
||||
TZ=${TZ} \
|
||||
DEVCONTAINER=true \
|
||||
NODE_OPTIONS=--max-old-space-size=4096 \
|
||||
POWERLEVEL9K_DISABLE_GITSTATUS=true
|
||||
|
||||
# Native build toolchain that gitnexus/postinstall needs. It compiles
|
||||
# tree-sitter native bindings, the vendored Dart/Proto/Swift grammars, and the
|
||||
# @ladybugdb/core N-API addon (a native Node add-on). python3, make, and g++ are
|
||||
# required. This mirrors the apt block in the existing Dockerfile.cli /
|
||||
# gitnexus/Dockerfile.test images.
|
||||
RUN apt-get update \
|
||||
&& apt-get install -y --no-install-recommends \
|
||||
python3 make g++ git curl ca-certificates bash unzip \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# Create and chown the named-volume mount points (~/.npm, ~/.local,
|
||||
# /commandhistory) up front. That way an empty volume inherits `node:node`
|
||||
# ownership the first time it is mounted. The three CLI config dirs (~/.claude,
|
||||
# ~/.codex, ~/.cursor) are bind-mounted from the host instead. A bind mount
|
||||
# completely hides the image-side ownership, so those paths need no chown here.
|
||||
RUN mkdir -p \
|
||||
/home/${USERNAME}/.npm \
|
||||
/home/${USERNAME}/.local/bin \
|
||||
/commandhistory \
|
||||
&& chown -R ${USERNAME}:${USERNAME} \
|
||||
/home/${USERNAME}/.npm \
|
||||
/home/${USERNAME}/.local \
|
||||
/commandhistory
|
||||
|
||||
USER ${USERNAME}
|
||||
|
||||
# Install Claude Code and the Codex CLI globally, as the `node` user. The base
|
||||
# image sets /usr/local/share/npm-global as the npm-global prefix and makes the
|
||||
# `npm` group writable by `node`. So `npm install -g` works without sudo. Both
|
||||
# versions come from build args. To upgrade, bump them in devcontainer.json and
|
||||
# rebuild.
|
||||
RUN npm install -g \
|
||||
@anthropic-ai/claude-code@${CLAUDE_CODE_VERSION} \
|
||||
@openai/codex@${CODEX_VERSION}
|
||||
|
||||
# Install the Cursor CLI. It is pinned and hash-verified, and we run no remote
|
||||
# script. The cursor.com/install script just detects os/arch, downloads a
|
||||
# versioned tarball from
|
||||
# downloads.cursor.com/lab/<version>/<os>/<arch>/agent-cli-package.tar.gz,
|
||||
# extracts it, and symlinks `agent`/`cursor-agent` into ~/.local/bin. We do that
|
||||
# ourselves against a PINNED version plus a per-arch sha256 hash. So the build
|
||||
# runs no unverified remote code. This matches how we pin the base image and npm
|
||||
# packages by digest (issue #1451). The download is fail-closed: if the hash
|
||||
# does not match, the build aborts.
|
||||
#
|
||||
# To bump: set CURSOR_VERSION and both CURSOR_SHA256_* in devcontainer.json
|
||||
# build.args. Get each arch's hash with:
|
||||
# curl -fSL https://downloads.cursor.com/lab/<ver>/linux/<x64|arm64>/agent-cli-package.tar.gz | sha256sum
|
||||
#
|
||||
# TARGETARCH is the per-platform build arg that BuildKit sets automatically. It
|
||||
# must be (re)declared in this stage to be visible. When the build is a
|
||||
# non-BuildKit `docker build`, TARGETARCH is unset, so we fall back to `dpkg
|
||||
# --print-architecture`.
|
||||
ARG TARGETARCH
|
||||
RUN set -eux; \
|
||||
arch="${TARGETARCH:-$(dpkg --print-architecture)}"; \
|
||||
case "$arch" in \
|
||||
amd64) cursor_arch=x64; cursor_sha="${CURSOR_SHA256_X64}";; \
|
||||
arm64) cursor_arch=arm64; cursor_sha="${CURSOR_SHA256_ARM64}";; \
|
||||
*) echo "unsupported architecture for Cursor: $arch" >&2; exit 1;; \
|
||||
esac; \
|
||||
url="https://downloads.cursor.com/lab/${CURSOR_VERSION}/linux/${cursor_arch}/agent-cli-package.tar.gz"; \
|
||||
curl -fSL --retry 3 --max-time 120 -o /tmp/cursor.tgz "$url"; \
|
||||
echo "${cursor_sha} /tmp/cursor.tgz" | sha256sum -c -; \
|
||||
dir="/home/${USERNAME}/.local/share/cursor-agent/versions/${CURSOR_VERSION}"; \
|
||||
install -d "$dir" "/home/${USERNAME}/.local/bin"; \
|
||||
tar --strip-components=1 -xzf /tmp/cursor.tgz -C "$dir"; \
|
||||
test -x "$dir/cursor-agent"; \
|
||||
ln -sf "$dir/cursor-agent" "/home/${USERNAME}/.local/bin/agent"; \
|
||||
ln -sf "$dir/cursor-agent" "/home/${USERNAME}/.local/bin/cursor-agent"; \
|
||||
rm -f /tmp/cursor.tgz
|
||||
|
||||
# Install Bun via the official remote installer, pinned by version. The first
|
||||
# positional arg to `bash` is the release tag (`bun-vX.Y.Z`), so a specific
|
||||
# version is fetched even though the install script itself is downloaded fresh
|
||||
# on every build. NOTE: this is the ONE remote script we run unverified in
|
||||
# this image — Cursor and the base image are pinned by sha256/digest. Hardening
|
||||
# path: switch to a pinned tarball + per-arch sha256 in the Cursor style
|
||||
# (artifacts at github.com/oven-sh/bun/releases). `BUN_INSTALL` is set in ENV
|
||||
# above so the binary lands at a known path regardless of any rc-file edits
|
||||
# the installer makes (which we ignore — we own the shell rc files).
|
||||
RUN set -eux; \
|
||||
curl -fsSL --retry 3 --max-time 120 https://bun.sh/install \
|
||||
| bash -s "bun-v${BUN_VERSION}"; \
|
||||
test -x "${BUN_INSTALL}/bin/bun"
|
||||
|
||||
# Put ~/.local/bin and Bun's bin dir on PATH for interactive shells and
|
||||
# lifecycle scripts. ~/.local/bin is where Cursor's installer drops the `agent`
|
||||
# and `cursor-agent` symlinks; ${BUN_INSTALL}/bin is where the Bun installer
|
||||
# drops `bun` / `bunx`.
|
||||
ENV PATH=/home/${USERNAME}/.local/bin:${BUN_INSTALL}/bin:${PATH}
|
||||
@@ -0,0 +1,364 @@
|
||||
# GitNexus Devcontainer
|
||||
|
||||
A cross-platform Dev Container that pre-installs Claude Code, OpenAI Codex CLI, Cursor CLI, and Bun alongside the GitNexus native build chain. Supported hosts: **macOS, Linux, Windows 11 (native), and Windows 11 via WSL2.** Windows-native needs a **one-time `HOME` env var setup** — handled automatically by the `initializeCommand` on first run (see [Windows 11 setup](#windows-11-setup)).
|
||||
|
||||
> ### ⚠️ Read this before using it on a work machine
|
||||
>
|
||||
> This devcontainer **does not write to your host AI-CLI config.** Your skills, agents, commands, plugins, memory, prompts, and rules are **copied once** from a read-only host stage into a per-container volume on first create; the container edits its own copy and can never write back. So a compromised workspace dependency running in the container **cannot** drop a malicious agent, command, skill, or plugin onto your host for your next host CLI session to load — the write-through vector earlier versions had is closed. Your **credentials** (Claude/Codex/Cursor logins, plus `gh`) likewise stay in per-container volumes and are never written back, and `~/.ssh`, `~/.aws`, `~/.azure`, and `~/.docker` are mounted **read-only**.
|
||||
>
|
||||
> What is **still** exposed: the read-only host stages (`/host/.claude`, `/host/.codex`, `/host/.cursor`, `/host/.claude-mem`) and the read-only credential mounts are all **readable** inside the container. A compromised dependency can therefore READ your host CLI config, memory, SSH/cloud credentials, and GitHub token — and there is **no egress firewall yet**, so it has the network to exfiltrate what it reads. Read-only protects you from tampering and write-back, not from disclosure.
|
||||
>
|
||||
> The trade-off of the copy model: host and container config **diverge after first create.** A skill or plugin you add on the host later won't appear in the container until you wipe the config volume and rebuild (see [§ Rebuild / reset](#rebuild--reset)). Edits you make inside the container persist across rebuilds but never reach the host.
|
||||
|
||||
## Quick start
|
||||
|
||||
1. Install [Docker Desktop](https://docs.docker.com/desktop/) (Windows/macOS) or Docker Engine (Linux).
|
||||
2. Install [VS Code](https://code.visualstudio.com/) with the [Dev Containers extension](https://marketplace.visualstudio.com/items?itemName=ms-vscode-remote.remote-containers).
|
||||
3. Install [Node.js](https://nodejs.org/) on the **host** (Node 18+). This is the only host-side toolchain dependency beyond Docker and VS Code — the devcontainer's `initializeCommand` runs `node .devcontainer/ensure-host-config-dirs.cjs` to set up the bind-mount source directories before container create. If you already use Claude Code or another Node-based CLI on the host, you're already set.
|
||||
4. Open the repo in VS Code → Command Palette → **Dev Containers: Reopen in Container**.
|
||||
5. Wait for the first build (~3–6 minutes) and `postCreateCommand` to finish installing workspace dependencies.
|
||||
6. Authenticate the three CLIs once — see [First-time CLI authentication](#first-time-cli-authentication) below.
|
||||
|
||||
## Windows 11 setup
|
||||
|
||||
### Windows-native (one-time setup, then "just works")
|
||||
|
||||
The host bind mounts use `${localEnv:HOME}/.claude` (and `.codex`, `.cursor`, `.ssh`, `.config/git`, `.config/gh`, `.gitconfig`). VS Code resolves `${localEnv:HOME}` by reading its own process env, and Windows doesn't set `HOME` by default — it uses `USERPROFILE`. So the bind mounts can't resolve until you tell Windows to also expose your profile as `HOME`.
|
||||
|
||||
The `initializeCommand` (`node .devcontainer/ensure-host-config-dirs.cjs`) handles this automatically:
|
||||
|
||||
1. **First time you Reopen in Container**, the script detects the missing `HOME`, runs `setx HOME "%USERPROFILE%"` (which writes to your user-level Windows env — no admin needed), prints a one-time setup banner, and exits.
|
||||
2. **Close all VS Code windows** (File → Exit) and reopen. VS Code picks up the new `HOME` at startup.
|
||||
3. **Reopen in Container again.** The script now sees `HOME=C:\Users\<you>`, skips the setup block, creates the bind-mount source dirs, and Docker brings the container up.
|
||||
|
||||
Subsequent rebuilds work normally with no extra steps. The `HOME` env var is set persistently in your Windows user environment, so it'll be there for every future VS Code session (and any other tool that wants `HOME`).
|
||||
|
||||
If you'd rather set it manually before opening the container:
|
||||
|
||||
```powershell
|
||||
setx HOME "%USERPROFILE%"
|
||||
# Close & reopen VS Code
|
||||
```
|
||||
|
||||
### Known trade-offs of Windows-native vs WSL2
|
||||
|
||||
Windows-native works, but Docker Desktop's Windows bind-mount layer has rough edges that WSL2 avoids:
|
||||
|
||||
- **File watchers can miss events.** Vite / jest `--watch` running inside the container watching workspace files mounted from `D:\...` may miss changes — chokidar polling (`CHOKIDAR_USEPOLLING=true`) is the usual workaround.
|
||||
- **`npm install` is 3-5× slower** through the Windows-to-Linux bind-mount translation than on a WSL2-native filesystem.
|
||||
- **Permission edge cases.** The husky `.husky/_/h` EPERM class we hit earlier in this PR is specific to Windows-side bind mounts changing UID ownership between container runs. `post-create.sh` clears the cache defensively to keep this from being fatal, but it's still a real source of friction.
|
||||
|
||||
If you hit any of those and want to migrate to WSL2 later, the steps are below.
|
||||
|
||||
### WSL2 (faster, fewer edge cases)
|
||||
|
||||
To clone and open the repo inside WSL2:
|
||||
|
||||
```bash
|
||||
# 1. Install WSL2 and a Linux distro if you haven't already.
|
||||
wsl --install -d Ubuntu
|
||||
|
||||
# 2. Enter WSL.
|
||||
wsl
|
||||
|
||||
# 3. Clone the repo inside your WSL2 home directory.
|
||||
cd ~
|
||||
git clone https://github.com/abhigyanpatwari/GitNexus.git
|
||||
cd GitNexus
|
||||
|
||||
# 4. Launch VS Code from inside WSL — this opens VS Code attached to the WSL2
|
||||
# filesystem, so `${localEnv:HOME}` resolves to the WSL user's home and
|
||||
# subsequent "Reopen in Container" uses the WSL2-side path.
|
||||
code .
|
||||
```
|
||||
|
||||
Then run **Dev Containers: Reopen in Container**. The workspace will be bind-mounted from `\\wsl$\Ubuntu\home\<user>\GitNexus`, which is fast and gives reliable file-system events. **Make sure Docker Desktop's WSL integration is enabled** for your distro: Docker Desktop → Settings → Resources → WSL Integration → toggle on the distro you cloned into.
|
||||
|
||||
## macOS
|
||||
|
||||
Open the repo folder in VS Code → **Reopen in Container**. The image is multi-arch; on Apple Silicon you'll pull the `linux/arm64` variant automatically.
|
||||
|
||||
## Linux
|
||||
|
||||
Same as macOS — open in VS Code and reopen in container. `updateRemoteUserUID: true` (default) shifts the container's `node` user UID/GID to match your host user, so bind-mounted files stay writable without extra setup.
|
||||
|
||||
## How CLI state flows from your host
|
||||
|
||||
### AI CLIs (Claude Code, Codex, Cursor): copy-once from a read-only host stage + per-container credentials
|
||||
|
||||
The three AI CLIs use a **copy-from-read-only-stage topology**: the host's `~/.<cli>` folders (and `~/.claude-mem`) are mounted **read-only** at `/host/.<cli>`, and `post-create.sh` copies out of them into per-container named volumes. Credentials, identity, and single config files are copied on **every** create; the shareable subdirs (plugins, skills, agents, memory, commands, prompts, rules) are copied **once** on first create and then owned by the container. Nothing is bind-mounted read-write into the host's CLI config, so the container can never modify your host setup. Session sub-paths overlay the config volume via their own named volumes (Docker mount precedence — more specific path wins).
|
||||
|
||||
| Mount | Source | Target | Mode | Purpose |
|
||||
| -------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------- | ------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Container Claude config dir | _named volume_ `claude-config-${devcontainerId}` | `/home/node/.claude` | rw | Per-container credentials + identity |
|
||||
| Container Codex config dir | _named volume_ `codex-config-${devcontainerId}` | `/home/node/.codex` | rw | Per-container credentials |
|
||||
| Container Cursor config dir | _named volume_ `cursor-config-${devcontainerId}` | `/home/node/.cursor` | rw | Per-container credentials |
|
||||
| Container gh config dir | _named volume_ `gh-config-${devcontainerId}` | `/home/node/.config/gh` | rw | Per-container `gh` auth (`hosts.yml`/`config.yml`) seeded from host stage; in-container login persists |
|
||||
| **Claude sessions** (overlay on the config volume) | _named volume_ `…-claude-sessions-${devcontainerId}` | `/home/node/.claude/projects` | rw | `--resume` transcripts; survives the `<cli>-config` volume wipe — see [Session resume](#session-resume-across-container-recreation) |
|
||||
| **Codex sessions** | _named volume_ `…-codex-sessions-${devcontainerId}` | `/home/node/.codex/sessions` | rw | `codex resume` rollouts; SQLite index backfills on recreation |
|
||||
| **Cursor sessions** | _named volumes_ `…-cursor-sessions-${devcontainerId}`, `…-cursor-projects-${devcontainerId}` | `/home/node/.cursor/chats`, `/home/node/.cursor/projects` | rw | `cursor-agent resume` store (best-effort — layout reverse-engineered) |
|
||||
| **claude-mem store** | _named volume_ `claude-mem-${devcontainerId}` | `/home/node/.claude-mem` | rw | claude-mem's SQLite DB + Chroma vector store; **seeded once** from `/host/.claude-mem`, then container-private — see note below |
|
||||
| Host Claude state, read-only stage | `$HOME/.claude` | `/host/.claude` | **read-only** | `post-create.sh` reads credentials + identity from here on container-create |
|
||||
| claude-mem store, read-only stage | `$HOME/.claude-mem` | `/host/.claude-mem` | **read-only** | `post-create.sh` seeds the claude-mem volume from here on first create |
|
||||
| Host Codex state, read-only stage | `$HOME/.codex` | `/host/.codex` | **read-only** | Same purpose for Codex |
|
||||
| Host Cursor state, read-only stage | `$HOME/.cursor` | `/host/.cursor` | **read-only** | Same purpose for Cursor |
|
||||
| **Claude shareable subdirs** | _seeded into the config volume from_ `$HOME/.claude/{plugins/marketplaces,plugins/cache,skills,agents,memory,commands}` | same under `/home/node/.claude/` | n/a (copy) | **Seed-once** copy from the read-only stage; container owns its copy after |
|
||||
| **Codex shareable subdirs** | _seeded from_ `$HOME/.codex/{plugins,prompts,memories,skills}` | same under `/home/node/.codex/` | n/a (copy) | **Seed-once** copy (whole `plugins/` dir — no path-bearing registry inside it) |
|
||||
| **Cursor shareable subdirs** | _seeded from_ `$HOME/.cursor/{plugins/marketplaces,plugins/local,rules,commands,agents,skills}` | same under `/home/node/.cursor/` | n/a (copy) | **Seed-once** copy of the Cursor 2.5 plugin/rules/commands surface |
|
||||
|
||||
**What gets seeded once from the host (copy, not bind):**
|
||||
|
||||
- **Claude**: `plugins/marketplaces`, `plugins/cache`, `skills/`, `agents/`, `memory/`, `commands/`
|
||||
- **Codex**: `plugins/` (whole dir), `prompts/`, `memories/`, `skills/`
|
||||
- **Cursor**: `plugins/marketplaces`, `plugins/local`, `rules/`, `commands/`, `agents/`, `skills/`
|
||||
|
||||
On the **first** container-create, `post-create.sh` copies each of these out of the read-only `/host/.<cli>` stage into the per-container config volume, then writes a `.devcontainer-shareable-seeded` marker. On every later rebuild the marker is present, so the copy is skipped and the container keeps whatever it has accumulated. A plugin/skill/agent you install **inside** the container persists across rebuilds; one you add on the **host** after first create won't appear in the container until you remove the config volume and rebuild (see [§ Rebuild / reset](#rebuild--reset)). Nothing here is writable back to the host — `/plugin marketplace add` inside the container installs into the container's own volume copy, not your host `~/.<cli>/plugins/`.
|
||||
|
||||
**Single config files are copied on container-create, not bind-mounted** — on Docker Desktop Windows a single-file bind is 9p while the named volume is ext4, and atomic config writes (`tmp` → rename onto target) trip EXDEV (this is what caused Codex's `config/batchWrite failed in TUI`). So these are synced from host on rebuild and the container rewrites its own copy until the next rebuild: `settings.json` + `$HOME/.claude.json` (Claude), `config.toml` (Codex), `cli-config.json` + `mcp.json` (Cursor). `hooks.json` (Cursor) is deliberately **not** synced — Cursor hooks execute shell commands, so sharing them would widen the supply-chain attack surface; add it yourself if you want it.
|
||||
|
||||
**Plugin registry files with absolute paths are translated, not copied verbatim** — Claude's `known_marketplaces.json` / `installed_plugins.json` / `plugin-catalog-cache.json` and Cursor's `installed_plugins.json` bake in `C:\Users\…` (Windows) or `/Users/…` (macOS) install paths. `post-create.sh` rewrites those to `/home/node/.<cli>/plugins/…` and writes the result into the named volume, so plugins resolve inside Linux instead of failing with `cache-miss`. This translation is **also seed-once per CLI** — it runs only for a CLI being seeded that create (`translate-plugin-registries.cjs claude cursor`), so it stays consistent with the seed-once `cache/` copy and won't overwrite a plugin you installed inside the container on a later rebuild. Codex needs no translation — its enablement registry is `config.toml` (git URLs + logical keys, no filesystem paths), so its whole `plugins/` dir is copied as-is.
|
||||
|
||||
**What stays per-container (in the named volume) and is synced from host on container-create:**
|
||||
|
||||
- `.credentials.json` (Claude OAuth tokens), `auth.json` (Codex), `cli-config.json` (Cursor) — credentials
|
||||
- `~/.claude/.claude.json` (Claude's identity-only file: `userID`, `oauthAccount`, migration tracking) — kept per-container so logging in via container doesn't overwrite host's stored identity
|
||||
|
||||
`post-create.sh` runs on every container-create, copies host's credentials into the volume if present, then container manages refresh from there. Sync is "always overwrite if host has the file, otherwise leave container alone". So:
|
||||
|
||||
- Host has credentials → container starts logged in.
|
||||
- Host has no credentials → `claude login` / `codex login --device-auth` / `cursor-agent login` inside container; credentials stay in the named volume across rebuilds (volume is keyed by `${devcontainerId}`, stable for the workspace path).
|
||||
- `claude logout` inside container clears volume credentials only; host is untouched.
|
||||
|
||||
**Why CLAUDE_CONFIG_DIR is intentionally NOT set:** Claude's default `~/.claude` matches the named-volume mount target, so the env var added no behavior — but setting it changed which file Claude reads `hasCompletedOnboarding` from. With it set, Claude reads `$CLAUDE_CONFIG_DIR/.claude.json` (the small identity-only file) and re-onboards every container; without it, Claude reads `$HOME/.claude.json` (copied from the read-only `/host/.claude.json` stage on container-create via `seed-claude-config.cjs`, with `hasCompletedOnboarding: true`).
|
||||
|
||||
**Host CLI config is protected from write-through.** The shareable dirs are copied out of a **read-only** stage into the container's own volume, so a compromised npm package in the workspace dep tree — running inside the container — **cannot** write a malicious agent, command, skill, or plugin back to `~/.claude/`, `~/.codex/`, or `~/.cursor/` on the host. The earlier design bind-mounted these read-write and accepted that write-through as the cost of live sync; this design closes it. An even earlier alternative (read-only stage + symlinks) made `/plugin marketplace add` inside the container fail with EROFS; copying into a writable volume avoids that, because the container writes to its own copy rather than a read-only mount. What a compromised dependency can still do is **read** the read-only host stages (`/host/.<cli>`, `/host/.claude-mem`) and the read-only credential mounts and exfiltrate them — there is [no egress firewall yet](#whats-not-included-yet). The cost of the copy model is **divergence**: host edits made after first create don't reach the container until you wipe the config volume and rebuild.
|
||||
|
||||
**Refresh-token divergence between rebuilds.** Container's credentials match host's at container-create time; after that, container manages its own refresh until the next rebuild. Anthropic rotates refresh tokens on every use, so an unattended container that hasn't talked to the API in weeks can hit a silent 401 if the host has refreshed since. Re-run `claude login` inside the container, or rebuild, to recover.
|
||||
|
||||
**claude-mem is seeded once, then container-private.** The [claude-mem](https://github.com/thedotmack/claude-mem) store (`$HOME/.claude-mem` — a multi-GB SQLite DB `claude-mem.db` + `-wal`/`-shm`, plus a Chroma vector store `chroma/chroma.sqlite3` and its HNSW index binaries) is the one shareable-looking folder that is **deliberately not a host bind**, for the same SQLite reason as sessions below: a multi-GB WAL database over the 9p/virtiofs bind risks unreliable `fcntl` locking and corruption — sharply so if claude-mem ran on the host and in the container against the same DB at once. So it gets its own per-container named volume (`claude-mem-${devcontainerId}`), and `post-create.sh` **seeds it once** from the read-only `/host/.claude-mem` stage _only when the volume has no DB yet_. The first container-create copies the host's store in (a one-time copy, possibly several GB); every later rebuild keeps whatever the container accumulated and skips the copy. The container's memory and the host's **diverge from that seed point** — writes do not flow back — which is the price of keeping SQLite off a shared bind. To re-seed from the host's current store, remove the volume (`docker volume rm claude-mem-<id>`) and rebuild. `ensure-host-config-dirs.cjs` creates an empty `~/.claude-mem` on hosts that never installed claude-mem, so the read-only stage bind always resolves; the seed then finds no DB and the container simply starts with empty memory.
|
||||
|
||||
### Session resume across container recreation
|
||||
|
||||
`claude --resume`, `codex resume`, and `cursor-agent resume` all read **local** transcript files. Those live _inside_ each CLI's config dir, which is a per-container named volume — so they already survive an ordinary **Rebuild Container**. What they did _not_ survive were the very things this README tells you to do: `docker volume rm <cli>-config-${devcontainerId}` to force a re-login or clear an `EACCES`, a `${devcontainerId}` change, or a full delete-and-recreate. Each of those drops the config volume and takes your session history with it.
|
||||
|
||||
So the resume/transcript directories get their **own** named volumes (mount group 6 in `devcontainer.json`), keyed like the `node_modules` volumes (`${localWorkspaceFolderBasename}-…-${devcontainerId}`) and mounted _over_ the config volume at the session sub-paths:
|
||||
|
||||
| Resume command | Persisted volume → target | What's stored |
|
||||
| -------------------------------- | --------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `claude --resume` / `--continue` | `…-claude-sessions-…` → `~/.claude/projects` | `<encoded-cwd>/<uuid>.jsonl` transcripts + `sessions-index.json`. Container cwd is always `/workspace`, so only that slice is stored. Pure JSONL/JSON — no SQLite. |
|
||||
| `codex resume` / `resume --last` | `…-codex-sessions-…` → `~/.codex/sessions` | `YYYY/MM/DD/rollout-*.jsonl`. The `state_5.sqlite` thread index stays on the config volume (a single WAL file we don't split out); when it's absent after a recreation, Codex rebuilds it from these rollouts on the next start (a one-time backfill). |
|
||||
| `cursor-agent resume` / `ls` | `…-cursor-sessions-…` → `~/.cursor/chats`; `…-cursor-projects-…` → `~/.cursor/projects` | `chats/{hash}/{uuid}/store.db` (one SQLite db per session, each in its own dir) + `projects/.../agent-transcripts`. cursor-agent's layout is reverse-engineered, so treat this as best-effort. |
|
||||
|
||||
Because these are **separate** volumes from `<cli>-config-${devcontainerId}`, the re-login fix (`docker volume rm claude-config-…`) no longer destroys your sessions — that was the point.
|
||||
|
||||
**Survives:** Rebuild Container, Rebuild Without Cache, a full delete-and-recreate of the container, and the `docker volume rm <cli>-config-…` re-login / `EACCES` fix.
|
||||
|
||||
**Does _not_ survive** (same durability tier as the `node_modules` volumes): `docker volume prune`, a `${devcontainerId}` change (moving the checkout to a new path, or switching between Windows-native and WSL2), or moving to a new machine. To deliberately wipe sessions, remove the session volumes too — see [Rebuild / reset](#rebuild--reset). Two checkouts with the **same folder name** on one host would share session volumes only if they also share a `${devcontainerId}`; they don't, so they stay separate.
|
||||
|
||||
**First rebuild after adopting this, one-time:** if a container created _before_ these volumes existed already had sessions on the config volume (`~/.claude/projects`, `~/.codex/sessions`, …), the new empty session volume mounts _over_ that sub-path and **masks** the old content — same Docker-precedence shadowing described for plugins above. The old sessions are hidden, not deleted. To carry them forward once, copy them out of the config volume into the session volume; or just start fresh — new sessions land on the session volume from then on.
|
||||
|
||||
**Why sessions are container-private and not even seeded from the host.** The shareable config dirs are _seeded once_ from the host (you want your skills/agents/plugins in the container). Sessions are deliberately _not_ seeded and never touch the host, because a transcript can contain anything you pasted or the agent read — API keys, file contents, connection strings. Binding or copying them to/from the host would (a) spill that to host disk, (b) add a write-through surface a compromised dependency can reach (there's still [no egress firewall](#whats-not-included-yet)), and (c) leak _every other project's_ transcripts into the container (Codex `sessions/` and Cursor `chats/` aren't project-scoped). Container-private volumes avoid all three while still surviving recreation. And Claude/Codex transcripts embed the container cwd (`/workspace`), so even if you _did_ bind them to the host, the host CLI wouldn't natively `--resume` them — its encoded-cwd folder differs.
|
||||
|
||||
**Opt in to host-shared sessions anyway.** If you want transcripts visible/portable on the host and accept the trade-offs above, uncomment the host-bind block in `devcontainer.json` (just below the group-6 volumes) and add the matching source dirs to `ensure-host-config-dirs.cjs`'s `DIRS` so Docker can resolve the binds. That block scopes Claude to `/workspace`'s encoded subdir to limit the cross-project leak; the Codex and Cursor stores can't be scoped that way, so they expose every project's transcripts.
|
||||
|
||||
### Other host bind mounts
|
||||
|
||||
| Container path | Host source | Mode | Why |
|
||||
| --------------- | ------------------- | ------------- | -------------------------------------------------------------------------------------- |
|
||||
| `~/.config/git` | `$HOME/.config/git` | **read-only** | XDG-style git config / ignore / attributes |
|
||||
| `~/.ssh` | `$HOME/.ssh` | **read-only** | SSH commit signing + git push over SSH |
|
||||
| `~/.config/gh` | `$HOME/.config/gh` | **copy → volume** | `gh` CLI auth (PR/issue create, checks) — seeded from your host login on create into a per-container volume; in-container `gh auth login` persists across rebuilds and never writes back to the host |
|
||||
| `~/.docker` | `$HOME/.docker` | **read-only** | Container registry auth + buildx config (inert until you add Docker CLI via a Feature) |
|
||||
| `~/.aws` | `$HOME/.aws` | **read-only** | AWS CLI / SDK credentials (forward-compat — empty by default) |
|
||||
| `~/.azure` | `$HOME/.azure` | **read-only** | Azure CLI credentials (forward-compat — empty by default) |
|
||||
|
||||
**Why `ssh`/`aws`/`azure`/`docker` are read-only, and why `gh` is copied into a volume:** `ssh`/`aws`/`azure` are consumed read-only by their clients (the SSH client and the AWS/Azure SDKs only read their credential files), so a one-way mount loses nothing. `docker` _can_ write its own state (`docker login` / buildx write `config.json`), but a read-write host bind would let a compromised in-container dependency rewrite your host `~/.docker/config.json` (point a `credHelper` at an attacker-controlled binary) — a credential-takeover vector. The common case is _reading_ an existing host login, so `docker` stays **read-only**: registry pulls/pushes using your host creds work, only a `docker login` inside the container won't persist back. `gh` used to be read-only for the same reason, but that meant an in-container `gh auth login` had nowhere to write and silently failed. So `gh` now uses the **copy-into-volume** model (the same one the AI-CLI credentials use): the host `~/.config/gh` is a read-only _stage_ at `/host/.config/gh`, and `post-create.sh` copies `hosts.yml`/`config.yml` out of it into the per-container `gh-config` volume on create. The container gets a **writable** copy — `gh auth login` / `gh auth refresh` inside the container now work and persist across rebuilds — while the read-only stage guarantees nothing is ever written back to the host's token. If you want `docker` to behave the same way, give it the same treatment (a `/host/.docker` stage + a docker-config volume + a copy step in `post-create.sh`).
|
||||
|
||||
`~/.gitconfig` is **not** bind-mounted — VS Code's Dev Containers extension auto-copies the host's gitconfig into the container at attach time (this is built-in behavior, not something this devcontainer configures). The bind-mount approach conflicts with that auto-copy mechanism, so we let VS Code own it. The end result is the same: your host's `user.name` / `user.email` are available inside the container.
|
||||
|
||||
If a host source dir doesn't exist when the container is first created, the `initializeCommand` (`node .devcontainer/ensure-host-config-dirs.cjs`) creates it empty — so the bind mount always has a valid source.
|
||||
|
||||
### Per-CLI quirks worth knowing
|
||||
|
||||
- **Claude Code on macOS** stores credentials in the system Keychain, not in `~/.claude/.credentials.json`. The sync silently no-ops; run `claude login` inside the container once and the named volume persists it.
|
||||
- **Codex on macOS / Linux with `cli_auth_credentials_store = "keyring"`** stores auth in the OS keyring (Keychain / Secret Service), so `~/.codex/auth.json` may not exist on host. Same fallback: `codex login --device-auth` inside the container.
|
||||
- **Cursor CLI inside containers** has [known upstream auth issues](https://forum.cursor.com/t/cursor-agent-authentication-issue-inside-docker/143995) — even with a correctly-synced `cli-config.json`, you may need to re-run `cursor-agent login` inside the container.
|
||||
- **Stale named volumes from old rebuilds can carry forward.** If you delete and re-create the same workspace, or if a prior container left interim state with a different `userID`, deleting the named volumes before rebuild guarantees a clean sync: `docker volume rm claude-config-${devcontainerId} codex-config-${devcontainerId} cursor-config-${devcontainerId}` (look them up with `docker volume ls | grep -config-`).
|
||||
- **User-scope MCP servers with absolute host paths won't resolve in-container.** `~/.claude.json` (Claude), `~/.codex/config.toml` (Codex), and `~/.cursor/mcp.json` (Cursor) are copied from host on container-create, so their user-scope `mcpServers` entries come along. But an entry whose `command` is an absolute host path (`C:\tools\foo.exe`, `/usr/local/bin/foo`) points at a binary that doesn't exist in the container — that server silently fails to launch. Only registry/`npx`-based servers (like this repo's `.mcp.json`, which uses `npx -y gitnexus@latest mcp`) and remote/URL servers work unchanged. The path-translation pass only rewrites `*/.<cli>/plugins/*` registry paths, **not** arbitrary `mcpServers` command paths (there's no correct container target for a host-local binary). Install such MCP servers inside the container, or use `npx`/remote ones.
|
||||
- **Host config is seeded once per devcontainer, then diverges — this now applies to everything.** A `mcpServers` entry, setting, plugin, skill, agent, or command you add **on the host after** the container was created is not visible in the container until you remove the config volume and rebuild. Single config files (`mcpServers`, `settings.json`, …) are copy-on-create; the shareable dirs (plugins/skills/agents/memory/commands/prompts/rules) are copy-on-**first**-create (they persist across ordinary rebuilds and aren't even re-copied). Both diverge from the host after their copy. To pull host-side changes in, wipe the relevant volume and rebuild (see [§ Rebuild / reset](#rebuild--reset)).
|
||||
- **Plugins/skills/agents installed in-container persist; they do not reach the host.** A `/plugin marketplace add` (or `codex plugin add`, or a new skill/agent) inside the container writes to the container's own config volume and survives ordinary rebuilds. It never appears on the host — the host dirs are read-only sources, not bind targets. To get a plugin onto the host, install it on the host (then wipe + rebuild to seed it into the container).
|
||||
- **No cross-checkout plugin contention.** Because each container copies plugins into its own per-`${devcontainerId}` volume rather than sharing one host bind source, two containers (or checkouts) installing plugins at the same time no longer interleave git clones/extractions against a shared host dir. Each writes only its own copy.
|
||||
|
||||
### What you still don't have inside the container
|
||||
|
||||
These are commonly-needed CLIs that aren't installed by default — adding them would be follow-up work, not in this PR's scope:
|
||||
|
||||
- **Docker CLI** (for `docker push` / `docker build` from inside the container). Add via `ghcr.io/devcontainers/features/docker-outside-of-docker:1` to the `features` block — `~/.docker/` is already mounted **read-only**, so your host `docker login` state works immediately for pulls/pushes; an in-container `docker login` won't persist to the host (drop `,readonly` on that mount if you need it to).
|
||||
- **AWS CLI / Azure CLI / gcloud / kubectl** — same pattern: add the matching Feature, the host config dirs already flow through.
|
||||
- **Private npm registry auth** (`~/.npmrc`) — you don't have a global one on this host. If you ever start using private packages, add `source=${localEnv:HOME}/.npmrc,target=/home/node/.npmrc,type=bind,readonly` to the mounts.
|
||||
|
||||
That means:
|
||||
|
||||
- **Authentication is shared.** If you're already logged in on the host (`claude login`, `codex login`, `cursor-agent login`, `gh auth login`), you're already logged in inside the container. No second login step.
|
||||
- **Plugins, skills, agents, memory, and commands are seeded from the host once, then container-private.** On first create the container copies your host's plugins/skills/agents/memory/commands (and Codex prompts/memories, Cursor rules) into its own volume. After that they're independent: install or edit inside the container and it stays in the container (persists across rebuilds); add a plugin or agent on the host and the container won't see it until you wipe the config volume and rebuild. Nothing the container does reaches the host. (`settings.json` and the user-scope `~/.claude.json` are copy-on-create the same way; `~/.claude/projects/` is container-local by design.)
|
||||
- **Git identity comes from the host.** Commits from inside the container use your host's `user.name` / `user.email` — VS Code's Dev Containers extension auto-copies your `~/.gitconfig` into the container at attach time. Any XDG-style config under `~/.config/git/` flows through via the read-only bind mount. To change git identity, edit `~/.gitconfig` on the host (container-side `git config --global` writes to a container-local file that's discarded on rebuild).
|
||||
- **SSH keys flow through (read-only).** Push over SSH remotes and SSH commit signing work inside the container using your host keys. The mount is read-only so container code can't exfiltrate or modify private keys — agent-perspective, this means you get git operations but the keys stay vendor-side.
|
||||
- **`gh` auth is shared, and in-container logins persist.** If you're logged in on the host, `gh pr create`, `gh pr checks`, `gh issue create` work inside the container without re-authenticating. If you're not, run `gh auth login` inside the container once — because `gh` config lives in a writable per-container volume (seeded from the host stage), that login persists across rebuilds and never touches the host's token.
|
||||
- **No per-workspace duplication.** All your devcontainers across all your projects see the same host CLI state, just like all your host shells do.
|
||||
|
||||
The bind mount source directories are guaranteed to exist by the `initializeCommand` (`node .devcontainer/ensure-host-config-dirs.cjs`), which runs on the host before container create. It's a Node script (not a shell one-liner) so the same command works on Windows `cmd.exe` and POSIX shells. It creates the top-level bind-mount source dirs — `~/.claude`, `~/.codex`, `~/.cursor`, `~/.claude-mem`, plus `~/.ssh`, `~/.docker`, `~/.aws`, `~/.azure`, `~/.config/{gh,git}`. It deliberately does **not** pre-create the shareable subdirs (skills/agents/plugins/…): those are no longer bind sources (they're copied out of the whole-`~/.<cli>` read-only stage), and pre-creating empty ones would needlessly write into the host of someone who never used that CLI.
|
||||
|
||||
### Trust boundary, concretely
|
||||
|
||||
Host and container share a single trust boundary by design — fine for personal-dev, but the consequence is concrete. Any malicious npm package or `postinstall` script in the workspace dep tree, running inside the container, has direct **read** access to:
|
||||
|
||||
- **Host AI CLI state** — the read-only stage at `/host/.claude`, `/host/.codex`, `/host/.cursor`, `/host/.claude-mem`, which exposes your **entire** host `~/.<cli>` tree (credentials, identity, AND the shareable skills/agents/plugins/memory/commands) for _reading_. The container copies what it needs out of this stage; a compromised dep can read all of it. It is read-only, so none of it can be written back
|
||||
- The **container's own credential snapshots** at `/home/node/.claude/.credentials.json` etc. (copied from host on container-create)
|
||||
- `~/.claude/memory/` / per-project memory (which may contain user-stored secrets if you've used the `/remember` skill)
|
||||
- The **current container's own session transcripts** (`~/.claude/projects`, `~/.codex/sessions`, `~/.cursor/chats`/`projects` — the group-6 volumes), which can hold anything pasted into or read during a session. These are container-private (see one-way note below), so this is read access to _this_ container's sessions only, not the host's or other projects'
|
||||
- Your **`gh` token** (`~/.config/gh`)
|
||||
- Your **SSH private keys** (`~/.ssh/`)
|
||||
- Docker registry tokens in **`~/.docker/config.json`** (if you've `docker login`-ed)
|
||||
- AWS/Azure CLI credentials if you've populated `~/.aws/` or `~/.azure/`
|
||||
|
||||
It does **not** have write-through to the host's CLI config. The shareable dirs are copied out of the read-only stage into the container's own volume, so a compromised in-container dep **cannot** write into your host `~/.claude/{plugins,agents,skills,commands,memory}/`, `~/.codex/{plugins,prompts,memories,skills}/`, or `~/.cursor/{plugins,rules,commands,agents,skills}/`. The persistence vector earlier versions had — drop a malicious auto-loaded agent/command/skill/rule onto the host, have it run in your next **host** session — is closed: there is no writable path from the container to those host folders. (Cursor's `hooks.json` is still additionally withheld from even the _container's_ copy, because hooks fire without an agent invoking them.) The boundary is now one-way for **all** of the host CLI config, not just credentials.
|
||||
|
||||
**What stays one-way (genuinely protected):** everything. Credentials never flow back to host — `.credentials.json` / `auth.json` / `cli-config.json` live only in the per-container named volumes, and the `/host/.<cli>` stage they're copied from is mounted **read-only**, so the snapshot can't be overwritten back. The shareable AI-CLI dirs (skills/agents/plugins/memory/commands/prompts/rules) are now copy-on-create from that same read-only stage, so they have the one-way property too — readable for the copy, never writable back. `~/.ssh`, `~/.config/git`, `~/.aws`, `~/.azure`, and **`~/.docker`** are read-only binds with the same property — a compromised dep can _read_ your registry tokens but cannot _rewrite_ them to hijack your future host auth. **`~/.config/gh`** is now a read-only _stage_ copied into a per-container volume, so it keeps that same one-way property: the container reads it once to seed its own writable copy, and the read-only stage means an in-container `gh auth login` can never overwrite your host token. **Session transcripts** live in per-workspace named volumes (mount group 6) and are never seeded from or written back to the host, and the container can't see any _other_ project's transcripts. The opt-in host-bind block in `devcontainer.json` reverses that for sessions only — enable it only if you accept transcripts on host disk; see [Session resume across container recreation](#session-resume-across-container-recreation).
|
||||
|
||||
**The egress firewall is the key compensating control that is still missing.** It's deferred (see "What's not included (yet)" below), so a compromised package currently has unrestricted outbound network to exfiltrate anything in the read list above. Until it lands, treat that read surface as exposed to any code you run in the container — don't use this devcontainer on a machine whose host credentials you couldn't afford to rotate. The isolated-volume setup below removes host AI-CLI config/credentials from that surface entirely.
|
||||
|
||||
**If a workspace dep is ever found compromised**, rotate credentials at the vendor side — local file deletion is insufficient because tokens may have already left:
|
||||
|
||||
- Anthropic: [console.anthropic.com → Settings → Keys](https://console.anthropic.com/settings/keys), revoke the OAuth session under Account
|
||||
- OpenAI / Codex: [platform.openai.com/api-keys](https://platform.openai.com/api-keys), revoke session under Profile
|
||||
- Cursor: dashboard → Integrations, rotate API key + revoke CLI session
|
||||
- GitHub: `gh auth refresh` or revoke the token at github.com/settings/tokens
|
||||
|
||||
For high-trust enterprise environments where the container should not even be able to **read** host CLI state, remove the three read-only stage binds (`/host/.claude`, `/host/.codex`, `/host/.cursor`) — plus `/host/.claude-mem` and `/host/.claude.json` — from `.devcontainer/devcontainer.json`. With no stage to copy from, `post-create.sh`'s seed and credential-sync steps quietly do nothing (their `[ -f ]` / `[ -d ]` guards), and each devcontainer starts with empty, fully isolated config and credentials (Anthropic's reference pattern). You give up seeding your host setup into the container in exchange for removing host config/credentials from the container's read surface entirely; log in inside each container instead.
|
||||
|
||||
## First-time CLI authentication
|
||||
|
||||
Each CLI works either way:
|
||||
|
||||
- **Log in on host first** → the container picks it up automatically on the next rebuild (`sync_from_host` copies the credential file into the named volume during `post-create.sh`). Host stays the source of truth.
|
||||
- **Log in inside the container** → credentials write to the named volume. They persist across ordinary rebuilds (volume is keyed by `${devcontainerId}`, which is stable for a given workspace folder). The host's credentials are untouched.
|
||||
|
||||
You can mix and match per-CLI. A common setup is "Claude logged in on host, Codex/Cursor logged in inside container".
|
||||
|
||||
### Claude Code
|
||||
|
||||
```bash
|
||||
claude login
|
||||
```
|
||||
|
||||
Opens a browser auth flow. VS Code's port forwarding handles the OAuth callback automatically. After auth, `~/.claude/` is populated and visible from both host and container. The `DISABLE_AUTOUPDATER=1` env var prevents the in-container CLI from auto-updating — rebuild the container to pick up a newer Claude Code.
|
||||
|
||||
### OpenAI Codex CLI
|
||||
|
||||
```bash
|
||||
codex login --device-auth
|
||||
```
|
||||
|
||||
The device-code flow prints a URL and a one-time code. Visit the URL on your host browser, paste the code, and the CLI authenticates without needing a callback listener — this is the most reliable path inside containers. Credentials land in `~/.codex/auth.json` (shared with host).
|
||||
|
||||
`codex login` (browser-callback variant) also works but can be flaky in some headless contexts; prefer `--device-auth`.
|
||||
|
||||
### Cursor CLI
|
||||
|
||||
```bash
|
||||
cursor-agent login
|
||||
```
|
||||
|
||||
Opens a browser auth flow; VS Code's port forwarding handles the callback. Credentials persist in `~/.cursor/cli-config.json` (shared with host).
|
||||
|
||||
Verify any time with `cursor-agent status`.
|
||||
|
||||
## Alternative: API key authentication (CI / headless)
|
||||
|
||||
For non-interactive use (CI runners, automated scripts), all three CLIs accept API keys via env vars:
|
||||
|
||||
| CLI | Env var | Where to get the key |
|
||||
| ----------- | ------------------- | --------------------------------------------- |
|
||||
| Claude Code | `ANTHROPIC_API_KEY` | <https://console.anthropic.com/settings/keys> |
|
||||
| Codex | `OPENAI_API_KEY` | <https://platform.openai.com/api-keys> |
|
||||
| Cursor | `CURSOR_API_KEY` | Cursor dashboard → Integrations |
|
||||
|
||||
These env vars are intentionally **not** injected into the container from the host. `${localEnv:VAR}` resolves an unset host variable to an empty string, and some CLIs (Cursor in particular) treat a set-but-empty key as "use this key" rather than "fall back to stored login" — which would silently break the login flow for everyone who hasn't pre-set the host var.
|
||||
|
||||
To use an API key inside the container, export it in your terminal session:
|
||||
|
||||
```bash
|
||||
export ANTHROPIC_API_KEY=sk-ant-...
|
||||
# or OPENAI_API_KEY, or CURSOR_API_KEY
|
||||
```
|
||||
|
||||
For persistence across container shells, carry the export via your VS Code [dotfiles repository](https://code.visualstudio.com/docs/devcontainers/containers#_personalizing-with-dotfile-repositories). VS Code clones the dotfiles repo into the container on attach and runs your install command, so the export lands in `~/.bashrc` / `~/.zshrc` per your own setup — and your API keys stay out of this repo's committed `devcontainer.json`.
|
||||
|
||||
A non-empty API key env var takes precedence over stored login credentials for each CLI.
|
||||
|
||||
## Port forwarding
|
||||
|
||||
| Port | Service | Notes |
|
||||
| ------ | -------------------------------- | ------------------------------------------------------------------------------------------------------ |
|
||||
| `5173` | Vite dev server (`gitnexus-web`) | Auto-forwarded with notification |
|
||||
| `4747` | `gitnexus serve` HTTP API | **Must not be remapped** — `gitnexus-web` hardcodes `http://localhost:4747` as the default backend URL |
|
||||
| `4173` | Static web (Vite preview) | Silently forwarded |
|
||||
|
||||
VS Code's Ports panel shows forwarded ports once their listener starts.
|
||||
|
||||
## Known gotchas
|
||||
|
||||
- **LadybugDB integration tests may fail in containers** (file-locking, `AGENTS.md` § Testing). Default to `npm run test:unit` inside the container; run integration tests on the host. Tracking issue: documented as a known limitation.
|
||||
- **Single-writer LadybugDB constraint** (`GUARDRAILS.md` § LadybugDB lock). Don't run `gitnexus analyze` on the host and inside the container against the same `.gitnexus/` directory simultaneously — the second writer will get `database busy`.
|
||||
- **Native grammar builds add ~30s to first install.** Tree-sitter Dart/Proto/Swift grammars build during `gitnexus`'s `postinstall`. To skip them (loses parsing for those three languages), set `GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1` in your shell or add it to `remoteEnv` and rebuild.
|
||||
- **`tree-sitter-kotlin` warnings on install** are expected (per `AGENTS.md`). Ignore them.
|
||||
- **`.mcp.json` works inside the container**: `npx -y gitnexus@latest mcp` resolves cleanly because npm registry is reachable and the workspace bind mount exposes the same `.mcp.json` the host sees.
|
||||
- **Husky pre-commit fires inside the container** without extra setup. The root `npm install` (run automatically in `postCreateCommand`) installs the hook via `package.json` `prepare`.
|
||||
|
||||
## Rebuild / reset
|
||||
|
||||
- **Rebuild Container** (Command Palette) — re-runs the Dockerfile build and `postCreateCommand` against the existing named volumes (auth, history, **and sessions** persist).
|
||||
- **Rebuild Container Without Cache** — fresh image layers, same volumes.
|
||||
- **To force a re-login / clear an `EACCES`** — remove the per-container _config_ volumes and rebuild. As of the session-volume change this **no longer drops your `--resume` history** (sessions are on separate volumes — see [Session resume](#session-resume-across-container-recreation)):
|
||||
```bash
|
||||
docker volume ls | grep -- -config- # the credential / identity volumes
|
||||
docker volume rm claude-config-<id> codex-config-<id> cursor-config-<id> gh-config-<id>
|
||||
```
|
||||
⚠️ Since the shareable dirs are now seeded into the config volume (not bind-mounted), wiping `<cli>-config` **also discards any plugin/skill/agent/command you installed _inside_ the container** and re-seeds those dirs from the host on the next rebuild. That is the intended way to pull host-side config changes in, but if you have in-container-only plugins you want to keep, reinstall them after the rebuild (or install them on the host first so the re-seed brings them along).
|
||||
- **To also wipe session history** (a true clean slate) — remove the session volumes too (`<name>` is your workspace folder name):
|
||||
```bash
|
||||
docker volume ls | grep -E -- '-(sessions|cursor-projects)-' # the group-6 volumes
|
||||
docker volume rm <name>-claude-sessions-<id> <name>-codex-sessions-<id> \
|
||||
<name>-cursor-sessions-<id> <name>-cursor-projects-<id>
|
||||
```
|
||||
Then rebuild.
|
||||
- **To re-seed claude-mem from the host** (the container's memory has diverged and you want the host's current store back) — remove the claude-mem volume and rebuild; `post-create.sh` copies the host store in again on the next create:
|
||||
```bash
|
||||
docker volume rm claude-mem-<id>
|
||||
```
|
||||
|
||||
## Bumping CLI versions
|
||||
|
||||
Bump the version pins in `.devcontainer/devcontainer.json` `build.args` and rebuild — all three are real, fail-loud pins. Claude Code installs via `npm install -g @anthropic-ai/claude-code@${CLAUDE_CODE_VERSION}` and Codex via `npm install -g @openai/codex@${CODEX_VERSION}`. **Cursor is pinned too:** bump `CURSOR_VERSION` **and** both `CURSOR_SHA256_X64` / `CURSOR_SHA256_ARM64` together — the Dockerfile downloads the pinned `downloads.cursor.com/lab/<version>/linux/<arch>/agent-cli-package.tar.gz` artifact directly (no remote install script) and fails the build on a sha256 mismatch. Re-hash each arch with `curl -fSL <url> | sha256sum`. To stop Cursor from auto-updating in the running container, don't call `cursor-agent update`.
|
||||
|
||||
## What's not included (yet)
|
||||
|
||||
- **Egress firewall — the most important hardening still outstanding.** The original plan included an opt-in iptables/ipset firewall adapted from Anthropic's reference devcontainer. It was deferred to a follow-up PR — `runArgs` is static in `devcontainer.json`, so toggling NET_ADMIN/NET_RAW capabilities cleanly requires either a separate `devcontainer-firewall.json` profile or an `initializeCommand`-generated overlay. Until it lands, the read surface in [§ Trust boundary](#trust-boundary-concretely) has no network containment — anything readable can be exfiltrated. Track at the project's issue tracker if you need this.
|
||||
- **Codespaces tuning.** The current config works in Codespaces incidentally (no privileged capabilities, no host-mount assumptions), but isn't actively tested there.
|
||||
- **Playwright e2e support.** `gitnexus-web`'s `npm run test:e2e` needs Chromium libs that the base image doesn't ship. Use the host for e2e until a Playwright layer is added.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
| Symptom | Likely cause | Fix |
|
||||
| -------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GitNexus devcontainer one-time Windows setup` banner from `initializeCommand` | First-time Windows-native Reopen-in-Container; `HOME` env var was missing | The script just ran `setx HOME "%USERPROFILE%"` for you. Close ALL VS Code windows (File → Exit) and reopen — see [Windows 11 setup](#windows-11-setup) |
|
||||
| `bind source path does not exist: /.claude` (or similar) from Docker | Windows-native `HOME` env var is still missing even after one rebuild — `setx` may have failed or VS Code wasn't fully restarted | Run `setx HOME "%USERPROFILE%"` in a Windows shell manually, fully exit VS Code (check Task Manager that no `Code.exe` remains), reopen |
|
||||
| `EACCES` / `EPERM` writing into `~/.claude`, `~/.codex`, or `~/.cursor` inside the container | Stale state from a previous container with a different effective UID | Move the affected dir aside and let the CLI rebuild it (`mv ~/.claude ~/.claude.bak` and log in again). Long-term: WSL2 setup, which doesn't hit this class of issue |
|
||||
| `EPERM: operation not permitted, copyfile ... '.husky/_/h'` in `postCreateCommand` | Leftover `.husky/_/` from a previous container run on a Windows-side bind mount | `post-create.sh` already runs `rm -rf .husky/_` defensively. If you hit this on an older config, delete `.husky/_/` on the host and rebuild. Long-term: clone in WSL2 |
|
||||
| Vite never hot-reloads | Repo cloned on Windows side, not WSL2 | Re-clone inside WSL2 |
|
||||
| `gitnexus-web` can't reach the backend | `4747` was remapped or backend isn't running | Verify the Ports panel shows `4747` forwarded with no remap; start the backend with `cd gitnexus && npx gitnexus serve` |
|
||||
| `npm install` fails on tree-sitter-swift / proto / dart | Native build toolchain missing | This shouldn't happen in the devcontainer — verify the apt layer installed `python3 make g++`. If iterating, set `GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1` to skip the vendored grammars |
|
||||
| Integration tests fail with `database busy` | LadybugDB single-writer constraint | Don't run host-side `gitnexus analyze` while the container is also analyzing the same repo; choose one writer |
|
||||
| API key env vars not visible inside the container | They are intentionally not auto-propagated from the host (so an empty/stale host var can't silently break `*-login` for everyone else) | `export ANTHROPIC_API_KEY=...` / `OPENAI_API_KEY=...` / `CURSOR_API_KEY=...` inside the container shell, or carry it via your VS Code [dotfiles repo](https://code.visualstudio.com/docs/devcontainers/containers#_personalizing-with-dotfile-repositories) for persistence |
|
||||
| `git commit` produces commits with empty author | `~/.gitconfig` is missing or empty on the host (VS Code's auto-copy had nothing to copy) | Set `git config --global user.name "Your Name"` and `git config --global user.email "you@example.com"` from the host shell, then rebuild the container |
|
||||
| `gh: not logged in` inside the container | Not logged in on the host (nothing to seed), or the `gh-config` volume is empty | Just run `gh auth login` **inside the container** — `gh` config lives in a writable per-container volume, so the login persists across rebuilds. (Logging in on the host instead also works: it seeds in on the next container create.) |
|
||||
@@ -0,0 +1,9 @@
|
||||
{
|
||||
"features": {
|
||||
"ghcr.io/devcontainers/features/github-cli:1": {
|
||||
"version": "1.1.0",
|
||||
"resolved": "ghcr.io/devcontainers/features/github-cli@sha256:d22f50b70ed75339b4eed1ba9ecde3a1791f90e88d37936517e3bace0bbad671",
|
||||
"integrity": "sha256:d22f50b70ed75339b4eed1ba9ecde3a1791f90e88d37936517e3bace0bbad671"
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,396 @@
|
||||
// Devcontainer for GitNexus. It pre-installs Claude Code, the OpenAI Codex
|
||||
// CLI, and the Cursor CLI, plus the Node.js native build chain. It works on
|
||||
// macOS, Linux, Windows via WSL2, and Windows native. Windows native needs a
|
||||
// one-time HOME setup. That setup runs automatically via initializeCommand.
|
||||
// See .devcontainer/README.md § Windows 11 setup. Open it with the VS Code
|
||||
// Dev Containers extension.
|
||||
//
|
||||
// For first-time setup, auth flows, and troubleshooting, see
|
||||
// .devcontainer/README.md.
|
||||
{
|
||||
"name": "GitNexus AI CLI Devcontainer",
|
||||
|
||||
"build": {
|
||||
"dockerfile": "Dockerfile",
|
||||
"context": ".",
|
||||
"args": {
|
||||
"CLAUDE_CODE_VERSION": "2.1.156",
|
||||
"CODEX_VERSION": "0.134.0",
|
||||
// Cursor: a pinned version plus one sha256 hash per CPU arch. The
|
||||
// Dockerfile checks the tarball against the hash at build time, so it
|
||||
// never runs a remote install script. Bump all three values together.
|
||||
// Re-hash each arch with:
|
||||
// curl -fSL https://downloads.cursor.com/lab/<ver>/linux/<x64|arm64>/agent-cli-package.tar.gz | sha256sum
|
||||
"CURSOR_VERSION": "2026.05.28-a70ca7c",
|
||||
"CURSOR_SHA256_X64": "7f8b6a09393e0b84b288cc6952b292fc98d15775f644cc01b0b9aa4f04b268df",
|
||||
"CURSOR_SHA256_ARM64": "05a0ab361e038729aba25fe7f407531b3e8432912e499d0bffdf1dda0e7833e9",
|
||||
// Bun: pinned by version. Installed by the official bun.sh/install
|
||||
// script, which accepts the release tag as its first positional arg
|
||||
// (`bash -s bun-vX.Y.Z`). UNLIKE Cursor, the install path runs an
|
||||
// unverified remote script — chosen at request time for simplicity.
|
||||
// To bump: pick a tag from github.com/oven-sh/bun/releases and update
|
||||
// this value.
|
||||
"BUN_VERSION": "1.3.14",
|
||||
"TZ": "${localEnv:TZ:UTC}"
|
||||
}
|
||||
},
|
||||
|
||||
// Runs on the HOST, not the container, before the container is created. We
|
||||
// write it as a single string on purpose. The spec treats the single-string
|
||||
// form as one command that each OS runs its own way. The object form means
|
||||
// "named parallel tasks", not per-OS dispatch. We run it with Node so the
|
||||
// same command works in cmd.exe on Windows and in bash/zsh on Linux, macOS,
|
||||
// and WSL. The script reads `os.homedir()`, which respects $HOME on
|
||||
// Linux/macOS and %USERPROFILE% on Windows. It then creates the host-side
|
||||
// bind mount source folders, and it is safe to re-run. Host prerequisite:
|
||||
// Node on PATH. That is the only host-side tool needed beyond Docker Desktop
|
||||
// and the VS Code Dev Containers extension.
|
||||
"initializeCommand": "node .devcontainer/ensure-host-config-dirs.cjs",
|
||||
|
||||
"features": {
|
||||
"ghcr.io/devcontainers/features/github-cli:1": {}
|
||||
},
|
||||
|
||||
"remoteUser": "node",
|
||||
"updateRemoteUserUID": true,
|
||||
|
||||
"workspaceMount": "source=${localWorkspaceFolder},target=/workspace,type=bind,consistency=delegated",
|
||||
"workspaceFolder": "/workspace",
|
||||
|
||||
// Mount topology, by group:
|
||||
//
|
||||
// 1. AI CLI host config — a READ-ONLY stage at /host/.<cli>. On container-
|
||||
// create, `post-create.sh` COPIES out of it: credentials, identity, and
|
||||
// single config files (always), plus the shareable subfolders (Claude
|
||||
// plugins/skills/agents/memory/commands; Codex plugins/prompts/memories/
|
||||
// skills; Cursor plugins/rules/commands/agents/skills) ONCE on first create.
|
||||
// Everything copied lands in the per-container named volume (role 2). It is
|
||||
// read-only so a container process can NEVER write back to the host — there
|
||||
// is no read-write bind into the host's CLI config at all. This protects the
|
||||
// host's on-disk setup: a compromised in-container dependency cannot drop a
|
||||
// skill, agent, command, or plugin onto the host for the next host session
|
||||
// to load. The cost is that host and container DIVERGE after the first
|
||||
// create — host edits don't reach the container until you wipe the config
|
||||
// volume and rebuild. See README § "Trust boundary, concretely".
|
||||
//
|
||||
// 2. AI CLI container config — one named volume per devcontainer. CODEX_HOME
|
||||
// points here. CLAUDE_CONFIG_DIR is left unset on purpose, so it resolves
|
||||
// to the default ~/.claude, which is this same path. Credentials,
|
||||
// identity, and single config files (.credentials.json,
|
||||
// ~/.claude/.claude.json, settings.json, config.toml, cli-config.json,
|
||||
// mcp.json) live here with correct Linux permissions. They are NOT
|
||||
// bind-mounted, because single-file binds break on Docker Desktop Windows
|
||||
// (the EXDEV error — see the SINGLE-FILE note below). Container-managed
|
||||
// state (sessions, history, caches, IDE locks) stays separate per
|
||||
// devcontainer. So two GitNexus checkouts on the same host can't corrupt
|
||||
// each other.
|
||||
//
|
||||
// 3. Other host config — read-only bind mounts for credential and identity
|
||||
// folders that lack the permission-flattening and onboarding-state
|
||||
// complications Claude Code has (ssh, aws, azure, git config, plus gh and
|
||||
// docker). gh and docker are read-only so a compromised dependency can't
|
||||
// rewrite the GitHub token or the Docker credHelper. See the inline note
|
||||
// at those mounts. `~/.gitconfig` is not mounted here. VS Code auto-copies
|
||||
// it separately.
|
||||
//
|
||||
// 4. Per-instance state — scoped by `${devcontainerId}`: shell history and
|
||||
// the npm cache. These survive rebuilds and stay separate between sibling
|
||||
// instances.
|
||||
//
|
||||
// 5. Per-workspace-name AND per-instance state — the workspace `node_modules`
|
||||
// volumes use both `${localWorkspaceFolderBasename}` (so you can spot them
|
||||
// in `docker volume ls`) and `${devcontainerId}` (so sibling instances of
|
||||
// the same repo never collide). This keeps tree-sitter native binaries and
|
||||
// onnxruntime off the workspace bind mount, which is faster on Windows and
|
||||
// macOS.
|
||||
//
|
||||
// 6. Per-workspace session state — dedicated named volumes for each CLI's
|
||||
// resume/transcript dirs (Claude projects/, Codex sessions/, Cursor chats/
|
||||
// + projects/). Same `${localWorkspaceFolderBasename}` + `${devcontainerId}`
|
||||
// keying as group 5, but SEPARATE volumes from the group-2 config volumes.
|
||||
// That separation is the point: the `docker volume rm <cli>-config-*`
|
||||
// re-login / EACCES fix (README § Rebuild/reset) no longer wipes sessions,
|
||||
// so `claude --resume`, `codex resume`, and `cursor-agent resume` survive a
|
||||
// rebuild, a full delete-and-recreate, AND that wipe. They overlay the
|
||||
// config volume at the session sub-paths (Docker precedence: more specific
|
||||
// path wins). Container-private by design — transcripts can hold pasted
|
||||
// secrets, and like the group-1 config (read-only stage, copy-once) these
|
||||
// add NO host write-through surface and leak no other projects' transcripts.
|
||||
// They do
|
||||
// NOT survive `docker volume prune`, a `${devcontainerId}` change (moving
|
||||
// the checkout, WSL vs native), or a new machine — same tier as group 5.
|
||||
// To make sessions host-visible/portable instead, see the commented
|
||||
// host-bind block below and README § "Session resume across recreation".
|
||||
"mounts": [
|
||||
// One named volume per container for credentials and identity state. Each
|
||||
// CLI's real `~/.<cli>` config folder lives in a volume. That keeps
|
||||
// credentials (with correct Linux 600 permissions) and per-container
|
||||
// session state separate from the host. Logging in inside the container
|
||||
// and logging in on the host are independent. The bind mounts BELOW these
|
||||
// volumes override the volume's contents at the paths they cover. Docker
|
||||
// mount precedence is: the more specific path wins.
|
||||
"source=claude-config-${devcontainerId},target=/home/node/.claude,type=volume",
|
||||
"source=codex-config-${devcontainerId},target=/home/node/.codex,type=volume",
|
||||
"source=cursor-config-${devcontainerId},target=/home/node/.cursor,type=volume",
|
||||
|
||||
// gh CLI config as a per-container named volume, same model as the AI CLI
|
||||
// configs above: post-create.sh COPIES hosts.yml/config.yml out of the
|
||||
// read-only /host/.config/gh stage into this volume on container-create.
|
||||
// The container then owns a WRITABLE copy, so `gh auth login` /
|
||||
// `gh auth refresh` run INSIDE the container persist across rebuilds — and
|
||||
// still never write back to the host (the stage is read-only). If the host
|
||||
// is logged in, that login seeds in; if not, an in-container login sticks.
|
||||
"source=gh-config-${devcontainerId},target=/home/node/.config/gh,type=volume",
|
||||
|
||||
// claude-mem store. UNLIKE the shareable dirs below (skills/agents/memory),
|
||||
// this is NOT a host bind. $HOME/.claude-mem is a large, multi-GB SQLite +
|
||||
// Chroma vector store (claude-mem.db + -wal/-shm, chroma/chroma.sqlite3, HNSW
|
||||
// index binaries). A read-write host bind would (a) push every byte over the
|
||||
// 9p/virtiofs share, and (b) expose those SQLite WAL files to unreliable
|
||||
// fcntl locking across that boundary — with a real corruption risk if
|
||||
// claude-mem ran on the host and in the container against the same DB at
|
||||
// once. So it gets its OWN per-container named volume here, same durability
|
||||
// tier as the config volumes (survives Rebuild Container and a
|
||||
// delete-and-recreate; keyed by ${devcontainerId}). post-create.sh SEEDS it
|
||||
// ONCE from the /host/.claude-mem read-only stage when the volume is empty,
|
||||
// then the container owns its copy — rebuilds never clobber it, and changes
|
||||
// do NOT flow back to the host. (Container and host memory diverge from the
|
||||
// seed point on; that is the price of safe SQLite.) Removed by the same
|
||||
// `docker volume rm` reset flow as the other volumes — see README.
|
||||
"source=claude-mem-${devcontainerId},target=/home/node/.claude-mem,type=volume",
|
||||
|
||||
// Per-workspace SESSION volumes (mount group 6). These OVERLAY the config
|
||||
// volumes above at the session sub-paths so "resume my last session"
|
||||
// survives container recreation the way the seeded config dirs do.
|
||||
// They are SEPARATE volumes from <cli>-config-${devcontainerId}, so the
|
||||
// README's `docker volume rm <cli>-config-${devcontainerId}` re-login fix
|
||||
// does not touch them. post-create.sh chowns each one explicitly (its
|
||||
// `find -xdev` stops at the config-volume filesystem boundary and won't
|
||||
// descend into these).
|
||||
//
|
||||
// KEEP IN SYNC: if you add/rename/remove a session sub-path, update all
|
||||
// three places that name it — (1) the mount line here, (2) the DIRS array
|
||||
// in post-create.sh (so its chown covers the volume), and (3) the mount
|
||||
// table + Session-resume section in README.md.
|
||||
//
|
||||
// Claude: projects/ holds <encoded-cwd>/<uuid>.jsonl transcripts plus the
|
||||
// sessions-index.json that the `/resume` picker reads. The container cwd is
|
||||
// always /workspace (encodes to the `-workspace` subdir), so this is the
|
||||
// container's own slice only. Pure JSONL/JSON — no SQLite/WAL, so a volume
|
||||
// here is clean. `claude --resume` / `--continue` read straight from it.
|
||||
"source=${localWorkspaceFolderBasename}-claude-sessions-${devcontainerId},target=/home/node/.claude/projects,type=volume",
|
||||
// Codex: sessions/ holds YYYY/MM/DD/rollout-*.jsonl transcripts. The thread
|
||||
// index (state_5.sqlite + -wal/-shm) stays at the ~/.codex root on the
|
||||
// config volume — it is a single WAL file we must NOT split onto a host
|
||||
// bind. On a recreation that drops the config volume, that index is cleanly
|
||||
// absent and Codex rebuilds it from these rollout files on the next start
|
||||
// (backfill). See README for the one-time-rebuild and corruption caveats.
|
||||
"source=${localWorkspaceFolderBasename}-codex-sessions-${devcontainerId},target=/home/node/.codex/sessions,type=volume",
|
||||
// Cursor: chats/{hash}/{uuid}/store.db is one SQLite db per session, each in
|
||||
// its own leaf dir — a DIRECTORY volume keeps each db beside its -wal/-shm
|
||||
// sidecar, so there is no cross-filesystem single-file hazard. projects/
|
||||
// (agent-transcripts) is added too. cursor-agent's on-disk layout is
|
||||
// community-reverse-engineered (LOW confidence), so this is best-effort;
|
||||
// keeping it container-private means a wrong guess can't corrupt host state.
|
||||
"source=${localWorkspaceFolderBasename}-cursor-sessions-${devcontainerId},target=/home/node/.cursor/chats,type=volume",
|
||||
"source=${localWorkspaceFolderBasename}-cursor-projects-${devcontainerId},target=/home/node/.cursor/projects,type=volume",
|
||||
//
|
||||
// OPT-IN: host-shared sessions (like the plugin/skill binds). Uncomment to
|
||||
// put transcripts on the host — fully visible and portable, but they then
|
||||
// land on host disk and become a write-through surface for a compromised
|
||||
// in-container dependency, and the whole-dir binds expose OTHER projects'
|
||||
// transcripts to the container. Claude is scoped to /workspace's encoded
|
||||
// subdir to limit that leak; Codex/Cursor stores are not project-scoped, so
|
||||
// they expose every project. If you enable these, also add the matching
|
||||
// source dirs to ensure-host-config-dirs.cjs — to its DIRS array (these are
|
||||
// directory binds), not FILES (which is only for single-file bind sources
|
||||
// like ~/.claude.json) — so Docker can resolve the binds. Read README
|
||||
// § "Session resume across recreation" first.
|
||||
// "source=${localEnv:HOME}/.claude/projects/-workspace,target=/home/node/.claude/projects/-workspace,type=bind",
|
||||
// "source=${localEnv:HOME}/.codex/sessions,target=/home/node/.codex/sessions,type=bind",
|
||||
// "source=${localEnv:HOME}/.cursor/chats,target=/home/node/.cursor/chats,type=bind",
|
||||
// "source=${localEnv:HOME}/.cursor/projects,target=/home/node/.cursor/projects,type=bind",
|
||||
|
||||
// Read-only host stage that post-create.sh copies FROM on container-create.
|
||||
// It is read-only so a container process can never write back to host CLI
|
||||
// state — that write-back is the attack vector we block. post-create.sh
|
||||
// reads two kinds of thing from here: (a) the credential + identity files
|
||||
// (copied into the volume always), and (b) the shareable dirs — skills,
|
||||
// agents, plugins, memory, commands, prompts, rules — which it copies into
|
||||
// the volume ONCE on first create (see step 3/4). Nothing here is bound
|
||||
// read-write into the container, so the host's on-disk setup is protected.
|
||||
"source=${localEnv:HOME}/.claude,target=/host/.claude,type=bind,readonly",
|
||||
"source=${localEnv:HOME}/.codex,target=/host/.codex,type=bind,readonly",
|
||||
"source=${localEnv:HOME}/.cursor,target=/host/.cursor,type=bind,readonly",
|
||||
// Read-only host stage for the claude-mem store. post-create.sh COPIES it
|
||||
// into the claude-mem named volume on first create (seed-once). Read-only so
|
||||
// the container can never write back to the host's live DB — the seed is a
|
||||
// one-way snapshot. ensure-host-config-dirs.cjs creates ~/.claude-mem on the
|
||||
// host so this bind resolves even when claude-mem was never installed there.
|
||||
"source=${localEnv:HOME}/.claude-mem,target=/host/.claude-mem,type=bind,readonly",
|
||||
|
||||
// NO read-write bind mounts for the shareable subfolders. They USED to be
|
||||
// bound here (Claude skills/agents/memory/commands/plugins; Codex plugins/
|
||||
// prompts/memories/skills; Cursor rules/commands/agents/skills/plugins) so
|
||||
// host and container shared one copy both ways. That bind was a write-through
|
||||
// hole: a compromised in-container dependency could drop a malicious skill,
|
||||
// agent, command, or plugin straight onto the host, which the next HOST
|
||||
// session would auto-load. To protect the host's on-disk setup, these are
|
||||
// now COPIED once from the read-only /host/.<cli> stage into the per-container
|
||||
// named volume by post-create.sh (step 3/4), exactly like claude-mem and the
|
||||
// session volumes. Trade-offs of the copy model:
|
||||
// - The container gets its OWN writable copy and can never write back to
|
||||
// the host. Host setup is protected.
|
||||
// - It is seed-ONCE: host edits made after first create don't reach the
|
||||
// container until you remove the config volume and rebuild. Container
|
||||
// edits persist across rebuilds. (See README § Rebuild/reset to re-seed.)
|
||||
// - The plugin REGISTRY JSONs (Claude known_marketplaces.json /
|
||||
// installed_plugins.json / plugin-catalog-cache.json; Cursor
|
||||
// installed_plugins.json) carry absolute OS-native paths, so they can't
|
||||
// be copied verbatim — post-create.sh translates their paths to the
|
||||
// container's Linux paths, also seed-once, alongside the cache/ copy so
|
||||
// the two stay consistent. Codex needs no translation (config.toml holds
|
||||
// git URLs, not paths), so its whole plugins/ dir is copied as-is.
|
||||
// - The old read-only-stage-plus-symlink design failed `/plugin marketplace
|
||||
// add` in the container with EROFS; copy-into-a-writable-volume avoids
|
||||
// that — the container writes to its own copy, not a read-only mount.
|
||||
//
|
||||
// SINGLE-FILE binds for settings.json, .claude.json, and config.toml are
|
||||
// deliberately ABSENT. On Docker Desktop Windows the named volume sits on
|
||||
// one filesystem (ext4, /dev/sdd) and a single-file bind from the host sits
|
||||
// on another (the 9p drvfs share). Apps save a config by writing `foo.tmp`
|
||||
// and renaming it over `foo`. That rename can't cross filesystems: it hits
|
||||
// the EXDEV error and fails with `Device or resource busy` or `inter-device
|
||||
// move failed`. Codex's TUI shows this as "config/batchWrite failed in
|
||||
// TUI"; Claude just silently loses the write the same way. Instead, we use
|
||||
// a read-only host stage at /host/.claude, and post-create.sh copies these
|
||||
// files into the named volume on every container-create. Host changes show
|
||||
// up on the next rebuild. Container changes stay inside the container until
|
||||
// a rebuild.
|
||||
"source=${localEnv:HOME}/.claude.json,target=/host/.claude.json,type=bind,readonly",
|
||||
"source=${localEnv:HOME}/.config/git,target=/home/node/.config/git,type=bind,readonly",
|
||||
"source=${localEnv:HOME}/.ssh,target=/home/node/.ssh,type=bind,readonly",
|
||||
// gh uses the COPY-INTO-VOLUME model (read-only host stage at
|
||||
// /host/.config/gh + the gh-config named volume above). post-create.sh seeds
|
||||
// hosts.yml/config.yml from this stage into the volume on create, so the
|
||||
// container has a writable copy: an in-container `gh auth login` persists
|
||||
// across rebuilds, and nothing is ever written back to the host because this
|
||||
// stage is read-only. docker stays a direct READ-ONLY bind: the container
|
||||
// reads your EXISTING host login (the common case), and a compromised
|
||||
// in-container dependency can't rewrite ~/.docker/config.json (the registry
|
||||
// credHelper, which points at a binary). A `docker login` run inside the
|
||||
// container won't persist back to the host — re-run it on the host, or give
|
||||
// docker the same copy-into-volume treatment as gh. See README § Trust boundary.
|
||||
"source=${localEnv:HOME}/.config/gh,target=/host/.config/gh,type=bind,readonly",
|
||||
"source=${localEnv:HOME}/.docker,target=/home/node/.docker,type=bind,readonly",
|
||||
"source=${localEnv:HOME}/.aws,target=/home/node/.aws,type=bind,readonly",
|
||||
"source=${localEnv:HOME}/.azure,target=/home/node/.azure,type=bind,readonly",
|
||||
"source=commandhistory-${devcontainerId},target=/commandhistory,type=volume",
|
||||
"source=npm-cache-${devcontainerId},target=/home/node/.npm,type=volume",
|
||||
"source=${localWorkspaceFolderBasename}-root-node-modules-${devcontainerId},target=/workspace/node_modules,type=volume",
|
||||
"source=${localWorkspaceFolderBasename}-gitnexus-node-modules-${devcontainerId},target=/workspace/gitnexus/node_modules,type=volume",
|
||||
"source=${localWorkspaceFolderBasename}-gitnexus-web-node-modules-${devcontainerId},target=/workspace/gitnexus-web/node_modules,type=volume",
|
||||
"source=${localWorkspaceFolderBasename}-gitnexus-shared-node-modules-${devcontainerId},target=/workspace/gitnexus-shared/node_modules,type=volume"
|
||||
],
|
||||
|
||||
// Interactive login is the default way to authenticate for all three CLIs.
|
||||
// Credentials live in the per-container named volumes (claude-config,
|
||||
// codex-config, cursor-config), NOT in the host bind mounts. They are copied
|
||||
// from the read-only /host/.<cli> stage into the volume on container-create.
|
||||
// Single-file binds would break on Docker Desktop Windows (the EXDEV error).
|
||||
// Shareable content (plugins, skills, agents, memory, commands) is NOT bound
|
||||
// read-write — it is copied once from the read-only /host/.<cli> stage into
|
||||
// the volume on first create, so the host's on-disk setup stays protected.
|
||||
// API keys (ANTHROPIC_API_KEY, OPENAI_API_KEY, CURSOR_API_KEY) are NOT
|
||||
// injected via containerEnv. `${localEnv:VAR}` turns an unset host var into
|
||||
// an empty string. Cursor in particular treats `CURSOR_API_KEY=""` as "use
|
||||
// this empty key" instead of "fall back to the stored login", which would
|
||||
// silently break `cursor-agent login`. If you need API-key auth, `export`
|
||||
// the var in your container shell, or carry it in your VS Code dotfiles repo
|
||||
// (see .devcontainer/README.md).
|
||||
// CLAUDE_CONFIG_DIR is left unset on purpose. The Claude default is
|
||||
// `$HOME/.claude` (= `/home/node/.claude`), which is exactly where the
|
||||
// claude-config named volume mounts. Setting the env var would change which
|
||||
// file Claude reads `hasCompletedOnboarding` from. With the var set, Claude
|
||||
// reads `$CLAUDE_CONFIG_DIR/.claude.json`, the small identity file. Without
|
||||
// it, Claude reads `$HOME/.claude.json`, the big onboarding-state file that
|
||||
// actually holds `hasCompletedOnboarding`, the user-scope MCP config, and
|
||||
// per-project trust. Leaving the var unset matches host behavior. It also
|
||||
// lets post-create.sh's sync of `$HOME/.claude.json` skip the setup wizard
|
||||
// on every container-create.
|
||||
//
|
||||
// CODEX_HOME is kept even though it matches the Codex default, as a canary.
|
||||
// If we ever move the Codex named volume target, this env var makes the
|
||||
// dependency explicit instead of silently following the default.
|
||||
"containerEnv": {
|
||||
"CODEX_HOME": "/home/node/.codex",
|
||||
"DISABLE_AUTOUPDATER": "1",
|
||||
// post-create.sh removes `installMethod` from the seeded ~/.claude.json so
|
||||
// the npm-global binary detects its own install method. This is a backup
|
||||
// safeguard for Claude Code issue #17289. The install-checks routine probes
|
||||
// ~/.local/bin/claude just because that directory EXISTS. It does exist
|
||||
// here, because Cursor drops agent and cursor-agent symlinks there. So even
|
||||
// when installMethod is non-native, the routine reports a false "claude
|
||||
// command not found at ~/.local/bin/claude". DISABLE_AUTOUPDATER does NOT
|
||||
// turn that routine off. DISABLE_INSTALLATION_CHECKS is its dedicated kill
|
||||
// switch.
|
||||
"DISABLE_INSTALLATION_CHECKS": "1",
|
||||
"HISTFILE": "/commandhistory/.zsh_history"
|
||||
},
|
||||
|
||||
"customizations": {
|
||||
"vscode": {
|
||||
"extensions": [
|
||||
"anthropic.claude-code",
|
||||
"dbaeumer.vscode-eslint",
|
||||
"esbenp.prettier-vscode",
|
||||
"eamodio.gitlens"
|
||||
],
|
||||
"settings": {
|
||||
"editor.formatOnSave": true,
|
||||
"editor.defaultFormatter": "esbenp.prettier-vscode",
|
||||
"editor.codeActionsOnSave": {
|
||||
"source.fixAll.eslint": "explicit"
|
||||
},
|
||||
"files.eol": "\n",
|
||||
"terminal.integrated.defaultProfile.linux": "zsh",
|
||||
"terminal.integrated.profiles.linux": {
|
||||
"bash": { "path": "bash", "icon": "terminal-bash" },
|
||||
"zsh": { "path": "zsh" }
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
|
||||
// Do not remap port 4747 (gitnexus serve). gitnexus-web hardcodes
|
||||
// http://localhost:4747 as its default backend URL.
|
||||
"forwardPorts": [5173, 4747, 4173],
|
||||
"portsAttributes": {
|
||||
"5173": {
|
||||
"label": "Vite dev (gitnexus-web)",
|
||||
"onAutoForward": "notify"
|
||||
},
|
||||
"4747": {
|
||||
"label": "gitnexus serve HTTP API",
|
||||
"onAutoForward": "notify",
|
||||
"requireLocalPort": true
|
||||
},
|
||||
"4173": {
|
||||
"label": "Static web (Vite preview)",
|
||||
"onAutoForward": "silent"
|
||||
}
|
||||
},
|
||||
|
||||
// Lifecycle split (from the Dev Container spec):
|
||||
// - `updateContentCommand` runs on container-create AND whenever the
|
||||
// workspace content changes, such as a lockfile update. It owns installing
|
||||
// the workspace dependencies. Re-installing on every container-create
|
||||
// wastes time when nothing changed, but it must re-run when deps change.
|
||||
// - `postCreateCommand` runs once on container-create. It owns syncing the
|
||||
// AI CLI credentials and identity from the host. That work should happen
|
||||
// exactly once per container instance, not on every content update.
|
||||
// Run both with an explicit `bash` so they don't depend on the script's
|
||||
// executable bit surviving the workspace bind mount.
|
||||
"updateContentCommand": "bash .devcontainer/install-deps.sh",
|
||||
"postCreateCommand": "bash .devcontainer/post-create.sh"
|
||||
}
|
||||
@@ -0,0 +1,149 @@
|
||||
// This runs on the HOST, not inside the container, before the dev container is
|
||||
// created. devcontainer.json calls it via `initializeCommand`. Its job is to
|
||||
// make sure the bind-mount source folders listed in devcontainer.json already
|
||||
// exist on the host. Docker rejects a bind mount when its source is missing,
|
||||
// which happens if a CLI has never been used.
|
||||
//
|
||||
// It works on every platform. `os.homedir()` returns the home folder ($HOME on
|
||||
// Mac/Linux, %USERPROFILE% on Windows). `fs.mkdirSync({recursive: true})`
|
||||
// creates folders. It is safe to run repeatedly: a path that already exists is
|
||||
// left alone. We deliberately do NOT handle `~/.gitconfig` here. VS Code's Dev
|
||||
// Containers extension copies the host gitconfig into the container when you
|
||||
// attach, and a bind mount fights with that, so it was removed.
|
||||
//
|
||||
// The path-creating logic is exported (ensurePaths/DIRS/FILES) so tests can use
|
||||
// it. The Windows HOME side effect only runs when this file is run directly as
|
||||
// the initializeCommand. That keeps tests able to drive it against a temp dir
|
||||
// without touching the real home or calling `setx`.
|
||||
//
|
||||
// Host prerequisite: Node.js must be on PATH. That is the only host requirement
|
||||
// beyond Docker Desktop and the VS Code Dev Containers extension. Everything
|
||||
// else runs inside the container.
|
||||
|
||||
'use strict';
|
||||
|
||||
const fs = require('fs');
|
||||
const os = require('os');
|
||||
const path = require('path');
|
||||
|
||||
// Folders that are bind-mount sources in devcontainer.json. Docker rejects a
|
||||
// bind mount whose source is missing, so we create each one.
|
||||
//
|
||||
// We create the TOP per-CLI folders (~/.claude, ~/.codex, ~/.cursor) and
|
||||
// ~/.claude-mem. These back the /host/.<cli> and /host/.claude-mem read-only
|
||||
// STAGE mounts that post-create.sh copies from on container-create. We do NOT
|
||||
// create the shareable subfolders (skills/agents/plugins/memory/commands/...)
|
||||
// here anymore: they used to be read-write bind sources, but they are now
|
||||
// copied once out of the read-only stage into the per-container volume, so they
|
||||
// are no longer bind sources and pre-creating empty ones would needlessly write
|
||||
// into the host of someone who never used that CLI. post-create.sh's seed step
|
||||
// simply skips any subfolder the host doesn't have. The read-only stage bind is
|
||||
// the whole ~/.<cli> dir, so whatever shareable subfolders DO exist are visible
|
||||
// to the seed without being listed here.
|
||||
const DIRS = [
|
||||
'.claude',
|
||||
// claude-mem store ($HOME/.claude-mem). A SEPARATE top-level folder from
|
||||
// ~/.claude, holding claude-mem's SQLite DB + Chroma vector store. It is NOT
|
||||
// bind-mounted (a multi-GB SQLite/WAL store is unsafe over a 9p bind on Docker
|
||||
// Desktop Windows). post-create.sh SEEDS it once into a per-container named
|
||||
// volume from the /host/.claude-mem read-only stage. We create the source here
|
||||
// so that stage bind resolves even for a host that never ran claude-mem
|
||||
// (Docker rejects a missing bind source); the seed then finds no DB to copy
|
||||
// and the container starts with empty memory.
|
||||
'.claude-mem',
|
||||
'.codex',
|
||||
'.cursor',
|
||||
'.ssh',
|
||||
'.docker',
|
||||
'.aws',
|
||||
'.azure',
|
||||
path.join('.config', 'gh'),
|
||||
path.join('.config', 'git'),
|
||||
];
|
||||
|
||||
// Files to pre-create. Only `~/.claude.json` is created here. It is the one
|
||||
// source that is bound as a single file (read-only at /host/.claude.json). If
|
||||
// that source is missing, Docker would create a FOLDER in its place, so it has
|
||||
// to exist as a file first. `~/.claude/settings.json` and
|
||||
// `~/.codex/config.toml` are NOT single-file binds. post-create.sh copies them
|
||||
// out of the /host/.<cli> read-only folder stage, and `sync_from_host` simply
|
||||
// does nothing when they are absent (the `[ -f ]` guard). Creating them here
|
||||
// would needlessly write to the host of someone who never ran that CLI, so we
|
||||
// don't.
|
||||
const FILES = ['.claude.json'];
|
||||
|
||||
// Create every folder and touch every file under `home`. Safe to run again:
|
||||
// an existing path is left untouched. The root is a parameter so tests can run
|
||||
// it against a temp dir.
|
||||
function ensurePaths(home, dirs = DIRS, files = FILES) {
|
||||
for (const dir of dirs) {
|
||||
const full = path.join(home, dir);
|
||||
if (!fs.existsSync(full)) {
|
||||
fs.mkdirSync(full, { recursive: true });
|
||||
}
|
||||
}
|
||||
for (const file of files) {
|
||||
const full = path.join(home, file);
|
||||
if (!fs.existsSync(full)) {
|
||||
fs.closeSync(fs.openSync(full, 'a'));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
module.exports = { ensurePaths, DIRS, FILES };
|
||||
|
||||
if (require.main === module) {
|
||||
// One-time setup for native Windows. VS Code fills in the bind-mount sources
|
||||
// using `${localEnv:HOME}`, which reads its own process environment. Windows
|
||||
// does not set `HOME` by default; it uses `USERPROFILE`. With no `HOME`, the
|
||||
// bind sources shrink to filesystem-root paths (`/.claude`, `/.codex`, ...)
|
||||
// and Docker rejects them with `bind source path does not exist`.
|
||||
//
|
||||
// The fix is to save `HOME=%USERPROFILE%` into the user's environment with
|
||||
// `setx`. `setx` writes to `HKCU\Environment`. Every process the user starts
|
||||
// after that inherits the new value, including VS Code once it restarts. The
|
||||
// current VS Code process can't see the change, because its environment was
|
||||
// set when it launched. So we tell the user to restart VS Code once.
|
||||
//
|
||||
// Later runs see that `HOME` is set, skip this block, and continue normally.
|
||||
// Mac, Linux, and WSL hosts already have `HOME` set by the shell, so this
|
||||
// block does nothing on those platforms.
|
||||
if (process.platform === 'win32' && !process.env.HOME) {
|
||||
const userprofile = process.env.USERPROFILE;
|
||||
if (userprofile) {
|
||||
try {
|
||||
require('child_process').execFileSync('setx', ['HOME', userprofile], {
|
||||
stdio: 'ignore',
|
||||
});
|
||||
console.error('');
|
||||
console.error('='.repeat(70));
|
||||
console.error(' GitNexus devcontainer one-time Windows setup');
|
||||
console.error('='.repeat(70));
|
||||
console.error('');
|
||||
console.error(`HOME has been set to %USERPROFILE% (${userprofile}).`);
|
||||
console.error("VS Code reads this at startup, so the current session can't pick it up.");
|
||||
console.error('');
|
||||
console.error(' 1. Close ALL VS Code windows (File > Exit, not just the window).');
|
||||
console.error(' 2. Reopen VS Code, open this folder, and re-run Reopen in Container.');
|
||||
console.error('');
|
||||
console.error('This is a one-time setup. Subsequent rebuilds work normally.');
|
||||
console.error('='.repeat(70));
|
||||
process.exit(1);
|
||||
} catch (err) {
|
||||
console.error('ERROR: failed to set HOME automatically: ' + err.message);
|
||||
console.error('');
|
||||
console.error('Run this in a Windows shell, then restart VS Code:');
|
||||
console.error(' setx HOME "%USERPROFILE%"');
|
||||
process.exit(1);
|
||||
}
|
||||
} else {
|
||||
console.error('ERROR: neither HOME nor USERPROFILE is set on this host.');
|
||||
console.error('');
|
||||
console.error('Set HOME to your user profile directory and restart VS Code:');
|
||||
console.error(' setx HOME "%USERPROFILE%"');
|
||||
process.exit(1);
|
||||
}
|
||||
}
|
||||
|
||||
ensurePaths(os.homedir());
|
||||
}
|
||||
@@ -0,0 +1,68 @@
|
||||
#!/usr/bin/env bash
|
||||
# Devcontainer updateContentCommand. The Dev Container spec runs this when the
|
||||
# container is created AND whenever workspace content changes (for example a
|
||||
# lockfile update). This script installs workspace dependencies only. Syncing AI
|
||||
# CLI state lives in post-create.sh, which runs once right after this.
|
||||
#
|
||||
# Why the split: updateContentCommand re-runs on content changes, but
|
||||
# postCreateCommand runs only at container-create. Keeping `npm install` here
|
||||
# means a rebuild after pulling new dependencies refreshes them. The AI CLI
|
||||
# credential and path-translation work does not re-run each time.
|
||||
|
||||
set -euo pipefail
|
||||
cd /workspace
|
||||
|
||||
echo "[install-deps] 1/4: chown workspace node_modules + npm cache mount points"
|
||||
# The named volumes (workspace/*/node_modules and ~/.npm) are created at first
|
||||
# mount. They inherit ownership from the image's UID before realignment. Then
|
||||
# `updateRemoteUserUID: true` shifts the `node` user's UID. Now the volumes are
|
||||
# owned by the old, stale UID and npm install cannot write to them. So we chown
|
||||
# again here, after realignment. Running it again later changes nothing.
|
||||
#
|
||||
# We use `find -xdev -exec chown -h` (the same idiom as post-create.sh) instead
|
||||
# of a plain `chown -R`. There are two separate guards. First, `-xdev` stops
|
||||
# find from descending past each volume's own filesystem, so it won't recurse
|
||||
# into a host folder mounted underneath. Second, `-h` makes chown change the
|
||||
# symlink itself instead of following it to its target. Without `-h`, a symlink
|
||||
# in the tree (one a dependency's postinstall drops, or a dangling
|
||||
# node_modules/.bin link) would either send the chown onto a target on another
|
||||
# filesystem, or fail to follow and abort the whole script under `set -e`. For
|
||||
# regular files and directories `-h` does nothing, so the ownership fix is the
|
||||
# same.
|
||||
for d in /workspace/node_modules \
|
||||
/workspace/gitnexus/node_modules \
|
||||
/workspace/gitnexus-web/node_modules \
|
||||
/workspace/gitnexus-shared/node_modules \
|
||||
/home/node/.npm; do
|
||||
sudo find "$d" -xdev -exec chown -h node:node {} +
|
||||
done
|
||||
|
||||
echo "[install-deps] 2/4: clear stale .husky/_ runtime cache"
|
||||
# On Docker Desktop for Windows, the bind-mount permission translation won't let
|
||||
# the new container's `node` user overwrite a `.husky/_/h` file that an earlier
|
||||
# container wrote under a different UID. So we delete it. `.husky/_` is a
|
||||
# gitignored runtime cache, and husky rebuilds it during the root `npm install`.
|
||||
# Husky upstream has no fix for this UID clash.
|
||||
rm -rf .husky/_
|
||||
|
||||
echo "[install-deps] 3/4: npm install at root, then gitnexus-shared (build required)"
|
||||
# Install order matters. Root goes first, for lint-staged, husky, and prettier.
|
||||
# Then gitnexus-shared, which must be built before installing gitnexus-web or
|
||||
# gitnexus. Both of those depend on it via `file:../gitnexus-shared`.
|
||||
npm install
|
||||
cd /workspace/gitnexus-shared
|
||||
npm install
|
||||
npm run build
|
||||
|
||||
echo "[install-deps] 4/4: npm install gitnexus-web, then gitnexus"
|
||||
# gitnexus-web goes before gitnexus. The gitnexus `prepare` script runs
|
||||
# scripts/build.js, which compiles gitnexus-web when that directory is present.
|
||||
# In the devcontainer the whole workspace is bind-mounted, so gitnexus-web/ is
|
||||
# present when gitnexus installs. The production Dockerfiles COPY only selected
|
||||
# files, so the directory is not present there.
|
||||
cd /workspace/gitnexus-web
|
||||
npm install
|
||||
cd /workspace/gitnexus
|
||||
npm install
|
||||
|
||||
echo "[install-deps] done"
|
||||
@@ -0,0 +1,310 @@
|
||||
#!/usr/bin/env bash
|
||||
# Devcontainer postCreate script. It runs once, right after the container is
|
||||
# created. devcontainer.json wires it up via `postCreateCommand`. Workspace
|
||||
# dependencies are installed elsewhere, in install-deps.sh (`updateContentCommand`).
|
||||
# That script runs BEFORE this one — that is the order the devcontainer spec
|
||||
# defines. This script does one job: sync the AI CLI credentials and identity
|
||||
# from the host.
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
echo "[post-create] 1/4: chown AI CLI named-volume mount points"
|
||||
# Fix ownership on the named volumes (~/.claude, ~/.codex, ~/.cursor,
|
||||
# /commandhistory). When they first mount, they take the user ID baked into the
|
||||
# image, before any realignment. Then `updateRemoteUserUID: true` shifts the
|
||||
# `node` user to a new ID. Now the volumes are owned by the old, stale ID, and
|
||||
# writes into them fail. (~/.local is a directory in the image, not a volume.
|
||||
# We chown it too, just to be safe.) install-deps.sh fixes the workspace side.
|
||||
# This script fixes the AI CLI side, so each lifecycle hook handles its own part.
|
||||
#
|
||||
# There are two separate guards here, and they do different things. `-xdev`
|
||||
# keeps find from descending into other filesystems. The shareable dirs (skills,
|
||||
# agents, plugins, memory, commands, prompts, rules) are no longer host bind
|
||||
# mounts — they now live INSIDE the config volume (seeded in step 3/4), so
|
||||
# `-xdev` correctly walks and chowns them as the container-private volume files
|
||||
# they are. What `-xdev` still stops at are the SESSION volumes (mount group 6),
|
||||
# which remain separate filesystems mounted at sub-paths (see below). `-h` tells
|
||||
# chown to act on a symlink ITSELF instead of following it, so it never lands on
|
||||
# a target across a filesystem boundary and never aborts on a broken symlink
|
||||
# under `set -e` (a legacy Option-B symlink could still exist on a carried-over
|
||||
# volume). For regular files and directories `-h` does nothing extra.
|
||||
#
|
||||
# The session volumes (mount group 6: .claude/projects, .codex/sessions,
|
||||
# .cursor/chats, .cursor/projects) are their OWN filesystems mounted at
|
||||
# sub-paths, so `-xdev` rooted at the config-volume parent deliberately skips
|
||||
# them. That is why each one is listed as its own root below: rooted there,
|
||||
# `-xdev` walks just that volume and chowns its top level, so the CLI's first
|
||||
# write doesn't hit EACCES on a stale image UID. These are container-private
|
||||
# volumes, not the host's own files — the read-only /host/.<cli> stages we copy
|
||||
# from are mounted elsewhere and are never chowned.
|
||||
DIRS=(
|
||||
/home/node/.claude
|
||||
/home/node/.claude/projects
|
||||
/home/node/.codex
|
||||
/home/node/.codex/sessions
|
||||
/home/node/.cursor
|
||||
/home/node/.cursor/chats
|
||||
/home/node/.cursor/projects
|
||||
/home/node/.config/gh
|
||||
/home/node/.local
|
||||
/commandhistory
|
||||
)
|
||||
# claude-mem volume: chown it ONLY on first create (its completion sentinel is
|
||||
# absent). The step-4/4 seed copies the store as the node user, so a populated
|
||||
# claude-mem volume is already node-owned on every later rebuild — a recursive
|
||||
# `find` over a multi-GB store (the 7GB+ DB plus the Chroma index) just to
|
||||
# re-stamp ownership that is already correct would add real latency to every
|
||||
# rebuild for nothing. On first create the volume is empty, so this chown of the
|
||||
# bare mount point is trivial and lets the seed write into it.
|
||||
[ -f /home/node/.claude-mem/.claude-mem-seeded ] || DIRS+=(/home/node/.claude-mem)
|
||||
for d in "${DIRS[@]}"; do
|
||||
# Skip a root that isn't present rather than aborting the whole run under
|
||||
# `set -e`. Docker creates every declared volume's mount point before this
|
||||
# script runs, so in the normal case all roots exist and this is a no-op.
|
||||
# The guard matters only if a session volume is later removed from
|
||||
# devcontainer.json without its matching DIRS entry being removed too — then
|
||||
# provisioning skips it instead of failing before credentials ever sync.
|
||||
[ -d "$d" ] || continue
|
||||
sudo find "$d" -xdev -exec chown -h node:node {} +
|
||||
done
|
||||
|
||||
echo "[post-create] 2/4: sync AI CLI credentials + identity from host"
|
||||
# Clean up after an older devcontainer design (Option B). Back then these paths
|
||||
# were symlinks pointing into the read-only host stage
|
||||
# (e.g. /home/node/.claude/plugins -> /host/.claude/plugins). A write through
|
||||
# such a symlink would land on a read-only host file and fail. Delete any that
|
||||
# survive on a carried-over volume. The shareable dirs are now real directories
|
||||
# in the named volume, seeded from the host in step 3/4 below.
|
||||
for p in plugins skills agents memory commands; do
|
||||
[ -L "/home/node/.claude/$p" ] && rm "/home/node/.claude/$p"
|
||||
done
|
||||
for p in plugins prompts memories skills config.toml; do
|
||||
[ -L "/home/node/.codex/$p" ] && rm "/home/node/.codex/$p"
|
||||
done
|
||||
for p in plugins rules commands agents skills; do
|
||||
[ -L "/home/node/.cursor/$p" ] && rm "/home/node/.cursor/$p"
|
||||
done
|
||||
mkdir -p /home/node/.claude/plugins /home/node/.cursor/plugins
|
||||
|
||||
# Shareable content (skills, agents, plugins, memory, commands, prompts, rules)
|
||||
# is NO LONGER bind-mounted. It is COPIED once from the read-only host stage into
|
||||
# the named volume in step 3/4 below, so a compromised in-container dependency
|
||||
# can't write through to the host's on-disk CLI setup. This step handles only the
|
||||
# credentials, identity, and single config files. Those stay per-container in the
|
||||
# named volume and are COPIED from the host once when the container is created:
|
||||
# - .credentials.json (Claude OAuth tokens)
|
||||
# - .claude/.claude.json (Claude identity: userID, oauthAccount, and
|
||||
# migration tracking — a different file from $HOME/.claude.json)
|
||||
# - settings.json (Claude), config.toml (Codex), mcp.json (Cursor). These are
|
||||
# single config files, and single files can't be bind-mounted on Windows
|
||||
# (the EXDEV error explained below).
|
||||
# - auth.json (Codex), cli-config.json (Cursor — which mixes auth and settings)
|
||||
# - the plugin registry JSONs that contain absolute paths (Claude + Cursor).
|
||||
# Those are translated below.
|
||||
#
|
||||
# How the sync behaves: it ALWAYS overwrites from the host when the container is
|
||||
# created. A fresh container then starts logged in as the host's user, if the
|
||||
# host had credentials. From that point the container manages its own login,
|
||||
# until the next rebuild copies the host files again. Logging out inside the
|
||||
# container does NOT log out the host. Per-container login is the goal, and
|
||||
# bind-mounting these files would instead make a logout shared between both.
|
||||
|
||||
sync_from_host() {
|
||||
local src=$1
|
||||
local dst=$2
|
||||
local mode=${3:-600}
|
||||
if [ -f "$src" ]; then
|
||||
rm -f "$dst"
|
||||
cp "$src" "$dst"
|
||||
chmod "$mode" "$dst"
|
||||
fi
|
||||
}
|
||||
|
||||
sync_from_host \
|
||||
/host/.claude/.credentials.json /home/node/.claude/.credentials.json
|
||||
sync_from_host \
|
||||
/host/.claude/.claude.json /home/node/.claude/.claude.json 644
|
||||
|
||||
# These config files are COPIED from the host, not bind-mounted. We tried
|
||||
# bind-mounting them as single files and it didn't work. On Docker Desktop for
|
||||
# Windows the named volume (ext4) and the host bind mount (9p drvfs) are
|
||||
# different filesystems. Apps save a config by writing a temp file and renaming
|
||||
# it over the real one, and that rename fails across filesystems (the "EXDEV" or
|
||||
# "Device or resource busy" error). So copy the host's version into the named
|
||||
# volume when the container is created. The container can then rewrite it freely
|
||||
# until the next rebuild copies the host version again.
|
||||
sync_from_host /host/.claude/settings.json /home/node/.claude/settings.json 644
|
||||
sync_from_host /host/.codex/config.toml /home/node/.codex/config.toml 644
|
||||
|
||||
# Seed $HOME/.claude.json from the host, but NOT as a straight copy. That file
|
||||
# mixes two kinds of state. Some is portable account and onboarding state we
|
||||
# want to keep: hasCompletedOnboarding, oauthAccount, userID, projects,
|
||||
# tipsHistory. The rest describes how Claude is installed on the host, and that
|
||||
# part is never valid here. This image installs Claude with `npm install -g`,
|
||||
# but the host's `installMethod` (for example "native") makes Claude look for
|
||||
# ~/.local/bin/claude and fail with
|
||||
# "claude command not found at /home/node/.local/bin/claude". The fix strips the
|
||||
# machine-specific fields and forces hasCompletedOnboarding, while handling a
|
||||
# host file that isn't a JSON object. That logic lives in seed-claude-config.cjs
|
||||
# so it can be unit-tested and prettier-checked
|
||||
# (translate-plugin-registries.test.cjs).
|
||||
node "$SCRIPT_DIR/seed-claude-config.cjs"
|
||||
|
||||
# Codex auth. Some hosts store credentials in the OS keyring instead of on disk
|
||||
# (`cli_auth_credentials_store = "keyring"`, the default on macOS). Those hosts
|
||||
# have no auth.json file, so the copy below quietly does nothing. In that case,
|
||||
# log in inside the container with `codex login --device-auth`.
|
||||
sync_from_host \
|
||||
/host/.codex/auth.json /home/node/.codex/auth.json
|
||||
|
||||
# Cursor CLI. Its cli-config.json holds both auth and settings in one file.
|
||||
# Cursor has known upstream problems authenticating inside Docker, even when the
|
||||
# config is copied correctly. If `cursor-agent` reports auth errors after the
|
||||
# copy, run `cursor-agent login` again inside the container. mcp.json (Cursor's
|
||||
# MCP server config) is also a single file, so it is copied on create rather
|
||||
# than bind-mounted, for the same EXDEV reason as above. hooks.json is left out
|
||||
# on purpose. Cursor hooks run shell commands, and sharing the host's hooks
|
||||
# would widen the supply-chain attack surface inside the container. Copy it in
|
||||
# yourself if you want the host's hooks in the container.
|
||||
sync_from_host \
|
||||
/host/.cursor/cli-config.json /home/node/.cursor/cli-config.json
|
||||
sync_from_host \
|
||||
/host/.cursor/mcp.json /home/node/.cursor/mcp.json 644
|
||||
|
||||
# gh CLI auth + settings. Same copy-into-volume model as the credentials above:
|
||||
# hosts.yml holds the GitHub token (mode 600), config.yml holds settings (644).
|
||||
# Copied from the read-only /host/.config/gh stage into the gh-config named
|
||||
# volume on create. Because the volume is writable, an in-container
|
||||
# `gh auth login` / `gh auth refresh` persists across rebuilds; because the
|
||||
# stage is read-only, nothing flows back to the host. If the host had no login,
|
||||
# both copies quietly no-op and whatever the container wrote is kept.
|
||||
sync_from_host /host/.config/gh/hosts.yml /home/node/.config/gh/hosts.yml
|
||||
sync_from_host /host/.config/gh/config.yml /home/node/.config/gh/config.yml 644
|
||||
|
||||
echo "[post-create] 3/4: seed shareable config dirs from host (first create only)"
|
||||
# The shareable dirs (Claude skills/agents/memory/commands/plugins; Codex
|
||||
# plugins/prompts/memories/skills; Cursor rules/commands/agents/skills/plugins)
|
||||
# used to be read-write host bind mounts, so a write inside the container landed
|
||||
# directly on the host's files. That exposed the host's on-disk CLI setup: a
|
||||
# compromised workspace dependency running in the container could drop a malicious
|
||||
# skill, agent, command, or plugin into the host's folders, which the next HOST
|
||||
# session would then auto-load. To protect the host, these are no longer bound.
|
||||
# Instead we COPY them once from the read-only /host/.<cli> stage into the
|
||||
# per-container named volume, exactly like claude-mem (step 4/4) and the session
|
||||
# volumes. The container gets its own writable copy and can NEVER write back to
|
||||
# the host. The container also avoids the old read-only-stage EROFS failure,
|
||||
# because it writes to its own volume copy, not a read-only mount.
|
||||
#
|
||||
# Seed-once, persist: a per-CLI marker file records that the copy has happened.
|
||||
# On the first container-create the marker is absent, so we copy; on every later
|
||||
# rebuild the marker is present, so we skip and keep whatever the container has
|
||||
# accumulated. Host edits made AFTER the first create do NOT reach the container
|
||||
# until you remove the config volume and rebuild (see README § Rebuild/reset).
|
||||
seed_shareable() {
|
||||
# seed_shareable <cli> <subdir>...: copy each /host/.<cli>/<subdir> into the
|
||||
# named volume, once. Skips a subdir the host doesn't have. We use `cp -r`,
|
||||
# NOT `cp -a`/`cp -p`: this script runs as the non-root node user, and the
|
||||
# host-stage files are owned by a different UID, so trying to preserve
|
||||
# ownership would fail with EPERM and abort the run under `set -e` (the same
|
||||
# reason sync_from_host uses plain cp). `cp -r` copies contents owned by node
|
||||
# — exactly what we want — and preserves symlinks as symlinks (GNU default).
|
||||
local cli=$1
|
||||
shift
|
||||
local marker="/home/node/.$cli/.devcontainer-shareable-seeded"
|
||||
[ -f "$marker" ] && return 0
|
||||
for sub in "$@"; do
|
||||
local src="/host/.$cli/$sub"
|
||||
local dst="/home/node/.$cli/$sub"
|
||||
[ -d "$src" ] || continue
|
||||
mkdir -p "$dst"
|
||||
cp -r "$src/." "$dst/"
|
||||
done
|
||||
}
|
||||
|
||||
# Decide which plugin registries to translate BEFORE seeding sets the markers.
|
||||
# We translate only a CLI being seeded this run, so a plugin installed inside the
|
||||
# container isn't overwritten by the host's registry on a later rebuild. Codex
|
||||
# has no path-bearing registry (config.toml holds git URLs), so it's never here.
|
||||
TRANSLATE_CLIS=()
|
||||
[ -f /home/node/.claude/.devcontainer-shareable-seeded ] || TRANSLATE_CLIS+=(claude)
|
||||
[ -f /home/node/.cursor/.devcontainer-shareable-seeded ] || TRANSLATE_CLIS+=(cursor)
|
||||
|
||||
seed_shareable claude skills agents memory commands plugins/marketplaces plugins/cache
|
||||
seed_shareable codex plugins prompts memories skills
|
||||
seed_shareable cursor rules commands agents skills plugins/marketplaces plugins/local
|
||||
|
||||
# Translate the path-bearing plugin registries (Claude + Cursor) for the CLIs we
|
||||
# just seeded. They store absolute, OS-native install paths
|
||||
# (`C:\Users\X\.claude\plugins\...` on Windows), which the Linux container can't
|
||||
# resolve — it would fail with `cache-miss`. translate-plugin-registries.cjs
|
||||
# rewrites those to the container's paths and writes the result into the volume.
|
||||
if [ "${#TRANSLATE_CLIS[@]}" -gt 0 ]; then
|
||||
node "$SCRIPT_DIR/translate-plugin-registries.cjs" "${TRANSLATE_CLIS[@]}"
|
||||
fi
|
||||
|
||||
# Record that each CLI's shareable surface is seeded, so later rebuilds keep the
|
||||
# container's copy. Touch even when the host had nothing to copy — an empty CLI
|
||||
# is still "seeded", and we don't want to re-scan the host on every rebuild.
|
||||
#
|
||||
# ORDERING INVARIANT — do NOT move these touches earlier (e.g. into
|
||||
# seed_shareable per-CLI). The markers must be written only AFTER the registry
|
||||
# translation above, because seed (cache copy) and translate (registry rewrite)
|
||||
# are logically atomic: a marker set between them would let a later rebuild skip
|
||||
# translation for an already-seeded CLI, leaving its cache/ in place but its
|
||||
# registry still pointing at host paths (`cache-miss`). Writing all markers here,
|
||||
# after translate, means any abort mid-seed leaves NO markers, so the next create
|
||||
# re-runs the whole seed+translate. The cost is re-copying an already-copied CLI
|
||||
# on retry; `cp -r` overwrites in place, so that is idempotent and cheap relative
|
||||
# to a broken plugin registry.
|
||||
for cli in claude codex cursor; do
|
||||
touch "/home/node/.$cli/.devcontainer-shareable-seeded"
|
||||
done
|
||||
|
||||
echo "[post-create] 4/4: seed claude-mem store from host (first create only)"
|
||||
# claude-mem keeps its memory in $HOME/.claude-mem — a SQLite DB (claude-mem.db
|
||||
# plus -wal/-shm) and a Chroma vector store (chroma/chroma.sqlite3 + HNSW index
|
||||
# binaries). It is mounted as a per-container named volume, NOT a host bind:
|
||||
# pushing a multi-GB SQLite/WAL store over the 9p/virtiofs bind risks unreliable
|
||||
# fcntl locking and corruption, especially if claude-mem ran on the host and in
|
||||
# the container against the same files at once (see devcontainer.json).
|
||||
#
|
||||
# So seed it ONCE, then let the container own its copy. On every later rebuild
|
||||
# we skip the copy and keep whatever the container has accumulated since —
|
||||
# rebuilds never clobber it. The container's memory and the host's diverge from
|
||||
# this seed point on; that is the deliberate cost of keeping SQLite off a shared
|
||||
# bind. To re-seed from the host, remove the volume (`docker volume rm
|
||||
# claude-mem-<id>`) and rebuild.
|
||||
#
|
||||
# The skip guard is a COMPLETION SENTINEL (.claude-mem-seeded), NOT the presence
|
||||
# of claude-mem.db. Keying on the DB file would be a trap: a multi-GB `cp -r` can
|
||||
# be interrupted (disk full, I/O error) and abort the script under `set -e`,
|
||||
# leaving a PARTIAL claude-mem.db behind. The next create would then see that
|
||||
# truncated file and treat the store as "already seeded", sticking the container
|
||||
# with a corrupt DB forever. With a sentinel touched only AFTER `cp` returns 0,
|
||||
# an interrupted seed leaves no sentinel; the next create clears the half-copied
|
||||
# store and retries cleanly. CONSISTENCY: copying a live WAL database is only
|
||||
# crash-consistent if claude-mem is NOT writing on the host during the copy — do
|
||||
# not run claude-mem on the host during a first-create or a re-seed rebuild.
|
||||
#
|
||||
# `cp -r` (not `cp -a`/`cp -p`) copies the DB together with its -wal/-shm
|
||||
# sidecars in one pass. We avoid preserving ownership for the same reason as the
|
||||
# shareable seed above: this runs as the non-root node user against host-owned
|
||||
# files, so `cp -a` would fail with EPERM and abort under `set -e`. `cp -r`
|
||||
# leaves the copies owned by node. The host stage is read-only, so this can
|
||||
# never write back to the host's live DB.
|
||||
if [ -f /host/.claude-mem/claude-mem.db ] && [ ! -f /home/node/.claude-mem/.claude-mem-seeded ]; then
|
||||
echo "[post-create] seeding ~/.claude-mem from host (one-time copy, may be several GB)"
|
||||
# Clear any partial store left by a previously-interrupted seed (mindepth 1
|
||||
# so the volume mount point itself is never removed), then copy and only then
|
||||
# write the sentinel. A partial store is node-owned (cp runs as node, and
|
||||
# step 1 re-chowns the volume whenever the sentinel is absent), so no sudo.
|
||||
find /home/node/.claude-mem -mindepth 1 -maxdepth 1 -exec rm -rf {} +
|
||||
cp -r /host/.claude-mem/. /home/node/.claude-mem/
|
||||
touch /home/node/.claude-mem/.claude-mem-seeded
|
||||
else
|
||||
echo "[post-create] skipping claude-mem seed (already seeded, or host has no store)"
|
||||
fi
|
||||
|
||||
echo "[post-create] done"
|
||||
@@ -0,0 +1,83 @@
|
||||
// Builds the container's $HOME/.claude.json from the host's copy. It does NOT
|
||||
// copy the host file verbatim. The host's ~/.claude.json holds two kinds of
|
||||
// data. Some is portable account and onboarding state: hasCompletedOnboarding,
|
||||
// oauthAccount, userID, projects, tipsHistory. We keep that. The rest tracks
|
||||
// how Claude was installed on the host machine, and that is never right inside
|
||||
// this container.
|
||||
//
|
||||
// Here is why the install fields break things. The image installs Claude with
|
||||
// `npm install -g`. But if the host's `installMethod` says something like
|
||||
// "native", Claude looks for ~/.local/bin/claude and fails with
|
||||
// "claude command not found at /home/node/.local/bin/claude". So we drop the
|
||||
// install and machine fields. With them gone, the npm-global binary detects its
|
||||
// own install method. We also force hasCompletedOnboarding so the setup wizard
|
||||
// is skipped, even when the host has never run Claude before.
|
||||
//
|
||||
// This logic was pulled out of a heredoc in post-create.sh. As its own file the
|
||||
// transform can be unit-tested and prettier-checked (see seed-claude-config.test
|
||||
// via the translate-plugin-registries test harness). DISABLE_AUTOUPDATER=1 in
|
||||
// containerEnv already stops runtime updates. This file only quiets the doctor
|
||||
// mismatch and the native-path probe.
|
||||
|
||||
'use strict';
|
||||
|
||||
const fs = require('fs');
|
||||
|
||||
// Fields that describe how Claude was installed on the host machine. They are
|
||||
// never valid in an `npm install -g` container. Removing them lets Claude
|
||||
// detect the npm-global install on its own.
|
||||
const MACHINE_FIELDS = [
|
||||
'installMethod',
|
||||
'autoUpdates',
|
||||
'autoUpdatesProtectedForNative',
|
||||
'shiftEnterKeyBindingInstalled',
|
||||
];
|
||||
|
||||
// Pure transform: take whatever the host file parsed to and return a config
|
||||
// object suitable for the container. It also guards against a host file that is
|
||||
// valid JSON but not an object. A bare number, string, or array would pass the
|
||||
// parse try/catch. Then the field deletes would do nothing, the
|
||||
// hasCompletedOnboarding assignment would silently fail, and onboarding would
|
||||
// trigger again on every rebuild. The guard replaces such a value with {}.
|
||||
function sanitizeClaudeConfig(parsed) {
|
||||
let cfg = parsed;
|
||||
if (cfg === null || typeof cfg !== 'object' || Array.isArray(cfg)) {
|
||||
cfg = {};
|
||||
}
|
||||
for (const k of MACHINE_FIELDS) {
|
||||
delete cfg[k];
|
||||
}
|
||||
cfg.hasCompletedOnboarding = true; // skip the wizard, even on a first-time host
|
||||
return cfg;
|
||||
}
|
||||
|
||||
function readHostConfig(src) {
|
||||
try {
|
||||
if (fs.existsSync(src) && fs.statSync(src).size > 0) {
|
||||
return JSON.parse(fs.readFileSync(src, 'utf8'));
|
||||
}
|
||||
} catch {
|
||||
// Host file is malformed or unreadable. Fall back to an empty config so the
|
||||
// container still gets a valid file that carries hasCompletedOnboarding.
|
||||
}
|
||||
return {};
|
||||
}
|
||||
|
||||
function main() {
|
||||
const src = process.argv[2] || '/host/.claude.json';
|
||||
const dst = process.argv[3] || '/home/node/.claude.json';
|
||||
const cfg = sanitizeClaudeConfig(readHostConfig(src));
|
||||
try {
|
||||
fs.writeFileSync(dst, JSON.stringify(cfg, null, 2));
|
||||
fs.chmodSync(dst, 0o644);
|
||||
} catch (err) {
|
||||
console.error(`[post-create] ERROR: failed to seed ${dst}: ${err && err.message}`);
|
||||
process.exit(1);
|
||||
}
|
||||
}
|
||||
|
||||
module.exports = { sanitizeClaudeConfig, readHostConfig, MACHINE_FIELDS };
|
||||
|
||||
if (require.main === module) {
|
||||
main();
|
||||
}
|
||||
@@ -0,0 +1,107 @@
|
||||
// Rewrites the host paths inside Claude and Cursor plugin-registry JSON files
|
||||
// so they point at the container's Linux paths, then writes the results into
|
||||
// the named volume.
|
||||
//
|
||||
// Why: both CLIs store absolute, OS-native install paths in their registry
|
||||
// JSONs. On Windows that looks like `C:\Users\X\.claude\plugins\...`; on macOS
|
||||
// like `/Users/X/.cursor/...`. The Linux container can't use those paths. If we
|
||||
// just bind-mounted the host files in, the CLI would try to resolve a Windows
|
||||
// path under Linux and fail with `cache-miss`. So for each CLI we read the host
|
||||
// registry, rewrite every absolute path ending in `/.<cli>/plugins/<rest>` to
|
||||
// `/home/node/.<cli>/plugins/<rest>`, and write the result into the named volume.
|
||||
//
|
||||
// Codex is left alone. Its registry is config.toml and holds git URLs, not
|
||||
// filesystem paths, so there's nothing to translate — its whole plugins/ dir is
|
||||
// copied as-is into the container volume instead (seeded once by post-create.sh).
|
||||
//
|
||||
// This code lived inside a post-create.sh heredoc. We pulled it out so the regex
|
||||
// and the deep rewrite can be unit-tested and prettier-checked. The regex has
|
||||
// had path-handling bugs before.
|
||||
|
||||
'use strict';
|
||||
|
||||
const fs = require('fs');
|
||||
const path = require('path');
|
||||
|
||||
// Build a regex that matches an absolute path containing
|
||||
// `<sep>.<cli><sep>plugins<sep><rest>`, where <sep> is `/` or `\`. It's anchored
|
||||
// at the start of the string. The lazy `.*?` eats the home prefix up to the
|
||||
// FIRST `.<cli>/plugins` segment.
|
||||
function buildRe(cliName) {
|
||||
return new RegExp(`^(?:[A-Za-z]:)?[\\\\/].*?[\\\\/]\\.${cliName}[\\\\/]plugins[\\\\/](.*)$`);
|
||||
}
|
||||
|
||||
// Walk `obj` and rewrite every string value that matches `re`. A match is
|
||||
// remapped under `ctr`, the container's plugins dir. Windows backslashes in the
|
||||
// matched part are switched to forward slashes.
|
||||
function rewriteDeep(obj, re, ctr) {
|
||||
if (Array.isArray(obj)) return obj.map((v) => rewriteDeep(v, re, ctr));
|
||||
if (obj && typeof obj === 'object') {
|
||||
const out = {};
|
||||
for (const [k, v] of Object.entries(obj)) out[k] = rewriteDeep(v, re, ctr);
|
||||
return out;
|
||||
}
|
||||
if (typeof obj === 'string') {
|
||||
return obj.replace(re, (_, rest) => `${ctr}/${rest.replace(/\\/g, '/')}`);
|
||||
}
|
||||
return obj;
|
||||
}
|
||||
|
||||
const REGISTRIES = [
|
||||
{
|
||||
cli: 'claude',
|
||||
host: '/host/.claude/plugins',
|
||||
ctr: '/home/node/.claude/plugins',
|
||||
files: ['known_marketplaces.json', 'installed_plugins.json', 'plugin-catalog-cache.json'],
|
||||
},
|
||||
{
|
||||
cli: 'cursor',
|
||||
host: '/host/.cursor/plugins',
|
||||
ctr: '/home/node/.cursor/plugins',
|
||||
files: ['installed_plugins.json'],
|
||||
},
|
||||
];
|
||||
|
||||
function translate(registries) {
|
||||
for (const reg of registries) {
|
||||
const re = buildRe(reg.cli);
|
||||
try {
|
||||
fs.mkdirSync(reg.ctr, { recursive: true });
|
||||
} catch (err) {
|
||||
console.error(`[post-create] ERROR: failed to create ${reg.ctr}: ${err && err.message}`);
|
||||
process.exit(1);
|
||||
}
|
||||
for (const name of reg.files) {
|
||||
const src = path.join(reg.host, name);
|
||||
const dst = path.join(reg.ctr, name);
|
||||
if (!fs.existsSync(src) || fs.statSync(src).size === 0) continue;
|
||||
let data;
|
||||
try {
|
||||
data = JSON.parse(fs.readFileSync(src, 'utf8'));
|
||||
} catch {
|
||||
continue; // Skip a malformed host registry instead of aborting.
|
||||
}
|
||||
try {
|
||||
fs.writeFileSync(dst, JSON.stringify(rewriteDeep(data, re, reg.ctr), null, 2));
|
||||
} catch (err) {
|
||||
console.error(`[post-create] ERROR: failed to write ${dst}: ${err && err.message}`);
|
||||
process.exit(1);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Filter the registry table by CLI name. post-create.sh passes the CLIs it is
|
||||
// seeding this run (e.g. `claude`), so a registry is only (re)generated on the
|
||||
// FIRST container-create for that CLI — never on a rebuild, where it would
|
||||
// clobber a plugin the user installed inside the container. An empty filter
|
||||
// (no args) means "translate every registry" — the original behavior.
|
||||
function selectRegistries(registries, only) {
|
||||
return only && only.length ? registries.filter((r) => only.includes(r.cli)) : registries;
|
||||
}
|
||||
|
||||
module.exports = { buildRe, rewriteDeep, REGISTRIES, translate, selectRegistries };
|
||||
|
||||
if (require.main === module) {
|
||||
translate(selectRegistries(REGISTRIES, process.argv.slice(2)));
|
||||
}
|
||||
@@ -0,0 +1,436 @@
|
||||
// Unit tests for the devcontainer host->container config transforms.
|
||||
//
|
||||
// This code used to live inside post-create.sh heredocs, where lint could not
|
||||
// see it and tests could not reach it. We test three things:
|
||||
// - plugin-registry path translation (buildRe + rewriteDeep + the real
|
||||
// filesystem translate() driver). Path handling here has had bugs before.
|
||||
// - the strip of machine-specific fields from $HOME/.claude.json
|
||||
// (sanitizeClaudeConfig + readHostConfig + the seed-claude-config main()
|
||||
// entry point)
|
||||
// - the host bind-source bootstrap (ensurePaths). One test guards against a
|
||||
// regression: ensurePaths must NOT pre-create settings.json / config.toml
|
||||
// on the host.
|
||||
//
|
||||
// We test both pure functions and code that touches the filesystem. The
|
||||
// filesystem tests use throwaway directories under os.tmpdir() and delete them
|
||||
// when done. So they run in CI with no mounts and never touch the real home dir.
|
||||
//
|
||||
// Run with the built-in Node test runner (no extra dependencies):
|
||||
// node --test .devcontainer/
|
||||
|
||||
'use strict';
|
||||
|
||||
const test = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const os = require('node:os');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
const { execFileSync } = require('node:child_process');
|
||||
|
||||
const {
|
||||
buildRe,
|
||||
rewriteDeep,
|
||||
translate,
|
||||
selectRegistries,
|
||||
} = require('./translate-plugin-registries.cjs');
|
||||
const { sanitizeClaudeConfig, readHostConfig } = require('./seed-claude-config.cjs');
|
||||
const { ensurePaths, DIRS, FILES } = require('./ensure-host-config-dirs.cjs');
|
||||
|
||||
const CLAUDE = '/home/node/.claude/plugins';
|
||||
const CURSOR = '/home/node/.cursor/plugins';
|
||||
|
||||
// Make a fresh throwaway directory under the OS temp root. mkdtemp picks a
|
||||
// unique name on every call, so we don't need Date.now() or random names.
|
||||
function tmp() {
|
||||
return fs.mkdtempSync(path.join(os.tmpdir(), 'gn-dc-'));
|
||||
}
|
||||
|
||||
function rw(value, cli, ctr) {
|
||||
return rewriteDeep(value, buildRe(cli), ctr);
|
||||
}
|
||||
|
||||
test('claude: Windows backslash absolute path -> container path', () => {
|
||||
assert.equal(
|
||||
rw('C:\\Users\\gergo\\.claude\\plugins\\cache\\x\\1.0', 'claude', CLAUDE),
|
||||
'/home/node/.claude/plugins/cache/x/1.0',
|
||||
);
|
||||
});
|
||||
|
||||
test('claude: Windows forward-slash absolute path -> container path', () => {
|
||||
assert.equal(
|
||||
rw('C:/Users/gergo/.claude/plugins/marketplaces/m', 'claude', CLAUDE),
|
||||
'/home/node/.claude/plugins/marketplaces/m',
|
||||
);
|
||||
});
|
||||
|
||||
test('claude: macOS POSIX path -> container path', () => {
|
||||
assert.equal(
|
||||
rw('/Users/alice/.claude/plugins/marketplaces/m', 'claude', CLAUDE),
|
||||
'/home/node/.claude/plugins/marketplaces/m',
|
||||
);
|
||||
});
|
||||
|
||||
test('claude: Linux POSIX path -> container path', () => {
|
||||
assert.equal(
|
||||
rw('/home/bob/.claude/plugins/cache/foo', 'claude', CLAUDE),
|
||||
'/home/node/.claude/plugins/cache/foo',
|
||||
);
|
||||
});
|
||||
|
||||
test('cursor: Windows path -> container cursor path', () => {
|
||||
assert.equal(
|
||||
rw('C:\\Users\\gergo\\.cursor\\plugins\\local\\myplug', 'cursor', CURSOR),
|
||||
'/home/node/.cursor/plugins/local/myplug',
|
||||
);
|
||||
});
|
||||
|
||||
test('cross-CLI isolation: claude regex leaves a .cursor path untouched', () => {
|
||||
const input = 'C:\\Users\\g\\.cursor\\plugins\\x';
|
||||
assert.equal(rw(input, 'claude', CLAUDE), input);
|
||||
});
|
||||
|
||||
test('non-path strings pass through unchanged', () => {
|
||||
assert.equal(rw('not-a-path', 'claude', CLAUDE), 'not-a-path');
|
||||
assert.equal(
|
||||
rw('https://github.com/EveryInc/x.git', 'claude', CLAUDE),
|
||||
'https://github.com/EveryInc/x.git',
|
||||
);
|
||||
});
|
||||
|
||||
test('non-string scalars pass through unchanged', () => {
|
||||
assert.equal(rw(42, 'claude', CLAUDE), 42);
|
||||
assert.equal(rw(null, 'claude', CLAUDE), null);
|
||||
assert.equal(rw(true, 'claude', CLAUDE), true);
|
||||
});
|
||||
|
||||
test('nested objects/arrays are rewritten deeply', () => {
|
||||
const input = {
|
||||
'compound-engineering@m': [
|
||||
{ installPath: 'C:\\Users\\g\\.claude\\plugins\\cache\\ce\\3.9.2', version: '3.9.2' },
|
||||
],
|
||||
nested: { installLocation: '/Users/g/.claude/plugins/marketplaces/m' },
|
||||
};
|
||||
const out = rw(input, 'claude', CLAUDE);
|
||||
assert.equal(
|
||||
out['compound-engineering@m'][0].installPath,
|
||||
'/home/node/.claude/plugins/cache/ce/3.9.2',
|
||||
);
|
||||
assert.equal(out['compound-engineering@m'][0].version, '3.9.2');
|
||||
assert.equal(out.nested.installLocation, '/home/node/.claude/plugins/marketplaces/m');
|
||||
});
|
||||
|
||||
test('sanitizeClaudeConfig: strips machine fields, forces hasCompletedOnboarding', () => {
|
||||
const out = sanitizeClaudeConfig({
|
||||
installMethod: 'native',
|
||||
autoUpdates: false,
|
||||
autoUpdatesProtectedForNative: true,
|
||||
shiftEnterKeyBindingInstalled: true,
|
||||
userID: 'abc',
|
||||
oauthAccount: { emailAddress: 'x@y.z' },
|
||||
});
|
||||
assert.equal(out.installMethod, undefined);
|
||||
assert.equal(out.autoUpdates, undefined);
|
||||
assert.equal(out.autoUpdatesProtectedForNative, undefined);
|
||||
assert.equal(out.shiftEnterKeyBindingInstalled, undefined);
|
||||
assert.equal(out.userID, 'abc');
|
||||
assert.equal(out.oauthAccount.emailAddress, 'x@y.z');
|
||||
assert.equal(out.hasCompletedOnboarding, true);
|
||||
});
|
||||
|
||||
test('sanitizeClaudeConfig: non-object inputs become a valid onboarding-bearing object', () => {
|
||||
for (const bad of [42, 'x', null, ['a'], true]) {
|
||||
const out = sanitizeClaudeConfig(bad);
|
||||
assert.equal(typeof out, 'object');
|
||||
assert.equal(Array.isArray(out), false);
|
||||
assert.equal(out.hasCompletedOnboarding, true);
|
||||
}
|
||||
});
|
||||
|
||||
test('sanitizeClaudeConfig: empty object still gets hasCompletedOnboarding', () => {
|
||||
assert.deepEqual(sanitizeClaudeConfig({}), { hasCompletedOnboarding: true });
|
||||
});
|
||||
|
||||
// --- readHostConfig: reading the file, and the fallbacks when it fails ------
|
||||
|
||||
test('readHostConfig: missing file -> {}', () => {
|
||||
const dir = tmp();
|
||||
try {
|
||||
assert.deepEqual(readHostConfig(path.join(dir, 'nope.json')), {});
|
||||
} finally {
|
||||
fs.rmSync(dir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test('readHostConfig: empty (zero-byte) file -> {}', () => {
|
||||
const dir = tmp();
|
||||
try {
|
||||
const f = path.join(dir, 'empty.json');
|
||||
fs.writeFileSync(f, '');
|
||||
assert.deepEqual(readHostConfig(f), {});
|
||||
} finally {
|
||||
fs.rmSync(dir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test('readHostConfig: malformed JSON -> {}', () => {
|
||||
const dir = tmp();
|
||||
try {
|
||||
const f = path.join(dir, 'bad.json');
|
||||
fs.writeFileSync(f, '{ not valid json');
|
||||
assert.deepEqual(readHostConfig(f), {});
|
||||
} finally {
|
||||
fs.rmSync(dir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test('readHostConfig: valid object is parsed through', () => {
|
||||
const dir = tmp();
|
||||
try {
|
||||
const f = path.join(dir, 'ok.json');
|
||||
fs.writeFileSync(f, JSON.stringify({ userID: 'u', hasCompletedOnboarding: false }));
|
||||
const out = readHostConfig(f);
|
||||
assert.equal(out.userID, 'u');
|
||||
assert.equal(out.hasCompletedOnboarding, false);
|
||||
} finally {
|
||||
fs.rmSync(dir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
// --- translate(): runs against real registry files on disk ------------------
|
||||
|
||||
test('translate: rewrites host absolute paths and writes into the ctr dir', () => {
|
||||
const hostDir = tmp();
|
||||
const ctrParent = tmp();
|
||||
const ctrDir = path.join(ctrParent, 'plugins'); // need not exist yet; translate creates it
|
||||
try {
|
||||
const reg = [{ cli: 'claude', host: hostDir, ctr: ctrDir, files: ['installed_plugins.json'] }];
|
||||
fs.writeFileSync(
|
||||
path.join(hostDir, 'installed_plugins.json'),
|
||||
JSON.stringify({ 'p@m': [{ installPath: 'C:\\Users\\g\\.claude\\plugins\\cache\\p\\1.0' }] }),
|
||||
);
|
||||
translate(reg);
|
||||
const out = JSON.parse(fs.readFileSync(path.join(ctrDir, 'installed_plugins.json'), 'utf8'));
|
||||
assert.equal(out['p@m'][0].installPath, `${ctrDir}/cache/p/1.0`);
|
||||
} finally {
|
||||
fs.rmSync(hostDir, { recursive: true, force: true });
|
||||
fs.rmSync(ctrParent, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test('translate: idempotent — a second run reproduces byte-identical output', () => {
|
||||
const hostDir = tmp();
|
||||
const ctrParent = tmp();
|
||||
const ctrDir = path.join(ctrParent, 'plugins');
|
||||
try {
|
||||
const reg = [{ cli: 'claude', host: hostDir, ctr: ctrDir, files: ['installed_plugins.json'] }];
|
||||
fs.writeFileSync(
|
||||
path.join(hostDir, 'installed_plugins.json'),
|
||||
JSON.stringify({ 'p@m': [{ installPath: 'C:\\Users\\g\\.claude\\plugins\\cache\\p\\1.0' }] }),
|
||||
);
|
||||
translate(reg);
|
||||
const first = fs.readFileSync(path.join(ctrDir, 'installed_plugins.json'), 'utf8');
|
||||
translate(reg);
|
||||
const second = fs.readFileSync(path.join(ctrDir, 'installed_plugins.json'), 'utf8');
|
||||
assert.equal(first, second);
|
||||
} finally {
|
||||
fs.rmSync(hostDir, { recursive: true, force: true });
|
||||
fs.rmSync(ctrParent, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test('translate: malformed host registry is skipped, dst not written', () => {
|
||||
const hostDir = tmp();
|
||||
const ctrParent = tmp();
|
||||
const ctrDir = path.join(ctrParent, 'plugins');
|
||||
try {
|
||||
const reg = [{ cli: 'claude', host: hostDir, ctr: ctrDir, files: ['installed_plugins.json'] }];
|
||||
fs.writeFileSync(path.join(hostDir, 'installed_plugins.json'), '{ broken');
|
||||
translate(reg);
|
||||
assert.equal(fs.existsSync(path.join(ctrDir, 'installed_plugins.json')), false);
|
||||
} finally {
|
||||
fs.rmSync(hostDir, { recursive: true, force: true });
|
||||
fs.rmSync(ctrParent, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test('translate: empty and missing host registries are skipped without error', () => {
|
||||
const hostDir = tmp();
|
||||
const ctrParent = tmp();
|
||||
const ctrDir = path.join(ctrParent, 'plugins');
|
||||
try {
|
||||
const reg = [
|
||||
{ cli: 'claude', host: hostDir, ctr: ctrDir, files: ['empty.json', 'missing.json'] },
|
||||
];
|
||||
fs.writeFileSync(path.join(hostDir, 'empty.json'), ''); // we never create missing.json
|
||||
translate(reg);
|
||||
assert.equal(fs.existsSync(path.join(ctrDir, 'empty.json')), false);
|
||||
assert.equal(fs.existsSync(path.join(ctrDir, 'missing.json')), false);
|
||||
} finally {
|
||||
fs.rmSync(hostDir, { recursive: true, force: true });
|
||||
fs.rmSync(ctrParent, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
// --- selectRegistries: the per-CLI filter post-create.sh drives translate with
|
||||
|
||||
test('selectRegistries: no filter -> all registries (original behavior)', () => {
|
||||
const regs = [{ cli: 'claude' }, { cli: 'cursor' }];
|
||||
assert.deepEqual(selectRegistries(regs, []), regs);
|
||||
assert.deepEqual(selectRegistries(regs, undefined), regs);
|
||||
});
|
||||
|
||||
test('selectRegistries: filter keeps only the named CLIs', () => {
|
||||
const regs = [{ cli: 'claude' }, { cli: 'cursor' }];
|
||||
assert.deepEqual(selectRegistries(regs, ['claude']), [{ cli: 'claude' }]);
|
||||
assert.deepEqual(selectRegistries(regs, ['cursor']), [{ cli: 'cursor' }]);
|
||||
assert.deepEqual(selectRegistries(regs, ['claude', 'cursor']), regs);
|
||||
});
|
||||
|
||||
test('selectRegistries: an unknown CLI name selects nothing', () => {
|
||||
const regs = [{ cli: 'claude' }, { cli: 'cursor' }];
|
||||
assert.deepEqual(selectRegistries(regs, ['codex']), []);
|
||||
});
|
||||
|
||||
test('selectRegistries: empty registry table stays empty under any filter', () => {
|
||||
assert.deepEqual(selectRegistries([], ['claude']), []);
|
||||
assert.deepEqual(selectRegistries([], []), []);
|
||||
});
|
||||
|
||||
// --- seed-claude-config main(): end-to-end, through the real CLI entry point
|
||||
|
||||
const SEED_SCRIPT = path.join(__dirname, 'seed-claude-config.cjs');
|
||||
|
||||
test('seed main: strips machine fields, keeps account, sets onboarding, chmod 644', () => {
|
||||
const dir = tmp();
|
||||
try {
|
||||
const src = path.join(dir, 'host.claude.json');
|
||||
const dst = path.join(dir, 'out.claude.json');
|
||||
fs.writeFileSync(
|
||||
src,
|
||||
JSON.stringify({
|
||||
installMethod: 'native',
|
||||
userID: 'abc',
|
||||
oauthAccount: { emailAddress: 'x@y.z' },
|
||||
}),
|
||||
);
|
||||
execFileSync(process.execPath, [SEED_SCRIPT, src, dst]);
|
||||
const out = JSON.parse(fs.readFileSync(dst, 'utf8'));
|
||||
assert.equal(out.installMethod, undefined);
|
||||
assert.equal(out.userID, 'abc');
|
||||
assert.equal(out.oauthAccount.emailAddress, 'x@y.z');
|
||||
assert.equal(out.hasCompletedOnboarding, true);
|
||||
if (process.platform !== 'win32') {
|
||||
assert.equal(fs.statSync(dst).mode & 0o777, 0o644);
|
||||
}
|
||||
} finally {
|
||||
fs.rmSync(dir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test('seed main: missing host file still writes a valid onboarding-bearing file', () => {
|
||||
const dir = tmp();
|
||||
try {
|
||||
const dst = path.join(dir, 'out.claude.json');
|
||||
execFileSync(process.execPath, [SEED_SCRIPT, path.join(dir, 'nope.json'), dst]);
|
||||
assert.deepEqual(JSON.parse(fs.readFileSync(dst, 'utf8')), { hasCompletedOnboarding: true });
|
||||
} finally {
|
||||
fs.rmSync(dir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test('seed main: chmodSync widens a pre-existing restrictive dst to 0o644', () => {
|
||||
// This checks the file's permission bits, which only exist on POSIX systems.
|
||||
//
|
||||
// The catch: CI's default umask is 022, so a plain writeFileSync already
|
||||
// creates files at mode 0o644. Asserting 0o644 right after a fresh write
|
||||
// would therefore NOT prove the explicit chmodSync did anything.
|
||||
//
|
||||
// So we pre-create dst at the stricter mode 0o600. Opening a file in 'w'
|
||||
// mode replaces its contents but KEEPS the mode of a file that already
|
||||
// exists. That means the only way dst can end up at 0o644 is the chmodSync
|
||||
// inside seed-claude-config.cjs. This pins the test to the chmod and not to
|
||||
// the umask: delete the chmodSync line and this test fails, while the other
|
||||
// seed test still passes.
|
||||
if (process.platform === 'win32') return;
|
||||
const dir = tmp();
|
||||
try {
|
||||
const src = path.join(dir, 'host.claude.json');
|
||||
const dst = path.join(dir, 'out.claude.json');
|
||||
fs.writeFileSync(src, JSON.stringify({ userID: 'u' }));
|
||||
fs.writeFileSync(dst, '{}');
|
||||
fs.chmodSync(dst, 0o600);
|
||||
execFileSync(process.execPath, [SEED_SCRIPT, src, dst]);
|
||||
assert.equal(fs.statSync(dst).mode & 0o777, 0o644);
|
||||
assert.equal(JSON.parse(fs.readFileSync(dst, 'utf8')).userID, 'u');
|
||||
} finally {
|
||||
fs.rmSync(dir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
// --- ensurePaths: sets up the host paths the bind mounts point at -----------
|
||||
|
||||
test('ensurePaths: creates every DIR and FILE under a temp home, idempotently', () => {
|
||||
const home = tmp();
|
||||
try {
|
||||
ensurePaths(home);
|
||||
for (const d of DIRS) {
|
||||
assert.equal(fs.statSync(path.join(home, d)).isDirectory(), true, `not a dir: ${d}`);
|
||||
}
|
||||
for (const f of FILES) {
|
||||
assert.equal(fs.statSync(path.join(home, f)).isFile(), true, `not a file: ${f}`);
|
||||
}
|
||||
// Running it again must not throw and must not overwrite existing content.
|
||||
fs.writeFileSync(path.join(home, '.claude.json'), '{"keep":true}');
|
||||
ensurePaths(home);
|
||||
assert.equal(fs.readFileSync(path.join(home, '.claude.json'), 'utf8'), '{"keep":true}');
|
||||
} finally {
|
||||
fs.rmSync(home, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test('ensurePaths: does NOT pre-create settings.json / config.toml (no gratuitous host mutation)', () => {
|
||||
const home = tmp();
|
||||
try {
|
||||
ensurePaths(home);
|
||||
assert.equal(fs.existsSync(path.join(home, '.claude', 'settings.json')), false);
|
||||
assert.equal(fs.existsSync(path.join(home, '.codex', 'config.toml')), false);
|
||||
} finally {
|
||||
fs.rmSync(home, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test('ensurePaths: does NOT pre-create the shareable subdirs (now copied, not bound)', () => {
|
||||
// The shareable dirs are seeded into the per-container volume from the
|
||||
// read-only /host stage, so they are no longer bind-mount sources. Pre-creating
|
||||
// empty ones would needlessly write into the host of someone who never used a
|
||||
// CLI. This pins the DIRS trim: re-adding any of these would fail the test.
|
||||
const home = tmp();
|
||||
const mustNotExist = [
|
||||
path.join('.claude', 'skills'),
|
||||
path.join('.claude', 'agents'),
|
||||
path.join('.claude', 'memory'),
|
||||
path.join('.claude', 'commands'),
|
||||
path.join('.claude', 'plugins'),
|
||||
path.join('.codex', 'plugins'),
|
||||
path.join('.codex', 'prompts'),
|
||||
path.join('.codex', 'memories'),
|
||||
path.join('.codex', 'skills'),
|
||||
path.join('.cursor', 'rules'),
|
||||
path.join('.cursor', 'commands'),
|
||||
path.join('.cursor', 'agents'),
|
||||
path.join('.cursor', 'skills'),
|
||||
path.join('.cursor', 'plugins'),
|
||||
];
|
||||
try {
|
||||
ensurePaths(home);
|
||||
for (const sub of mustNotExist) {
|
||||
assert.equal(fs.existsSync(path.join(home, sub)), false, `should not pre-create: ${sub}`);
|
||||
}
|
||||
// The top-level stage roots that ARE still bind sources must exist.
|
||||
for (const top of ['.claude', '.codex', '.cursor', '.claude-mem']) {
|
||||
assert.equal(fs.statSync(path.join(home, top)).isDirectory(), true, `missing root: ${top}`);
|
||||
}
|
||||
} finally {
|
||||
fs.rmSync(home, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
@@ -0,0 +1,19 @@
|
||||
description = "GitNexus production-readiness PR swarm review (Solo mode)"
|
||||
|
||||
prompt = """
|
||||
You are the GitNexus PR review coordinator. Review this pull request: {{args}}
|
||||
(a PR URL or number for https://github.com/abhigyanpatwari/GitNexus). If no target was
|
||||
given, ask for one.
|
||||
|
||||
Read `pr-swarm-review/orchestration.md` in this repository and follow it exactly. It is the
|
||||
canonical, CLI-neutral review contract (lanes, classifications, output structure, finding
|
||||
format, hidden-Unicode checks, behavior rules).
|
||||
|
||||
Run in **Solo mode**: you are a single agent, so perform all seven lanes yourself in
|
||||
dependency order, adopting each persona in `pr-swarm-review/personas/0N-*.md` in turn
|
||||
(lanes 1-2 first, then 3-6, then lane 7). Keep every lane's findings in context. Lane 7
|
||||
(synthesis critic) is a hard gate: do not emit the final review until its "Required
|
||||
corrections before posting" section is empty — revise and re-run it otherwise.
|
||||
|
||||
Stay strictly read-only: investigate and report; never edit files, commit, or post to GitHub.
|
||||
"""
|
||||
@@ -1,2 +1,17 @@
|
||||
* text=auto eol=lf
|
||||
.husky/* text eol=lf
|
||||
|
||||
# Shell scripts: force LF unconditionally so devcontainer scripts
|
||||
# (e.g. anything COPYed into a Linux container) execute correctly when
|
||||
# checked out on Windows hosts with core.autocrlf=true.
|
||||
*.sh text eol=lf
|
||||
*.bash text eol=lf
|
||||
|
||||
# Native and binary assets shouldn't be treated as text under any
|
||||
# auto-detection or eol normalization.
|
||||
*.node binary
|
||||
*.wasm binary
|
||||
*.onnx binary
|
||||
*.so binary
|
||||
*.dll binary
|
||||
*.dylib binary
|
||||
|
||||
@@ -0,0 +1,19 @@
|
||||
---
|
||||
description: 'GitNexus production-readiness PR swarm review (Solo mode)'
|
||||
mode: 'agent'
|
||||
---
|
||||
|
||||
You are the GitNexus PR review coordinator. Review the pull request the user names (a PR URL
|
||||
or number for `https://github.com/abhigyanpatwari/GitNexus`). If none was given, ask for one.
|
||||
|
||||
Read `pr-swarm-review/orchestration.md` in this repository and follow it exactly — it is the
|
||||
canonical, CLI-neutral review contract (lanes, classifications, output structure, finding
|
||||
format, hidden-Unicode checks, behavior rules).
|
||||
|
||||
Run in **Solo mode**: you are a single agent, so perform all seven lanes yourself in
|
||||
dependency order, adopting each persona in `pr-swarm-review/personas/0N-*.md` in turn
|
||||
(lanes 1–2 first, then 3–6, then lane 7). Keep every lane's findings in context. Lane 7
|
||||
(synthesis critic) is a hard gate: do not emit the final review until its "Required
|
||||
corrections before posting" section is empty.
|
||||
|
||||
Stay strictly read-only: investigate and report; never edit files, commit, or post to GitHub.
|
||||
@@ -332,6 +332,111 @@ def vendored_drift_summary(
|
||||
}
|
||||
|
||||
|
||||
# ── Assert mode (CI gate) ─────────────────────────────────────────────────
|
||||
|
||||
|
||||
def assert_current() -> int:
|
||||
"""Assert every grammar's ABI is loadable by the CURRENT runtime.
|
||||
|
||||
Unlike the readiness report (which probes the npm registry + upstream
|
||||
main for the *target* runtime), this mode is hermetic and offline: it
|
||||
reads only what's checked out / installed locally and asserts each
|
||||
grammar's compiled ABI lies within the current runtime's
|
||||
``RUNTIME_ABI_RANGES`` window. It is the static half of the #1922 ABI
|
||||
gate; the runtime load-smoke (`parser-loader-abi.test.ts`) is the
|
||||
dynamic half.
|
||||
|
||||
Coverage, reusing the existing helpers:
|
||||
- npm-installed grammars: ABI from node_modules/<name>/<parser.c>.
|
||||
- vendored grammars (dart/proto/swift): ABI via ``vendored_drift_summary``.
|
||||
- Swift is prebuilt-only (no parser.c) → not introspectable here;
|
||||
treated as "covered by the runtime load-smoke", not asserted.
|
||||
- INTENTIONAL_PINS are honored: a pinned grammar is expected to sit at
|
||||
an ABI the current runtime loads (that's *why* it's pinned), so it is
|
||||
asserted like any other rather than skipped.
|
||||
|
||||
Returns 0 when every introspectable grammar is in range, 1 otherwise.
|
||||
Prints a plain-text (non-Markdown) report so CI logs stay readable.
|
||||
"""
|
||||
current_runtime = read_current_runtime()
|
||||
abi_range = RUNTIME_ABI_RANGES.get(current_runtime)
|
||||
if abi_range is None:
|
||||
print(
|
||||
f"FAIL: RUNTIME_ABI_RANGES has no entry for current runtime "
|
||||
f"{current_runtime!r}; add it before asserting.",
|
||||
)
|
||||
return 1
|
||||
lo, hi = abi_range
|
||||
pinned_versions = read_pinned_grammar_versions()
|
||||
|
||||
print(
|
||||
f"Asserting all grammar ABIs load on tree-sitter@{current_runtime}.x "
|
||||
f"(ABI {lo}–{hi})."
|
||||
)
|
||||
|
||||
failures: list[str] = []
|
||||
checked = 0
|
||||
skipped: list[str] = []
|
||||
|
||||
for name, (upstream_repo, upstream_branch, parser_path) in sorted(GRAMMARS.items()):
|
||||
pinned_spec = pinned_versions.get(name, "—")
|
||||
pin_note = f" [intentional pin: {pinned_spec}]" if name in INTENTIONAL_PINS else ""
|
||||
|
||||
if is_vendored_pin(pinned_spec):
|
||||
v = vendored_drift_summary(name, upstream_repo, upstream_branch, parser_path)
|
||||
abi = v["vendored_abi"]
|
||||
if abi is None:
|
||||
# Prebuilt-only vendor (e.g. tree-sitter-swift): no parser.c to
|
||||
# introspect. The runtime load-smoke covers it instead.
|
||||
skipped.append(f"{name} (vendored, prebuilt — covered by load-smoke)")
|
||||
continue
|
||||
checked += 1
|
||||
if lo <= abi <= hi:
|
||||
print(f" OK {name}: vendored ABI {abi} in range{pin_note}")
|
||||
else:
|
||||
msg = (
|
||||
f"{name}: vendored ABI {abi} outside current runtime range "
|
||||
f"{lo}..{hi}{pin_note}"
|
||||
)
|
||||
print(f" FAIL {msg}")
|
||||
failures.append(msg)
|
||||
continue
|
||||
|
||||
installed_parser = GITNEXUS_DIR / "node_modules" / name / parser_path
|
||||
if not installed_parser.is_file():
|
||||
installed_parser = GITNEXUS_DIR / "node_modules" / name / "src" / "parser.c"
|
||||
abi = extract_language_version(installed_parser)
|
||||
if abi is None:
|
||||
skipped.append(f"{name} (not installed / no parser.c — covered by load-smoke)")
|
||||
continue
|
||||
checked += 1
|
||||
if lo <= abi <= hi:
|
||||
print(f" OK {name}: installed ABI {abi} in range{pin_note}")
|
||||
else:
|
||||
msg = (
|
||||
f"{name}: installed ABI {abi} outside current runtime range "
|
||||
f"{lo}..{hi}{pin_note}"
|
||||
)
|
||||
print(f" FAIL {msg}")
|
||||
failures.append(msg)
|
||||
|
||||
print("")
|
||||
if skipped:
|
||||
print("Not statically introspectable (asserted via runtime load-smoke):")
|
||||
for s in skipped:
|
||||
print(f" - {s}")
|
||||
print("")
|
||||
|
||||
if failures:
|
||||
print(f"RESULT: FAIL — {len(failures)} grammar(s) out of range, {checked} checked.")
|
||||
for f in failures:
|
||||
print(f" - {f}")
|
||||
return 1
|
||||
|
||||
print(f"RESULT: OK — all {checked} introspectable grammar ABIs in range.")
|
||||
return 0
|
||||
|
||||
|
||||
# ── Main ────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
@@ -806,4 +911,9 @@ if __name__ == "__main__":
|
||||
sys.stdout.reconfigure(encoding="utf-8") # type: ignore[attr-defined]
|
||||
except Exception:
|
||||
pass
|
||||
# `--assert-current` is the offline CI gate (#1922): assert every grammar's
|
||||
# ABI loads on the CURRENT runtime. Bare invocation keeps the original
|
||||
# target-runtime readiness report behaviour.
|
||||
if "--assert-current" in sys.argv[1:]:
|
||||
sys.exit(assert_current())
|
||||
sys.exit(main())
|
||||
|
||||
@@ -0,0 +1,127 @@
|
||||
name: Devcontainer Smoke
|
||||
|
||||
# Smoke-tests .devcontainer/ whenever it changes. Two things happen here.
|
||||
# First, unit tests run on the pure host->container config transforms: the
|
||||
# plugin-registry path translation, and the strip of the machine field from
|
||||
# $HOME/.claude.json. Second, the devcontainer image is built through the
|
||||
# standard @devcontainers/cli path. That CLI reads build.args from
|
||||
# devcontainer.json, so the version pin there stays the single source of truth.
|
||||
on:
|
||||
push:
|
||||
branches: [main]
|
||||
paths:
|
||||
- '.devcontainer/**'
|
||||
- '.github/workflows/ci-devcontainer.yml'
|
||||
pull_request:
|
||||
paths:
|
||||
- '.devcontainer/**'
|
||||
- '.github/workflows/ci-devcontainer.yml'
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
# Concurrency convention: see CONTRIBUTING.md → "GitHub Actions — Concurrency Convention".
|
||||
# Grouped per branch or tag. Cancel a PR run when a newer one replaces it.
|
||||
# Never cancel a push-to-main run.
|
||||
concurrency:
|
||||
group: ${{ github.workflow }}-${{ github.ref }}
|
||||
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
|
||||
|
||||
jobs:
|
||||
config-transforms:
|
||||
name: Config-transform unit tests
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 5
|
||||
steps:
|
||||
# persist-credentials: false — this job only reads (tests and syntax
|
||||
# checks) and never pushes. The setting keeps GITHUB_TOKEN out of
|
||||
# .git/config, which zizmor flags as the "artipacked" issue.
|
||||
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
||||
with:
|
||||
persist-credentials: false
|
||||
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
|
||||
with:
|
||||
node-version: 22
|
||||
- name: Unit-test the host->container config transforms
|
||||
run: node --test .devcontainer/translate-plugin-registries.test.cjs
|
||||
- name: Syntax-check the lifecycle shell scripts
|
||||
run: |
|
||||
bash -n .devcontainer/install-deps.sh
|
||||
bash -n .devcontainer/post-create.sh
|
||||
|
||||
build:
|
||||
name: Build devcontainer image
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 30
|
||||
steps:
|
||||
# persist-credentials: false — this is a read-only build smoke that
|
||||
# never pushes. The setting keeps GITHUB_TOKEN out of .git/config,
|
||||
# which zizmor flags as the "artipacked" issue.
|
||||
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
||||
with:
|
||||
persist-credentials: false
|
||||
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
|
||||
with:
|
||||
node-version: 22
|
||||
# Builds the image the same way a developer's "Reopen in Container" does.
|
||||
# @devcontainers/cli reads devcontainer.json (jsonc format), resolves
|
||||
# build.args (the CLAUDE_CODE_VERSION / CODEX_VERSION pins), and runs the
|
||||
# Dockerfile. This smoke catches Dockerfile regressions and any drift from
|
||||
# the canonical version pins. The lifecycle hooks (post-create.sh) do not
|
||||
# run here. They need the host config mounts, and CI has none.
|
||||
#
|
||||
# ARCH COVERAGE: this runs on an x64 runner with no --platform or QEMU, so
|
||||
# it builds only the amd64 Cursor branch (CURSOR_SHA256_X64). The arm64
|
||||
# branch (CURSOR_SHA256_ARM64 plus the arm64 tarball URL) is pinned by a
|
||||
# sha256 checked against the published artifact, but it is not BUILT here.
|
||||
# Cursor's extract-and-symlink step does not depend on the architecture, so
|
||||
# the only remaining gap is a stale arm64 URL or hash. If that becomes a
|
||||
# concern, add a linux/arm64 matrix leg (docker/setup-qemu-action plus
|
||||
# `--platform`).
|
||||
#
|
||||
# The @devcontainers/cli version is pinned on purpose. A bare
|
||||
# `npx --yes @devcontainers/cli` would resolve @latest at run time. A
|
||||
# breaking or malicious publish could then change CI behavior, or change
|
||||
# how devcontainer.json is read, with no diff to show for it. Bump this pin
|
||||
# deliberately, alongside the Dockerfile and devcontainer.json pins.
|
||||
#
|
||||
# @devcontainers/cli wraps the Dockerfile with `# syntax=docker/dockerfile:1`,
|
||||
# which BuildKit resolves from Docker Hub. Hub blips surface as
|
||||
# `DeadlineExceeded` / `i/o timeout` on the syntax frontend (see run
|
||||
# 26797815133). Build retry (2 attempts, 45s backoff) matches
|
||||
# `.github/actions/docker-build-push-retry` (docker/build-push-action#1422).
|
||||
# Pre-pull of docker/dockerfile:1 is extra hardening; best-effort so the
|
||||
# build retry still runs if Hub is flaky only during pull.
|
||||
- name: Pre-pull BuildKit Dockerfile frontend (retry)
|
||||
continue-on-error: true
|
||||
run: |
|
||||
set -euo pipefail
|
||||
img="docker/dockerfile:1"
|
||||
for attempt in 1 2 3; do
|
||||
if docker pull "$img"; then
|
||||
exit 0
|
||||
fi
|
||||
echo "::warning::docker pull ${img} attempt ${attempt} failed"
|
||||
if [ "$attempt" -lt 3 ]; then
|
||||
sleep $((attempt * 15))
|
||||
fi
|
||||
done
|
||||
echo "::warning::failed to pre-pull ${img} after 3 attempts; continuing — build step may still succeed"
|
||||
exit 1
|
||||
- name: Build devcontainer via @devcontainers/cli
|
||||
run: |
|
||||
set -euo pipefail
|
||||
for attempt in 1 2; do
|
||||
if npx --yes @devcontainers/cli@0.87.0 build --workspace-folder .; then
|
||||
if [ "$attempt" -eq 2 ]; then
|
||||
echo "::notice::devcontainer build retry succeeded (attempt 2); investigate if this recurs across runs."
|
||||
fi
|
||||
exit 0
|
||||
fi
|
||||
if [ "$attempt" -eq 2 ]; then
|
||||
echo "::error::devcontainer build failed after 2 attempts"
|
||||
exit 1
|
||||
fi
|
||||
echo "::warning::devcontainer build attempt ${attempt} failed; retrying in 45s…"
|
||||
sleep 45
|
||||
done
|
||||
@@ -12,7 +12,13 @@ jobs:
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 25
|
||||
steps:
|
||||
# persist-credentials: false — this job runs tests and uploads a
|
||||
# test-reports artifact (if: always()). The default-persisted token in
|
||||
# .git/config must not be capturable through that upload (zizmor
|
||||
# credential-persistence / artipacked audit). The job never pushes.
|
||||
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
||||
with:
|
||||
persist-credentials: false
|
||||
- uses: ./.github/actions/setup-gitnexus
|
||||
with:
|
||||
build: 'true'
|
||||
@@ -72,7 +78,11 @@ jobs:
|
||||
runs-on: ${{ matrix.os }}
|
||||
timeout-minutes: 20
|
||||
steps:
|
||||
# persist-credentials: false — runs tests only, never pushes (zizmor
|
||||
# credential-persistence / artipacked audit).
|
||||
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
||||
with:
|
||||
persist-credentials: false
|
||||
- uses: ./.github/actions/setup-gitnexus
|
||||
with:
|
||||
build: 'true'
|
||||
@@ -80,6 +90,34 @@ jobs:
|
||||
run: npx tsx scripts/run-cross-platform.ts
|
||||
working-directory: gitnexus
|
||||
|
||||
# Tree-sitter ABI gate (#1922). Two halves, both blocking:
|
||||
# 1. Static, offline: assert every grammar's compiled ABI loads on the
|
||||
# pinned runtime (check-tree-sitter-upgrade-readiness.py --assert-current).
|
||||
# 2. Dynamic: run the parser-loader ABI load-smoke on the OS matrix so an
|
||||
# ABI-incompatible prebuilt (esp. the binary-only Swift vendor, which the
|
||||
# static check can't introspect) fails on the platform it ships to.
|
||||
abi-assert:
|
||||
name: tree-sitter ABI (${{ matrix.os }})
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
os: [ubuntu-latest, windows-latest, macos-latest]
|
||||
runs-on: ${{ matrix.os }}
|
||||
timeout-minutes: 20
|
||||
steps:
|
||||
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
||||
- uses: ./.github/actions/setup-gitnexus
|
||||
with:
|
||||
build: 'true'
|
||||
|
||||
- name: Assert installed + vendored grammar ABIs (static)
|
||||
shell: bash
|
||||
run: python3 .github/scripts/check-tree-sitter-upgrade-readiness.py --assert-current
|
||||
|
||||
- name: Run parser-loader ABI load-smoke (dynamic)
|
||||
run: npx vitest run test/unit/parser-loader-abi.test.ts
|
||||
working-directory: gitnexus
|
||||
|
||||
# End-to-end smoke test for the #1728 packaging fix: pack the published
|
||||
# tarball, install it globally into a temp prefix, and assert no junction
|
||||
# creation (the EPERM root cause) plus working CLI plus vendor cleanliness
|
||||
@@ -181,3 +219,60 @@ jobs:
|
||||
else
|
||||
"$PREFIX/bin/gitnexus" --version
|
||||
fi
|
||||
|
||||
# ── Dedicated benchmark gate ─────────────────────────────────────
|
||||
# The cross-language `*-pipeline-benchmark.test.ts` suites are gated behind
|
||||
# GITNEXUS_BENCH (they generate synthetic codebases at scale), so the main
|
||||
# coverage job above SKIPS them — their O(n^2) scaling guards never ran in CI.
|
||||
# Run them here with GITNEXUS_BENCH=1, alongside the Python scope-capture and
|
||||
# import-resolution fingerprint + scaling guards (PR #1918 P2a).
|
||||
#
|
||||
# `--no-file-parallelism` is REQUIRED: these suites measure wall-clock and peak
|
||||
# heap, so parallel forks both skew the timings and OOM the worker pool — they
|
||||
# must run one file at a time.
|
||||
#
|
||||
# go-pipeline-benchmark.test.ts is deliberately NOT included: its
|
||||
# worker-pool (#1848) suite spins a real worker pool that exits unexpectedly
|
||||
# under vitest's fork pool (reproduced in validation), which would make this
|
||||
# gate flaky. Go is already guarded by its non-gated O(n^2) tripwire (runs in
|
||||
# the main coverage job) plus its golden capture-parity test.
|
||||
benchmarks:
|
||||
name: benchmarks (GITNEXUS_BENCH)
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 25
|
||||
steps:
|
||||
# persist-credentials: false — this job only runs npm + vitest benchmarks
|
||||
# and never pushes; the default-persisted token in .git/config would be at
|
||||
# risk of leaking through an artifact upload (zizmor credential-persistence
|
||||
# / artipacked audit). Mirrors the packaged-install-smoke job below.
|
||||
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
||||
with:
|
||||
persist-credentials: false
|
||||
- uses: ./.github/actions/setup-gitnexus
|
||||
with:
|
||||
build: 'true'
|
||||
|
||||
- name: Python scope-capture + import-resolution fingerprint / scaling guards
|
||||
run: |
|
||||
node --import tsx bench/python-scope/measure.mjs --check
|
||||
node --import tsx bench/python-scope/import-target-fingerprint.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: Cross-language scope-capture fingerprint + scaling guards
|
||||
# Build-free: asserts emit<Lang>ScopeCaptures output is unchanged
|
||||
# (fingerprint) and stays linear (scaling < 1.5) for go/csharp/rust/php/
|
||||
# ruby/cobol. Catches an O(n^2) re-regression without the worker pool.
|
||||
run: node --import tsx bench/scope-capture/measure.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: Cross-language pipeline benchmarks (GITNEXUS_BENCH, serial)
|
||||
env:
|
||||
GITNEXUS_BENCH: '1'
|
||||
run: >-
|
||||
npx vitest run --no-file-parallelism
|
||||
test/integration/cobol-pipeline-benchmark.test.ts
|
||||
test/integration/csharp-pipeline-benchmark.test.ts
|
||||
test/integration/rust-pipeline-benchmark.test.ts
|
||||
test/integration/php-pipeline-benchmark.test.ts
|
||||
test/integration/ruby-pipeline-benchmark.test.ts
|
||||
working-directory: gitnexus
|
||||
|
||||
@@ -112,6 +112,11 @@ jobs:
|
||||
shell: bash
|
||||
env:
|
||||
QUALITY: ${{ needs.quality.result }}
|
||||
# The tree-sitter ABI gate (#1922) runs as the `abi-assert` job
|
||||
# inside the `tests` reusable workflow. A failed job fails the
|
||||
# reusable workflow, so `needs.tests.result` below blocks the merge
|
||||
# on an ABI mismatch. (`jobs.<id>.result` cannot be exposed as a
|
||||
# workflow_call output, so the gate is enforced transitively here.)
|
||||
TESTS: ${{ needs.tests.result }}
|
||||
E2E: ${{ needs.e2e.result }}
|
||||
SCOPE_PARITY: ${{ needs.scope-parity.result }}
|
||||
@@ -120,9 +125,12 @@ jobs:
|
||||
echo "Tests: $TESTS"
|
||||
echo "E2E: $E2E"
|
||||
echo "Scope parity: $SCOPE_PARITY"
|
||||
# A failed `abi-assert` job (#1922) inside the tests reusable
|
||||
# workflow makes TESTS != success, so this clause also blocks the
|
||||
# merge on a tree-sitter ABI mismatch.
|
||||
if [[ "$QUALITY" != "success" ]] ||
|
||||
[[ "$TESTS" != "success" ]]; then
|
||||
echo "::error::Quality or test jobs failed"
|
||||
echo "::error::Quality or test jobs failed (includes the tree-sitter ABI gate, #1922)"
|
||||
exit 1
|
||||
fi
|
||||
if [[ "$E2E" != "success" && "$E2E" != "skipped" ]]; then
|
||||
|
||||
+4
-2
@@ -91,11 +91,13 @@ gitnexus/vendor/**/node_modules/
|
||||
|
||||
.claude-flow/
|
||||
|
||||
.claude/agents/
|
||||
.claude/agents/*
|
||||
!.claude/agents/gitnexus-*.md
|
||||
.claude/commands/
|
||||
.claude/helpers
|
||||
.claude/skills/
|
||||
.claude/skills/*
|
||||
!.claude/skills/gitnexus/
|
||||
!.claude/skills/gitnexus-pr-swarm-review/
|
||||
|
||||
.history/
|
||||
|
||||
|
||||
@@ -44,6 +44,18 @@ Commands and gotchas live under **Repo reference** below and in **[CONTRIBUTING.
|
||||
- **Cursor:** `.cursor/index.mdc` (always-on); `.cursor/rules/*.mdc` (glob-scoped). Legacy `.cursorrules` deprecated.
|
||||
- **GitNexus:** skills in `.claude/skills/gitnexus/`; MCP rules in `gitnexus:start` block below.
|
||||
|
||||
## PR Swarm Review (cross-CLI)
|
||||
|
||||
To run a production-readiness review of a GitNexus pull request from **any** AI CLI, follow
|
||||
the canonical, CLI-neutral spec **[`pr-swarm-review/orchestration.md`](pr-swarm-review/orchestration.md)**
|
||||
(seven read-only review personas under `pr-swarm-review/personas/`). It defines two
|
||||
execution modes with the same output contract: **Swarm mode** (parallel subagents, e.g.
|
||||
Claude Code) and **Solo mode** (one agent runs all lanes sequentially — Codex, Gemini,
|
||||
Cursor, Copilot, or any agent reading this file). Per-CLI entrypoints are thin wrappers
|
||||
listed in [`pr-swarm-review/README.md`](pr-swarm-review/README.md); edit review logic only
|
||||
in the canonical files, never in the wrappers. The review is read-only — it never edits,
|
||||
commits, or posts.
|
||||
|
||||
## Changelog
|
||||
|
||||
| Date | Version | Change |
|
||||
@@ -65,7 +77,7 @@ Commands and gotchas live under **Repo reference** below and in **[CONTRIBUTING.
|
||||
|
||||
This project is indexed by GitNexus as **GitNexus** (26675 symbols, 35395 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
|
||||
|
||||
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
|
||||
> Index stale? Run `node .gitnexus/run.cjs analyze` from the project root — it auto-selects an available runner. No `.gitnexus/run.cjs` yet? `npx gitnexus analyze` (npm 11 crash → `npm i -g gitnexus`; #1939).
|
||||
|
||||
## Always Do
|
||||
|
||||
|
||||
@@ -58,7 +58,7 @@ See the `<!-- gitnexus:start --> … <!-- gitnexus:end -->` block in **[AGENTS.m
|
||||
|
||||
This project is indexed by GitNexus as **GitNexus** (26675 symbols, 35395 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
|
||||
|
||||
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
|
||||
> Index stale? Run `node .gitnexus/run.cjs analyze` from the project root — it auto-selects an available runner. No `.gitnexus/run.cjs` yet? `npx gitnexus analyze` (npm 11 crash → `npm i -g gitnexus`; #1939).
|
||||
|
||||
## Always Do
|
||||
|
||||
|
||||
@@ -18,6 +18,10 @@ This project uses the [PolyForm Noncommercial License 1.0.0](https://polyformpro
|
||||
3. **Web UI (if needed):** `cd gitnexus-web && npm install`
|
||||
4. Run tests as described in [TESTING.md](TESTING.md).
|
||||
|
||||
### Containerized development (optional)
|
||||
|
||||
If you prefer an isolated environment with Claude Code, OpenAI Codex CLI, and Cursor CLI pre-installed, open the repo in VS Code with the [Dev Containers extension](https://marketplace.visualstudio.com/items?itemName=ms-vscode-remote.remote-containers) and run **Dev Containers: Reopen in Container**. See [`.devcontainer/README.md`](.devcontainer/README.md) for first-time auth flows and Windows WSL2 setup.
|
||||
|
||||
## Branch and pull requests
|
||||
|
||||
- Use short-lived branches off the default branch of the repo you are targeting.
|
||||
|
||||
@@ -104,6 +104,14 @@ npx gitnexus analyze
|
||||
|
||||
That's it. This indexes the codebase, installs agent skills, registers Claude Code hooks, and creates `AGENTS.md` / `CLAUDE.md` context files — all in one command.
|
||||
|
||||
> **On npm 11.x?** `npx` can crash during install with `Cannot destructure property 'package' of 'node.target'` (an npm/arborist bug, before GitNexus runs). Use pnpm instead — it builds the native deps explicitly:
|
||||
>
|
||||
> ```bash
|
||||
> pnpm --allow-build=@ladybugdb/core --allow-build=gitnexus --allow-build=tree-sitter dlx gitnexus@latest analyze
|
||||
> ```
|
||||
>
|
||||
> Or install globally (`npm install -g gitnexus@latest`) and run `gitnexus analyze`. See [#1939](https://github.com/abhigyanpatwari/GitNexus/issues/1939).
|
||||
|
||||
To configure MCP for your editor, run `npx gitnexus setup` once — or set it up manually below.
|
||||
|
||||
> **Faster install (no C++ toolchain needed):** set `GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1` before `npm install -g gitnexus` to skip vendored grammar materialize/build (`tree-sitter-dart`, `tree-sitter-proto`, `tree-sitter-swift`). Dart/Proto/Swift files won't be parsed, but install completes in seconds without `python3`/`make`/`g++`. Strict `=1` only — any other value falls through to the rebuild.
|
||||
@@ -219,7 +227,7 @@ gitnexus analyze --skills # Generate repo-specific skill files from detec
|
||||
gitnexus analyze --skip-embeddings # Skip embedding generation (faster)
|
||||
gitnexus analyze --skip-agents-md # Preserve custom AGENTS.md/CLAUDE.md gitnexus section edits
|
||||
gitnexus analyze --skip-git # Index folders that are not Git repositories
|
||||
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
|
||||
gitnexus analyze --embeddings [limit] # Enable embedding generation (slower, better search)
|
||||
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
|
||||
gitnexus analyze --worker-timeout 60 # Increase worker idle timeout for slow parses
|
||||
gitnexus analyze --wal-checkpoint-threshold 67108864 # 64 MiB. Control LadybugDB WAL auto-checkpoint threshold (default: 67108864 = 64 MiB; -1 keeps Ladybug stock ~16 MiB)
|
||||
@@ -248,6 +256,23 @@ gitnexus group status <name> # Check staleness of repos in a group
|
||||
|
||||
If `analyze` reports a worker parse timeout on a large or unusual repository, it keeps running and falls back safely. To give slow worker jobs more time, use `gitnexus analyze --worker-timeout 60` or set `GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS=60000`. For very large files, `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` controls the worker job byte budget.
|
||||
|
||||
#### Embeddings node limit
|
||||
|
||||
`gitnexus analyze --embeddings` generates semantic search vectors with a default 50,000-node safety cap to protect memory on large repositories. Override the cap when you know the host has enough memory for a larger graph, or disable it entirely for a one-off full embeddings run.
|
||||
|
||||
```bash
|
||||
# Generate embeddings with the default 50,000 node safety cap
|
||||
gitnexus analyze --embeddings
|
||||
|
||||
# Disable the safety cap entirely
|
||||
gitnexus analyze --embeddings 0
|
||||
|
||||
# Use a custom cap
|
||||
gitnexus analyze --embeddings 100000
|
||||
```
|
||||
|
||||
If embeddings are skipped on a large repository, the indexed graph likely exceeds the default safety cap. Re-run with `gitnexus analyze --embeddings 0` to remove the cap, or `gitnexus analyze --embeddings <n>` to choose a higher limit while still keeping memory bounded.
|
||||
|
||||
#### Environment variables
|
||||
|
||||
Most `analyze` knobs are also CLI flags (`--workers`, `--worker-timeout`, `--max-file-size`, `--verbose`). Use the env-var form when you'd otherwise repeat the same flag every run, or when invoking GitNexus from a long-running host (MCP server, eval-server, CI shell) that already manages its own environment. CLI flags take precedence over env vars; env vars take precedence over built-in defaults.
|
||||
@@ -660,6 +685,14 @@ UPSTREAM (what depends on this):
|
||||
|
||||
Options: `maxDepth`, `minConfidence`, `relationTypes` (`CALLS`, `IMPORTS`, `EXTENDS`, `IMPLEMENTS`), `includeTests`, `limit` (max symbols per depth, default 100), `offset` (pagination start per depth), `summaryOnly` (counts and risk only, omits symbol list)
|
||||
|
||||
**Disambiguation** — when several symbols share the target name, `impact` returns a ranked `ambiguous` candidate list instead of guessing. Narrow it with `target_uid` (exact, zero-ambiguity), `file_path`, or `kind` (`Function`, `Class`, `Method`, …). From the CLI these are `--uid`, `--file`, and `--kind`, matching `gitnexus context`:
|
||||
|
||||
```bash
|
||||
gitnexus impact get_embeddings # → ambiguous: lists ranked candidates
|
||||
gitnexus impact get_embeddings --file src/embed.py # → resolves to the one in that file
|
||||
gitnexus impact get_embeddings --uid "Function:src/embed.py:get_embeddings" # exact
|
||||
```
|
||||
|
||||
### Process-Grouped Search
|
||||
|
||||
```
|
||||
|
||||
@@ -16,6 +16,7 @@ const path = require('path');
|
||||
const { spawnSync } = require('child_process');
|
||||
const { acquireHookSlot } = require('./hook-lock.js');
|
||||
const { hasGitNexusDbLockedByGitNexusServer } = require('./hook-db-lock-probe.cjs');
|
||||
const { formatAnalyzeCommand } = require('./resolve-analyze-cmd.cjs');
|
||||
|
||||
/**
|
||||
* Read JSON input from stdin synchronously.
|
||||
@@ -345,7 +346,7 @@ function handlePostToolUse(input) {
|
||||
// If HEAD matches last indexed commit, no reindex needed
|
||||
if (currentHead && currentHead === lastCommit) return;
|
||||
|
||||
const analyzeCmd = `npx gitnexus analyze${hadEmbeddings ? ' --embeddings' : ''}`;
|
||||
const analyzeCmd = formatAnalyzeCommand({ embeddings: hadEmbeddings });
|
||||
sendHookResponse(
|
||||
'PostToolUse',
|
||||
`GitNexus index is stale (last indexed: ${lastCommit ? lastCommit.slice(0, 7) : 'never'}). ` +
|
||||
|
||||
@@ -0,0 +1,323 @@
|
||||
/**
|
||||
* Single source of truth for how docs, hooks, and warnings invoke gitnexus.
|
||||
*
|
||||
* Automatically selects a working invocation path:
|
||||
* 1. Global `gitnexus` on PATH (best — no install step)
|
||||
* 2. npm 11+ with pnpm on PATH → `pnpm --allow-build=… dlx` (avoids the npx
|
||||
* arborist crash *and* pnpm 10+ ignored-build-script failures, #1939)
|
||||
* 3. npm < 11 with npm on PATH → `npx` (works; simpler than pnpm dlx)
|
||||
* 4. pnpm-only → `pnpm --allow-build=… dlx`
|
||||
* 5. Last resort → `npx` (warned on npm 11+ from analyze.ts)
|
||||
*
|
||||
* The `--allow-build` flags MUST precede the `dlx` token. pnpm < 10.14 keeps
|
||||
* `dlx` in its argv escape list, so flags placed *after* `dlx` are parsed as
|
||||
* package specs (ERR_PNPM_SPEC_NOT_SUPPORTED). The pre-`dlx` position parses
|
||||
* into dlx's allow-build option and has been honored since pnpm 10.2.0 (#1939).
|
||||
*
|
||||
* This stays self-contained CJS because the Claude/Antigravity hooks run as
|
||||
* standalone files copied into the user's hook dir, where no package import is
|
||||
* available. The CLI reuses this module from src/cli/resolve-invocation.ts via
|
||||
* createRequire rather than re-implementing it. Two committed copies must stay
|
||||
* byte-identical (enforced by resolve-invocation.test.ts) — edit both together:
|
||||
* gitnexus/hooks/claude/ (the canonical copy the CLI and `gitnexus setup` read)
|
||||
* and gitnexus-claude-plugin/hooks/. A THIRD copy is written at runtime to
|
||||
* `<repo>/.gitnexus/run.cjs` by `gitnexus analyze` (ai-context.ts) so docs can
|
||||
* reference it directly via the `require.main === module` exec tail below; that
|
||||
* copy is gitignored and refreshed on every analyze, so it cannot drift for long.
|
||||
*/
|
||||
|
||||
const { execFileSync } = require('child_process');
|
||||
const fs = require('fs');
|
||||
const path = require('path');
|
||||
|
||||
const NPX_REF = 'gitnexus@latest';
|
||||
|
||||
// Native packages whose postinstall must run under pnpm 10+ (blocked by default).
|
||||
const PNPM_ALLOW_BUILD_BASE = ['@ladybugdb/core', 'gitnexus', 'tree-sitter'];
|
||||
const PNPM_ALLOW_BUILD_EMBEDDINGS = ['onnxruntime-node'];
|
||||
|
||||
// Version-probe timeout, kept under Claude Code's 10s hook budget. PATH presence
|
||||
// detection is now spawn-free (resolveOnPath scans PATH directly), so the only
|
||||
// subprocesses left are the version probes: in a linked worktree the stale-index
|
||||
// hook first runs `git rev-parse --git-common-dir` (~2s) and `git rev-parse HEAD`
|
||||
// (~3s); the pnpm path then adds up to two 1s `--version` probes (npm, pnpm), so
|
||||
// the worst case is ~7s — within budget. A healthy `--version` returns in well
|
||||
// under a second, so the realistic cost is far lower.
|
||||
const PROBE_TIMEOUT_MS = 1000;
|
||||
|
||||
/**
|
||||
* Absolute path to `command` on PATH, or null — a pure-Node, spawn-free lookup
|
||||
* that mirrors how a shell resolves a bare command name: each PATH dir × the
|
||||
* platform's executable extensions (PATHEXT on Windows; the bare name + X_OK on
|
||||
* POSIX). This replaces the former `where`/`which` subprocess (#1938 "Option A"):
|
||||
* it is byte-for-byte identical on every OS, with no dependency on the probe
|
||||
* binary being reachable (a sanitized PATH that drops System32 / `/usr/bin` no
|
||||
* longer defeats detection), no shell-spawn surface (CVE-2024-27980), and no
|
||||
* spawn timeout to tune. On Windows it matches PATHEXT extensions ONLY — exactly
|
||||
* what `where`/cmd.exe resolve — so neither an un-spawnable `.ps1`-only shim (not
|
||||
* in default PATHEXT) nor a bare extensionless file (which the shell cannot launch
|
||||
* as `command`) is a false positive. `preferExecExt` returns a recognized
|
||||
* `.cmd`/`.bat`/`.exe` shim ahead of an exotic PATHEXT hit (e.g. `.COM`) when both
|
||||
* match, matching what a user would actually launch. Pure (platform/env injectable)
|
||||
* so it is unit-testable without touching the host PATH.
|
||||
*/
|
||||
function resolveOnPath(
|
||||
command,
|
||||
preferExecExt = false,
|
||||
{ platform = process.platform, env = process.env } = {},
|
||||
) {
|
||||
const pathValue = env.PATH || env.Path || env.path || '';
|
||||
if (!pathValue) return null;
|
||||
const isWin = platform === 'win32';
|
||||
const exts = isWin
|
||||
? (env.PATHEXT || '.COM;.EXE;.BAT;.CMD')
|
||||
.split(';')
|
||||
.map((e) => e.trim())
|
||||
.filter(Boolean)
|
||||
.map((e) => (e.startsWith('.') ? e : `.${e}`))
|
||||
: [''];
|
||||
let weakHit = null;
|
||||
// Split on the host's PATH delimiter. `platform` is injected only to choose the
|
||||
// extension/exec-bit rules; the PATH string is always host-format, so it must
|
||||
// split on the host delimiter (`path.delimiter`) — in production `platform` IS
|
||||
// the host, so they coincide. (Deriving the delimiter from an injected platform
|
||||
// would split a Windows drive-letter path `C:\…` at its colon under a POSIX
|
||||
// injection.)
|
||||
for (const dir of pathValue.split(path.delimiter).filter(Boolean)) {
|
||||
for (const ext of exts) {
|
||||
const candidate = path.join(dir, `${command}${ext}`);
|
||||
try {
|
||||
if (!fs.statSync(candidate).isFile()) continue;
|
||||
if (!isWin) fs.accessSync(candidate, fs.constants.X_OK);
|
||||
// Prefer a runnable .cmd/.bat/.exe shim; remember an exotic PATHEXT hit
|
||||
// (e.g. .COM) only as a last resort if nothing better turns up.
|
||||
if (isWin && preferExecExt && !/\.(cmd|bat|exe)$/i.test(ext)) {
|
||||
weakHit = weakHit || candidate;
|
||||
continue;
|
||||
}
|
||||
return candidate;
|
||||
} catch {
|
||||
/* not a runnable file here — try the next candidate */
|
||||
}
|
||||
}
|
||||
}
|
||||
return weakHit;
|
||||
}
|
||||
|
||||
// One spawn of `<command> --version` → { major, minor } (each null when
|
||||
// unreadable). Version injection happens at the resolver seam (getNpmMajorVersion
|
||||
// / formatPnpmAllowBuildArgs), so this stays a pure real-process probe.
|
||||
function probeVersion(command) {
|
||||
try {
|
||||
const output = execFileSync(command, ['--version'], {
|
||||
encoding: 'utf-8',
|
||||
timeout: PROBE_TIMEOUT_MS,
|
||||
stdio: ['ignore', 'pipe', 'ignore'],
|
||||
windowsHide: true,
|
||||
// On Windows, npm/pnpm resolve to `.cmd` shims; execFileSync does no
|
||||
// PATHEXT resolution and Node refuses to spawn `.cmd`/`.bat` without a
|
||||
// shell (CVE-2024-27980), so a bare `<command> --version` ENOENTs and the
|
||||
// probe would wrongly report a present tool as absent. A shell lets the OS
|
||||
// resolve the shim. POSIX needs no shell (direct PATH lookup works).
|
||||
shell: process.platform === 'win32',
|
||||
});
|
||||
// Find the first line that starts with a version token (`MAJOR.MINOR`,
|
||||
// optional `v` prefix) rather than splitting the whole output — pnpm/npm
|
||||
// under Corepack or with an update notice can print a banner line on stdout
|
||||
// before the version (stderr is already dropped via the stdio config).
|
||||
const versionLine = output
|
||||
.split('\n')
|
||||
.map((l) => l.trim())
|
||||
.find((l) => /^v?\d+\.\d+/.test(l));
|
||||
const match = versionLine ? versionLine.match(/^v?(\d+)\.(\d+)/) : null;
|
||||
return {
|
||||
major: match ? Number(match[1]) : null,
|
||||
minor: match ? Number(match[2]) : null,
|
||||
};
|
||||
} catch {
|
||||
return { major: null, minor: null };
|
||||
}
|
||||
}
|
||||
|
||||
// `deps` is the single injection seam: an explicitly provided key — including a
|
||||
// `null` value, detected via `in` — is honored as-is so tests can simulate an
|
||||
// absent tool without spawning; an absent key falls through to the real probe.
|
||||
function getNpmMajorVersion(deps = {}) {
|
||||
return 'npmMajor' in deps ? deps.npmMajor : probeVersion('npm').major;
|
||||
}
|
||||
|
||||
/**
|
||||
* `--allow-build` flags for the pre-`dlx` position. Emitted for pnpm >= 10.2
|
||||
* (where the flag exists, and pnpm 10+ blocks build scripts by default). Omitted
|
||||
* below 10.2: pnpm < 10 runs build scripts anyway, and pnpm 10.0/10.1 lack the
|
||||
* flag (it would be rejected as an unknown option). `alwaysAllowBuild` forces the
|
||||
* flags for committed documentation, which cannot probe the reader's pnpm.
|
||||
*/
|
||||
function formatPnpmAllowBuildArgs(options = {}, deps = {}) {
|
||||
if (!options.alwaysAllowBuild) {
|
||||
const { major, minor } =
|
||||
'pnpmMajor' in deps
|
||||
? { major: deps.pnpmMajor, minor: 'pnpmMinor' in deps ? deps.pnpmMinor : null }
|
||||
: probeVersion('pnpm');
|
||||
const lacksAllowBuild =
|
||||
major !== null && (major < 10 || (major === 10 && minor !== null && minor < 2));
|
||||
if (lacksAllowBuild) return [];
|
||||
}
|
||||
const pkgs = [...PNPM_ALLOW_BUILD_BASE];
|
||||
if (options.embeddings) pkgs.push(...PNPM_ALLOW_BUILD_EMBEDDINGS);
|
||||
return pkgs.map((p) => `--allow-build=${p}`);
|
||||
}
|
||||
|
||||
/** Fixed install-free command for committed AGENTS.md / SKILL.md (pnpm >= 10.2). */
|
||||
function formatDocumentationDlxCommand(gitnexusArgs, options = {}) {
|
||||
const flags = formatPnpmAllowBuildArgs({ ...options, alwaysAllowBuild: true }).join(' ');
|
||||
const prefix = flags ? `${flags} ` : '';
|
||||
return `pnpm ${prefix}dlx ${NPX_REF} ${gitnexusArgs}`;
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve `gitnexus` | `pnpm` | `npx`. `GITNEXUS_INVOCATION` forces a mode
|
||||
* (test/escape hatch). `probe` is injectable so the preference order can be
|
||||
* unit-tested without spawning; it defaults to the real PATH probe. `deps` can
|
||||
* inject `{ npmMajor, pnpmMajor }` for tests.
|
||||
*/
|
||||
function resolveInvocationMode(probe = resolveOnPath, deps = {}) {
|
||||
const forced = process.env.GITNEXUS_INVOCATION?.trim().toLowerCase();
|
||||
if (forced === 'gitnexus' || forced === 'pnpm' || forced === 'npx') {
|
||||
return forced;
|
||||
}
|
||||
if (probe('gitnexus', true)) return 'gitnexus';
|
||||
|
||||
const npmMajor = getNpmMajorVersion(deps);
|
||||
// pnpm presence: prefer an explicit `pnpmPresent` flag (set by
|
||||
// formatAnalyzeCommand, which falls back to a PATH probe when the version is
|
||||
// unreadable) so a present-but-unparseable pnpm — slow probe, Corepack
|
||||
// banner — still selects pnpm instead of the npx crash path. Otherwise an
|
||||
// injected version (a successful `pnpm --version` proves presence)
|
||||
// short-circuits the `which pnpm` probe; failing both, fall back to PATH.
|
||||
const hasPnpm =
|
||||
'pnpmPresent' in deps
|
||||
? deps.pnpmPresent
|
||||
: 'pnpmMajor' in deps
|
||||
? deps.pnpmMajor !== null
|
||||
: Boolean(probe('pnpm'));
|
||||
|
||||
// npm 11+ npx install crash (#1939) — prefer pnpm dlx when available.
|
||||
if (hasPnpm && npmMajor !== null && npmMajor >= 11) return 'pnpm';
|
||||
// npm 10 and earlier: npx works; prefer it over pnpm dlx when npm is present.
|
||||
if (npmMajor !== null && npmMajor < 11) return 'npx';
|
||||
// npm absent or unreadable — use pnpm if present (with allow-build flags).
|
||||
if (hasPnpm) return 'pnpm';
|
||||
|
||||
return 'npx';
|
||||
}
|
||||
|
||||
function formatPnpmDlxCommand(gitnexusArgs, options = {}, deps = {}) {
|
||||
const flags = formatPnpmAllowBuildArgs(options, deps).join(' ');
|
||||
const prefix = flags ? `${flags} ` : '';
|
||||
return `pnpm ${prefix}dlx ${NPX_REF} ${gitnexusArgs}`;
|
||||
}
|
||||
|
||||
function formatAnalyzeCommand(options = {}, deps = {}) {
|
||||
const suffix = options.embeddings ? ' --embeddings' : '';
|
||||
// Keep the stale-index hook budget tight by querying each tool at most once.
|
||||
// The memoized `probe` is a spawn-free PATH scan (resolveOnPath) shared with
|
||||
// resolveInvocationMode, so `gitnexus` is scanned only once and no subprocess
|
||||
// is spawned for presence. pnpm's *version* is still captured by a single
|
||||
// `pnpm --version` (the allow-build gate needs the number), which also proves
|
||||
// presence; the memoized scan only re-checks pnpm when that version is
|
||||
// unreadable. Injected deps (tests) and forced/global modes skip the pnpm probe.
|
||||
const cache = new Map();
|
||||
const probe = (command, gitnexusWrapper) => {
|
||||
const key = `${command}:${gitnexusWrapper ? 1 : 0}`;
|
||||
if (!cache.has(key)) cache.set(key, resolveOnPath(command, gitnexusWrapper));
|
||||
return cache.get(key);
|
||||
};
|
||||
let resolved = deps;
|
||||
if (!('pnpmMajor' in deps)) {
|
||||
const forced = process.env.GITNEXUS_INVOCATION?.trim().toLowerCase();
|
||||
// pnpm is only consulted when no non-pnpm mode is already certain: forced
|
||||
// gitnexus/npx never use pnpm, and a present global gitnexus wins outright.
|
||||
const mightUsePnpm = forced === 'pnpm' || (forced !== 'gitnexus' && forced !== 'npx');
|
||||
if (mightUsePnpm && (forced === 'pnpm' || !probe('gitnexus', true))) {
|
||||
const { major, minor } = probeVersion('pnpm');
|
||||
// Carry presence separately from version: when the version probe fails
|
||||
// (timeout, Corepack banner) but pnpm is on PATH, still treat it as
|
||||
// present so mode resolution picks pnpm over the npx crash path. The
|
||||
// PATH probe is memoized and only runs when the version is unreadable.
|
||||
const pnpmPresent = major !== null || Boolean(probe('pnpm'));
|
||||
resolved = { ...deps, pnpmMajor: major, pnpmMinor: minor, pnpmPresent };
|
||||
}
|
||||
}
|
||||
const mode = resolveInvocationMode(probe, resolved);
|
||||
if (mode === 'gitnexus') return `gitnexus analyze${suffix}`;
|
||||
if (mode === 'pnpm') return `${formatPnpmDlxCommand(`analyze${suffix}`, options, resolved)}`;
|
||||
return `npx ${NPX_REF} analyze${suffix}`;
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve `mode` into a concrete { program, args } pair for a set of gitnexus
|
||||
* subcommand arguments. Shared by the direct-exec entrypoint below; pure (no
|
||||
* spawn) so it is unit-testable. `--embeddings` widens the pnpm allow-build set.
|
||||
*/
|
||||
function buildRunnerArgv(mode, gitnexusArgs, deps = {}) {
|
||||
// Match both the space form (`--embeddings`) and the equals form
|
||||
// (`--embeddings=5000`) Commander accepts, so the pnpm allow-build set still
|
||||
// widens to onnxruntime-node when a user hand-types the equals form.
|
||||
const embeddings = gitnexusArgs.some(
|
||||
(a) => a === '--embeddings' || a.startsWith('--embeddings='),
|
||||
);
|
||||
if (mode === 'gitnexus') return { program: 'gitnexus', args: [...gitnexusArgs] };
|
||||
if (mode === 'pnpm') {
|
||||
return {
|
||||
program: 'pnpm',
|
||||
args: [...formatPnpmAllowBuildArgs({ embeddings }, deps), 'dlx', NPX_REF, ...gitnexusArgs],
|
||||
};
|
||||
}
|
||||
return { program: 'npx', args: [NPX_REF, ...gitnexusArgs] };
|
||||
}
|
||||
|
||||
module.exports = {
|
||||
formatAnalyzeCommand,
|
||||
formatDocumentationDlxCommand,
|
||||
formatPnpmAllowBuildArgs,
|
||||
formatPnpmDlxCommand,
|
||||
resolveInvocationMode,
|
||||
buildRunnerArgv,
|
||||
resolveOnPath,
|
||||
getNpmMajorVersion,
|
||||
NPX_REF,
|
||||
PNPM_ALLOW_BUILD_BASE,
|
||||
};
|
||||
|
||||
// Direct-exec entrypoint (#1945): `node run.cjs <gitnexus args…>` resolves the
|
||||
// best available runner (global `gitnexus` → `pnpm dlx` → `npx`) at call time and
|
||||
// runs it, inheriting stdio and propagating the child's exit code. This lets the
|
||||
// committed skills and generated AGENTS.md/CLAUDE.md reference ONE stable,
|
||||
// CLI-neutral command without baking in a package-manager assumption. `gitnexus
|
||||
// analyze` drops a copy of this file at `.gitnexus/run.cjs`. Skipped on require()
|
||||
// (the CLI and tests reuse the exports above), so it runs only when invoked as a
|
||||
// script.
|
||||
if (require.main === module) {
|
||||
const gitnexusArgs = process.argv.slice(2);
|
||||
const { program, args } = buildRunnerArgv(resolveInvocationMode(), gitnexusArgs);
|
||||
try {
|
||||
execFileSync(program, args, {
|
||||
stdio: 'inherit',
|
||||
windowsHide: true,
|
||||
// On Windows, `npx`/`pnpm`/`gitnexus` resolve to `.cmd`/`.ps1`/`.exe`
|
||||
// shims (npm, Volta, Corepack, scoop). execFileSync does not do PATHEXT
|
||||
// resolution and Node refuses to spawn `.cmd`/`.bat` without a shell
|
||||
// (CVE-2024-27980), so a bare program name ENOENTs. A shell lets the OS
|
||||
// resolve the shim; POSIX needs no shell (direct PATH lookup works).
|
||||
shell: process.platform === 'win32',
|
||||
});
|
||||
} catch (err) {
|
||||
// Make spawn failures (resolved program absent from PATH) self-explanatory
|
||||
// instead of a silent exit 1, then propagate the runner's own exit code.
|
||||
if (typeof err.status !== 'number') {
|
||||
process.stderr.write(`gitnexus runner: could not launch \`${program}\` — ${err.message}\n`);
|
||||
}
|
||||
process.exit(typeof err.status === 'number' ? err.status : 1);
|
||||
}
|
||||
}
|
||||
@@ -5,14 +5,16 @@ description: "Use when the user needs to run GitNexus CLI commands like analyze/
|
||||
|
||||
# GitNexus CLI Commands
|
||||
|
||||
All commands work via `npx` — no global install required.
|
||||
Commands below use `node .gitnexus/run.cjs <command>` — the project-local runner `gitnexus analyze` drops next to the index. It auto-selects an available runner at call time (global `gitnexus`, else `pnpm dlx`, else `npx`), so no package-manager assumption and no global install is required.
|
||||
|
||||
> **Not analyzed yet, or `node .gitnexus/run.cjs` reports `Cannot find module`** (the gitignored runner is absent — e.g. a fresh clone or `git clean`)? (Re)generate it with `npx gitnexus analyze` from the project root. On **npm 11.x**, if `npx` crashes during install (`node.target is null`), install once with `npm i -g gitnexus` (then `gitnexus analyze`) or use `pnpm --allow-build=@ladybugdb/core --allow-build=gitnexus --allow-build=tree-sitter dlx gitnexus@latest analyze`. See [#1939](https://github.com/abhigyanpatwari/GitNexus/issues/1939).
|
||||
|
||||
## Commands
|
||||
|
||||
### analyze — Build or refresh the index
|
||||
|
||||
```bash
|
||||
npx gitnexus analyze
|
||||
node .gitnexus/run.cjs analyze
|
||||
```
|
||||
|
||||
Run from the project root. This parses all source files, builds the knowledge graph, writes it to `.gitnexus/`, and generates CLAUDE.md / AGENTS.md context files.
|
||||
@@ -28,7 +30,7 @@ Run from the project root. This parses all source files, builds the knowledge gr
|
||||
### status — Check index freshness
|
||||
|
||||
```bash
|
||||
npx gitnexus status
|
||||
node .gitnexus/run.cjs status
|
||||
```
|
||||
|
||||
Shows whether the current repo has a GitNexus index, when it was last updated, and symbol/relationship counts. Use this to check if re-indexing is needed.
|
||||
@@ -36,7 +38,7 @@ Shows whether the current repo has a GitNexus index, when it was last updated, a
|
||||
### clean — Delete the index
|
||||
|
||||
```bash
|
||||
npx gitnexus clean
|
||||
node .gitnexus/run.cjs clean
|
||||
```
|
||||
|
||||
Deletes the `.gitnexus/` directory and unregisters the repo from the global registry. Use before re-indexing if the index is corrupt or after removing GitNexus from a project.
|
||||
@@ -49,7 +51,7 @@ Deletes the `.gitnexus/` directory and unregisters the repo from the global regi
|
||||
### wiki — Generate documentation from the graph
|
||||
|
||||
```bash
|
||||
npx gitnexus wiki
|
||||
node .gitnexus/run.cjs wiki
|
||||
```
|
||||
|
||||
Generates repository documentation from the knowledge graph using an LLM. Requires an API key (saved to `~/.gitnexus/config.json` on first use).
|
||||
@@ -68,7 +70,7 @@ Generates repository documentation from the knowledge graph using an LLM. Requir
|
||||
### list — Show all indexed repos
|
||||
|
||||
```bash
|
||||
npx gitnexus list
|
||||
node .gitnexus/run.cjs list
|
||||
```
|
||||
|
||||
Lists all repositories registered in `~/.gitnexus/registry.json`. The MCP `list_repos` tool provides the same information.
|
||||
|
||||
@@ -22,7 +22,7 @@ description: "Use when the user is debugging a bug, tracing an error, or asking
|
||||
4. gitnexus_cypher({query: "MATCH path..."}) → Custom traces if needed
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
||||
## Checklist
|
||||
|
||||
|
||||
@@ -23,7 +23,7 @@ description: "Use when the user asks how code works, wants to understand archite
|
||||
5. READ gitnexus://repo/{name}/process/{name} → Trace full execution flow
|
||||
```
|
||||
|
||||
> If step 2 says "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
> If step 2 says "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
||||
## Checklist
|
||||
|
||||
|
||||
@@ -15,7 +15,7 @@ For any task involving code understanding, debugging, impact analysis, or refact
|
||||
2. **Match your task to a skill below** and **read that skill file**
|
||||
3. **Follow the skill's workflow and checklist**
|
||||
|
||||
> If step 1 warns the index is stale, run `npx gitnexus analyze` in the terminal first.
|
||||
> If step 1 warns the index is stale, run `node .gitnexus/run.cjs analyze` in the terminal first.
|
||||
|
||||
## Skills
|
||||
|
||||
|
||||
@@ -23,7 +23,7 @@ description: "Use when the user wants to know what will break if they change som
|
||||
4. Assess risk and report to user
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
||||
## Checklist
|
||||
|
||||
|
||||
@@ -26,7 +26,7 @@ description: "Use when the user wants to review a pull request, understand what
|
||||
6. Summarize findings with risk assessment
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal before reviewing.
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal before reviewing.
|
||||
|
||||
## Checklist
|
||||
|
||||
|
||||
@@ -22,7 +22,7 @@ description: "Use when the user wants to rename, extract, split, move, or restru
|
||||
4. Plan update order: interfaces → implementations → callers → tests
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
||||
## Checklists
|
||||
|
||||
|
||||
@@ -40,7 +40,7 @@ If you already have a `.cursor/hooks.json`, merge the `hooks.postToolUse` array
|
||||
|
||||
### Verify
|
||||
|
||||
1. Index the project: `npx gitnexus analyze`
|
||||
1. Index the project: `npx gitnexus analyze` (on npm 11.x, `npx` can crash during install — use `pnpm --allow-build=@ladybugdb/core --allow-build=gitnexus --allow-build=tree-sitter dlx gitnexus@latest analyze` instead; see [#1939](https://github.com/abhigyanpatwari/GitNexus/issues/1939))
|
||||
2. Reload the Cursor window so it picks up the new hook config.
|
||||
3. Ask the agent something that triggers `Read` / `Grep` / `Shell rg`. You should see a `[GitNexus]` block appended to the tool result.
|
||||
4. Diagnose silent no-ops by setting `GITNEXUS_DEBUG=1` in your shell environment — the hook will write Cursor's raw event payload to stderr so you can verify field names.
|
||||
|
||||
@@ -21,7 +21,7 @@ description: Trace bugs through call chains using knowledge graph
|
||||
4. gitnexus_cypher({query: "MATCH path..."}) → Custom traces if needed
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
||||
## Checklist
|
||||
|
||||
|
||||
@@ -22,7 +22,7 @@ description: Navigate unfamiliar code using GitNexus knowledge graph
|
||||
5. READ gitnexus://repo/{name}/process/{name} → Trace full execution flow
|
||||
```
|
||||
|
||||
> If step 2 says "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
> If step 2 says "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
||||
## Checklist
|
||||
|
||||
|
||||
@@ -22,7 +22,7 @@ description: Analyze blast radius before making code changes
|
||||
4. Assess risk and report to user
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
||||
## Checklist
|
||||
|
||||
|
||||
@@ -26,7 +26,7 @@ description: "Use when the user wants to review a pull request, understand what
|
||||
6. Summarize findings with risk assessment
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal before reviewing.
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal before reviewing.
|
||||
|
||||
## Checklist
|
||||
|
||||
|
||||
@@ -21,7 +21,7 @@ description: Plan safe refactors using blast radius and dependency mapping
|
||||
4. Plan update order: interfaces → implementations → callers → tests
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
||||
## Checklists
|
||||
|
||||
|
||||
+118
-67
@@ -24,27 +24,35 @@ npx gitnexus analyze
|
||||
|
||||
That's it. This indexes the codebase, installs agent skills, registers Claude Code hooks, and creates `AGENTS.md` / `CLAUDE.md` context files — all in one command.
|
||||
|
||||
> **On npm 11.x?** `npx` can crash during install (`Cannot destructure property 'package' of 'node.target'`). Use the pnpm form instead:
|
||||
>
|
||||
> ```bash
|
||||
> pnpm --allow-build=@ladybugdb/core --allow-build=gitnexus --allow-build=tree-sitter dlx gitnexus@latest analyze
|
||||
> ```
|
||||
>
|
||||
> See [Troubleshooting → `npx gitnexus` crashes with `node.target is null` (npm 11)](#cannot-destructure-property-package-of-nodetarget-as-it-is-null) for the full matrix (global install, npm downgrade).
|
||||
|
||||
To configure MCP for your editor, run `npx gitnexus setup` once — or set it up manually below.
|
||||
|
||||
`gitnexus setup` auto-detects your editors and writes the correct global MCP config. You only need to run it once.
|
||||
|
||||
### Editor Support
|
||||
|
||||
| Editor | MCP | Skills | Hooks (auto-augment) | Support |
|
||||
|--------|-----|--------|---------------------|---------|
|
||||
| **Claude Code** | Yes | Yes | Yes (PreToolUse) | **Full** |
|
||||
| **Cursor** | Yes | Yes | Yes (postToolUse, [manual install](../gitnexus-cursor-integration/README.md#hook-install)) | **Full** |
|
||||
| **Antigravity** (Google) | Yes | Yes | Yes (AfterTool, [Gemini CLI hooks schema](https://geminicli.com/docs/hooks/reference/)) | **Full** |
|
||||
| **Codex** | Yes | Yes | — | MCP + Skills |
|
||||
| **Windsurf** | Yes | — | — | MCP |
|
||||
| **OpenCode** | Yes | Yes | — | MCP + Skills |
|
||||
| Editor | MCP | Skills | Hooks (auto-augment) | Support |
|
||||
| ------------------------ | --- | ------ | ------------------------------------------------------------------------------------------ | ------------ |
|
||||
| **Claude Code** | Yes | Yes | Yes (PreToolUse) | **Full** |
|
||||
| **Cursor** | Yes | Yes | Yes (postToolUse, [manual install](../gitnexus-cursor-integration/README.md#hook-install)) | **Full** |
|
||||
| **Antigravity** (Google) | Yes | Yes | Yes (AfterTool, [Gemini CLI hooks schema](https://geminicli.com/docs/hooks/reference/)) | **Full** |
|
||||
| **Codex** | Yes | Yes | — | MCP + Skills |
|
||||
| **Windsurf** | Yes | — | — | MCP |
|
||||
| **OpenCode** | Yes | Yes | — | MCP + Skills |
|
||||
|
||||
> **Claude Code** gets the deepest integration: MCP tools + agent skills + PreToolUse hooks that automatically enrich grep/glob/bash calls with knowledge graph context.
|
||||
|
||||
### Community Integrations
|
||||
|
||||
| Agent | Install | Source |
|
||||
|-------|---------|--------|
|
||||
| Agent | Install | Source |
|
||||
| -------------------- | ---------------------------- | ------------------------------------------------------- |
|
||||
| [pi](https://pi.dev) | `pi install npm:pi-gitnexus` | [pi-gitnexus](https://github.com/tintinweb/pi-gitnexus) |
|
||||
|
||||
## MCP Setup (manual)
|
||||
@@ -116,36 +124,36 @@ The result is a **LadybugDB graph database** stored locally in `.gitnexus/` with
|
||||
|
||||
Your AI agent gets these tools automatically:
|
||||
|
||||
| Tool | What It Does | `repo` Param |
|
||||
|------|-------------|--------------|
|
||||
| `list_repos` | Discover all indexed repositories | — |
|
||||
| `query` | Process-grouped hybrid search (BM25 + semantic + RRF) | Optional |
|
||||
| `context` | 360-degree symbol view — categorized refs, process participation | Optional |
|
||||
| `impact` | Blast radius analysis with depth grouping and confidence | Optional |
|
||||
| `detect_changes` | Git-diff impact — maps changed lines to affected processes | Optional |
|
||||
| `rename` | Multi-file coordinated rename with graph + text search | Optional |
|
||||
| `cypher` | Raw Cypher graph queries | Optional |
|
||||
| Tool | What It Does | `repo` Param |
|
||||
| ---------------- | ---------------------------------------------------------------- | ------------ |
|
||||
| `list_repos` | Discover all indexed repositories | — |
|
||||
| `query` | Process-grouped hybrid search (BM25 + semantic + RRF) | Optional |
|
||||
| `context` | 360-degree symbol view — categorized refs, process participation | Optional |
|
||||
| `impact` | Blast radius analysis with depth grouping and confidence | Optional |
|
||||
| `detect_changes` | Git-diff impact — maps changed lines to affected processes | Optional |
|
||||
| `rename` | Multi-file coordinated rename with graph + text search | Optional |
|
||||
| `cypher` | Raw Cypher graph queries | Optional |
|
||||
|
||||
> With one indexed repo, the `repo` param is optional. With multiple, specify which: `query({query: "auth", repo: "my-app"})`.
|
||||
|
||||
## MCP Resources
|
||||
|
||||
| Resource | Purpose |
|
||||
|----------|---------|
|
||||
| `gitnexus://repos` | List all indexed repositories (read first) |
|
||||
| `gitnexus://repo/{name}/context` | Codebase stats, staleness check, and available tools |
|
||||
| `gitnexus://repo/{name}/clusters` | All functional clusters with cohesion scores |
|
||||
| `gitnexus://repo/{name}/cluster/{name}` | Cluster members and details |
|
||||
| `gitnexus://repo/{name}/processes` | All execution flows |
|
||||
| `gitnexus://repo/{name}/process/{name}` | Full process trace with steps |
|
||||
| `gitnexus://repo/{name}/schema` | Graph schema for Cypher queries |
|
||||
| Resource | Purpose |
|
||||
| --------------------------------------- | ---------------------------------------------------- |
|
||||
| `gitnexus://repos` | List all indexed repositories (read first) |
|
||||
| `gitnexus://repo/{name}/context` | Codebase stats, staleness check, and available tools |
|
||||
| `gitnexus://repo/{name}/clusters` | All functional clusters with cohesion scores |
|
||||
| `gitnexus://repo/{name}/cluster/{name}` | Cluster members and details |
|
||||
| `gitnexus://repo/{name}/processes` | All execution flows |
|
||||
| `gitnexus://repo/{name}/process/{name}` | Full process trace with steps |
|
||||
| `gitnexus://repo/{name}/schema` | Graph schema for Cypher queries |
|
||||
|
||||
## MCP Prompts
|
||||
|
||||
| Prompt | What It Does |
|
||||
|--------|-------------|
|
||||
| `detect_impact` | Pre-commit change analysis — scope, affected processes, risk level |
|
||||
| `generate_map` | Architecture documentation from the knowledge graph with mermaid diagrams |
|
||||
| Prompt | What It Does |
|
||||
| --------------- | ------------------------------------------------------------------------- |
|
||||
| `detect_impact` | Pre-commit change analysis — scope, affected processes, risk level |
|
||||
| `generate_map` | Architecture documentation from the knowledge graph with mermaid diagrams |
|
||||
|
||||
## CLI Commands
|
||||
|
||||
@@ -170,6 +178,13 @@ gitnexus clean --all --force # Delete all indexes
|
||||
gitnexus wiki [path] # Generate LLM-powered docs from knowledge graph
|
||||
gitnexus wiki --model <model> # Wiki with custom LLM model (default: gpt-4o-mini)
|
||||
|
||||
# Direct graph queries — the same tools the MCP server exposes, no MCP daemon needed
|
||||
gitnexus query "<concept>" # Process-grouped hybrid search
|
||||
gitnexus context <symbol> [--uid <uid> | --file <path>] # 360° symbol view; flags disambiguate a shared name
|
||||
gitnexus impact <symbol> [--uid <uid> | --file <path> | --kind <kind>] # Blast radius; flags disambiguate a shared name
|
||||
gitnexus detect-changes # Map the working-tree diff to affected symbols and execution flows
|
||||
gitnexus cypher "<query>" # Run a raw Cypher query against the knowledge graph
|
||||
|
||||
# Repository groups (multi-repo / monorepo service tracking)
|
||||
gitnexus group create <name> # Create a repository group
|
||||
gitnexus group add <group> <groupPath> <registryName> # Add a repo to a group. <groupPath> is a hierarchy path (e.g. hr/hiring/backend); <registryName> is the repo's name from the registry (see `gitnexus list`)
|
||||
@@ -205,21 +220,21 @@ TypeScript, JavaScript, Python, Java, C, C++, C#, Go, Rust, PHP, Kotlin, Swift,
|
||||
|
||||
### Language Feature Matrix
|
||||
|
||||
| Language | Imports | Named Bindings | Exports | Heritage | Type Annotations | Constructor Inference | Config | Frameworks | Entry Points |
|
||||
|----------|---------|----------------|---------|----------|-----------------|---------------------|--------|------------|-------------|
|
||||
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| JavaScript | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
|
||||
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| Java | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| Kotlin | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| Go | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| Rust | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| PHP | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| Ruby | ✓ | — | ✓ | ✓ | — | ✓ | — | ✓ | ✓ |
|
||||
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| Language | Imports | Named Bindings | Exports | Heritage | Type Annotations | Constructor Inference | Config | Frameworks | Entry Points |
|
||||
| ---------- | ------- | -------------- | ------- | -------- | ---------------- | --------------------- | ------ | ---------- | ------------ |
|
||||
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| JavaScript | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
|
||||
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| Java | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| Kotlin | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| Go | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| Rust | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| PHP | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| Ruby | ✓ | — | ✓ | ✓ | — | ✓ | — | ✓ | ✓ |
|
||||
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
|
||||
**Imports** — cross-file import resolution · **Named Bindings** — `import { X as Y }` / re-export tracking · **Exports** — public/exported symbol detection · **Heritage** — class inheritance, interfaces, mixins · **Type Annotations** — explicit type extraction for receiver resolution · **Constructor Inference** — infer receiver type from constructor calls (`self`/`this` resolution included for all languages) · **Config** — language toolchain config parsing (tsconfig, go.mod, etc.) · **Frameworks** — AST-based framework pattern detection · **Entry Points** — entry point scoring heuristics
|
||||
|
||||
@@ -266,22 +281,58 @@ for the full list; stable `latest` is unaffected.
|
||||
|
||||
### `Cannot destructure property 'package' of 'node.target' as it is null`
|
||||
|
||||
This crash was caused by a dependency URL format that is incompatible with
|
||||
certain npm/arborist versions ([npm/cli#8126](https://github.com/npm/cli/issues/8126)).
|
||||
It is fixed in **gitnexus v1.6.2+**. Upgrade to the latest version:
|
||||
This error comes from **npm 11.x's arborist** while installing gitnexus (often via `npx`), before gitnexus code runs. It is triggered by platform-filtered `optionalDependencies` in native packages such as `onnxruntime-node` / `@huggingface/transformers` (used when indexing with `--embeddings`). GitNexus cannot catch it at runtime — use one of these workarounds:
|
||||
|
||||
```bash
|
||||
npx gitnexus@latest analyze # always uses the newest release
|
||||
# — or —
|
||||
npm install -g gitnexus@latest # upgrade a global install
|
||||
pnpm --allow-build=@ladybugdb/core --allow-build=gitnexus --allow-build=tree-sitter dlx gitnexus@latest analyze # auto-selected when pnpm + npm 11+
|
||||
npm install -g gitnexus@latest # global install avoids per-run npx reify
|
||||
gitnexus analyze # if already installed globally
|
||||
```
|
||||
|
||||
If you still hit npm install issues after upgrading, these generic workarounds
|
||||
may help:
|
||||
On **pnpm 10+**, lifecycle scripts are blocked unless explicitly allowed — the resolver adds `--allow-build` for `@ladybugdb/core`, `gitnexus`, and `tree-sitter` automatically when it picks `pnpm dlx`.
|
||||
|
||||
If you must stay on npm 11.x without pnpm, downgrade npm toolchain-wide (last resort):
|
||||
|
||||
```bash
|
||||
npm install -g npm@latest # update npm itself
|
||||
npm cache clean --force # clear a possibly corrupt cache
|
||||
npm install -g npm@10.9.0
|
||||
```
|
||||
|
||||
See [#1939](https://github.com/abhigyanpatwari/GitNexus/issues/1939) and the original [#819](https://github.com/abhigyanpatwari/GitNexus/issues/819) thread. An older variant of this crash (tree-sitter-dart tarball URL) was fixed in gitnexus v1.6.2+ ([#820](https://github.com/abhigyanpatwari/GitNexus/pull/820)); if you still see install failures after upgrading, clear cache:
|
||||
|
||||
```bash
|
||||
npm cache clean --force
|
||||
npx gitnexus@latest analyze
|
||||
```
|
||||
|
||||
### `ERR_DLOPEN_FAILED` / `lbugjs.node` missing (pnpm dlx, pnpx)
|
||||
|
||||
GitNexus depends on `@ladybugdb/core`, whose native database addon
|
||||
(`lbugjs.node`) is placed by a postinstall script. `pnpm dlx`, `pnpx`, and any
|
||||
install run with `--ignore-scripts` skip lifecycle scripts, so the addon is
|
||||
never put in place and the runtime crashes with `ERR_DLOPEN_FAILED`:
|
||||
|
||||
```
|
||||
Error: dlopen(.../@ladybugdb/core/lbugjs.node, ...): tried: '...' (no such file)
|
||||
code: 'ERR_DLOPEN_FAILED'
|
||||
```
|
||||
|
||||
Options that run install scripts:
|
||||
|
||||
```bash
|
||||
# pnpm dlx with explicit build permission (one-off, no global install required)
|
||||
pnpm --allow-build=@ladybugdb/core --allow-build=gitnexus --allow-build=tree-sitter \
|
||||
dlx gitnexus@latest serve
|
||||
|
||||
# npm: global install (recommended on npm 11+; bare npx may crash — see section above)
|
||||
npm install -g gitnexus@latest
|
||||
gitnexus serve
|
||||
|
||||
# npx (npm < 11, or after upgrading npm)
|
||||
npx gitnexus@latest serve
|
||||
|
||||
# pnpm: global install with build scripts allowed (pnpm 10.2+; no approve-builds -g on pnpm 11+)
|
||||
pnpm add -g --allow-build=@ladybugdb/core --allow-build=gitnexus --allow-build=tree-sitter gitnexus
|
||||
gitnexus serve
|
||||
```
|
||||
|
||||
### Installation fails with native module errors
|
||||
@@ -305,11 +356,11 @@ GitNexus uses optional DuckDB extensions for BM25 and vector search. The `gitnex
|
||||
|
||||
Configure the behavior with two environment variables:
|
||||
|
||||
| Variable | Values | Default | Effect |
|
||||
|----------|--------|---------|--------|
|
||||
| `GITNEXUS_LBUG_EXTENSION_INSTALL` | `auto`, `load-only`, `never` | `auto` | `auto` runs one bounded INSTALL if LOAD fails. `load-only` only uses already-installed extensions (recommended for offline / firewalled environments). `never` skips optional extensions entirely. |
|
||||
| `GITNEXUS_LBUG_EXTENSION_INSTALL_TIMEOUT_MS` | positive integer | `15000` | Wall-clock budget for the out-of-process `INSTALL` child before it is killed. |
|
||||
| `GITNEXUS_WAL_CHECKPOINT_THRESHOLD` | integer `>= -1` | `67108864` (64 MiB) | LadybugDB WAL auto-checkpoint threshold during analyze (bytes). Auto-checkpoint remains enabled; `-1` keeps Ladybug's stock ~16 MiB. Larger thresholds reduce checkpoint frequency but increase the WAL size at rotation time — choose a smaller value on disk-constrained environments. |
|
||||
| Variable | Values | Default | Effect |
|
||||
| -------------------------------------------- | ---------------------------- | ------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GITNEXUS_LBUG_EXTENSION_INSTALL` | `auto`, `load-only`, `never` | `auto` | `auto` runs one bounded INSTALL if LOAD fails. `load-only` only uses already-installed extensions (recommended for offline / firewalled environments). `never` skips optional extensions entirely. |
|
||||
| `GITNEXUS_LBUG_EXTENSION_INSTALL_TIMEOUT_MS` | positive integer | `15000` | Wall-clock budget for the out-of-process `INSTALL` child before it is killed. |
|
||||
| `GITNEXUS_WAL_CHECKPOINT_THRESHOLD` | integer `>= -1` | `67108864` (64 MiB) | LadybugDB WAL auto-checkpoint threshold during analyze (bytes). Auto-checkpoint remains enabled; `-1` keeps Ladybug's stock ~16 MiB. Larger thresholds reduce checkpoint frequency but increase the WAL size at rotation time — choose a smaller value on disk-constrained environments. |
|
||||
|
||||
```bash
|
||||
# Offline/airgapped: never reach the network for extensions
|
||||
@@ -366,11 +417,11 @@ For repositories with very large source files, `GITNEXUS_WORKER_SUB_BATCH_MAX_BY
|
||||
|
||||
Three env vars expose the pool's resilience layers (respawn budget, cumulative-timeout cap, circuit breaker). Defaults are tuned for typical repos; bump them when an analyze legitimately needs more retries, or lower them to fail-fast on a known-bad shape.
|
||||
|
||||
| Variable | Default | Effect |
|
||||
| ------------------------------------------------- | ------------------------- | --------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT` | `3` | Max replacement spawns per slot before the slot is dropped from the active rotation. |
|
||||
| `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` | `5 × subBatchTimeoutMs` | Total retry wall-time budget per job before quarantining. Bounds exponentially-growing retry waits. |
|
||||
| `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD` | `max(3, poolSize)` | Per-slot consecutive deaths before the pool's circuit breaker trips. After tripping, dispatches require a fresh pool. |
|
||||
| Variable | Default | Effect |
|
||||
| ----------------------------------------------- | ----------------------- | --------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT` | `3` | Max replacement spawns per slot before the slot is dropped from the active rotation. |
|
||||
| `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` | `5 × subBatchTimeoutMs` | Total retry wall-time budget per job before quarantining. Bounds exponentially-growing retry waits. |
|
||||
| `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD` | `max(3, poolSize)` | Per-slot consecutive deaths before the pool's circuit breaker trips. After tripping, dispatches require a fresh pool. |
|
||||
|
||||
## Privacy
|
||||
|
||||
|
||||
@@ -0,0 +1 @@
|
||||
06687dff942d531c4d453b5906a8666c90db4867eb43ed18304aa59a8a93ef9d
|
||||
@@ -0,0 +1 @@
|
||||
d51ea9edd1902fc20dd888f3fc51907c8119a4692b92c32a427801ac7f41d451
|
||||
@@ -0,0 +1,202 @@
|
||||
/**
|
||||
* Resolver-output correctness fingerprint for `resolvePythonImportTarget`
|
||||
* (ce-optimize: python-scope-capture, hypothesis H2).
|
||||
*
|
||||
* H2 replaces the per-import O(files) suffix scan + candidate scan in
|
||||
* import-target.ts with a memoized index. The index MUST reproduce the exact
|
||||
* resolution result (including the deterministic tie-break and the
|
||||
* false-positive gating) for every input. This harness pins that: it runs
|
||||
* resolvePythonImportTarget over an exhaustive branch matrix PLUS a large
|
||||
* deterministic fuzz (varied repo layouts that force collisions / multi-match
|
||||
* tie-breaks), and prints an order-independent sha256 over every
|
||||
* `fromFile | targetRaw | result` triple.
|
||||
*
|
||||
* Build-free via tsx (static .ts import). Run:
|
||||
* node --import tsx bench/python-scope/import-target-fingerprint.mjs
|
||||
*
|
||||
* The fingerprint before and after the H2 change MUST be identical.
|
||||
*/
|
||||
import crypto from 'node:crypto';
|
||||
import fs from 'node:fs';
|
||||
import path from 'node:path';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
import { resolvePythonImportTarget } from '../../src/core/ingestion/languages/python/import-target.ts';
|
||||
|
||||
function mkImport(targetRaw) {
|
||||
return { kind: 'absolute', targetRaw, isRelative: false, names: [] };
|
||||
}
|
||||
|
||||
function resolve(fromFile, files, targetRaw) {
|
||||
const ctx = { fromFile, allFilePaths: new Set(files) };
|
||||
return resolvePythonImportTarget(mkImport(targetRaw), ctx);
|
||||
}
|
||||
|
||||
const lines = [];
|
||||
let nonNull = 0;
|
||||
function record(fromFile, files, targetRaw) {
|
||||
const r = resolve(fromFile, files, targetRaw);
|
||||
if (r !== null) nonNull++;
|
||||
lines.push(`${fromFile}\t${targetRaw}\t${r === null ? 'NULL' : r}`);
|
||||
}
|
||||
|
||||
// ---- 1. Exhaustive branch matrix ----------------------------------------
|
||||
|
||||
// direct root hit
|
||||
record('app/main.py', ['services/sync.py', 'services/__init__.py'], 'services.sync');
|
||||
// direct package (__init__) hit
|
||||
record('app/main.py', ['services/__init__.py'], 'services');
|
||||
// ancestor walk hit
|
||||
record('backend/routers/cron.py', ['backend/services/sync.py'], 'services.sync');
|
||||
// ancestor pkg hit
|
||||
record('backend/routers/cron.py', ['backend/services/__init__.py'], 'services');
|
||||
// suffix fallback single match (nested vendor layout)
|
||||
record('app/main.py', ['pkg/__init__.py', 'vendor/pkg/thing.py'], 'pkg.thing');
|
||||
// suffix fallback to __init__ (package)
|
||||
record('app/main.py', ['pkg/__init__.py', 'x/pkg/subpkg/__init__.py'], 'pkg.subpkg');
|
||||
// suffix multi-match tie-break: fewest segments wins
|
||||
record('app/main.py', ['pkg/__init__.py', 'a/pkg/models.py', 'b/c/pkg/models.py'], 'pkg.models');
|
||||
// suffix multi-match tie-break at SAME depth: lexicographic
|
||||
record('app/main.py', ['pkg/__init__.py', 'z/pkg/models.py', 'a/pkg/models.py'], 'pkg.models');
|
||||
// suffix file vs pkg same name, mixed — both candidate forms present
|
||||
record(
|
||||
'app/main.py',
|
||||
['pkg/__init__.py', 'q/pkg/models.py', 'r/pkg/models/__init__.py'],
|
||||
'pkg.models',
|
||||
);
|
||||
// hasRepoCandidate FALSE — external dotted import w/ colliding local basename (django.apps guard)
|
||||
record('app/main.py', ['accounts/apps.py'], 'django.apps');
|
||||
// hasRepoCandidate TRUE via top-level package, but no concrete file -> null
|
||||
record('app/main.py', ['pkg/__init__.py'], 'pkg.ghost');
|
||||
// hasRepoCandidate via nested ancestor namespace package
|
||||
record('backend/routers/cron.py', ['backend/services/sync.py'], 'services.helpers.util');
|
||||
// collision: accounts.models must NOT match billing/models.py
|
||||
record('app/main.py', ['accounts/__init__.py', 'billing/models.py'], 'accounts.models');
|
||||
// relative imports
|
||||
record('app/main.py', ['app/sibling.py'], '.sibling');
|
||||
record('app/pkg/mod.py', ['app/sibling.py'], '..sibling');
|
||||
record('app/main.py', ['app/sibling.py'], '...way.too.far');
|
||||
// single-segment bare import (no '/'): skips candidate gate
|
||||
record('app/main.py', ['mod.py'], 'mod');
|
||||
record('app/main.py', ['lib/mod.py'], 'mod');
|
||||
record('app/pkg/main.py', ['app/pkg/local.py'], 'local');
|
||||
// empty / dynamic
|
||||
record('app/main.py', ['a.py'], '');
|
||||
// windows-style backslash paths in the set
|
||||
record('app\\main.py', ['svc\\sync.py', 'svc\\__init__.py'], 'svc.sync');
|
||||
|
||||
// ---- 2. Deterministic fuzz ----------------------------------------------
|
||||
// LCG (no Math.random — deterministic + reproducible).
|
||||
let seed = 0x9e3779b9;
|
||||
function rnd() {
|
||||
seed = (seed * 1664525 + 1013904223) >>> 0;
|
||||
return seed / 0x100000000;
|
||||
}
|
||||
function pick(arr) {
|
||||
return arr[Math.floor(rnd() * arr.length)];
|
||||
}
|
||||
|
||||
const DIRS = ['', 'a', 'b', 'a/b', 'b/c', 'x/y/z', 'vendor', 'src', 'src/app', 'pkg'];
|
||||
const SEGS = [
|
||||
'pkg',
|
||||
'services',
|
||||
'models',
|
||||
'sync',
|
||||
'util',
|
||||
'core',
|
||||
'apps',
|
||||
'sub',
|
||||
'thing',
|
||||
'helpers',
|
||||
];
|
||||
|
||||
function randPath() {
|
||||
const dir = pick(DIRS);
|
||||
const base = pick(SEGS);
|
||||
const isPkg = rnd() < 0.3;
|
||||
const file = isPkg ? `${base}/__init__.py` : `${base}.py`;
|
||||
return dir ? `${dir}/${file}` : file;
|
||||
}
|
||||
function randDotted() {
|
||||
const n = 1 + Math.floor(rnd() * 3);
|
||||
const parts = [];
|
||||
for (let i = 0; i < n; i++) parts.push(pick(SEGS));
|
||||
const rel = rnd() < 0.15 ? '.'.repeat(1 + Math.floor(rnd() * 2)) : '';
|
||||
return rel + parts.join('.');
|
||||
}
|
||||
|
||||
for (let repo = 0; repo < 400; repo++) {
|
||||
const fileCount = 3 + Math.floor(rnd() * 14);
|
||||
const files = [];
|
||||
for (let i = 0; i < fileCount; i++) files.push(randPath());
|
||||
const fromFile = randPath();
|
||||
for (let imp = 0; imp < 25; imp++) {
|
||||
record(fromFile, files, randDotted());
|
||||
}
|
||||
}
|
||||
|
||||
// ---- 3. Absolute-path coverage (PR #1918 review P3a) --------------------
|
||||
// Production paths are repo-relative, but the index's prefix gating must
|
||||
// reproduce the old `f.startsWith(prefix)` semantics for absolute paths too.
|
||||
// The reviewer's exact case + a fuzz over leading-`/` file sets and absolute
|
||||
// importer paths lock the absolute-path behavior end to end.
|
||||
|
||||
// The flagged case: an absolute file under the importer's own root.
|
||||
record('/repo/app/main.py', ['/repo/svc/x.py'], 'svc.x');
|
||||
record('/repo/app/main.py', ['/repo/svc/__init__.py', '/repo/svc/x.py'], 'svc.x');
|
||||
// Absolute file NOT under the importer root — gate must not pass it.
|
||||
record('/repo/app/main.py', ['/other/svc/x.py'], 'svc.x');
|
||||
// Absolute vendored layout reachable only by suffix.
|
||||
record('/repo/app/main.py', ['/repo/pkg/__init__.py', '/repo/vendor/pkg/thing.py'], 'pkg.thing');
|
||||
// Absolute tie-break.
|
||||
record(
|
||||
'/repo/app/main.py',
|
||||
['/repo/pkg/__init__.py', '/a/pkg/models.py', '/b/c/pkg/models.py'],
|
||||
'pkg.models',
|
||||
);
|
||||
// Mixed absolute/relative file set.
|
||||
record('/repo/app/main.py', ['/repo/pkg/__init__.py', 'pkg/models.py'], 'pkg.models');
|
||||
|
||||
function randAbsPath() {
|
||||
// Reuse the relative generator under one of a few absolute roots.
|
||||
const root = pick(['/repo', '/srv/app', '/']);
|
||||
const rel = randPath();
|
||||
return root === '/' ? `/${rel}` : `${root}/${rel}`;
|
||||
}
|
||||
|
||||
for (let repo = 0; repo < 200; repo++) {
|
||||
const fileCount = 3 + Math.floor(rnd() * 12);
|
||||
const files = [];
|
||||
for (let i = 0; i < fileCount; i++) files.push(randAbsPath());
|
||||
const fromFile = randAbsPath();
|
||||
for (let imp = 0; imp < 20; imp++) {
|
||||
record(fromFile, files, randDotted());
|
||||
}
|
||||
}
|
||||
|
||||
const fingerprint = crypto
|
||||
.createHash('sha256')
|
||||
.update([...lines].sort().join('\n'))
|
||||
.digest('hex');
|
||||
const result = { fingerprint, cases: lines.length, non_null: nonNull };
|
||||
|
||||
if (!process.argv.includes('--check')) {
|
||||
process.stdout.write(JSON.stringify(result) + '\n');
|
||||
} else {
|
||||
// CI gate: resolver output unchanged (fingerprint == committed baseline).
|
||||
// Re-baseline a legitimate resolution change by running without --check and
|
||||
// committing the new baseline-import-target-fingerprint.txt deliberately.
|
||||
const __dirname = path.dirname(fileURLToPath(import.meta.url));
|
||||
const baseline = fs
|
||||
.readFileSync(path.resolve(__dirname, 'baseline-import-target-fingerprint.txt'), 'utf8')
|
||||
.trim();
|
||||
process.stdout.write(JSON.stringify(result) + '\n');
|
||||
if (result.fingerprint !== baseline) {
|
||||
process.stderr.write(
|
||||
`[import-target-fingerprint --check] FAIL: resolver fingerprint drift: got ` +
|
||||
`${result.fingerprint}, expected ${baseline} (resolvePythonImportTarget output changed — ` +
|
||||
`re-baseline intentionally if expected)\n`,
|
||||
);
|
||||
process.exit(1);
|
||||
}
|
||||
process.stderr.write('[import-target-fingerprint --check] PASS (resolver fingerprint)\n');
|
||||
}
|
||||
@@ -0,0 +1,220 @@
|
||||
/**
|
||||
* Build-free measurement harness for `emitPythonScopeCaptures`
|
||||
* (ce-optimize: python-scope-capture).
|
||||
*
|
||||
* This is Python's counterpart to `bench/scope-capture/measure.mjs` (which
|
||||
* covers go/csharp/rust/php/ruby/cobol). Python lives here, NOT in that unified
|
||||
* harness, because this one ALSO covers import resolution
|
||||
* (`import-target-fingerprint.mjs`). Python's capture-scaling guard therefore
|
||||
* runs via `python-scope/measure.mjs --check`, not the unified harness — don't
|
||||
* remove either thinking the other covers Python.
|
||||
*
|
||||
* Mirrors the Go scope-capture harness (#1848). Imports the `.ts` hotpath
|
||||
* directly through tsx (`node --import tsx bench/python-scope/measure.mjs`):
|
||||
* a static `.ts` import works; a top-level `await import()` breaks tsx's lexer.
|
||||
*
|
||||
* Emits ONE JSON object on stdout with:
|
||||
* - elapsed_ms_250 / elapsed_ms_800: median wall-clock (ms) of
|
||||
* emitPythonScopeCaptures over a synthetic DAO-style source at that many
|
||||
* top-level entities (warmed up first). 800/250 ~ 3.2x input; an O(n^2)
|
||||
* path scales ~quadratically, an O(n) path ~linearly.
|
||||
* - scaling_ratio: (t800/t250)/(800/250). ~3.2 = quadratic, ~1.0 = linear.
|
||||
* - capture_groups_250 / capture_groups_800: match counts (a fast-but-empty
|
||||
* regression can't pass — counts must stay > 0).
|
||||
* - fingerprint: order-independent sha256 over emitPythonScopeCaptures output
|
||||
* across the whole lang-resolution/python-* fixture corpus + a fixed
|
||||
* 20-entity synthetic DAO. This is the CORRECTNESS gate: any change to the
|
||||
* captures changes the fingerprint. Entity-count-fixed so it is comparable
|
||||
* across experiments regardless of the timing sizes.
|
||||
* - capture_groups_fp / fixture_count: corpus sanity.
|
||||
*/
|
||||
import fs from 'node:fs';
|
||||
import path from 'node:path';
|
||||
import crypto from 'node:crypto';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
import { emitPythonScopeCaptures } from '../../src/core/ingestion/languages/python/captures.ts';
|
||||
|
||||
const __dirname = path.dirname(fileURLToPath(import.meta.url));
|
||||
const FIXTURE_ROOT = path.resolve(__dirname, '..', '..', 'test', 'fixtures', 'lang-resolution');
|
||||
|
||||
// ---- correctness fingerprint (order-independent, mirrors the Go golden) ----
|
||||
|
||||
function canonicalizeMatch(match) {
|
||||
const parts = [];
|
||||
for (const tag of Object.keys(match)) {
|
||||
const cap = match[tag];
|
||||
const r = cap.range;
|
||||
parts.push(`${tag}|${cap.text}|${r.startLine}:${r.startCol}-${r.endLine}:${r.endCol}`);
|
||||
}
|
||||
parts.sort();
|
||||
return parts.join(';');
|
||||
}
|
||||
|
||||
function digestCaptures(matches) {
|
||||
const matchStrings = matches.map(canonicalizeMatch).sort();
|
||||
return crypto.createHash('sha256').update(matchStrings.join('\n')).digest('hex');
|
||||
}
|
||||
|
||||
/** All `.py` files under `lang-resolution/python-*`, sorted by repo-relative key. */
|
||||
function collectPythonFixtures() {
|
||||
const out = [];
|
||||
for (const entry of fs.readdirSync(FIXTURE_ROOT, { withFileTypes: true })) {
|
||||
if (!entry.isDirectory() || !entry.name.startsWith('python-')) continue;
|
||||
const stack = [path.join(FIXTURE_ROOT, entry.name)];
|
||||
while (stack.length) {
|
||||
const dir = stack.pop();
|
||||
for (const c of fs.readdirSync(dir, { withFileTypes: true })) {
|
||||
const p = path.join(dir, c.name);
|
||||
if (c.isDirectory()) stack.push(p);
|
||||
else if (c.name.endsWith('.py')) {
|
||||
out.push({ key: path.relative(FIXTURE_ROOT, p).split(path.sep).join('/'), absPath: p });
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
out.sort((a, b) => a.key.localeCompare(b.key));
|
||||
return out;
|
||||
}
|
||||
|
||||
/**
|
||||
* Synthetic DAO-style source: top-level imports + N classes (each with methods,
|
||||
* exercising @scope.function + @declaration.function + receiver binding) + N
|
||||
* module functions. Maximizes top-level children (rootChildren) AND function
|
||||
* matches, which is exactly the O(matches x rootChildren) shape #1848 hit.
|
||||
*/
|
||||
function generatePyDao(entityCount) {
|
||||
const lines = [];
|
||||
for (let i = 0; i < 12; i++) {
|
||||
lines.push(`from pkg.mod${i} import alpha${i}, beta${i}, gamma${i} as g${i}`);
|
||||
lines.push(`import top.level.module${i}`);
|
||||
}
|
||||
lines.push('');
|
||||
// Shared base + mixin so every Entity is heritage-bearing — exercises the
|
||||
// @reference.inherits synth (#1951) at scale (single + multiple inheritance),
|
||||
// not just the base capture loop.
|
||||
lines.push('class Base:', ' pass', '', 'class Mixin:', ' pass', '');
|
||||
for (let i = 0; i < entityCount; i++) {
|
||||
const n = String(i).padStart(4, '0');
|
||||
lines.push(
|
||||
`class Entity${n}(Base, Mixin):`,
|
||||
` def __init__(self, id: int, name: str):`,
|
||||
` self.id = id`,
|
||||
` self.name = name`,
|
||||
` def get_id(self) -> int:`,
|
||||
` return self.id`,
|
||||
` def set_name(self, name: str) -> None:`,
|
||||
` self.name = name`,
|
||||
` @classmethod`,
|
||||
` def make(cls, id: int):`,
|
||||
` return cls(id, "x")`,
|
||||
'',
|
||||
`def build_entity${n}(id: int, name: str) -> Entity${n}:`,
|
||||
` return Entity${n}(id, name)`,
|
||||
'',
|
||||
);
|
||||
}
|
||||
return lines.join('\n');
|
||||
}
|
||||
|
||||
// ---- timing ----
|
||||
|
||||
function timeOnce(src, filePath) {
|
||||
const start = process.hrtime.bigint();
|
||||
const matches = emitPythonScopeCaptures(src, filePath);
|
||||
const end = process.hrtime.bigint();
|
||||
return { ms: Number(end - start) / 1e6, count: matches.length };
|
||||
}
|
||||
|
||||
function median(xs) {
|
||||
const s = [...xs].sort((a, b) => a - b);
|
||||
const m = Math.floor(s.length / 2);
|
||||
return s.length % 2 ? s[m] : (s[m - 1] + s[m]) / 2;
|
||||
}
|
||||
|
||||
function measureSize(entityCount, reps) {
|
||||
const src = generatePyDao(entityCount);
|
||||
// Warm up parser/query JIT (not counted).
|
||||
timeOnce(src, 'warmup.py');
|
||||
const samples = [];
|
||||
let count = 0;
|
||||
for (let i = 0; i < reps; i++) {
|
||||
const r = timeOnce(src, `bench-${entityCount}.py`);
|
||||
samples.push(r.ms);
|
||||
count = r.count;
|
||||
}
|
||||
return { ms: median(samples), count };
|
||||
}
|
||||
|
||||
// ---- run ----
|
||||
|
||||
function computeFingerprint() {
|
||||
let groups = 0;
|
||||
const perFixtureDigests = [];
|
||||
for (const { key, absPath } of collectPythonFixtures()) {
|
||||
const src = fs.readFileSync(absPath, 'utf8');
|
||||
const matches = emitPythonScopeCaptures(src, absPath);
|
||||
groups += matches.length;
|
||||
perFixtureDigests.push(`${key}\t${matches.length}\t${digestCaptures(matches)}`);
|
||||
}
|
||||
// Fixed 20-entity synthetic source so the fingerprint is comparable across
|
||||
// experiments independent of the timing sizes.
|
||||
const daoMatches = emitPythonScopeCaptures(generatePyDao(20), 'synthetic-dao-20.py');
|
||||
groups += daoMatches.length;
|
||||
perFixtureDigests.push(`synthetic:dao-20\t${daoMatches.length}\t${digestCaptures(daoMatches)}`);
|
||||
const fingerprint = crypto
|
||||
.createHash('sha256')
|
||||
.update(perFixtureDigests.sort().join('\n'))
|
||||
.digest('hex');
|
||||
return { fingerprint, groups, fixtureCount: perFixtureDigests.length };
|
||||
}
|
||||
|
||||
// Higher rep count keeps the median stable on noisy shared CI runners.
|
||||
const REPS = 7;
|
||||
const SCALING_BUDGET = 1.5; // ~3.2 (quadratic) vs ~1.0 (linear); 1.5 has headroom.
|
||||
const CHECK = process.argv.includes('--check');
|
||||
|
||||
const fp = computeFingerprint();
|
||||
const small = measureSize(250, REPS);
|
||||
const large = measureSize(800, REPS);
|
||||
const scalingRatio = small.ms > 0 ? large.ms / small.ms / (800 / 250) : 0;
|
||||
|
||||
const result = {
|
||||
elapsed_ms_250: Number(small.ms.toFixed(2)),
|
||||
elapsed_ms_800: Number(large.ms.toFixed(2)),
|
||||
scaling_ratio: Number(scalingRatio.toFixed(3)),
|
||||
capture_groups_250: small.count,
|
||||
capture_groups_800: large.count,
|
||||
fingerprint: fp.fingerprint,
|
||||
capture_groups_fp: fp.groups,
|
||||
fixture_count: fp.fixtureCount,
|
||||
};
|
||||
|
||||
if (!CHECK) {
|
||||
process.stdout.write(JSON.stringify(result) + '\n');
|
||||
} else {
|
||||
// CI gate: capture output unchanged (fingerprint == committed baseline) AND
|
||||
// the path is still linear (scaling ratio under budget). Re-baseline a
|
||||
// legitimate capture change with `node --import tsx measure.mjs` (no --check)
|
||||
// and commit the new baseline-fingerprint.txt deliberately.
|
||||
const baselinePath = path.resolve(__dirname, 'baseline-fingerprint.txt');
|
||||
const baseline = fs.readFileSync(baselinePath, 'utf8').trim();
|
||||
const failures = [];
|
||||
if (result.fingerprint !== baseline) {
|
||||
failures.push(
|
||||
`capture fingerprint drift: got ${result.fingerprint}, expected ${baseline} ` +
|
||||
`(emitPythonScopeCaptures output changed — re-baseline intentionally if expected)`,
|
||||
);
|
||||
}
|
||||
if (result.scaling_ratio >= SCALING_BUDGET) {
|
||||
failures.push(
|
||||
`capture scaling ratio ${result.scaling_ratio} >= ${SCALING_BUDGET} ` +
|
||||
`(possible O(n^2) regression; 250->800ms ${result.elapsed_ms_250}->${result.elapsed_ms_800})`,
|
||||
);
|
||||
}
|
||||
process.stdout.write(JSON.stringify(result) + '\n');
|
||||
if (failures.length > 0) {
|
||||
for (const f of failures) process.stderr.write(`[measure --check] FAIL: ${f}\n`);
|
||||
process.exit(1);
|
||||
}
|
||||
process.stderr.write('[measure --check] PASS (capture fingerprint + scaling)\n');
|
||||
}
|
||||
@@ -0,0 +1,81 @@
|
||||
{
|
||||
"_comment": "Per-language baselines for bench/scope-capture/measure.mjs --check. fingerprint = order-independent sha256 over the lang-resolution/<lang>-* fixture corpus + a 20-entity synthetic source (correctness gate; re-baseline intentionally on a legitimate capture change). scaling_budget = max allowed (t800/t250)/(800/250); ~1.0 is linear, ~3.2 is quadratic. The synthetic source is now HERITAGE-BEARING for every language (each Entity extends/implements/embeds/uses-trait/conforms-to a shared base) so the #1951 @reference.inherits synth is gated at scale, not just the base capture loop. All languages thread the tree-sitter captured node instead of re-deriving it with findNodeAtRange(tree.rootNode,...) per match, so all are linear (go #1915, python #1918, ruby/php/rust/csharp #1951, java #1956).",
|
||||
"go": {
|
||||
"fingerprint": "a909c197b07921f974de1a8a47cc997b9580153db2d400ef13d205b6c1de5865",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined": "#1966: Go structural interface implementation detection changes capture output (method_elem, return-type, pointer receiver raw form, struct_type/interface_type containers)."
|
||||
},
|
||||
"cobol": {
|
||||
"fingerprint": "68ee0e95eb9f86f2d92ca35f730f4c2d4d83abc1b5241ae767ff3437780ec8d1",
|
||||
"scaling_budget": 1.5,
|
||||
"_note": "Updated for F17-F23 fixes (P2: TIMES guard, ADD GIVING, SQL AS alias). See PR #1959."
|
||||
},
|
||||
"c": {
|
||||
"fingerprint": "0de009bdbfe095f530fa87eb32bce6ab83092c904f26b3c8fe8d8ab587cf6dc9",
|
||||
"scaling_budget": 1.5,
|
||||
"_added": "#1956: c added to the scope-capture bench (was UNBENCHED). C has no inheritance \u2014 flat scale source. Adding it exposed + fixed a pre-existing O(n^2) findNodeAtRange root-walk in c/captures.ts (threaded c.node, byte-identical over c-* fixtures); scaling 3.475 -> 0.96."
|
||||
},
|
||||
"cpp": {
|
||||
"fingerprint": "931bf7af55dc1480d1a5d3c479ea3803003a6a2e2c4406447bd96f3e312e88de",
|
||||
"scaling_budget": 1.5,
|
||||
"_added": "#1956: cpp added to the scope-capture bench (was UNBENCHED). Heritage-bearing scale source (: public Base, public Mixin) drives emitCppInheritanceCaptures at scale. Adding it exposed + fixed a pre-existing O(n^2) findNodeAtRange root-walk in cpp/captures.ts (~12 sites, threaded c.node, byte-identical over 263 cpp-* fixtures); scaling 2.30 -> 1.12.",
|
||||
"_rebaselined": "#1965 / #1923 F4: uninitialized non-leading multi-declarators now emit @declaration.variable captures; cpp-adl-inner-callable-outer-noncallable data::Pair a, b adds the legitimate fixture drift. Linear (~1.06).",
|
||||
"_note": "#1975: + cpp-out-of-line-class fixture (out-of-line struct Outer::Inner / Other::Inner). Pure fixture-corpus drift — the fix is the legacy structure-query qualified_identifier arm, NOT the cpp scope-extractor; existing fixtures' captures byte-identical. fixture_count 263->265."
|
||||
},
|
||||
"csharp": {
|
||||
"_rebaselined": "#1956 synth-widening: + csharp-qualified-base fixture; the synth now walks record_declaration + struct_declaration base_lists and handles alias_qualified_name (matching the #1940 legacy leg), so record/struct heritage now emits. csharp-record-base gains a record inherits capture. (record->record SAME-namespace EXTENDS is a separate registry resolution gap, tracked as follow-up.) Linear (~1.00). (Earlier #1956: heritage-bearing scale source.)",
|
||||
"fingerprint": "68ef32c126d5c6de5d8184c6ad0a6104043036daf9805947db8b21741b883f43",
|
||||
"scaling_budget": 1.5
|
||||
},
|
||||
"rust": {
|
||||
"fingerprint": "3c4b8e0a707299cc5db0af2528c72a99457859104589a7ef3cd1f377da01793e",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined": "#1956 tri-review U1: rust-qualified-trait fixture (scoped + generic-of-scoped impl trait paths); bareTypeIdentifier now resolves scoped_type_identifier bases by their name: tail (additive, no existing-fixture drift); linear (~1.04).",
|
||||
"_note": "#1975: + rust-scoped-impl fixture (impl a::Inner / b::Inner inherent scoped impls). Pure fixture-corpus drift — the fix is the legacy @definition.impl scoped arm + findEnclosingClassInfo inherent-impl scoped target, NOT the rust scope-extractor; existing fixtures' captures byte-identical. fixture_count 120->121."
|
||||
},
|
||||
"php": {
|
||||
"fingerprint": "f9c8eaf6d1084f9b95a9fb97ccce5e618a24d936c85fb8af4b96c73a560f7a7f",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined": "#1956: heritage-bearing scale source (class extends Base + use trait); both forms gated at scale; linear (~1.04)."
|
||||
},
|
||||
"ruby": {
|
||||
"fingerprint": "ee81145cf0af796878e8e048192b87c8c8dc445a3e3fcdff6c6e26c179e97232",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined": "#1956 synth-widening: + ruby-qualified-base fixture; synth now reduces a scope_resolution superclass (class C < Mod::Super) to its trailing constant (matching the #1940 legacy leg), at parity. Linear (~1.03). (Earlier #1956: heritage-bearing scale source.)",
|
||||
"_note": "F62: + scope_resolution class/module declaration captures — fixture count 78→81, fingerprint drift expected. #1975: + ruby-tail-collision fixture (Foo::Bar vs Baz::Bar stay distinct nodes) — pure fixture-corpus drift, scope-extractor captures unchanged; 81→82."
|
||||
},
|
||||
"swift": {
|
||||
"fingerprint": "53325c6345161c5a495f997297af5a24fb718fd3e6647040160f8ab2a2c8e4c0",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined": "#1956: swift-qualified-base fixture + heritage-bearing scale source (class: Base, Serviceable \u2014 extends + protocol conformance); linear (~1.03)."
|
||||
},
|
||||
"dart": {
|
||||
"fingerprint": "a9e882b537765e8fd0ddfcd33b38b253dd86fc5ddffa6e4bf5a85ed8ee615eaa",
|
||||
"scaling_budget": 1.5,
|
||||
"_added": "#939: dart added to the scope-capture bench with the registry-primary migration. Heritage-bearing scale source (Entity extends Base implements Marker) gates the @reference.inherits synth + the postfix-chain reference walk at scale. emitDartScopeCaptures threads tree-sitter captured nodes (no findNodeAtRange root-walk), so it is linear (~1.0).",
|
||||
"_rebaselined": "#1970 review + tri-review follow-ups: constructor-call retag, cascade calls, built-in suppression, enum scope, #1926 F24/F25, named-ctor dedup (crash fix), container-name binding suppression; heritage file-affinity resolution. Fixtures: member-call-contexts, constructor-body, named-constructor-body, heritage-name-collision, construct-cascade."
|
||||
},
|
||||
"java": {
|
||||
"fingerprint": "b63f9be458f7ece854e7b007159d7bf65b4b66a86e83a6c0656fc93ebd5d83da",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined": "#1956 synth-widening: + java-iface-extends fixture; synthesizeJavaInheritanceReferences now ALSO walks interface_declaration extends_interfaces (interface IA extends IB, IC<T>), matching the #1940 legacy leg. (Earlier U2+review: java-qualified-base fixture covers 2- AND 3-segment qualified bases guarding the legacy end-anchor; synth tail-resolves scoped bases.) Linear (~1.03). (Earliest: java added to bench, exposed+fixed the O(n^2) findNodeAtRange root-walk; 3.09 -> ~0.99.)"
|
||||
},
|
||||
"typescript": {
|
||||
"fingerprint": "3f44a4a6892698df2d145c8ff2812c3b318807648983c88aca28fbd694f172f9",
|
||||
"scaling_budget": 1.5,
|
||||
"_rebaselined": "#1962: F44 (class scope@), F85 (enum member declarations), F87 (optional_parameter type annotations) add new captures \u2014 fingerprint drift expected.",
|
||||
"_note": "#1968: F44, F85, F87 \u2014 fingerprint drift expected."
|
||||
},
|
||||
"javascript": {
|
||||
"fingerprint": "a8ddfb15620ae55e50651fc21ab14c4a1f874d9b19e208cc6cbf0a8daac8ec5b",
|
||||
"scaling_budget": 1.5,
|
||||
"_added": "#1951: bench coverage added (was ungated); scale source heritage-bearing (extends Base); js/kotlin O(n^2) findNodeAtRange-per-match fixed to threaded captured node, now linear.",
|
||||
"_rebaselined": "#1956 synth-widening: + javascript-qualified-base fixture; synthesizeJsInheritanceReferences now handles a member_expression base (class S extends ns.Base -> Base), matching the #1940 legacy leg + the TS terminalTsTypeNameNode property_identifier case, at parity. Linear (~1.05)."
|
||||
},
|
||||
"kotlin": {
|
||||
"fingerprint": "5121a11855cd9cc44a357ae3ff50953de80cdd743f00e8924c31503b132bcd84",
|
||||
"scaling_budget": 1.5,
|
||||
"_added": "#1951: bench coverage added (was ungated); scale source heritage-bearing (: Base()); js/kotlin O(n^2) findNodeAtRange-per-match fixed to threaded captured node, now linear.",
|
||||
"_rebaselined": "#1956 synth-widening: + kotlin-qualified-base fixture; synthesizeKotlinInheritanceReferences now handles the explicit_delegation form (class F : Iface by d -> Iface), matching the #1940 legacy leg, at parity. Linear (~0.87)."
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,409 @@
|
||||
/**
|
||||
* Unified build-free scope-capture measurement harness for every currently
|
||||
* benchmarked language (the ones with a `*-pipeline-benchmark.test.ts`):
|
||||
* go, csharp, rust, php, ruby, cobol, swift — plus python lives in its own
|
||||
* `bench/python-scope/` harness (richer: it also covers import resolution).
|
||||
*
|
||||
* For each language it:
|
||||
* - times `emit<Lang>ScopeCaptures` on a synthetic DAO-style source at two
|
||||
* sizes (250 / 800 top-level entities), reporting elapsed_ms + a scaling
|
||||
* ratio `(t_large/t_small)/(800/250)`: ~1.0 is linear, ~3.2 is quadratic
|
||||
* (the O(matches × rootChildren) shape #1848 hit in Go);
|
||||
* - computes an order-independent sha256 fingerprint over the whole
|
||||
* `lang-resolution/<lang>-*` fixture corpus + a fixed 20-entity synthetic
|
||||
* source, as the correctness gate.
|
||||
*
|
||||
* Build-free: imports the `.ts` hotpaths through tsx
|
||||
* (`node --import tsx bench/scope-capture/measure.mjs`). Static `.ts` imports
|
||||
* work; a top-level `await import()` breaks tsx's lexer.
|
||||
*
|
||||
* Without args: prints one JSON object per language.
|
||||
* With `--check`: asserts each language's fingerprint == its committed baseline
|
||||
* (baselines.json) AND scaling_ratio < that language's recorded budget; exits
|
||||
* non-zero on any drift/regression.
|
||||
*/
|
||||
import fs from 'node:fs';
|
||||
import path from 'node:path';
|
||||
import crypto from 'node:crypto';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
|
||||
import { emitGoScopeCaptures } from '../../src/core/ingestion/languages/go/index.ts';
|
||||
import { emitCsharpScopeCaptures } from '../../src/core/ingestion/languages/csharp/index.ts';
|
||||
import { emitRustScopeCaptures } from '../../src/core/ingestion/languages/rust/index.ts';
|
||||
import { emitPhpScopeCaptures } from '../../src/core/ingestion/languages/php/index.ts';
|
||||
import { emitRubyScopeCaptures } from '../../src/core/ingestion/languages/ruby/index.ts';
|
||||
import { emitCobolScopeCaptures } from '../../src/core/ingestion/languages/cobol/index.ts';
|
||||
import { emitSwiftScopeCaptures } from '../../src/core/ingestion/languages/swift/index.ts';
|
||||
import { emitTsScopeCaptures } from '../../src/core/ingestion/languages/typescript/index.ts';
|
||||
import { emitJsScopeCaptures } from '../../src/core/ingestion/languages/javascript/index.ts';
|
||||
import { emitKotlinScopeCaptures } from '../../src/core/ingestion/languages/kotlin/index.ts';
|
||||
import { emitJavaScopeCaptures } from '../../src/core/ingestion/languages/java/index.ts';
|
||||
import { emitCScopeCaptures } from '../../src/core/ingestion/languages/c/index.ts';
|
||||
import { emitCppScopeCaptures } from '../../src/core/ingestion/languages/cpp/index.ts';
|
||||
import { emitDartScopeCaptures } from '../../src/core/ingestion/languages/dart/index.ts';
|
||||
|
||||
const __dirname = path.dirname(fileURLToPath(import.meta.url));
|
||||
const FIXTURE_ROOT = path.resolve(__dirname, '..', '..', 'test', 'fixtures', 'lang-resolution');
|
||||
const BASELINE_PATH = path.resolve(__dirname, 'baselines.json');
|
||||
|
||||
// ---- correctness fingerprint (order-independent; mirrors python harness) ----
|
||||
|
||||
function canonicalizeMatch(match) {
|
||||
const parts = [];
|
||||
for (const tag of Object.keys(match)) {
|
||||
const cap = match[tag];
|
||||
if (cap === undefined || cap === null || cap.range === undefined) {
|
||||
parts.push(`${tag}|<no-range>`);
|
||||
continue;
|
||||
}
|
||||
const r = cap.range;
|
||||
parts.push(`${tag}|${cap.text}|${r.startLine}:${r.startCol}-${r.endLine}:${r.endCol}`);
|
||||
}
|
||||
parts.sort();
|
||||
return parts.join(';');
|
||||
}
|
||||
|
||||
function digestCaptures(matches) {
|
||||
return crypto
|
||||
.createHash('sha256')
|
||||
.update(matches.map(canonicalizeMatch).sort().join('\n'))
|
||||
.digest('hex');
|
||||
}
|
||||
|
||||
/** All fixture files for a language, sorted by repo-relative key. */
|
||||
function collectFixtures(prefix, exts) {
|
||||
const out = [];
|
||||
for (const entry of fs.readdirSync(FIXTURE_ROOT, { withFileTypes: true })) {
|
||||
if (!entry.isDirectory() || !entry.name.startsWith(`${prefix}-`)) continue;
|
||||
const stack = [path.join(FIXTURE_ROOT, entry.name)];
|
||||
while (stack.length) {
|
||||
const dir = stack.pop();
|
||||
for (const c of fs.readdirSync(dir, { withFileTypes: true })) {
|
||||
const p = path.join(dir, c.name);
|
||||
if (c.isDirectory()) stack.push(p);
|
||||
else if (exts.some((e) => c.name.endsWith(e))) {
|
||||
out.push({ key: path.relative(FIXTURE_ROOT, p).split(path.sep).join('/'), absPath: p });
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
out.sort((a, b) => a.key.localeCompare(b.key));
|
||||
return out;
|
||||
}
|
||||
|
||||
// ---- per-language config: synthetic DAO generators + fixture globs ----
|
||||
|
||||
const LANGS = [
|
||||
{
|
||||
name: 'go',
|
||||
emit: emitGoScopeCaptures,
|
||||
fixturePrefix: 'go',
|
||||
exts: ['.go'],
|
||||
file: 'bench.go',
|
||||
// Heritage-bearing: each Entity embeds Base (Go inheritance = struct
|
||||
// embedding) so the @reference.inherits synth (#1951) is driven at scale.
|
||||
header:
|
||||
'package generated\n\ntype Base struct{}\n\nfunc (b *Base) BaseMethod() string { return "base" }\n\n',
|
||||
unit: (n) =>
|
||||
`type Entity${n} struct {\n\tBase\n\tid int64\n\tname string\n}\n\n` +
|
||||
`func (e *Entity${n}) GetID() int64 { return e.id }\n` +
|
||||
`func (e *Entity${n}) SetName(v string) { e.name = v }\n\n`,
|
||||
},
|
||||
{
|
||||
name: 'csharp',
|
||||
emit: emitCsharpScopeCaptures,
|
||||
fixturePrefix: 'csharp',
|
||||
exts: ['.cs'],
|
||||
file: 'bench.cs',
|
||||
// Heritage-bearing: extends Base + implements IEntity (both forms) so the
|
||||
// @reference.inherits synth (#1951) is driven at scale, not just the base loop.
|
||||
header:
|
||||
'namespace Generated;\n\npublic class Base { }\n\npublic interface IEntity {\n long GetId();\n}\n\n',
|
||||
unit: (n) =>
|
||||
`public class Entity${n} : Base, IEntity {\n` +
|
||||
` public long Id;\n public string Name;\n` +
|
||||
` public long GetId() { return Id; }\n` +
|
||||
` public void SetName(string v) { Name = v; }\n}\n\n`,
|
||||
},
|
||||
{
|
||||
name: 'rust',
|
||||
emit: emitRustScopeCaptures,
|
||||
fixturePrefix: 'rust',
|
||||
exts: ['.rs'],
|
||||
file: 'bench.rs',
|
||||
// Heritage-bearing: `impl Shape for Entity_n` (Rust inheritance lives on
|
||||
// impl_item) so the @reference.inherits trait-impl synth (#1951) is driven
|
||||
// at scale. The two methods move into the trait impl to keep unit size flat.
|
||||
header: 'trait Shape {\n fn area(&self) -> i64;\n fn name(&self) -> String;\n}\n\n',
|
||||
unit: (n) =>
|
||||
`struct Entity${n} {\n id: i64,\n name: String,\n}\n\n` +
|
||||
`impl Shape for Entity${n} {\n` +
|
||||
` fn area(&self) -> i64 { self.id }\n` +
|
||||
` fn name(&self) -> String { self.name.clone() }\n}\n\n`,
|
||||
},
|
||||
{
|
||||
name: 'php',
|
||||
emit: emitPhpScopeCaptures,
|
||||
fixturePrefix: 'php',
|
||||
exts: ['.php'],
|
||||
file: 'bench.php',
|
||||
// Heritage-bearing: extends Base + uses a trait (both forms) so the
|
||||
// @reference.inherits synth (#1951) is driven at scale.
|
||||
header:
|
||||
'<?php\n\nclass Base {}\n\ntrait Auditable {\n public function audit() { return true; }\n}\n\n',
|
||||
unit: (n) =>
|
||||
`class Entity${n} extends Base {\n` +
|
||||
` use Auditable;\n` +
|
||||
` public $id;\n public $name;\n` +
|
||||
` function getId() { return $this->id; }\n` +
|
||||
` function setName($v) { $this->name = $v; }\n}\n\n`,
|
||||
},
|
||||
{
|
||||
name: 'ruby',
|
||||
emit: emitRubyScopeCaptures,
|
||||
fixturePrefix: 'ruby',
|
||||
exts: ['.rb'],
|
||||
file: 'bench.rb',
|
||||
// Heritage-bearing: `< Base` superclass + `include Trackable` mixin (both
|
||||
// forms) so the @reference.inherits synth (#1951) is driven at scale.
|
||||
header:
|
||||
'class Base\n def base_id\n @id\n end\nend\n\nmodule Trackable\n def track\n @tracked = true\n end\nend\n\n',
|
||||
unit: (n) =>
|
||||
`class Entity${n} < Base\n` +
|
||||
` include Trackable\n` +
|
||||
` def get_id\n @id\n end\n` +
|
||||
` def set_name(v)\n @name = v\n end\nend\n\n`,
|
||||
},
|
||||
{
|
||||
name: 'cobol',
|
||||
emit: emitCobolScopeCaptures,
|
||||
fixturePrefix: 'cobol',
|
||||
exts: ['.cbl', '.cpy'],
|
||||
file: 'bench.cbl',
|
||||
header:
|
||||
' IDENTIFICATION DIVISION.\n' +
|
||||
' PROGRAM-ID. BENCH.\n' +
|
||||
' PROCEDURE DIVISION.\n',
|
||||
unit: (n) => ` PARA-${String(n).padStart(5, '0')}.\n DISPLAY "P${n}".\n`,
|
||||
},
|
||||
{
|
||||
name: 'c',
|
||||
emit: emitCScopeCaptures,
|
||||
fixturePrefix: 'c',
|
||||
exts: ['.c', '.h'],
|
||||
file: 'bench.c',
|
||||
// C has no inheritance construct — flat scale source. Added (was unbenched);
|
||||
// adding it exposed + fixed the same O(n²) findNodeAtRange root-walk (#1956).
|
||||
header: '#include <stdint.h>\n#include <stddef.h>\n\ntypedef int64_t id_t;\n\n',
|
||||
unit: (n) =>
|
||||
`typedef struct Entity${n} {\n id_t id;\n const char *name;\n} Entity${n};\n\n` +
|
||||
`id_t entity_${n}_get_id(Entity${n} *e) { return e->id; }\n` +
|
||||
`void entity_${n}_set_name(Entity${n} *e, const char *v) { e->name = v; }\n\n`,
|
||||
},
|
||||
{
|
||||
name: 'cpp',
|
||||
emit: emitCppScopeCaptures,
|
||||
fixturePrefix: 'cpp',
|
||||
exts: ['.cpp', '.cc', '.cxx', '.hpp', '.h'],
|
||||
file: 'bench.cpp',
|
||||
// Heritage-bearing: `: public Base, public Mixin` (single + multiple
|
||||
// inheritance) drives emitCppInheritanceCaptures (#1951) at scale. Added
|
||||
// (was unbenched); adding it exposed + fixed the same O(n²) root-walk (#1956).
|
||||
header:
|
||||
'#include <string>\n\nclass Base {\n public:\n long baseId() const { return 0; }\n};\n\nclass Mixin {\n public:\n void mix() {}\n};\n\n',
|
||||
unit: (n) =>
|
||||
`class Entity${n} : public Base, public Mixin {\n public:\n long id;\n std::string name;\n` +
|
||||
` long getId() const { return id; }\n` +
|
||||
` void setName(std::string v) { name = v; }\n};\n\n`,
|
||||
},
|
||||
{
|
||||
name: 'swift',
|
||||
emit: emitSwiftScopeCaptures,
|
||||
fixturePrefix: 'swift',
|
||||
exts: ['.swift'],
|
||||
file: 'bench.swift',
|
||||
// Heritage-bearing: inherits Base + conforms to Serviceable (both forms) so
|
||||
// the @reference.inherits synth (#1951) is driven at scale.
|
||||
header:
|
||||
'class Base {\n func ping() -> String { return "base" }\n}\n\nprotocol Serviceable {\n func serve() -> String\n}\n\n',
|
||||
unit: (n) =>
|
||||
`class Entity${n}: Base, Serviceable {\n` +
|
||||
` var id: Int64 = 0\n var name: String = ""\n` +
|
||||
` func getId() -> Int64 { return self.id }\n` +
|
||||
` func serve() -> String { return self.name }\n}\n\n`,
|
||||
},
|
||||
{
|
||||
name: 'dart',
|
||||
emit: emitDartScopeCaptures,
|
||||
fixturePrefix: 'dart',
|
||||
exts: ['.dart'],
|
||||
file: 'bench.dart',
|
||||
// Heritage-bearing: `extends Base` (generic @reference.inherits pre-pass)
|
||||
// + `implements Marker` (Dart `implements <class>` → IMPLEMENTS marker) so
|
||||
// both heritage paths and the postfix-chain reference walk run at scale.
|
||||
header:
|
||||
'class Base {\n String ping() { return "base"; }\n}\n\nabstract class Marker {\n String mark();\n}\n\n',
|
||||
unit: (n) =>
|
||||
`class Entity${n} extends Base implements Marker {\n` +
|
||||
` int id = 0;\n String name = '';\n` +
|
||||
` int getId() { return this.id; }\n` +
|
||||
` String mark() { return this.name; }\n}\n\n`,
|
||||
},
|
||||
{
|
||||
name: 'java',
|
||||
emit: emitJavaScopeCaptures,
|
||||
fixturePrefix: 'java',
|
||||
exts: ['.java'],
|
||||
file: 'bench.java',
|
||||
// Java was previously unbenched. Heritage-bearing: extends Base + implements
|
||||
// Marker (both forms) so the @reference.inherits synth (#1951) is driven at scale.
|
||||
header: 'package generated;\n\nclass Base {}\n\ninterface Marker {}\n\n',
|
||||
unit: (n) =>
|
||||
`class Entity${n} extends Base implements Marker {\n` +
|
||||
` long id = 0L;\n String name = "";\n` +
|
||||
` public long getId() { return this.id; }\n` +
|
||||
` public void setName(String v) { this.name = v; }\n}\n\n`,
|
||||
},
|
||||
{
|
||||
name: 'typescript',
|
||||
emit: emitTsScopeCaptures,
|
||||
fixturePrefix: 'typescript',
|
||||
exts: ['.ts', '.tsx'],
|
||||
file: 'bench.ts',
|
||||
// Inheritance-bearing units so the @reference.inherits synth pass (#1951)
|
||||
// is exercised at scale, not just the base capture loop.
|
||||
header: 'class Base {}\n\n',
|
||||
unit: (n) =>
|
||||
`class Entity${n} extends Base {\n` +
|
||||
` id: number = 0;\n name: string = '';\n` +
|
||||
` getId(): number { return this.id; }\n` +
|
||||
` setName(v: string): void { this.name = v; }\n}\n\n`,
|
||||
},
|
||||
{
|
||||
name: 'javascript',
|
||||
emit: emitJsScopeCaptures,
|
||||
fixturePrefix: 'javascript',
|
||||
exts: ['.js', '.jsx', '.mjs', '.cjs'],
|
||||
file: 'bench.js',
|
||||
header: 'class Base {}\n\n',
|
||||
unit: (n) =>
|
||||
`class Entity${n} extends Base {\n` +
|
||||
` getId() { return this.id; }\n` +
|
||||
` setName(v) { this.name = v; }\n}\n\n`,
|
||||
},
|
||||
{
|
||||
name: 'kotlin',
|
||||
emit: emitKotlinScopeCaptures,
|
||||
fixturePrefix: 'kotlin',
|
||||
exts: ['.kt', '.kts'],
|
||||
file: 'bench.kt',
|
||||
header: 'open class Base\n\n',
|
||||
unit: (n) =>
|
||||
`class Entity${n} : Base() {\n` +
|
||||
` var id: Long = 0\n var name: String = ""\n` +
|
||||
` fun getId(): Long { return id }\n` +
|
||||
` fun setName(v: String) { name = v }\n}\n\n`,
|
||||
},
|
||||
];
|
||||
|
||||
function generate(lang, entityCount) {
|
||||
let src = lang.header;
|
||||
for (let i = 0; i < entityCount; i++) src += lang.unit(i);
|
||||
return src;
|
||||
}
|
||||
|
||||
// ---- timing ----
|
||||
|
||||
function median(xs) {
|
||||
const s = [...xs].sort((a, b) => a - b);
|
||||
const m = Math.floor(s.length / 2);
|
||||
return s.length % 2 ? s[m] : (s[m - 1] + s[m]) / 2;
|
||||
}
|
||||
|
||||
function timeEmit(emit, src, file, reps) {
|
||||
emit(src, `warmup-${file}`); // warm parser/query JIT (not counted)
|
||||
const samples = [];
|
||||
let count = 0;
|
||||
for (let i = 0; i < reps; i++) {
|
||||
const start = process.hrtime.bigint();
|
||||
const out = emit(src, file);
|
||||
samples.push(Number(process.hrtime.bigint() - start) / 1e6);
|
||||
count = out.length;
|
||||
}
|
||||
return { ms: median(samples), count };
|
||||
}
|
||||
|
||||
const SMALL = 250;
|
||||
const LARGE = 800;
|
||||
const REPS = 7;
|
||||
|
||||
function measureLang(lang) {
|
||||
// Correctness fingerprint over the fixture corpus + a fixed 20-entity source.
|
||||
const perFixture = [];
|
||||
let groups = 0;
|
||||
for (const { key, absPath } of collectFixtures(lang.fixturePrefix, lang.exts)) {
|
||||
const matches = lang.emit(fs.readFileSync(absPath, 'utf8'), absPath);
|
||||
groups += matches.length;
|
||||
perFixture.push(`${key}\t${matches.length}\t${digestCaptures(matches)}`);
|
||||
}
|
||||
const daoMatches = lang.emit(generate(lang, 20), `synthetic-dao-20${path.extname(lang.file)}`);
|
||||
groups += daoMatches.length;
|
||||
perFixture.push(`synthetic:dao-20\t${daoMatches.length}\t${digestCaptures(daoMatches)}`);
|
||||
const fingerprint = crypto
|
||||
.createHash('sha256')
|
||||
.update(perFixture.sort().join('\n'))
|
||||
.digest('hex');
|
||||
|
||||
// Scaling.
|
||||
const small = timeEmit(lang.emit, generate(lang, SMALL), lang.file, REPS);
|
||||
const large = timeEmit(lang.emit, generate(lang, LARGE), lang.file, REPS);
|
||||
const scalingRatio = small.ms > 0 ? large.ms / small.ms / (LARGE / SMALL) : 0;
|
||||
|
||||
return {
|
||||
language: lang.name,
|
||||
elapsed_ms_small: Number(small.ms.toFixed(2)),
|
||||
elapsed_ms_large: Number(large.ms.toFixed(2)),
|
||||
scaling_ratio: Number(scalingRatio.toFixed(3)),
|
||||
capture_groups_small: small.count,
|
||||
capture_groups_large: large.count,
|
||||
fingerprint,
|
||||
capture_groups_fp: groups,
|
||||
fixture_count: perFixture.length,
|
||||
};
|
||||
}
|
||||
|
||||
// ---- run ----
|
||||
|
||||
const CHECK = process.argv.includes('--check');
|
||||
const results = LANGS.map(measureLang);
|
||||
|
||||
if (!CHECK) {
|
||||
for (const r of results) process.stdout.write(JSON.stringify(r) + '\n');
|
||||
} else {
|
||||
const baselines = JSON.parse(fs.readFileSync(BASELINE_PATH, 'utf8'));
|
||||
const failures = [];
|
||||
for (const r of results) {
|
||||
const base = baselines[r.language];
|
||||
if (base === undefined) {
|
||||
failures.push(`${r.language}: no baseline recorded`);
|
||||
continue;
|
||||
}
|
||||
if (r.fingerprint !== base.fingerprint) {
|
||||
failures.push(
|
||||
`${r.language}: capture fingerprint drift (got ${r.fingerprint}, expected ${base.fingerprint})`,
|
||||
);
|
||||
}
|
||||
if (r.scaling_ratio >= base.scaling_budget) {
|
||||
failures.push(
|
||||
`${r.language}: scaling ratio ${r.scaling_ratio} >= budget ${base.scaling_budget} ` +
|
||||
`(${SMALL}->${LARGE} ms ${r.elapsed_ms_small}->${r.elapsed_ms_large})`,
|
||||
);
|
||||
}
|
||||
process.stdout.write(JSON.stringify(r) + '\n');
|
||||
}
|
||||
if (failures.length > 0) {
|
||||
for (const f of failures) process.stderr.write(`[scope-capture --check] FAIL: ${f}\n`);
|
||||
process.exit(1);
|
||||
}
|
||||
process.stderr.write(`[scope-capture --check] PASS (${results.length} languages)\n`);
|
||||
}
|
||||
@@ -25,6 +25,7 @@ const path = require('path');
|
||||
const { spawnSync } = require('child_process');
|
||||
const { acquireHookSlot } = require('./hook-lock.cjs');
|
||||
const { hasGitNexusDbLockedByGitNexusServer } = require('./hook-db-lock-probe.cjs');
|
||||
const { formatAnalyzeCommand } = require('./resolve-analyze-cmd.cjs');
|
||||
|
||||
function readInput() {
|
||||
try {
|
||||
@@ -315,7 +316,7 @@ function buildStaleIndexHint(gitNexusDir, cwd) {
|
||||
|
||||
if (currentHead === lastCommit) return '';
|
||||
|
||||
const analyzeCmd = `npx gitnexus analyze${hadEmbeddings ? ' --embeddings' : ''}`;
|
||||
const analyzeCmd = formatAnalyzeCommand({ embeddings: hadEmbeddings });
|
||||
return (
|
||||
`[GitNexus] index is stale (last indexed: ${lastCommit ? lastCommit.slice(0, 7) : 'never'}). ` +
|
||||
`Run \`${analyzeCmd}\` to refresh the knowledge graph.`
|
||||
|
||||
@@ -16,6 +16,7 @@ const path = require('path');
|
||||
const { spawnSync } = require('child_process');
|
||||
const { acquireHookSlot } = require('./hook-lock.cjs');
|
||||
const { hasGitNexusDbLockedByGitNexusServer } = require('./hook-db-lock-probe.cjs');
|
||||
const { formatAnalyzeCommand } = require('./resolve-analyze-cmd.cjs');
|
||||
|
||||
/**
|
||||
* Read JSON input from stdin synchronously.
|
||||
@@ -340,7 +341,7 @@ function handlePostToolUse(input) {
|
||||
// If HEAD matches last indexed commit, no reindex needed
|
||||
if (currentHead && currentHead === lastCommit) return;
|
||||
|
||||
const analyzeCmd = `npx gitnexus analyze${hadEmbeddings ? ' --embeddings' : ''}`;
|
||||
const analyzeCmd = formatAnalyzeCommand({ embeddings: hadEmbeddings });
|
||||
sendHookResponse(
|
||||
'PostToolUse',
|
||||
`GitNexus index is stale (last indexed: ${lastCommit ? lastCommit.slice(0, 7) : 'never'}). ` +
|
||||
|
||||
@@ -0,0 +1,323 @@
|
||||
/**
|
||||
* Single source of truth for how docs, hooks, and warnings invoke gitnexus.
|
||||
*
|
||||
* Automatically selects a working invocation path:
|
||||
* 1. Global `gitnexus` on PATH (best — no install step)
|
||||
* 2. npm 11+ with pnpm on PATH → `pnpm --allow-build=… dlx` (avoids the npx
|
||||
* arborist crash *and* pnpm 10+ ignored-build-script failures, #1939)
|
||||
* 3. npm < 11 with npm on PATH → `npx` (works; simpler than pnpm dlx)
|
||||
* 4. pnpm-only → `pnpm --allow-build=… dlx`
|
||||
* 5. Last resort → `npx` (warned on npm 11+ from analyze.ts)
|
||||
*
|
||||
* The `--allow-build` flags MUST precede the `dlx` token. pnpm < 10.14 keeps
|
||||
* `dlx` in its argv escape list, so flags placed *after* `dlx` are parsed as
|
||||
* package specs (ERR_PNPM_SPEC_NOT_SUPPORTED). The pre-`dlx` position parses
|
||||
* into dlx's allow-build option and has been honored since pnpm 10.2.0 (#1939).
|
||||
*
|
||||
* This stays self-contained CJS because the Claude/Antigravity hooks run as
|
||||
* standalone files copied into the user's hook dir, where no package import is
|
||||
* available. The CLI reuses this module from src/cli/resolve-invocation.ts via
|
||||
* createRequire rather than re-implementing it. Two committed copies must stay
|
||||
* byte-identical (enforced by resolve-invocation.test.ts) — edit both together:
|
||||
* gitnexus/hooks/claude/ (the canonical copy the CLI and `gitnexus setup` read)
|
||||
* and gitnexus-claude-plugin/hooks/. A THIRD copy is written at runtime to
|
||||
* `<repo>/.gitnexus/run.cjs` by `gitnexus analyze` (ai-context.ts) so docs can
|
||||
* reference it directly via the `require.main === module` exec tail below; that
|
||||
* copy is gitignored and refreshed on every analyze, so it cannot drift for long.
|
||||
*/
|
||||
|
||||
const { execFileSync } = require('child_process');
|
||||
const fs = require('fs');
|
||||
const path = require('path');
|
||||
|
||||
const NPX_REF = 'gitnexus@latest';
|
||||
|
||||
// Native packages whose postinstall must run under pnpm 10+ (blocked by default).
|
||||
const PNPM_ALLOW_BUILD_BASE = ['@ladybugdb/core', 'gitnexus', 'tree-sitter'];
|
||||
const PNPM_ALLOW_BUILD_EMBEDDINGS = ['onnxruntime-node'];
|
||||
|
||||
// Version-probe timeout, kept under Claude Code's 10s hook budget. PATH presence
|
||||
// detection is now spawn-free (resolveOnPath scans PATH directly), so the only
|
||||
// subprocesses left are the version probes: in a linked worktree the stale-index
|
||||
// hook first runs `git rev-parse --git-common-dir` (~2s) and `git rev-parse HEAD`
|
||||
// (~3s); the pnpm path then adds up to two 1s `--version` probes (npm, pnpm), so
|
||||
// the worst case is ~7s — within budget. A healthy `--version` returns in well
|
||||
// under a second, so the realistic cost is far lower.
|
||||
const PROBE_TIMEOUT_MS = 1000;
|
||||
|
||||
/**
|
||||
* Absolute path to `command` on PATH, or null — a pure-Node, spawn-free lookup
|
||||
* that mirrors how a shell resolves a bare command name: each PATH dir × the
|
||||
* platform's executable extensions (PATHEXT on Windows; the bare name + X_OK on
|
||||
* POSIX). This replaces the former `where`/`which` subprocess (#1938 "Option A"):
|
||||
* it is byte-for-byte identical on every OS, with no dependency on the probe
|
||||
* binary being reachable (a sanitized PATH that drops System32 / `/usr/bin` no
|
||||
* longer defeats detection), no shell-spawn surface (CVE-2024-27980), and no
|
||||
* spawn timeout to tune. On Windows it matches PATHEXT extensions ONLY — exactly
|
||||
* what `where`/cmd.exe resolve — so neither an un-spawnable `.ps1`-only shim (not
|
||||
* in default PATHEXT) nor a bare extensionless file (which the shell cannot launch
|
||||
* as `command`) is a false positive. `preferExecExt` returns a recognized
|
||||
* `.cmd`/`.bat`/`.exe` shim ahead of an exotic PATHEXT hit (e.g. `.COM`) when both
|
||||
* match, matching what a user would actually launch. Pure (platform/env injectable)
|
||||
* so it is unit-testable without touching the host PATH.
|
||||
*/
|
||||
function resolveOnPath(
|
||||
command,
|
||||
preferExecExt = false,
|
||||
{ platform = process.platform, env = process.env } = {},
|
||||
) {
|
||||
const pathValue = env.PATH || env.Path || env.path || '';
|
||||
if (!pathValue) return null;
|
||||
const isWin = platform === 'win32';
|
||||
const exts = isWin
|
||||
? (env.PATHEXT || '.COM;.EXE;.BAT;.CMD')
|
||||
.split(';')
|
||||
.map((e) => e.trim())
|
||||
.filter(Boolean)
|
||||
.map((e) => (e.startsWith('.') ? e : `.${e}`))
|
||||
: [''];
|
||||
let weakHit = null;
|
||||
// Split on the host's PATH delimiter. `platform` is injected only to choose the
|
||||
// extension/exec-bit rules; the PATH string is always host-format, so it must
|
||||
// split on the host delimiter (`path.delimiter`) — in production `platform` IS
|
||||
// the host, so they coincide. (Deriving the delimiter from an injected platform
|
||||
// would split a Windows drive-letter path `C:\…` at its colon under a POSIX
|
||||
// injection.)
|
||||
for (const dir of pathValue.split(path.delimiter).filter(Boolean)) {
|
||||
for (const ext of exts) {
|
||||
const candidate = path.join(dir, `${command}${ext}`);
|
||||
try {
|
||||
if (!fs.statSync(candidate).isFile()) continue;
|
||||
if (!isWin) fs.accessSync(candidate, fs.constants.X_OK);
|
||||
// Prefer a runnable .cmd/.bat/.exe shim; remember an exotic PATHEXT hit
|
||||
// (e.g. .COM) only as a last resort if nothing better turns up.
|
||||
if (isWin && preferExecExt && !/\.(cmd|bat|exe)$/i.test(ext)) {
|
||||
weakHit = weakHit || candidate;
|
||||
continue;
|
||||
}
|
||||
return candidate;
|
||||
} catch {
|
||||
/* not a runnable file here — try the next candidate */
|
||||
}
|
||||
}
|
||||
}
|
||||
return weakHit;
|
||||
}
|
||||
|
||||
// One spawn of `<command> --version` → { major, minor } (each null when
|
||||
// unreadable). Version injection happens at the resolver seam (getNpmMajorVersion
|
||||
// / formatPnpmAllowBuildArgs), so this stays a pure real-process probe.
|
||||
function probeVersion(command) {
|
||||
try {
|
||||
const output = execFileSync(command, ['--version'], {
|
||||
encoding: 'utf-8',
|
||||
timeout: PROBE_TIMEOUT_MS,
|
||||
stdio: ['ignore', 'pipe', 'ignore'],
|
||||
windowsHide: true,
|
||||
// On Windows, npm/pnpm resolve to `.cmd` shims; execFileSync does no
|
||||
// PATHEXT resolution and Node refuses to spawn `.cmd`/`.bat` without a
|
||||
// shell (CVE-2024-27980), so a bare `<command> --version` ENOENTs and the
|
||||
// probe would wrongly report a present tool as absent. A shell lets the OS
|
||||
// resolve the shim. POSIX needs no shell (direct PATH lookup works).
|
||||
shell: process.platform === 'win32',
|
||||
});
|
||||
// Find the first line that starts with a version token (`MAJOR.MINOR`,
|
||||
// optional `v` prefix) rather than splitting the whole output — pnpm/npm
|
||||
// under Corepack or with an update notice can print a banner line on stdout
|
||||
// before the version (stderr is already dropped via the stdio config).
|
||||
const versionLine = output
|
||||
.split('\n')
|
||||
.map((l) => l.trim())
|
||||
.find((l) => /^v?\d+\.\d+/.test(l));
|
||||
const match = versionLine ? versionLine.match(/^v?(\d+)\.(\d+)/) : null;
|
||||
return {
|
||||
major: match ? Number(match[1]) : null,
|
||||
minor: match ? Number(match[2]) : null,
|
||||
};
|
||||
} catch {
|
||||
return { major: null, minor: null };
|
||||
}
|
||||
}
|
||||
|
||||
// `deps` is the single injection seam: an explicitly provided key — including a
|
||||
// `null` value, detected via `in` — is honored as-is so tests can simulate an
|
||||
// absent tool without spawning; an absent key falls through to the real probe.
|
||||
function getNpmMajorVersion(deps = {}) {
|
||||
return 'npmMajor' in deps ? deps.npmMajor : probeVersion('npm').major;
|
||||
}
|
||||
|
||||
/**
|
||||
* `--allow-build` flags for the pre-`dlx` position. Emitted for pnpm >= 10.2
|
||||
* (where the flag exists, and pnpm 10+ blocks build scripts by default). Omitted
|
||||
* below 10.2: pnpm < 10 runs build scripts anyway, and pnpm 10.0/10.1 lack the
|
||||
* flag (it would be rejected as an unknown option). `alwaysAllowBuild` forces the
|
||||
* flags for committed documentation, which cannot probe the reader's pnpm.
|
||||
*/
|
||||
function formatPnpmAllowBuildArgs(options = {}, deps = {}) {
|
||||
if (!options.alwaysAllowBuild) {
|
||||
const { major, minor } =
|
||||
'pnpmMajor' in deps
|
||||
? { major: deps.pnpmMajor, minor: 'pnpmMinor' in deps ? deps.pnpmMinor : null }
|
||||
: probeVersion('pnpm');
|
||||
const lacksAllowBuild =
|
||||
major !== null && (major < 10 || (major === 10 && minor !== null && minor < 2));
|
||||
if (lacksAllowBuild) return [];
|
||||
}
|
||||
const pkgs = [...PNPM_ALLOW_BUILD_BASE];
|
||||
if (options.embeddings) pkgs.push(...PNPM_ALLOW_BUILD_EMBEDDINGS);
|
||||
return pkgs.map((p) => `--allow-build=${p}`);
|
||||
}
|
||||
|
||||
/** Fixed install-free command for committed AGENTS.md / SKILL.md (pnpm >= 10.2). */
|
||||
function formatDocumentationDlxCommand(gitnexusArgs, options = {}) {
|
||||
const flags = formatPnpmAllowBuildArgs({ ...options, alwaysAllowBuild: true }).join(' ');
|
||||
const prefix = flags ? `${flags} ` : '';
|
||||
return `pnpm ${prefix}dlx ${NPX_REF} ${gitnexusArgs}`;
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve `gitnexus` | `pnpm` | `npx`. `GITNEXUS_INVOCATION` forces a mode
|
||||
* (test/escape hatch). `probe` is injectable so the preference order can be
|
||||
* unit-tested without spawning; it defaults to the real PATH probe. `deps` can
|
||||
* inject `{ npmMajor, pnpmMajor }` for tests.
|
||||
*/
|
||||
function resolveInvocationMode(probe = resolveOnPath, deps = {}) {
|
||||
const forced = process.env.GITNEXUS_INVOCATION?.trim().toLowerCase();
|
||||
if (forced === 'gitnexus' || forced === 'pnpm' || forced === 'npx') {
|
||||
return forced;
|
||||
}
|
||||
if (probe('gitnexus', true)) return 'gitnexus';
|
||||
|
||||
const npmMajor = getNpmMajorVersion(deps);
|
||||
// pnpm presence: prefer an explicit `pnpmPresent` flag (set by
|
||||
// formatAnalyzeCommand, which falls back to a PATH probe when the version is
|
||||
// unreadable) so a present-but-unparseable pnpm — slow probe, Corepack
|
||||
// banner — still selects pnpm instead of the npx crash path. Otherwise an
|
||||
// injected version (a successful `pnpm --version` proves presence)
|
||||
// short-circuits the `which pnpm` probe; failing both, fall back to PATH.
|
||||
const hasPnpm =
|
||||
'pnpmPresent' in deps
|
||||
? deps.pnpmPresent
|
||||
: 'pnpmMajor' in deps
|
||||
? deps.pnpmMajor !== null
|
||||
: Boolean(probe('pnpm'));
|
||||
|
||||
// npm 11+ npx install crash (#1939) — prefer pnpm dlx when available.
|
||||
if (hasPnpm && npmMajor !== null && npmMajor >= 11) return 'pnpm';
|
||||
// npm 10 and earlier: npx works; prefer it over pnpm dlx when npm is present.
|
||||
if (npmMajor !== null && npmMajor < 11) return 'npx';
|
||||
// npm absent or unreadable — use pnpm if present (with allow-build flags).
|
||||
if (hasPnpm) return 'pnpm';
|
||||
|
||||
return 'npx';
|
||||
}
|
||||
|
||||
function formatPnpmDlxCommand(gitnexusArgs, options = {}, deps = {}) {
|
||||
const flags = formatPnpmAllowBuildArgs(options, deps).join(' ');
|
||||
const prefix = flags ? `${flags} ` : '';
|
||||
return `pnpm ${prefix}dlx ${NPX_REF} ${gitnexusArgs}`;
|
||||
}
|
||||
|
||||
function formatAnalyzeCommand(options = {}, deps = {}) {
|
||||
const suffix = options.embeddings ? ' --embeddings' : '';
|
||||
// Keep the stale-index hook budget tight by querying each tool at most once.
|
||||
// The memoized `probe` is a spawn-free PATH scan (resolveOnPath) shared with
|
||||
// resolveInvocationMode, so `gitnexus` is scanned only once and no subprocess
|
||||
// is spawned for presence. pnpm's *version* is still captured by a single
|
||||
// `pnpm --version` (the allow-build gate needs the number), which also proves
|
||||
// presence; the memoized scan only re-checks pnpm when that version is
|
||||
// unreadable. Injected deps (tests) and forced/global modes skip the pnpm probe.
|
||||
const cache = new Map();
|
||||
const probe = (command, gitnexusWrapper) => {
|
||||
const key = `${command}:${gitnexusWrapper ? 1 : 0}`;
|
||||
if (!cache.has(key)) cache.set(key, resolveOnPath(command, gitnexusWrapper));
|
||||
return cache.get(key);
|
||||
};
|
||||
let resolved = deps;
|
||||
if (!('pnpmMajor' in deps)) {
|
||||
const forced = process.env.GITNEXUS_INVOCATION?.trim().toLowerCase();
|
||||
// pnpm is only consulted when no non-pnpm mode is already certain: forced
|
||||
// gitnexus/npx never use pnpm, and a present global gitnexus wins outright.
|
||||
const mightUsePnpm = forced === 'pnpm' || (forced !== 'gitnexus' && forced !== 'npx');
|
||||
if (mightUsePnpm && (forced === 'pnpm' || !probe('gitnexus', true))) {
|
||||
const { major, minor } = probeVersion('pnpm');
|
||||
// Carry presence separately from version: when the version probe fails
|
||||
// (timeout, Corepack banner) but pnpm is on PATH, still treat it as
|
||||
// present so mode resolution picks pnpm over the npx crash path. The
|
||||
// PATH probe is memoized and only runs when the version is unreadable.
|
||||
const pnpmPresent = major !== null || Boolean(probe('pnpm'));
|
||||
resolved = { ...deps, pnpmMajor: major, pnpmMinor: minor, pnpmPresent };
|
||||
}
|
||||
}
|
||||
const mode = resolveInvocationMode(probe, resolved);
|
||||
if (mode === 'gitnexus') return `gitnexus analyze${suffix}`;
|
||||
if (mode === 'pnpm') return `${formatPnpmDlxCommand(`analyze${suffix}`, options, resolved)}`;
|
||||
return `npx ${NPX_REF} analyze${suffix}`;
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve `mode` into a concrete { program, args } pair for a set of gitnexus
|
||||
* subcommand arguments. Shared by the direct-exec entrypoint below; pure (no
|
||||
* spawn) so it is unit-testable. `--embeddings` widens the pnpm allow-build set.
|
||||
*/
|
||||
function buildRunnerArgv(mode, gitnexusArgs, deps = {}) {
|
||||
// Match both the space form (`--embeddings`) and the equals form
|
||||
// (`--embeddings=5000`) Commander accepts, so the pnpm allow-build set still
|
||||
// widens to onnxruntime-node when a user hand-types the equals form.
|
||||
const embeddings = gitnexusArgs.some(
|
||||
(a) => a === '--embeddings' || a.startsWith('--embeddings='),
|
||||
);
|
||||
if (mode === 'gitnexus') return { program: 'gitnexus', args: [...gitnexusArgs] };
|
||||
if (mode === 'pnpm') {
|
||||
return {
|
||||
program: 'pnpm',
|
||||
args: [...formatPnpmAllowBuildArgs({ embeddings }, deps), 'dlx', NPX_REF, ...gitnexusArgs],
|
||||
};
|
||||
}
|
||||
return { program: 'npx', args: [NPX_REF, ...gitnexusArgs] };
|
||||
}
|
||||
|
||||
module.exports = {
|
||||
formatAnalyzeCommand,
|
||||
formatDocumentationDlxCommand,
|
||||
formatPnpmAllowBuildArgs,
|
||||
formatPnpmDlxCommand,
|
||||
resolveInvocationMode,
|
||||
buildRunnerArgv,
|
||||
resolveOnPath,
|
||||
getNpmMajorVersion,
|
||||
NPX_REF,
|
||||
PNPM_ALLOW_BUILD_BASE,
|
||||
};
|
||||
|
||||
// Direct-exec entrypoint (#1945): `node run.cjs <gitnexus args…>` resolves the
|
||||
// best available runner (global `gitnexus` → `pnpm dlx` → `npx`) at call time and
|
||||
// runs it, inheriting stdio and propagating the child's exit code. This lets the
|
||||
// committed skills and generated AGENTS.md/CLAUDE.md reference ONE stable,
|
||||
// CLI-neutral command without baking in a package-manager assumption. `gitnexus
|
||||
// analyze` drops a copy of this file at `.gitnexus/run.cjs`. Skipped on require()
|
||||
// (the CLI and tests reuse the exports above), so it runs only when invoked as a
|
||||
// script.
|
||||
if (require.main === module) {
|
||||
const gitnexusArgs = process.argv.slice(2);
|
||||
const { program, args } = buildRunnerArgv(resolveInvocationMode(), gitnexusArgs);
|
||||
try {
|
||||
execFileSync(program, args, {
|
||||
stdio: 'inherit',
|
||||
windowsHide: true,
|
||||
// On Windows, `npx`/`pnpm`/`gitnexus` resolve to `.cmd`/`.ps1`/`.exe`
|
||||
// shims (npm, Volta, Corepack, scoop). execFileSync does not do PATHEXT
|
||||
// resolution and Node refuses to spawn `.cmd`/`.bat` without a shell
|
||||
// (CVE-2024-27980), so a bare program name ENOENTs. A shell lets the OS
|
||||
// resolve the shim; POSIX needs no shell (direct PATH lookup works).
|
||||
shell: process.platform === 'win32',
|
||||
});
|
||||
} catch (err) {
|
||||
// Make spawn failures (resolved program absent from PATH) self-explanatory
|
||||
// instead of a silent exit 1, then propagate the runner's own exit code.
|
||||
if (typeof err.status !== 'number') {
|
||||
process.stderr.write(`gitnexus runner: could not launch \`${program}\` — ${err.message}\n`);
|
||||
}
|
||||
process.exit(typeof err.status === 'number' ? err.status : 1);
|
||||
}
|
||||
}
|
||||
Generated
+4
-3
@@ -28,6 +28,7 @@
|
||||
"jsonc-parser": "^3.3.1",
|
||||
"lru-cache": "^11.0.0",
|
||||
"mnemonist": "^0.40.3",
|
||||
"node-addon-api": "8.8.0",
|
||||
"onnxruntime-node": "^1.24.0",
|
||||
"pandemonium": "^2.4.0",
|
||||
"pino": "^10.3.1",
|
||||
@@ -3799,9 +3800,9 @@
|
||||
}
|
||||
},
|
||||
"node_modules/node-addon-api": {
|
||||
"version": "8.7.0",
|
||||
"resolved": "https://registry.npmjs.org/node-addon-api/-/node-addon-api-8.7.0.tgz",
|
||||
"integrity": "sha512-9MdFxmkKaOYVTV+XVRG8ArDwwQ77XIgIPyKASB1k3JPq3M8fGQQQE3YpMOrKm6g//Ktx8ivZr8xo1Qmtqub+GA==",
|
||||
"version": "8.8.0",
|
||||
"resolved": "https://registry.npmjs.org/node-addon-api/-/node-addon-api-8.8.0.tgz",
|
||||
"integrity": "sha512-c5Ko1fZJIJmzhFIkhRN76WTq+fC6tWnGy9CXA0fA+XygsWZmEwG8vmbkNqxMyoaa0Tin4djul49NzdVcJJcjeA==",
|
||||
"license": "MIT",
|
||||
"engines": {
|
||||
"node": "^18 || ^20 || >= 21"
|
||||
|
||||
@@ -77,7 +77,7 @@
|
||||
"pandemonium": "^2.4.0",
|
||||
"pino": "^10.3.1",
|
||||
"pino-pretty": "^13.1.3",
|
||||
"tree-sitter": "^0.21.1",
|
||||
"tree-sitter": "0.21.1",
|
||||
"tree-sitter-c": "0.21.4",
|
||||
"tree-sitter-c-sharp": "0.23.1",
|
||||
"tree-sitter-cpp": "0.23.2",
|
||||
|
||||
@@ -30,6 +30,7 @@ const PLATFORM_LOGIC = [
|
||||
'test/unit/setup-jsonc.test.ts',
|
||||
'test/unit/setup-codex.test.ts',
|
||||
'test/unit/setup-antigravity.test.ts',
|
||||
'test/unit/resolve-invocation.test.ts',
|
||||
'test/unit/platform-capabilities.test.ts',
|
||||
'test/unit/worker-pool-windows-quarantine.test.ts',
|
||||
'test/unit/lbug-pool-win-fts-probe.test.ts',
|
||||
@@ -84,6 +85,7 @@ const SPAWN_CLI = [
|
||||
'test/integration/setup-antigravity.test.ts',
|
||||
'test/integration/antigravity-hook-e2e.test.ts',
|
||||
'test/unit/local-cli-subprocess.test.ts',
|
||||
'test/unit/runner-exec-tail.test.ts',
|
||||
];
|
||||
|
||||
// Worker threads tests — exercise real worker_threads which have
|
||||
@@ -101,6 +103,7 @@ const NATIVE_ADDON_SMOKE = [
|
||||
'test/integration/pipeline.test.ts',
|
||||
'test/integration/pipeline-graph-golden.test.ts',
|
||||
'test/unit/parser-loader.test.ts',
|
||||
'test/unit/parser-loader-abi.test.ts',
|
||||
];
|
||||
|
||||
// Filesystem behavior tests — exercise operations that vary across
|
||||
|
||||
@@ -5,14 +5,16 @@ description: "Use when the user needs to run GitNexus CLI commands like analyze/
|
||||
|
||||
# GitNexus CLI Commands
|
||||
|
||||
All commands work via `npx` — no global install required.
|
||||
Commands below use `node .gitnexus/run.cjs <command>` — the project-local runner `gitnexus analyze` drops next to the index. It auto-selects an available runner at call time (global `gitnexus`, else `pnpm dlx`, else `npx`), so no package-manager assumption and no global install is required.
|
||||
|
||||
> **Not analyzed yet, or `node .gitnexus/run.cjs` reports `Cannot find module`** (the gitignored runner is absent — e.g. a fresh clone or `git clean`)? (Re)generate it with `npx gitnexus analyze` from the project root. On **npm 11.x**, if `npx` crashes during install (`node.target is null`), install once with `npm i -g gitnexus` (then `gitnexus analyze`) or use `pnpm --allow-build=@ladybugdb/core --allow-build=gitnexus --allow-build=tree-sitter dlx gitnexus@latest analyze`. See [#1939](https://github.com/abhigyanpatwari/GitNexus/issues/1939).
|
||||
|
||||
## Commands
|
||||
|
||||
### analyze — Build or refresh the index
|
||||
|
||||
```bash
|
||||
npx gitnexus analyze
|
||||
node .gitnexus/run.cjs analyze
|
||||
```
|
||||
|
||||
Run from the project root. This parses all source files, builds the knowledge graph, writes it to `.gitnexus/`, and generates CLAUDE.md / AGENTS.md context files.
|
||||
@@ -28,7 +30,7 @@ Run from the project root. This parses all source files, builds the knowledge gr
|
||||
### status — Check index freshness
|
||||
|
||||
```bash
|
||||
npx gitnexus status
|
||||
node .gitnexus/run.cjs status
|
||||
```
|
||||
|
||||
Shows whether the current repo has a GitNexus index, when it was last updated, and symbol/relationship counts. Use this to check if re-indexing is needed.
|
||||
@@ -36,7 +38,7 @@ Shows whether the current repo has a GitNexus index, when it was last updated, a
|
||||
### clean — Delete the index
|
||||
|
||||
```bash
|
||||
npx gitnexus clean
|
||||
node .gitnexus/run.cjs clean
|
||||
```
|
||||
|
||||
Deletes the `.gitnexus/` directory and unregisters the repo from the global registry. Use before re-indexing if the index is corrupt or after removing GitNexus from a project.
|
||||
@@ -49,7 +51,7 @@ Deletes the `.gitnexus/` directory and unregisters the repo from the global regi
|
||||
### wiki — Generate documentation from the graph
|
||||
|
||||
```bash
|
||||
npx gitnexus wiki
|
||||
node .gitnexus/run.cjs wiki
|
||||
```
|
||||
|
||||
Generates repository documentation from the knowledge graph using an LLM. Requires an API key (saved to `~/.gitnexus/config.json` on first use).
|
||||
@@ -66,7 +68,7 @@ Generates repository documentation from the knowledge graph using an LLM. Requir
|
||||
### list — Show all indexed repos
|
||||
|
||||
```bash
|
||||
npx gitnexus list
|
||||
node .gitnexus/run.cjs list
|
||||
```
|
||||
|
||||
Lists all repositories registered in `~/.gitnexus/registry.json`. The MCP `list_repos` tool provides the same information.
|
||||
|
||||
@@ -22,7 +22,7 @@ description: "Use when the user is debugging a bug, tracing an error, or asking
|
||||
4. gitnexus_cypher({query: "MATCH path..."}) → Custom traces if needed
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
||||
## Checklist
|
||||
|
||||
|
||||
@@ -23,7 +23,7 @@ description: "Use when the user asks how code works, wants to understand archite
|
||||
5. READ gitnexus://repo/{name}/process/{name} → Trace full execution flow
|
||||
```
|
||||
|
||||
> If step 2 says "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
> If step 2 says "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
||||
## Checklist
|
||||
|
||||
|
||||
@@ -15,7 +15,7 @@ For any task involving code understanding, debugging, impact analysis, or refact
|
||||
2. **Match your task to a skill below** and **read that skill file**
|
||||
3. **Follow the skill's workflow and checklist**
|
||||
|
||||
> If step 1 warns the index is stale, run `npx gitnexus analyze` in the terminal first.
|
||||
> If step 1 warns the index is stale, run `node .gitnexus/run.cjs analyze` in the terminal first.
|
||||
|
||||
## Skills
|
||||
|
||||
|
||||
@@ -23,7 +23,7 @@ description: "Use when the user wants to know what will break if they change som
|
||||
4. Assess risk and report to user
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
||||
## Checklist
|
||||
|
||||
|
||||
@@ -26,7 +26,7 @@ description: "Use when the user wants to review a pull request, understand what
|
||||
6. Summarize findings with risk assessment
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal before reviewing.
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal before reviewing.
|
||||
|
||||
## Checklist
|
||||
|
||||
|
||||
@@ -22,7 +22,7 @@ description: "Use when the user wants to rename, extract, split, move, or restru
|
||||
4. Plan update order: interfaces → implementations → callers → tests
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
||||
## Checklists
|
||||
|
||||
|
||||
@@ -89,13 +89,17 @@ async function findGroupsContainingRegistryName(registryName: string): Promise<s
|
||||
return hits;
|
||||
}
|
||||
|
||||
function generateGitNexusContent(
|
||||
export function generateGitNexusContent(
|
||||
projectName: string,
|
||||
stats: RepoStats,
|
||||
generatedSkills?: GeneratedSkillInfo[],
|
||||
groupNames?: string[],
|
||||
noStats?: boolean,
|
||||
skipSkills?: boolean,
|
||||
// Project-relative path to the runner `gitnexus analyze` drops next to the
|
||||
// index (#1945). Referenced by docs so a single CLI-neutral command resolves
|
||||
// the available runner (global `gitnexus` → `pnpm dlx` → `npx`) at call time.
|
||||
runnerPath: string = '.gitnexus/run.cjs',
|
||||
): string {
|
||||
const generatedRows =
|
||||
generatedSkills && generatedSkills.length > 0
|
||||
@@ -127,13 +131,22 @@ function generateGitNexusContent(
|
||||
|------|---------------------|
|
||||
${tableBody}`
|
||||
: '';
|
||||
// Docs reference the project-local runner `gitnexus analyze` writes (#1945):
|
||||
// a single, CLI-neutral, machine-independent command (no per-machine churn,
|
||||
// #1706) that auto-selects the available runner at call time. Kept terse to
|
||||
// stay under the CLAUDE.md block token budget (#856); the cli skill carries the
|
||||
// full bootstrap + npm-11 fallback (`node.target is null` npx install crash).
|
||||
const runner = `node ${runnerPath}`;
|
||||
const bootstrapNote =
|
||||
`No \`${runnerPath}\` yet? \`npx gitnexus analyze\` ` +
|
||||
'(npm 11 crash → `npm i -g gitnexus`; #1939).';
|
||||
|
||||
return `${GITNEXUS_START_MARKER}
|
||||
# GitNexus — Code Intelligence
|
||||
|
||||
This project is indexed by GitNexus as **${projectName}**${noStats ? '' : ` (${stats.nodes || 0} symbols, ${stats.edges || 0} relationships, ${stats.processes || 0} execution flows)`}. Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
|
||||
|
||||
> If any GitNexus tool warns the index is stale, run \`npx gitnexus analyze\` in terminal first.
|
||||
> Index stale? Run \`${runner} analyze\` from the project root — it auto-selects an available runner. ${bootstrapNote}
|
||||
|
||||
## Always Do
|
||||
|
||||
@@ -163,7 +176,7 @@ ${
|
||||
groupNames && groupNames.length > 0
|
||||
? `## Cross-Repo Groups
|
||||
|
||||
This repository is listed under GitNexus **group(s): ${groupNames.join(', ')}** (see \`~/.gitnexus/groups/\`). For cross-repo analysis, use MCP tools \`impact\`, \`query\`, and \`context\` with \`repo\` set to \`@<groupName>\` or \`@<groupName>/<memberPath>\` (paths match keys in that group’s \`group.yaml\`). Use \`group_list\` / \`group_sync\` for membership and sync. From the terminal: \`npx gitnexus group list\`, \`npx gitnexus group sync <name>\`, \`npx gitnexus group impact <name> --target <symbol> --repo <group-path>\`.
|
||||
This repository is listed under GitNexus **group(s): ${groupNames.join(', ')}** (see \`~/.gitnexus/groups/\`). For cross-repo analysis, use MCP tools \`impact\`, \`query\`, and \`context\` with \`repo\` set to \`@<groupName>\` or \`@<groupName>/<memberPath>\` (paths match keys in that group’s \`group.yaml\`). Use \`group_list\` / \`group_sync\` for membership and sync. From the project root: \`${runner} group list\`, \`${runner} group sync <name>\`, \`${runner} group impact <name> --target <symbol> --repo <group-path>\` (the \`${runnerPath}\` path is repo-root-relative).
|
||||
|
||||
`
|
||||
: ''
|
||||
@@ -379,13 +392,36 @@ Use GitNexus tools to accomplish this task.
|
||||
*/
|
||||
export async function generateAIContextFiles(
|
||||
repoPath: string,
|
||||
_storagePath: string,
|
||||
storagePath: string,
|
||||
projectName: string,
|
||||
stats: RepoStats,
|
||||
generatedSkills?: GeneratedSkillInfo[],
|
||||
options?: AIContextOptions,
|
||||
): Promise<{ files: string[] }> {
|
||||
const groupNames = await findGroupsContainingRegistryName(projectName);
|
||||
|
||||
// Drop a project-local runner next to the index (#1945) so the generated docs
|
||||
// can reference one CLI-neutral command that resolves the available runner at
|
||||
// call time. It is a copy of the canonical self-contained resolver, which the
|
||||
// CLI and hooks already share; failure to copy is non-fatal (docs carry a
|
||||
// bootstrap fallback). `runnerPath` is project-relative with POSIX separators
|
||||
// so the emitted command is identical across platforms.
|
||||
const runnerPath = path.relative(repoPath, path.join(storagePath, 'run.cjs')).replace(/\\/g, '/');
|
||||
try {
|
||||
const runnerSrc = path.join(
|
||||
__dirname,
|
||||
'..',
|
||||
'..',
|
||||
'hooks',
|
||||
'claude',
|
||||
'resolve-analyze-cmd.cjs',
|
||||
);
|
||||
await fs.mkdir(storagePath, { recursive: true });
|
||||
await fs.copyFile(runnerSrc, path.join(storagePath, 'run.cjs'));
|
||||
} catch (err) {
|
||||
logger.warn(`Could not write GitNexus runner to ${runnerPath}: ${String(err)}`);
|
||||
}
|
||||
|
||||
const content = generateGitNexusContent(
|
||||
projectName,
|
||||
stats,
|
||||
@@ -393,6 +429,7 @@ export async function generateAIContextFiles(
|
||||
groupNames,
|
||||
options?.noStats,
|
||||
options?.skipSkills,
|
||||
runnerPath,
|
||||
);
|
||||
const createdFiles: string[] = [];
|
||||
|
||||
|
||||
@@ -35,6 +35,7 @@ import fs from 'fs/promises';
|
||||
import { cliError } from './cli-message.js';
|
||||
import { formatElapsed } from './format-elapsed.js';
|
||||
import { isHfDownloadFailure } from '../core/embeddings/hf-env.js';
|
||||
import { warnIfNpm11NpxRisk } from './resolve-invocation.js';
|
||||
|
||||
// Capture stderr.write at module load BEFORE anything (LadybugDB native
|
||||
// init, progress bar, console redirection) can monkey-patch it. The
|
||||
@@ -617,6 +618,11 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
|
||||
// a stack trace and a non-zero exit code instead of a silent exit 0.
|
||||
installFatalHandlers();
|
||||
|
||||
// npm-11 npx-crash nudge (#1939). Runs here, after the heap re-exec guard,
|
||||
// so it fires once in the working process and never on the lazy-startup path
|
||||
// of other commands (e.g. `gitnexus mcp`).
|
||||
warnIfNpm11NpxRisk();
|
||||
|
||||
// Snapshot the GITNEXUS_* env vars that the impl writes for downstream
|
||||
// consumption, so they don't leak across `analyzeCommand` invocations in
|
||||
// programmatic callers (tests, long-running hosts). `process.exit(0)` on
|
||||
@@ -1096,6 +1102,17 @@ const analyzeCommandImpl = async (inputPath?: string, options?: AnalyzeOptions):
|
||||
);
|
||||
console.log(` ${repoPath}`);
|
||||
|
||||
// Persistent (non-scrolling) warning when FTS indexing was skipped — the
|
||||
// progress-bar log() that fired mid-run has already scrolled away, so the
|
||||
// degraded-search state must also appear in the final summary (#1161).
|
||||
if (result.ftsSkipped) {
|
||||
console.log(
|
||||
`\n Warning: full-text/BM25 search is disabled — the LadybugDB FTS extension was unavailable.\n` +
|
||||
` Install it once with network access (GITNEXUS_LBUG_EXTENSION_INSTALL=auto) then rerun, or\n` +
|
||||
` run \`gitnexus analyze --repair-fts\` when connected. Run \`gitnexus doctor\` for details.`,
|
||||
);
|
||||
}
|
||||
|
||||
try {
|
||||
await fs.access(getGlobalRegistryPath());
|
||||
} catch {
|
||||
|
||||
@@ -2,6 +2,7 @@ import { getRuntimeCapabilities, getRuntimeFingerprint } from '../core/platform/
|
||||
import { resolveEmbeddingConfig } from '../core/embeddings/config.js';
|
||||
import { isHttpMode } from '../core/embeddings/http-client.js';
|
||||
import { checkLbugNative } from '../core/lbug/native-check.js';
|
||||
import { getExtensionInstallPolicy } from '../core/lbug/extension-loader.js';
|
||||
import { t } from './i18n/index.js';
|
||||
|
||||
function isCombiningMark(codePoint: number): boolean {
|
||||
@@ -74,6 +75,17 @@ export const doctorCommand = async () => {
|
||||
console.log(` ${label('doctor.labels.fullTextSearch', 18)}${capabilities.fts}`);
|
||||
console.log(` ${label('doctor.labels.vectorIndex', 18)}${capabilities.vector}`);
|
||||
console.log(` ${label('doctor.labels.semanticMode', 18)}${capabilities.semanticMode}`);
|
||||
// Surface the optional-extension install policy so offline users can see
|
||||
// whether analyze/query will reach the network (extension.ladybugdb.com).
|
||||
// Literal label (like the 'native' line) to avoid adding i18n keys.
|
||||
const installPolicy = getExtensionInstallPolicy();
|
||||
const policyHint =
|
||||
installPolicy === 'load-only'
|
||||
? ' (offline; load only, no network install)'
|
||||
: installPolicy === 'never'
|
||||
? ' (optional extensions disabled)'
|
||||
: ' (installs missing extensions over network)';
|
||||
console.log(` ${padDisplayEnd('Ext install:', 18)}${installPolicy}${policyHint}`);
|
||||
console.log(
|
||||
` ${label('doctor.labels.exactScanLimit', 18)}${t('doctor.chunks', { count: capabilities.exactScanLimit })}`,
|
||||
);
|
||||
|
||||
@@ -101,6 +101,9 @@ const OPTION_DESCRIPTION_KEYS = {
|
||||
'context|--content': 'help.option.content',
|
||||
'impact|-d, --direction <dir>': 'help.option.impact.direction',
|
||||
'impact|-r, --repo <name>': 'help.option.repo.target',
|
||||
'impact|-u, --uid <uid>': 'help.option.context.uid',
|
||||
'impact|-f, --file <path>': 'help.option.context.file',
|
||||
'impact|--kind <kind>': 'help.option.impact.kind',
|
||||
'impact|--depth <n>': 'help.option.impact.depth',
|
||||
'impact|--include-tests': 'help.option.impact.includeTests',
|
||||
'impact|--limit <n>': 'help.option.impact.limit',
|
||||
|
||||
@@ -43,8 +43,11 @@ export const en = {
|
||||
'tool.noIndexed': 'GitNexus: No indexed repositories found. Run: gitnexus analyze',
|
||||
'tool.usage.query': 'Usage: gitnexus query <search_query>',
|
||||
'tool.usage.context': 'Usage: gitnexus context <symbol_name> [--uid <uid>] [--file <path>]',
|
||||
'tool.usage.impact': 'Usage: gitnexus impact <symbol_name> [--direction upstream|downstream]',
|
||||
'tool.usage.impact':
|
||||
'Usage: gitnexus impact <symbol_name> [--uid <uid>] [--file <path>] [--kind <kind>] [--direction upstream|downstream]',
|
||||
'tool.usage.cypher': 'Usage: gitnexus cypher <cypher_query>',
|
||||
'tool.warn.unknownKind':
|
||||
"--kind '{{kind}}' is not a known symbol kind (e.g. Function, Class, Method); it will not narrow the result.",
|
||||
'tool.detectChanges.noChanges': 'No changes detected.',
|
||||
'tool.detectChanges.changesSummary': 'Changes: {{files}} files, {{symbols}} symbols',
|
||||
'tool.detectChanges.affectedProcesses': 'Affected processes: {{count}}',
|
||||
@@ -213,6 +216,8 @@ export const en = {
|
||||
'help.option.repo.target': 'Target repository',
|
||||
'help.option.context.uid': 'Direct symbol UID (zero-ambiguity lookup)',
|
||||
'help.option.context.file': 'File path to disambiguate common names',
|
||||
'help.option.impact.kind':
|
||||
'Kind filter to disambiguate common names (e.g. Function, Class, Method)',
|
||||
'help.option.impact.direction': 'upstream (dependants) or downstream (dependencies)',
|
||||
'help.option.impact.depth': 'Max relationship depth (default: 3)',
|
||||
'help.option.impact.includeTests': 'Include test files in results',
|
||||
|
||||
@@ -47,8 +47,11 @@ export const zhCN = {
|
||||
'tool.noIndexed': 'GitNexus:未找到已索引仓库。请运行:gitnexus analyze',
|
||||
'tool.usage.query': '用法:gitnexus query <搜索词>',
|
||||
'tool.usage.context': '用法:gitnexus context <符号名> [--uid <uid>] [--file <路径>]',
|
||||
'tool.usage.impact': '用法:gitnexus impact <符号名> [--direction upstream|downstream]',
|
||||
'tool.usage.impact':
|
||||
'用法:gitnexus impact <符号名> [--uid <uid>] [--file <路径>] [--kind <类型>] [--direction upstream|downstream]',
|
||||
'tool.usage.cypher': '用法:gitnexus cypher <Cypher 查询>',
|
||||
'tool.warn.unknownKind':
|
||||
"--kind '{{kind}}' 不是已知的符号类型(如 Function、Class、Method),不会用于缩小结果范围。",
|
||||
'tool.detectChanges.noChanges': '未检测到变更。',
|
||||
'tool.detectChanges.changesSummary': '变更:{{files}} 个文件,{{symbols}} 个符号',
|
||||
'tool.detectChanges.affectedProcesses': '受影响流程:{{count}}',
|
||||
@@ -199,6 +202,7 @@ export const zhCN = {
|
||||
'help.option.repo.target': '目标仓库',
|
||||
'help.option.context.uid': '直接符号 UID(零歧义查找)',
|
||||
'help.option.context.file': '用于消除常见名称歧义的文件路径',
|
||||
'help.option.impact.kind': '用于消除常见名称歧义的类型过滤(如 Function、Class、Method)',
|
||||
'help.option.impact.direction': 'upstream(依赖它的项)或 downstream(它依赖的项)',
|
||||
'help.option.impact.depth': '最大关系遍历深度(默认:3)',
|
||||
'help.option.impact.includeTests': '在结果中包含测试文件',
|
||||
|
||||
@@ -219,10 +219,16 @@ program
|
||||
.action(createLbugLazyAction(() => import('./tool.js'), 'contextCommand'));
|
||||
|
||||
program
|
||||
.command('impact <target>')
|
||||
.command('impact [target]')
|
||||
.description('Blast radius analysis: what breaks if you change a symbol')
|
||||
.option('-d, --direction <dir>', 'upstream (dependants) or downstream (dependencies)', 'upstream')
|
||||
.option('-r, --repo <name>', 'Target repository')
|
||||
.option('-u, --uid <uid>', 'Direct symbol UID (zero-ambiguity lookup)')
|
||||
.option('-f, --file <path>', 'File path to disambiguate common names')
|
||||
.option(
|
||||
'--kind <kind>',
|
||||
'Kind filter to disambiguate common names (e.g. Function, Class, Method)',
|
||||
)
|
||||
.option('--depth <n>', 'Max relationship depth (default: 3)')
|
||||
.option('--include-tests', 'Include test files in results')
|
||||
.option('--limit <n>', 'Max symbols per depth level (default: 100)')
|
||||
|
||||
@@ -0,0 +1,98 @@
|
||||
/**
|
||||
* npm 11.x npx-install-crash nudge for the `analyze` command (#1939).
|
||||
*
|
||||
* The gitnexus/pnpm/npx selection itself lives in the canonical hook helper
|
||||
* (hooks/claude/resolve-analyze-cmd.cjs) — self-contained CJS because the copied
|
||||
* hook runtime cannot import from the package. We reuse it here via createRequire
|
||||
* instead of re-implementing it, so there is one source of truth for the
|
||||
* invocation decision. This module adds only the npm-version probe and the
|
||||
* warning, which are CLI-only. The relative path resolves identically from
|
||||
* src/cli/ (tsx, vitest) and dist/cli/ (shipped), since both sit one level under
|
||||
* the package root and `hooks/` is published.
|
||||
*/
|
||||
|
||||
import { execFileSync } from 'node:child_process';
|
||||
import { createRequire } from 'node:module';
|
||||
|
||||
type InvocationMode = 'gitnexus' | 'pnpm' | 'npx';
|
||||
|
||||
interface InvocationResolver {
|
||||
// `probe` is injectable in the cjs (defaults to the real PATH probe) so the
|
||||
// preference order is unit-testable without spawning; the CLI calls it with
|
||||
// no argument.
|
||||
resolveInvocationMode: (
|
||||
probe?: (command: string, gitnexusWrapper?: boolean) => string | null,
|
||||
) => InvocationMode;
|
||||
formatDocumentationDlxCommand: (
|
||||
gitnexusArgs: string,
|
||||
options?: { embeddings?: boolean },
|
||||
) => string;
|
||||
NPX_REF: string;
|
||||
}
|
||||
|
||||
const { resolveInvocationMode, formatDocumentationDlxCommand, NPX_REF } = createRequire(
|
||||
import.meta.url,
|
||||
// `require()` returns `any`; go through `unknown` so the cast reads as an
|
||||
// explicit narrowing to the subset this module uses, not a claim that the
|
||||
// cjs's full export shape is known here. The drift guard below verifies it.
|
||||
)('../../hooks/claude/resolve-analyze-cmd.cjs') as unknown as InvocationResolver;
|
||||
|
||||
// Fail loud at module load if the canonical cjs export shape drifts (e.g. a
|
||||
// renamed export), rather than as a late TypeError inside warnIfNpm11NpxRisk.
|
||||
if (
|
||||
typeof resolveInvocationMode !== 'function' ||
|
||||
typeof formatDocumentationDlxCommand !== 'function' ||
|
||||
typeof NPX_REF !== 'string'
|
||||
) {
|
||||
throw new Error(
|
||||
'resolve-analyze-cmd.cjs must export resolveInvocationMode (function), formatDocumentationDlxCommand (function), and NPX_REF (string)',
|
||||
);
|
||||
}
|
||||
|
||||
export { NPX_REF };
|
||||
|
||||
// Re-implemented here (rather than reusing the cjs export) so vitest's
|
||||
// `vi.mock('node:child_process')` intercepts it — the cjs uses bare
|
||||
// `require('child_process')`, which the mock cannot reach. Timeout matches the
|
||||
// cjs PROBE_TIMEOUT_MS (1s) so this CLI probe shares the same hook-budget cap;
|
||||
// `npm --version` is a sub-second local call.
|
||||
export function getNpmMajorVersion(): number | null {
|
||||
try {
|
||||
const output = execFileSync('npm', ['--version'], {
|
||||
encoding: 'utf-8',
|
||||
timeout: 1000,
|
||||
stdio: ['ignore', 'pipe', 'ignore'],
|
||||
windowsHide: true,
|
||||
// Windows `npm` is a `.cmd` shim; without a shell execFileSync ENOENTs
|
||||
// (CVE-2024-27980) and the npm-11 npx-crash warning below would never
|
||||
// fire on Windows. Mirrors probeVersion in resolve-analyze-cmd.cjs.
|
||||
shell: process.platform === 'win32',
|
||||
});
|
||||
// Read the first version-shaped line so a Corepack/update banner on stdout
|
||||
// doesn't defeat the parse (mirrors the cjs probeVersion hardening).
|
||||
const major = output
|
||||
.split('\n')
|
||||
.map((l) => l.trim())
|
||||
.find((l) => /^v?\d+\./.test(l))
|
||||
?.match(/^v?(\d+)\./);
|
||||
return major ? Number(major[1]) : null;
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* One-line stderr nudge when an npm 11+ user is on the npx install path (#1939).
|
||||
* Skipped when a global `gitnexus` or `pnpm` is already preferred, so it never
|
||||
* nags users who are not exposed to the npx/arborist crash.
|
||||
*/
|
||||
export function warnIfNpm11NpxRisk(): void {
|
||||
if (resolveInvocationMode() !== 'npx') return;
|
||||
const major = getNpmMajorVersion();
|
||||
if (major === null || major < 11) return;
|
||||
process.stderr.write(
|
||||
`Warning: npm ${major}.x can crash while installing gitnexus via npx ` +
|
||||
`(npm/arborist "node.target is null"). Prefer: ${formatDocumentationDlxCommand('analyze')} ` +
|
||||
`or npm install -g ${NPX_REF}. See https://github.com/abhigyanpatwari/GitNexus/issues/1939\n`,
|
||||
);
|
||||
}
|
||||
+127
-49
@@ -33,7 +33,44 @@ if (typeof _pkg.version !== 'string' || !_pkg.version) {
|
||||
'gitnexus/package.json#version is missing or not a string — cannot generate MCP fallback config.',
|
||||
);
|
||||
}
|
||||
const NPX_REF = `gitnexus@${_pkg.version}`;
|
||||
// Version-pinned ref for the persisted MCP entry — deliberately distinct from
|
||||
// the cjs's exported `gitnexus@latest` hint ref (resolve-analyze-cmd.cjs); the
|
||||
// two are not unified (see the comment above and that file's MCP_PINNED_REF).
|
||||
const MCP_PINNED_REF = `gitnexus@${_pkg.version}`;
|
||||
|
||||
/**
|
||||
* Build the `command` string written into an editor's hook settings, which the
|
||||
* editor shell-evaluates. `hookPath` is already forward-slash-normalized.
|
||||
*
|
||||
* On POSIX, single-quote the path: a single-quoted shell string expands nothing,
|
||||
* so spaces and metacharacters ($, backtick, ;, |, &, newline, parens) in the
|
||||
* install path cannot run as commands. The only character needing escaping
|
||||
* inside single quotes is the single quote, via the standard `'\''` idiom
|
||||
* (close, literal-quote, reopen). The previous double-quoted `node "..."` form
|
||||
* left $/backtick live — a code-execution risk for an adversarial $HOME.
|
||||
*
|
||||
* On Windows, filenames cannot contain these POSIX metacharacters and the path
|
||||
* is forward-slashed, so keep the double-quoted form with backslash-then-quote
|
||||
* escaping (CodeQL js/incomplete-sanitization safe ordering).
|
||||
*/
|
||||
export function formatHookCommand(
|
||||
hookPath: string,
|
||||
isWindows = process.platform === 'win32',
|
||||
): string {
|
||||
if (isWindows) {
|
||||
const escaped = hookPath.replace(/\\/g, '\\\\').replace(/"/g, '\\"');
|
||||
return `node "${escaped}"`;
|
||||
}
|
||||
return `node '${hookPath.replace(/'/g, "'\\''")}'`;
|
||||
}
|
||||
|
||||
// The exact source line each hook adapter ships, rewritten at install time to
|
||||
// point cliPath at the installed CLI. Kept as a named constant so the install
|
||||
// patch and its drift guard reference one string — if the adapter source ever
|
||||
// changes this literal, the guard records an actionable error instead of
|
||||
// silently shipping a hook with an unresolved relative cliPath.
|
||||
const CLI_PATH_SOURCE_LITERAL =
|
||||
"let cliPath = path.resolve(__dirname, '..', '..', 'dist', 'cli', 'index.js');";
|
||||
|
||||
interface SetupResult {
|
||||
configured: string[];
|
||||
@@ -99,12 +136,12 @@ function getMcpEntry() {
|
||||
if (process.platform === 'win32') {
|
||||
return {
|
||||
command: 'cmd',
|
||||
args: ['/c', 'npx', '-y', NPX_REF, 'mcp'],
|
||||
args: ['/c', 'npx', '-y', MCP_PINNED_REF, 'mcp'],
|
||||
};
|
||||
}
|
||||
return {
|
||||
command: 'npx',
|
||||
args: ['-y', NPX_REF, 'mcp'],
|
||||
args: ['-y', MCP_PINNED_REF, 'mcp'],
|
||||
};
|
||||
}
|
||||
|
||||
@@ -120,9 +157,9 @@ function getOpenCodeMcpEntry() {
|
||||
}
|
||||
|
||||
if (process.platform === 'win32') {
|
||||
return { type: 'local', command: ['cmd', '/c', 'npx', '-y', NPX_REF, 'mcp'] };
|
||||
return { type: 'local', command: ['cmd', '/c', 'npx', '-y', MCP_PINNED_REF, 'mcp'] };
|
||||
}
|
||||
return { type: 'local', command: ['npx', '-y', NPX_REF, 'mcp'] };
|
||||
return { type: 'local', command: ['npx', '-y', MCP_PINNED_REF, 'mcp'] };
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -334,6 +371,48 @@ async function mergeHooksJsonc(
|
||||
return true;
|
||||
}
|
||||
|
||||
const HOOK_HELPERS = [
|
||||
'hook-lock.cjs',
|
||||
'hook-db-lock-probe.cjs',
|
||||
'win-rm-list-json.ps1',
|
||||
'resolve-analyze-cmd.cjs',
|
||||
] as const;
|
||||
|
||||
// win-rm-list-json.ps1 is best-effort: it is read (not require()'d) by
|
||||
// hook-db-lock-probe.cjs only on Windows, and that probe fails open when the
|
||||
// script is absent. Every other helper is top-level require()'d by the adapters,
|
||||
// so its absence crashes the installed hook — those are the ones a failed copy
|
||||
// must gate hook registration on (see copyHookHelpers' return value).
|
||||
const BEST_EFFORT_HOOK_HELPERS = new Set<string>(['win-rm-list-json.ps1']);
|
||||
|
||||
/**
|
||||
* Copy the shared hook helpers from `srcDir` into `destDir`. The adapters
|
||||
* top-level `require()` the `.cjs` helpers, so a missing required helper makes
|
||||
* the installed hook crash with MODULE_NOT_FOUND. A failed copy is recorded as a
|
||||
* setup error, and the names of any failed REQUIRED helpers are returned so the
|
||||
* caller can fail closed (skip hook registration) instead of registering a hook
|
||||
* that crashes at runtime. `win-rm-list-json.ps1` is best-effort — its absence is
|
||||
* recorded but does not gate registration. Both the Claude and Antigravity
|
||||
* install paths copy this same list from hooks/claude/ (the canonical source).
|
||||
*/
|
||||
export async function copyHookHelpers(
|
||||
srcDir: string,
|
||||
destDir: string,
|
||||
label: string,
|
||||
result: SetupResult,
|
||||
): Promise<string[]> {
|
||||
const failedRequired: string[] = [];
|
||||
for (const helper of HOOK_HELPERS) {
|
||||
try {
|
||||
await fs.copyFile(path.join(srcDir, helper), path.join(destDir, helper));
|
||||
} catch {
|
||||
result.errors.push(`${label}: failed to copy ${helper} — hook may crash at runtime`);
|
||||
if (!BEST_EFFORT_HOOK_HELPERS.has(helper)) failedRequired.push(helper);
|
||||
}
|
||||
}
|
||||
return failedRequired;
|
||||
}
|
||||
|
||||
/**
|
||||
* Install GitNexus hooks to ~/.claude/settings.json for Claude Code.
|
||||
* Merges hook config without overwriting existing hooks, preserving
|
||||
@@ -361,49 +440,44 @@ async function installClaudeCodeHooks(result: SetupResult): Promise<void> {
|
||||
const resolvedCli = path.join(__dirname, '..', 'cli', 'index.js');
|
||||
const normalizedCli = path.resolve(resolvedCli).replace(/\\/g, '/');
|
||||
const jsonCli = JSON.stringify(normalizedCli);
|
||||
content = content.replace(
|
||||
"let cliPath = path.resolve(__dirname, '..', '..', 'dist', 'cli', 'index.js');",
|
||||
`let cliPath = ${jsonCli};`,
|
||||
);
|
||||
if (!content.includes(CLI_PATH_SOURCE_LITERAL)) {
|
||||
result.errors.push(
|
||||
'Claude Code hooks: gitnexus-hook.cjs no longer contains the cliPath literal to patch — the installed hook may fail to resolve the CLI. Update CLI_PATH_SOURCE_LITERAL in setup.ts.',
|
||||
);
|
||||
}
|
||||
content = content.replace(CLI_PATH_SOURCE_LITERAL, `let cliPath = ${jsonCli};`);
|
||||
await fs.writeFile(dest, content, 'utf-8');
|
||||
} catch {
|
||||
// Script not found in source — skip
|
||||
}
|
||||
|
||||
// Fail closed: registering the hook without its adapter would crash on every
|
||||
// tool invocation. Mirrors the Antigravity adapter guard below (this path
|
||||
// previously registered regardless of whether the adapter wrote).
|
||||
try {
|
||||
await fs.copyFile(
|
||||
path.join(pluginHooksPath, 'hook-lock.cjs'),
|
||||
path.join(destHooksDir, 'hook-lock.cjs'),
|
||||
);
|
||||
await fs.access(dest);
|
||||
} catch {
|
||||
// Helper not found in source — skip
|
||||
result.errors.push(
|
||||
'Claude Code hooks: adapter script was not installed — skipping hook registration',
|
||||
);
|
||||
return;
|
||||
}
|
||||
|
||||
try {
|
||||
await fs.copyFile(
|
||||
path.join(pluginHooksPath, 'hook-db-lock-probe.cjs'),
|
||||
path.join(destHooksDir, 'hook-db-lock-probe.cjs'),
|
||||
const failedRequired = await copyHookHelpers(
|
||||
pluginHooksPath,
|
||||
destHooksDir,
|
||||
'Claude Code hooks',
|
||||
result,
|
||||
);
|
||||
if (failedRequired.length > 0) {
|
||||
result.errors.push(
|
||||
`Claude Code hooks: required helper(s) ${failedRequired.join(', ')} failed to copy — skipping hook registration`,
|
||||
);
|
||||
} catch {
|
||||
// Helper not found in source — skip
|
||||
}
|
||||
|
||||
try {
|
||||
await fs.copyFile(
|
||||
path.join(pluginHooksPath, 'win-rm-list-json.ps1'),
|
||||
path.join(destHooksDir, 'win-rm-list-json.ps1'),
|
||||
);
|
||||
} catch {
|
||||
// Helper not found in source — skip
|
||||
return;
|
||||
}
|
||||
|
||||
const hookPath = path.join(destHooksDir, 'gitnexus-hook.cjs').replace(/\\/g, '/');
|
||||
// Escape backslashes FIRST, then quotes (CodeQL js/incomplete-sanitization).
|
||||
// The previous shape `replace(/"/g, '\\"')` alone would let `path\with"quote`
|
||||
// become `path\with\"quote`, where the trailing `\` before `"` could
|
||||
// unescape the quote inside the surrounding double-quoted shell context.
|
||||
const escapedHookPath = hookPath.replace(/\\/g, '\\\\').replace(/"/g, '\\"');
|
||||
const hookCmd = `node "${escapedHookPath}"`;
|
||||
const hookCmd = formatHookCommand(hookPath);
|
||||
|
||||
// Check which hook events need entries (idempotent: skip if already registered)
|
||||
const parsed = await (async () => {
|
||||
@@ -566,10 +640,12 @@ async function installAntigravityHooks(result: SetupResult): Promise<void> {
|
||||
const resolvedCli = path.join(__dirname, '..', 'cli', 'index.js');
|
||||
const normalizedCli = path.resolve(resolvedCli).replace(/\\/g, '/');
|
||||
const jsonCli = JSON.stringify(normalizedCli);
|
||||
content = content.replace(
|
||||
"let cliPath = path.resolve(__dirname, '..', '..', 'dist', 'cli', 'index.js');",
|
||||
`let cliPath = ${jsonCli};`,
|
||||
);
|
||||
if (!content.includes(CLI_PATH_SOURCE_LITERAL)) {
|
||||
result.errors.push(
|
||||
'Antigravity hooks: gitnexus-antigravity-hook.cjs no longer contains the cliPath literal to patch — the installed hook may fail to resolve the CLI. Update CLI_PATH_SOURCE_LITERAL in setup.ts.',
|
||||
);
|
||||
}
|
||||
content = content.replace(CLI_PATH_SOURCE_LITERAL, `let cliPath = ${jsonCli};`);
|
||||
await fs.writeFile(adapterDest, content, 'utf-8');
|
||||
} catch {
|
||||
// Adapter not found in source — skip
|
||||
@@ -591,19 +667,21 @@ async function installAntigravityHooks(result: SetupResult): Promise<void> {
|
||||
// required by hook-db-lock-probe.cjs on Windows — without it, the MCP
|
||||
// server ownership probe silently fails open and the hook may contend
|
||||
// with the MCP server on the LadybugDB.
|
||||
for (const helper of ['hook-lock.cjs', 'hook-db-lock-probe.cjs', 'win-rm-list-json.ps1']) {
|
||||
try {
|
||||
await fs.copyFile(path.join(pluginClaudeDir, helper), path.join(destHooksDir, helper));
|
||||
} catch {
|
||||
result.errors.push(
|
||||
`Antigravity hooks: failed to copy ${helper} — hook may crash at runtime`,
|
||||
);
|
||||
}
|
||||
const failedRequired = await copyHookHelpers(
|
||||
pluginClaudeDir,
|
||||
destHooksDir,
|
||||
'Antigravity hooks',
|
||||
result,
|
||||
);
|
||||
if (failedRequired.length > 0) {
|
||||
result.errors.push(
|
||||
`Antigravity hooks: required helper(s) ${failedRequired.join(', ')} failed to copy — skipping hook registration`,
|
||||
);
|
||||
return;
|
||||
}
|
||||
|
||||
const hookPath = path.join(destHooksDir, 'gitnexus-antigravity-hook.cjs').replace(/\\/g, '/');
|
||||
const escapedHookPath = hookPath.replace(/\\/g, '\\\\').replace(/"/g, '\\"');
|
||||
const hookCmd = `node "${escapedHookPath}"`;
|
||||
const hookCmd = formatHookCommand(hookPath);
|
||||
|
||||
const parsed = await (async () => {
|
||||
try {
|
||||
|
||||
@@ -16,8 +16,8 @@
|
||||
*/
|
||||
|
||||
import { writeSync } from 'node:fs';
|
||||
import { LocalBackend } from '../mcp/local/local-backend.js';
|
||||
import { cliErrorKey } from './cli-message.js';
|
||||
import { LocalBackend, VALID_NODE_LABELS } from '../mcp/local/local-backend.js';
|
||||
import { cliErrorKey, cliWarnKey } from './cli-message.js';
|
||||
import { formatDetectChangesResult } from './detect-changes-format.js';
|
||||
|
||||
let _backend: LocalBackend | null = null;
|
||||
@@ -94,6 +94,11 @@ export async function contextCommand(
|
||||
content?: boolean;
|
||||
},
|
||||
): Promise<void> {
|
||||
// Reject a `--`-prefixed uid swallowed from a following flag (see impactCommand).
|
||||
if (options?.uid?.startsWith('--')) {
|
||||
cliErrorKey('tool.usage.context');
|
||||
process.exit(1);
|
||||
}
|
||||
if (!name?.trim() && !options?.uid) {
|
||||
cliErrorKey('tool.usage.context');
|
||||
process.exit(1);
|
||||
@@ -111,10 +116,13 @@ export async function contextCommand(
|
||||
}
|
||||
|
||||
export async function impactCommand(
|
||||
target: string,
|
||||
target?: string,
|
||||
options?: {
|
||||
direction?: string;
|
||||
repo?: string;
|
||||
uid?: string;
|
||||
file?: string;
|
||||
kind?: string;
|
||||
depth?: string;
|
||||
includeTests?: boolean;
|
||||
limit?: string;
|
||||
@@ -122,10 +130,25 @@ export async function impactCommand(
|
||||
summaryOnly?: boolean;
|
||||
},
|
||||
): Promise<void> {
|
||||
if (!target?.trim()) {
|
||||
// A `--`-prefixed uid means Commander swallowed a following flag as the uid
|
||||
// value (e.g. `impact --uid --file x` → uid === '--file'). Reject it rather
|
||||
// than forwarding a garbage uid that would silently resolve to not-found.
|
||||
if (options?.uid?.startsWith('--')) {
|
||||
cliErrorKey('tool.usage.impact');
|
||||
process.exit(1);
|
||||
}
|
||||
// Target is an optional positional: a uid alone is enough to resolve (parity
|
||||
// with `context [name]`). Only error when neither a target nor a uid is given.
|
||||
if (!target?.trim() && !options?.uid) {
|
||||
cliErrorKey('tool.usage.impact');
|
||||
process.exit(1);
|
||||
}
|
||||
// Soft-validate --kind: an unknown kind is a no-op hint (the backend scores
|
||||
// it but it matches nothing), so warn and proceed rather than rejecting —
|
||||
// parity with the lenient MCP surface and forward-compatible with new labels.
|
||||
if (options?.kind && !VALID_NODE_LABELS.has(options.kind)) {
|
||||
cliWarnKey('tool.warn.unknownKind', { kind: options.kind });
|
||||
}
|
||||
|
||||
try {
|
||||
const backend = await getBackend();
|
||||
@@ -134,7 +157,10 @@ export async function impactCommand(
|
||||
const parsedLimit = Number.isFinite(rawLimit) ? rawLimit : undefined;
|
||||
const parsedOffset = Number.isFinite(rawOffset) ? rawOffset : undefined;
|
||||
const result = await backend.callTool('impact', {
|
||||
target,
|
||||
target: target || undefined,
|
||||
target_uid: options?.uid,
|
||||
file_path: options?.file,
|
||||
kind: options?.kind,
|
||||
direction: options?.direction || 'upstream',
|
||||
maxDepth: options?.depth ? parseInt(options.depth, 10) : undefined,
|
||||
includeTests: options?.includeTests ?? false,
|
||||
|
||||
@@ -43,20 +43,38 @@ import {
|
||||
STALE_HASH_SENTINEL,
|
||||
} from '../lbug/schema.js';
|
||||
import { loadVectorExtension } from '../lbug/lbug-adapter.js';
|
||||
import type { ExtensionInstallPolicy } from '../lbug/extension-loader.js';
|
||||
import { getExactScanLimit } from '../platform/capabilities.js';
|
||||
import { logger } from '../logger.js';
|
||||
|
||||
const isDev = process.env.NODE_ENV === 'development';
|
||||
|
||||
const vectorUnavailableMessage =
|
||||
'VECTOR extension is unavailable for this LadybugDB runtime; semantic search will use exact scan when embeddings exist.';
|
||||
'VECTOR extension unavailable; semantic embeddings fall back to exact scan. ' +
|
||||
'To enable vector search, install it once with network access ' +
|
||||
'(GITNEXUS_LBUG_EXTENSION_INSTALL=auto), or pre-install it for offline use. ' +
|
||||
'Set GITNEXUS_LBUG_EXTENSION_INSTALL=never to skip installs and silence this.';
|
||||
|
||||
/**
|
||||
* Resolve the extension-install policy for the embedding WRITE path (analyze).
|
||||
*
|
||||
* Generating embeddings is an explicit opt-in to a feature that requires the
|
||||
* VECTOR extension, so when the operator has NOT pinned a policy we default to
|
||||
* `auto` (one bounded, out-of-process INSTALL) — matching the documented
|
||||
* "auto = default for analyze" intent in extension-loader.ts. An explicit
|
||||
* GITNEXUS_LBUG_EXTENSION_INSTALL=load-only|never|auto always wins, so an
|
||||
* offline or locked-down operator is never silently forced onto the network
|
||||
* (the #1153 regression caused by hard-coding `auto` here). Read on every call
|
||||
* (not memoized) so test env stubbing works.
|
||||
*/
|
||||
export const resolveEmbeddingInstallPolicy = (): ExtensionInstallPolicy => {
|
||||
const raw = process.env.GITNEXUS_LBUG_EXTENSION_INSTALL;
|
||||
if (raw === 'load-only' || raw === 'never' || raw === 'auto') return raw;
|
||||
return 'auto';
|
||||
};
|
||||
|
||||
const ensureVectorExtensionAvailable = async (): Promise<boolean> => {
|
||||
const vectorReady = await loadVectorExtension();
|
||||
if (!vectorReady) {
|
||||
return false;
|
||||
}
|
||||
return true;
|
||||
return loadVectorExtension(undefined, { policy: resolveEmbeddingInstallPolicy() });
|
||||
};
|
||||
/**
|
||||
* Bump this when the embedding text template changes in a way that should
|
||||
@@ -257,7 +275,7 @@ export const runEmbeddingPipeline = async (
|
||||
|
||||
try {
|
||||
const vectorAvailable = await ensureVectorExtensionAvailable();
|
||||
if (!vectorAvailable && isDev) {
|
||||
if (!vectorAvailable) {
|
||||
logger.warn(vectorUnavailableMessage);
|
||||
}
|
||||
|
||||
@@ -584,7 +602,11 @@ export const semanticSearch = async (
|
||||
string,
|
||||
{ distance: number; chunkIndex: number; startLine: number; endLine: number }
|
||||
>();
|
||||
if (await loadVectorExtension()) {
|
||||
// Query/read path: NEVER spawn a network INSTALL on a user query. If the
|
||||
// VECTOR extension was not pre-installed, fall back to exact scan rather than
|
||||
// blocking the query on a download (offline-first; see extension-loader.ts
|
||||
// "load-only" — used by all serve/MCP query paths).
|
||||
if (await loadVectorExtension(undefined, { policy: 'load-only' })) {
|
||||
try {
|
||||
bestChunks = await collectBestChunks(k, async (fetchLimit) => {
|
||||
const vectorQuery = `
|
||||
|
||||
@@ -188,6 +188,16 @@ function makeContract(
|
||||
|
||||
export interface ProtoServiceInfo {
|
||||
package: string;
|
||||
/**
|
||||
* Optional. Value of `option java_package = "..."` declared in the
|
||||
* same `.proto` file, when present and different from `package`.
|
||||
* Empty string when the option is absent or equals `package`. Used by
|
||||
* `detectionToContract()` to translate a Java import path back to the
|
||||
* proto package whenever the proto explicitly publishes its generated
|
||||
* Java code under a different namespace (a common pattern in
|
||||
* Google-style protobuf projects).
|
||||
*/
|
||||
javaPackage: string;
|
||||
serviceName: string;
|
||||
methods: string[];
|
||||
protoPath: string;
|
||||
@@ -207,6 +217,19 @@ function extractProtoImports(content: string): string[] {
|
||||
return imports;
|
||||
}
|
||||
|
||||
/**
|
||||
* Extract `option java_package = "..."` from a `.proto` file, if any.
|
||||
* The Java code generator places generated `XxxGrpc.java` classes under
|
||||
* this package (instead of the proto `package` declaration) when the
|
||||
* option is set. Real-world projects (Google Cloud Java APIs, internal
|
||||
* shaded SDKs) routinely use this to publish their Java artifacts under
|
||||
* a corporate namespace different from the wire-protocol package.
|
||||
*/
|
||||
function extractJavaPackageOption(content: string): string {
|
||||
const m = content.match(/^\s*option\s+java_package\s*=\s*"([\w.]+)"\s*;/m);
|
||||
return m?.[1] ?? '';
|
||||
}
|
||||
|
||||
function longestSharedSegmentRun(aPath: string, bPath: string): number {
|
||||
const a = aPath.split('/').filter(Boolean);
|
||||
const b = bPath.split('/').filter(Boolean);
|
||||
@@ -228,8 +251,18 @@ function longestSharedSegmentRun(aPath: string, bPath: string): number {
|
||||
async function buildProtoContext(repoPath: string): Promise<{
|
||||
packagesByProto: Map<string, string>;
|
||||
servicesByName: Map<string, ProtoServiceInfo[]>;
|
||||
/**
|
||||
* Reverse index: `option java_package` value → ProtoServiceInfo[]
|
||||
* declared in `.proto` files that ship under that Java namespace.
|
||||
* Only populated when `java_package` is set AND differs from
|
||||
* `package`. Lets `detectionToContract()` translate an import-derived
|
||||
* Java package back to its source proto package whenever the proto
|
||||
* is in the same repository.
|
||||
*/
|
||||
servicesByJavaPackage: Map<string, ProtoServiceInfo[]>;
|
||||
}> {
|
||||
const servicesByName = new Map<string, ProtoServiceInfo[]>();
|
||||
const servicesByJavaPackage = new Map<string, ProtoServiceInfo[]>();
|
||||
// `.gitnexusignore` / `.gitignore` honoured via the shared IgnoreService —
|
||||
// see `filesystem-walker.ts` for the canonical pattern. Replaces a
|
||||
// hardcoded `[node_modules, .git, vendor]` array; those names plus the
|
||||
@@ -292,6 +325,13 @@ async function buildProtoContext(repoPath: string): Promise<{
|
||||
const content = contents.get(normalizedRel);
|
||||
if (!content) continue;
|
||||
const pkg = resolvePackage(normalizedRel);
|
||||
const javaPkgOption = extractJavaPackageOption(content);
|
||||
// Only retain `javaPackage` when it actively diverges from `pkg`.
|
||||
// When equal (or absent), the import-derived path produces the
|
||||
// same FQN as the proto-derived path, so no translation is needed
|
||||
// and we keep the field empty to avoid populating the reverse
|
||||
// index with redundant entries.
|
||||
const javaPackage = javaPkgOption && javaPkgOption !== pkg ? javaPkgOption : '';
|
||||
|
||||
const serviceBlocks = extractServiceBlocks(content);
|
||||
for (const block of serviceBlocks) {
|
||||
@@ -303,6 +343,7 @@ async function buildProtoContext(repoPath: string): Promise<{
|
||||
}
|
||||
const info: ProtoServiceInfo = {
|
||||
package: pkg,
|
||||
javaPackage,
|
||||
serviceName: block.name,
|
||||
methods,
|
||||
protoPath: normalizedRel,
|
||||
@@ -310,10 +351,16 @@ async function buildProtoContext(repoPath: string): Promise<{
|
||||
const existing = servicesByName.get(block.name) ?? [];
|
||||
existing.push(info);
|
||||
servicesByName.set(block.name, existing);
|
||||
|
||||
if (javaPackage) {
|
||||
const byJava = servicesByJavaPackage.get(javaPackage) ?? [];
|
||||
byJava.push(info);
|
||||
servicesByJavaPackage.set(javaPackage, byJava);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return { packagesByProto, servicesByName };
|
||||
return { packagesByProto, servicesByName, servicesByJavaPackage };
|
||||
}
|
||||
|
||||
export async function buildProtoMap(repoPath: string): Promise<Map<string, ProtoServiceInfo[]>> {
|
||||
@@ -377,6 +424,7 @@ export class GrpcExtractor implements ContractExtractor {
|
||||
const out: ExtractedContract[] = [];
|
||||
const protoContext = await buildProtoContext(repoPath);
|
||||
const protoMap = protoContext.servicesByName;
|
||||
const javaPackageMap = protoContext.servicesByJavaPackage;
|
||||
|
||||
// ─── Proto files — definitive provider source ─────────────────
|
||||
// When tree-sitter-proto is available, .proto files are handled by
|
||||
@@ -435,7 +483,7 @@ export class GrpcExtractor implements ContractExtractor {
|
||||
continue;
|
||||
}
|
||||
for (const d of detections) {
|
||||
const contract = this.detectionToContract(d, rel, protoMap);
|
||||
const contract = this.detectionToContract(d, rel, protoMap, javaPackageMap);
|
||||
if (contract) out.push(contract);
|
||||
}
|
||||
}
|
||||
@@ -449,12 +497,163 @@ export class GrpcExtractor implements ContractExtractor {
|
||||
* either a service-level (`grpc::pkg.Svc/*`) or method-level
|
||||
* (`grpc::pkg.Svc/Method`) contract id, and selecting confidence
|
||||
* based on whether the proto map had an entry.
|
||||
*
|
||||
* Resolution order for the package prefix:
|
||||
*
|
||||
* 1. **Java-package translation** (when detection
|
||||
* supplied a `protoPackage` from a Java import).
|
||||
* A `.proto` in the SAME repo may set `option
|
||||
* java_package = "..."` to publish its generated
|
||||
* Java classes under a namespace different from
|
||||
* the proto `package`. Real-world projects (e.g.
|
||||
* Google Cloud Java APIs) routinely do this.
|
||||
* When the import-derived package matches that
|
||||
* `java_package` value, translate back to the
|
||||
* proto `package` so the resulting contract id
|
||||
* is wire-correct rather than Java-namespace.
|
||||
*
|
||||
* 2. **Per-repo proto map check** (when the same
|
||||
* service name has `.proto` candidates in this
|
||||
* repo). The proto file is the authoritative
|
||||
* source. If the proto's `package` agrees with
|
||||
* the import's `protoPackage`, both paths produce
|
||||
* the same FQN — emit it. If they DISAGREE (e.g.
|
||||
* a typo'd Java import, or a mismatched
|
||||
* java_package the reverse index didn't catch),
|
||||
* trust the proto map and warn — the import
|
||||
* MUST NOT silently overwrite an authoritative
|
||||
* proto package.
|
||||
*
|
||||
* 3. **Import-derived FQN fallback** (when neither
|
||||
* a `java_package` translation nor a proto map
|
||||
* candidate exists in this repo). Typical for the
|
||||
* "client-jar" pattern, where a consumer repo
|
||||
* depends on a published stub jar and never
|
||||
* carries the originating `.proto`. Use the
|
||||
* import path verbatim as the proto package. Note
|
||||
* the known limitation: when the published proto
|
||||
* sets `option java_package` differing from
|
||||
* `package`, the resulting FQN reflects the Java
|
||||
* namespace rather than the proto namespace and
|
||||
* will not match a provider repo's contract id —
|
||||
* we cannot translate without sight of the proto.
|
||||
*
|
||||
* 4. **Per-repo proto map (no import)** — the legacy
|
||||
* path. Used when the plugin didn't supply
|
||||
* `protoPackage` (no import statement, wildcard
|
||||
* import only, or non-Java languages that haven't
|
||||
* been retrofitted yet).
|
||||
*
|
||||
* 5. **Short-name fallback** — when none of the
|
||||
* above resolves a package, emit a service-only
|
||||
* short-name contract id (`grpc::Svc/*`),
|
||||
* preserving the pre-fix behaviour.
|
||||
*/
|
||||
private detectionToContract(
|
||||
d: GrpcDetection,
|
||||
filePath: string,
|
||||
protoMap: Map<string, ProtoServiceInfo[]>,
|
||||
javaPackageMap: Map<string, ProtoServiceInfo[]>,
|
||||
): ExtractedContract | null {
|
||||
if (d.protoPackage) {
|
||||
// Step 1: java_package translation. The import-derived package
|
||||
// may be the `option java_package` value of a `.proto` in the
|
||||
// SAME repo. Look it up and, if found for the same service name,
|
||||
// use the underlying proto `package` to build a wire-correct
|
||||
// contract id.
|
||||
const javaCandidates = javaPackageMap.get(d.protoPackage) ?? [];
|
||||
const javaTranslated = javaCandidates.find((p) => p.serviceName === d.serviceName);
|
||||
if (javaTranslated) {
|
||||
const cid = d.methodName
|
||||
? contractId(javaTranslated.package, d.serviceName, d.methodName)
|
||||
: serviceContractId(javaTranslated.package, d.serviceName);
|
||||
const meta: Record<string, unknown> = {
|
||||
service: d.serviceName,
|
||||
source: d.source,
|
||||
package: javaTranslated.package,
|
||||
protoPackageSource: 'import-translated',
|
||||
};
|
||||
if (d.methodName) meta.method = d.methodName;
|
||||
return makeContract(cid, d.role, filePath, d.symbolName, d.confidenceWithProto, meta);
|
||||
}
|
||||
|
||||
// Step 2: proto map cross-check. When this repo also carries a
|
||||
// `.proto` defining the same short service name, the proto is
|
||||
// authoritative and decides the package. The import is only used
|
||||
// to disambiguate among same-short-name candidates when the
|
||||
// resolution heuristic can't pick a unique winner on path alone.
|
||||
const candidates = protoMap.get(d.serviceName) ?? [];
|
||||
if (candidates.length > 0) {
|
||||
const proto = resolveProtoConflict(d.serviceName, filePath, candidates);
|
||||
if (proto === null) {
|
||||
// Ambiguous proto resolution; resolveProtoConflict already warned.
|
||||
return null;
|
||||
}
|
||||
const protoPkg = proto.package;
|
||||
if (protoPkg === d.protoPackage) {
|
||||
// Both paths agree.
|
||||
const cid = d.methodName
|
||||
? contractId(protoPkg, d.serviceName, d.methodName)
|
||||
: serviceContractId(protoPkg, d.serviceName);
|
||||
const meta: Record<string, unknown> = {
|
||||
service: d.serviceName,
|
||||
source: d.source,
|
||||
package: protoPkg,
|
||||
protoPackageSource: 'import',
|
||||
};
|
||||
if (d.methodName) meta.method = d.methodName;
|
||||
return makeContract(cid, d.role, filePath, d.symbolName, d.confidenceWithProto, meta);
|
||||
}
|
||||
// Disagreement. Trust the proto file and emit a warning so
|
||||
// operators can investigate the import. This protects against
|
||||
// the symmetric Finding 2 case: a stale or typo'd Java import
|
||||
// silently corrupting the contract id of a service whose
|
||||
// `.proto` lives in the same repo.
|
||||
logger.warn(
|
||||
`[grpc-extractor] Java import package "${d.protoPackage}" for service ` +
|
||||
`"${d.serviceName}" disagrees with local proto package "${protoPkg}" at ` +
|
||||
`${filePath}; using proto package as authoritative source`,
|
||||
);
|
||||
const cid = d.methodName
|
||||
? contractId(protoPkg, d.serviceName, d.methodName)
|
||||
: serviceContractId(protoPkg, d.serviceName);
|
||||
const meta: Record<string, unknown> = {
|
||||
service: d.serviceName,
|
||||
source: d.source,
|
||||
package: protoPkg,
|
||||
protoPackageSource: 'proto-override',
|
||||
importPackage: d.protoPackage,
|
||||
};
|
||||
if (d.methodName) meta.method = d.methodName;
|
||||
return makeContract(cid, d.role, filePath, d.symbolName, d.confidenceWithProto, meta);
|
||||
}
|
||||
|
||||
// Step 3: import-derived fallback. No `.proto` in this repo
|
||||
// names the service, and no `java_package` reverse-lookup
|
||||
// matched. Emit the FQN with the import-derived package. This
|
||||
// is the typical client-jar consumer path.
|
||||
//
|
||||
// Known limitation: when the published proto sets
|
||||
// `option java_package` to a value that differs from
|
||||
// `package`, this path produces a contract id that reflects
|
||||
// the Java namespace, not the proto namespace, and will not
|
||||
// match a provider repo. Resolving that case requires
|
||||
// group-level proto knowledge, which is intentionally out of
|
||||
// scope for this fix.
|
||||
const cid = d.methodName
|
||||
? contractId(d.protoPackage, d.serviceName, d.methodName)
|
||||
: serviceContractId(d.protoPackage, d.serviceName);
|
||||
const meta: Record<string, unknown> = {
|
||||
service: d.serviceName,
|
||||
source: d.source,
|
||||
package: d.protoPackage,
|
||||
protoPackageSource: 'import',
|
||||
};
|
||||
if (d.methodName) meta.method = d.methodName;
|
||||
return makeContract(cid, d.role, filePath, d.symbolName, d.confidenceWithProto, meta);
|
||||
}
|
||||
|
||||
// Steps 4 + 5: legacy per-repo proto map resolution (no import).
|
||||
const candidates = protoMap.get(d.serviceName) ?? [];
|
||||
const proto = resolveProtoConflict(d.serviceName, filePath, candidates);
|
||||
// If there were proto candidates but resolution was ambiguous, skip
|
||||
|
||||
@@ -78,6 +78,33 @@ const STUB_PATTERNS = compilePatterns({
|
||||
],
|
||||
} satisfies LanguagePatterns<Record<string, never>>);
|
||||
|
||||
// `import <pkg>.<XxxGrpc>;` — captures the proto package of the
|
||||
// imported gRPC class (e.g. `cn.unipus.ucf.admin.proto.client.service`
|
||||
// for `import cn.unipus.ucf.admin.proto.client.service.ContentRpcServiceGrpc`).
|
||||
// Used by `scan` to build a per-file `XxxGrpc → fullPackage` map so
|
||||
// consumer-side detections can carry a fully-qualified contract id
|
||||
// even when the consumer repo does not contain any `.proto` files.
|
||||
//
|
||||
// `import static …` is excluded by tree-sitter shape: the `name:`
|
||||
// field is only present on the non-static form. `import w.x.*;` is
|
||||
// also excluded for the same reason — wildcard imports have an
|
||||
// `asterisk` child instead of a named identifier.
|
||||
const GRPC_CLASS_IMPORT_PATTERNS = compilePatterns({
|
||||
name: 'java-grpc-class-import',
|
||||
language: Java,
|
||||
patterns: [
|
||||
{
|
||||
meta: {},
|
||||
query: `
|
||||
(import_declaration
|
||||
(scoped_identifier
|
||||
scope: (_) @import_pkg
|
||||
name: (identifier) @import_name (#match? @import_name "Grpc$")))
|
||||
`,
|
||||
},
|
||||
],
|
||||
} satisfies LanguagePatterns<Record<string, never>>);
|
||||
|
||||
/**
|
||||
* Check whether a `class_declaration` node has a `@GrpcService`
|
||||
* annotation in its modifiers list. In tree-sitter-java, class-level
|
||||
@@ -118,6 +145,39 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
|
||||
const out: GrpcDetection[] = [];
|
||||
const emittedClassIds = new Set<number>();
|
||||
|
||||
// ─── Build per-file gRPC class import map ───────────────────────
|
||||
// Maps `XxxGrpc` (short class name) → fully-qualified proto package
|
||||
// (e.g. `cn.unipus.ucf.admin.proto.client.service`). Used below to
|
||||
// tag both provider and consumer detections with a `protoPackage`
|
||||
// so the orchestrator can build a fully-qualified contract id
|
||||
// without depending on the current repo carrying any `.proto`
|
||||
// files. This is the key fix for client-jar consumer repos.
|
||||
//
|
||||
// Same-short-name disambiguation: when two distinct `import` lines
|
||||
// bring different `XxxGrpc` classes from different packages into
|
||||
// the same file (rare for grpc — the second import would be a
|
||||
// compile error in Java), the last one wins. Java's compiler
|
||||
// forbids that case so we don't bother modelling it.
|
||||
const grpcClassImports = new Map<string, string>();
|
||||
for (const match of runCompiledPatterns(GRPC_CLASS_IMPORT_PATTERNS, tree)) {
|
||||
const pkgNode = match.captures.import_pkg;
|
||||
const nameNode = match.captures.import_name;
|
||||
if (!pkgNode || !nameNode) continue;
|
||||
grpcClassImports.set(nameNode.text, pkgNode.text);
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve the fully-qualified proto package for a short service
|
||||
* name in this file. Looks up `<serviceName>Grpc` in the import
|
||||
* map; returns `undefined` when the class is referenced via a
|
||||
* fully-qualified name on every call site (no import line) or
|
||||
* when only a wildcard import is present. The orchestrator falls
|
||||
* back to the per-repo proto map in that case, preserving the
|
||||
* pre-fix behaviour.
|
||||
*/
|
||||
const protoPackageFor = (serviceName: string): string | undefined =>
|
||||
grpcClassImports.get(`${serviceName}Grpc`);
|
||||
|
||||
// ─── Providers: scoped form (`...Grpc.XxxImplBase`) ─────────────
|
||||
for (const match of runCompiledPatterns(SCOPED_IMPL_BASE_PATTERNS, tree)) {
|
||||
const classNode = match.captures.class;
|
||||
@@ -127,6 +187,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
|
||||
if (!serviceName) continue;
|
||||
emittedClassIds.add(classNode.id);
|
||||
const annotated = hasGrpcServiceAnnotation(classNode);
|
||||
const protoPackage = protoPackageFor(serviceName);
|
||||
out.push({
|
||||
role: 'provider',
|
||||
serviceName,
|
||||
@@ -134,6 +195,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
|
||||
source: annotated ? 'java_grpc_service' : 'java_impl_base',
|
||||
confidenceWithProto: 0.8,
|
||||
confidenceWithoutProto: 0.65,
|
||||
...(protoPackage ? { protoPackage } : {}),
|
||||
});
|
||||
}
|
||||
|
||||
@@ -147,6 +209,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
|
||||
if (!serviceName) continue;
|
||||
emittedClassIds.add(classNode.id);
|
||||
const annotated = hasGrpcServiceAnnotation(classNode);
|
||||
const protoPackage = protoPackageFor(serviceName);
|
||||
out.push({
|
||||
role: 'provider',
|
||||
serviceName,
|
||||
@@ -154,6 +217,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
|
||||
source: annotated ? 'java_grpc_service' : 'java_impl_base',
|
||||
confidenceWithProto: 0.8,
|
||||
confidenceWithoutProto: 0.65,
|
||||
...(protoPackage ? { protoPackage } : {}),
|
||||
});
|
||||
}
|
||||
|
||||
@@ -164,6 +228,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
|
||||
const grpcMatch = GRPC_SUFFIX_RE.exec(grpcClsNode.text);
|
||||
if (!grpcMatch) continue;
|
||||
const serviceName = grpcMatch[1];
|
||||
const protoPackage = protoPackageFor(serviceName);
|
||||
out.push({
|
||||
role: 'consumer',
|
||||
serviceName,
|
||||
@@ -171,6 +236,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
|
||||
source: 'java_stub',
|
||||
confidenceWithProto: 0.75,
|
||||
confidenceWithoutProto: 0.55,
|
||||
...(protoPackage ? { protoPackage } : {}),
|
||||
});
|
||||
}
|
||||
|
||||
|
||||
@@ -86,16 +86,33 @@ const NEW_QUALIFIED_CTOR_SPEC: PatternSpec<Record<string, never>> = {
|
||||
// proto loader). Matches either a bare call or an `obj.loadPackageDefinition(...)`
|
||||
// call. Plugin gates the qualified-constructor consumer on this —
|
||||
// structural check avoids materializing `tree.rootNode.text` for every file.
|
||||
const LOAD_PACKAGE_DEFINITION_SPEC: PatternSpec<Record<string, never>> = {
|
||||
meta: {},
|
||||
query: `
|
||||
(call_expression
|
||||
function: [
|
||||
(identifier) @fn (#eq? @fn "loadPackageDefinition")
|
||||
(member_expression property: (property_identifier) @fn (#eq? @fn "loadPackageDefinition"))
|
||||
])
|
||||
`,
|
||||
};
|
||||
//
|
||||
// These are TWO separate specs, NOT one `function: [ (identifier) ... (member_expression) ... ]`
|
||||
// alternation. Under the pinned tree-sitter@0.21.1 binding a top-level alternation
|
||||
// whose branches reuse the same capture name (`@fn`) collapses to one pattern with
|
||||
// a shared predicate bucket; the second branch's `@fn` is left unbound and its
|
||||
// `#eq?` is never enforced, so the member-expression branch would match EVERY
|
||||
// `obj.method(...)` call (e.g. `console.log(...)`) — turning this gate always-on
|
||||
// and emitting spurious qualified-constructor consumers. Two specs compile to two
|
||||
// queries with independent predicate buckets; `runCompiledPatterns` concatenates
|
||||
// their matches, so the `.length > 0` gate still means "either form is present".
|
||||
const LOAD_PACKAGE_DEFINITION_SPECS: PatternSpec<Record<string, never>>[] = [
|
||||
{
|
||||
meta: {},
|
||||
query: `
|
||||
(call_expression
|
||||
function: (identifier) @fn (#eq? @fn "loadPackageDefinition"))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {},
|
||||
query: `
|
||||
(call_expression
|
||||
function: (member_expression
|
||||
property: (property_identifier) @fn (#eq? @fn "loadPackageDefinition")))
|
||||
`,
|
||||
},
|
||||
];
|
||||
|
||||
interface NodeGrpcPatternBundle {
|
||||
grpcMethod: CompiledPatterns<Record<string, never>>;
|
||||
@@ -107,11 +124,14 @@ interface NodeGrpcPatternBundle {
|
||||
}
|
||||
|
||||
function compileBundle(language: unknown, name: string): NodeGrpcPatternBundle {
|
||||
const mk = (spec: PatternSpec<Record<string, never>>, suffix: string) =>
|
||||
const mk = (
|
||||
spec: PatternSpec<Record<string, never>> | PatternSpec<Record<string, never>>[],
|
||||
suffix: string,
|
||||
) =>
|
||||
compilePatterns({
|
||||
name: `${name}-${suffix}`,
|
||||
language,
|
||||
patterns: [spec],
|
||||
patterns: Array.isArray(spec) ? spec : [spec],
|
||||
} satisfies LanguagePatterns<Record<string, never>>);
|
||||
return {
|
||||
grpcMethod: mk(GRPC_METHOD_SPEC, 'grpc-method'),
|
||||
@@ -119,7 +139,7 @@ function compileBundle(language: unknown, name: string): NodeGrpcPatternBundle {
|
||||
getService: mk(GET_SERVICE_SPEC, 'get-service'),
|
||||
newSimpleCtor: mk(NEW_SIMPLE_CTOR_SPEC, 'new-simple-ctor'),
|
||||
newQualifiedCtor: mk(NEW_QUALIFIED_CTOR_SPEC, 'new-qualified-ctor'),
|
||||
loadPackageDefinition: mk(LOAD_PACKAGE_DEFINITION_SPEC, 'load-package-definition'),
|
||||
loadPackageDefinition: mk(LOAD_PACKAGE_DEFINITION_SPECS, 'load-package-definition'),
|
||||
};
|
||||
}
|
||||
|
||||
|
||||
@@ -18,7 +18,8 @@ import type { GrpcDetection, GrpcLanguagePlugin } from './types.js';
|
||||
*
|
||||
* The grammar is vendored in `vendor/tree-sitter-proto/` with
|
||||
* parser.c regenerated against tree-sitter-cli 0.24 (ABI version 14)
|
||||
* so it is compatible with the project's tree-sitter 0.25 runtime.
|
||||
* so it is compatible with the project's tree-sitter 0.21.1 runtime
|
||||
* (which loads ABI 13–14).
|
||||
*/
|
||||
|
||||
const _require = createRequire(import.meta.url);
|
||||
|
||||
@@ -36,6 +36,18 @@ export interface GrpcDetection {
|
||||
confidenceWithProto: number;
|
||||
/** Confidence when the proto map has no entry. */
|
||||
confidenceWithoutProto: number;
|
||||
/**
|
||||
* Optional. Fully-qualified proto package the detection's service
|
||||
* belongs to (e.g. `cn.unipus.ucf.admin.proto.client.service`),
|
||||
* derived directly from the source file's import statements when
|
||||
* available. When set, the orchestrator uses this package to build
|
||||
* the contract id INSTEAD of consulting the per-repo proto map —
|
||||
* letting consumer repos that don't carry `.proto` files (the
|
||||
* client-jar architecture used by most Java gRPC microservices)
|
||||
* still emit a fully-qualified contract id that matches the
|
||||
* provider repo's contract id verbatim.
|
||||
*/
|
||||
protoPackage?: string;
|
||||
}
|
||||
|
||||
/**
|
||||
|
||||
@@ -8,7 +8,13 @@ import { PYTHON_HTTP_PLUGIN } from './python.js';
|
||||
import { PHP_HTTP_PLUGIN } from './php.js';
|
||||
import { JAVASCRIPT_HTTP_PLUGIN, TYPESCRIPT_HTTP_PLUGIN, TSX_HTTP_PLUGIN } from './node.js';
|
||||
|
||||
export type { HttpDetection, HttpLanguagePlugin, HttpRole } from './types.js';
|
||||
export type {
|
||||
HttpDetection,
|
||||
HttpFileDetections,
|
||||
HttpLanguagePlugin,
|
||||
HttpRole,
|
||||
HttpScanInput,
|
||||
} from './types.js';
|
||||
|
||||
/**
|
||||
* File-extension → HTTP language plugin registry. The top-level
|
||||
|
||||
@@ -6,19 +6,31 @@ import {
|
||||
unquoteLiteral,
|
||||
type LanguagePatterns,
|
||||
} from '../tree-sitter-scanner.js';
|
||||
import type { HttpDetection, HttpLanguagePlugin } from './types.js';
|
||||
import type {
|
||||
HttpDetection,
|
||||
HttpFileDetections,
|
||||
HttpLanguagePlugin,
|
||||
HttpScanInput,
|
||||
} from './types.js';
|
||||
|
||||
/**
|
||||
* Java HTTP plugin. Handles:
|
||||
* - Spring `@RequestMapping` class prefixes + `@(Get|Post|...)Mapping` method annotations
|
||||
* - Spring `RestTemplate.getForObject/...`, `WebClient.method(HttpMethod.X, ...)`
|
||||
* - Spring `RestTemplate.getForObject/...`, `exchange(...)`
|
||||
* - Spring `WebClient.method(HttpMethod.X, ...)`, `WebClient.get().uri(...)`
|
||||
* - OkHttp `new Request.Builder().url("...")`
|
||||
* - OpenFeign interfaces with Spring MVC method annotations or
|
||||
* native `@RequestLine("METHOD /path")` annotations
|
||||
* - Java / Apache HttpClient literal request construction
|
||||
*
|
||||
* The plugin runs two pattern bundles: one to collect class-level
|
||||
* `@RequestMapping` prefixes keyed by the enclosing class node, and a
|
||||
* second to match method-level annotations. The `scan` function walks
|
||||
* up from each matched annotation to find its enclosing class and
|
||||
* combines the prefix with the method path.
|
||||
* Every route-defining annotation (class/interface `@RequestMapping`
|
||||
* prefixes, `@FeignClient(path)` prefixes, `@(Get|...)Mapping` method
|
||||
* routes and native `@RequestLine`s) is matched by a single consolidated
|
||||
* query (`JAVA_ROUTE_ANNOTATION_PATTERNS`) in one pass via
|
||||
* `scanRouteAnnotations`. The `scan` function then walks up from each
|
||||
* matched method to its enclosing class/interface to combine the prefix
|
||||
* with the method path. Call-site consumers (RestTemplate, WebClient,
|
||||
* OkHttp, Java/Apache HttpClient) keep their own focused queries.
|
||||
*/
|
||||
|
||||
const METHOD_ANNOTATION_TO_HTTP: Record<string, string> = {
|
||||
@@ -29,93 +41,178 @@ const METHOD_ANNOTATION_TO_HTTP: Record<string, string> = {
|
||||
PatchMapping: 'PATCH',
|
||||
};
|
||||
|
||||
// ─── Provider: Spring class-level @RequestMapping prefix ──────────────
|
||||
// Two patterns are needed because the AST shape differs depending on
|
||||
// whether the annotation uses a positional argument or a named one:
|
||||
// Each route-defining annotation has two AST shapes — a positional argument
|
||||
// and a named one — that must both be matched:
|
||||
// @RequestMapping("/api") → (annotation_argument_list (string_literal))
|
||||
// @RequestMapping(path = "/api") → (annotation_argument_list (element_value_pair key:(identifier) value:(string_literal)))
|
||||
// @RequestMapping(value = "/api") → same as above
|
||||
// For named arguments only the route member keys (`path`/`value`) carry a URL;
|
||||
// non-route attributes (`produces`, `consumes`, `headers`, `name`, `params`)
|
||||
// would otherwise be mis-extracted (e.g. `produces = "application/json"` would
|
||||
// corrupt every route). That key filtering is done in `isRouteMemberKey`, and
|
||||
// all of these annotations are matched by the one `JAVA_ROUTE_ANNOTATION_PATTERNS`
|
||||
// query below (see its header for why the filtering lives in JS, not the query).
|
||||
interface SpringRouteBinding {
|
||||
method: string;
|
||||
path: string;
|
||||
}
|
||||
|
||||
interface SpringMethodInfo {
|
||||
name: string;
|
||||
routes: SpringRouteBinding[];
|
||||
}
|
||||
|
||||
interface SpringTypeInfo {
|
||||
filePath: string;
|
||||
kind: 'class' | 'interface';
|
||||
name: string;
|
||||
classPrefix: string;
|
||||
implementedInterfaces: string[];
|
||||
isController: boolean;
|
||||
methods: SpringMethodInfo[];
|
||||
}
|
||||
|
||||
// ─── Route-defining annotations (one generic query, one pass) ─────────
|
||||
// Every Java route-mapper annotation shares one shape: an annotation carrying a
|
||||
// single string argument — positional `"..."` or named `key = "..."` — on a
|
||||
// class, interface, or method. This SINGLE query matches that shape generically;
|
||||
// `scanRouteAnnotations` then reads the annotation NAME (`@ann`) and declaration
|
||||
// kind (`@node.type`) in its for-loop to decide what each match means. Adding a
|
||||
// new framework annotation that follows this single-string-argument shape is a
|
||||
// change to that loop (and the lookup maps), not to this query. Annotations with
|
||||
// a different argument shape — e.g. an array value `@RequestMapping({"/a","/b"})`
|
||||
// — are out of scope here (as they were for the prior queries) and would need a
|
||||
// new branch.
|
||||
//
|
||||
// The named-argument pattern MUST constrain the `key` field to the route
|
||||
// member names (`path`/`value`); without it, the query also captures
|
||||
// non-route attributes such as `produces`, `consumes`, `headers`, `name`,
|
||||
// `params` (their right-hand string literals would be mis-extracted as
|
||||
// route prefixes — e.g. `produces = "application/json"` would corrupt
|
||||
// every method route under that controller). The sibling
|
||||
// `topic-patterns/java.ts` uses the same `key:` constraint approach.
|
||||
const SPRING_CLASS_PREFIX_PATTERNS = compilePatterns({
|
||||
name: 'java-spring-class-prefix',
|
||||
// Captures (shared across all branches; intentionally framework-agnostic):
|
||||
// @ann → the annotation name identifier (RequestMapping, GetMapping, RequestLine, …)
|
||||
// @node → the enclosing declaration (class_declaration | interface_declaration | method_declaration)
|
||||
// @value → the string-literal argument
|
||||
// @key → the named-argument member key (absent for the positional shape)
|
||||
// @member → the method name (method_declaration branches only)
|
||||
//
|
||||
// The query carries NO `#eq?` / `#match?` predicates. Under the pinned
|
||||
// tree-sitter 0.21.x binding a top-level `[ ... ]` alternation compiles to one
|
||||
// pattern whose text predicates share a single bucket keyed by capture name, and
|
||||
// a `#match?` against a capture absent from the matched branch evaluates FALSE —
|
||||
// silently dropping sibling-branch matches. Keeping the query predicate-free
|
||||
// sidesteps that hazard entirely; all name/key discrimination lives in the
|
||||
// for-loop, where it reads as straight-line code.
|
||||
const JAVA_ROUTE_ANNOTATION_PATTERNS = compilePatterns({
|
||||
name: 'java-route-annotation',
|
||||
language: Java,
|
||||
patterns: [
|
||||
{
|
||||
meta: {},
|
||||
query: `
|
||||
(class_declaration
|
||||
(modifiers
|
||||
(annotation
|
||||
name: (identifier) @ann (#eq? @ann "RequestMapping")
|
||||
arguments: (annotation_argument_list (string_literal) @prefix)))) @class
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {},
|
||||
query: `
|
||||
(class_declaration
|
||||
(modifiers
|
||||
(annotation
|
||||
name: (identifier) @ann (#eq? @ann "RequestMapping")
|
||||
arguments: (annotation_argument_list
|
||||
(element_value_pair
|
||||
key: (identifier) @key (#match? @key "^(path|value)$")
|
||||
value: (string_literal) @prefix))))) @class
|
||||
[
|
||||
(class_declaration
|
||||
(modifiers
|
||||
(annotation
|
||||
name: (identifier) @ann
|
||||
arguments: (annotation_argument_list (string_literal) @value)))) @node
|
||||
(interface_declaration
|
||||
(modifiers
|
||||
(annotation
|
||||
name: (identifier) @ann
|
||||
arguments: (annotation_argument_list (string_literal) @value)))) @node
|
||||
(class_declaration
|
||||
(modifiers
|
||||
(annotation
|
||||
name: (identifier) @ann
|
||||
arguments: (annotation_argument_list
|
||||
(element_value_pair
|
||||
key: (identifier) @key
|
||||
value: (string_literal) @value))))) @node
|
||||
(interface_declaration
|
||||
(modifiers
|
||||
(annotation
|
||||
name: (identifier) @ann
|
||||
arguments: (annotation_argument_list
|
||||
(element_value_pair
|
||||
key: (identifier) @key
|
||||
value: (string_literal) @value))))) @node
|
||||
(method_declaration
|
||||
(modifiers
|
||||
(annotation
|
||||
name: (identifier) @ann
|
||||
arguments: (annotation_argument_list (string_literal) @value)))
|
||||
name: (identifier) @member) @node
|
||||
(method_declaration
|
||||
(modifiers
|
||||
(annotation
|
||||
name: (identifier) @ann
|
||||
arguments: (annotation_argument_list
|
||||
(element_value_pair
|
||||
key: (identifier) @key
|
||||
value: (string_literal) @value))))
|
||||
name: (identifier) @member) @node
|
||||
]
|
||||
`,
|
||||
},
|
||||
],
|
||||
} satisfies LanguagePatterns<Record<string, never>>);
|
||||
|
||||
// ─── Provider: Spring @(Get|Post|...)Mapping method annotations ───────
|
||||
// Same dual-pattern approach: positional vs named argument. The named
|
||||
// pattern restricts the annotation member name to `path`/`value` to
|
||||
// avoid capturing unrelated string-valued attributes
|
||||
// (`produces`, `consumes`, `headers`, `name`, `params`, ...).
|
||||
const SPRING_METHOD_ROUTE_PATTERNS = compilePatterns({
|
||||
name: 'java-spring-method-route',
|
||||
const SPRING_TYPE_DECLARATION_PATTERNS = compilePatterns({
|
||||
name: 'java-spring-type-declaration',
|
||||
language: Java,
|
||||
patterns: [
|
||||
{
|
||||
meta: {},
|
||||
query: `
|
||||
(method_declaration
|
||||
(modifiers
|
||||
(annotation
|
||||
name: (identifier) @ann (#match? @ann "^(Get|Post|Put|Delete|Patch)Mapping$")
|
||||
arguments: (annotation_argument_list (string_literal) @path)))
|
||||
name: (identifier) @method_name) @method
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {},
|
||||
query: `
|
||||
(method_declaration
|
||||
(modifiers
|
||||
(annotation
|
||||
name: (identifier) @ann (#match? @ann "^(Get|Post|Put|Delete|Patch)Mapping$")
|
||||
arguments: (annotation_argument_list
|
||||
(element_value_pair
|
||||
key: (identifier) @key (#match? @key "^(path|value)$")
|
||||
value: (string_literal) @path))))
|
||||
name: (identifier) @method_name) @method
|
||||
[
|
||||
(class_declaration name: (identifier) @type_name) @type
|
||||
(interface_declaration name: (identifier) @type_name) @type
|
||||
]
|
||||
`,
|
||||
},
|
||||
],
|
||||
} satisfies LanguagePatterns<Record<string, never>>);
|
||||
|
||||
// ─── Consumer: OpenFeign `@RequestLine("METHOD /path")` parsing ───────
|
||||
// OpenFeign's native annotation pairs an HTTP method and path in a single
|
||||
// string literal — see https://github.com/OpenFeign/feign#interface-annotations.
|
||||
// It is method-level only and is mutually exclusive with Spring MVC
|
||||
// `@GetMapping` / `@PostMapping` etc. on the same method (mixing them
|
||||
// requires a different Feign Contract — they are not combined). The match
|
||||
// itself comes from `JAVA_ROUTE_ANNOTATION_PATTERNS`; this regex splits the
|
||||
// verb from the path of the captured literal.
|
||||
//
|
||||
// Examples:
|
||||
// @RequestLine("GET /users/{id}")
|
||||
// @RequestLine("POST /users?status=active")
|
||||
const REQUEST_LINE_VERB_RE = /^\s*(GET|POST|PUT|DELETE|PATCH|HEAD|OPTIONS)\s+(\S.*?)\s*$/i;
|
||||
|
||||
/**
|
||||
* Parse a Feign `@RequestLine` value into a method + path pair.
|
||||
*
|
||||
* `@RequestLine("METHOD /path[?query]")` packs both fields in one string;
|
||||
* the query portion is dropped because contract IDs are method+path only
|
||||
* (consistent with how other consumers like RestTemplate/WebClient drop
|
||||
* query strings when their values are inline literals).
|
||||
*
|
||||
* Returns null if the value is not a recognized HTTP verb followed by a
|
||||
* path beginning with `/`.
|
||||
*/
|
||||
function parseRequestLine(raw: string): { method: string; path: string } | null {
|
||||
const match = REQUEST_LINE_VERB_RE.exec(raw);
|
||||
if (!match) return null;
|
||||
const [, verb, rest] = match;
|
||||
if (typeof verb !== 'string' || typeof rest !== 'string') return null;
|
||||
const queryIdx = rest.indexOf('?');
|
||||
const pathOnly = (queryIdx >= 0 ? rest.slice(0, queryIdx) : rest).trim();
|
||||
if (!pathOnly.startsWith('/')) return null;
|
||||
return { method: verb.toUpperCase(), path: pathOnly };
|
||||
}
|
||||
|
||||
// ─── Consumer: Spring RestTemplate (object-named + method-named) ──────
|
||||
// RestTemplate.getForObject / getForEntity → GET
|
||||
// RestTemplate.postForObject / postForEntity → POST
|
||||
// RestTemplate.put → PUT
|
||||
// RestTemplate.delete → DELETE
|
||||
// RestTemplate.patchForObject → PATCH
|
||||
// Source-scan only: receiver must be named exactly `restTemplate`.
|
||||
// Fields, `this.restTemplate`, aliases, and other injection names are deferred.
|
||||
const REST_TEMPLATE_TO_HTTP: Record<string, string> = {
|
||||
getForObject: 'GET',
|
||||
getForEntity: 'GET',
|
||||
@@ -146,22 +243,48 @@ const REST_TEMPLATE_PATTERNS = compilePatterns({
|
||||
],
|
||||
} satisfies LanguagePatterns<RestTemplateMeta>);
|
||||
|
||||
// ─── Consumer: Spring WebClient — webClient.method(HttpMethod.X, "path") ─
|
||||
const WEB_CLIENT_PATTERNS = compilePatterns({
|
||||
name: 'java-web-client',
|
||||
const REST_TEMPLATE_EXCHANGE_PATTERNS = compilePatterns({
|
||||
name: 'java-rest-template-exchange',
|
||||
language: Java,
|
||||
patterns: [
|
||||
{
|
||||
meta: { framework: 'spring-rest-template' },
|
||||
query: `
|
||||
(method_invocation
|
||||
object: (identifier) @obj (#eq? @obj "restTemplate")
|
||||
name: (identifier) @method (#eq? @method "exchange")
|
||||
arguments: (argument_list
|
||||
. (string_literal) @path
|
||||
(field_access
|
||||
object: (identifier) @httpMethodCls (#eq? @httpMethodCls "HttpMethod")
|
||||
field: (identifier) @http_method)))
|
||||
`,
|
||||
},
|
||||
],
|
||||
} satisfies LanguagePatterns<RestTemplateMeta>);
|
||||
|
||||
const WEB_CLIENT_SHORT_TO_HTTP: Record<string, string> = {
|
||||
get: 'GET',
|
||||
post: 'POST',
|
||||
put: 'PUT',
|
||||
delete: 'DELETE',
|
||||
patch: 'PATCH',
|
||||
};
|
||||
|
||||
const WEB_CLIENT_SHORT_FORM_PATTERNS = compilePatterns({
|
||||
name: 'java-web-client-short-form',
|
||||
language: Java,
|
||||
patterns: [
|
||||
{
|
||||
meta: {},
|
||||
query: `
|
||||
(method_invocation
|
||||
object: (identifier) @obj (#eq? @obj "webClient")
|
||||
name: (identifier) @method (#eq? @method "method")
|
||||
arguments: (argument_list
|
||||
(field_access
|
||||
object: (identifier) @httpMethodCls (#eq? @httpMethodCls "HttpMethod")
|
||||
field: (identifier) @http_method)
|
||||
(string_literal) @path))
|
||||
object: (method_invocation
|
||||
object: (identifier) @obj (#eq? @obj "webClient")
|
||||
name: (identifier) @verb (#match? @verb "^(get|post|put|delete|patch)$")
|
||||
arguments: (argument_list))
|
||||
name: (identifier) @uri_method (#eq? @uri_method "uri")
|
||||
arguments: (argument_list . (string_literal) @path))
|
||||
`,
|
||||
},
|
||||
],
|
||||
@@ -188,10 +311,58 @@ const OK_HTTP_PATTERNS = compilePatterns({
|
||||
],
|
||||
} satisfies LanguagePatterns<Record<string, never>>);
|
||||
|
||||
const JAVA_HTTP_CLIENT_PATTERNS = compilePatterns({
|
||||
name: 'java-http-client',
|
||||
language: Java,
|
||||
patterns: [
|
||||
{
|
||||
meta: {},
|
||||
query: `
|
||||
(method_invocation
|
||||
object: (method_invocation
|
||||
object: (method_invocation
|
||||
object: (identifier) @builderCls (#eq? @builderCls "HttpRequest")
|
||||
name: (identifier) @newBuilder (#eq? @newBuilder "newBuilder")
|
||||
arguments: (argument_list))
|
||||
name: (identifier) @uri_method (#eq? @uri_method "uri")
|
||||
arguments: (argument_list
|
||||
(method_invocation
|
||||
object: (identifier) @uriCls (#eq? @uriCls "URI")
|
||||
name: (identifier) @create (#eq? @create "create")
|
||||
arguments: (argument_list . (string_literal) @path))))
|
||||
name: (identifier) @http_method (#match? @http_method "^(GET|POST|PUT|DELETE)$"))
|
||||
`,
|
||||
},
|
||||
],
|
||||
} satisfies LanguagePatterns<Record<string, never>>);
|
||||
|
||||
const APACHE_HTTP_CLIENT_TO_HTTP: Record<string, string> = {
|
||||
HttpGet: 'GET',
|
||||
HttpPost: 'POST',
|
||||
HttpPut: 'PUT',
|
||||
HttpDelete: 'DELETE',
|
||||
HttpPatch: 'PATCH',
|
||||
};
|
||||
|
||||
const APACHE_HTTP_CLIENT_PATTERNS = compilePatterns({
|
||||
name: 'java-apache-http-client',
|
||||
language: Java,
|
||||
patterns: [
|
||||
{
|
||||
meta: {},
|
||||
query: `
|
||||
(object_creation_expression
|
||||
type: (type_identifier) @type (#match? @type "^Http(Get|Post|Put|Delete|Patch)$")
|
||||
arguments: (argument_list . (string_literal) @path))
|
||||
`,
|
||||
},
|
||||
],
|
||||
} satisfies LanguagePatterns<Record<string, never>>);
|
||||
|
||||
/**
|
||||
* Find the nearest enclosing class_declaration ancestor for a node, or
|
||||
* null if the node is top-level. Tree-sitter's SyntaxNode.parent walks
|
||||
* one level at a time.
|
||||
* Find the nearest enclosing class/interface declaration ancestor for
|
||||
* a node, or null if the node is top-level. Tree-sitter's
|
||||
* SyntaxNode.parent walks one level at a time.
|
||||
*/
|
||||
function findEnclosingClass(node: Parser.SyntaxNode): Parser.SyntaxNode | null {
|
||||
let cur: Parser.SyntaxNode | null = node.parent;
|
||||
@@ -202,6 +373,15 @@ function findEnclosingClass(node: Parser.SyntaxNode): Parser.SyntaxNode | null {
|
||||
return null;
|
||||
}
|
||||
|
||||
function findEnclosingInterface(node: Parser.SyntaxNode): Parser.SyntaxNode | null {
|
||||
let cur: Parser.SyntaxNode | null = node.parent;
|
||||
while (cur) {
|
||||
if (cur.type === 'interface_declaration') return cur;
|
||||
cur = cur.parent;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Join a class-level prefix and a method-level path into a single URL
|
||||
* path. Mirrors the semantics of the original regex implementation:
|
||||
@@ -215,45 +395,356 @@ function joinPath(prefix: string, methodPath: string): string {
|
||||
return `/${cleanPrefix}/${cleanSub}`;
|
||||
}
|
||||
|
||||
function getNodeName(node: Parser.SyntaxNode): string | null {
|
||||
return node.childForFieldName('name')?.text ?? null;
|
||||
}
|
||||
|
||||
function hasAnnotation(node: Parser.SyntaxNode, names: string | readonly string[]): boolean {
|
||||
const modifiers = node.namedChildren.find((child) => child.type === 'modifiers');
|
||||
if (!modifiers) return false;
|
||||
const allowed = new Set(typeof names === 'string' ? [names] : names);
|
||||
const stack = [...modifiers.namedChildren];
|
||||
while (stack.length > 0) {
|
||||
const cur = stack.pop()!;
|
||||
const annotationName = cur.childForFieldName('name')?.text ?? '';
|
||||
const simpleName = annotationName.split('.').pop() ?? annotationName;
|
||||
if (
|
||||
(cur.type === 'annotation' || cur.type === 'marker_annotation') &&
|
||||
(allowed.has(annotationName) || allowed.has(simpleName))
|
||||
) {
|
||||
return true;
|
||||
}
|
||||
stack.push(...cur.namedChildren);
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/**
|
||||
* A named annotation argument contributes a route only when its member key is
|
||||
* `path` or `value`; a positional argument (no key node) always qualifies.
|
||||
* This is the JS-side replacement for the in-query `^(path|value)$` filter and
|
||||
* drops Spring's non-route string attributes (`produces`, `consumes`,
|
||||
* `headers`, `name`, `params`) that would otherwise be mis-read as routes.
|
||||
*/
|
||||
function isRouteMemberKey(keyNode: Parser.SyntaxNode | undefined): boolean {
|
||||
if (!keyNode) return true;
|
||||
return keyNode.text === 'path' || keyNode.text === 'value';
|
||||
}
|
||||
|
||||
interface MethodRouteAnnotation {
|
||||
methodNode: Parser.SyntaxNode;
|
||||
methodName: string | null;
|
||||
httpMethod: string;
|
||||
rawPath: string;
|
||||
}
|
||||
|
||||
interface RequestLineAnnotation {
|
||||
methodNode: Parser.SyntaxNode;
|
||||
methodName: string | null;
|
||||
parsed: { method: string; path: string };
|
||||
}
|
||||
|
||||
interface RouteAnnotationScan {
|
||||
/** Spring `@RequestMapping` URL prefix per class/interface node id (last write wins). */
|
||||
prefixByTypeId: Map<number, string>;
|
||||
/** OpenFeign interface prefix per interface node id; `@FeignClient(path)` wins over `@RequestMapping`. */
|
||||
feignPrefixByInterfaceId: Map<number, string>;
|
||||
/** One entry per resolved Spring `@(Get|...)Mapping` route — a method with N mappings yields N entries. */
|
||||
methodRoutes: MethodRouteAnnotation[];
|
||||
/** One entry per OpenFeign `@RequestLine` whose value parses to a verb + path. */
|
||||
requestLines: RequestLineAnnotation[];
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve every Java route-defining annotation in a single tree-sitter pass.
|
||||
*
|
||||
* The generic `JAVA_ROUTE_ANNOTATION_PATTERNS` query yields one match per
|
||||
* annotation-carrying-a-string-argument on any class / interface / method. This
|
||||
* loop reads the annotation name and declaration kind to decide what each match
|
||||
* means, ignoring annotations it does not recognise. The HTTP verb map
|
||||
* (`METHOD_ANNOTATION_TO_HTTP`) and the `path`/`value` key filter
|
||||
* (`isRouteMemberKey`) live here rather than in the query (see its header).
|
||||
*/
|
||||
function scanRouteAnnotations(tree: Parser.Tree): RouteAnnotationScan {
|
||||
const matches = runCompiledPatterns(JAVA_ROUTE_ANNOTATION_PATTERNS, tree);
|
||||
|
||||
// The two prefix maps intentionally diverge for the same interface node:
|
||||
// `prefixByTypeId` feeds the Spring *provider* path (class prefix +
|
||||
// collectSpringTypes cross-file inheritance), while `feignPrefixByInterfaceId`
|
||||
// feeds the OpenFeign *consumer* path in scan(). An interface carrying both
|
||||
// `@RequestMapping` and `@FeignClient(path)` lands a different value in each.
|
||||
const prefixByTypeId = new Map<number, string>();
|
||||
const feignPrefixByInterfaceId = new Map<number, string>();
|
||||
const methodRoutes: MethodRouteAnnotation[] = [];
|
||||
const requestLines: RequestLineAnnotation[] = [];
|
||||
// Interface `@RequestMapping` prefixes rank below `@FeignClient(path)`;
|
||||
// collect them and apply only after the FeignClient pass below.
|
||||
const interfaceRequestMappingPrefixes: Array<{ id: number; prefix: string }> = [];
|
||||
|
||||
for (const { captures } of matches) {
|
||||
const annNode = captures.ann;
|
||||
const node = captures.node;
|
||||
const valueNode = captures.value;
|
||||
if (!annNode || !node || !valueNode) continue;
|
||||
const ann = annNode.text;
|
||||
const keyNode = captures.key; // undefined for the positional shape
|
||||
|
||||
if (node.type === 'method_declaration') {
|
||||
// Method-level: a Spring `@(Get|...)Mapping` route, or native `@RequestLine`.
|
||||
const httpMethod = METHOD_ANNOTATION_TO_HTTP[ann];
|
||||
if (httpMethod) {
|
||||
if (!isRouteMemberKey(keyNode)) continue;
|
||||
const rawPath = unquoteLiteral(valueNode.text);
|
||||
if (rawPath !== null) {
|
||||
methodRoutes.push({
|
||||
methodNode: node,
|
||||
methodName: captures.member?.text ?? null,
|
||||
httpMethod,
|
||||
rawPath,
|
||||
});
|
||||
}
|
||||
} else if (ann === 'RequestLine') {
|
||||
// Feign packs verb + path in one literal; its only named argument is `value`.
|
||||
if (keyNode && keyNode.text !== 'value') continue;
|
||||
const raw = unquoteLiteral(valueNode.text);
|
||||
const parsed = raw !== null ? parseRequestLine(raw) : null;
|
||||
if (parsed) {
|
||||
requestLines.push({
|
||||
methodNode: node,
|
||||
methodName: captures.member?.text ?? null,
|
||||
parsed,
|
||||
});
|
||||
}
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
// Type-level (class or interface): a Spring `@RequestMapping` URL prefix, or
|
||||
// — on an interface — an OpenFeign `@FeignClient(path = "...")` prefix.
|
||||
if (ann === 'RequestMapping') {
|
||||
if (!isRouteMemberKey(keyNode)) continue;
|
||||
const prefix = unquoteLiteral(valueNode.text);
|
||||
if (prefix !== null) {
|
||||
prefixByTypeId.set(node.id, prefix);
|
||||
if (node.type === 'interface_declaration') {
|
||||
interfaceRequestMappingPrefixes.push({ id: node.id, prefix });
|
||||
}
|
||||
}
|
||||
} else if (ann === 'FeignClient' && node.type === 'interface_declaration') {
|
||||
// Feign's `name`/`value` identify a service, not a path — only `path` is a prefix.
|
||||
if (!keyNode || keyNode.text !== 'path') continue;
|
||||
const prefix = unquoteLiteral(valueNode.text);
|
||||
if (prefix !== null && !feignPrefixByInterfaceId.has(node.id)) {
|
||||
feignPrefixByInterfaceId.set(node.id, prefix);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
for (const { id, prefix } of interfaceRequestMappingPrefixes) {
|
||||
if (!feignPrefixByInterfaceId.has(id)) feignPrefixByInterfaceId.set(id, prefix);
|
||||
}
|
||||
|
||||
return { prefixByTypeId, feignPrefixByInterfaceId, methodRoutes, requestLines };
|
||||
}
|
||||
|
||||
function collectDirectMethods(typeNode: Parser.SyntaxNode): Parser.SyntaxNode[] {
|
||||
const out: Parser.SyntaxNode[] = [];
|
||||
const visit = (node: Parser.SyntaxNode): void => {
|
||||
for (const child of node.namedChildren) {
|
||||
if (child.type === 'method_declaration') {
|
||||
out.push(child);
|
||||
continue;
|
||||
}
|
||||
if (
|
||||
child !== typeNode &&
|
||||
(child.type === 'class_declaration' || child.type === 'interface_declaration')
|
||||
) {
|
||||
continue;
|
||||
}
|
||||
visit(child);
|
||||
}
|
||||
};
|
||||
visit(typeNode);
|
||||
return out;
|
||||
}
|
||||
|
||||
function collectImplementedInterfaces(typeNode: Parser.SyntaxNode): string[] {
|
||||
const interfacesNode = typeNode.childForFieldName('interfaces');
|
||||
if (!interfacesNode) return [];
|
||||
const out: string[] = [];
|
||||
const visit = (node: Parser.SyntaxNode): void => {
|
||||
if (node.type === 'type_identifier' || node.type === 'scoped_type_identifier') {
|
||||
out.push(node.text.split('.').pop() ?? node.text);
|
||||
return;
|
||||
}
|
||||
for (const child of node.namedChildren) visit(child);
|
||||
};
|
||||
visit(interfacesNode);
|
||||
return out;
|
||||
}
|
||||
|
||||
function collectSpringTypes(filePath: string, tree: Parser.Tree): SpringTypeInfo[] {
|
||||
const { prefixByTypeId, methodRoutes } = scanRouteAnnotations(tree);
|
||||
const routesByMethodId = new Map<number, SpringRouteBinding[]>();
|
||||
for (const route of methodRoutes) {
|
||||
const routes = routesByMethodId.get(route.methodNode.id) ?? [];
|
||||
routes.push({ method: route.httpMethod, path: route.rawPath });
|
||||
routesByMethodId.set(route.methodNode.id, routes);
|
||||
}
|
||||
const out: SpringTypeInfo[] = [];
|
||||
|
||||
for (const match of runCompiledPatterns(SPRING_TYPE_DECLARATION_PATTERNS, tree)) {
|
||||
const typeNode = match.captures.type;
|
||||
const typeNameNode = match.captures.type_name;
|
||||
if (!typeNode || !typeNameNode) continue;
|
||||
const kind = typeNode.type === 'interface_declaration' ? 'interface' : 'class';
|
||||
const methods = collectDirectMethods(typeNode)
|
||||
.map((methodNode) => ({
|
||||
name: getNodeName(methodNode),
|
||||
routes: routesByMethodId.get(methodNode.id) ?? [],
|
||||
}))
|
||||
.filter((method): method is SpringMethodInfo => method.name !== null);
|
||||
|
||||
out.push({
|
||||
filePath,
|
||||
kind,
|
||||
name: typeNameNode.text,
|
||||
classPrefix: prefixByTypeId.get(typeNode.id) ?? '',
|
||||
implementedInterfaces: kind === 'class' ? collectImplementedInterfaces(typeNode) : [],
|
||||
isController: kind === 'class' && hasAnnotation(typeNode, ['RestController', 'Controller']),
|
||||
methods,
|
||||
});
|
||||
}
|
||||
|
||||
return out;
|
||||
}
|
||||
|
||||
function scanSpringProject(files: readonly HttpScanInput[]): HttpFileDetections[] {
|
||||
const types = files.flatMap((file) => collectSpringTypes(file.filePath, file.tree));
|
||||
const interfaceRoutes = new Map<string, Map<string, SpringRouteBinding[]> | null>();
|
||||
|
||||
for (const type of types) {
|
||||
if (type.kind !== 'interface') continue;
|
||||
if (interfaceRoutes.has(type.name)) {
|
||||
interfaceRoutes.set(type.name, null);
|
||||
continue;
|
||||
}
|
||||
const methodMap = new Map<string, SpringRouteBinding[]>();
|
||||
for (const method of type.methods) {
|
||||
const routes = method.routes.map((route) => ({
|
||||
method: route.method,
|
||||
path: type.classPrefix ? joinPath(type.classPrefix, route.path) : route.path,
|
||||
}));
|
||||
if (routes.length > 0) methodMap.set(method.name, routes);
|
||||
}
|
||||
interfaceRoutes.set(type.name, methodMap);
|
||||
}
|
||||
|
||||
const detectionsByFile = new Map<string, HttpDetection[]>();
|
||||
for (const type of types) {
|
||||
if (type.kind !== 'class' || !type.isController) continue;
|
||||
for (const method of type.methods) {
|
||||
if (method.routes.length > 0) continue;
|
||||
const inheritedRoutes = type.implementedInterfaces.flatMap((interfaceName) => {
|
||||
const routeMap = interfaceRoutes.get(interfaceName);
|
||||
if (!routeMap) return [];
|
||||
const routes = routeMap.get(method.name) ?? [];
|
||||
return routes.map((route) => ({
|
||||
method: route.method,
|
||||
path: joinPath(type.classPrefix, route.path),
|
||||
}));
|
||||
});
|
||||
|
||||
for (const route of inheritedRoutes) {
|
||||
const detections = detectionsByFile.get(type.filePath) ?? [];
|
||||
detections.push({
|
||||
role: 'provider',
|
||||
framework: 'spring',
|
||||
method: route.method,
|
||||
path: route.path,
|
||||
name: method.name,
|
||||
confidence: 0.8,
|
||||
});
|
||||
detectionsByFile.set(type.filePath, detections);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return [...detectionsByFile.entries()].map(([filePath, detections]) => ({
|
||||
filePath,
|
||||
detections,
|
||||
}));
|
||||
}
|
||||
|
||||
export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
|
||||
name: 'java-http',
|
||||
language: Java,
|
||||
scan(tree) {
|
||||
const out: HttpDetection[] = [];
|
||||
|
||||
// ─── Providers: Spring class prefix + method annotations ────────
|
||||
const prefixByClassId = new Map<number, string>();
|
||||
for (const match of runCompiledPatterns(SPRING_CLASS_PREFIX_PATTERNS, tree)) {
|
||||
const prefixNode = match.captures.prefix;
|
||||
const classNode = match.captures.class;
|
||||
if (!prefixNode || !classNode) continue;
|
||||
const prefix = unquoteLiteral(prefixNode.text);
|
||||
if (prefix !== null) prefixByClassId.set(classNode.id, prefix);
|
||||
}
|
||||
// ─── Spring providers + OpenFeign consumers (one query pass) ────
|
||||
// `scanRouteAnnotations` resolves every route-defining annotation —
|
||||
// class/interface prefixes, method `@(Get|...)Mapping`s and native
|
||||
// `@RequestLine`s — from a single `matches()` pass over the tree.
|
||||
const { prefixByTypeId, feignPrefixByInterfaceId, methodRoutes, requestLines } =
|
||||
scanRouteAnnotations(tree);
|
||||
|
||||
for (const match of runCompiledPatterns(SPRING_METHOD_ROUTE_PATTERNS, tree)) {
|
||||
const annNode = match.captures.ann;
|
||||
const pathNode = match.captures.path;
|
||||
const nameNode = match.captures.method_name;
|
||||
const methodNode = match.captures.method;
|
||||
if (!annNode || !pathNode || !methodNode) continue;
|
||||
const httpMethod = METHOD_ANNOTATION_TO_HTTP[annNode.text];
|
||||
if (!httpMethod) continue;
|
||||
const rawPath = unquoteLiteral(pathNode.text);
|
||||
if (rawPath === null) continue;
|
||||
const enclosingClass = findEnclosingClass(methodNode);
|
||||
const prefix = enclosingClass ? (prefixByClassId.get(enclosingClass.id) ?? '') : '';
|
||||
const fullPath = joinPath(prefix, rawPath);
|
||||
// A `@(Get|...)Mapping` inside a `@FeignClient` interface is an OpenFeign
|
||||
// *consumer* (it describes a remote call); the same annotation inside a
|
||||
// class is a Spring *provider*. A mapping on a non-Feign interface has no
|
||||
// enclosing class and is dropped here — interface→controller inheritance is
|
||||
// handled by `scanProject`.
|
||||
for (const route of methodRoutes) {
|
||||
const enclosingInterface = findEnclosingInterface(route.methodNode);
|
||||
if (enclosingInterface && hasAnnotation(enclosingInterface, 'FeignClient')) {
|
||||
const prefix = feignPrefixByInterfaceId.get(enclosingInterface.id) ?? '';
|
||||
out.push({
|
||||
role: 'consumer',
|
||||
framework: 'openfeign',
|
||||
method: route.httpMethod,
|
||||
path: joinPath(prefix, route.rawPath),
|
||||
name: route.methodName,
|
||||
confidence: 0.7,
|
||||
});
|
||||
continue;
|
||||
}
|
||||
const enclosingClass = findEnclosingClass(route.methodNode);
|
||||
if (!enclosingClass) continue;
|
||||
const prefix = prefixByTypeId.get(enclosingClass.id) ?? '';
|
||||
out.push({
|
||||
role: 'provider',
|
||||
framework: 'spring',
|
||||
method: httpMethod,
|
||||
path: fullPath,
|
||||
name: nameNode?.text ?? null,
|
||||
method: route.httpMethod,
|
||||
path: joinPath(prefix, route.rawPath),
|
||||
name: route.methodName,
|
||||
confidence: 0.8,
|
||||
});
|
||||
}
|
||||
|
||||
// Native OpenFeign `@RequestLine("METHOD /path")`. Method-level only and
|
||||
// always declared on an interface (Feign builds a proxy from the interface).
|
||||
// We do NOT require an enclosing `@FeignClient`: `@RequestLine` is a core
|
||||
// `feign.*` annotation used with `Feign.builder()`, whereas `@FeignClient`
|
||||
// is the Spring Cloud variant that uses Spring MVC annotations instead — the
|
||||
// two are effectively mutually exclusive, so requiring `@FeignClient` here
|
||||
// would miss the annotation's primary use. The `RequestLine` name is itself
|
||||
// a strong, framework-specific signal, so a structural interface check is
|
||||
// enough to keep false positives away. A `@FeignClient(path=...)` prefix is
|
||||
// still applied when present (rare, but harmless).
|
||||
for (const requestLine of requestLines) {
|
||||
const enclosingInterface = findEnclosingInterface(requestLine.methodNode);
|
||||
if (!enclosingInterface) continue;
|
||||
const prefix = feignPrefixByInterfaceId.get(enclosingInterface.id) ?? '';
|
||||
out.push({
|
||||
role: 'consumer',
|
||||
framework: 'openfeign',
|
||||
method: requestLine.parsed.method,
|
||||
path: joinPath(prefix, requestLine.parsed.path),
|
||||
name: requestLine.methodName,
|
||||
confidence: 0.75,
|
||||
});
|
||||
}
|
||||
|
||||
// ─── Consumers: RestTemplate ────────────────────────────────────
|
||||
for (const match of runCompiledPatterns(REST_TEMPLATE_PATTERNS, tree)) {
|
||||
const methodNode = match.captures.method;
|
||||
@@ -273,8 +764,7 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
|
||||
});
|
||||
}
|
||||
|
||||
// ─── Consumers: WebClient.method(HttpMethod.X, "path") ──────────
|
||||
for (const match of runCompiledPatterns(WEB_CLIENT_PATTERNS, tree)) {
|
||||
for (const match of runCompiledPatterns(REST_TEMPLATE_EXCHANGE_PATTERNS, tree)) {
|
||||
const httpMethodNode = match.captures.http_method;
|
||||
const pathNode = match.captures.path;
|
||||
if (!httpMethodNode || !pathNode) continue;
|
||||
@@ -282,7 +772,7 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
|
||||
if (path === null) continue;
|
||||
out.push({
|
||||
role: 'consumer',
|
||||
framework: 'spring-web-client',
|
||||
framework: 'spring-rest-template',
|
||||
method: httpMethodNode.text.toUpperCase(),
|
||||
path,
|
||||
name: null,
|
||||
@@ -290,6 +780,28 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
|
||||
});
|
||||
}
|
||||
|
||||
// ─── Consumers: WebClient.get().uri("path") short form ─────────
|
||||
// Source-scan only: receiver must be named exactly `webClient`.
|
||||
// The real long-form chain `webClient.method(HttpMethod.X).uri("/x")`
|
||||
// needs multi-hop chain analysis and is intentionally deferred.
|
||||
for (const match of runCompiledPatterns(WEB_CLIENT_SHORT_FORM_PATTERNS, tree)) {
|
||||
const verbNode = match.captures.verb;
|
||||
const pathNode = match.captures.path;
|
||||
if (!verbNode || !pathNode) continue;
|
||||
const httpMethod = WEB_CLIENT_SHORT_TO_HTTP[verbNode.text];
|
||||
if (!httpMethod) continue;
|
||||
const path = unquoteLiteral(pathNode.text);
|
||||
if (path === null) continue;
|
||||
out.push({
|
||||
role: 'consumer',
|
||||
framework: 'spring-web-client',
|
||||
method: httpMethod,
|
||||
path,
|
||||
name: null,
|
||||
confidence: 0.7,
|
||||
});
|
||||
}
|
||||
|
||||
// ─── Consumers: OkHttp Request.Builder().url("path") ────────────
|
||||
for (const match of runCompiledPatterns(OK_HTTP_PATTERNS, tree)) {
|
||||
const pathNode = match.captures.path;
|
||||
@@ -306,6 +818,45 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
|
||||
});
|
||||
}
|
||||
|
||||
// ─── Consumers: Java HttpClient request builder ─────────────────
|
||||
// Java's builder exposes GET/POST/PUT/DELETE helpers. PATCH uses
|
||||
// `.method("PATCH", body)`, which is intentionally deferred.
|
||||
for (const match of runCompiledPatterns(JAVA_HTTP_CLIENT_PATTERNS, tree)) {
|
||||
const httpMethodNode = match.captures.http_method;
|
||||
const pathNode = match.captures.path;
|
||||
if (!httpMethodNode || !pathNode) continue;
|
||||
const path = unquoteLiteral(pathNode.text);
|
||||
if (path === null) continue;
|
||||
out.push({
|
||||
role: 'consumer',
|
||||
framework: 'java-http-client',
|
||||
method: httpMethodNode.text.toUpperCase(),
|
||||
path,
|
||||
name: null,
|
||||
confidence: 0.65,
|
||||
});
|
||||
}
|
||||
|
||||
// ─── Consumers: Apache HttpClient request constructors ──────────
|
||||
for (const match of runCompiledPatterns(APACHE_HTTP_CLIENT_PATTERNS, tree)) {
|
||||
const typeNode = match.captures.type;
|
||||
const pathNode = match.captures.path;
|
||||
if (!typeNode || !pathNode) continue;
|
||||
const httpMethod = APACHE_HTTP_CLIENT_TO_HTTP[typeNode.text];
|
||||
if (!httpMethod) continue;
|
||||
const path = unquoteLiteral(pathNode.text);
|
||||
if (path === null) continue;
|
||||
out.push({
|
||||
role: 'consumer',
|
||||
framework: 'apache-http-client',
|
||||
method: httpMethod,
|
||||
path,
|
||||
name: null,
|
||||
confidence: 0.65,
|
||||
});
|
||||
}
|
||||
|
||||
return out;
|
||||
},
|
||||
scanProject: scanSpringProject,
|
||||
};
|
||||
|
||||
@@ -17,18 +17,22 @@ import type { HttpDetection, HttpLanguagePlugin } from './types.js';
|
||||
* named annotation arguments (`@GetMapping(value = "/x")` and
|
||||
* `@GetMapping(path = "/x")`) are supported.
|
||||
*
|
||||
* **Consumers** (this PR) — three call-site patterns common in Kotlin
|
||||
* **Consumers** — four call-site patterns common in Kotlin
|
||||
* Spring projects:
|
||||
*
|
||||
* 1. `restTemplate.getForObject("/x", ...)` and friends
|
||||
* 2. `webClient.get().uri("/x")` (short form, 1 verb hop + 1 uri hop)
|
||||
* 3. `Request.Builder().url("/x")` (OkHttp)
|
||||
* 1. `restTemplate.getForObject("/x", ...)` and friends (#1855)
|
||||
* 2. `webClient.get().uri("/x")` — short form (#1855)
|
||||
* 3. `Request.Builder().url("/x")` — OkHttp (#1855)
|
||||
* 4. `webClient.method(HttpMethod.X).uri("/y")` — long form (this PR)
|
||||
*
|
||||
* The long-form `webClient.method(HttpMethod.X).uri("/y")` chain is
|
||||
* intentionally deferred to a follow-up: it requires walk-up logic
|
||||
* to recover the verb from a sibling `call_expression`, and we can
|
||||
* land 80% of real-world Kotlin Spring consumer coverage with the
|
||||
* three simpler patterns above.
|
||||
* The long form puts the verb on a sibling `call_expression` two hops
|
||||
* away from the path. Rather than introducing imperative walk-up logic,
|
||||
* we use a single deeper tree-sitter query that matches the full chain
|
||||
* structurally — see `WEB_CLIENT_LONG_PATTERNS` below. The verb is
|
||||
* captured directly as the `simple_identifier` of `HttpMethod.X`, so
|
||||
* variable-bound verbs (`val verb = HttpMethod.PATCH; webClient.method(verb)...`)
|
||||
* are intentionally NOT picked up — those need a graph-aware resolver
|
||||
* and are out of scope for source-scan.
|
||||
*
|
||||
* tree-sitter-kotlin (fwcd) AST shapes used here:
|
||||
* class_declaration
|
||||
@@ -109,6 +113,16 @@ const WEB_CLIENT_SHORT_TO_HTTP: Record<string, string> = {
|
||||
patch: 'PATCH',
|
||||
};
|
||||
|
||||
/**
|
||||
* Allowed HTTP verbs for the WebClient long-form path
|
||||
* `webClient.method(HttpMethod.X).uri("/y")`. Compiled once at module
|
||||
* load (instead of inside the scan loop) per maintainer feedback on
|
||||
* PR #1884. Mirrors the keys of `WEB_CLIENT_SHORT_TO_HTTP` above —
|
||||
* keeping HEAD/OPTIONS/TRACE intentionally excluded for symmetry
|
||||
* with the short form and the Java plugin.
|
||||
*/
|
||||
const WEB_CLIENT_LONG_VERB_RE = /^(GET|POST|PUT|DELETE|PATCH)$/;
|
||||
|
||||
/**
|
||||
* Build the plugin only if the Kotlin grammar is available. Compiling
|
||||
* the queries against a null grammar would throw at module load time
|
||||
@@ -265,8 +279,9 @@ function buildKotlinPlugin(language: unknown): HttpLanguagePlugin {
|
||||
// - outer call's first value_argument is a string literal
|
||||
//
|
||||
// The long-form `webClient.method(HttpMethod.GET).uri("/x")` chain
|
||||
// uses an extra navigation hop and an enum field access — it's
|
||||
// intentionally out of scope here (see file header).
|
||||
// uses an extra navigation hop and an enum field access — handled
|
||||
// by `WEB_CLIENT_LONG_PATTERNS` below, separately so each query is
|
||||
// straightforward to reason about.
|
||||
const WEB_CLIENT_SHORT_PATTERNS = compilePatterns({
|
||||
name: 'kotlin-web-client-short',
|
||||
language,
|
||||
@@ -290,6 +305,59 @@ function buildKotlinPlugin(language: unknown): HttpLanguagePlugin {
|
||||
],
|
||||
} satisfies LanguagePatterns<Record<string, never>>);
|
||||
|
||||
// ─── Consumer: Spring WebClient (long form) ───────────────────────────
|
||||
// The fluent long form passes the verb as a `HttpMethod.X` enum field
|
||||
// access through `.method(...)`, then carries the path on a separate
|
||||
// `.uri(...)` hop further down the chain:
|
||||
//
|
||||
// webClient.method(HttpMethod.GET).uri("/x").retrieve().awaitBody<T>()
|
||||
//
|
||||
// Compared to the short form there are two extra structural hops:
|
||||
// - the inner `.method(...)` `call_expression` has a `value_argument`
|
||||
// whose payload is itself a `navigation_expression` (HttpMethod → .GET)
|
||||
// - the outer `.uri(...)` is reached via one more
|
||||
// `navigation_expression` wrapping that inner call
|
||||
//
|
||||
// We capture the verb at the `simple_identifier` under `HttpMethod`'s
|
||||
// `navigation_suffix`. That `simple_identifier` is the literal field
|
||||
// name (`GET`, `POST`, ...) used in source — Kotlin enum fields by
|
||||
// convention are upper-case, matching `HttpMethod` from
|
||||
// `org.springframework.http`. We forward the captured text as-is.
|
||||
//
|
||||
// Variable-bound verbs (`val verb = HttpMethod.PATCH; webClient.method(verb)...`)
|
||||
// do NOT match — they fail the `(navigation_expression ...)` shape
|
||||
// because the value_argument carries a bare `simple_identifier` instead
|
||||
// of a `HttpMethod.X` field access. This is intentional: source-scan
|
||||
// can't follow the binding without graph context. Pinned by an
|
||||
// anti-overreach test in the consumer suite.
|
||||
const WEB_CLIENT_LONG_PATTERNS = compilePatterns({
|
||||
name: 'kotlin-web-client-long',
|
||||
language,
|
||||
patterns: [
|
||||
{
|
||||
meta: {},
|
||||
query: `
|
||||
(call_expression
|
||||
(navigation_expression
|
||||
(call_expression
|
||||
(navigation_expression
|
||||
(simple_identifier) @obj (#eq? @obj "webClient")
|
||||
(navigation_suffix
|
||||
(simple_identifier) @method_call (#eq? @method_call "method")))
|
||||
(call_suffix
|
||||
(value_arguments
|
||||
. (value_argument
|
||||
(navigation_expression
|
||||
(simple_identifier) @httpMethodCls (#eq? @httpMethodCls "HttpMethod")
|
||||
(navigation_suffix (simple_identifier) @verb))))))
|
||||
(navigation_suffix (simple_identifier) @uri (#eq? @uri "uri")))
|
||||
(call_suffix
|
||||
(value_arguments . (value_argument . (string_literal) @path))))
|
||||
`,
|
||||
},
|
||||
],
|
||||
} satisfies LanguagePatterns<Record<string, never>>);
|
||||
|
||||
// ─── Consumer: OkHttp Request.Builder().url("/x") ─────────────────────
|
||||
// Kotlin parses `Request.Builder()` as a `call_expression` whose
|
||||
// callee is a `navigation_expression` (Request → .Builder), NOT as
|
||||
@@ -437,6 +505,33 @@ function buildKotlinPlugin(language: unknown): HttpLanguagePlugin {
|
||||
});
|
||||
}
|
||||
|
||||
// ─── Consumers: WebClient long form (.method(HttpMethod.X) → .uri) ─
|
||||
for (const match of runCompiledPatterns(WEB_CLIENT_LONG_PATTERNS, tree)) {
|
||||
const verbNode = match.captures.verb;
|
||||
const pathNode = match.captures.path;
|
||||
if (!verbNode || !pathNode) continue;
|
||||
// The captured text is the literal `HttpMethod.X` field name.
|
||||
// Spring's `org.springframework.http.HttpMethod` defines GET,
|
||||
// POST, PUT, DELETE, PATCH, HEAD, OPTIONS, TRACE — we only
|
||||
// emit for the five verbs we already handle elsewhere, so
|
||||
// exotic ones are silently skipped (consistent with the
|
||||
// short form's WEB_CLIENT_SHORT_TO_HTTP guard). The accepted
|
||||
// verb regex is hoisted to module scope (see
|
||||
// `WEB_CLIENT_LONG_VERB_RE` near the top of this file).
|
||||
const verbText = verbNode.text;
|
||||
if (!WEB_CLIENT_LONG_VERB_RE.test(verbText)) continue;
|
||||
const path = unquoteLiteral(pathNode.text);
|
||||
if (path === null) continue;
|
||||
out.push({
|
||||
role: 'consumer',
|
||||
framework: 'spring-web-client',
|
||||
method: verbText,
|
||||
path,
|
||||
name: null,
|
||||
confidence: 0.7,
|
||||
});
|
||||
}
|
||||
|
||||
// ─── Consumers: OkHttp Request.Builder().url("path") ────────────
|
||||
for (const match of runCompiledPatterns(OK_HTTP_PATTERNS, tree)) {
|
||||
const pathNode = match.captures.path;
|
||||
|
||||
@@ -622,8 +622,9 @@ interface PythonRepoContext {
|
||||
|
||||
/** Strip `.py` and return the bare basename (e.g. `api/users.py` → `users`). */
|
||||
function fileShortKey(rel: string): string {
|
||||
const slash = rel.lastIndexOf('/');
|
||||
const file = slash >= 0 ? rel.slice(slash + 1) : rel;
|
||||
const normalized = rel.replace(/\\/g, '/');
|
||||
const slash = normalized.lastIndexOf('/');
|
||||
const file = slash >= 0 ? normalized.slice(slash + 1) : normalized;
|
||||
return file.endsWith('.py') ? file.slice(0, -3) : file;
|
||||
}
|
||||
|
||||
@@ -633,7 +634,8 @@ function fileShortKey(rel: string): string {
|
||||
* case callers should fall back to the short key.
|
||||
*/
|
||||
function fileLongKey(rel: string): string {
|
||||
const noExt = rel.endsWith('.py') ? rel.slice(0, -3) : rel;
|
||||
const normalized = rel.replace(/\\/g, '/');
|
||||
const noExt = normalized.endsWith('.py') ? normalized.slice(0, -3) : normalized;
|
||||
const lastSlash = noExt.lastIndexOf('/');
|
||||
if (lastSlash < 0) return '';
|
||||
const beforeLast = noExt.slice(0, lastSlash);
|
||||
|
||||
@@ -40,6 +40,16 @@ export interface HttpDetection {
|
||||
confidence: number;
|
||||
}
|
||||
|
||||
export interface HttpScanInput {
|
||||
filePath: string;
|
||||
tree: Parser.Tree;
|
||||
}
|
||||
|
||||
export interface HttpFileDetections {
|
||||
filePath: string;
|
||||
detections: HttpDetection[];
|
||||
}
|
||||
|
||||
/**
|
||||
* One language-scoped HTTP plugin. The plugin owns the tree-sitter
|
||||
* grammar and the `scan` function that translates a parsed tree into
|
||||
@@ -95,4 +105,10 @@ export interface HttpLanguagePlugin {
|
||||
* single-file plugins can keep their unary `scan(tree)` shape.
|
||||
*/
|
||||
scan(tree: Parser.Tree, repoContext?: RepoContext, fileRel?: string): HttpDetection[];
|
||||
/**
|
||||
* Optional project-level scan hook for language rules that require
|
||||
* multiple files, such as Java controllers inheriting Spring mappings
|
||||
* from annotated interfaces.
|
||||
*/
|
||||
scanProject?(files: readonly HttpScanInput[]): HttpFileDetections[];
|
||||
}
|
||||
|
||||
@@ -6,7 +6,13 @@ import type { ContractExtractor, CypherExecutor } from '../contract-extractor.js
|
||||
import type { ExtractedContract, RepoHandle } from '../types.js';
|
||||
import { readSafe } from './fs-utils.js';
|
||||
import { parseSourceSafe } from '../../tree-sitter/safe-parse.js';
|
||||
import { getPluginForFile, HTTP_SCAN_GLOB, type HttpDetection } from './http-patterns/index.js';
|
||||
import {
|
||||
getPluginForFile,
|
||||
HTTP_SCAN_GLOB,
|
||||
type HttpDetection,
|
||||
type HttpLanguagePlugin,
|
||||
type HttpScanInput,
|
||||
} from './http-patterns/index.js';
|
||||
|
||||
/**
|
||||
* Language-agnostic orchestrator for HTTP route (provider + consumer)
|
||||
@@ -160,6 +166,12 @@ export class HttpRouteExtractor implements ContractExtractor {
|
||||
// both graph-assisted enrichment and source-scan emission.
|
||||
const parser = new Parser();
|
||||
const cachedDetections = new Map<string, HttpDetection[]>();
|
||||
const cachedInputs = new Map<
|
||||
string,
|
||||
{ plugin: HttpLanguagePlugin; input: HttpScanInput; repoContext: unknown } | null
|
||||
>();
|
||||
const projectDetections = new Map<string, HttpDetection[]>();
|
||||
let projectScanComplete = false;
|
||||
|
||||
// Per-plugin cross-file context (e.g. Python's FastAPI router →
|
||||
// include_router(prefix=...) map). Built lazily on first
|
||||
@@ -189,32 +201,50 @@ export class HttpRouteExtractor implements ContractExtractor {
|
||||
}
|
||||
};
|
||||
|
||||
const getDetections = async (rel: string): Promise<HttpDetection[]> => {
|
||||
const cached = cachedDetections.get(rel);
|
||||
if (cached) return cached;
|
||||
const getScanInput = async (
|
||||
rel: string,
|
||||
): Promise<{
|
||||
plugin: HttpLanguagePlugin;
|
||||
input: HttpScanInput;
|
||||
repoContext: unknown;
|
||||
} | null> => {
|
||||
if (cachedInputs.has(rel)) return cachedInputs.get(rel) ?? null;
|
||||
const plugin = getPluginForFile(rel);
|
||||
if (!plugin) {
|
||||
cachedDetections.set(rel, []);
|
||||
return [];
|
||||
cachedInputs.set(rel, null);
|
||||
return null;
|
||||
}
|
||||
const repoContext = await ensureRepoContext(plugin);
|
||||
const content = readSafe(repoPath, rel);
|
||||
if (!content) {
|
||||
cachedDetections.set(rel, []);
|
||||
return [];
|
||||
cachedInputs.set(rel, null);
|
||||
return null;
|
||||
}
|
||||
try {
|
||||
parser.setLanguage(plugin.language);
|
||||
const tree = parseSourceSafe(parser, content);
|
||||
const detections = plugin.scan(tree, repoContext, rel);
|
||||
cachedDetections.set(rel, detections);
|
||||
return detections;
|
||||
const input = { filePath: rel, tree };
|
||||
const item = { plugin, input, repoContext };
|
||||
cachedInputs.set(rel, item);
|
||||
return item;
|
||||
} catch {
|
||||
cachedDetections.set(rel, []);
|
||||
return [];
|
||||
cachedInputs.set(rel, null);
|
||||
return null;
|
||||
}
|
||||
};
|
||||
|
||||
const getDetections = async (rel: string): Promise<HttpDetection[]> => {
|
||||
const cached = cachedDetections.get(rel);
|
||||
if (cached) return cached;
|
||||
const scanInput = await getScanInput(rel);
|
||||
const ownDetections = scanInput
|
||||
? scanInput.plugin.scan(scanInput.input.tree, scanInput.repoContext, rel)
|
||||
: [];
|
||||
const detections = [...ownDetections, ...(projectDetections.get(rel) ?? [])];
|
||||
cachedDetections.set(rel, detections);
|
||||
return detections;
|
||||
};
|
||||
|
||||
// Glob the source-scan file list at most once per extract() —
|
||||
// both provider and consumer fallback paths share the same list.
|
||||
let scannedFiles: string[] | null = null;
|
||||
@@ -224,20 +254,46 @@ export class HttpRouteExtractor implements ContractExtractor {
|
||||
return scannedFiles;
|
||||
};
|
||||
|
||||
const collectProjectDetections = async (files: string[]): Promise<void> => {
|
||||
if (projectScanComplete) return;
|
||||
projectScanComplete = true;
|
||||
const byPlugin = new Map<HttpLanguagePlugin, HttpScanInput[]>();
|
||||
for (const rel of files) {
|
||||
const scanInput = await getScanInput(rel);
|
||||
if (!scanInput?.plugin.scanProject) continue;
|
||||
const items = byPlugin.get(scanInput.plugin) ?? [];
|
||||
items.push(scanInput.input);
|
||||
byPlugin.set(scanInput.plugin, items);
|
||||
}
|
||||
|
||||
for (const [plugin, inputs] of byPlugin) {
|
||||
const results = plugin.scanProject?.(inputs) ?? [];
|
||||
for (const result of results) {
|
||||
const existing = projectDetections.get(result.filePath) ?? [];
|
||||
projectDetections.set(result.filePath, [...existing, ...result.detections]);
|
||||
}
|
||||
}
|
||||
|
||||
cachedDetections.clear();
|
||||
};
|
||||
|
||||
const files = await getScannedFiles();
|
||||
await collectProjectDetections(files);
|
||||
|
||||
const graphProviders =
|
||||
dbExecutor != null ? await this.extractProvidersGraph(dbExecutor, getDetections) : [];
|
||||
// Source scan always runs to capture routes in languages/files not covered
|
||||
// by graph edges; the glob and per-file parse results are cached above.
|
||||
const providers = this.mergeGraphAndSourceContracts(
|
||||
graphProviders,
|
||||
await this.extractProvidersSourceScan(await getScannedFiles(), getDetections),
|
||||
await this.extractProvidersSourceScan(files, getDetections),
|
||||
);
|
||||
|
||||
const graphConsumers =
|
||||
dbExecutor != null ? await this.extractConsumersGraph(dbExecutor, getDetections) : [];
|
||||
const consumers = this.mergeGraphAndSourceContracts(
|
||||
graphConsumers,
|
||||
await this.extractConsumersSourceScan(await getScannedFiles(), getDetections),
|
||||
await this.extractConsumersSourceScan(files, getDetections),
|
||||
);
|
||||
|
||||
return [...providers, ...consumers];
|
||||
|
||||
@@ -40,6 +40,8 @@ import { getLanguageFromFilename, SupportedLanguages } from 'gitnexus-shared';
|
||||
import { isRegistryPrimary } from './registry-primary-flag.js';
|
||||
import { isVerboseIngestionEnabled } from './utils/verbose.js';
|
||||
import {
|
||||
ALWAYS_ON_SLOW_FILE_WARN_THROTTLE_MS,
|
||||
alwaysOnSlowFileWarnMs,
|
||||
deferredCallFileSlowMs,
|
||||
deferredCallLogEveryN,
|
||||
getDeferredProfileDroppedCount,
|
||||
@@ -408,7 +410,7 @@ const findEnclosingFunction = (
|
||||
|
||||
while (current) {
|
||||
if (FUNCTION_NODE_TYPES.has(current.type)) {
|
||||
const efnResult = provider.methodExtractor?.extractFunctionName?.(current);
|
||||
const efnResult = provider.methodExtractor?.extractFunctionName?.(current, filePath);
|
||||
const funcName = efnResult?.funcName ?? genericFuncName(current);
|
||||
const label = efnResult?.label ?? inferFunctionLabel(current.type);
|
||||
|
||||
@@ -923,6 +925,7 @@ export const processCalls = async (
|
||||
const importedReturnTypes = importedReturnTypesMap?.get(file.path);
|
||||
const importedRawReturnTypes = importedRawReturnTypesMap?.get(file.path);
|
||||
const typeEnv = buildTypeEnv(tree, language, {
|
||||
filePath: file.path,
|
||||
model: ctx.model,
|
||||
parentMap,
|
||||
importedBindings,
|
||||
@@ -1033,15 +1036,18 @@ export const processCalls = async (
|
||||
? { declaredType: routedFieldInfo.type }
|
||||
: {}),
|
||||
});
|
||||
const relId = generateId('DEFINES', `${fileId}->${nodeId}`);
|
||||
graph.addRelationship({
|
||||
id: relId,
|
||||
sourceId: fileId,
|
||||
targetId: nodeId,
|
||||
type: 'DEFINES',
|
||||
confidence: 1.0,
|
||||
reason: '',
|
||||
});
|
||||
// Only emit File -> Property DEFINES for top-level properties (issue #1944).
|
||||
if (!propEnclosingClassId) {
|
||||
const relId = generateId('DEFINES', `${fileId}->${nodeId}`);
|
||||
graph.addRelationship({
|
||||
id: relId,
|
||||
sourceId: fileId,
|
||||
targetId: nodeId,
|
||||
type: 'DEFINES',
|
||||
confidence: 1.0,
|
||||
reason: '',
|
||||
});
|
||||
}
|
||||
if (propEnclosingClassId) {
|
||||
graph.addRelationship({
|
||||
id: generateId('HAS_PROPERTY', `${propEnclosingClassId}->${nodeId}`),
|
||||
@@ -1291,7 +1297,8 @@ export const processCalls = async (
|
||||
while (p) {
|
||||
if (FUNCTION_NODE_TYPES.has(p.type)) {
|
||||
const funcName =
|
||||
provider.methodExtractor?.extractFunctionName?.(p)?.funcName ?? genericFuncName(p);
|
||||
provider.methodExtractor?.extractFunctionName?.(p, file.path)?.funcName ??
|
||||
genericFuncName(p);
|
||||
if (funcName) {
|
||||
scope = `${funcName}@${p.startIndex}`;
|
||||
break;
|
||||
@@ -2930,6 +2937,15 @@ export const processCallsFromExtracted = async (
|
||||
const logEveryN = profileCalls ? deferredCallLogEveryN() : 0;
|
||||
let skippedRegistryPrimaryFiles = 0;
|
||||
|
||||
// Always-on slow-file watchdog (#1741). Independent of the verbose/profile
|
||||
// gate above: even a plain `analyze` run surfaces ONE actionable warning
|
||||
// when a single file's call resolution is pathologically slow — turning the
|
||||
// silent "stuck at Resolving calls (N/M)" symptom into a named culprit.
|
||||
// Throttled so a genuinely slow repo can't produce a warn storm.
|
||||
const alwaysSlowFileMs = alwaysOnSlowFileWarnMs();
|
||||
let lastSlowFileWarnAt = 0;
|
||||
let suppressedSlowFileWarnings = 0;
|
||||
|
||||
// Fresh dropped-log counter per analyze run — the module-private counter
|
||||
// in deferred-resolution-profile.ts is process-lived, so without a reset
|
||||
// here it would accumulate across consecutive analyze invocations in the
|
||||
@@ -2941,12 +2957,16 @@ export const processCallsFromExtracted = async (
|
||||
// denominator stays stable as the loop iterates. Otherwise `${totalFiles -
|
||||
// skippedRegistryPrimaryFiles}` drifts upward — files iterated before later
|
||||
// registry-primary skips have been seen carry an inflated denominator, and
|
||||
// the ratio only self-corrects after every file has been classified. Pre-
|
||||
// count runs only on the enabled path so the disabled path stays free of
|
||||
// the extra Map iteration. Defaults to 0 on the disabled path; the live log
|
||||
// gate is also disabled there, so the value is never read.
|
||||
// the ratio only self-corrects after every file has been classified.
|
||||
//
|
||||
// Runs whenever its result will actually be read: on the profile path (the
|
||||
// live deferred-profile log) OR when the always-on slow-file watchdog is
|
||||
// active (#1741) — the watchdog's warning prints `${resolvedFiles}/${resolvedTotal}`
|
||||
// unconditionally, so leaving resolvedTotal at 0 on a plain run produced a
|
||||
// bogus "Resolved N/0 files" denominator on exactly the unprofiled runs the
|
||||
// watchdog exists for. When both gates are off, skip the extra Map pass.
|
||||
let resolvedTotal = 0;
|
||||
if (profileCalls) {
|
||||
if (profileCalls || alwaysSlowFileMs > 0) {
|
||||
for (const filePath of byFile.keys()) {
|
||||
const lang = getLanguageFromFilename(filePath);
|
||||
if (!lang || !isRegistryPrimary(lang)) resolvedTotal++;
|
||||
@@ -2970,6 +2990,9 @@ export const processCallsFromExtracted = async (
|
||||
|
||||
resolvedFiles++;
|
||||
const tFile = startTimer(profileCalls);
|
||||
// Always-on timer (cheap: one hrtime read) feeding the slow-file watchdog
|
||||
// below. Distinct from `tFile`, which is null unless profiling is on.
|
||||
const tFileAlways = alwaysSlowFileMs > 0 ? process.hrtime.bigint() : null;
|
||||
|
||||
if (profileCalls && (resolvedFiles === 1 || resolvedFiles % logEveryN === 0)) {
|
||||
logDeferredProfile(
|
||||
@@ -3143,6 +3166,30 @@ export const processCallsFromExtracted = async (
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
// Always-on slow-file watchdog (#1741) — fires regardless of verbose.
|
||||
if (tFileAlways !== null) {
|
||||
const elapsedAlways = profileElapsedMs(tFileAlways);
|
||||
if (elapsedAlways >= alwaysSlowFileMs) {
|
||||
const now = Date.now();
|
||||
if (now - lastSlowFileWarnAt >= ALWAYS_ON_SLOW_FILE_WARN_THROTTLE_MS) {
|
||||
lastSlowFileWarnAt = now;
|
||||
const suppressedNote =
|
||||
suppressedSlowFileWarnings > 0
|
||||
? ` (+${suppressedSlowFileWarnings} more slow files since the last warning)`
|
||||
: '';
|
||||
logger.warn(
|
||||
`⏳ Call resolution for ${filePath} took ${(elapsedAlways / 1000).toFixed(1)}s ` +
|
||||
`(${calls.length} call sites, ${fileLanguage ?? 'unknown'}). The run is not frozen — ` +
|
||||
`this file is unusually expensive to resolve. Resolved ${resolvedFiles}/${resolvedTotal} ` +
|
||||
`files so far.${suppressedNote} Pass -v for per-file deferred-resolution timing.`,
|
||||
);
|
||||
suppressedSlowFileWarnings = 0;
|
||||
} else {
|
||||
suppressedSlowFileWarnings++;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if (profileCalls) {
|
||||
|
||||
@@ -25,6 +25,8 @@ import {
|
||||
import { expandCopies } from './cobol/cobol-copy-expander.js';
|
||||
import { processJclFiles } from './cobol/jcl-processor.js';
|
||||
|
||||
import { logger } from '../logger.js';
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// File detection
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -60,6 +62,7 @@ export interface CobolProcessResult {
|
||||
sets: number;
|
||||
inspects: number;
|
||||
initializes: number;
|
||||
arithmeticOps: number;
|
||||
}
|
||||
|
||||
/** Returns true if the file is a COBOL or copybook file. */
|
||||
@@ -114,6 +117,7 @@ export const processCobol = (
|
||||
sets: 0,
|
||||
inspects: 0,
|
||||
initializes: 0,
|
||||
arithmeticOps: 0,
|
||||
};
|
||||
|
||||
// ── 1. Separate programs, copybooks, and JCL ───────────────────────
|
||||
@@ -150,16 +154,40 @@ export const processCobol = (
|
||||
const entry = copybookMap.get(name.toUpperCase());
|
||||
return entry ? entry.path : null;
|
||||
};
|
||||
// Memoize preprocessed copybook content for the duration of this
|
||||
// processCobol call. A single copybook is COPYed by many programs (and at
|
||||
// many COPY sites within a program); without this cache
|
||||
// preprocessCobolSource would re-run once per COPY site —
|
||||
// O(programs × copybooks) preprocessing passes over the same content.
|
||||
// Keyed by the resolved copybook path. REPLACING is applied later by the
|
||||
// expander on the returned (pre-REPLACING) content (see
|
||||
// cobol-copy-expander.ts readFile→applyReplacing), so caching the
|
||||
// pre-REPLACING preprocessed text here is safe and per-call-scoped.
|
||||
const preprocessedCopyCache = new Map<string, string>();
|
||||
const readCopy = (copyPath: string): string | null => {
|
||||
const cached = preprocessedCopyCache.get(copyPath);
|
||||
if (cached !== undefined) return cached;
|
||||
const content = copybookByPath.get(copyPath);
|
||||
return content ? preprocessCobolSource(content) : null;
|
||||
if (!content) return null; // preserves original falsy→null (missing/empty)
|
||||
const preprocessed = preprocessCobolSource(content);
|
||||
preprocessedCopyCache.set(copyPath, preprocessed);
|
||||
return preprocessed;
|
||||
};
|
||||
|
||||
// Track module names for cross-program CALL resolution
|
||||
const moduleNodeIds = new Map<string, string>(); // uppercase program name -> node id
|
||||
|
||||
// ── 3. Process each COBOL program ──────────────────────────────────
|
||||
const raw = parseInt(process.env.GITNEXUS_MAX_COBOL_FILE_SIZE_BYTES ?? '', 10);
|
||||
const MAX_COBOL_FILE_SIZE = Number.isFinite(raw) && raw > 0 ? raw : 5 * 1024 * 1024;
|
||||
for (const file of programs) {
|
||||
// File-size guard: skip excessively large files to prevent OOM
|
||||
if (file.content.length > MAX_COBOL_FILE_SIZE) {
|
||||
logger.warn(
|
||||
`[cobol-processor] Skipping oversized file (${(file.content.length / 1024 / 1024).toFixed(1)}MB > ${(MAX_COBOL_FILE_SIZE / 1024 / 1024).toFixed(0)}MB): ${file.path}`,
|
||||
);
|
||||
continue;
|
||||
}
|
||||
const fileNodeId = generateId('File', file.path);
|
||||
// Skip if file node doesn't exist (structure-processor creates it)
|
||||
if (!graph.getNode(fileNodeId)) continue;
|
||||
@@ -199,6 +227,7 @@ export const processCobol = (
|
||||
result.sets += extracted.sets.length;
|
||||
result.inspects += extracted.inspects.length;
|
||||
result.initializes += extracted.initializes.length;
|
||||
result.arithmeticOps += extracted.arithmeticOps.length;
|
||||
}
|
||||
|
||||
// ── 4. Second pass: resolve cross-program CALL targets ─────────────
|
||||
@@ -1211,7 +1240,9 @@ function mapToGraph(
|
||||
|
||||
// ── MOVE data flow -> ACCESSES edges (read/write) ──────────────
|
||||
for (const move of extracted.moves) {
|
||||
const fromPropId = dataItemMap.get(move.from.toUpperCase());
|
||||
// Strip any subscript from the source name for data item lookup
|
||||
const fromBase = stripMoveSubscript(move.from);
|
||||
const fromPropId = dataItemMap.get(fromBase.toUpperCase());
|
||||
const callerId = scopedCallerLookup(move.caller, move.line);
|
||||
|
||||
// One read edge per MOVE (regardless of number of targets)
|
||||
@@ -1228,7 +1259,8 @@ function mapToGraph(
|
||||
|
||||
// One write edge per target
|
||||
for (const target of move.targets) {
|
||||
const toPropId = dataItemMap.get(target.toUpperCase());
|
||||
const toBase = stripMoveSubscript(target);
|
||||
const toPropId = dataItemMap.get(toBase.toUpperCase());
|
||||
if (toPropId) {
|
||||
graph.addRelationship({
|
||||
id: generateId('ACCESSES', `${callerId}->write->${target}:L${move.line}`),
|
||||
@@ -1242,6 +1274,39 @@ function mapToGraph(
|
||||
}
|
||||
}
|
||||
|
||||
// ── Arithmetic operations -> ACCESSES edges ──────────────────
|
||||
for (const arith of extracted.arithmeticOps) {
|
||||
const callerId = scopedCallerLookup(arith.caller, arith.line);
|
||||
// Write edge to target variable
|
||||
const targetBase = stripMoveSubscript(arith.target);
|
||||
const targetPropId = dataItemMap.get(targetBase.toUpperCase());
|
||||
if (targetPropId) {
|
||||
graph.addRelationship({
|
||||
id: generateId('ACCESSES', `${callerId}->arith-write->${arith.target}:L${arith.line}`),
|
||||
type: 'ACCESSES',
|
||||
sourceId: callerId,
|
||||
targetId: targetPropId,
|
||||
confidence: 0.9,
|
||||
reason: 'cobol-arithmetic-write',
|
||||
});
|
||||
}
|
||||
// Read edge for each source operand
|
||||
for (const src of arith.sources) {
|
||||
const srcBase = stripMoveSubscript(src);
|
||||
const srcPropId = dataItemMap.get(srcBase.toUpperCase());
|
||||
if (srcPropId) {
|
||||
graph.addRelationship({
|
||||
id: generateId('ACCESSES', `${callerId}->arith-read->${src}:L${arith.line}`),
|
||||
type: 'ACCESSES',
|
||||
sourceId: callerId,
|
||||
targetId: srcPropId,
|
||||
confidence: 0.9,
|
||||
reason: 'cobol-arithmetic-read',
|
||||
});
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// ── File declarations -> Record nodes ──────────────────────────
|
||||
for (const fd of extracted.fileDeclarations) {
|
||||
const fdId = generateId('Record', `${filePath}:${fd.selectName}`);
|
||||
@@ -1383,6 +1448,11 @@ function mapToGraph(
|
||||
// Helpers
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/** Strip parenthesized subscript/reference-modification suffixes */
|
||||
function stripMoveSubscript(name: string): string {
|
||||
return name.replace(/\([^)]*\)/g, '').trim();
|
||||
}
|
||||
|
||||
/** Find the enclosing program name for a given line number (innermost wins). */
|
||||
function findOwningProgramName(
|
||||
lineNum: number,
|
||||
|
||||
@@ -183,6 +183,19 @@ export interface CobolRegexResults {
|
||||
|
||||
// Phase 4.1: INITIALIZE
|
||||
initializes: Array<{ target: string; line: number; caller: string | null }>;
|
||||
|
||||
// Phase 4.2: Arithmetic operations (COMPUTE, ADD, SUBTRACT, MULTIPLY, DIVIDE)
|
||||
arithmeticOps: Array<{
|
||||
verb: 'COMPUTE' | 'ADD' | 'SUBTRACT' | 'MULTIPLY' | 'DIVIDE';
|
||||
/** Target variable (written to) */
|
||||
target: string;
|
||||
/** Source operand variables (read from) */
|
||||
sources: string[];
|
||||
line: number;
|
||||
caller: string | null;
|
||||
/** For ADD/SUBTRACT/MULTIPLY/DIVIDE with GIVING: the GIVING target */
|
||||
givingTarget?: string;
|
||||
}>;
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -306,36 +319,37 @@ const RE_SECTION =
|
||||
/\b(WORKING-STORAGE|LINKAGE|FILE|LOCAL-STORAGE|SCREEN|INPUT-OUTPUT|CONFIGURATION)\s+SECTION\b/i;
|
||||
|
||||
// IDENTIFICATION DIVISION
|
||||
const RE_PROGRAM_ID = /\bPROGRAM-ID\.\s*([A-Z][A-Z0-9-]*)(?:\s+IS\s+COMMON)?/i;
|
||||
const RE_END_PROGRAM = /\bEND\s+PROGRAM\s+([A-Z][A-Z0-9-]*)\s*\./i;
|
||||
const RE_PROGRAM_ID = /\bPROGRAM-ID\.\s*([A-Z0-9][A-Z0-9-]*)(?:\s+IS\s+COMMON)?/i;
|
||||
const RE_END_PROGRAM = /\bEND\s+PROGRAM\s+([A-Z0-9][A-Z0-9-]*)\s*\./i;
|
||||
const RE_AUTHOR = /^\s+AUTHOR\.\s*(.+)/i;
|
||||
const RE_DATE_WRITTEN = /^\s+DATE-WRITTEN\.\s*(.+)/i;
|
||||
const RE_DATE_COMPILED = /^\s+DATE-COMPILED\.\s*(.+)/i;
|
||||
const RE_INSTALLATION = /^\s+INSTALLATION\.\s*(.+)/i;
|
||||
|
||||
// ENVIRONMENT DIVISION — SELECT
|
||||
const RE_SELECT_START = /\bSELECT\s+(?:OPTIONAL\s+)?([A-Z][A-Z0-9-]+)/i;
|
||||
const RE_SELECT_START = /\bSELECT\s+(?:OPTIONAL\s+)?([A-Z0-9][A-Z0-9-]+)/i;
|
||||
|
||||
// DATA DIVISION
|
||||
// ^\s* (not ^\s+) to support both fixed-format (indented) and free-format (trimmed)
|
||||
const RE_FD = /^\s*(?:FD|SD|RD)\s+([A-Z][A-Z0-9-]+)/i;
|
||||
const RE_DATA_ITEM = /^\s*(\d{1,2})\s+([A-Z][A-Z0-9-]+)\s*(.*)/i;
|
||||
const RE_ANONYMOUS_REDEFINES = /^\s*(\d{1,2})\s+REDEFINES\s+([A-Z][A-Z0-9-]+)/i;
|
||||
const RE_88_LEVEL = /^\s*88\s+([A-Z][A-Z0-9-]+)\s+VALUES?\s+(?:ARE\s+)?(.+)/i;
|
||||
const RE_FD = /^\s*(?:FD|SD|RD)\s+([A-Z0-9][A-Z0-9-]+)/i;
|
||||
const RE_DATA_ITEM = /^\s*(\d{1,2})\s+([A-Z0-9][A-Z0-9-]+)\s*(.*)/i;
|
||||
const RE_ANONYMOUS_REDEFINES = /^\s*(\d{1,2})\s+REDEFINES\s+([A-Z0-9][A-Z0-9-]+)/i;
|
||||
const RE_88_LEVEL = /^\s*88\s+([A-Z0-9][A-Z0-9-]+)\s+VALUES?\s+(?:ARE\s+)?(.+)/i;
|
||||
|
||||
// PROCEDURE DIVISION
|
||||
// These patterns support both fixed-format (7 leading spaces) and free-format (any indentation)
|
||||
const RE_PROC_SECTION = /^\s*([A-Z][A-Z0-9-]+)\s+SECTION(?:\s+\d+)?\.\s*$/i;
|
||||
const RE_PROC_PARAGRAPH = /^\s*([A-Z][A-Z0-9-]+)\.\s*$/i;
|
||||
const RE_PERFORM = /\bPERFORM\s+([A-Z][A-Z0-9-]+)(?:\s+(?:THRU|THROUGH)\s+([A-Z][A-Z0-9-]+))?/gi;
|
||||
const RE_PROC_SECTION = /^\s*([A-Z0-9][A-Z0-9-]+)\s+SECTION(?:\s+\d+)?\.\s*$/i;
|
||||
const RE_PROC_PARAGRAPH = /^\s*([A-Z0-9][A-Z0-9-]+)\.\s*$/i;
|
||||
const RE_PERFORM =
|
||||
/\bPERFORM\s+([A-Z0-9][A-Z0-9-]+)(?:\s+(?:THRU|THROUGH)\s+([A-Z0-9][A-Z0-9-]+))?/gi;
|
||||
|
||||
// ALL DIVISIONS
|
||||
// Both double-quoted ("PROG") and single-quoted ('PROG') targets are valid COBOL.
|
||||
// Use separate alternation groups so quotes must match (prevents "PROG' false-matches).
|
||||
const RE_CALL = /\bCALL\s+(?:"([^"]+)"|'([^']+)')/gi;
|
||||
// Dynamic CALL via data item (no quotes): CALL WS-PROGRAM-NAME
|
||||
const RE_CALL_DYNAMIC = /(?<![A-Z0-9-])\bCALL\s+([A-Z][A-Z0-9-]+)(?=\s|\.|$)/gi;
|
||||
const RE_COPY_UNQUOTED = /\bCOPY\s+([A-Z][A-Z0-9-]+)(?:\s|\.)/i;
|
||||
const RE_CALL_DYNAMIC = /(?<![A-Z0-9-])\bCALL\s+([A-Z0-9][A-Z0-9-]+)(?=\s|\.|$)/gi;
|
||||
const RE_COPY_UNQUOTED = /\bCOPY\s+([A-Z0-9][A-Z0-9-]+)(?:\s|\.)/i;
|
||||
const RE_COPY_QUOTED = /\bCOPY\s+(?:"([^"]+)"|'([^']+)')(?:\s|\.)/i;
|
||||
|
||||
// EXEC blocks
|
||||
@@ -346,32 +360,32 @@ const RE_END_EXEC = /\bEND-EXEC\b/i;
|
||||
// GO TO — control flow transfer (same graph semantics as PERFORM)
|
||||
// GO TO — captures first target; GO TO p1 p2 p3 DEPENDING ON x handled below
|
||||
const RE_GOTO =
|
||||
/\bGO\s+TO\s+([A-Z][A-Z0-9-]+(?:\s+[A-Z][A-Z0-9-]+)*?)(?:\s+DEPENDING\s+ON\s+[A-Z][A-Z0-9-]+)?(?:\s*\.|$)/i;
|
||||
/\bGO\s+TO\s+([A-Z0-9][A-Z0-9-]+(?:\s+[A-Z0-9][A-Z0-9-]+)*?)(?:\s+DEPENDING\s+ON\s+[A-Z0-9][A-Z0-9-]+)?(?:\s*\.|$)/i;
|
||||
|
||||
// SORT/MERGE file references
|
||||
const RE_SORT = /\bSORT\s+([A-Z][A-Z0-9-]+)/i;
|
||||
const RE_MERGE = /\bMERGE\s+([A-Z][A-Z0-9-]+)/i;
|
||||
const RE_SORT = /\bSORT\s+([A-Z0-9][A-Z0-9-]+)/i;
|
||||
const RE_MERGE = /\bMERGE\s+([A-Z0-9][A-Z0-9-]+)/i;
|
||||
|
||||
// SEARCH — table access
|
||||
const RE_SEARCH = /\bSEARCH\s+(?:ALL\s+)?([A-Z][A-Z0-9-]+)/i;
|
||||
const RE_SEARCH = /\bSEARCH\s+(?:ALL\s+)?([A-Z0-9][A-Z0-9-]+)/i;
|
||||
|
||||
// CANCEL — program lifecycle
|
||||
const RE_CANCEL = /\bCANCEL\s+(?:"([^"]+)"|'([^']+)')/gi;
|
||||
const RE_CANCEL_DYNAMIC = /(?<![A-Z0-9-])\bCANCEL\s+([A-Z][A-Z0-9-]+)(?=\s|\.|$)/gi;
|
||||
const RE_CANCEL_DYNAMIC = /(?<![A-Z0-9-])\bCANCEL\s+([A-Z0-9][A-Z0-9-]+)(?=\s|\.|$)/gi;
|
||||
|
||||
// Level 66 RENAMES
|
||||
const RE_66_LEVEL = /^\s*66\s+([A-Z][A-Z0-9-]+)\s+RENAMES\s+([A-Z][A-Z0-9-]+)/i;
|
||||
const RE_66_LEVEL = /^\s*66\s+([A-Z0-9][A-Z0-9-]+)\s+RENAMES\s+([A-Z0-9][A-Z0-9-]+)/i;
|
||||
|
||||
// DECLARATIVES boundary and USE AFTER EXCEPTION
|
||||
const RE_DECLARATIVES_START = /^\s*DECLARATIVES\s*\.\s*$/i;
|
||||
const RE_DECLARATIVES_END = /^\s*END\s+DECLARATIVES\s*\.\s*$/i;
|
||||
const RE_USE_AFTER =
|
||||
/\bUSE\s+(?:AFTER\s+)?(?:STANDARD\s+)?(?:EXCEPTION|ERROR)\s+ON\s+([A-Z][A-Z0-9-]+|INPUT|OUTPUT|I-O|EXTEND)\b/i;
|
||||
/\bUSE\s+(?:AFTER\s+)?(?:STANDARD\s+)?(?:EXCEPTION|ERROR)\s+ON\s+([A-Z0-9][A-Z0-9-]+|INPUT|OUTPUT|I-O|EXTEND)\b/i;
|
||||
|
||||
// SET statement (condition, index)
|
||||
//
|
||||
// Catastrophic-backtracking note (CodeQL js/redos): the previous shape
|
||||
// `((?:[A-Z][A-Z0-9-]+(?:\s+OF\s+[A-Z][A-Z0-9-]+)?\s+)+)TO\s+TRUE`
|
||||
// `((?:[A-Z0-9][A-Z0-9-]+(?:\s+OF\s+[A-Z0-9][A-Z0-9-]+)?\s+)+)TO\s+TRUE`
|
||||
// nested `\s+` quantifiers across alternations and was exponential on
|
||||
// inputs like "SET a OF a OF a ... TO TRUE". Replaced with a lazy
|
||||
// dot-match bounded by the explicit `\s+TO\s+TRUE` suffix — `.+?` is
|
||||
@@ -382,7 +396,7 @@ const RE_USE_AFTER =
|
||||
// pathological-input timing assertion exercises the production regex
|
||||
// instead of an inline copy that drifts.
|
||||
export const RE_SET_TO_TRUE = /\bSET\s+(.+?)\s+TO\s+TRUE\b/i;
|
||||
export const RE_SET_INDEX = /\bSET\s+(.+?)\s+(TO|UP\s+BY|DOWN\s+BY)\s+(\d+|[A-Z][A-Z0-9-]+)/i;
|
||||
export const RE_SET_INDEX = /\bSET\s+(.+?)\s+(TO|UP\s+BY|DOWN\s+BY)\s+(\d+|[A-Z0-9][A-Z0-9-]+)/i;
|
||||
|
||||
// INITIALIZE statement — data reset (captures targets before REPLACING/WITH clause)
|
||||
const RE_INITIALIZE = /\bINITIALIZE\s+([\s\S]*?)(?=\bREPLACING\b|\bWITH\b|\.\s*$|$)/i;
|
||||
@@ -409,7 +423,8 @@ const RE_PROC_USING = /\bPROCEDURE\s+DIVISION\s+USING\s+([\s\S]*?)(?:\.|$)/i;
|
||||
const RE_ENTRY = /\bENTRY\s+(?:"([^"]+)"|'([^']+)')(?:\s+USING\s+([\s\S]*?))?(?:\.|$)/i;
|
||||
|
||||
// MOVE statement — captures everything after TO for multi-target extraction
|
||||
const RE_MOVE = /\bMOVE\s+((?:CORRESPONDING|CORR)\s+)?([A-Z][A-Z0-9-]+)\s+TO\s+(.+)/i;
|
||||
const RE_MOVE =
|
||||
/\bMOVE\s+((?:CORRESPONDING|CORR)\s+)?([A-Z0-9][A-Z0-9-]+(?:\([^)]*\))*)\s+TO\s+(.+)/i;
|
||||
const MOVE_SKIP = new Set([
|
||||
'SPACES',
|
||||
'ZEROS',
|
||||
@@ -449,7 +464,7 @@ function extractMoveTargets(afterTo: string): string[] {
|
||||
skipNext = true;
|
||||
continue;
|
||||
}
|
||||
if (/^[A-Z][A-Z0-9-]+$/i.test(token) && !MOVE_SKIP.has(token.toUpperCase())) {
|
||||
if (/^[A-Z0-9][A-Z0-9-]+$/i.test(token) && !MOVE_SKIP.has(token.toUpperCase())) {
|
||||
targets.push(token);
|
||||
}
|
||||
}
|
||||
@@ -605,14 +620,14 @@ function parseDataItemClauses(rest: string): {
|
||||
}
|
||||
|
||||
// REDEFINES <name>
|
||||
const redefMatch = text.match(/\bREDEFINES\s+([A-Z][A-Z0-9-]+)/i);
|
||||
const redefMatch = text.match(/\bREDEFINES\s+([A-Z0-9][A-Z0-9-]+)/i);
|
||||
if (redefMatch) {
|
||||
result.redefines = redefMatch[1];
|
||||
}
|
||||
|
||||
// OCCURS <n> [TO <m>] [TIMES] [DEPENDING ON <field>]
|
||||
const occursMatch = text.match(
|
||||
/\bOCCURS\s+(\d+)(?:\s+TO\s+(\d+))?\s*(?:TIMES\s*)?(?:DEPENDING\s+ON\s+([A-Z][A-Z0-9-]+(?:\s*\([^)]*\))?))?/i,
|
||||
/\bOCCURS\s+(\d+)(?:\s+TO\s+(\d+))?\s*(?:TIMES\s*)?(?:DEPENDING\s+ON\s+([A-Z0-9][A-Z0-9-]+(?:\s*\([^)]*\))?))?/i,
|
||||
);
|
||||
if (occursMatch) {
|
||||
result.occurs = parseInt(occursMatch[1], 10);
|
||||
@@ -653,7 +668,7 @@ function parseDataItemClauses(rest: string): {
|
||||
result.value = numMatch[1];
|
||||
} else {
|
||||
// Try figurative constant or identifier
|
||||
const identMatch = afterValue.match(/^([A-Z][A-Z0-9-]*)/i);
|
||||
const identMatch = afterValue.match(/^([A-Z0-9][A-Z0-9-]*)/i);
|
||||
if (identMatch) result.value = identMatch[1].toUpperCase();
|
||||
}
|
||||
}
|
||||
@@ -721,7 +736,7 @@ function parseSelectStatement(stmt: string, startLine: number): FileDeclaration
|
||||
// Normalize whitespace
|
||||
const text = stmt.replace(/\s+/g, ' ').trim();
|
||||
|
||||
const nameMatch = text.match(/^SELECT\s+(?:OPTIONAL\s+)?([A-Z][A-Z0-9-]+)/i);
|
||||
const nameMatch = text.match(/^SELECT\s+(?:OPTIONAL\s+)?([A-Z0-9][A-Z0-9-]+)/i);
|
||||
if (!nameMatch) return null;
|
||||
|
||||
const result: FileDeclaration = {
|
||||
@@ -730,7 +745,7 @@ function parseSelectStatement(stmt: string, startLine: number): FileDeclaration
|
||||
line: startLine,
|
||||
};
|
||||
|
||||
const assignMatch = text.match(/\bASSIGN\s+(?:TO\s+)?("([^"]+)"|([A-Z][A-Z0-9-]*))/i);
|
||||
const assignMatch = text.match(/\bASSIGN\s+(?:TO\s+)?("([^"]+)"|([A-Z0-9][A-Z0-9-]*))/i);
|
||||
if (assignMatch) {
|
||||
result.assignTo = assignMatch[2] || assignMatch[3] || '';
|
||||
}
|
||||
@@ -747,19 +762,21 @@ function parseSelectStatement(stmt: string, startLine: number): FileDeclaration
|
||||
result.access = accessMatch[1].toUpperCase();
|
||||
}
|
||||
|
||||
const keyMatch = text.match(/\bRECORD\s+KEY\s+(?:IS\s+)?([A-Z][A-Z0-9-]+)/i);
|
||||
const keyMatch = text.match(/\bRECORD\s+KEY\s+(?:IS\s+)?([A-Z0-9][A-Z0-9-]+)/i);
|
||||
if (keyMatch) {
|
||||
result.recordKey = keyMatch[1];
|
||||
}
|
||||
|
||||
// ALTERNATE RECORD KEY
|
||||
const altKeyMatches = text.matchAll(/\bALTERNATE\s+RECORD\s+KEY\s+(?:IS\s+)?([A-Z][A-Z0-9-]+)/gi);
|
||||
const altKeyMatches = text.matchAll(
|
||||
/\bALTERNATE\s+RECORD\s+KEY\s+(?:IS\s+)?([A-Z0-9][A-Z0-9-]+)/gi,
|
||||
);
|
||||
const alternateKeys: string[] = [];
|
||||
for (const m of altKeyMatches) alternateKeys.push(m[1]);
|
||||
if (alternateKeys.length > 0) result.alternateKeys = alternateKeys;
|
||||
|
||||
// FILE STATUS IS / STATUS IS
|
||||
const statusMatch = text.match(/\b(?:FILE\s+)?STATUS\s+(?:IS\s+)?([A-Z][A-Z0-9-]+)/i);
|
||||
const statusMatch = text.match(/\b(?:FILE\s+)?STATUS\s+(?:IS\s+)?([A-Z0-9][A-Z0-9-]+)/i);
|
||||
if (statusMatch) {
|
||||
result.fileStatus = statusMatch[1];
|
||||
}
|
||||
@@ -789,11 +806,12 @@ function parseExecSqlBlock(
|
||||
block: string,
|
||||
line: number,
|
||||
): CobolRegexResults['execSqlBlocks'][number] {
|
||||
// Strip EXEC SQL ... END-EXEC wrapper
|
||||
// Strip EXEC SQL ... END-EXEC wrapper and trailing period
|
||||
const body = block
|
||||
.replace(/\bEXEC\s+SQL\b/i, '')
|
||||
.replace(/\bEND-EXEC\b/i, '')
|
||||
.replace(/\s+/g, ' ')
|
||||
.replace(/\.\s*$/, '')
|
||||
.trim();
|
||||
|
||||
// Determine operation from first SQL keyword
|
||||
@@ -823,18 +841,24 @@ function parseExecSqlBlock(
|
||||
// Extract table names from FROM, INTO (INSERT), UPDATE, DELETE FROM, JOIN
|
||||
const tables: string[] = [];
|
||||
const tablePatterns = [
|
||||
/\bFROM\s+([A-Z][A-Z0-9_]+)/gi,
|
||||
/\bINSERT\s+INTO\s+([A-Z][A-Z0-9_]+)/gi,
|
||||
/\bUPDATE\s+([A-Z][A-Z0-9_]+)/gi,
|
||||
/\bJOIN\s+([A-Z][A-Z0-9_]+)/gi,
|
||||
// FROM table1 [AS alias], table2 [AS alias] … — handle comma-separated
|
||||
// lists with optional AS keyword, terminated by SQL clause keywords
|
||||
// (WHERE, JOIN, GROUP, ON, ORDER, HAVING, UNION, SET, INTO, VALUES, FETCH, FOR, LIMIT, OFFSET, WITH).
|
||||
/\bFROM\s+([A-Z0-9][A-Z0-9_]+(?:\s+(?:AS\s+)?[A-Z0-9][A-Z0-9_]*)?(?:\s*,\s*[A-Z0-9][A-Z0-9_]+(?:\s+(?:AS\s+)?[A-Z0-9][A-Z0-9_]*)?)*)(?:\s+(?:WHERE|JOIN|GROUP|ON|ORDER|HAVING|UNION|SET|INTO|VALUES|FETCH|FOR|LIMIT|OFFSET|WITH)\b|$)/gi,
|
||||
/\bINSERT\s+INTO\s+([A-Z0-9][A-Z0-9_]+)/gi,
|
||||
/\bUPDATE\s+([A-Z0-9][A-Z0-9_]+)/gi,
|
||||
/\bJOIN\s+([A-Z0-9][A-Z0-9_]+)/gi,
|
||||
];
|
||||
for (const re of tablePatterns) {
|
||||
let m: RegExpExecArray | null;
|
||||
while ((m = re.exec(body)) !== null) {
|
||||
const name = m[1].toUpperCase();
|
||||
// Skip host variables and SQL keywords
|
||||
if (!name.startsWith(':') && !tables.includes(name)) {
|
||||
tables.push(name);
|
||||
// Split comma-separated table list and strip aliases
|
||||
const names = m[1].split(',').map((n) => n.trim().split(/\s+/)[0].toUpperCase());
|
||||
for (const name of names) {
|
||||
// Skip host variables and SQL keywords
|
||||
if (!name.startsWith(':') && !tables.includes(name)) {
|
||||
tables.push(name);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -849,7 +873,7 @@ function parseExecSqlBlock(
|
||||
|
||||
// Extract host variables: :VARIABLE-NAME (strip the colon)
|
||||
const hostVariables: string[] = [];
|
||||
const hostRe = /:([A-Z][A-Z0-9-]+)/gi;
|
||||
const hostRe = /:([A-Z0-9][A-Z0-9-]+)/gi;
|
||||
let hm: RegExpExecArray | null;
|
||||
while ((hm = hostRe.exec(body)) !== null) {
|
||||
const name = hm[1];
|
||||
@@ -910,24 +934,24 @@ function parseExecCicsBlock(
|
||||
const result: CobolRegexResults['execCicsBlocks'][number] = { line, command };
|
||||
|
||||
// MAP name: MAP('name') or MAP("name") or MAP(IDENTIFIER)
|
||||
const mapMatch = body.match(/\bMAP\s*\(\s*(?:['"]([^'"]+)['"]|([A-Z][A-Z0-9-]+))\s*\)/i);
|
||||
const mapMatch = body.match(/\bMAP\s*\(\s*(?:['"]([^'"]+)['"]|([A-Z0-9][A-Z0-9-]+))\s*\)/i);
|
||||
if (mapMatch) result.mapName = mapMatch[1] ?? mapMatch[2];
|
||||
|
||||
// PROGRAM name: PROGRAM('name') or PROGRAM("name") or PROGRAM(VARIABLE)
|
||||
const progMatch = body.match(/\bPROGRAM\s*\(\s*(?:['"]([^'"]+)['"]|([A-Z][A-Z0-9-]+))\s*\)/i);
|
||||
const progMatch = body.match(/\bPROGRAM\s*\(\s*(?:['"]([^'"]+)['"]|([A-Z0-9][A-Z0-9-]+))\s*\)/i);
|
||||
if (progMatch) {
|
||||
result.programName = progMatch[1] ?? progMatch[2];
|
||||
result.programIsLiteral = !!progMatch[1];
|
||||
}
|
||||
|
||||
// TRANSID: TRANSID('name') or TRANSID("name") or TRANSID(VARIABLE)
|
||||
const transMatch = body.match(/\bTRANSID\s*\(\s*(?:['"]([^'"]+)['"]|([A-Z][A-Z0-9-]+))\s*\)/i);
|
||||
const transMatch = body.match(/\bTRANSID\s*\(\s*(?:['"]([^'"]+)['"]|([A-Z0-9][A-Z0-9-]+))\s*\)/i);
|
||||
if (transMatch) result.transId = transMatch[1] ?? transMatch[2];
|
||||
|
||||
// FILE/DATASET: FILE('name') or DATASET('name') or FILE(VARIABLE)
|
||||
// Used in CICS READ, WRITE, REWRITE, DELETE, STARTBR, READNEXT, READPREV, ENDBR
|
||||
const fileMatch = body.match(
|
||||
/\b(?:FILE|DATASET)\s*\(\s*(?:['"]([^'"]+)['"]|([A-Z][A-Z0-9-]+))\s*\)/i,
|
||||
/\b(?:FILE|DATASET)\s*\(\s*(?:['"]([^'"]+)['"]|([A-Z0-9][A-Z0-9-]+))\s*\)/i,
|
||||
);
|
||||
if (fileMatch) {
|
||||
result.fileName = fileMatch[1] ?? fileMatch[2];
|
||||
@@ -935,19 +959,19 @@ function parseExecCicsBlock(
|
||||
}
|
||||
|
||||
// QUEUE: QUEUE('name') — used in WRITEQ/READQ TS/TD
|
||||
const queueMatch = body.match(/\bQUEUE\s*\(\s*(?:['"]([^'"]+)['"]|([A-Z][A-Z0-9-]+))\s*\)/i);
|
||||
const queueMatch = body.match(/\bQUEUE\s*\(\s*(?:['"]([^'"]+)['"]|([A-Z0-9][A-Z0-9-]+))\s*\)/i);
|
||||
if (queueMatch) result.queueName = queueMatch[1] ?? queueMatch[2];
|
||||
|
||||
// HANDLE ABEND LABEL(paragraph-name) — error handler target
|
||||
const labelMatch = body.match(/\bLABEL\s*\(\s*([A-Z][A-Z0-9-]+)\s*\)/i);
|
||||
const labelMatch = body.match(/\bLABEL\s*\(\s*([A-Z0-9][A-Z0-9-]+)\s*\)/i);
|
||||
if (labelMatch) result.labelName = labelMatch[1];
|
||||
|
||||
// INTO(data-area) — data target (READ INTO, RECEIVE INTO, RETRIEVE INTO, READQ INTO)
|
||||
const intoMatch = body.match(/\bINTO\s*\(\s*([A-Z][A-Z0-9-]+)\s*\)/i);
|
||||
const intoMatch = body.match(/\bINTO\s*\(\s*([A-Z0-9][A-Z0-9-]+)\s*\)/i);
|
||||
if (intoMatch) result.intoField = intoMatch[1];
|
||||
|
||||
// FROM(data-area) — data source (WRITE FROM, SEND FROM, WRITEQ FROM, START FROM)
|
||||
const fromMatch = body.match(/\bFROM\s*\(\s*([A-Z][A-Z0-9-]+)\s*\)/i);
|
||||
const fromMatch = body.match(/\bFROM\s*\(\s*([A-Z0-9][A-Z0-9-]+)\s*\)/i);
|
||||
if (fromMatch) result.fromField = fromMatch[1];
|
||||
|
||||
return result;
|
||||
@@ -972,16 +996,16 @@ function parseExecDliBlock(
|
||||
const pcbMatch = body.match(/\bUSING\s+PCB\s*\(\s*(\d+)\s*\)/i);
|
||||
if (pcbMatch) result.pcbNumber = parseInt(pcbMatch[1], 10);
|
||||
|
||||
const segMatch = body.match(/\bSEGMENT\s*\(\s*([A-Z][A-Z0-9-]*)\s*\)/i);
|
||||
const segMatch = body.match(/\bSEGMENT\s*\(\s*([A-Z0-9][A-Z0-9-]*)\s*\)/i);
|
||||
if (segMatch) result.segmentName = segMatch[1];
|
||||
|
||||
const intoMatch = body.match(/\bINTO\s*\(\s*([A-Z][A-Z0-9-]+)\s*\)/i);
|
||||
const intoMatch = body.match(/\bINTO\s*\(\s*([A-Z0-9][A-Z0-9-]+)\s*\)/i);
|
||||
if (intoMatch) result.intoField = intoMatch[1];
|
||||
|
||||
const fromMatch = body.match(/\bFROM\s*\(\s*([A-Z][A-Z0-9-]+)\s*\)/i);
|
||||
const fromMatch = body.match(/\bFROM\s*\(\s*([A-Z0-9][A-Z0-9-]+)\s*\)/i);
|
||||
if (fromMatch) result.fromField = fromMatch[1];
|
||||
|
||||
const psbMatch = body.match(/\bPSB\s*\(\s*([A-Z][A-Z0-9-]+)\s*\)/i);
|
||||
const psbMatch = body.match(/\bPSB\s*\(\s*([A-Z0-9][A-Z0-9-]+)\s*\)/i);
|
||||
if (psbMatch) result.psbName = psbMatch[1];
|
||||
|
||||
return result;
|
||||
@@ -1028,6 +1052,7 @@ export function extractCobolSymbolsWithRegex(
|
||||
sets: [],
|
||||
inspects: [],
|
||||
initializes: [],
|
||||
arithmeticOps: [],
|
||||
};
|
||||
|
||||
// --- State ---
|
||||
@@ -1451,7 +1476,7 @@ export function extractCobolSymbolsWithRegex(
|
||||
}
|
||||
} else if (
|
||||
currentDivision === 'procedure' &&
|
||||
/(?<![A-Z0-9-])\bCALL\s+(?:"[^"]+"|'[^']+'|[A-Z][A-Z0-9-]+)/i.test(line)
|
||||
/(?<![A-Z0-9-])\bCALL\s+(?:"[^"]+"|'[^']+'|[A-Z0-9][A-Z0-9-]+)/i.test(line)
|
||||
) {
|
||||
// Check if this is a complete single-line CALL (ends with period or END-CALL)
|
||||
if (/\.\s*$/.test(line) || /\bEND-CALL\b/i.test(line)) {
|
||||
@@ -1587,7 +1612,9 @@ export function extractCobolSymbolsWithRegex(
|
||||
.trim()
|
||||
.split(/\s+/)
|
||||
.map((f) => f.replace(/\.$/, ''))
|
||||
.filter((f) => /^[A-Z][A-Z0-9-]+$/i.test(f) && !SORT_CLAUSE_NOISE.has(f.toUpperCase())),
|
||||
.filter(
|
||||
(f) => /^[A-Z0-9][A-Z0-9-]+$/i.test(f) && !SORT_CLAUSE_NOISE.has(f.toUpperCase()),
|
||||
),
|
||||
);
|
||||
}
|
||||
if (givingIdx >= 0) {
|
||||
@@ -1597,16 +1624,18 @@ export function extractCobolSymbolsWithRegex(
|
||||
.trim()
|
||||
.split(/\s+/)
|
||||
.map((f) => f.replace(/\.$/, ''))
|
||||
.filter((f) => /^[A-Z][A-Z0-9-]+$/i.test(f) && !SORT_CLAUSE_NOISE.has(f.toUpperCase())),
|
||||
.filter(
|
||||
(f) => /^[A-Z0-9][A-Z0-9-]+$/i.test(f) && !SORT_CLAUSE_NOISE.has(f.toUpperCase()),
|
||||
),
|
||||
);
|
||||
}
|
||||
// INPUT PROCEDURE IS / OUTPUT PROCEDURE IS → control-flow targets (like PERFORM)
|
||||
// Supports optional THRU/THROUGH range: INPUT PROCEDURE IS proc-start THRU proc-end
|
||||
const inputProcMatch = fullSort.match(
|
||||
/\bINPUT\s+PROCEDURE\s+(?:IS\s+)?([A-Z][A-Z0-9-]+)(?:\s+(?:THRU|THROUGH)\s+([A-Z][A-Z0-9-]+))?/i,
|
||||
/\bINPUT\s+PROCEDURE\s+(?:IS\s+)?([A-Z0-9][A-Z0-9-]+)(?:\s+(?:THRU|THROUGH)\s+([A-Z0-9][A-Z0-9-]+))?/i,
|
||||
);
|
||||
const outputProcMatch = fullSort.match(
|
||||
/\bOUTPUT\s+PROCEDURE\s+(?:IS\s+)?([A-Z][A-Z0-9-]+)(?:\s+(?:THRU|THROUGH)\s+([A-Z][A-Z0-9-]+))?/i,
|
||||
/\bOUTPUT\s+PROCEDURE\s+(?:IS\s+)?([A-Z0-9][A-Z0-9-]+)(?:\s+(?:THRU|THROUGH)\s+([A-Z0-9][A-Z0-9-]+))?/i,
|
||||
);
|
||||
if (inputProcMatch) {
|
||||
result.performs.push({
|
||||
@@ -1632,7 +1661,7 @@ export function extractCobolSymbolsWithRegex(
|
||||
function flushInspect(): void {
|
||||
if (inspectAccum === null) return;
|
||||
const text = inspectAccum;
|
||||
const fieldMatch = text.match(/\bINSPECT\s+([A-Z][A-Z0-9-]+)/i);
|
||||
const fieldMatch = text.match(/\bINSPECT\s+([A-Z0-9][A-Z0-9-]+)/i);
|
||||
if (!fieldMatch) {
|
||||
inspectAccum = null;
|
||||
return;
|
||||
@@ -1643,7 +1672,7 @@ export function extractCobolSymbolsWithRegex(
|
||||
/\bTALLYING\b([\s\S]+?)(?:\bREPLACING\b|\bCONVERTING\b|\.\s*$)/i,
|
||||
);
|
||||
if (tallySection) {
|
||||
const counterRe = /([A-Z][A-Z0-9-]+)\s+FOR\b/gi;
|
||||
const counterRe = /([A-Z0-9][A-Z0-9-]+)\s+FOR\b/gi;
|
||||
let cm: RegExpExecArray | null;
|
||||
while ((cm = counterRe.exec(tallySection[1])) !== null) {
|
||||
counters.push(cm[1]);
|
||||
@@ -1693,10 +1722,10 @@ export function extractCobolSymbolsWithRegex(
|
||||
(s) =>
|
||||
s.length > 0 &&
|
||||
!CALL_USING_FILTER.has(s.toUpperCase()) &&
|
||||
/^[A-Z][A-Z0-9-]+$/i.test(s),
|
||||
/^[A-Z0-9][A-Z0-9-]+$/i.test(s),
|
||||
)
|
||||
: undefined;
|
||||
const retMatch = afterCall.match(/\bRETURNING\s+([A-Z][A-Z0-9-]+)/i);
|
||||
const retMatch = afterCall.match(/\bRETURNING\s+([A-Z0-9][A-Z0-9-]+)/i);
|
||||
const returning = retMatch ? retMatch[1] : undefined;
|
||||
result.calls.push({
|
||||
target: callTarget,
|
||||
@@ -1720,10 +1749,10 @@ export function extractCobolSymbolsWithRegex(
|
||||
(s) =>
|
||||
s.length > 0 &&
|
||||
!CALL_USING_FILTER.has(s.toUpperCase()) &&
|
||||
/^[A-Z][A-Z0-9-]+$/i.test(s),
|
||||
/^[A-Z0-9][A-Z0-9-]+$/i.test(s),
|
||||
)
|
||||
: undefined;
|
||||
const dynRetMatch = afterDynCall.match(/\bRETURNING\s+([A-Z][A-Z0-9-]+)/i);
|
||||
const dynRetMatch = afterDynCall.match(/\bRETURNING\s+([A-Z0-9][A-Z0-9-]+)/i);
|
||||
const dynReturning = dynRetMatch ? dynRetMatch[1] : undefined;
|
||||
result.calls.push({
|
||||
target: dynCallMatch[1],
|
||||
@@ -1933,11 +1962,14 @@ export function extractCobolSymbolsWithRegex(
|
||||
const target = perfMatch[1];
|
||||
// Skip COBOL inline-perform keywords that are not paragraph names
|
||||
if (!PERFORM_KEYWORD_SKIP.has(target.toUpperCase())) {
|
||||
// Also check for "PERFORM identifier TIMES" — the identifier is a
|
||||
// data item count, not a paragraph name (fundamental regex ambiguity).
|
||||
const matchEnd = perfMatch.index! + perfMatch[0].length;
|
||||
const afterTarget = line.substring(matchEnd).trim();
|
||||
if (!/^TIMES\b/i.test(afterTarget)) {
|
||||
// Check for inline PERFORM ... TIMES pattern where the target IS
|
||||
// the counter variable itself (e.g., PERFORM WS-COUNT TIMES).
|
||||
// Out-of-line PERFORM target count TIMES (e.g., PERFORM 2000-PROCESS 3 TIMES)
|
||||
// IS a real paragraph call — do NOT suppress it.
|
||||
const hasTimesClause = /^\s*TIMES\b/i.test(afterTarget);
|
||||
if (!hasTimesClause) {
|
||||
result.performs.push({
|
||||
caller: currentParagraph,
|
||||
target,
|
||||
@@ -1976,7 +2008,7 @@ export function extractCobolSymbolsWithRegex(
|
||||
// MOVE CORRESPONDING is always single-target per COBOL standard
|
||||
const targets = isCorresponding
|
||||
? [moveMatch[3].replace(/\..*$/, '').trim().split(/\s+/)[0]].filter((t) =>
|
||||
/^[A-Z][A-Z0-9-]+$/i.test(t),
|
||||
/^[A-Z0-9][A-Z0-9-]+$/i.test(t),
|
||||
)
|
||||
: extractMoveTargets(moveMatch[3]);
|
||||
|
||||
@@ -1992,13 +2024,169 @@ export function extractCobolSymbolsWithRegex(
|
||||
}
|
||||
}
|
||||
|
||||
// Arithmetic statements — COMPUTE, ADD, SUBTRACT, MULTIPLY, DIVIDE
|
||||
// All extract target (written) and source operands (read) for ACCESSES edges
|
||||
// Mask quoted strings before matching to avoid false positives from
|
||||
// arithmetic keywords inside string literals (e.g., DISPLAY "COMPUTE").
|
||||
const lineForArith = line.replace(/"[^"]*"/g, ' ').replace(/'[^']*'/g, ' ');
|
||||
const arithMatch = lineForArith.match(/\b(COMPUTE|ADD|SUBTRACT|MULTIPLY|DIVIDE)\s+(.+)/i);
|
||||
if (arithMatch) {
|
||||
const verb = arithMatch[1].toUpperCase() as
|
||||
| 'COMPUTE'
|
||||
| 'ADD'
|
||||
| 'SUBTRACT'
|
||||
| 'MULTIPLY'
|
||||
| 'DIVIDE';
|
||||
const rest = arithMatch[2].replace(/\..*$/, '').trim();
|
||||
let target = '';
|
||||
const sources: string[] = [];
|
||||
let givingTarget: string | undefined;
|
||||
|
||||
switch (verb) {
|
||||
case 'COMPUTE': {
|
||||
// COMPUTE target = expression
|
||||
const eqIdx = rest.indexOf('=');
|
||||
if (eqIdx > 0) {
|
||||
target = rest.substring(0, eqIdx).trim().split(/\s+/)[0] || '';
|
||||
const expr = rest.substring(eqIdx + 1).trim();
|
||||
// Extract identifiers from expression (skip literals and operators)
|
||||
const idRe = /[A-Z0-9][A-Z0-9-]*/gi;
|
||||
let idMatch: RegExpExecArray | null;
|
||||
while ((idMatch = idRe.exec(expr)) !== null) {
|
||||
const name = idMatch[0];
|
||||
if (
|
||||
!/^(?:AND|OR|NOT|IN|OF|BY|TO|FROM|DIVIDED|INTO|GIVING|TIMES|PLUS|MINUS|MULTIPLIED)$/i.test(
|
||||
name,
|
||||
)
|
||||
) {
|
||||
if (!sources.includes(name)) sources.push(name);
|
||||
}
|
||||
}
|
||||
}
|
||||
break;
|
||||
}
|
||||
case 'ADD': {
|
||||
// ADD a TO b [GIVING c] — target is after TO or GIVING.
|
||||
// If no TO, try GIVING directly (ADD a GIVING b).
|
||||
const addGiving = rest.match(
|
||||
/\bTO\s+([A-Z0-9][A-Z0-9-]+)(?:\s+GIVING\s+([A-Z0-9][A-Z0-9-]+))?/i,
|
||||
);
|
||||
if (addGiving) {
|
||||
target = addGiving[2] ?? addGiving[1];
|
||||
if (addGiving[2]) givingTarget = addGiving[2];
|
||||
// Everything before TO is sources
|
||||
const beforeTo = rest.substring(0, rest.toUpperCase().indexOf(' TO '));
|
||||
beforeTo.replace(/\b([A-Z0-9][A-Z0-9-]+)\b/gi, (m: string) => {
|
||||
if (!/^(?:ADD|CORRESPONDING|CORR)$/i.test(m) && !sources.includes(m)) {
|
||||
sources.push(m);
|
||||
}
|
||||
return m;
|
||||
});
|
||||
// Non-GIVING ADD A TO B: the TO operand (B) is both read and written
|
||||
// (the existing value is read, added, then stored back). Add B as a
|
||||
// source so both ACCESSES edges are created.
|
||||
if (!addGiving[2]) {
|
||||
if (!sources.includes(target)) sources.push(target);
|
||||
}
|
||||
} else {
|
||||
// No TO — try GIVING directly: ADD a GIVING b
|
||||
const addOnlyGiving = rest.match(/\bGIVING\s+([A-Z0-9][A-Z0-9-]+)/i);
|
||||
if (addOnlyGiving) {
|
||||
target = addOnlyGiving[1];
|
||||
givingTarget = addOnlyGiving[1];
|
||||
// Everything before GIVING is sources
|
||||
const beforeGiving = rest.substring(0, rest.toUpperCase().indexOf(' GIVING '));
|
||||
beforeGiving.replace(/\b([A-Z0-9][A-Z0-9-]+)\b/gi, (m: string) => {
|
||||
if (!/^(?:ADD|CORRESPONDING|CORR)$/i.test(m) && !sources.includes(m)) {
|
||||
sources.push(m);
|
||||
}
|
||||
return m;
|
||||
});
|
||||
}
|
||||
}
|
||||
break;
|
||||
}
|
||||
case 'SUBTRACT': {
|
||||
// SUBTRACT a FROM b [GIVING c] — target is after FROM or GIVING
|
||||
const subGiving = rest.match(
|
||||
/\bFROM\s+([A-Z0-9][A-Z0-9-]+)(?:\s+GIVING\s+([A-Z0-9][A-Z0-9-]+))?/i,
|
||||
);
|
||||
if (subGiving) {
|
||||
target = subGiving[2] ?? subGiving[1];
|
||||
if (subGiving[2]) givingTarget = subGiving[2];
|
||||
const beforeFrom = rest.substring(0, rest.toUpperCase().indexOf(' FROM '));
|
||||
beforeFrom.replace(/\b([A-Z0-9][A-Z0-9-]+)\b/gi, (m: string) => {
|
||||
if (!/^(?:SUBTRACT|CORRESPONDING|CORR)$/i.test(m) && !sources.includes(m)) {
|
||||
sources.push(m);
|
||||
}
|
||||
return m;
|
||||
});
|
||||
}
|
||||
break;
|
||||
}
|
||||
case 'MULTIPLY': {
|
||||
// MULTIPLY a BY b [GIVING c] — target is after BY or GIVING
|
||||
const mulGiving = rest.match(
|
||||
/\bBY\s+([A-Z0-9][A-Z0-9-]+)(?:\s+GIVING\s+([A-Z0-9][A-Z0-9-]+))?/i,
|
||||
);
|
||||
if (mulGiving) {
|
||||
target = mulGiving[2] ?? mulGiving[1];
|
||||
if (mulGiving[2]) givingTarget = mulGiving[2];
|
||||
const beforeBy = rest.substring(0, rest.toUpperCase().indexOf(' BY '));
|
||||
beforeBy.replace(/\b([A-Z0-9][A-Z0-9-]+)\b/gi, (m: string) => {
|
||||
if (!/^(?:MULTIPLY|CORRESPONDING|CORR)$/i.test(m) && !sources.includes(m)) {
|
||||
sources.push(m);
|
||||
}
|
||||
return m;
|
||||
});
|
||||
}
|
||||
break;
|
||||
}
|
||||
case 'DIVIDE': {
|
||||
// DIVIDE a INTO b [GIVING c] or DIVIDE a BY b [GIVING c]
|
||||
const divInto = rest.match(
|
||||
/\bINTO\s+([A-Z0-9][A-Z0-9-]+)(?:\s+GIVING\s+([A-Z0-9][A-Z0-9-]+))?/i,
|
||||
);
|
||||
const divBy = !divInto
|
||||
? rest.match(/\bBY\s+([A-Z0-9][A-Z0-9-]+)(?:\s+GIVING\s+([A-Z0-9][A-Z0-9-]+))?/i)
|
||||
: null;
|
||||
const divMatch = divInto ?? divBy;
|
||||
if (divMatch) {
|
||||
target = divMatch[2] ?? divMatch[1];
|
||||
if (divMatch[2]) givingTarget = divMatch[2];
|
||||
const beforeKeyword = divInto
|
||||
? rest.substring(0, rest.toUpperCase().indexOf(' INTO '))
|
||||
: rest.substring(0, rest.toUpperCase().indexOf(' BY '));
|
||||
beforeKeyword.replace(/\b([A-Z0-9][A-Z0-9-]+)\b/gi, (m: string) => {
|
||||
if (!/^(?:DIVIDE|CORRESPONDING|CORR)$/i.test(m) && !sources.includes(m)) {
|
||||
sources.push(m);
|
||||
}
|
||||
return m;
|
||||
});
|
||||
}
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
if (target) {
|
||||
result.arithmeticOps.push({
|
||||
verb,
|
||||
target,
|
||||
sources,
|
||||
line: lineNum,
|
||||
caller: currentParagraph,
|
||||
givingTarget,
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
// GO TO — control flow transfer (handles GO TO p1 p2 p3 DEPENDING ON x)
|
||||
const gotoMatch = line.match(RE_GOTO);
|
||||
if (gotoMatch) {
|
||||
const targets = gotoMatch[1]
|
||||
.trim()
|
||||
.split(/\s+/)
|
||||
.filter((t) => /^[A-Z][A-Z0-9-]+$/i.test(t));
|
||||
.filter((t) => /^[A-Z0-9][A-Z0-9-]+$/i.test(t));
|
||||
for (const target of targets) {
|
||||
result.gotos.push({ caller: currentParagraph, target, line: lineNum });
|
||||
}
|
||||
@@ -2046,7 +2234,7 @@ export function extractCobolSymbolsWithRegex(
|
||||
}
|
||||
}
|
||||
}
|
||||
const inspectMatch = line.match(/\bINSPECT\s+([A-Z][A-Z0-9-]+)/i);
|
||||
const inspectMatch = line.match(/\bINSPECT\s+([A-Z0-9][A-Z0-9-]+)/i);
|
||||
if (inspectMatch && inspectAccum === null) {
|
||||
inspectAccum = line;
|
||||
inspectStartLine = lineNum;
|
||||
@@ -2079,7 +2267,7 @@ export function extractCobolSymbolsWithRegex(
|
||||
const targets = setTrueMatch[1]
|
||||
.trim()
|
||||
.split(/\s+/)
|
||||
.filter((t) => /^[A-Z][A-Z0-9-]+$/i.test(t) && t.toUpperCase() !== 'OF');
|
||||
.filter((t) => /^[A-Z0-9][A-Z0-9-]+$/i.test(t) && t.toUpperCase() !== 'OF');
|
||||
if (targets.length > 0) {
|
||||
result.sets.push({ targets, form: 'to-true', line: lineNum, caller: currentParagraph });
|
||||
}
|
||||
@@ -2089,7 +2277,7 @@ export function extractCobolSymbolsWithRegex(
|
||||
const targets = setIdxMatch[1]
|
||||
.trim()
|
||||
.split(/\s+/)
|
||||
.filter((t) => /^[A-Z][A-Z0-9-]+$/i.test(t));
|
||||
.filter((t) => /^[A-Z0-9][A-Z0-9-]+$/i.test(t));
|
||||
const mode = setIdxMatch[2].toUpperCase();
|
||||
const form =
|
||||
mode === 'TO'
|
||||
@@ -2114,7 +2302,8 @@ export function extractCobolSymbolsWithRegex(
|
||||
.trim()
|
||||
.split(/\s+/)
|
||||
.filter(
|
||||
(t) => /^[A-Z][A-Z0-9-]+$/i.test(t) && !INITIALIZE_CLAUSE_KEYWORDS.has(t.toUpperCase()),
|
||||
(t) =>
|
||||
/^[A-Z0-9][A-Z0-9-]+$/i.test(t) && !INITIALIZE_CLAUSE_KEYWORDS.has(t.toUpperCase()),
|
||||
);
|
||||
for (const target of targets) {
|
||||
result.initializes.push({ target, line: lineNum, caller: currentParagraph });
|
||||
|
||||
@@ -0,0 +1,154 @@
|
||||
/**
|
||||
* Pure predicates gating C# `using` suffix-fallback resolution so BCL usings
|
||||
* (e.g. `System.Threading.Tasks`) can't match a coincidentally-named local
|
||||
* file (#1881).
|
||||
*
|
||||
* Lives in the shared `ingestion/` layer — NOT under `languages/csharp/` — so
|
||||
* BOTH the registry-primary scope resolver (`languages/csharp/import-target.ts`)
|
||||
* and the legacy DAG resolver (`import-resolvers/csharp.ts`) can import it
|
||||
* without an `import-resolvers/ -> languages/` dependency inversion (#5).
|
||||
*/
|
||||
|
||||
import type { CSharpNamespaceEvidence } from './language-config.js';
|
||||
|
||||
/**
|
||||
* Top-level namespace segments that clearly belong to the BCL / runtime / a
|
||||
* ubiquitous third-party package — i.e. roots a normal repo does NOT declare.
|
||||
* These stay gated even when the namespace scan is truncated, so a single
|
||||
* unreadable file / capped subtree can't silently re-enable BCL→local suffix
|
||||
* matches repo-wide (#1881). A repo that legitimately declares one of these
|
||||
* roots is still allowed via the alignment escape hatch below.
|
||||
*/
|
||||
const CSHARP_EXTERNAL_ROOTS: ReadonlySet<string> = new Set([
|
||||
// .NET BCL / runtime
|
||||
'System',
|
||||
'Microsoft',
|
||||
'Windows',
|
||||
'Mono',
|
||||
// ubiquitous third-party NuGet roots
|
||||
'Newtonsoft',
|
||||
'Serilog',
|
||||
'AutoMapper',
|
||||
'MediatR',
|
||||
'Polly',
|
||||
'FluentValidation',
|
||||
'Grpc',
|
||||
'Google',
|
||||
'Azure',
|
||||
'Amazon',
|
||||
'AWSSDK',
|
||||
// common test frameworks
|
||||
'Xunit',
|
||||
'NUnit',
|
||||
'Moq',
|
||||
'FluentAssertions',
|
||||
'NSubstitute',
|
||||
'Shouldly',
|
||||
]);
|
||||
|
||||
/** Whether `targetRaw`'s top-level segment is a clearly-external root. */
|
||||
function isExternalRoot(targetRaw: string): boolean {
|
||||
const dot = targetRaw.indexOf('.');
|
||||
const top = dot === -1 ? targetRaw : targetRaw.slice(0, dot);
|
||||
return CSHARP_EXTERNAL_ROOTS.has(top);
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether the unanchored suffix fallback may run for `targetRaw`.
|
||||
*
|
||||
* Fails OPEN when the namespace scan was truncated (large repos must not
|
||||
* silently lose legitimate edges, #1881 #11) and when no evidence was
|
||||
* threaded at all (preserves legacy permissive behavior). The truncation
|
||||
* fail-open is carved out for clearly-external roots (BCL / well-known
|
||||
* packages) that the repo does not declare, so one incomplete scan can't
|
||||
* re-open the #1881 hole repo-wide. Otherwise defers to
|
||||
* {@link importAlignsWithDeclaredNamespaces}.
|
||||
*/
|
||||
export function csharpSuffixFallbackAllowed(
|
||||
targetRaw: string,
|
||||
evidence: CSharpNamespaceEvidence | undefined,
|
||||
): boolean {
|
||||
if (evidence === undefined) return true;
|
||||
if (evidence.truncated) {
|
||||
// Keep clearly-external roots blocked through truncation UNLESS the repo
|
||||
// actually declares an aligning namespace (the alignment check is the
|
||||
// escape hatch — a repo that declares `namespace System;` still resolves).
|
||||
if (
|
||||
isExternalRoot(targetRaw) &&
|
||||
!importAlignsWithDeclaredNamespaces(
|
||||
targetRaw,
|
||||
evidence.declaredNamespaces,
|
||||
evidence.rootNamespaces,
|
||||
)
|
||||
) {
|
||||
return false;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
return importAlignsWithDeclaredNamespaces(
|
||||
targetRaw,
|
||||
evidence.declaredNamespaces,
|
||||
evidence.rootNamespaces,
|
||||
);
|
||||
}
|
||||
|
||||
/** True when `targetRaw` plausibly refers to a namespace declared in-repo. */
|
||||
export function importAlignsWithDeclaredNamespaces(
|
||||
targetRaw: string,
|
||||
declaredNamespaces: ReadonlySet<string> | undefined,
|
||||
rootNamespaces?: ReadonlySet<string>,
|
||||
): boolean {
|
||||
if (declaredNamespaces === undefined || declaredNamespaces.size === 0) return false;
|
||||
|
||||
// Exact: the import IS a declared in-repo namespace.
|
||||
if (declaredNamespaces.has(targetRaw)) return true;
|
||||
|
||||
// Child-of: the import's IMMEDIATE parent namespace is declared in-repo.
|
||||
// Anchoring on the direct parent — not "any declared prefix" — is what stops
|
||||
// a declared BCL prefix from green-lighting an unrelated BCL using: a repo
|
||||
// that declares `namespace System;` must NOT make `using
|
||||
// System.Threading.Tasks;` resolve to a coincidental local `Tasks.cs`,
|
||||
// because the import's parent `System.Threading` is not itself declared
|
||||
// (#1881). The case this still allows is a type / `using static` import under
|
||||
// a declared namespace laid out without its full path on disk, e.g.
|
||||
// `using static MyApp.Utils.Logger;` when `MyApp.Utils` is declared.
|
||||
const lastDot = targetRaw.lastIndexOf('.');
|
||||
if (lastDot > 0 && declaredNamespaces.has(targetRaw.slice(0, lastDot))) return true;
|
||||
|
||||
// Ancestor-of: the import is a strict prefix of some declared namespace
|
||||
// (e.g. `using MyApp;` when `MyApp.Models` is declared). Only honored when
|
||||
// the import also sits at or above an in-repo root namespace, so a BCL prefix
|
||||
// can't qualify merely because a file declares something deeper under it
|
||||
// (e.g. `System.Threading.Tasks.Extensions`) (#1881).
|
||||
const childPrefix = targetRaw + '.';
|
||||
for (const ns of declaredNamespaces) {
|
||||
if (ns.startsWith(childPrefix)) {
|
||||
return isAtOrAboveInRepoRoot(targetRaw, declaredNamespaces, rootNamespaces);
|
||||
}
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
function isAtOrAboveInRepoRoot(
|
||||
targetRaw: string,
|
||||
declaredNamespaces: ReadonlySet<string>,
|
||||
rootNamespaces: ReadonlySet<string> | undefined,
|
||||
): boolean {
|
||||
const descendantPrefix = targetRaw + '.';
|
||||
if (rootNamespaces !== undefined && rootNamespaces.size > 0) {
|
||||
for (const root of rootNamespaces) {
|
||||
// targetRaw equals a root, or is an ancestor of one (e.g. `using MyApp;`
|
||||
// for csproj RootNamespace `MyApp.Core`).
|
||||
if (root === targetRaw || root.startsWith(descendantPrefix)) return true;
|
||||
}
|
||||
return false;
|
||||
}
|
||||
// No explicit roots (e.g. no csproj): treat the top-level segment of each
|
||||
// declared namespace as the implied root.
|
||||
for (const ns of declaredNamespaces) {
|
||||
const dot = ns.indexOf('.');
|
||||
const top = dot === -1 ? ns : ns.slice(0, dot);
|
||||
if (top === targetRaw) return true;
|
||||
}
|
||||
return false;
|
||||
}
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user