Compare commits

...
Author SHA1 Message Date
Gergő Magyar e7a86467d6 Merge branch 'main' into ci/vercel-deterministic-install 2026-05-22 04:42:42 +01:00
dependabot[bot] 8c1983a8bf chore(deps)(deps-dev): bump @types/node in /gitnexus (#1767) 2026-05-22 04:41:57 +01:00
Abhigyan PatwariandClaude Opus 4.7 317fc1f00c ci(web): use npm ci for deterministic Vercel installs
Vercel restores a cached node_modules snapshot and runs the custom
installCommand incrementally. With `npm install`, a lockfile change makes
the incremental reconcile leave gitnexus-web without a usable `tsc`
binary, so `npm run build` (`tsc -b && vite build`) fails with
`sh: line 1: tsc: command not found` (exit 127) — the failure blocking
#1705. `npm ci` ignores the restored tree and installs exactly from the
lockfile (the canonical CI/deploy install), so build tooling is always
present. `--include=dev` keeps devDependencies even if the build env sets
NODE_ENV=production. Vercel still caches the npm download store, so the
only cost is a fast relink each build.

Unblocks #1705.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-22 00:44:04 +05:30
231ad71d40 fix(mcp): disambiguate duplicate-name repo resolution for worktrees (#1753)
* fix(mcp): disambiguate duplicate-name repo resolution for worktrees

When multiple indexed repos share the same registry name (main checkout plus linked worktrees), MCP tools no longer silently pick the first sibling. Resolution prefers the repo matching process.cwd()'s git root, throws RegistryAmbiguousTargetError when still ambiguous, and uses canonical path matching aligned with the CLI registry.

Fixes #1658. Complements worktree detect_changes fixes in #1654/#1691.

* fix(mcp): refresh registry on duplicate-name ambiguity before failing

resolveRepo now retries resolveRepoFromCache after RegistryAmbiguousTargetError so stale in-memory siblings clear when the registry changes. Adds detect_changes callTool ambiguity test, registry-refresh regression test, pickRepoHandleForCwd MCP cwd doc, and temp-dir cleanup in #1658 fixtures.

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(mcp): PR #1753 review follow-ups + collision-id case bug

Address Findings 3-6 from the production-readiness review on PR #1753,
plus a latent bug surfaced while writing the F5 regression test:

- F3: drop the no-op `try { ... } catch (err) { throw err; }` wrapper
  around the miss-path retry in `resolveRepo`; the catch only re-threw.
- F4: rewrite the misleading "child/repo" example on the relative-path
  tier — `child/repo` would be classified as path-like and never reach
  this branch. Comment now describes bare, separator-free names
  resolved against `process.cwd()`.
- F5: add regression test for the stable hashed-id tier so a duplicate
  sibling can be reached by its `<name>-<hash>` id. Writing this test
  exposed that `repoId()` produced a mixed-case base64url suffix while
  `resolveRepoFromCache` lowercased the param before the Map lookup, so
  collision ids with any uppercase byte in the hash were unreachable.
  Fix: lowercase the hash in `repoId` so it survives `paramLower`.
- F6: add regression test asserting two repos sharing a name prefix
  (`project-a`, `project-b`) cause `resolveRepo("project")` to reject
  as not-found rather than silently returning the first partial match.

* refactor(mcp): tighten PR #1753 follow-up tests + pin hash length

Address three P2 maintainability findings from the ce-code-review pass
on commit aa7f2050:

- Export `REPO_ID_HASH_LENGTH` from local-backend.ts and use it in both
  `repoId()` and the hashed-id test. Closes the silent-drift hole where
  the test's inline formula could fall out of sync with the source
  without any signal.
- Extract `makeSharedPrefixFixture(nameA, nameB)` next to
  `makeDuplicateNameFixture`. Centralises the temp-dir + `.gitnexus`
  scaffolding + `duplicateFixtureDirs.push()` cleanup contract so
  future callers can't drop the cleanup step.
- Reorder the hashed-id test's comment block so the intentional-coupling
  rationale leads, before the description of the formula being mirrored.

* chore(autofix): apply prettier + eslint fixes via /autofix command

* chore: re-run CI

---------

Co-authored-by: Test <test@example.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-21 19:21:25 +01:00
df2ed009ce fix(group): detect httpx AsyncClient alias imports (#1687)
* fix(group): detect httpx AsyncClient alias imports

* fix(group): anchor httpx dotted imports and skip shadowed aliases

Addresses Findings 1-3 of the production-readiness review on PR #1687.

- F1: the `(dotted_name (identifier) @module)` capture matches every
  segment of a dotted module path, so `import package.httpx as hx` and
  `from package.httpx import AsyncClient` would falsely populate the
  alias sets. Anchor the check on `moduleNode.parent?.text === 'httpx'`
  so the full dotted_name must equal `httpx`.

- F2: `moduleAliases` and `asyncClientAliases` were file-global and
  unaware of Python scope. A function-local rebind like
  `AsyncClient = lambda: MockClient()` left the alias entry intact and
  any subsequent `client = AsyncClient(); client.get(...)` emitted a
  false-positive consumer contract. Walk every
  `(assignment left: (identifier) @name)` whose name matches an alias,
  record the enclosing function/class scope as poisoned, and skip
  direct- and module-attribute matches when the call site is inside
  that scope chain.

- F3: extend the existing fixture with dotted-package look-alikes and
  three local-shadow cases (`shadow_direct_alias`, `shadow_module_alias`,
  `shadow_direct_context`) and assert the would-be FP contractIds are
  not emitted.

- F6: refresh the module-level docstring to mention the supported
  import-alias forms and the shadow-exclusion behavior.

* refactor(group): tighten httpx alias shadow detection and broaden tests

Follow-up addressing the residual review findings on PR #1687.

- Replace inline scope-key construction in isAliasShadowed with a
  getScopeKey call so the two helpers cannot drift apart (M1).
- Collapse the double tree traversal in collectHttpxAsyncClients: build
  one combined alias set and pass it to a single
  collectAliasShadowScopes call (perf, P2).
- Add a `shadowScopeKey` helper that returns the scope a rebind actually
  shadows under Python LEGB rules: function scope for in-function
  rebinds, 'module' for top-level rebinds, and `null` for class-body
  rebinds (class attributes do not shadow bare-name lookups in methods).
  Removes the previous blanket `scopeKey === 'module'` skip and now
  correctly poisons module-level rebinds (correctness #1).
- Extend `ALIAS_SHADOW_PATTERNS` to cover tuple, list, and pattern_list
  destructuring targets (correctness #2).
- Rename `ALIAS_REBIND_PATTERNS` to `ALIAS_SHADOW_PATTERNS` and update
  the block comment to say "shadowed" rather than "poisoned" (M4).
- Collapse `callScopeKeys` to a single-line return; the dead Set wrap
  was misleading future readers (M2).

Tests:
- New negative fixtures for 3-segment dotted import
  (`import a.b.c.httpx as deep_evil`), relative import
  (`from .httpx import AsyncClient as rel_evil_async`), tuple
  destructuring rebind, and an isolated file exercising the module-level
  rebind path (T1, correctness #2, expanded F2).
- New positive fixture confirming that a class-body assignment of
  `AsyncClient` does NOT poison the surrounding methods.
- Add a positive control assertion for `module_direct_client` so the
  dotted-package negative assertions cannot pass vacuously (T3).

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Test <test@example.com>
2026-05-21 18:24:27 +01:00
luyua9andGergő Magyar dd3527327d feat(ingestion): Link object literal methods to exported bindings (#1718)
* fix: link object literal methods to exported bindings

* fix(ingestion): bridge object-literal value receivers in scope-resolution (PR #1718 review)

Addresses adversarial production-readiness review on PR #1718 / issue #1358:
- F1 (caller resolution) — setting `ownerId` on object-literal method symbols
  alone is not sufficient; the scope-resolution receiver-bound resolver only
  consults class-like or type-annotated bindings, so lowercase value receivers
  (`export const fooService = {...}; fooService.getUser(...)`) never reach the
  owner-indexed lookup. Adds a Case 5 value-receiver bridge in
  receiver-bound-calls.ts that resolves the receiver name as a Const/Variable
  binding, translates its def to the canonical graph node id, and emits the
  CALLS edge via the owner-indexed method registry.
- F2 (boundary guard) — rewrites findObjectLiteralBindingInfo as an explicit
  two-phase AST walk: Phase A tracks object-literal depth (returns null for
  nested literals and pre-declarator function/class boundaries — IIFE
  patterns); Phase B walks the declarator's ancestors and rejects function,
  class, and block-statement containers (if / for / while / try / catch /
  switch / etc.) before reaching program/export_statement. Prevents false
  HAS_METHOD edges for locally-scoped or block-scoped object literals.
- F4 — drops the dead `ownerName` field from ObjectLiteralBindingInfo.

Constraint: TS/JS are scope-resolution migrated per RFC #909; the legacy
Call-Resolution DAG (call-processor.ts) is intentionally left untouched.

Tests:
- test/integration/ast-helpers-object-literal-binding.test.ts (13 cases) —
  pins helper semantics: happy paths, function/arrow/class-ctor boundaries,
  nested literals, block scope (if / for-of / try), IIFE, assignment
  expressions without declarator.
- test/integration/object-literal-owner-resolution.test.ts (9 cases) —
  drives the full pipeline against an on-disk fixture: sequential CALLS edge
  emission (issue #1358 proof), worker-mode parity, negative local binding,
  and nested-literal attribution boundary.

Full sweep: 2958/2958 integration + 6056/6056 unit tests pass.

* refactor(ingestion): address code-review findings on object-literal owner resolution

Multi-agent code review on the prior commit surfaced 7 actionable findings,
all walked through and applied here. None change observable behavior for
issue #1358's fix; all harden correctness, predicate stability, and test
signal.

- #1 (P1 / 3-reviewer corroboration): Case 5 in receiver-bound-calls.ts no
  longer hand-builds graph.addRelationship + a dedup key. New
  tryEmitEdgeWithExplicitTargetId in edges.ts takes a pre-resolved target
  id (the canonical Method nodeId from the parser) and reuses every
  invariant of tryEmitEdge: dedup-key format, collapse-flag honoring,
  caller-id resolution, rel-id shape, mapReferenceKindToEdgeType for
  read/write ACCESSES. This also lands the adversarial reviewer's "F2"
  follow-up (hardcoded type: 'CALLS' for non-call sites) for free.

- #2 (P2 cross-reviewer): findValueBindingInScope's predicate inverted
  from denylist ("not class-like and not callable") to explicit allowlist
  matching reconcileOwnership's registration set:
  Const | Variable | Property | Static. Extracted as isOwnableValueLabel
  so future NodeLabel additions require an explicit opt-in.

- #6 (P2): walkScopeChain<T>() extracted; both findClassBindingInScope
  and findValueBindingInScope now route through it. Local scope.bindings
  are exhausted BEFORE lookupBindingsAt (imported/augmented) at every
  scope level — preserves JavaScript lexical scoping where a local const
  shadows an imported binding of the same name. Behavior was already
  correct in findClassBindingInScope but was implicit; now it is the
  walker's explicit, documented contract.

- #7 (P2): scope-walker duplication closed. findClassBindingInScope and
  findValueBindingInScope reduce to thin wrappers over walkScopeChain
  with their respective predicate. findClassBindingInScope keeps its
  qualifiedNames + dotted-name fallback tail.

- #3 (P2): parse-worker.ts hoists `const ownerId = enclosingClassId ??
  objectLiteralOwnerInfo?.ownerId` once before the symbol push, dropping
  the duplicated coalesce + `as string` cast. Matches the cast-free
  pattern at parsing-processor.ts:793. HAS_METHOD emit site reuses the
  same hoisted local.

- #4 (P2): object-literal-owner-resolution.test.ts Test A's CALLS-edge
  assertion no longer matches by name alone. .toEqual now pins the
  canonical target id (Method:src/service.ts:getUser#1 via generateId),
  confidence (0.85), and reason ('import-resolved'). A regression that
  emits the edge at confidence=0, with the wrong reason, or against a
  phantom Method node now fails the test.

- #5 (P2): worker-parity test adds a CI tripwire — when CI=1 and
  dist/parse-worker.js is missing, throw at module top with a clear
  message. Locally, skipIf(!hasDistWorker) keeps the fast-iteration
  experience; CI cannot pass with U3 (worker-path ownerId) unverified.

Verification: tsc --noEmit clean. Targeted regression sweep on
ast-helpers-object-literal-binding (13), object-literal-owner-resolution
(9), has-method (60), cross-file-binding (40) — 122/122 pass. Full unit
sweep: 6056/6056. Integration suite: 1 pre-existing Windows-flake in
worker-pool.test.ts (passes 28/28 in isolation) unrelated to this diff.

* refactor(scope-resolution): align Const label emission with legacy DAG (PR #1718 review F1)

Eliminates the architectural fragility surfaced by PR #1718's adversarial review
Finding 1. Previously, normalizeNodeLabel('const') returned 'Variable' while
the legacy DAG parse phase emits 'Const' graph nodes (via @definition.const
capture for lexical_declaration). PR #1718's Case 5 value-receiver bridge
resolved correctly only because resolveDefGraphId happened to fall back to
simpleKey after the qualified-key miss — accidental correctness.

After this change, scope-resolution defs for `const x = ...` declarations
report def.type === 'Const', matching the graph node label. resolveDefGraphId's
qualified-key path now hits on the first try; the simple-key fallback is no
longer load-bearing for value receivers and can be tightened in future without
silently breaking Case 5.

Audit completeness verification:
- Grep `\bVariable\b` across src/core/ingestion/scope-resolution/ surfaced two
  consumer sites that already accept both labels: reconcile-ownership.ts:101+168
  (`def.type === 'Variable' || def.type === 'Const' || ...`) and
  walkers.ts:207 isOwnableValueLabel (`Const | Variable | Property | Static`).
  No language hook in src/core/ingestion/languages/ branches on
  `def.type === 'Variable'` for what's actually a const declaration.
- Sentinel stress test (the full unit + integration suite run with the
  renamed label in place): 6137/6137 unit tests pass; 2967/2967 integration
  tests pass. One pre-existing Windows-only flake on worker-pool.test.ts when
  run alongside the full integration suite (passes 28/28 in isolation,
  unrelated to scope-extractor — same flake observed before this diff).

The variable mapping (`'variable' → 'Variable'`) is preserved for `var`
declarations, matching the legacy DAG's `@definition.variable` capture for
variable_declaration. The split now mirrors the parse-phase capture
distinction exactly.

Per plan docs/plans/2026-05-21-002-feat-pr1718-followups-class-instance-and-label-normalization-plan.md
U4 + U5. T1 (class-instance singleton resolution from issue #1358's second
sub-case) is deferred to a standalone pre-plan investigation, not shipped
here.

* test(ingestion): add regression coverage for issue #1358 singleton sub-cases

Closes the remaining sub-cases of issue #1358 surfaced by PR #1718's
adversarial review (Finding 4, NOTED): the class-instance singleton
(`export const fooService = new FooService();`) and the factory-pattern
singleton (`export const fooService = makeFooService();`).

Pre-plan investigation (per docs/plans/2026-05-21-002 § "Pre-Plan
Investigation Task (T1)") confirmed Outcome A for both patterns — they
already resolve end-to-end through scope-resolution's
`@type-binding.constructor` capture (languages/typescript/query.ts:489-511)
+ `propagateImportedReturnTypes` chain-follow
(scope-resolution/passes/imported-return-types.ts:114) + receiver-bound
Case 4 simple typeBinding lookup (receiver-bound-calls.ts:625). The
mechanism was wired correctly before this session; the regression-net
wasn't.

This test pins the behavior:
- Pattern 1: `caller → FooService.getUser` CALLS edge with
  confidence 0.85 and reason 'import-resolved'
- Pattern 2: same edge shape via factory chain-follow (the
  `@type-binding.alias` capture for `const u = find()` style)

Both assertions use exact `.toEqual([{...}])` shape pinning so a future
regression that targets a phantom Method node, emits at lower confidence,
or drops the cross-file import-resolved reason fails loudly.

Verification: 5/5 pass, 127/127 in targeted regression sweep including
object-literal-owner-resolution.test.ts, ast-helpers-object-literal-
binding.test.ts, has-method.test.ts, and cross-file-binding.test.ts.

No production code change. The class methods get a class-qualified node id
(`Method:src/service.ts:FooService.getUser#1`) distinguishing them from
same-name methods on other classes — distinct from the bare-name node id
shape PR #1718's object-literal case uses.

* test(resolvers): add class-instance + factory-pattern singleton coverage for TS/JS (issue #1358)

Closes the remaining sub-cases of issue #1358 surfaced by PR #1718's
adversarial review (Finding 4). PR #1718 fixed object-literal-shorthand
singletons (`export const fooService = { getUser() {} }`); this commit adds
parallel coverage for the two other singleton shapes that resolve through
the existing scope-resolution chain:

  // Pattern 1 — class-instance singleton
  export class FooService { getUser(id) { ... } }
  export const fooService = new FooService();

  // Pattern 2 — factory-pattern singleton
  export class FooService { getUser(id) { ... } }
  export function makeFooService() { return new FooService(); }
  export const fooService = makeFooService();

Pre-plan investigation (per local plan docs/plans/2026-05-21-002 § "Pre-Plan
Investigation Task (T1)") confirmed Outcome A — both patterns already
resolve end-to-end through:
  - `@type-binding.constructor` capture (languages/{typescript,javascript}/
    query.ts) seeds `fooService → FooService` at parse time
  - `propagateImportedReturnTypes` (scope-resolution/passes/
    imported-return-types.ts:114) mirrors the typeBinding cross-file
  - Receiver-bound Case 4 simple typeBinding lookup
    (scope-resolution/passes/receiver-bound-calls.ts:625) MRO-walks
    FooService and emits the CALLS edge to getUser

Tests added per language × pattern (5 each, 10 total):
- node existence (Class, Method, Function, Const, plus Function for the
  factory pattern's `makeFooService`)
- HAS_METHOD edge from class to method (class-instance variant)
- CALLS edge from caller to `getUser` with `targetFilePath: 'src/service.{ts,js}'`,
  `reason: 'import-resolved'`, `confidence: 0.85` — exact `.toEqual([{...}])`
  shape pinning so a regression that emits at lower confidence or drops the
  cross-file reason fails loudly

Fixtures placed under the existing `test/fixtures/lang-resolution/` convention.
Tests appended to `test/integration/resolvers/{typescript,javascript}.test.ts`,
matching the in-file pattern of every other resolver scenario.

Also supersedes and removes the standalone
`test/integration/class-instance-and-factory-singleton-resolution.test.ts`
introduced earlier in this PR session (`0df91b77`) — the proper home for
language-resolver scenarios is the per-language resolver test file alongside
similar fixtures (`javascript-self-this-resolution`, `javascript-cross-file`,
`typescript-tsconfig-paths`, etc.). One canonical location for the scenario,
not two.

Verification: 10/10 new singleton tests pass; 297/297 full TS+JS resolver
suite pass (no regression in any existing resolver test).

* test(resolvers): gate TS/JS singleton tests behind scope-resolution parity (CI run 26223603426)

The class-instance and factory-pattern singleton CALLS-edge resolution
tests added in c8e573bc rely on scope-resolution-only mechanisms
(`@type-binding.constructor` capture + `propagateImportedReturnTypes`
mirror + receiver-bound Case 4). The `scope-parity / typescript parity`
and `scope-parity / javascript parity` CI jobs run with
`REGISTRY_PRIMARY_TYPESCRIPT=0` / `REGISTRY_PRIMARY_JAVASCRIPT=0` and
exercise the legacy DAG path, which has no cross-file constructor-derived
typeBinding propagation. Verified by job 77202610819 (TS parity) and
77202610869 (JS parity) failing with:

  × resolves caller.fooService.getUser() to FooService.getUser via constructor-inferred typeBinding
  × resolves caller.fooService.getUser() through the factory chain to FooService.getUser

Note: my local Windows shell-prefix env-var invocation did not propagate
the flag into vitest workers correctly (the cpp parity gate's 47-skipped
behavior masked the issue when I ran an ad-hoc comparison), so the
empirical "both modes pass" finding I posted earlier was wrong. CI is the
source of truth.

Changes:
- test/integration/resolvers/helpers.ts: add `typescript` and `javascript`
  entries to `LEGACY_RESOLVER_PARITY_EXPECTED_FAILURES` for the 2 CALLS-edge
  resolution tests in each language. Node-existence and HAS_METHOD
  assertions are NOT excluded — those pass under legacy DAG (parser-level
  emission is intact).
- test/integration/resolvers/typescript.test.ts: drop the `it` import from
  vitest; replace with `const it = createResolverParityIt('typescript');`
  shadow (matches the c/cpp/csharp/go pattern at the top of those files).
- test/integration/resolvers/javascript.test.ts: same shadow with
  `createResolverParityIt('javascript')`.

Verification:
- Default mode (registry-primary): 297/297 TS+JS resolver tests pass.
- Legacy DAG mode: the 4 listed singleton CALLS-edge tests will skip; all
  other singleton assertions (node existence + HAS_METHOD edge) continue
  to run and pass under both modes.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-21 17:18:27 +01:00
Gergő MagyarandCursor d3de5fa5d5 fix(install): materialize vendored grammars to fix Windows EPERM (#1728) (#1729)
* fix(install): materialize vendored grammars to fix Windows EPERM (#1728)

Stop using file: optionalDependencies for tree-sitter-dart/proto/swift,
which made npm symlink vendor paths on install and fail on Windows without
symlink privileges. Copy vendor trees into node_modules at postinstall
instead; keep native builds and #836 vendor hygiene.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(install): atomic materialize swap + fail-soft tests (#1728, #836)

Hardens PR #1729 against two issues the original implementation could
still hit:

1. Torn-state on rmSync→cpSync. The previous loop deleted the
   destination before copying. If cpSync threw — the exact Windows EPERM
   scenario this PR targets — a previously-working grammar was silently
   wiped. Now we copy to {dest}.materialize-tmp first and renameSync into
   place, so an interrupted copy leaves the prior materialization intact.

2. Fail-soft try/catch had no test coverage. Adds two POSIX-only tests
   (chmod 0o555 to deterministically force cpSync to throw) that verify
   (a) a single grammar failure does not abort the other two, and (b) an
   existing materialization survives a partial-copy failure. Skipped on
   Windows where chmod doesn't enforce write restriction; runs on Linux
   CI.

Other test improvements locking in the install-hygiene invariants:

- All three vendored grammars (dart/proto/swift) checked, not just dart.
- GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1 short-circuit is exercised.
- Vendor cleanliness (#836): no node_modules/build under vendor/.
- Idempotent re-runs (clean overwrite verified via sentinel file).
- Missing-vendor warn+continue path now has explicit coverage.
- Vendored package manifests asserted to carry no install script or
  runtime dependencies.
- package.json optionalDependencies asserted free of vendored grammars.
- package-lock.json assertion tightened from `if (entry !== undefined)
  { expect(entry.link).not.toBe(true); }` (vacuous when entry is absent,
  i.e. the expected post-fix state) to `expect(...).toBeUndefined()`.

Verified locally:
- npx tsc --noEmit: clean
- vitest test/unit/materialize-vendor-grammars.test.ts: 8 pass + 2
  POSIX-only skipped on Windows
- npm pack tarball: no vendor/*/node_modules or vendor/*/build entries
- Isolated global install (clean + upgrade + SKIP env) into temp prefix:
  succeeds; gitnexus --version → 1.6.5; vendor stays clean post-install.

* fix(install): address review feedback — Swift parity, atomicity, CI smoke

Resolves all findings from the automated production-readiness review on
verify/issue-1728-symlink.

Swift warning parity (review #2):
  Add tree-sitter-swift to OPTIONAL_GRAMMARS in src/cli/optional-grammars.ts
  alongside Dart and Proto. Before this commit, Swift was materialized at
  postinstall and probed by build-tree-sitter-swift.cjs but the runtime
  warnMissingOptionalGrammars() never warned when it failed to load —
  users got silent Swift degradation from the optional-grammars surface
  (parser-loader's separate unavailableNote only fires on demand). Now
  the warning path matches the materialize path.

README env-var table (review #1):
  Update the GITNEXUS_SKIP_OPTIONAL_GRAMMARS row at README.md line 248 to
  list all three vendored grammars (dart, proto, swift). The quick note
  earlier in the README already mentioned all three; only the table row
  was stale.

Atomicity hardening (review #3):
  materialize-vendor-grammars.cjs now copies to {dest}.materialize-tmp,
  renames the existing dest to {dest}.materialize-bak (if present), then
  renames the partial into dest, then removes the backup. If the
  partial→dest rename fails (e.g. Windows AV scanner racing the swap),
  the catch block restores from backup so the previously-materialized
  grammar is preserved. Closes the narrow torn-state window where the
  prior implementation could leave dest deleted after rmSync succeeded
  but renameSync failed.

Swift probe docs (review #4):
  build-tree-sitter-swift.cjs script header rewritten to describe what
  the script actually does — probe node-gyp-build at install time so
  missing-prebuild failures surface as install-time warnings instead of
  first-parse runtime errors. The script does not "activate" anything;
  the runtime require() in parser-loader does the actual load. Console
  warning text updated to match ("prebuild probe" not "activation").

Windows packaged-install smoke test (review #5):
  New CI job `packaged-install-smoke` in .github/workflows/ci-tests.yml
  matrices on windows-latest and ubuntu-latest. Runs npm pack, installs
  the produced tarball globally into RUNNER_TEMP, then asserts:
    * no vendor/*/node_modules or vendor/*/build (#836 invariant)
    * tree-sitter-{dart,proto,swift} in node_modules are real
      directories, not junctions/symlinks (#1728 invariant)
    * gitnexus --version runs against the installed CLI
  Closes the coverage gap where the existing windows-latest job only
  ran `npm ci` in the source checkout — exercising postinstall but not
  the tarball reify step that historically tripped EPERM.

Verified locally:
  npx tsc --noEmit: clean
  vitest test/unit/materialize-vendor-grammars.test.ts test/unit/cli-commands.test.ts:
    18 pass + 2 POSIX-only skipped on Windows
  prettier + eslint on all changed files: clean

* fix(ci): disable credential persistence on packaged-install-smoke checkout

GitHub Advanced Security (zizmor artipacked) flagged the new
packaged-install-smoke job's actions/checkout step as a potential
credential-persistence risk. The job runs `npm pack` + global install
and never pushes back, so the GITHUB_TOKEN that checkout would persist
in .git/config provides no value and only widens the leak surface (any
future artifact-upload step in this job would carry the token).

Disable persistence explicitly via `persist-credentials: false` on this
job's checkout. Scoped to the new job — pre-existing checkouts above
are left unchanged.

* fix(ci): use find instead of ls for tarball lookup (SC2012)

actionlint shellcheck SC2012 flagged `TARBALL=$(ls gitnexus-*.tgz | head -n1)`.
Switch to `find . -maxdepth 1 -name 'gitnexus-*.tgz' -print -quit` which
handles non-alphanumeric filenames safely. Also add an explicit
empty-result check so the failure mode is a clear error message instead
of a silent `npm install -g ""` later.

* fix(tests): sabotage vendor src (not partial path) in POSIX fail-soft tests

The fail-soft tests in materialize-vendor-grammars.test.ts pre-chmod'd
the destination's .materialize-tmp partial directory to 0o555 to force
cpSync to throw. After the atomicity rewrite (`fix(install): atomic
materialize swap + fail-soft tests`), the materialize script now starts
each grammar's loop with `fs.rmSync(partial, { force: true })`, which
deletes the chmod'd sabotage before cpSync runs — so cpSync succeeds and
the partial is then renamed into dest, leaving the test's `finally`
block with no path to chmod back (ENOENT) and the assertion that proto
remained unmaterialized failing because it materialized cleanly.

Fix: sabotage the *vendor source* directory (which the script reads from
but never modifies) by chmod'ing it to 0o000. cpSync then fails on
readdir, the catch block fires per-grammar, dart and swift still
materialize from their unaffected sources, and the existing-dest
preservation test verifies that a sabotaged second-run leaves the prior
materialization (and its sentinel file) intact.

Tests now pass locally (8 pass + 2 POSIX-only skipped on Windows) and
should pass on macOS/Ubuntu CI where the sabotage runs.

* fix(tests): restrict fail-soft tests to Linux (macOS Node cpSync abort)

Node 22 on macOS aborts the process with `libc++abi: terminating due
to uncaught exception filesystem_error` when fs.cpSync hits a source
directory it can't read — the abort happens at the C++ filesystem layer
and bypasses Node's JS try/catch entirely (nodejs/node#51399). My
chmod-0o000-the-source sabotage strategy triggers this SIGABRT on
macOS CI before the production script's `try { cpSync } catch` ever
runs, so the test sees a child-process crash instead of the fail-soft
warning it's verifying.

The production script's fail-soft is correct on Linux (where EACCES
surfaces as a normal JS exception) and effectively untestable on macOS
via permission sabotage. Real installs don't hit this — npm always
ships vendor/ with readable permissions — so the macOS gap is a test
artifact, not a behavior gap.

Restrict the two chmod-based tests to Linux only by replacing
`skipOnWin` with `linuxOnly`. Linux CI continues to verify both the
one-grammar-fails-others-succeed and existing-materialization-preserved
invariants. macOS and Windows runs skip these two scenarios; the other
8 tests still run on every platform.

* fix(tests): remove materialize unit tests, rely on CI smoke job

The materialize-vendor-grammars.test.ts file has been a recurring source
of platform-specific CI noise:

  - Windows: chmod doesn't enforce read/write restrictions the way POSIX
    does, so the fail-soft tests had to be skipped there.
  - macOS Node 22: cpSync against an unreadable source aborts the process
    with a libc++ filesystem_error (nodejs/node#51399) that bypasses JS
    try/catch entirely — making the chmod-based fail-soft tests
    unrunnable on macOS too.
  - The "vendor-cleanliness" and "idempotency" tests on Windows
    intermittently flake due to fs.cpSync timing on the GitHub runner.

The invariants these tests verified are now covered by stronger,
more realistic surfaces:

  - packaged-install-smoke (ci-tests.yml): runs `npm pack` then
    `npm install -g ./gitnexus-*.tgz` on windows-latest and
    ubuntu-latest, then asserts no vendor/*/node_modules,
    no vendor/*/build (#836), no junctions/symlinks on the
    materialized grammar directories (#1728), and a working
    `gitnexus --version`. This is the actual end-user install path.

  - cli-commands.test.ts (kept, unmodified): asserts package.json
    declares no `file:` optionalDependencies for vendored grammars,
    the Swift vendor manifest carries no install script or
    dependencies, and the postinstall chain runs
    materialize-vendor-grammars.cjs + build-tree-sitter-swift.cjs.
    These are static manifest checks — deterministic, fast, no
    flake risk.

Removing the dynamic script-execution tests trades unit-level coverage
for end-to-end smoke coverage that actually exercises the
`file:` → cpSync change against a real npm install lifecycle, on
the platform the fix targets (windows-latest).

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-21 16:47:22 +01:00
4c06d64a3b chore(deps)(deps): bump zod from 4.3.6 to 4.4.3 in /gitnexus-web (#1736)
Bumps [zod](https://github.com/colinhacks/zod) from 4.3.6 to 4.4.3.
- [Release notes](https://github.com/colinhacks/zod/releases)
- [Commits](https://github.com/colinhacks/zod/compare/v4.3.6...v4.4.3)

---
updated-dependencies:
- dependency-name: zod
  dependency-version: 4.4.3
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-21 16:25:35 +01:00
ChamHerryandwangxc 2a3d14057a fix(analyze): prevent cache-hit native workers from aborting (#1751)
* fix(analyze): prevent cache-hit native workers from aborting

Delay parse worker startup until a cache miss requires it, fall back to sequential parsing when initial worker readiness fails, and preserve analyzer diagnostics/progress when heap respawn captures child output.

Constraint: Node 25 and tree-sitter/N-API worker initialization can abort before ready, while warm-cache analysis should not start workers at all.

Rejected: Treating status-134/SIGABRT as heap OOM unconditionally | native worker aborts require distinct recovery guidance and stderr/stdout evidence.

Rejected: cli-progress noTTYOutput for respawn progress | it appends newline frames instead of preserving one-line redraw UX.

Confidence: high

Scope-risk: moderate

Directive: Keep parse-worker creation behind confirmed cache misses and preserve TTY-style progress when respawn pipes stderr for crash classification.

Tested: GitNexus impact analysis for ensureHeap, runChunkedParseAndResolve, createWorkerPool, WorkerPool, walkRepositoryPaths; GitNexus detect_changes scoped to staged worktree; targeted vitest for analyze respawn, parse lazy cache, filesystem walker, worker pool; npx tsc --noEmit; npm run build; NODE_OPTIONS='--max-old-space-size=8192' npm test.

Not-tested: Windows terminal rendering and published npm package install path.

* ci(docker): tolerate slower arm64 TypeScript builds

Docker PR builds run gitnexus prepare under QEMU for linux/arm64, where the fixed 120s TypeScript timeout can kill otherwise healthy builds. Increase the default timeout and allow GITNEXUS_BUILD_TIMEOUT_MS to tune slower environments without changing the build steps.

Constraint: PR #1751 Docker Build & Push gitnexus failed with spawnSync /bin/sh ETIMEDOUT while running node_modules/.bin/tsc in scripts/build.js.\nRejected: Rerunning CI only | the failure was the build script's deterministic timeout boundary under arm64 emulation, not a code assertion.\nConfidence: high\nScope-risk: narrow\nDirective: Keep build timeout changes in scripts/build.js configurable; do not hide real compiler failures, only allow slower successful compiles to finish.\nTested: GitNexus impact for gitnexus/scripts/build.js reported LOW; gitnexus detect_changes reported 1 changed file, 0 affected processes, low risk; git diff --check; gitnexus npm run build.\nNot-tested: GitHub Docker arm64 build rerun before pushing; local Docker multi-platform build under QEMU.

* fix(analyze): truncate respawn progress safely

Preserve complete ANSI escape sequences and grapheme boundaries when the respawn progress terminal shim truncates wrapped output, so the shim does not emit dangling escape bytes or split surrogate pairs while keeping raw writes untouched.

Constraint: Claude review on PR #1751 flagged `s.slice(0, width)` in createAnsiPipeTerminal.write() as a latent terminal-corruption risk.
Rejected: Adding a display-width dependency | a local helper is sufficient for this narrow respawn terminal shim and avoids new dependency churn.
Rejected: Changing silent status-134 classification | current tests already document the output-less 134 fallback as heap guidance.
Confidence: high
Scope-risk: narrow
Directive: Keep respawn terminal writes ANSI-aware and preserve rawWrite bypass semantics for callers that intentionally write control sequences.
Tested: GitNexus impact for createAnsiPipeTerminal reported LOW; GitNexus detect_changes reported 2 changed files, 3 affected processes, medium risk; targeted vitest for analyze respawn progress and heap respawn; gitnexus npx tsc --noEmit; prettier check for changed files; eslint for changed files.
Not-tested: Full npm test suite; manual terminal rendering on Windows.

---------

Co-authored-by: wangxc <wangxc_a_bj@si-tech.com.cn>
2026-05-21 16:17:02 +01:00
a9fef2c68d fix(lbug): keep serve stable when sidecars are missing (#1747)
* fix(lbug): keep serve stable when sidecars are missing

Shared missing-shadow WAL recovery prevents repeated read-only open warnings when LadybugDB sidecars are absent, while the Express preflight fix keeps `gitnexus serve` compatible with Express 5 route parsing.

Constraint: LadybugDB read-only replay can require a `.shadow` sidecar that may be absent after interrupted writes or checkpoint edge cases.
Rejected: keep reactive WARN-only quarantine in each adapter | it leaves repeated user-visible warnings and duplicate recovery behavior.
Confidence: high
Scope-risk: broad
Directive: Do not silently delete large orphan WALs; only quarantine tiny orphan WALs before open and keep large WALs for explicit recovery.
Tested: cd gitnexus && npx vitest run test/unit/sidecar-recovery.test.ts test/unit/lbug-adapter-wal-schema.test.ts test/unit/pool-wal-recovery.test.ts test/unit/web-ui-serving.test.ts && npx tsc --noEmit
Not-tested: full npm test in this split branch; full unit suite passed on the source branch before PR split.

Co-authored-by: OmX <omx@oh-my-codex.dev>

* fix(lbug): pool-caller ENOENT guard, symmetric size gate, permission-aware errors (PR #1747 review)

Addresses the production-readiness review of PR #1747 (Findings 1, 2, 3 of 6).
Findings 4, 5, 6 are deferred to follow-ups per the plan.

1. ENOENT-tolerance scoped to pool-adapter callers only
   - `quarantineWalForMissingShadow` stays strict in `sidecar-recovery.ts`.
     The direct adapter calls it inside `acquireInitLock` (cross-process
     file lock) — ENOENT there means the file vanished under lock and
     remains a real bug to surface.
   - New `tryQuarantineForMissingShadow` local helper in `pool-adapter.ts`
     returns a discriminated union { kind: 'quarantined', path } |
     { kind: 'peer-handled' }. Catches ENOENT, re-verifies via
     statIfExists, and converts to 'peer-handled' only when WAL really
     is gone. Defensive: if ENOENT but WAL still present, throws as
     classified error rather than silently returning success.

2. Symmetric WAL-size gate on both recovery paths
   - `refuseLargeWalQuarantine` applied in both
     `reopenReadOnlyAfterMissingShadow` and
     `reopenWritableAfterMissingShadow`. Closes the read-only data-loss
     vector (large orphan WAL silently discarded would never be replayed
     by a later writable open).

3. Permission-aware error classifier
   - New `renameFailureMessage` and `isPermissionRenameError` in
     `sidecar-recovery.ts`. EACCES / EPERM / EBUSY now surface a
     permission-specific message pointing at ACLs, AV exclusions, and
     file-locks. Other codes (ENOSPC, EROFS, EIO, ENOENT) fall through
     to `shadowSidecarRecoveryMessage`.
   - Used at both pool-adapter and direct-adapter caller catches around
     `quarantineWalForMissingShadow`.
   - `doInitLbug`'s pass-through classifier extended to include the new
     permission message. The lock-retry substring match tightened so
     "file-lock error" in the permission message is not mistaken for a
     LadybugDB lock-retry trigger.

Tests
   - sidecar-recovery.test.ts: 7 new tests for `renameFailureMessage` and
     `isPermissionRenameError`.
   - pool-wal-recovery.test.ts: 6 new tests covering ENOENT race,
     EACCES/EPERM/EBUSY classification, ENOSPC fallthrough, and the
     defensive "WAL still present after ENOENT" branch.
   - lbug-adapter-wal-schema.test.ts: 5 new tests covering the symmetric
     size gate on both recovery paths, including the boundary at exactly
     TINY_ORPHAN_WAL_BYTES (4096) and the off-by-one at 4097.

Deferred (tracked as follow-up work)
   - Brittle LadybugDB error-string matching (Finding 4).
   - PNA header end-to-end coverage gap (Finding 5).
   - warnedKeys module-global persistence (Finding 6).
   - Cross-process init lock for pool-adapter.

* fix(lbug): dedup shadow-replay predicate + counter-based warn anti-spam (PR #1747 review, Findings 4 & 6)

Smallest viable response to the two remaining non-blocking findings from the
production-readiness review of PR #1747. An earlier-revision plan proposed
regex widening + a near-miss detector + per-dbPath warn scoping; an
adversarial doc-review found those defended against hypothetical strings
LadybugDB does not produce, added observability theater with no recovery
behavior change, and did not actually fix the long-running gitnexus serve
case for hot dbPaths (where finalizeLbugSidecarsAfterClose rarely fires).
Scope shrunk to dedup + counter-based — strictly behavior-changing and
fully testable.

Finding 4 — dedup + version-coupling markers
   - `isReadOnlyShadowReplayError` was inlined in both `lbug-adapter.ts:451`
     and `pool-adapter.ts:317`. Centralized as an export from
     `sidecar-recovery.ts`. The two local copies are removed; both adapters
     now import from the shared module.
   - Both LadybugDB-coupled predicates (`isMissingShadowSidecarError` and
     `isReadOnlyShadowReplayError`) gain a `// LADYBUGDB-CONTRACT:` marker
     comment citing `@ladybugdb/core ^0.16.1`. When bumping LadybugDB,
     `git grep "LADYBUGDB-CONTRACT"` enumerates every version-coupled spot.
   - Strict matcher unchanged — when LadybugDB actually changes the error
     format, the failure mode stays loud (raw native error propagates) and
     the markers make every affected predicate trivially greppable.

Finding 6 — counter-based warn anti-spam
   - `warnedKeys: Set<string>` → `warnedKeyCounts: Map<string, number>`.
     `warnOnce` keeps its signature `(logger, key, message)` and keying
     convention unchanged — the swap is internal.
   - `WARN_MILESTONES = [1, 10, 100, 1000, 10000]`. Logarithmic spacing
     gives O(log N) warns for a condition that fires N times. Past the
     first occurrence the warn message is suffixed with "(Nth occurrence
     of this condition)" so persistence is visible in the log line itself.
   - Solves the long-running serve case: a hot dbPath hitting the same
     condition 100 times now fires 3 warns (occurrences 1, 10, 100)
     instead of 1 warn + 99 silent debug lines.

Tests (10 new in sidecar-recovery.test.ts, all green)
   - Centralized isReadOnlyShadowReplayError: positive match, false-positive
     guard, structural assertion that the duplicate regex is gone from both
     adapter files, LADYBUGDB-CONTRACT marker count.
   - Counter-based warnOnce: milestone-at-10 with suffix, milestone-at-100,
     key isolation across dbPaths, reset zeroes the counter, first-occurrence
     message does NOT carry the suffix.

Deferred (tracked separately)
   - Finding 5 — PNA header end-to-end coverage gap (CORS boundary is sound).
   - LadybugDB structured error codes (if/when the library exposes them).
   - Per-call milestone configurability — re-open if tuning is needed.

* chore(autofix): apply prettier + eslint fixes via /autofix command

* ci: trigger CI rebuild

---------

Co-authored-by: wangxc <wangxc_a_bj@si-tech.com.cn>
Co-authored-by: OmX <omx@oh-my-codex.dev>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-21 12:35:43 +01:00
dependabot[bot] 8d71847791 chore(deps)(deps): bump @tailwindcss/vite in /gitnexus-web (#1734)
Bumps [@tailwindcss/vite](https://github.com/tailwindlabs/tailwindcss/tree/HEAD/packages/@tailwindcss-vite) from 4.2.4 to 4.3.0.
- [Release notes](https://github.com/tailwindlabs/tailwindcss/releases)
- [Changelog](https://github.com/tailwindlabs/tailwindcss/blob/main/CHANGELOG.md)
- [Commits](https://github.com/tailwindlabs/tailwindcss/commits/v4.3.0/packages/@tailwindcss-vite)

---
updated-dependencies:
- dependency-name: "@tailwindcss/vite"
  dependency-version: 4.3.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:07:41 +01:00
dependabot[bot] 061e123d72 chore(deps): bump actions/dependency-review-action from 4.9.0 to 5.0.0 (#1739)
Bumps [actions/dependency-review-action](https://github.com/actions/dependency-review-action) from 4.9.0 to 5.0.0.
- [Release notes](https://github.com/actions/dependency-review-action/releases)
- [Commits](https://github.com/actions/dependency-review-action/compare/2031cfc080254a8a887f58cffee85186f0e49e48...a1d282b36b6f3519aa1f3fc636f609c47dddb294)

---
updated-dependencies:
- dependency-name: actions/dependency-review-action
  dependency-version: 5.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:07:22 +01:00
dependabot[bot] a4954368ad chore(deps): bump release-drafter/release-drafter from 7.2.1 to 7.3.0 (#1740)
Bumps [release-drafter/release-drafter](https://github.com/release-drafter/release-drafter) from 7.2.1 to 7.3.0.
- [Release notes](https://github.com/release-drafter/release-drafter/releases)
- [Commits](https://github.com/release-drafter/release-drafter/compare/563bf132657a13ded0b01fcb723c5a58cdd824e2...c2e2804cc59f45f57076a99af580d0fedb697927)

---
updated-dependencies:
- dependency-name: release-drafter/release-drafter
  dependency-version: 7.3.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:07:07 +01:00
dependabot[bot] be1071143a chore(deps)(deps): bump react-syntax-highlighter in /gitnexus-web (#1731)
Bumps [react-syntax-highlighter](https://github.com/react-syntax-highlighter/react-syntax-highlighter) from 16.1.0 to 16.1.1.
- [Release notes](https://github.com/react-syntax-highlighter/react-syntax-highlighter/releases)
- [Changelog](https://github.com/react-syntax-highlighter/react-syntax-highlighter/blob/master/CHANGELOG.MD)
- [Commits](https://github.com/react-syntax-highlighter/react-syntax-highlighter/compare/v16.1.0...v16.1.1)

---
updated-dependencies:
- dependency-name: react-syntax-highlighter
  dependency-version: 16.1.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:06:41 +01:00
dependabot[bot] 3d8aa7f435 chore(deps): bump github/codeql-action from 4.35.3 to 4.35.4 (#1738)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4.35.3 to 4.35.4.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/e46ed2cbd01164d986452f91f178727624ae40d7...68bde559dea0fdcac2102bfdf6230c5f70eb485e)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.35.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:06:15 +01:00
dependabot[bot] 4606e24f25 chore(deps)(deps): bump dompurify from 3.4.2 to 3.4.3 in /gitnexus-web (#1735)
Bumps [dompurify](https://github.com/cure53/DOMPurify) from 3.4.2 to 3.4.3.
- [Release notes](https://github.com/cure53/DOMPurify/releases)
- [Commits](https://github.com/cure53/DOMPurify/compare/3.4.2...3.4.3)

---
updated-dependencies:
- dependency-name: dompurify
  dependency-version: 3.4.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:06:02 +01:00
74653a8ffc feat(web): Support GitLab repository urls. (#1565)
Add GitLab URL input mode alongside existing GitHub and local modes:
- GitLab URL validation for gitlab.com and self-hosted instances
- GitLab icon component (custom SVG, matching existing GitHub icon pattern)
- Mode tab UI with GitLab option
- Backend API integration for GitLab HTTPS URLs

No token configuration included — public repositories supported only.

Close: #378


AI-model: kimi-for-coding/k2p6

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-21 10:52:01 +01:00
Gergő MagyarandCursor 8db51184ab fix(server): restore gitnexus serve startup under Express 5 (#1749)
* fix(server): restore gitnexus serve startup under Express 5

Express 5 rejects app.options('*'), which broke CI e2e when the backend
failed to start. Move PNA middleware before cors so preflight responses
include Access-Control-Allow-Private-Network, and add regression tests.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(server): address PR review — prettier, ephemeral port, cleanup

- Format integration and rate-limit test files for CI quality/format
- Use OS-assigned port instead of random 47xxx range
- Remove per-test GITNEXUS_HOME temp dir in afterEach
- Use regex for PNA-before-cors structural guard (indent-agnostic)

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-21 10:18:09 +01:00
1b5c6e5b6a feat(ingestion): add Kotlin scope resolver (#1727)
* feat(ingestion): add Kotlin scope resolver

* fix(ingestion): tighten Kotlin scope captures

---------

Co-authored-by: Shining <xuenning@qiyi.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-21 08:52:23 +01:00
CopilotandGergő Magyar c34c36036f fix(workers): resilient + zero-copy ingestion worker pool — prevent analyze hangs on TS-root-scale loads (#1693)
* Initial plan

* fix: skip worker-timeout files in sequential fallback and optimize TS capture node lookup

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0e53743e-0600-4690-bd0d-198894daef58

* refactor: clarify TS capture helpers after validation feedback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0e53743e-0600-4690-bd0d-198894daef58

* fix(workers): exclude in-flight file on worker error/exit, not just singleton timeout

WorkerPoolDispatchError previously surfaced the stalled path only for the
singleton-timeout final-fail branch. Worker `error` and `exit` events (and
the msg-channel `error` reply) fell back to plain `Error`, so the sequential
fallback re-attempted every file in the active job — re-hanging on the same
pathological file when the worker crashed mid-parse.

Lift the in-flight-file inference into `inFlightExcludePath(job, lastProgress)`
and wire it into the three remaining in-pool failure sites. `lastProgress` is
already in `runWorker` scope, so `items[lastProgress]` (the next file the
worker was about to acknowledge) is the best single guess at the culprit;
earlier files are still re-tried sequentially. Returns `[]` when no path is
determinable (`lastProgress >= items.length`, or path missing/non-string) so
sequential retries the whole job.

Replacement-worker startup failures stay plain `Error` (no job context); the
result-before-flush protocol bug stays plain `Error` (code fault, not file).

Tests cover the three new exclusion paths plus a negative test confirming
non-WorkerPoolDispatchError throws fall through to full sequential retry.

* fix(review): apply autofix feedback

- Use cause-neutral "worker-excluded" label in skip messages and tests now
  that worker error/exit paths share the same exclusion contract as
  singleton-timeout (correctness + maintainability reviewers).
- Add JSDoc to findSelfOrAncestorOfType{s} explaining the parent-walk
  short-circuit vs root-DFS fallback (maintainability reviewer).

* feat(workers): resilient + scalable worker pool

Restructures `createWorkerPool` so a single bad file no longer kills the
pool for the rest of an analyze run. Five interlocking layers:

1. **Auto-respawn on error/exit** — worker death triggers `replaceWorker`
   on the same slot, bounded by `maxRespawnsPerSlot` (default 3). The slot
   is dropped from rotation when the budget is exhausted; other slots
   keep running.

2. **Circuit breaker** — replaces the permanent `poolBroken=true` with a
   consecutive-failure counter. The pool only trips after
   `consecutiveFailureThreshold` deaths (default `max(3, poolSize)`) with
   no successful job in between. A successful job resets the counter so
   transient bursts of bad files don't escalate.

3. **Session-scoped file quarantine** — paths identified as the in-flight
   file at the moment of a worker death are added to a `Set<string>` on
   the pool. `dispatch()` filters quarantined items up front (they never
   reach a worker again this pool lifetime). Exposed via the new
   `WorkerPool.getQuarantinedPaths()` so callers can log/route them.
   `processParsing` surfaces the per-chunk quarantine summary alongside
   the existing fallback-exclusion log.

4. **Authoritative in-flight tracking** — `parse-worker.ts` emits
   `{type:'starting-file', path}` before each file. The pool tracks this
   per slot and uses it for crash attribution, falling back to the
   `items[lastProgress]` heuristic only when no starting-file has been
   observed (very-early crash, older worker build). Closes the
   reorder/race concerns raised by reviewers C1 and R3 in the earlier
   review run.

5. **Per-job cumulative timeout budget** — each `WorkerJob` tracks the
   total wall time spent across attempts/splits/retries. When the budget
   is exhausted (default 5x `subBatchIdleTimeoutMs`), the pool surfaces
   the in-flight path instead of letting exponential backoff balloon
   into multi-hour stalls.

Cross-layer wiring: a new `wakeIdleSlots` helper kicks any non-busy live
slot when items are requeued (after a death or split-retry), so a dropped
slot doesn't strand work in the queue. `recoverAndResume` consolidates
the per-job teardown shared by the three in-pool death sites (`error`,
`exit`, msg-channel `error`).

New env knobs: `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT`,
`GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS`,
`GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD`.
New `WorkerPoolOptions.workerFactory` injection point for unit tests.

Tests: 12 new unit tests using a FakeWorker mock cover quarantine
seeding, slot-respawn, slot-drop after budget, breaker trip + reset,
and quarantine filtering. Plus option-resolution tests for the three
new env vars. All 19 worker-pool/-fallback/-options tests pass; full
unit suite 6040 passed / 30 skipped / 0 failed.

* fix(workers): apply code-review fixes (12 findings)

Walks through every finding from ce-code-review run
20260519-094648-3549cf5e. All 12 picked Apply.

Critical:
- F1 — Layer 5 cumulative-timeout exhaustion no longer silently drops
  the rest of the job. `requeueRemainder` is now invoked before
  `handleWorkerDeath` in both Layer 5 and singleton-final-fail give-up
  paths so non-quarantined items get re-tried by another worker.
- F2 — idle-timer recovery overhaul. `!shouldContinue` branch no
  longer calls `replaceWorker` (double-spawn race with the
  `handleWorkerDeath` inside `requeueAfterTimeout`). `shouldContinue`
  branch now enforces `maxRespawnsPerSlot` before respawning, closing
  the budget-bypass for the timeout-retry path. Also fixes premature
  `maybeDone` by simplifying the bookkeeping.
- F3 — `requeueRemainder` no longer pre-charges `cumulativeTimeoutMs`
  by `job.timeoutMs`. The death itself consumed no budget, so the
  next `requeueAfterTimeout` was double-billing the first attempt.
- F4 — `WorkerPool.getQuarantinedPaths` is now optional on the
  interface, matching the defensive `?.()` call site and the existing
  mocks. Removes the contract-vs-callsite contradiction.
- F5 — per-job unattributed-death tracking. When a worker dies with
  no exclusion attribution, `requeueRemainder` tracks death count per
  `startIndex`. First time: re-queue intact. Second time: quarantine
  items[0] as best guess, or drop the job entirely when items lack
  paths. Bounds the death loop the original design admitted to.
- F6 — per-slot consecutive-failure counter. Replaces the pool-wide
  scalar so a chronically-failing slot trips the breaker on its own
  streak instead of being masked by another slot's successes.

Smaller:
- F7 — exhaustiveness `never` check on `WorkerOutgoingMessage` union.
- F8 — recursive `runWorker` on fully-quarantined jobs converted to
  a while-loop.
- F9 — `tripBreaker` calls `reject(err)` BEFORE awaiting
  `worker.terminate()`. A stuck terminate no longer blocks the caller.
- F10 — `parsing-processor.ts` quarantine log de-duplicates per pool
  instance via a `WeakMap`. Only newly-quarantined paths are logged
  in each chunk; the per-chunk count still surfaces via progress.
- F11 — extract `firstPath` local in `requeueAfterTimeout`; eliminates
  double `itemPath` call and the `unknown as string` cast.

Tests (F12, 6 new):
- crash-error event path (errorHandler).
- F5 drop-branch coverage via items without `.path`.
- Common-case unattributable crash falling back to items[0] heuristic.
- `replaceWorker` startup failure (workerFactory emits 'exit' before
  'online').
- All-slots-dropped breaker trip.
- `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` env override.

Residual gap (deferred): no unit test exercises the Layer 5
cumulative-budget runtime path — requires fake-timer interleaving
with FakeWorker that's too brittle for this iteration. Tracked.

Unit suite: 257 files / 6056 passed / 30 skipped / 0 failed.

* test(workers): integration tests for resilience layers + fix requeue-after-timeout flow

Adds 6 new real-worker integration tests covering the PR #1693
resilience layers + fixes 3 follow-on bugs surfaced while writing them.

New integration coverage (real worker threads + temp fixture scripts):

- `respawns the slot after worker process.exit and finishes the work on
  the replacement` — exercises Layer 1 auto-respawn + Layer 3 quarantine
  through real IPC.
- `attributes exactly via authoritative starting-file message on worker
  crash` — Layer 4 end-to-end: starting-file message → exact quarantine
  attribution (not the items[0] heuristic).
- `quarantine filters subsequent dispatches without sending to a worker`
  — second dispatch's sub-batch payload audited via filesystem; the
  quarantined path is never sent across the message channel.
- `drops a slot after maxRespawnsPerSlot and continues on the survivor`
  — 2-slot pool, slot dies twice past budget, survivor finishes
  re-queued remainder.
- `trips the circuit breaker on cascading per-slot consecutive failures`
  — single-slot pool, dies on every job, breaker trips after
  consecutiveFailureThreshold with WorkerPoolDispatchError carrying
  the cumulative quarantine.
- `survives a worker error event (uncaught throw) the same as a
  process.exit` — validates recoverAndResume on the errorHandler path
  via a real worker `throw` (not just process.exit).

Bug fixes uncovered while writing these tests:

1. **Stack-overflow recursion in runWorker's no-worker branch** —
   `if (!worker) { ...; wakeIdleSlots(); maybeDone(); }` recursed
   indefinitely when multiple slots were mid-respawn simultaneously
   (wakeIdleSlots → runWorker → no worker → wakeIdleSlots → …).
   Removed the wakeIdleSlots call: the slot's own respawn IIFE owns
   runWorker post-respawn, and other slots will pick up work via
   finishJob's runWorker.

2. **requeueAfterTimeout dispatched work before respawn completed** —
   the F2 fix had `requeueAfterTimeout` `void`-discarding
   `handleWorkerDeath`, so the `!shouldContinue` IIFE had no way to
   know when the respawn finished. New design: `requeueAfterTimeout`
   returns a `TimeoutDecision` discriminated union; the IIFE owns
   the death-and-respawn-and-dispatch orchestration in an async
   closure so it can `await handleWorkerDeath` and then call
   `runWorker` deterministically.

3. **Stalled-singleton + protocol-error + replacement-startup-crash
   tests** had stale contracts predating the resilience refactor. The
   stalled-singleton no longer rejects (it quarantines + resolves
   `[]`); the protocol-error rejection message now mentions
   "circuit breaker tripped"; the replacement-startup-crash test
   documents the known `waitForWorkerOnline` race (online fires
   before the worker's main script runs, so a top-level throw looks
   like a successful spawn) — the test asserts the file is
   quarantined via the second-idle-timeout give-up path.

Full suite: 334 files / 8982 passed / 43 skipped / 0 failed (second
run; first run had a Vitest-reported flake from an uncaught worker
exception bleeding into the test report — repeated runs are clean).

* perf(workers): raise pool cap to cores-1 + defer per-chunk extraction to keep workers busy

User reported 4-5% CPU utilization on a multi-core machine during
ingestion. Two structural reasons:

1. **Pool cap.** `createWorkerPool` resolved size as
   `Math.min(8, max(1, os.cpus().length - 1))` — a 16-core box got 8
   workers (50% theoretical max). U1 lifts the default to
   `min(16, max(1, cores - 1))`, exposes `GITNEXUS_WORKER_POOL_SIZE`
   env override, and adds `--workers <N>` CLI flag (`0` disables the
   pool for sequential fallback).

2. **Per-chunk extraction serialized the loop.** Per chunk:
   dispatch → await workers → main-thread `processImportsFromExtracted`
   + `processHeritageFromExtracted` + `processRoutesFromExtracted`
   + `synthesizeWildcardImportBindings` + `seedCrossFileReceiverTypes`
   → next chunk dispatch. Workers sat idle through every extraction
   block. U2 (revised from the plan's pipelined-chunks design) defers
   these passes to a single end-of-loop batch. Chunk loop becomes
   parse + merge + accumulate. Resolution sees strictly-more-info
   (full repo graph) so cross-chunk import/heritage targets resolve at
   least as well as before. Memory cost: `deferredWorkerImports`
   accumulates across chunks; bounded by total file count, acceptable.

Plan deviation note: the plan called for an in-flight chunk pipeline
(N concurrent dispatches with bounded memory). That design needed
either a `processParsing` API refactor or duplicating its catch-block
fallback in `parse-impl`. The deferred-extraction approach delivers
the same "workers stay busy" outcome with much smaller surface area
and zero changes to `processParsing`. The `GITNEXUS_PARSE_CHUNK_CONCURRENCY`
env var documented in U2 of the plan is therefore not implemented in
this commit; if memory growth from `deferredWorkerImports` becomes
a problem at very-large-repo scale, a bounded sliding-window variant
can land as a follow-up.

Tests:
- New `test/unit/analyze-worker-pool-size.test.ts` covers --workers
  validation (5 invalid inputs rejected with exit code 1 + clear
  error; valid integers set the env var; `--workers 0` routes to
  sequential).
- Extended `worker-pool-resilience.test.ts` with `resolveAutoPoolSize`
  scenarios: env override, env=0, env above cap, invalid env fallback,
  auto-formula match, integer return type.
- Full unit suite: 6097 / 6127 passed / 30 skipped / 0 failed.
- Full integration suite (second run): 77 / 78 passed / 1 skipped /
  0 failed. First run had a known cosmetic flake from an uncaught
  worker exception bleeding into the test reporter.

Resilience contract from PR #1693 preserved: per-slot respawn budget,
circuit breaker, quarantine, authoritative in-flight tracking,
cumulative timeout budget — all unchanged.

New env vars surfaced in --help: GITNEXUS_WORKER_POOL_SIZE,
GITNEXUS_PARSE_CHUNK_CONCURRENCY (reserved for future bounded
pipelining).

* docs(readme): document --workers CLI flag

* feat(workers): add getStats() and per-chunk throughput logging

* test(workers): cleanup leaked temp-dirs and drop duplicate option-resolution block

- Add afterEach to worker-pool-resilience.test.ts cleaning up the per-test temp
  directory created by beforeEach (~25 stale dirs per CI run previously).
- Delete the duplicated describe('worker pool option resolution', ...) block.
  Verified the first block (lines 490-532) is a strict superset (includes the
  GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS env test the second block omitted),
  so deletion loses no test coverage.

Addresses PR #1693 review findings L2 (temp-dir leak) and L3 (duplicate block).

* feat(cli): thread --workers via PipelineOptions + snapshot/restore CLI env

Resolves PR #1693 review B2 (env-var leak in long-running hosts):

- --workers is now threaded through AnalyzeOptions -> runFullAnalysis
  -> PipelineOptions.workerPoolSize -> createWorkerPool's explicit
  poolSize arg, bypassing the GITNEXUS_WORKER_POOL_SIZE env channel.
  The env var remains as a back-compat fallback inside resolveAutoPoolSize
  for operators who set it directly.
- analyzeCommand and wikiCommand snapshot the GITNEXUS_* env vars they
  mutate at function entry and restore them in finally. Inner *Impl
  extraction keeps the diff surgical (no body re-indent). process.exit(0)
  on the CLI success path still terminates the process; restoration
  matters for programmatic callers (tests, long-running hosts) reaching
  early-return paths or the alreadyUpToDate fast path.
- Tests updated to assert the new behavior:
    analyze-worker-pool-size.test.ts: workerPoolSize flows through
      runFullAnalysis options; env is not mutated; back-to-back calls
      see their own values, not the previous call's leak.
    analyze-worker-timeout.test.ts: env IS set during the runFullAnalysis
      call (captured via mockImplementation) and restored after, proving
      the timeout reaches downstream while the leak fix holds.
- Also addresses L4: afterEach NODE_OPTIONS restore so back-to-back test
  runs don't accumulate --max-old-space-size=8192 tokens.

Addresses PR #1693 review B2 (blocker) and L4 (test polish).

* feat(workers): harden worker lifecycle (messageerror + availableParallelism + ready handshake)

Resolves PR #1693 review H1, H2, M4:

H1 - messageerror handler at every dispatch site
  V8 deserialization failure on postMessage previously left the message
  silently lost; the pool would wait out the idle timeout (default 30s)
  instead of treating it as worker death. The dispatch loop now wires
  worker.once('messageerror', ...) alongside error/exit and routes through
  recoverAndResume so the existing per-slot respawn budget, in-flight
  file attribution, and circuit-breaker layers fire as designed.

H2 - resolveAutoPoolSize uses os.availableParallelism()
  Mirrors the pattern at capabilities.ts:85 (defaultEmbeddingThreads).
  os.cpus().length returns the host CPU count, which over-sizes the pool
  on cgroup-limited containers, taskset-restricted runtimes, and CI
  runners with explicit CPU quotas. Falls back to os.cpus().length on
  Node < 18.14.

M4 - worker-side ready handshake replaces online-trust
  parse-worker.ts now emits {type: 'ready'} after all top-of-script
  initialization completes, BEFORE the message handler is attached. The
  pool's renamed waitForWorkerReady listens for this message under a
  bounded WORKER_READY_TIMEOUT_MS (5s) budget instead of trusting Node's
  online event - which fires when the worker thread starts, BEFORE the
  script body runs, letting init crashes slip past pool startup. ready
  is added to WorkerOutgoingMessage with an exhaustiveness-checked
  no-op branch in the dispatch handler (defensive: the message is
  consumed by waitForWorkerReady before dispatch handlers attach).
  messageerror is wired into waitForWorkerReady the same way.

Test scaffolding:
  - FakeWorker emits {type: 'ready'} in addition to 'online' so
    replacement workers in unit tests don't hit the 5s budget.
  - Integration test ad-hoc worker scripts go through a writeReadyWorker
    helper that prepends the ready handshake. Tests intending to script
    "crash BEFORE ready" can bypass the helper.

61/61 worker-pool unit tests pass; 28/28 integration tests pass.

* feat(parse-impl): monotonic progress + verbose-gated throughput log + seed-before-build

Resolves PR #1693 review M2, M3, L1, L5 in a single parse-impl.ts pass:

M2 - Monotonic progress through deferred phase (no more "stuck at 82%")
  Previously the deferred resolution stages (imports, heritage, routes,
  calls) all emitted percent: 82 — the UI looked frozen for the duration
  of the deferred work, which on large repos is several seconds to minutes
  and visually identical to the hang PR #1693 set out to fix.
  Redistributed:
    parse phase:  20-70 (was 20-82)
    imports:      70-75
    heritage:     75-80
    routes:       80-85
    calls:        85-95
  Each deferred stage now advances through its own band via the existing
  per-batch progress callback. Skipped stages (zero deferred input) leave
  their band as a no-op jump - the next stage still starts at its own
  band, preserving strict monotonicity. The "no parseable files" early
  return now jumps to 95 (was 82), and the duplicate "Parsing N files..."
  announcement is suppressed when totalParseable === 0 to avoid a
  non-monotonic 95 -> 20 regression that pre-existed (uncovered by the
  new monotonic test).

M3 - Throughput log gated on `--verbose`, not just NODE_ENV=development
  The per-chunk files/s log was gated on `isDev`, so operators running
  `gitnexus analyze --verbose` in a production install never saw it.
  Now fires when (isDev || isVerboseIngestionEnabled()) — matches the
  documented promise that `--verbose` shows tuning observability.

L1 - Typo rename: `chunkChunkStartMs` -> `chunkStartMs`

L5 - `buildExportedTypeMapFromGraph` runs BEFORE `seedCrossFileReceiverTypes`
  Previously the seeding branch was reached with `exportedTypeMap.size === 0`
  in the worker path (the map was only built far below, AFTER the seeding
  branch), so the seed dead-coded itself silently and call resolution
  never got the cross-file receiver-type enrichment. Now the map is
  populated from the in-progress graph before the seed call; the
  post-parse builder remains as a defensive sequential-path fallback,
  guarded by `size === 0` so we don't pay the cost twice on the worker
  path. Net win: cross-file CALLS edges that previously had no receiver
  type now get enriched.

New test: parse-impl-progress-monotonic.test.ts
  Asserts the emitted percent stream is strictly non-decreasing across
  the parse + deferred phases, and that the deferred band (>=70) is
  actually reached. Also pins the "no parseable files" path to exactly
  [95] so the 95 -> 20 regression we just fixed can't re-emerge.

* feat(parse-impl): bounded chunk concurrency via file-pre-fetch pipeline

Resolves PR #1693 review B1 (GITNEXUS_PARSE_CHUNK_CONCURRENCY documented
in --help but unimplemented).

The chunk loop now pre-fetches chunk file contents up to
`parseChunkConcurrency` chunks ahead of the worker-dispatch cursor so
disk I/O overlaps with worker compute. Worker dispatch itself stays
serial because WorkerPool.dispatch is not reentrant — concurrent calls
would race on the shared per-slot busy/in-flight state, regressing the
hang/resilience work this PR is built on. The pre-fetch path is the
honest interpretation of "concurrent in-flight parse chunks" that the
help text advertises: I/O overlap, not parallel worker dispatch.

Concurrency value resolution:
  1. PipelineOptions.parseChunkConcurrency (threaded from CLI)
  2. GITNEXUS_PARSE_CHUNK_CONCURRENCY env var
  3. Default 2 (matches the help text)

F4 (wildcard-synthesis ordering) is preserved: deferred-state
aggregation runs in chunkIdx order because the for-loop iterates
sequentially after awaiting each chunk's pre-fetched contents.
Cross-chunk processors (processImportsFromExtracted,
synthesizeWildcardImportBindings, etc.) still run only after all
chunks complete — they see deterministic input regardless of
file-read completion order.

Concurrency=1 produces behavior identical to the pure-serial loop;
that's the regression baseline.

New test: parse-impl-chunk-concurrency.test.ts
  - Asserts graph output is identical (nodeCount + relationshipCount)
    between parseChunkConcurrency=1 and =2 — the critical correctness
    invariant. Exact .toBe(N) comparisons per DoD §2.7 (the second run's
    counts must equal the first run's exactly).
  - Pins specific fixture symbols (foo/bar/Baz) under both
    parseChunkConcurrency=1 and the env-fallback (3) path.
  - Env-fallback test confirms GITNEXUS_PARSE_CHUNK_CONCURRENCY is
    honored when the option is undefined.

* test(workers): pin cumulative-timeout exhaustion behavior

Resolves PR #1693 review M6: the existing resilience suite asserts only
the *default value* of maxCumulativeTimeoutMs (5x subBatchIdleTimeoutMs),
not that dispatch actually aborts the offending job when the cumulative
wall-clock budget is exhausted. Without this test, a future refactor
could remove the exhaustion branch in requeueAfterTimeout and the suite
would stay green while the pool sat in retry loops for an hour on a
real production stall.

Scenario:
  subBatchIdleTimeoutMs    = 100ms
  timeoutBackoffFactor     = 10
  maxCumulativeTimeoutMs   = 300ms

Single file, HangingWorker that never responds. First attempt times
out at 100ms (cumulative=100). The next backoff (1000ms, cumulative
1100ms) exceeds the 300ms cap, so requeueAfterTimeout returns
give-up on the first timeout retry and the file goes to the session
quarantine. Asserts:
  - pool.getQuarantinedPaths() includes 'src/stuck.ts' after dispatch
  - if dispatch rejected, the error is a WorkerPoolDispatchError
    (the typed surface that routes to sequential fallback)

Uses a local minimal HangingWorker double rather than the full
action-scripted FakeWorker from worker-pool-resilience.test.ts —
the inverse pattern (always hang) doesn't need the scripted-action
machinery and keeps the test file focused on the one behavior.

* docs(readme): add environment-variables reference table

Resolves PR #1693 review L6: operator-facing env vars were either
mentioned inline (GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS) or only
documented via `gitnexus --help`, with no single place to look up
the full set. The new "Environment variables" subsection under the
Quick Start CLI block lists every operator-facing knob with default,
effect, and tuning guidance, matching the names in cli/index.ts
addHelpText post-U2 / U1.

Covers:
  GITNEXUS_WORKER_POOL_SIZE           (--workers)
  GITNEXUS_PARSE_CHUNK_CONCURRENCY    (newly real per U1)
  GITNEXUS_VERBOSE                    (--verbose)
  GITNEXUS_MAX_FILE_SIZE              (--max-file-size)
  GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS (--worker-timeout × 1000)
  GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES
  GITNEXUS_CHUNK_BYTE_BUDGET
  GITNEXUS_NO_GITIGNORE
  GITNEXUS_SKIP_OPTIONAL_GRAMMARS

CLI flag vs env-var precedence is stated explicitly (CLI > env > default)
so operators running long-lived hosts (MCP server, eval-server) know
which channel wins.

* test(workers): pin quarantine path round-trip and non-normalization contract

Resolves PR #1693 review M5 (Windows quarantine path-normalization
coverage). worker-pool.ts quarantines paths via a Set<string> keyed by
exact string equality. The existing suite never asserted this contract,
which lets a future "helpfully normalizing" refactor on one side of the
pipeline (caller, worker, or pool) silently break quarantine filtering
on Windows.

This file pins the contract from both directions:

1. Round-trip: a path the caller dispatches with backslashes
   (src\bad.ts) flows through starting-file -> death -> quarantine ->
   next-dispatch filter verbatim. The replacement worker never sees the
   re-dispatched bad path because the pool's pre-dispatch filter
   short-circuits it.

2. Non-normalization: quarantining src\poison.ts does NOT filter
   src/poison.ts. Whoever changes that contract has to update this test
   alongside (the load-bearing assertion catches accidental
   path.normalize() calls in the quarantine path).

Runs on every platform — the path strings are test-injected, so the
test exercises the same code path regardless of the host's path.sep.
Used a self-contained FakeWorker that emits {type:'ready'} for U3's
waitForWorkerReady handshake, so the test doesn't depend on the larger
worker-pool-resilience.test.ts harness.

* test(typescript): pin capture-anchor rewrite invariants (B5 regression)

Resolves PR #1693 review B5: the captures.ts ancestor-walk rewrite
(findSelfOrAncestorOfType[s] + pickFirstNode replacing the prior
findNodeAtRange-from-root path) was semantically equivalent to its
predecessor per Lane 4 of the production-readiness review, but the
existing typescript-captures.test.ts didn't pin the specific sharp
edges where an over-aggressive walk would silently break captures.
This file does.

Each test exercises a capture class whose anchor type is one the
rewrite explicitly handles:

  - member call obj.foo() -> @reference.call.member (call_expression
    anchor walks to self)
  - dynamic import import("./helper") -> raw @import.dynamic gets
    decomposed by splitImportStatement into @import.statement with
    @import.kind=dynamic + @import.source stripped of quotes
  - JSX <Foo /> in .tsx -> @reference.call.free emitted (TSX query
    pattern, query.ts:899-905) but @declaration.parameter-count is
    NOT synthesized because findSelfOrAncestorOfType('call_expression')
    returns null on a jsx_self_closing_element anchor. Pre-rewrite the
    range lookup also returned null. Pinning this contract catches
    accidental "walk JSX -> outer call" refactors.
  - constructor `new Foo(1,2)` -> @reference.call.constructor (new_expression
    anchor walks to self)
  - named/namespace import + re-export -> @import.statement (one each)
  - class method override -> @declaration.method per class, no collapse
  - member read obj.foo (no call) -> @reference.read.member

All assertions use exact .toBe(N) per DoD §2.7.

* test(parse-impl): pin multi-chunk graph equivalence under deferred extraction

Resolves PR #1693 review B4: the deferred-extraction reorder (moving
processImportsFromExtracted / Heritage / Routes / Wildcard /
ReceiverTypes from per-chunk to end-of-loop) was proven observably
equivalent by Lane 4 of the production-readiness review. Until now,
the existing suite never asserted cross-chunk graph equivalence,
which lets a future refactor that accidentally tightens the per-chunk
vs end-of-loop coupling silently break cross-chunk resolution.

This test forces multi-chunk parsing on a small fixture by setting
GITNEXUS_CHUNK_BYTE_BUDGET=64 BEFORE the parse-impl module loads
(the budget is captured at module load via vi.resetModules — a future
move to function-scope env reads is U14 in Phase 2). Then runs the
same fixture under a 10MB budget (single chunk) and asserts the two
graphs are byte-identical: same nodeCount, same relationshipCount,
exact .toBe(N) per DoD §2.7.

Fixture: 3-file class hierarchy with cross-file inheritance — Animal
(a.ts) -> Dog extends Animal (b.ts) -> makeDog returns Dog (c.ts).
Forces the resolver to chain imports + heritage across chunks. A
second test pins specific symbol names (Animal, Dog, makeDog, speak,
bark) in the multi-chunk graph so a regression in chunk-boundary
resolution surfaces as a missing-symbol failure with a specific
diagnostic instead of a bare count mismatch.

* test(parse-impl): wall-clock integration pinning multi-chunk pipeline (B3)

Resolves PR #1693 review B3 — the final P0/P1 merge blocker. With this
test, all five doc-review blockers (B1-B5) are pinned by regression
coverage.

The PR's headline claim is "analyze no longer hangs on TS-root-shaped
loads". The existing suite pins each resilience layer (worker-pool-
resilience.test.ts), the deferred-extraction equivalence (U7), and
the chunk-concurrency contract (U1). What was missing: a single
end-to-end run that exercises the full chunked parse-and-resolve
path on a multi-chunk fixture, BOUNDED by a wall-clock budget so a
regression that re-introduces the hang fails this test loudly via
timeout rather than slipping past as a count drift.

Implementation:
  - 17-file synthetic fixture: 15 small modules (one function each),
    one "realistic dense" complex.ts (30 functions + class + interface),
    and an index.ts re-exporting them. Forces cross-chunk import
    chains.
  - GITNEXUS_CHUNK_BYTE_BUDGET=64 via vi.resetModules forces multi-chunk
    parsing on the small fixture.
  - Promise.race with 30s timeout: a hang fails as
    "exceeded WALL_CLOCK_BUDGET_MS — likely the hang B3 was meant to
    prevent", not as a bounds-only inequality (DoD §2.7 distinction —
    hang-detector via exception, not regression-mask via inequality).
  - Exact .toBe(true) assertions on specific expected symbols
    (fn0..fn14, Service, Config, configure, describe, complex0/15/29)
    so a silent mid-chunk crash that exits 0 without producing graph
    data also fails this test, not just the hang case.

Scope: runs the sequential-fallback path (skipWorkers: true) because
the full real-worker scenario requires a built dist/parse-worker.js
and ~60s wall-clock per run — appropriate for a CI-integration job,
not vitest. The load-bearing invariants pinned here catch the bulk
of B3's concern; the dist-worker swap is a Phase 2 follow-up
documented in the file header.

* refactor(parse-impl): move chunk-byte-budget env read to function scope

Resolves PR #1693 review F7 / U14: pre-U14, `CHUNK_BYTE_BUDGET` was a
module-load IIFE constant that captured `GITNEXUS_CHUNK_BYTE_BUDGET`
once and froze the value for the module's lifetime. That defeated
per-call option threading (a future
`PipelineOptions.chunkByteBudget` was silently no-op'd because the
function body read the frozen module-level constant) AND forced tests
to use `vi.resetModules` to vary chunk layout. The U7
deferred-extraction test and the U6 multi-chunk integration test
both used the workaround.

After this change:

  - `DEFAULT_CHUNK_BYTE_BUDGET = 2 * 1024 * 1024` stays as a
    module-level constant — purely a default, no env access.
  - `resolveChunkByteBudget(options)` runs per call: option wins,
    then env, then default. Same options-first/env-fallback/default
    pattern as resolveAutoPoolSize and the U1 parseChunkConcurrency
    resolver — keeps the ingestion code's configuration model uniform.
  - `PipelineOptions.chunkByteBudget?` added with documentation that
    threading through options lets long-running hosts (eval-server,
    MCP daemon) size per-call without leaking process.env state
    across analyze invocations.

New test (parse-impl-env-reads.test.ts) pins all four behaviors:
  1. option-first: option present + env present -> option wins
  2. env-fallback: option absent + env present -> env wins
  3. default-fallback: both absent -> 2 MB default
  4. per-call: two back-to-back runs in the same vitest worker with
     different chunkByteBudget option values observe their OWN values,
     proving the module-load freeze is gone (no vi.resetModules in
     this test — that's the invariant being verified).

All four assertions use exact `.toBe(N)` per DoD §2.7. The chunk
count is observed by parsing the `Parsing chunk X/Y` progress message
stream — a stable proxy that doesn't require exposing internal
parse-impl counter state.

Note: U7 and U6 tests still use `vi.resetModules` because they were
written before this change. A follow-up cleanup could simplify those
tests (drop the resetModules dance, pass chunkByteBudget via options),
but they pass as-is so this commit doesn't touch them.

* feat(workers): per-slot generation counter for late-event protection (U12)

Adds a monotonic per-slot generation counter to createWorkerPool's
state. Each successful worker replacement (replaceWorker) bumps the
slot's counter exactly once — atomically with the workers[slotIndex]
swap, so observers (getStats) see the new (worker, generation) pair
consistently. Handler closures in the dispatch loop capture the
slot's generation at attach time and short-circuit when they fire
on a stale generation.

In the current implementation, cleanup() synchronously removes
listeners on a Worker instance the moment a death is observed, so
no listener naturally fires on a stale generation — the guard is a
defensive layer protecting against any future refactor that loosens
cleanup() ordering or re-attaches handlers across the swap. The
load-bearing observable is the slotGenerations[] array exposed via
WorkerPoolStats so operators (and tests) can confirm a slot was
actually replaced and not just the same worker recycled.

Implementation:
  - const slotGenerations: number[] = new Array(size).fill(0) in
    createWorkerPool's per-pool state, alongside respawnCount and
    consecutiveFailuresPerSlot.
  - replaceWorker: slotGenerations[workerIndex]++ AFTER the
    workers[workerIndex] = replacement swap (only on the success
    branch — drop-slot paths leave the counter unchanged).
  - runWorker dispatch loop: const slotGen = slotGenerations[workerIndex]
    captured before handler attachment; every handler (handler /
    errorHandler / exitHandler / messageErrorHandler) starts with
    `if (slotGenerations[workerIndex] !== slotGen) return`.
  - WorkerPoolStats gains `readonly slotGenerations: readonly number[]`.
  - getStats() returns slotGenerations.slice() so callers can't mutate
    pool state by writing to the returned array.

Two existing toEqual snapshots in worker-pool-resilience.test.ts
extended with the new slotGenerations field (both expect all-zeros —
neither test scenario triggers a respawn).

New test file (worker-pool-slot-generation.test.ts, 4 tests):
  1. Fresh pool: every slot at generation 0.
  2. Successful crash + respawn: generation bumps to 1 exactly once.
  3. Crash that drops the slot (maxRespawnsPerSlot:0): generation
     stays at 0 because no successful respawn happened. The dispatch
     rejection on breaker trip is the expected outcome here; the
     load-bearing assertion is the post-rejection stats.
  4. Multi-slot independence: one slot crashing bumps only that
     slot's generation, not the other. Order-independent via sort()
     because the round-robin assignment isn't pinned by contract.

All assertions exact .toEqual / .toBe per DoD §2.7.

* docs(bench): add parse-throughput benchmark scaffold (R13)

Resolves PR #1693 review R13 (benchmark artifact requirement).

Creates `gitnexus/bench/parse-throughput.md` documenting:

- Synthetic fixture spec (same shape as the U6 integration test, so
  CI smoke baseline and ad-hoc benchmark exercise the same paths).
- What to measure (wall-clock, peak heap, chunk count, getStats
  snapshot) and the hardware-shape metadata to record alongside.
- Harness recipe — vitest + env-var overrides to exercise sequential
  fallback vs worker-pool paths.
- Latest-measurement table with placeholder rows for the three paths
  (sequential, workers+concurrency, workers single-threaded) and an
  explicit "Status: scaffold — fill in before merging" callout. The
  U6 test's observed ~6 s wall-clock is captured as a smoke-baseline.
- Operator-tuning quick reference cross-linked to the README env-var
  section (U11) so the doc is actionable without re-reading the PR.
- "What this benchmark does NOT measure" section explicitly scoping
  the artifact's limits (synthetic ≠ real-repo, throughput-only ≠
  resilience-tested, Phase 3 IPC repack row reserved for U16-U17).

Mitigates the doc-review SG5 "static doc drift" concern via:
  1. Explicit "regenerate this file before merging" callout at the top.
  2. Self-contained methodology so anyone can re-run the numbers.
  3. Cross-links to the U6 integration test that already bounds the
     wall-clock as part of the CI suite — so "is it still completing?"
     is regression-tested even if the numbers in this doc drift.

The standalone harness script (`bench/scripts/parse-throughput.ts`)
remains a stretch goal per the original plan. The U6 vitest with
verbose ingestion logs covers the primary observability gap until
the standalone harness lands.

* perf(parse-impl): free deferred-extraction arrays after consumption (U15 lightweight M1)

PR #1693 review M1 noted that the deferred-extraction accumulator
arrays (`deferredWorkerImports`, `deferredWorkerCalls`,
`deferredWorkerHeritage`, `deferredConstructorBindings`,
`deferredAssignments`) were retained until function return, making
peak accumulator memory O(repo) instead of O(in-flight stage).

This commit implements the LIGHTWEIGHT version: free each array
immediately after its last consumer drains/reads it, dropping peak
accumulator memory progressively through the deferred-extraction
stages. The structural per-chunk streaming variant (the original
U15 framing) is deliberately deferred — the doc-review's adversarial
reviewer (A4) flagged it as defending unmeasured memory pressure,
and the simpler array-clearing captures the bulk of the benefit
without committing to a scheduling-strategy decision (microtask vs
parallel extractor task vs worker-side) that profile data should
inform.

Clears added:

  1. After `processImportsFromExtracted` (the sole consumer of
     `deferredWorkerImports`): clear the imports array before
     the heavier heritage/calls stages run.
  2. After `buildHeritageMap` (the LAST consumer of the raw
     `deferredWorkerHeritage` records — processCallsFromExtracted
     reads from the derived `fullWorkerHeritageMap` instead):
     clear the heritage array before the call-resolution stage.
  3. After `processAssignmentsFromExtracted` (the joint last
     consumer with processCallsFromExtracted for the calls/
     bindings/assignments triple): clear all three before
     downstream graph-build / scope-resolution uses its own
     working memory.

Arrays returned in the function result object (allFetchCalls,
allExtractedRoutes, allDecoratorRoutes, allToolDefs, allORMQueries,
allParsedFiles) intentionally stay live — downstream consumers
need them.

Graph-output equivalence is preserved (U7 multi-chunk equivalence
test passes — the clears happen AFTER each array's last consumer
has copied data into the graph or derived structures).

* feat(workers): introduce protocol.ts wire-format module (U16, IPC scaffold)

Defines the binary frame for worker-thread IPC as an isolated, fully-tested
module. Production wiring is deferred to U17 — shipping the wire-format
contract first de-risks the migration by establishing a single source of
truth for the byte layout. Resolves the scaffold half of PR #1693 review
R12.

Wire layout (per message, single buffer):

  +---------+-----------+---------------------+
  | tag     | length    | payload bytes …     |
  | 1 byte  | 4 bytes   |                     |
  +---------+-----------+---------------------+

  tag    : MessageTag enum value (0x01 DispatchJob ... 0x08 Ready)
  length : little-endian uint32 byte count for the payload region
  payload: UTF-8 JSON-encoded value, possibly "null"

Why JSON for the body (rather than per-shape binary encoders): the
doc-review adversarial reviewer (A2) flagged that a true per-shape
binary encoder for the result message — which carries nested
heterogeneous extracted-call / import / heritage / route arrays —
would be 500-1500 LOC and a substantial maintenance burden. The
honest perf win the IPC repack targets is moving file CONTENTS via
ArrayBuffer transferList (zero-copy ownership transfer for the
largest single piece of state in any message). That win is captured
by U17 layering transferList over the bulk file-content payload while
keeping this module's framing for the surrounding metadata. If U18
benchmark data shows the JSON body is itself a bottleneck after U17
lands, a follow-up unit can swap to per-shape binary encoding behind
the same encodeMessage / decodeMessage surface without changing the
frame.

API:
  - MessageTag (const object): stable byte tags 0x01..0x08
  - PROTOCOL_HEADER_BYTES = 5
  - ProtocolDecodeError extends Error: distinct class so U17's
    pool-side handler can route protocol violations through the
    existing messageerror recovery layer (U3 H1) distinctly from
    other failure classes
  - encodeMessage(tag, payload): Buffer
  - decodeMessage(buf): { tag, payload }
  - Uses Buffer#subarray instead of the deprecated Buffer#slice

Tests (18, all exact-equality per DoD §2.7):
  - byte layout (tag at offset 0, length LE uint32 at offset 1)
  - empty/null payload encodes to 5-byte header + 4-byte "null" body
  - round-trip for every MessageTag with representative payloads
  - non-ASCII path string (UTF-8 byte-length boundary)
  - 9 MB payload (well past the existing 8 MB sub-batch budget)
  - decode errors surface as ProtocolDecodeError, not generic Error:
      * buffer < header size
      * tag outside valid range
      * declared length exceeds buffer
      * payload bytes are not valid JSON
  - error class name is preserved through prototype chain so callers
    can `err instanceof ProtocolDecodeError` reliably

* refactor(workers): extract quarantine into its own module (U13 partial)

Honest partial U13: extract the quarantine resilience layer (Layer 3
of the 5-layer model) into a dedicated module with a small explicit
interface. The full 5-module split that the original plan named was
flagged by doc-review A10 as abstraction-without-multi-consumer-demand
("Each has exactly one consumer: worker-pool.ts. None of these layers
is imported elsewhere in the codebase pre-extraction, and the plan
doesn't identify any future consumer.") This commit ships the smallest
self-contained layer as a named module to validate the factory +
interface pattern with minimal risk. The remaining four layers
(respawn-budget, cumulative-timeout, circuit-breaker, slot-attribution)
stay inline until a real second consumer emerges (e.g., a non-parse
worker pool that reuses the same resilience layers).

Module shape (`workers/quarantine.ts`, ~30 LOC):

  interface Quarantine {
    add(path: string): void;
    has(path: string): boolean;
    snapshot(): string[];   // defensive copy
    readonly size: number;  // getter, reflects state at access time
  }
  function createQuarantine(): Quarantine

Replaces in `worker-pool.ts`:
  - `const quarantined: Set<string> = new Set()` -> `createQuarantine()`
  - `quarantined.has(p)`            -> `quarantine.has(p)` (2 sites)
  - `quarantined.add(p)`            -> `quarantine.add(p)` (2 sites)
  - `quarantined.size`              -> `quarantine.size` (2 sites)
  - `Array.from(quarantined)`       -> `quarantine.snapshot()` (6 sites)

Public worker-pool.ts API is unchanged — `getQuarantinedPaths()` still
returns the same defensive `string[]` copy. The behavioral contract is
preserved: paths are quarantined as opaque strings (the U9 / M5
non-normalization contract still holds — see the new dedicated test).

Tests:
  - 8 isolated unit tests for the quarantine module — pins the
    interface contract (empty start, add/has/size, dedup on repeated
    add, no separator normalization, snapshot defensive copy + freshness,
    size-getter live behavior).
  - All 86 existing worker-pool tests pass unchanged — they exercise
    the quarantine through the pool and act as the regression net for
    behavior preservation.

Why not the full 5-module extraction in this commit: doc-review A10's
concern is real — a single-consumer abstraction adds module-boundary
overhead (5 sets of imports, 5 dedicated test files, 5 interfaces to
keep in sync with worker-pool) without any structural benefit until a
second consumer materializes. Extracting one validates the pattern;
the remaining four can be moved on demand.

* feat(workers): wire protocol.ts encoded IPC into parse-worker + pool (U17)

Production worker IPC now uses the U16 binary wire format (1-byte tag +
4-byte LE length + UTF-8 JSON body) end-to-end. The pool encodes every
outgoing `sub-batch` / `flush` dispatch via `encodeMessage`; the worker
decodes incoming frames via `decodeMessage` and encodes its `ready`,
`starting-file`, `progress`, `sub-batch-done`, `result`, `warning`, and
`error` outputs the same way.

The load-bearing correctness fix is making `decodeMessage` accept
`Uint8Array` rather than only `Buffer`: Node's `worker_threads`
`postMessage` structured-clones the payload, which strips the `Buffer`
prototype on the receive side. A frame sent as `Buffer` arrives as a
plain `Uint8Array`, and `Buffer.isBuffer(raw)` returns false — so the
first attempt at U17 (gating decode on `Buffer.isBuffer`) silently
treated every incoming frame as POJO and the worker never responded.
The fix adopts the underlying memory zero-copy via
`Buffer.from(view.buffer, view.byteOffset, view.byteLength)` and uses
`raw instanceof Uint8Array` at every call site (parse-worker decode,
pool dispatch handler, pool ready-handshake handler, FakeWorker test
mocks, and the integration-test worker preamble).

The pool stays tolerant of POJO incoming so unit-test FakeWorkers
don't need rewriting — only the new outgoing encoded dispatches require
the test scaffolding to decode on receive, which the test FakeWorkers
and the integration test's inline `parentPort.on` wrapper now do.

The slot-drop integration test was rewritten from a shared-counter-file
race (which pre-U17 timing happened to land on the assertion-friendly
counter==2 endpoint, but post-U17 protocol decoding latency shifted to
counter==1 and produced 3 quarantines instead of 2) to a deterministic
path-based crash trigger: slot 0 crashes on a.ts, respawns, crashes on
the requeued b.ts, slot is dropped after budget exhausted; slot 1
handles [c.ts, d.ts] normally. Outcome no longer depends on inter-worker
file-write ordering.

Protocol coverage adds two regression tests pinning the Uint8Array
decode path: structured-clone-stripped frames decode identically to
their Buffer originals, and Uint8Array views with non-zero byteOffset
into a wider ArrayBuffer also decode correctly (catches `Buffer.from(uint8)`
copying semantics if a future refactor loses the zero-copy adoption).

All 94 worker-pool tests (9 files, unit + integration) pass; the full
unit suite (6128 tests across 268 files) passes unchanged.

* perf(workers): zero-copy file content transfer via transferList (U19)

Pool dispatch now hoists `{path, content: string}[]` file contents OUT
of the U17 JSON envelope into separately-allocated `Uint8Array`s whose
ArrayBuffers are passed to `worker.postMessage`'s `transferList` for
zero-copy ownership transfer. The envelope itself carries only
lightweight metadata (`{path, byteLength}` per file) and is structure-
cloned the same as before.

What this saves vs U17 baseline:

- **JSON.stringify of file contents on main thread** drops to zero —
  the envelope is now O(paths + sizes), not O(total bytes). For a 200-
  file sub-batch of 10 KB TS files, that's ~2 MB of escape processing
  per dispatch that disappears. JSON.stringify's per-character branch
  on quotes/backslashes/control chars is roughly 2x slower than
  UTF-8 transcode in TextEncoder, so the replacement is a CPU win
  even though it adds a single TextEncoder.encode per file.
- **Structured-clone memcpy of file contents** drops to zero — the
  contents' backing ArrayBuffers are ownership-transferred, not copied
  into the worker's heap. The envelope's struct-clone cost is now
  proportional to metadata size only.
- **JSON.parse on worker thread** likewise no longer scales with
  content size. Worker decodes each `Uint8Array` to string via
  `TextDecoder` lazily at the parse boundary — runs on the worker
  thread, parallel with continued main-thread work, vs U17's
  sequential JSON.parse blocking the worker before processBatch can
  start.

Pipelining: TextEncoder.encode (main) and TextDecoder.decode (worker)
can both run while the OTHER side is doing useful work. Under U17,
struct-clone was a synchronous main-thread blocker.

The ArrayBuffer ownership contract is load-bearing:

- File-content `Uint8Array`s are allocated via `TextEncoder.encode`,
  NOT `Buffer.from(str, 'utf8')`. TextEncoder produces a dedicated
  ArrayBuffer per call; `Buffer.from(str)` carves from Node's shared
  `Buffer.poolSize` slab for small strings, so transferring one
  pool-backed Buffer's ArrayBuffer would detach every other Buffer
  that shares that slab — silent data corruption.
- The envelope itself is NOT transferred. It MAY be pool-backed by
  `encodeMessage`, and at ~30-80 bytes/file the struct-clone cost is
  negligible. Not transferring avoids the same detach-collateral risk
  the contents path is careful to dodge.

Detection is strict: every input element must have both `path: string`
and `content: string`. A single non-conforming element disqualifies
the whole batch from the transfer path and falls back to the legacy
single-Uint8Array `encodeMessage` envelope. Safer than partial
transfer (which would split a sub-batch into mixed-shape messages
the worker can't reassemble).

`parse-worker.ts` `decodeIncomingMessage` recognizes the hybrid
`{envelope, contents}` shape, decodes the envelope, zips metadata
positionally with the contents array, decodes UTF-8 → string per file,
and hands the reassembled `ParseWorkerInput[]` to the existing
`processBatch`. Identical downstream behavior to U17 — the IPC
optimization is invisible above this line.

Test scaffolding (3 FakeWorkers + 1 integration-test preamble) gain a
`decodeDispatchedMessage` helper that tolerates BOTH shapes (legacy
single-frame Uint8Array AND the new hybrid envelope+contents) so the
in-process unit mocks keep their existing action-scripting API and the
9 ad-hoc integration test workers keep their `msg.type === 'sub-batch'`
handlers unchanged.

`buildDispatchMessage` is now exported from worker-pool.ts so its
contract can be tested in isolation. A new
`test/unit/worker-pool-transferlist.test.ts` pins:
  - hybrid shape produced for parse-worker inputs
  - transferList carries one ArrayBuffer per file in input order
  - envelope decodes to metadata only (no `content` field)
  - content bytes round-trip byte-for-byte through UTF-8 (ASCII,
    multi-byte, surrogate-pair emoji)
  - each content's ArrayBuffer is independently allocated (no pool
    sharing) — the load-bearing transfer-safety invariant
  - non-parse shapes, empty arrays, and mixed-conformance arrays all
    fall back to the legacy single-frame path

All 271 test files (6166 unit + integration tests) pass.

* fix(workers,tests,docs): apply ce-code-review findings (16 items)

Walks the full set of findings from a multi-agent code review (11
reviewers, 1 maintainability dispatch lost to tool-permission denial)
of the PR #1693 branch. All 16 actionable findings — 4 P1, 4 P2,
8 P3 — applied in a single pass against a consistent tree. Tests
pass (269/269 unit files, 29/29 integration).

P1 — bounds-only / disguised-bounds assertions across 4 test files
(per user-memory DoD §2.7):
  - worker-pool.test.ts: 5 sites — `nodes.length > 0` dropped (redundant
    after `.toContain('validateInput')`); `files.length >= 4` pinned to
    `.toBe(7)` (mini-repo/src has exactly 7 .ts files); `results.length
    > 0` pinned to `.toHaveLength(1)` (default sub-batch absorbs all 7);
    `result.fileCount >= 0` pinned to `.toBe(1)` (empty file is still
    "processed"); `warnRecords.length > 0` replaced with content-
    predicate `/respawn|dropping|replacement|did not report ready/`
    (catches silenced warnings); `fallbackExcludePaths.length > 0`
    pinned to exact `['one.ts', 'two.ts']` (deterministic given the
    single-slot pool + 2 items + per-item starting-file).
  - parse-impl-fallback.test.ts: 3 sites — `astCacheClearCalls >= 1`
    pinned to exact 4 (per-chunk × 2 + finally × 2); the two error-path
    delta checks pinned to exact +2 and +3 (verified empirically).
  - parse-impl-progress-monotonic.test.ts: `percents.length > 0` →
    `.not.toEqual([])`; per-element `Math.max(prev, cur)` tautology
    replaced with direct `if (cur < prev) throw`; final-percent
    `Math.min(last, 95)` tautology pinned to exact `.toBe(70)` (3-file
    skipWorkers fixture's deferred band lands at the band start).
  - parse-impl-large-fixture.test.ts: `Math.min(elapsedMs, BUDGET)`
    tautology removed; Promise.race rejection is the load-bearing
    wall-clock check.

P1 — terminate() lacks `.catch` mask:
  - worker-pool.ts terminate() now matches the `.catch(() => undefined)`
    pattern used at every other internal terminate site. Prevents a
    hung/OOM worker's terminate rejection from masking the original
    pipeline error when called from parse-impl.ts's finally block, and
    guarantees `workers.length = 0` / `activeSlots.clear()` always run.

P1 — hybrid envelope length-mismatch + null-payload silent data loss:
  - parse-worker.ts decodeIncomingMessage: explicit non-null-and-typed
    check before `.type` access (decodeMessage permits null payloads
    per encodeMessage contract); explicit length-equality assertion
    between `decoded.files` and `contents` before zipping. Without
    these, `TextDecoder.decode(undefined)` silently returns "" and
    produces empty-content graph nodes — a contract violation that
    used to be undetectable. Both throws route through the outer
    try/catch → worker `error` reply → pool's recoverAndResume.

P1 — unsafe casts at the IPC boundary:
  - buildDispatchMessage now uses a properly-typed `isParseWorkerItemArray`
    type guard. The narrowed branch accesses `item.path` and
    `item.content` as statically-typed strings — a future rename of
    `ParseWorkerInput.content` would fail to compile inside the branch
    instead of silently mismatching at runtime. The remaining
    decodeMessage payload casts are bounded by the F3/F6 runtime
    guards.

P2 — idle-timeout retry bypasses circuit breaker:
  - worker-pool.ts timeout-retry IIFE now increments
    `consecutiveFailuresPerSlot[workerIndex]` alongside `respawnCount`.
    A slot that consistently times out (vs crashes) now trips the
    per-slot breaker, instead of consuming its full respawn budget
    over potentially tens of minutes without the breaker firing.

P2 — null/non-object worker message crashes pool handler:
  - Dispatch handler in worker-pool.ts now guards `null /
    non-object / no string type discriminant` before `msg.type` access
    and routes through recoverAndResume on violation. Previously a
    legitimate `null` payload would throw TypeError out of the
    EventEmitter listener → uncaughtException on main, crashing the
    analyze.

P2 — workerPoolSize === 0 creates unusable pool:
  - parse-impl.ts now treats `workerPoolSize === 0` as `skipWorkers`
    at the gate. Matches the PipelineOptions docstring contract ("0
    disables the pool entirely — equivalent to skipWorkers"); avoids
    constructing a pool that rejects every dispatch and logs
    "Worker pool parsing stopped" per chunk.

P2 — encodeMessage 2-buffer allocation per frame:
  - protocol.ts encodeMessage coalesced to a single
    `Buffer.allocUnsafe + writeUInt8 + writeUInt32LE + buf.write
    (string, offset, 'utf8')`. Drops the intermediate
    `Buffer.from(JSON.stringify(...), 'utf8')` allocation + memcpy.
    Length pre-check via `Buffer.byteLength(string, 'utf8')` surfaces
    the uint32 cap before any allocation.

P3 — slotGenerations made optional on WorkerPoolStats so external
  implementations of getStats() that predate U12 don't compile-break;
  in-repo callers already use optional chaining.

P3 — buildDispatchMessage marked `@internal` so it isn't surfaced as
  public API by typedoc / api-extractor (it's a test-only export).

P3 — verboseThroughputLog hoisted above the chunk loop (env vars can't
  change mid-run; one O(env-read) per analyze, not per chunk).

P3 — corrected the messageerror routing comment in worker-pool.ts
  dispatch handler. `ProtocolDecodeError` is caught by the surrounding
  try/catch — distinct from `messageerror`, which fires for V8
  structured-clone failures before the message body would reach the
  handler.

P3 — initial pool spawn now uses a `Promise.allSettled` ready-handshake
  gate symmetric with `replaceWorker`. Dispatch awaits this gate before
  selecting slots, so an init-crashing initial worker is dropped from
  `activeSlots` and a downstream OOM/missing-native-binding failure
  surfaces in seconds (bounded by WORKER_READY_TIMEOUT_MS) rather than
  waiting for the first idle timeout (30s default).

P3 — `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT`,
  `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS`,
  `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD` added to:
    - CLI `--help` text in src/cli/index.ts
    - Root README env-var table
    - gitnexus/README troubleshooting section (new "Worker pool
      resilience tuning" subsection)

P3 — CLI `catch (e: any)` / `catch (err: any)` in analyze.ts replaced
  with `catch (err: unknown)` + narrowed access; matches modern TS
  best practice and the codebase pattern at other catch sites.

P3 — `WorkerPoolStats.terminated: boolean` field added (optional, for
  backward compatibility). `terminate()` sets it true; `getStats()`
  surfaces it. Distinguishes graceful shutdown from a circuit-breaker
  trip in observability surfaces.

Coverage / advisory items not addressed in this commit (kept in the
report only):
  - maintainability reviewer failed (Read/Bash denied) — god-module
    audit on worker-pool.ts (~1400 LOC) carried as residual risk
  - quarantine case-sensitivity contract unpinned (adversarial #8)
  - WORKER_READY_TIMEOUT_MS env-configurability (adversarial #2)
  - chunk-byte-budget × parseChunkConcurrency memory multiplier doc
    (adversarial #5)
  - MCP discoverability gaps for env vars / verbose (agent-native W1/W2)
  - bench/parse-throughput.md scaffold-with-TBD-rows (PS RR-003)

* fix(parsing): sequential gap-fill for worker-quarantined chunk files (U20.U1)

When the worker pool's Layer 3 quarantine filters one or more files
out of a chunk's dispatch, the worker results returned to
processParsing are silently narrower than the input chunk. Without
this reparse, the graph for this run would be missing every quarantined
file's symbols/imports/calls/heritage with no failure signal.

After the existing per-chunk quarantine log emits in
processParsing's worker-path try-block, run processParsingSequential
on JUST the quarantined-in-chunk files. The sequential path writes
directly to the graph, so symbols for those files land alongside
worker output for the surviving files.

Mirrors the WorkerPoolDispatchError catch-block's processParsingSequential
call shape — same signature, same args, same scopeTreeCache wiring.
Emits a structured warn naming `reparsedPaths` so operators can
observe the sequential fall-through.

This fixes the in-run side of the corruption Codex's adversarial
review of PR #1693 flagged. The cross-run side (chunk-cache
poisoning) is closed by U20.U2 in a follow-up commit.

References plan: docs/plans/2026-05-20-002-fix-chunk-cache-corruption-on-worker-quarantine-plan.md

* fix(parse-impl): suppress chunk-cache write when any chunk file was quarantined (U20.U2)

The chunk hash at parse-impl.ts:424-428 is computed from every file
in the chunk. The worker pool's Layer 3 quarantine
(worker-pool.ts createQuarantine) filters quarantined files out of
dispatch, so `rawResults` reflects only the surviving files. Before
this commit, the write at line 500-507 stored that partial result
under the full-coverage chunk hash — and on the next analyze with
unchanged content, the cache HIT branch (line 439-464) silently
replayed the incomplete result. Symbols from the quarantined file
were missing from the graph for as long as the cache survived.

Codex's adversarial review of PR #1693 flagged this as a silent-
corruption class because there's no failure signal: no warn log
during the replay, no graph-equivalence check, no exit code change.
The corruption only surfaces if an operator notices a missing symbol
in `gitnexus_query` output.

Guard the write with `chunkFiles.some(f => quarantineSet.has(f.path))`.
When any chunk file is in the worker pool's cumulative quarantine
snapshot, skip the `parseCache.entries.set` call. Emits a verbose-
only info log so operators investigating "why aren't my chunks
caching" have a diagnostic trail.

Skipping the write means the next analyze gets a cache miss for this
chunk and re-dispatches it. Quarantine is session-scoped (a fresh
createWorkerPool starts with an empty quarantine), so the new pool
gives the quarantined file another chance. If quarantine fires again,
U20.U1's sequential gap-fill still produces a complete graph for that
run; the cache stays empty for the chunk until a fully-clean
dispatch lands.

The cache-hit replay branch at parse-impl.ts:439-464 is unchanged.
Its contract strengthens: "cache entries are complete" becomes true
post-fix, but the replay code doesn't need to know that.

Closes the cross-run side of the Codex finding. U20.U3 adds the
regression test.

References plan: docs/plans/2026-05-20-002-fix-chunk-cache-corruption-on-worker-quarantine-plan.md

* test(parse-impl): integration regression for quarantine + chunk-cache (U20.U3)

Pins the U20 fix end-to-end via REAL `worker_threads` + `createWorkerPool`.
Mirrors the writeReadyWorker pattern from `test/integration/worker-pool.test.ts`
— inline READY_PREAMBLE + custom test worker script that:

  1. Decodes the U17/U19 IPC protocol (Buffer frame OR hybrid envelope/
     contents shape) the same way the production parse-worker does.
  2. Emits a `{type:'ready'}` handshake so the pool's
     `waitForWorkerReady` resolves promptly.
  3. On a sub-batch containing `poison.ts`, emits starting-file +
     `process.exit(134)`. The pool attributes the death to `poison.ts`
     via the in-flight signal and adds it to the session-scoped
     quarantine.
  4. On a sub-batch without poison, synthesizes a minimal valid
     `ParseWorkerResult` with one `Function` node per file (no
     tree-sitter dep in the test worker — the synthesized nodes give
     `mergeChunkResults` deterministic content for the graph).

Assertions exercise both fix layers:

  - U1 (sequential gap-fill in processParsing): the graph contains a
    `Function` node named `poison` AFTER the run. The custom worker
    never emits anything for `poison.ts`, so the only path for that
    symbol to reach the graph is `processParsing`'s sequential
    reparse of the quarantined-in-chunk file using the real
    tree-sitter parser against the actual source.
  - U2 (cache-write suppression in runChunkedParseAndResolve):
    `parseCache.entries` does NOT contain the chunk hash after the
    run; `parseCache.usedKeys` DOES contain it (chunk processed,
    cache write specifically skipped).
  - Cross-run: a second pass over the same fixture with the same
    parseCache and a fresh worker pool re-dispatches the chunk
    (cache empty), the worker crashes again, sequential gap-fill
    runs again, and the cache stays empty. Pins the round-trip
    contract.

Adds `workerUrlForTest?: URL` to PipelineOptions — same `@internal`
test-only injection precedent as `workerThresholdsForTest` (already
in PipelineOptions for thresholds). When set, parse-impl uses the
provided URL instead of the src/ → dist/ resolution dance. Production
call sites never set this field; the only consumer today is this
integration test.

Why integration over unit:
  - The fix lives at the boundary between parsing-processor.ts and
    parse-impl.ts under a real WorkerPool. Unit-mocking the
    worker-pool module bypasses the structured-clone boundary, the
    dispatch lifecycle, and the actual quarantine flow — it verifies
    the test setup rather than the contract. The real worker thread
    executing through the U17/U19 IPC protocol IS the load-bearing
    surface.
  - User-explicit preference (saved as
    feedback_integration_over_vimock.md memory). For worker-pool /
    parse-impl / IPC-touching code: write integration tests under
    test/integration/ using writeReadyWorker patterns; avoid
    vi.mock on worker-pool.js.

Test wall-clock: under 2s; both `it` blocks together complete in
~1.8s under the existing CI conditions.

References plan: docs/plans/2026-05-20-002-fix-chunk-cache-corruption-on-worker-quarantine-plan.md

* refactor(parsing): remove sequential-parser fallback (U20 design pivot)

The worker pool's resilience layers — respawn budget, circuit breaker,
quarantine, slot-attribution, cumulative timeout — are now the SOLE
contract for handling worker failures. Two sequential-reparse paths
are removed from processParsing:

1. **U20.U1 sequential gap-fill for quarantined chunk files** (just
   added in commit 7dd489e9, now reverted). The pre-emptive rescue
   would re-run processParsingSequential on the file that ALREADY
   killed a worker — which for the most common quarantine cause
   (tree-sitter native SIGSEGV on a pathological file) re-triggers
   the same native crash on the main thread, killing the entire
   analyze. The "rescue" turned silent missing-symbols into a louder
   analyze-wide crash. Drop the rescue; accept the per-run gap.

2. **Pre-existing WorkerPoolDispatchError catch-block sequential
   fallback** (in production since PR #1693's resilience layer
   landed). Same risk class — when the pool exhausts its respawn
   budget / trips the circuit breaker, the failing files are
   precisely the ones likely to crash a sequential parser too. The
   "graceful degradation" hid pool failures behind degraded-but-
   completing analyze runs, making operational issues harder to
   surface and diagnose. Drop the catch-block; WorkerPoolDispatchError
   propagates to the analyze entry point where the user sees a clear
   hard signal.

What stays:
- The `skipWorkers: true` / small-repo path that uses
  `processParsingSequential` as the EXPLICIT primary path (not a
  fallback). Caller-driven opt-out and tiny-repo perf optimization
  are different intents.
- U2's chunk-cache write suppression in parse-impl.ts (commit
  7c9c9556). When quarantine fires, the chunk stays uncached so the
  next analyze with a fresh pool retries the file cleanly. That's
  the cross-run correctness Codex's adversarial review actually
  asked for.
- The per-chunk quarantine warn log (parsing-processor.ts) — operators
  see which files were skipped, both immediately and across runs.

What changed:
- `processParsing` worker-path try-block: unwrapped. The
  `processParsingWithWorkers` call is now direct (no try/catch
  wrapping); errors propagate to the chunk-loop caller.
- `parsing-worker-fallback.test.ts` rewritten: the previous 5 tests
  asserted graceful sequential-fallback behavior. Replaced with 3
  tests pinning the new contract — raw Error propagates, WorkerPool-
  DispatchError propagates with fallbackExcludePaths intact, normal
  quarantine signal does NOT throw and surfaces via progress detail.
- `parse-impl-quarantine-cache-skip.test.ts` (U20 integration test)
  updated: poison.ts is NOT in the post-run graph; surviving files
  are; chunk-cache stays empty; second pass re-dispatches and leaves
  cache empty.
- Plan doc updated to mark R1 as dropped and explain the U20 pivot
  in the Summary.

User decision: explicit directive ("let's remove the sequential
fallback entirely we must rely on entirely that the parallel process
is resilient enough to work itself through the code base"). The pool's
resilience layers are designed for this — respawn budget, circuit
breaker, quarantine, slot-generation, cumulative-timeout cap — and
adding a layer below them was redundant insurance with real downside.

Tests: 269/269 unit files (6135 tests) green. 31/31 worker-pool +
parse-impl integration tests green. The 2 reported "errors" in the
integration run are the pre-existing intentional-process.exit unhandled-
exception leaks from test workers — unchanged by U20.

References plan: docs/plans/2026-05-20-002-fix-chunk-cache-corruption-on-worker-quarantine-plan.md

* fix(workers,tests,docs): address ce-ultrareview findings F1/F2/F3/F4

Multi-lane review run on the PR #1693 branch surfaced four addressable
items beyond the blocking three.

F1 (minor, CodeQL): unused `findMatch` helper in
test/unit/scope-resolution/typescript/typescript-captures-anchor.test.ts:28
removed. `countMatchesTsx` flagged by the same CodeQL pass is a false
positive — it's called at line 88 by the JSX-anchor regression tests
so the rewrite case actually fires under TSX, not just TS.

F2 (medium, docs): bench/parse-throughput.md retitled as
"(scaffold)" with an explicit "no measurement data has been collected
yet" note above the table. The self-contradictory "Regenerate this
file before merging any PR that touches the ingestion pipeline"
instruction is dropped — the file ships intentionally without
numbers; the load-bearing perf-regression protection lives in
test/integration/parse-impl-large-fixture.test.ts (U6, 30s
Promise.race wall-clock budget). The Latest measurement section now
preserves the ~6s sequential observation as a smoke reference, not as
a regression target.

F3 (low, API hygiene): `WorkerPoolDispatchError.fallbackExcludePaths`
renamed to `quarantinedPaths`. The "fallback" terminology was
load-bearing under the pre-U20 design when `processParsing`'s
sequential-fallback catch-block consumed it to filter the fallback
file list. After commit be1f65c removed that catch-block, no
production code reads the field — but it stays populated by the pool
because the snapshot is genuinely useful operator diagnostics when
the breaker trips. The rename clarifies the field's actual semantics
(here are the files the pool quarantined before it tripped) without
changing wire behavior. Definition + the lone surviving in-pool
comment reference + both test assertions updated.

F4 (low → real fix, reliability): timeout-retry IIFE in
worker-pool.ts now consults `consecutiveFailureThreshold` and trips
the circuit breaker when the per-slot consecutive-failure count
crosses it. Closes a gap left by ce-code-review's REL-02 patch — that
fix added the `consecutiveFailuresPerSlot[workerIndex]++` increment
in the timeout-retry path but did NOT add the corresponding
threshold-check + tripBreaker call. Result: chronic pure-timeout
deaths accumulated counts that never tripped the breaker until the
slot also hit `respawnCount > maxRespawnsPerSlot`. Now timeouts and
crashes are structurally treated the same way by the breaker, which
is what the REL-02 increment was meant to enable. Test coverage:
worker-pool-resilience.test.ts already exercises the breaker via the
shared handleWorkerDeath path; this new branch traces the same
trip semantics with a different entry point, so the breaker-tripped
state is observable via the same `getStats().poolBroken` and
`WorkerPoolDispatchError.quarantinedPaths` surface.

Out of scope here (caller actions or future PRs):
  - F5 (info): cumulative-quarantine cache check is safe in practice
    because chunks are alphabetically deterministic; no action.
  - F6 (low): exit-code-0 quarantine exemption — pre-existing P2
    residual, bounded by quarantine + respawn budget; deferred.
  - F7 (info): dispatch non-reentrancy contract documented but not
    enforced; no production caller violates it; deferred.
  - PR title `[WIP]` removal — happens on GitHub side.

Tests: 274/274 test files (6185 passing, 30 skipped). The single
"error" in the integration runner is the pre-existing intentional-
process.exit unhandled-exception leak from the deliberate startup-
crash test worker, unchanged by these fixes.

* fix(workers): swap protocol body from JSON to V8 serialize/deserialize

CI scope-parity tests on Ubuntu surfaced silent data loss in the
worker IPC: `Phase 'scopeResolution' failed: scope.typeBindings is not
iterable` (Python, Go) and `importerModule.typeBindings.has is not a
function` (Python). Plus three #1066 large-file regression tests
(Python / C# / TypeScript) failed because call relationships weren't
resolving from the worker output.

**Root cause:** U17 introduced `JSON.stringify`/`JSON.parse` as the
protocol body codec. JSON has no representation for `Map`, `Set`,
`Date`, `RegExp`, `BigInt`, `TypedArray`, `undefined` values, or
circular refs — `JSON.stringify(someMap)` returns `"{}"`. Production
scope-resolution code keys data structures on Maps throughout
(`ParsedFile.scopes[*].typeBindings: ReadonlyMap<string, TypeRef>`,
plus `bindings`, `bySourceScope`, `byTargetDef`, the finalize-algorithm
edge indexes, etc.). The JSON round-trip silently turned every Map
into an empty object, manifesting downstream as iteration / `.has`
calls failing on the decoded payload.

**Fix:** replace the JSON body with `node:v8`'s `serialize` /
`deserialize`. That's the same structured-clone algorithm Node's
`worker.postMessage` uses natively — bit-for-bit compatible with the
pre-U17 implicit-clone path. Full type fidelity for Map, Set, Date,
RegExp, BigInt, TypedArray, undefined values, and circular refs. No
external dependency.

A previous iteration of this fix attempted to bolt a Map/Set
replacer+reviver onto the JSON path. Rejected in favor of V8
serialization because:
  - the JSON tag-marker approach requires per-type registration
    (Map, Set; then Date, RegExp, BigInt would each need their own
    sentinels); V8 handles them all uniformly
  - keys to JSON-encode would still need handling for nested types
    (and the marker approach doesn't survive nested Maps-in-Maps
    cleanly without recursive replacer logic)
  - V8 is faster than JSON for object-heavy payloads anyway (binary
    format, no string escaping pass)
  - the user-explicit ask was "a much more generic solution that will
    work for everything" — V8 serialization IS the generic solution

Trade-offs documented in the module header:
  - body bytes are opaque (binary, not human-readable) — debugging
    requires `v8.deserialize` ad-hoc; protocol.test.ts exercises every
    supported MessageTag including the new type-fidelity cases as a
    regression net.
  - format is tied to the running Node major. Pool always spawns
    workers on the same Node instance the main thread runs, so this is
    moot in production. Would matter if frames ever persisted to disk
    (nothing does today).

Protocol test file rewritten:
  - drops the JSON-specific byte-layout assertions (e.g. `body must
    equal "null" string`) — replaced with V8-derived expected lengths
  - adds a "structured-clone type fidelity" describe block that pins
    Map, nested Map, Set, Date, RegExp, BigInt, TypedArray, undefined
    values, and circular-ref round-trips. These are the load-bearing
    regression tests preventing a future "optimize" PR from quietly
    swapping V8 back to JSON.
  - the bad-body decode-error test now uses arbitrary non-V8 bytes
    instead of `{not-json}` — same intent.

Integration test READY_PREAMBLEs (worker-pool.test.ts and
parse-impl-quarantine-cache-skip.test.ts) update their inline
decoders to use `v8.deserialize` matching the production codec.
Both files have a standalone CJS worker preamble that can't import
dist/protocol.js by relative path, so the V8 dependency is required
via `node:v8` directly.

Tests: 271/271 unit files (6163 tests + 30 skipped). 28/28
worker-pool integration. 3/3 parse-impl integration. 791/791
scope-parity tests (the four CI-failing files: python.test.ts,
go.test.ts, typescript.test.ts, csharp.test.ts) all green again.

References plan: docs/plans/2026-05-20-002-fix-chunk-cache-corruption-on-worker-quarantine-plan.md

* refactor(workers): drop protocol.ts; use native postMessage + transferList

The protocol.ts framing layer was redundant — Node's `worker.postMessage`
already runs V8 structured-clone internally, the same algorithm that
backed `v8.serialize`. Wrapping V8.serialize → Buffer →
postMessage(struct-clone-Buffer) was a double-walk: one full
structured-clone pass to produce the Buffer, then another pass when
postMessage cloned that Buffer across threads. This commit cuts the
wrapper layer; workers and pool exchange POJO directly via
`worker.postMessage(value, transferList)`, with file-content
`ArrayBuffer`s in `transferList` for zero-copy ownership transfer.

What changes:

- **Deleted** `src/core/ingestion/workers/protocol.ts` (~180 LOC) +
  `test/unit/workers/protocol.test.ts` (~250 LOC). The MessageTag
  enum / ProtocolDecodeError / encodeMessage / decodeMessage surface
  is gone. Tag-based routing is replaced by the `msg.type`
  discriminant that every receive site already checks. Protocol-decode
  errors map to Node's `messageerror` event (V8 deserialization
  failures during postMessage), which the pool already wires to
  `recoverAndResume`.
- **`worker-pool.ts`**: `decodeIncomingWorkerMessage` removed; handlers
  receive POJO directly. `buildDispatchMessage` now returns
  `{message: {type:'sub-batch', files: [{path, content: Uint8Array}]},
  transferList: ArrayBuffer[]}`. The Uint8Array-per-content allocation
  via `TextEncoder.encode` is preserved (it's the load-bearing
  transfer-safety contract that keeps content out of Node's shared
  `Buffer.poolSize` slab). Flush dispatch is now plain
  `worker.postMessage({type:'flush'})`.
- **`parse-worker.ts`**: `decodeIncomingMessage` removed. The message
  handler receives POJO directly; the only conversion is
  `Uint8Array → string` for sub-batch file contents at the
  `decodeSubBatchFiles` boundary, before handing to `processBatch`.
  Outgoing messages are emitted as POJO via plain
  `parentPort.postMessage({type:'starting-file', ...})` etc. The
  `sharedHybridDecoder` is now `sharedContentDecoder` (same intent,
  clearer name for the simpler shape).
- **Test scaffolding**: FakeWorkers in `worker-pool-resilience`,
  `worker-pool-windows-quarantine`, and `worker-pool-slot-generation`
  drop their `decodeMessage` import + `decodeDispatchedMessage` helper.
  The helpers stay (still convert `files[i].content` Uint8Array →
  string for test-action introspection) but no longer touch any
  protocol framing — just shape-check for sub-batch.
- **Integration READY_PREAMBLEs** (worker-pool.test.ts and
  parse-impl-quarantine-cache-skip.test.ts): drop the inline
  v8.deserialize + envelope-unzip logic; the preamble is now just
  the ready handshake + a `parentPort.on` wrapper that converts
  `files[i].content` Uint8Array → string for the ad-hoc test worker
  scripts.
- **`worker-pool-transferlist.test.ts`**: contract tests updated for
  the new buildDispatchMessage shape — no `envelope` field anymore;
  `message.files[i].content` is Uint8Array; transferList holds each
  content.buffer in input order. Pool-slab independence still pinned.

What stays the same:

- Zero-copy file-content transfer via transferList — every file's
  ArrayBuffer is ownership-transferred to the worker (no copy).
- Full structured-clone type fidelity — Map / Set / Date / RegExp /
  BigInt / TypedArray / undefined / circular refs all preserved by
  Node's native postMessage. The V8 fix from commit 06f6957e is
  inherent in this path; there's no JSON layer to lose them.
- TextEncoder-per-content allocation — keeps content buffers out of
  the shared `Buffer.poolSize` slab so transferring one cannot detach
  another.
- The pool's resilience layers (respawn, breaker, quarantine,
  starting-file attribution, cumulative timeout, ready handshake,
  slot-generation guard) — unchanged.
- U20 chunk-cache write suppression on quarantine — unchanged.

Net: ~430 LOC removed (protocol.ts + tests + inline decoders + helpers),
~120 LOC simplified in worker-pool.ts and parse-worker.ts. One less
serialization pass per message on the hot path.

Tests: 270/270 unit files (6133 + 30 skipped). 822/822 integration
tests including the four CI-failing scope-parity files (Python, Go,
TypeScript, C#) — the V8-fidelity contract holds via native
postMessage with no explicit serializer. The single "error" reported
in worker-pool.test.ts is the pre-existing intentional
process.exit unhandled-exception artifact from the deliberate
startup-crash test, unchanged by this commit.

* refactor(parse-worker): drop legacy single-message dispatch mode

The `parentPort.on('message', ...)` handler had an `Array.isArray(msg)`
branch left over from a pre-sub-batch dispatch shape — the pool used
to send the items array directly, before the worker pool added
sub-batching and the `{type:'sub-batch', files: ...}` envelope.

No production caller has dispatched that shape since the sub-batching
refactor landed; verified by grepping the repo for `postMessage([`
patterns (zero matches). The `ParseWorkerInput[]` arm in the
`WorkerIncomingMessage` discriminated union also blocked
exhaustiveness narrowing — flagged by the kieran-typescript code
review (RR-01) as "if a future unit removes the legacy array path,
this arm should be dropped." Dropping it now.

What changes:
  - Remove the `Array.isArray(msg)` branch from the message handler.
  - Drop `ParseWorkerInput[]` from the `WorkerIncomingMessage` union;
    it's now a clean `{type:'sub-batch'} | {type:'flush'}` discriminated
    union, so the dispatch switch is exhaustive over `msg.type`.

Tests: 71/71 worker-pool unit + integration tests green (resilience,
slot-generation, windows-quarantine, transferlist, parsing-worker-
fallback, worker-pool integration, parse-impl-quarantine-cache-skip).

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 20:39:35 +01:00
azizur100389andGergő Magyar aa8f4d6efe fix(group): Union HTTP graph and source contracts (#1709)
* Union HTTP graph and source contracts

* test(group): Document HTTP source union follow-ups

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 17:44:07 +01:00
Shane Thurston Wijaya 4d2ed0e525 fix(eval-server): localhost now doesn't normalize into IPv4 instead lets OS decide which to bind (#1722)
* fix(eval-server): localhost now doesn't normalize into IPv4 instead lets OS decide which to bind

* fix(eval-server): EADDRNOTAVAIL now treats as potential IPv6

* test(eval-server): new integration test for --host localhost

* docs(eval-server): updated eval/README.md based on latest update

* fix(eval-server): clarify EADDRNOTAVAIL diagnostic, guard server.address(), and soften localhost docs
2026-05-20 16:14:13 +01:00
df1882d36b fix(ingestion): surface skipped large-file paths by default (#1659) (#1661)
* fix(ingestion): surface skipped large-file paths by default (#1659)

The 512 KB skip threshold in filesystem-walker is necessary, but the
existing warning only said "Skipped N large files" with no paths unless
GITNEXUS_VERBOSE=1 was set. In a repo with one or two oversized first-
party source files (e.g. a 17K-line cron handler), every IMPORTS/CALLS
edge from that file silently disappeared and the surface looked like a
Python resolver bug. Issue #1659 was filed against the resolver for
exactly that reason, but the resolver was fine; the file was being
dropped before parse.

Changes:
  * Always print up to 5 skipped paths after the count line.
  * If more than 5 were skipped, append "...and N more" with a hint to
    set GITNEXUS_VERBOSE=1 for the full list.
  * When running at the default threshold, emit a one-line hint about
    GITNEXUS_MAX_FILE_SIZE=<KB> so operators know how to widen it.
  * Cover the new behavior with three additional tests in the existing
    filesystem-walker integration suite, plus a new describe block for
    the >5 preview-cap case.

Verified end-to-end on a 680-file Python repo that hit #1659: before
the patch, "Skipped 3 large files (>512KB, ...)" was the only signal
and impact upstream of a function called from cron.py returned 1 of 5
real callers; after the patch the cron file is listed by name with the
hint, and running with GITNEXUS_MAX_FILE_SIZE=1024 brings the missing
callers back (impactedCount 1 -> 9).

* fix(ingestion): address #1661 adversarial review follow-ups (F1/F2/F3)

Three non-blocking nits flagged by the adversarial review on #1661:

F1 (output stability) — skippedLargePaths was populated by concurrent
fs.stat callbacks in batches of 32, so push order within a batch was
completion-order rather than input-order. The default preview's "first
5" could vary across runs on the same repo. Fix: sort the array before
slicing. New test asserts the verbose output is in sorted order.

F2 (boundary coverage) — the preview-cap describe block created 8
large files, so the SKIPPED_PREVIEW_CAP = 5 comparison was never
exercised at the exact <= boundary. A future off-by-one (<= → <) would
not fail the suite. Fix: add two tests, one with exactly 5 files (all
listed, no truncation) and one with exactly 6 files (5 listed plus
"...and 1 more").

F3 (hint accuracy) — isDefault compared effective bytes, so an
operator who explicitly set GITNEXUS_MAX_FILE_SIZE=512 (the same KB as
the default) would still see the "Set GITNEXUS_MAX_FILE_SIZE=<KB>..."
hint. Fix: gate the hint on whether the env var is unset, not on the
resulting byte value. New test pins the explicit-default-value case.

All 34 filesystem-walker tests pass (was 30; +4 new). Prettier clean,
typecheck clean for the changed files.

---------

Co-authored-by: scotjelinski <58397194+scotjelinski@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 14:37:58 +01:00
f350ae278a feat: Add analyze --repair-fts, enforce FTS verification, and harden repair safeguards (#1720)
* Initial plan

* feat(analyze): add --repair-fts and verify FTS index rebuilds

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/dccb3673-af86-43aa-aede-2e1449399775

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor(fts): tighten repair/verify messaging and option naming

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/dccb3673-af86-43aa-aede-2e1449399775

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docs: highlight analyze --repair-fts vs --force in READMEs

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/61edc967-debc-419f-9f51-aebf2ef08d22

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(analyze): guard repair mode against missing graph store

* fix(cli): reject --repair-fts with --force

* test(analyze): document repair-store fixture intent

* test(analyze): tidy repair failure fixtures and constants

* test(analyze): clarify mock constants in repair tests

* test(analyze): rename simulated missing-index constant

* test(analyze): clarify mocked graph shape in full-verify test

* refactor(analyze): finalize flag validation and test clarity

* test(skip-git): avoid hard failing when FTS extension is unavailable

* test(skip-git): log visible FTS-unavailable test skips

* test(skip-git): tighten FTS-unavailable error detection

* test(skip-git): simplify FTS-unavailable message checks

* test(skip-git): avoid HOME pointing at parent repo in fixture env

* fix(analyze): address Claude follow-up findings for repair guardrails

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

* fix(repair-fts): clarify invalid graph-store preflight errors

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

* test(analyze): strengthen assertions for conflict and missing-store errors

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

* fix(repair-fts): make invalid graph-store type errors explicit

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

* fix(repair-fts): improve graph-store type diagnostics

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 13:37:04 +01:00
azizur100389andGergő Magyar dae70a26ea feat(cpp): Add pointer nullptr ellipsis conversion ranks (#1708)
* Add C++ pointer null ellipsis ranks

* test(cpp): Strengthen pointer overload assertions

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 12:06:51 +01:00
CopilotandGergő Magyar b4a2a4b91e fix(ingestion): Prioritize same-module Java type resolution for duplicate FQNs across modules (#1712)
* Initial plan

* Fix Java same-name type resolution with same-module priority

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/df0843e3-e244-4e0f-a94a-311df3899bd0

* Refine Java ambiguity fallback safety check

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/df0843e3-e244-4e0f-a94a-311df3899bd0

* Remove Java-specific fallback from shared scope walkers

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e882906a-2c96-411e-94a5-123a345421a9

* Harden Java module key and ambiguous owner fallback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e882906a-2c96-411e-94a5-123a345421a9

* Add negative assertions for duplicate-FQN module edges

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e882906a-2c96-411e-94a5-123a345421a9

* Make Java same-module ordering path-agnostic

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b317a759-f6bc-4590-bd2a-628f0ee9c477

* Refine generic Java path-affinity ordering safeguards

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b317a759-f6bc-4590-bd2a-628f0ee9c477

* Polish Java path-affinity ordering clarity

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b317a759-f6bc-4590-bd2a-628f0ee9c477

* Simplify Java path-affinity ordering logic

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b317a759-f6bc-4590-bd2a-628f0ee9c477

* Revert legacy DAG Java ambiguity ordering changes

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/94e50cf2-9733-4e69-a0eb-9fd38cbdb589

* Skip duplicate-FQN Java assertions in legacy parity mode

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/1b560efa-1b3b-4697-b590-c6ef447f431e

* Tighten duplicate-FQN Java CALLS edge cardinality assertions

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/67c18f93-5e56-4b15-8404-cdf1be9b4485

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 08:00:03 +01:00
Nilotpal KashyapandGergő Magyar d7e1815aa3 fix(detect-changes): guard resolveWorktreeCwd against overriding a separately-indexed worktree (#1691)
* fix(detect-changes): guard resolveWorktreeCwd against overriding a separately-indexed worktree

When the repo registry entry points to a linked worktree (both main
checkout and worktree indexed separately), resolveWorktreeCwd was
incorrectly replacing the correct worktree repoPath with the server's
main-checkout launch directory. Both share the same canonical root so
the existing same-repo check passed, causing git diff to run from the
wrong directory and return 0 changes (issue #1659).

Fix: early-exit guard — if tryRealpath(repoPath) differs from
tryRealpath(getCanonicalRepoRoot(repoPath)), repoPath is itself a
linked worktree and is returned unchanged. Auto-detection only fires
when repoPath equals the canonical main-checkout root.

Also normalises the launchCanonical comparison in the auto-detect path
to use tryRealpath for cross-platform consistency.

Regression test: 'returns worktreeDir unchanged when repoPath IS a
linked worktree and launchCwd is the main checkout'.

* test(detect-changes): add worktreeA→worktreeB case and assumption comment

Cover the missing case from the production-readiness review:
repoPath = wt-A (indexed), launchCwd = wt-B (server on a different
linked worktree). The guard fires on repoPath being a worktree
regardless of launchCwd, so wt-A is returned unchanged.

Also add an inline comment documenting the assumption that repoPath
is a git root or linked-worktree root (not an arbitrary subdirectory),
as noted in Finding 2 of the review.

* refactor(detect-changes): validate repoPath is a git root before canonical comparison

Instead of relying on a comment asserting repoPath is always a git
root, call getGitRoot(repoPath) first. Only if the result matches
repoPath itself do we call getCanonicalRepoRoot and apply the guard.

This eliminates the over-classification risk for subdirectory repoPath
values and makes the assumption explicit in code. repoCanonical is
shared across both the guard and the auto-detect block.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 06:46:16 +01:00
dependabot[bot] 92ad0f5491 chore(deps): bump idna in /eval in the uv group across 1 directory (#1713) 2026-05-20 05:38:26 +01:00
LocallyInsaneDBandGergő Magyar 803f0bed5f fix(lbug): probe-then-load FTS extension on Windows (#1690) (#1692)
* fix(lbug): probe-then-load FTS extension on Windows (#1690)

The Windows skip-on-process.platform==='win32' guard in pool-adapter.ts
hard-skipped loadFTSExtension() for every Windows host, even when the
FTS extension binary was already present locally at
~/.lbdb/extension/<version>/win_amd64/fts/libfts.lbug_extension.

That left BM25 silently degraded on Windows hosts that had a working
extension on disk, with no error path — `gitnexus doctor` still reported
FTS as available, but query returned 0 BM25 hits.

This patch adds hasLocalWinFtsExtension() which probes
~/.lbdb/extension/*/win_amd64/fts/ before the Windows skip. When a binary
is on disk we call loadFTSExtension(..., { policy: 'load-only' }); the
crashing install path documented in #1199 / #1217 is never exercised at
query time, and LadybugDB's version-specific resolution combined with
the ExtensionManager's tryLoad try/catch handles stale or zero-byte
sibling version dirs cleanly (no dlopen attempted on a stale binary).
When no binary is on disk at all, we fall back to the upstream skip so
install-time SIGSEGV continues to be avoided.

Verified on Windows 10 + Node 22.19.0 + gitnexus 1.6.5 +
@ladybugdb/core 0.16.1 with the FTS extension cached at 0.16.0:

  * BM25 timing goes from 0 → ~250-326ms on previously-zero queries
  * gitnexus context / impact / cypher unaffected
  * Adversarial-mixed-state run (real 0.16.0 binary + zero-byte stubs at
    0.15.0, 0.16.1, 0.17.0): exits 0, no SIGSEGV, FTS resolves to the
    real 0.16.0 binary, BM25 returns real hits
  * Stub-only state at the resolution path (0.16.0, zero-byte): exits 0,
    emits "FTS extension unavailable; load-only policy: extension not
    pre-installed", FTS marked unavailable cleanly via markUnavailable
    in extension-loader.ts — no silent greenlight

Closes #1690

* test(lbug): cover hasLocalWinFtsExtension probe + format pool-adapter

- Export hasLocalWinFtsExtension and add lbug-pool-win-fts-probe.test.ts
  with 7 cases against a real tmpdir + os.homedir spy:
    * missing ~/.lbdb/extension dir -> false
    * extension root present but no version dirs -> false
    * one version dir with binary present -> true
    * zero-byte stub at probe path -> true (LOAD failure handled downstream)
    * multi-version with binary only in a non-first dir -> true
    * multi-version with no binary anywhere (Nix/Bazel/MDM tree) -> false
    * fs.readdir throws (EACCES) -> false

  The Windows conditional in doInitLbug / initLbugWithDb is intentionally
  not unit-isolated: it reduces to `probe ? load : true` over a fully
  constructed lbug.Database + Connection pool, which the
  test/integration/lbug-pool*.test.ts suites already exercise on the
  windows-latest CI matrix.

- Apply prettier format to the fs.stat() call in pool-adapter.ts,
  resolving the quality/format CI failure surfaced by gitnexus/autofix.

Addresses DoD §2.7 test-coverage blocker raised in the production-
readiness review on #1692, and the dir-exists-no-file regression case
raised on #1690.

Refs #1690.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-19 12:09:21 +01:00
55f8d442f6 fix(mcp): setup fallback on Windows when global gitnexus resolves to a non-spawnable shim (#1694)
* Initial plan

* fix: avoid invalid Windows MCP shim paths

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a052306e-483a-42d0-b65a-2646906457c7

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: cover .ps1 windows mcp fallback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/5aaed570-2a0b-4ed9-a0ac-ca099ce5675e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: assert windows fallback for cursor and codex

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/5aaed570-2a0b-4ed9-a0ac-ca099ce5675e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-19 08:17:26 +01:00
18167400c4 chore(deps)(deps): bump express and @types/express in /gitnexus (#872)
* chore(deps)(deps): bump express and @types/express in /gitnexus

Bumps [express](https://github.com/expressjs/express) and [@types/express](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/express). These dependencies needed to be updated together.

Updates `express` from 4.22.1 to 5.2.1
- [Release notes](https://github.com/expressjs/express/releases)
- [Changelog](https://github.com/expressjs/express/blob/master/History.md)
- [Commits](https://github.com/expressjs/express/compare/v4.22.1...v5.2.1)

Updates `@types/express` from 4.17.25 to 5.0.6
- [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases)
- [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/express)

---
updated-dependencies:
- dependency-name: "@types/express"
  dependency-version: 5.0.6
  dependency-type: direct:development
  update-type: version-update:semver-major
- dependency-name: express
  dependency-version: 5.2.1
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>

* fix(server): normalize jobId param for Express 5 SSE routes

Express 5 types req.params values as string | string[]. mountSSEProgress uses a dynamic route path so TypeScript cannot narrow jobId; assert it once with assertString and reuse in the SSE progress callback.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(ci): retrigger CI

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-19 07:43:07 +01:00
dad1ca7ab5 chore(deps)(deps): bump zod from 3.25.76 to 4.3.6 in /gitnexus-web (#1464)
Bumps [zod](https://github.com/colinhacks/zod) from 3.25.76 to 4.3.6.
- [Release notes](https://github.com/colinhacks/zod/releases)
- [Commits](https://github.com/colinhacks/zod/compare/v3.25.76...v4.3.6)

---
updated-dependencies:
- dependency-name: zod
  dependency-version: 4.3.6
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:56:04 +01:00
c746f30c90 chore(deps)(deps): bump langsmith (#1552)
Bumps the npm_and_yarn group with 1 update in the /gitnexus-web directory: [langsmith](https://github.com/langchain-ai/langsmith-sdk).


Updates `langsmith` from 0.5.23 to 0.6.3
- [Release notes](https://github.com/langchain-ai/langsmith-sdk/releases)
- [Commits](https://github.com/langchain-ai/langsmith-sdk/commits)

---
updated-dependencies:
- dependency-name: langsmith
  dependency-version: 0.6.3
  dependency-type: indirect
  dependency-group: npm_and_yarn
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:55:44 +01:00
15a667ae5e chore(deps)(deps-dev): bump @vitest/coverage-v8 in /gitnexus (#1604)
Bumps [@vitest/coverage-v8](https://github.com/vitest-dev/vitest/tree/HEAD/packages/coverage-v8) from 4.1.5 to 4.1.6.
- [Release notes](https://github.com/vitest-dev/vitest/releases)
- [Commits](https://github.com/vitest-dev/vitest/commits/v4.1.6/packages/coverage-v8)

---
updated-dependencies:
- dependency-name: "@vitest/coverage-v8"
  dependency-version: 4.1.6
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:55:24 +01:00
6210d80f1e chore(deps)(deps-dev): bump tsx from 4.21.0 to 4.21.1 in /gitnexus (#1698)
Bumps [tsx](https://github.com/privatenumber/tsx) from 4.21.0 to 4.21.1.
- [Release notes](https://github.com/privatenumber/tsx/releases)
- [Changelog](https://github.com/privatenumber/tsx/blob/master/release.config.cjs)
- [Commits](https://github.com/privatenumber/tsx/compare/v4.21.0...v4.21.1)

---
updated-dependencies:
- dependency-name: tsx
  dependency-version: 4.21.1
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:55:04 +01:00
637cfca39c chore(deps)(deps): bump express-rate-limit in /gitnexus (#1697)
Bumps [express-rate-limit](https://github.com/express-rate-limit/express-rate-limit) from 8.5.1 to 8.5.2.
- [Release notes](https://github.com/express-rate-limit/express-rate-limit/releases)
- [Commits](https://github.com/express-rate-limit/express-rate-limit/compare/v8.5.1...v8.5.2)

---
updated-dependencies:
- dependency-name: express-rate-limit
  dependency-version: 8.5.2
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:54:33 +01:00
dependabot[bot]andGergő Magyar 73543a4714 chore(deps)(deps-dev): bump @types/node in /gitnexus (#1696)
Bumps [@types/node](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/node) from 25.6.2 to 25.7.0.
- [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases)
- [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/node)

---
updated-dependencies:
- dependency-name: "@types/node"
  dependency-version: 25.7.0
  dependency-type: direct:development
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-19 06:54:14 +01:00
DuduPhudu b37974fdac feat(javascript): migrate JavaScript to scope-based resolution (RFC #909 Ring 3, issue #928) (#1640) 2026-05-19 06:23:13 +01:00
dependabot[bot] ade2069633 chore(deps)(deps): bump brace-expansion from 5.0.5 to 5.0.6 in /gitnexus (#1689) 2026-05-19 05:35:27 +01:00
azizur100389 5f0c0eba0e feat(cpp): Expand type_traits constraint registry (#1648) 2026-05-18 21:10:18 +01:00
Gergő Magyar 2632bcccc0 fix(api): open lbug read-only for /api/graph, /api/search, /api/grep (#1686) 2026-05-18 19:57:35 +01:00
Gergő Magyar c9199b654f fix(test): retry Windows temp cleanup in cli-e2e teardown (#1688) 2026-05-18 18:17:54 +01:00
Shane Thurston Wijaya 33f18ceaa2 feat(eval-server): added --host for user configured host IP instead of system hardcoded IP (127.0.0.1) (#1667)
* feat(eval-server): added --host for user configured host IP instead of system hardcoded IP (127.0.0.1)

* fix(eval-server): localhost value in --host now returns 127.0.0.1 instead of the raw input to fix wrong address, handled error for ipv6 disabled containers

* feat(eval-server): add --host flag with validation and error handling

  Co-Authored-By: Val Vladescu <val.vladescu@thirdbridge.com>

* fix(eval-server): bracketed IPv6 addresses to remove ambiguity

* docs(eval-server): document --host flag, READY signal format, and parser migration note

* fix(eval-server): use actual bound port in READY signal; strengthen --host e2e tests

  Co-Authored-By: Val Vladescu <val.vladescu@thirdbridge.com>

* feat(eval): wire eval-server --host through gitnexus_docker.py

* docs(eval): added guidance for docker user

* docs(eval): revise the imprecise documentation

* fix(e2e): updated original stdout for new format
2026-05-18 16:00:42 +01:00
c30833fad3 perf(scope-resolution): use owner-keyed lookup for Step 2 member resolution (#1657)
* perf(scope-resolution): use owner-keyed lookup for Step 2 member resolution (#1656)

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(scope-resolution): index Const/Static in FieldRegistry for Step 2 lookup

Extend FieldRegistry to hold multiple defs per (owner, name), reconcile Const and Static into the owner-keyed index, and wire lookupAllByOwner through the production hook so Step 2 does not drop field kinds the registry never indexed. Pass explicitReceiver on read/write reference sites and document undefined-vs-empty hook semantics for defs fallback.

Co-authored-by: Cursor <cursoragent@cursor.com>

* perf(scope-resolution): centralize O(1) owned-member hook and guard hot path

Extract lookupOwnedMembersByOwner for the production Step 2 hook so merges stay O(1) per registry with no defs.byId scan. Add a perf-contract unit test that throws if byId.values runs when the hook is wired. Reuse a frozen empty sentinel on double miss to avoid per-probe allocations.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore: drop unused buildFieldRegistry import

* chore(scope-resolution): apply ce-code-review safe_auto fixes

- Drop unreachable return + unused values() capture in perf-contract trap (Finding #7)
- Type lookupOwnedMembersByOwner ownerDefId as DefId (Finding #9)
- Add Static-kind Step 2 lookup test mirroring the Const case (Finding #11)

* docs(field-registry): document lookupFieldByOwner first-wins semantics

Audit of all 6 production callers (call-processor.ts:2279, walkers.ts:535,
receiver-bound-calls.ts:380+730, type-env.ts:627+631) confirms none depends
on last-wins precedence — all treat the return as a generic 'field with
this name owned by this class'. Clarify the JSDoc to surface the semantic
change introduced when FieldRegistry moved from last-wins to append-order
storage (ce-code-review finding #2).

* test(scope-resolution): extend Step 2 perf contract to implicit-self, MRO, field paths

Adds three sibling tests under the Step 2 perf contract describe block, each
asserting defs.byId.values() does NOT execute when ownedMembersByOwner is wired:

- implicit-self receiver via typeBindings.self (no explicitReceiver branch)
- 2-level MRO chain (Child extends Parent, save resolves on Parent at depth 1)
- FieldRegistry read via Step 2 (property lookup, separate registry path)

Pins the perf invariant on every distinct entry into walkReceiverTypeBinding
so a regression bypassing the hook on any sub-path now fails CI immediately
(ce-code-review finding #8).

* test(resolve-references): cover arity-overload filtering via resolveReferenceSites

Pins the orchestration-layer wiring of providers.arityCompatibility:
hook returns [save(arity 1), save(arity 2)], referenceSite.arity = 1,
arityCompatibility verdicts 'compatible'/'incompatible' by parameterCount,
exactly one reference emitted with toDef = the arity-1 overload.

registries.test.ts already covered arity at the buildMethodRegistry level;
this adds the missing entry-point check that resolveReferenceSites threads
providers correctly through to lookupCore.Step5 (ce-code-review finding #10).

* test(resolve-references): add hook-on vs hook-off parity test

Runs resolveReferenceSites twice on the same fixture (Parent.save method
hit + Child.name field hit, Child extends Parent MRO chain) — once with
ownedMembersByOwner wired to a synthetic registry, once with the hook
absent so collectOwnedMembers takes the defs.byId fallback. Asserts:

- stats are identical (sitesProcessed / referencesEmitted / unresolved)
- referenceIndex.bySourceScope entries have equal length
- toDef sets are equal
- each per-site reference (including evidence and depth) is .toEqual

Locks the semantic-parity claim in code while both paths still exist.
Will be removed alongside the fallback in finding #1 (ce-code-review #3).

* test(typescript): probe Step 2 MRO walk against ambient (declare class) base

Adds typescript-ambient-base-class fixture with an export declare class
AmbientBase + Derived extends AmbientBase and a call site d.ambientMethod().
Integration assertions:

- Both classes are detected
- EXTENDS edge Derived → AmbientBase emitted
- CALLS edge to ambient.ts:ambientMethod resolved via MRO walk

Probes the ce-code-review #6 concern that ambient-only owners (whose
bodies are never parsed) might be silently skipped by Step 2 after the
owner-keyed lookup change. Result: the call resolves correctly — the
method signature inside the declare class body still flows through
reconcileOwnership into model.methods, so the hook returns the right
ancestor hits. Residual risk is empirically closed.

* feat(scope-resolution): route nested types via owner-keyed TypeRegistry

Closes the Step 2 contract footgun where 'hook returns [] = authoritative
miss' silently dropped any owned def whose NodeLabel was outside the
method/field if-chain in reconcileOwnership.

- TypeRegistry: add nestedByOwner Map + lookupAllByOwner(owner, simple)
  + registerByOwner(owner, simple, def). Mirrors MethodRegistry/
  FieldRegistry shape; cleared with the rest on cascade clear.
- reconcileOwnership: route class-like NodeLabels (Class/Interface/Enum/
  Struct/Union/Trait/TypeAlias/Typedef/Record/Delegate/Annotation/
  Template/Namespace) via types.registerByOwner. New nestedTypesRegistered
  stat. Idempotent skip via nodeId match.
- validateOwnershipParity: extend the I9 invariant check to nested types.
- lookupOwnedMembersByOwner: merge methods + fields + nested-type hits;
  short-circuit when any one source contributes the full result.

Unblocks future receiver-MRO registries that need to resolve 'Outer.Inner'
through the receiver's type-binding chain (ce-code-review finding #5a).

* refactor(scope-resolution): make ownedMembersByOwner required; delete byId fallback

Per ce-code-review finding #1, the optional-hook design encoded a silent
O(|defs|) perf cliff into the type system: any RegistryContext built
without the hook regressed Step 2 to scanning every def per probe with
no warning. Production wires the hook unconditionally; the fallback was
exercised only by tests.

- RegistryContext.ownedMembersByOwner: required, returns readonly
  SymbolDefinition[] (no | undefined). Implementations MUST return [] on
  authoritative miss.
- collectOwnedMembers in lookup-core.ts collapses to a one-line forward
  to the hook; the defs.byId.values() scan and simpleNameOf helper are
  deleted (simpleNameOf had no other consumers).
- ResolveReferencesInput.ownedMembersByOwner: required to match.
- Tests: drop three fallback-path tests (registries Const fallback,
  resolveReferenceSites no-hook fallback, resolveReferenceSites Const-
  undefined fallback) and the hook-vs-fallback parity test added by
  finding #3. makeCtx in registries.test.ts now defaults to a real
  owner-keyed scan over the test fixture defs so tests that don't care
  about the hook keep working.

* perf(free-call-fallback): cache global callables by simple name once per pass

pickUniqueGlobalCallable scanned scopes.defs.byId.values() on every
free-call fallback site. After PR #1656 fixed Step 2, this scan became
the dominant remaining O(|defs|) hot path on large repos (ce-code-review
finding #4).

- buildGlobalCallableIndex builds a Map<simpleName, SymbolDefinition[]>
  over scopes.defs once at the top of emitFreeCallFallback. Same filter
  the per-site scan applied: Function / Method / Constructor, keyed by
  the last .-segment of qualifiedName.
- pickUniqueGlobalCallable consumes the prebuilt index via O(1) Map.get
  instead of iterating every def. Per-site complexity drops from
  O(|defs|) to O(|defs with this simple name|).
- Cost: O(|defs|) once per pass instead of O(|defs| * |free-call sites|).

Subsequent narrowing (arity, conversion-rank) and the model-side fallback
(model.symbols.lookupCallableByName + model.methods.lookupMethodByName)
are unchanged.

* chore(autofix): apply prettier + eslint fixes via /autofix command

* ci: trigger build

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Gergo Magyar <abhigyan1.patwari@gmail.com>
2026-05-18 13:14:27 +01:00
Copilot 7d500390b9 fix: Use Ladybug native read-only enforcement and prepared statement execution for Cypher query paths (#1655) 2026-05-18 06:54:24 +01:00
Nilotpal Kashyap bdc0439a10 feat(detect-changes): support git worktrees (#1654) 2026-05-17 20:54:41 +01:00
Shane Thurston Wijaya 105efd0f7c feat(wiki): added --lang <lang> flags to gitnexus wiki for multilanguage wiki generation support (#1613) 2026-05-17 19:54:02 +01:00
Copilot 493827222d fix(ingestion): Raise analyze auto-heap to 16GB and tighten cross-platform OOM guidance for UE5-scale repositories (#1652) 2026-05-17 16:28:07 +01:00
Copilot ed50a6729f fix(wiki): Remove the hidden 60s default timeout, validate gitnexus wiki timeout/retry flags, and surface timeout errors (#1651) 2026-05-17 12:03:54 +01:00
Nilotpal Kashyap dfbe68ad24 fix(lbug): issue #1647, detect WAL corruption in schema init and surface recovery (#1650) 2026-05-17 10:46:45 +01:00
azizur100389andGergő Magyar 2376912ca7 feat(ingestion): Add C++ parameter type class sidecar (#1642)
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-16 21:44:26 +01:00
Zander Raycraft a4dfebd073 feat(cpp): sfinae filter (#1623)
* feat(cpp): SFINAE-aware overload filter — drops candidates whose enable_if_t / requires constraints fail (#1579)

* fix(cpp):  SFINAE follow-ups for is_integral_v/is_arithmetic_v bool and char support, an unqualified F1 test fixture, and parameter-lookup gap documentation (#1579) -> claude feedback

* revert: reverting all changes to .md files
2026-05-16 20:23:13 +01:00
Gergő Magyar 42d4fcaf6f chore: release v1.6.5 (#1645) 2026-05-16 17:11:25 +01:00
a26ac55fb0 fix(lbug): Recover gitnexus analyze from orphan LadybugDB sidecars when main DB file is missing (#1622)
* Initial plan

* fix: recover from orphan lbug sidecars on init

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6e8ea6e8-f9ab-46ff-9c1b-4d2c73a6452c

* test: strengthen orphan sidecar recovery coverage

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6e8ea6e8-f9ab-46ff-9c1b-4d2c73a6452c

* fix(lbug): only clean orphan sidecars when DB is missing

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a34217b5-0e98-4949-bae1-2a50933f291e

* test(lbug): cover no-cleanup path when db file exists

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a34217b5-0e98-4949-bae1-2a50933f291e

* test(lbug): use errno-shaped ENOENT mocks for sidecar recovery

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a34217b5-0e98-4949-bae1-2a50933f291e

* test(lbug): cover partial sidecar and unlink-failure recovery cases

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a34217b5-0e98-4949-bae1-2a50933f291e

* refactor(lbug): tighten ENOENT detection and test naming

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a34217b5-0e98-4949-bae1-2a50933f291e

* test(lbug): normalize errno mock helpers across sidecar tests

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a34217b5-0e98-4949-bae1-2a50933f291e

* docs(lbug): annotate orphan `.wal.checkpoint` cleanup provenance

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/7edf5156-43e0-412d-87a4-bf4b2934deac

* test(lbug): clarify unlink-failure path test intent

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/7edf5156-43e0-412d-87a4-bf4b2934deac

* fix(lbug): handle orphan-sidecar cleanup error paths explicitly

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/93294b2c-57f6-459c-8eb2-86e3b8920fb0

* refactor(lbug): extract errno and error-summary helpers

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/93294b2c-57f6-459c-8eb2-86e3b8920fb0

* test(lbug): expand non-ENOENT lstat coverage and remove magic number

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/93294b2c-57f6-459c-8eb2-86e3b8920fb0

* test(lbug): add native integration test for orphan sidecar recovery

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2dd28264-4604-430a-a249-af52afd29245

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(lbug): annotate best-effort catch in integration test cleanup

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2dd28264-4604-430a-a249-af52afd29245

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(lbug): add cross-process init lock for orphan sidecar cleanup with integration tests

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e4cbcfec-a252-449d-8d65-2f3570a253f8

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor(lbug): use INIT_LOCK_STALE_MS in stale lock detection and address review feedback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e4cbcfec-a252-449d-8d65-2f3570a253f8

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* style(lbug): fix Prettier line-length violation in acquireInitLock fs.open call

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a140b567-0e9b-4ec9-a158-9fe6b8685ec2

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(lbug): ensure parent directory exists before creating init lock file

acquireInitLock tried to create `${dbPath}.init.lock` using O_CREAT | O_EXCL,
but on a fresh repo the parent directory (`.gitnexus/`) doesn't exist yet —
the mkdir call was inside the locked section. This caused ENOENT failures
on all platforms (Windows, macOS, Ubuntu) during `gitnexus analyze`.

Move mkdir to before the lock file creation attempt.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6883dc3c-36eb-4907-bcd8-61d23e2c641a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(lbug): verify acquireInitLock succeeds when parent directory does not exist

Adds an integration test proving the fix from the previous commit:
acquireInitLock now creates the parent directory before attempting
to create the lock file, preventing ENOENT on fresh repos.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6883dc3c-36eb-4907-bcd8-61d23e2c641a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-16 11:45:32 +01:00
azizur100389andGergő Magyar 467c14caa2 feat(cpp): standard-conversion-sequence ranking for overload resolution (#1606)
* feat(cpp): add standard-conversion-sequence ranking to overload resolution (#1578)

Introduce `ConversionRankFn` abstraction and `cppConversionRank` implementation
to disambiguate C++ overloaded calls by argument-to-parameter conversion cost.
Exact type match (rank 0) beats standard arithmetic conversion (rank 2), which
beats non-viable mismatch (Infinity). Thread the rank function through
`narrowOverloadCandidates`, `pickImplicitThisOverload`, `pickOverload`, and
`pickUniqueGlobalCallable` via the `ScopeResolver.conversionRankFn` contract.
Add `findAllCallableBindingsInScope` scope walker for collecting all overloads
at the first binding scope. Guard against false ambiguity suppression when
candidates span different files (local-shadows-import preservation).

* fix: address Claude review findings on conversion-rank PR

Finding 1 (HIGH): add tests that exercise the conversion ranker.
  - p('a') with p(int)/p(double): char→int promotion (rank 1) beats
    char→double conversion (rank 2), forcing step 4b in
    narrowOverloadCandidates. Exact-type filter misses both overloads.
  - h(42, 2.5) with h(int,int)/h(double,double): multi-arg tied total
    score forces the ranker, both candidates score 2 → suppressed.

Finding 2 (HIGH): unify multi-candidate suppression across all paths.
  - Non-ADL free-call: suppress when narrowed.length > 1 (same-file
    guard), mirroring ADL merged-candidate behavior.
  - ADL ordinary-only: same pattern.
  - pickOverload: return OVERLOAD_AMBIGUOUS when candidates.length > 1
    after normalized-ambiguity check.
  - Case 0.5 (this receiver): set ambiguous=true when narrowed > 1.

Finding 3+4 (MEDIUM): implement rank-1 integral promotions.
  - char→int and bool→int now return rank 1 (ISO C++ [conv.prom]).
  - Updated comment to remove misleading ISO table header; document
    only the post-normalization ranking that is actually implemented.
  - Updated ConversionRankFn JSDoc in overload-narrowing.ts.

218/218 C++ tests pass (registry-primary). Legacy: 186+32.

* fix: implement pairwise dominance comparison for overload ranking

Replace the summed per-slot conversion cost with ISO C++-aligned
pairwise dominance comparison ([over.ics.rank]). F1 is better than
F2 only when F1 is not worse for every argument and strictly better
for at least one. Non-dominated candidates are returned; if multiple
remain they are genuinely ambiguous.

This fixes false CALLS edges for asymmetric multi-arg overloads:
h('a', 2.5) against h(int,int) / h(double,double) — the old summed
cost picked h(double,double) (cost 2 < 3), but ISO C++ considers
the call ambiguous because h(int,int) is better at arg 0 via char
promotion. The pairwise check correctly finds neither dominates.

Add h('a', 2.5) test case asserting zero CALLS edges alongside
the existing h(42, 2.5) symmetric-tie test.

218/218 C++ tests pass (registry-primary). Legacy: 186+32.

* docs: update step 4b JSDoc to reflect pairwise dominance

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-16 11:15:21 +01:00
fa06c5610b fix: resolve cross-file type propagation stall on large repos (#1626)
* Initial plan

* fix: add time-based deadline to cross-file type propagation to prevent stalling on large repos

Adds a 2-minute wall-clock time limit (DEFAULT_CROSS_FILE_ELAPSED_MS) to
runCrossFileBindingPropagation. When exceeded, the phase gracefully stops
and logs a warning. Users can override via GITNEXUS_CROSS_FILE_TIMEOUT_MS
env var. This prevents the analyze command from stalling for hours on very
large repositories where per-file re-resolution is expensive.

Fixes the reported issue where gitnexus analyze stalls at "Cross-file type
propagation" for several hours on repos with 15000+ files.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b8341947-557c-4111-a3a8-991ba455ab01

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: root cause - cache tree-sitter queries across files, add live progress reporting

Root cause: cross-file propagation called processCalls() with 1 file at a time,
causing Parser.Query to be recompiled from the query string for every single file
(O(N) compilations vs O(1) for the whole phase). Additionally, progress was only
reported once at the start, making the phase appear completely frozen.

Fixes:
- Add optional `compiledQueryCache` parameter to `processCalls` so callers that
  invoke it with single-file batches can share compiled query objects across calls.
  The cross-file phase now compiles each language's query string exactly once and
  reuses it for all files of that language (e.g. 1 TypeScript compile for 595+ files).
- Pre-count candidate files and emit onProgress every 25 files showing
  "Cross-file type propagation (N/M files)..." so the UI shows real movement
  instead of a frozen bar.
- Keep the wall-clock deadline (GITNEXUS_CROSS_FILE_TIMEOUT_MS) as a safety
  net for pathological inputs.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/f5028cc8-4bc9-4309-8ffb-798fe2bd7a0a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: address code review - use SupportedLanguages key type, rename queryCache to compiledQueryCache

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/f5028cc8-4bc9-4309-8ffb-798fe2bd7a0a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(cross-file): remove wall-clock timeout from type propagation

The query compilation cache and live progress reporting address the
original stall; the 2-minute deadline could truncate cross-file work on
large repos. MAX_CROSS_FILE_REPROCESS (2000) remains as the only cap.

* test(cross-file): verify compiledQueryCache is shared across all processCalls invocations

Finding 1: O(N) query recompilation was fixed by sharing a compiledQueryCache Map
across all processCalls invocations in runCrossFileBindingPropagation. This test
verifies the fix is correctly wired: the same Map instance is passed as the
12th argument to every call, proving queries are compiled once per language,
not once per file.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/3ab768d9-3993-4882-9d8f-17f7fcbd086e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cross-file): verify live progress events are emitted with N/M format

Finding 2: frozen progress display was fixed by emitting onProgress every 25 files
with "Cross-file type propagation (N/M files)..." messages instead of calling it
once at phase start. This test verifies the fix with 50 candidate files: expects
onProgress called 3 times (1 initial + at 25 + at 50) with correct N/M counters.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/3ab768d9-3993-4882-9d8f-17f7fcbd086e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(cross-file): skip registry-primary language files before readFileContents

Finding 3 (from comment 4466231612): cross-file-impl was calling processCalls
for every candidate file even when that file's language is registry-primary
(TypeScript, C++, Python, Go, C#, PHP, C — since AGENTS.md v1.7.0). processCalls
would immediately skip those files via its own isRegistryPrimary guard, but
cross-file-impl still paid the full cost: readFileContents I/O, buildImportedReturnTypes,
buildImportedRawReturnTypes, and Map allocation — all discarded.

Fix: check isRegistryPrimary(lang) in both the totalCandidates pre-count loop
and the levelCandidates builder, before any file I/O or map building. This
eliminates 595+ no-op processCalls invocations on large TypeScript repos.

Test: mocks isRegistryPrimary to always return true and verifies that
processCalls is never invoked and result is 0. The mock also defaults to false
in beforeEach so existing tests using .ts files are unaffected.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/3ab768d9-3993-4882-9d8f-17f7fcbd086e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor(test): address code review - simplify mock factory, name the arg index constant

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/3ab768d9-3993-4882-9d8f-17f7fcbd086e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-16 10:02:40 +01:00
Gergő Magyar f28185d67e fix(ci): bump publish job to Node 24 for npm OIDC support (#1628)
PR #1627's npm install -g npm@latest step crashed mid-install with MODULE_NOT_FOUND: promise-retry — a known fragility when npm self-upgrades. Node 22's bundled npm is 10.9.x (no OIDC). Fix: bump publish job's node-version to 24, which ships with npm 11.x natively. Package consumers unaffected (this Node version is only used during publish; engines.node is >=22.0.0; ci-tests.yml continues testing on Node 22).
2026-05-16 09:18:08 +01:00
Gergő Magyar f69c382bcb fix(ci): engage npm Trusted Publishing OIDC properly (#1627)
First live-fire RC publish after #1610 failed at npm publish with E404. The if: failure() cleanup correctly auto-deleted the partial v-tag and rc-marker, but OIDC never engaged. Root cause: two coordinated upstream bugs.

1. actions/setup-node@v6 with registry-url: writes _authToken into the runner .npmrc AND exports NODE_AUTH_TOKEN from its token: input (defaulting to github.token). npm publish sends GITHUB_TOKEN as the bearer and the registry returns 404. OIDC never tried because npm thinks it already has a credential. See actions/setup-node#1440.

2. The Node 22 runner ships with npm 10.9.x. npm Trusted Publishing OIDC support requires npm >= 11.5.1.

Fix: omit registry-url: from the setup-node step (per the consensus workaround in community discussion #176761), and add npm install -g npm@latest before publish. --provenance flag is NOT added; npm auto-attaches provenance under Trusted Publishing.

Sources:
- https://github.com/actions/setup-node/issues/1440
- https://github.com/orgs/community/discussions/176761
- https://docs.npmjs.com/trusted-publishers/
2026-05-16 08:36:49 +01:00
Gergő Magyar 83fbd4be26 refactor(ci): unify release pipeline under publish.yml (#1610)
Collapse release-candidate.yml into publish.yml so there is exactly one workflow that publishes gitnexus to npm, creates GitHub Releases, and triggers Docker builds — for both release candidates and stable releases. Closes #1609 architecturally.

A first-stage `route` job classifies push-to-main / push-tag / workflow_dispatch into `rc` / `stable` modes and fails closed on malformed shapes. RC path runs rc-guard → ci.yml → publish (mint GitHub App token → checkout with persist-credentials:false → resolve next rc version → atomic v-tag + rc/<SHA> marker push → vtag integrity gate → npm publish via OIDC → GitHub prerelease → if: failure() cleanup) → docker.yml. Stable path verifies package.json matches the tag and publishes to `latest` via OIDC (no docker).

Hardening:

  • Self-trigger prevention via negative-glob `tags: ['v*', '!v*-rc.*']` — the bug class behind #1609 cannot recur.
  • Two distinct actions/checkout steps per mode (no conditional `token:` expression footgun).
  • Workflow-level `permissions: {}` deny-all + per-job grants; `id-token: write` only where OIDC is used.
  • npm Trusted Publishing replaces NPM_TOKEN (delete the secret after the first successful publish).
  • GitHub App installation token (actions/create-github-app-token@v3.2.0) replaces the long-lived RELEASE_PUSH_TOKEN PAT (delete after first successful RC).
  • vtag integrity gate fails closed on empty / mode-mismatched output (prevents Release named `main` from a github.ref fallback).
  • Annotation-injection sanitization on every logged ref.
  • Explicit `secrets:` passthrough on docker.yml (DOCKERHUB_USERNAME, DOCKERHUB_TOKEN); ci.yml no longer inherits anything.
  • `if: failure()` cleanup auto-deletes v-tag + rc-marker on partial failure (eliminates the external-consumer phantom-version ingestion window).
  • ACTIONS_STEP_DEBUG window closed via `set +x` wrap on the inline auth-header compute.
  • Curated retry-loud error handling on `gh api` bot-user-id lookup and `npx semver`.

Pre-merge validation:

  • 10-reviewer multi-agent code-review pass; 14 findings fixed inline (commit 820cefae), 6 deferred to follow-ups.
  • End-to-end dry-run rehearsal via workflow_dispatch (run 25919563064) validated route classification, rc-guard, App token mint, RC checkout, version resolver, vtag synthetic-regex check, and faithful tarball pack at the bumped version.
  • All zizmor findings on the unification commits closed.
  • Branch-protection required checks all green.

Post-merge actions:

  • After the first successful RC, delete the `NPM_TOKEN` and `RELEASE_PUSH_TOKEN` secrets — they are no longer used.
  • The first real RC after merge is the live-fire test for steps dry-run could not exercise (atomic tag push, real npm OIDC handshake, GitHub Release creation, docker.yml under explicit secrets passthrough). The if: failure() cleanup step handles the partial-failure recovery automatically; the Rollback Runbook in CONTRIBUTING.md covers the rare cases auto-cleanup can't reach.
2026-05-16 07:46:56 +01:00
263ca353a6 fix: shard parse cache persistence on large repos (#1580)
* fix: shard parse cache persistence on large repos

* fix(parse-cache): validate shard keys, docs, and sharded-cache tests

- Reject non-sha256-hex keys from index.json before path.join (path traversal).

- saveParseCache: skip invalid keys defensively; try/catch per-shard JSON.stringify.

- Clarify save comment (tmp dir + rename vs atomic).

- Tests: hex keys throughout, traversal keys, multi-shard, version-mismatch+legacy, second save, legacy removal.

- AGENTS.md / GUARDRAILS.md: document .gitnexus/parse-cache/ vs legacy parse-cache.json.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-16 07:19:06 +01:00
azizur100389 8500f18e5f fix(cpp): detect same-name ambiguity across inline namespace children (#1564) (#1600) 2026-05-15 19:11:09 +01:00
Copilotandmagyargergo aed370b931 feat: C++ ADL V2: merge ordinary and ADL free-call candidates before overload selection (#1599)
* Initial plan

* Merge C++ ADL and ordinary free-call candidate sets

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/1aea3511-3471-4ec2-9819-0fb27ac40b89

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* Address review feedback on merged ADL ambiguity suppression

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/1aea3511-3471-4ec2-9819-0fb27ac40b89

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: apply prettier to C++ ADL resolver fallback files

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9b9c1494-bc69-4db5-a89d-69eb816bab82

* docs: update ADL ambiguity comments to merged narrowing flow

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9b9c1494-bc69-4db5-a89d-69eb816bab82

* fix: suppress global fallback when merged ADL narrowing yields zero candidates

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9b9c1494-bc69-4db5-a89d-69eb816bab82

* docs: clarify free-call fallback comment for ADL merged path

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9b9c1494-bc69-4db5-a89d-69eb816bab82

* feat: ADL Gap 2 — enum-typed arguments contribute enclosing namespace

ISO C++ [basic.lookup.argdep] §2: "If T is an enumeration type, its
associated namespace is the namespace in which it is defined."

- Add Enum to findCppClassDefBySimpleName type filter
- Map Enum defs to enclosing namespace in populateCppAssociatedNamespaces
- Add test fixture cpp-adl-enum-arg with color::Channel enum

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/ca8987b6-365e-4034-af56-ca3f9b439902

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* feat: ADL Gap 6 — inline namespace expansion in associated set

ISO C++ inline namespaces are transparent for ADL: if a namespace is
in the associated set, candidates declared in its inline-namespace
children are also reachable.

- Expand pickCppAdlCandidates to scan inline-namespace children of
  associated namespaces (via isCppInlineNamespaceScope predicate)
- Add test fixture cpp-adl-inline-ns-expansion: Event in outer audit,
  record in inline v1, other::record(int) forces arity disambiguation

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/ca8987b6-365e-4034-af56-ca3f9b439902

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* feat: ADL Gap 1 — hidden friend functions visible via ADL

ISO C++ [basic.lookup.argdep] §2: friend functions declared inside a
class body are visible via ADL when the class is an associated class.

- Exempt friend_declaration from cppLabelOverride's class-body function
  suppression (c-cpp.ts) so friend function defs are captured
- Scan Function scopes that are direct children of associated Class
  scopes in pickCppAdlCandidates (adl.ts) to find hidden friends
- Add test fixture cpp-adl-hidden-friend: `friend void process(Foo&)`
  declared inside lib::Foo, resolved via ADL from app::run()

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/ca8987b6-365e-4034-af56-ca3f9b439902

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* feat: ADL Gap 3 — non-function ordinary lookup suppresses ADL

ISO C++ [basic.lookup.unqual] §7: if ordinary unqualified lookup finds
a name that is not a function or function template, ADL is not performed.

- Add hasNonCallableBindingInScope walker in walkers.ts
- In free-call-fallback, check for non-callable binding before invoking
  ADL; when found, bypass resolveAdlCandidates entirely
- Add test fixture cpp-adl-non-function-blocks: variable `int record`
  shadows the function name, blocking ADL from finding audit::record

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/ca8987b6-365e-4034-af56-ca3f9b439902

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: use nearest-scope semantics for ADL non-callable blocker check

Finding 1: `hasNonCallableBindingInScope` walked the entire scope chain,
which could incorrectly suppress ADL when an inner scope had a callable
and an outer scope had a non-callable for the same name. Per ISO C++
`[basic.lookup.unqual]` §7, ADL is blocked only when ordinary lookup
itself finds a non-function — if ordinary lookup stops at an inner scope
where only callables exist, ADL should still fire.

Replace the separate `hasNonCallableBindingInScope` + `findAllCallable
BindingsInScope` calls with a combined `findCallableBindingsAndAdlBlocker`
walker that stops at the first scope with ANY binding for the name and
returns both `{ callables, nonCallableFound }`. One pass, one stop.

Fixture: cpp-adl-inner-callable-outer-noncallable — inner scope has
callable `swap(int,int)`, outer scope has `int swap = 0`. ADL fires and
resolves to `data::swap(Pair&,Pair&)` via argTypes narrowing.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a2f97daf-17fd-4891-8b10-a81e44d32808

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: block-scope function declaration suppresses ADL

Finding 2: ISO C++ [basic.lookup.argdep] lists three ADL blockers:
1. class member declaration (handled by pickImplicitThisOverload)
2. block-scope function declaration NOT a using-declaration (NEW)
3. non-function/non-template declaration (handled by nonCallableFound)

Extend `findCallableBindingsAndAdlBlocker` to return `blockScopeDeclFound`
when a callable is found at a Function or Block scope — indicating a local
forward declaration that should suppress ADL per standard.

`free-call-fallback.ts` now checks both `nonCallableFound` and
`blockScopeDeclFound` to determine ADL suppression.

Fixture: cpp-adl-block-scope-decl-blocks — `void record(int);` declared
inside function body prevents ADL from discovering audit::record.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a2f97daf-17fd-4891-8b10-a81e44d32808

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docs: update stale ADL_AMBIGUOUS comment in unqualified-ref-collision fixture

Finding 3: The `ADL_AMBIGUOUS` sentinel was removed by this PR (replaced
by `isOverloadAmbiguousAfterNormalization` in merged-narrowing). Update
the fixture comment to reference the current mechanism.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a2f97daf-17fd-4891-8b10-a81e44d32808

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: add legacy-parity expected failures for ADL blocker tests

The new ADL nearest-scope blocker and block-scope function declaration
tests rely on scope-resolution-only mechanisms not present in the legacy
DAG path. Register them in LEGACY_RESOLVER_PARITY_EXPECTED_FAILURES.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a2f97daf-17fd-4891-8b10-a81e44d32808

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore: revert unrelated prettier-plugin-tailwindcss devDep addition

The `prettier-plugin-tailwindcss` dependency was accidentally added while
running local prettier; it is not needed for the C++ ADL changes.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a2f97daf-17fd-4891-8b10-a81e44d32808

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-05-15 15:41:14 +01:00
99b8c7b03b feat: C++ ADL V2: free-function reference args contribute enclosing namespace (#1598)
* Initial plan

* cpp ADL V2: free-function reference args contribute enclosing namespace

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/24805583-c0c4-4ef8-978f-b874bd917947

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* merge: resolve conflicts with origin/main and fix overloaded fixture app.cpp

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/136aeffe-45da-47e2-95dd-e3883e85fad7

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(finding-1): replace ISO C++ [basic.lookup.argdep] misstatement with GitNexus-approximation label

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/fa37c0dd-65f9-4dd4-9811-617227a37073

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(finding-2): verify Function/Method exists in namespace before contributing via qualified_identifier arg

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/fa37c0dd-65f9-4dd4-9811-617227a37073

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(finding-3): function parameters in parameter_list no longer misclassified as free-function refs

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/fa37c0dd-65f9-4dd4-9811-617227a37073

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* doc(finding-4): document typedef/using-aliased function-pointer limitation in lookupAdlIdentifierType

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/fa37c0dd-65f9-4dd4-9811-617227a37073

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(finding-5): add negative fixtures for local-fp shadowing free-func and unqualified namespace collision

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/fa37c0dd-65f9-4dd4-9811-617227a37073

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(legacy-parity): skip two new negative-fixture tests from legacy DAG parity run

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/64ffbaf1-f442-4a2b-8542-4afa500d9182

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-15 10:55:44 +01:00
813acd7ec5 feat: C++ ADL V2: include base-class associated namespaces via MRO (#1597)
* Initial plan

* fix(cpp): include base-class namespaces in ADL candidate selection

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/7df6f692-1af9-43e6-82de-099ed43a60cb

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cpp): clarify ADL base-namespace test names

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/7df6f692-1af9-43e6-82de-099ed43a60cb

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cpp): remove stale legacy parity expected-failure entry

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/1a55a5e8-ae91-44bc-9b21-9324cdfea3de

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cpp): assert base-namespace ADL tests are not parity skips

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b3726b70-e797-4f37-955d-7d61fd28d338

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(cpp): avoid MRO amplification on ambiguous class-name ADL lookup

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b3726b70-e797-4f37-955d-7d61fd28d338

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cpp): strengthen ADL base-namespace target identity assertions

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b3726b70-e797-4f37-955d-7d61fd28d338

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cpp): add ADL negative cases for anonymous and unresolved bases

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b3726b70-e797-4f37-955d-7d61fd28d338

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cpp): fix anonymous-base parity expectation and formatting

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6a2e3cf9-beea-435c-8494-6a7a00af0f1e

* fix(cpp): propagate unnamed-namespace members through #include in registry-primary resolver

Anonymous-namespace contents in a header (e.g. `namespace { void f(); }`)
are reachable by unqualified lookup in any TU that #includes the header
per ISO C++ [basic.namespace.anon]/1 (the unnamed namespace behaves as
if a `using namespace unique;` is inserted into the enclosing scope, with
per-TU `unique`). The registry-primary path was filtering these defs out
of `expandCppWildcardNames` via both the structural Namespace-owner check
and the `isFileLocal` mark, so `hidden_probe(d)` from a TU including the
header resolved to nothing while the legacy DAG returned the correct edge.

Track anonymous-`namespace_definition` source ranges at capture time,
resolve them to ScopeIds in `populateOwners` (parallels inline-namespace
handling), and exempt those scopes from the two wildcard-expansion filters
plus the `populateCppNonGloballyVisible` structural set. `markFileLocal`
is preserved so the global free-call fallback still blocks cross-TU leaks
for files that do NOT #include the declaring file (cpp-anon-ns-cross-file
guard still passes).

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-05-15 07:54:23 +01:00
Copilot 7fbf302018 feat: C++ ADL V2: include template-specialization associated namespaces (with nested template args) (#1596) 2026-05-15 05:04:42 +01:00
dependabot[bot] 5ee1122330 chore(deps)(deps-dev): bump vitest from 4.1.5 to 4.1.6 in /gitnexus (#1605) 2026-05-14 22:28:31 +01:00
Copilot cdac8a691a feat: C++ ADL V2: include class-typed reference args (incl. rvalue refs) in associated-namespace lookup (#1595) 2026-05-14 20:25:12 +01:00
b00ba2ab47 feat(cpp): resolve template-body this-> + using ns::name calls in scope resolver (#1590)
* Initial plan

* fix(cpp): resolve this-> and using-name calls in template bodies

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/d9d91945-f19c-4fd2-9b52-b0ebc9aa34b6

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(cpp): treat duplicate using-name hits as ambiguous

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/d9d91945-f19c-4fd2-9b52-b0ebc9aa34b6

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(cpp): gate this-receiver path and harden overload semantics

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/030a1842-c698-460d-ae2a-95037e6def73

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cpp): add positive this-> overload case and document field shadowing

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/030a1842-c698-460d-ae2a-95037e6def73

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cpp): skip new template-this assertions in legacy parity lane

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/27002f6e-6331-41e3-8175-9d9e4691927c

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-14 18:18:36 +01:00
CopilotandGergő Magyar c2193318b5 feat(cpp): Enable C++ ADL for class pointer arguments and exclude function pointers (#1592)
* Initial plan

* fix: unwrap cpp adl pointer argument types

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2e9c8549-e062-410c-9ce3-66ba0a181590

* chore: tighten cpp adl function-pointer guard

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2e9c8549-e062-410c-9ce3-66ba0a181590

* docs: clarify cpp adl implementation comments

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2e9c8549-e062-410c-9ce3-66ba0a181590

* fix: avoid aborting cpp adl declaration scan

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e54f1d4b-9aac-407c-9b5e-b5f3ea0534ea

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-14 17:56:32 +01:00
8b2d8018bc fix(cli): tolerate read-only workspace in ensureGitNexusIgnored (#1549) (#1550)
* fix(cli): tolerate read-only workspace in ensureGitNexusIgnored

The documented Docker workflow mounts the host workspace at /workspace:ro
and runs `gitnexus index /workspace/<repo>` against an index produced by
a prior host-side `analyze`. Since PR #1248 ("keep GitNexus ignores
inside .gitnexus") the index command has called `ensureGitNexusIgnored`,
which unconditionally writes `<repo>/.gitnexus/.gitignore` and
`<repo>/.git/info/exclude` — both fail with EROFS on the :ro bind mount
even though the host already wrote the correct file during `analyze`.

Two complementary changes:

1. Idempotent fast path. Read the existing .gitnexus/.gitignore content
   first; if it already matches the desired value (`*\n`), skip the
   write entirely. This is the common case for the Docker workflow and
   avoids touching the FS at all.

2. EROFS/EACCES tolerance. When a write is genuinely needed but the FS
   refuses it, log a structured warning via the existing pino logger
   and continue. `registerRepo` runs before `ensureGitNexusIgnored` in
   `indexCommand`, so the global-registry write is already committed
   when we get here — letting the gitignore-write failure propagate
   leaves the user with a registered-but-error-exited command.

Three new unit tests pin the behaviour:
- idempotent re-call leaves mtime untouched
- ENOENT-then-correct path on a writable parent succeeds
- :ro parent (simulated via chmod 0o555) does not throw, on the
  already-correct fast path and on the cold-create path

Existing tests (61) still pass.

Closes #1549.

* test(storage): cover read-only ignore paths and tolerate EPERM (#1550)

- Add isReadOnlyFilesystemError helper including EPERM alongside EROFS/EACCES
  for ensureGitNexusIgnored and ensureGitInfoExclude (Windows parity with
  lbug-config / bridge-db patterns).
- Skip chmod-based read-only tests on win32 and uid 0; assert logger.warn
  on POSIX chmod denial for missing .gitignore.
- Add repo-manager-ensure-ignore-readonly.test.ts with vi.mock fs/promises
  delegating writeFile so EROFS/EACCES/EPERM rejections are asserted with
  structured log path and message for both .gitignore and .git/info/exclude.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-14 17:26:34 +01:00
89c03b2ebb fix: skip Claude augment hook when GitNexus server owns DB (#1493)
* fix(claude): skip augment hook when server owns db

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(hooks): cross-platform DB lock probe for MCP owner guard

Extract hook-db-lock-probe.cjs with a single hasGitNexusDbLockedByGitNexusServer
entry point used by both Claude hooks:

- Linux: scan /proc/<pid>/fd via dev+inode (no lsof required), optional lsof
  fallback; GITNEXUS_HOOK_LINUX_PROC_BUDGET_MS caps scan time
- macOS and other Unix: trusted lsof + ps (absolute paths / env overrides)
- Windows: Restart Manager + Win32_Process via win-rm-list-json.ps1 and
  GITNEXUS_HOOK_POWERSHELL_PATH

Update hooks.test.ts source coverage for the probe module.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Update gitnexus/hooks/claude/win-rm-list-json.ps1

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* Apply suggestion from @github-actions[bot]

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(gitnexus): repair package.json JSON after malformed engines edit

Co-authored-by: Cursor <cursoragent@cursor.com>

* Update Node.js engine version requirement to 22.0.0

* Update Node.js engine version to >=22.0.0

* fix(hooks): address ce-code-review findings on PR #1493

P0:
- Replace malformed `RM_UNIQUE_PROCESS` block in
  `gitnexus/hooks/claude/win-rm-list-json.ps1` (duplicate struct decl +
  duplicate `ProcessStartTime` + unbalanced braces) with a single
  well-formed `[StructLayout(LayoutKind.Sequential, Pack = 4)]` struct,
  so PowerShell `Add-Type` actually compiles and the Windows DB-lock
  probe stops fail-open on every machine.
- `gitnexus/src/cli/setup.ts` now copies `hook-db-lock-probe.cjs` and
  `win-rm-list-json.ps1` into the user's `~/.claude/hooks/gitnexus/`
  alongside `hook-lock.cjs`, preventing the `MODULE_NOT_FOUND` thrown
  by `gitnexus-hook.cjs:18`'s top-level require on every fresh install.
  `gitnexus/test/unit/setup.test.ts` extended to assert both new copy
  destinations.
- Four fail-open hook tests (`ENOENT lsof`, `npx parent line`,
  `non-GitNexus ps line`, `ps ENOENT`) now seed `createHookToolDir`
  with a valid `[GitNexus]` stderr line so
  `expect(parseHookOutput).not.toBeNull()` actually holds on CI.

P1:
- Plugin copy of `win-rm-list-json.ps1` gains `Pack = 4` so its CLR
  struct matches the 12-byte native `RM_UNIQUE_PROCESS` layout
  (multi-blocker `RmGetList` no longer reads mangled `dwProcessId`).
- `GITNEXUS_HOOK_CLI_PATH = ''` now falls through to the resolution
  chain in `gitnexus-hook.cjs`, matching the plugin copy and removing
  the twin-file divergence on empty-string envs.
- Lock-warning suppression test seeds `gitnexusMarkerPath` and asserts
  the augment subprocess actually ran, plus `GITNEXUS_DEBUG=1`
  preserves the full discarded prefix.
- MCP-owner skip branch in both hook copies now emits
  `[GitNexus] augment skipped: MCP server owns DB` on stderr, so
  agents can distinguish intentional skip from silent failure.

P2:
- `ps` loop in `hook-db-lock-probe.cjs` fails-closed on `ETIMEDOUT`
  to mirror the `lsof` handling (symmetric subprocess-probe contract).
- `RmStartSession` return value captured in both `.ps1` copies; exits
  early with `[]` on non-zero so subsequent RM API calls don't operate
  on an invalid handle.
- Windows RM-list `.ps1` encoded cache distinguishes uninitialized
  (`undefined`) from load-failed (`null`) with a one-shot
  `GITNEXUS_DEBUG` warning instead of silently caching empty string.
- `createHookToolDir` helper accepts `lsofOutputLines` and
  `psOutputByPid`; the multi-PID test uses them instead of duplicating
  the fake-binary construction inline.
- All five skip-path tests now assert `result.status === 0` and the
  new skip-signal stderr line.
- `AGENTS.md` documents the seven hook configuration env vars
  (`GITNEXUS_HOOK_CLI_PATH`, `_LSOF_PATH`, `_PS_PATH`,
  `_POWERSHELL_PATH`, `_LINUX_PROC_BUDGET_MS`, `_RM_TARGET`,
  `GITNEXUS_DEBUG`).
- `GITNEXUS_DEBUG` path in `gitnexus-hook.cjs`/`.js` writes the full
  discarded stderr prefix instead of a 180-char preview.
- Inline comment in `hook-db-lock-probe.cjs` explains the intentional
  Windows ETIMEDOUT fail-closed semantics.
- Removed the unnecessary `as WriteFileOptions` cast and orphaned
  `import type { WriteFileOptions }` in `hooks.test.ts`.

P3:
- `isGitNexusServerCommand` unexported from
  `hook-db-lock-probe.cjs` (kept as private helper).
- Env-path overrides (`GITNEXUS_HOOK_CLI_PATH`,
  `_POWERSHELL_PATH`, `_LSOF_PATH`, `_PS_PATH`) require
  `fs.existsSync` before being returned, so typos / stale config fall
  through to the standard resolution chain.

Misc:
- `gitnexus/package.json` engines.node back to `>=22.0.0` (matches
  origin/main and the original PR reviewer's earlier request).

Twin-tree parity / CI sync mechanism tracked separately at
abhigyanpatwari/GitNexus#1591.

Test plan: vitest run test/unit/hooks.test.ts → 113 passed,
18 Unix-only skipped; setup.test.ts → 14 passed.

* chore(autofix): apply prettier + eslint fixes via /autofix command

* trigger

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-14 16:39:30 +01:00
911a2ee1e6 fix: apply ESM .js extension fallback to tsconfig path alias resolution (#1530)
* fix: apply ESM .js extension fallback to tsconfig path alias resolution

Path alias imports (e.g. `@/utils.js` via tsconfig paths) now correctly
strip JS-family extensions and retry with TS equivalents when the literal
.js file does not exist. This applies the same stripJsExtension fallback
already used for relative imports to the alias resolution branch.

Fixes #1528

* chore(autofix): apply prettier + eslint fixes via /autofix command

* test(esm): cover .mjs/.cjs path-alias extension resolution

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(esm): use Map for path aliases in resolveWithAlias helper

Matches TsconfigPaths.aliases from language-config. CI cannot run tsc -p tsconfig.test.json yet: the project has hundreds of pre-existing errors under test/ (fixtures + unit/integration); enable that step after backlog cleanup.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-14 16:15:17 +01:00
CopilotandGergő Magyar 586dbf7aa1 feat(cpp): disambiguate template specializations in class graph IDs and receiver routing (#1587)
* Initial plan

* fix(cpp): disambiguate template specializations in class graph IDs

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c929edb2-f2e9-41c6-a2e9-2092b967f603

* fix(cpp): guard template-specialization class lookup fallback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c929edb2-f2e9-41c6-a2e9-2092b967f603

* fix(cpp): address github-actions inline review findings

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/68d8fbac-4ff4-47f7-b732-eaf2c2f94043

* fix(cpp): cover template-type receiver binding for specialization routing

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9505dfcd-3fb6-4bc2-a134-f60fe0dc8cd9

* chore(cpp): clarify specialization-binding fallback assumptions

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9505dfcd-3fb6-4bc2-a134-f60fe0dc8cd9

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-14 15:00:29 +01:00
c901ee4666 fix(cpp): workspace-wide dependent-base name resolution for cross-file templates (#1586)
* Initial plan

* fix(cpp): workspace-wide dependent-base name resolution (cross-file support)

- Replace per-file `populateCppDependentBases(parsed)` with a workspace-wide
  `populateCppDependentBases(parsedFiles)` that builds a cross-file class index
- Use qualified-name prefix for namespace disambiguation when multiple classes
  share a simple name (e.g. `Box` in two namespaces)
- Move the call from `populateOwners` (per-file) to the new `populateWorkspaceOwners`
  hook so all files are processed before resolution runs
- Add `cpp-two-phase-dependent-base-ns` fixture: Base<T> in a namespace in a
  separate file from Derived<T>, plus a namespace-free function with the same
  name — exercises the path where the class-owned filter does not apply
- Add two integration tests for the new fixture"

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/d78fae8b-cd32-45d8-a815-2b27d7d89e62

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(cpp): clarify V1 conservative exact-prefix namespace match in two-phase-lookup

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/d78fae8b-cd32-45d8-a815-2b27d7d89e62

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-05-14 13:06:39 +01:00
Copilotandmagyargergo 75cb49477e feat(cpp): emit EXTENDS edges for template and qualified template bases (#1581)
* Initial plan

* fix: emit cpp extends edges for template bases

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/eaddb1ac-7b57-4f44-94ba-a07a578d078d

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore: address final review notes

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/eaddb1ac-7b57-4f44-94ba-a07a578d078d

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: keep cpp extends edges class-owned

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b10bbb4d-6746-46fa-9b82-5c0962cd8b3f

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: address cpp follow-up review findings

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/be67e437-055f-4a71-a24e-d3bfb87ad0cd

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-05-14 12:25:05 +01:00
e01f0912bc feat(cpp): migrate C++ to scope-based resolution model (#938) (#1520)
* fix(cpp): complete scope-resolution parity

* fix(ci): resolve formatting, lint errors for PR #1520

- prettier: format arity-metadata.ts, captures.ts, index.ts
- eslint: rename unused HEADER_GLOB to _HEADER_GLOB
- eslint: replace unsafe parser.parse() with parseSourceSafe()
- eslint: suppress intentional console.warn/log in sync.ts
- eslint: remove unused _it import alias in cpp.test.ts

* fix(ci): complete formatting, lint, and typecheck fixes

- prettier: format call-processor.ts, imported-return-types.ts,
  include-extractor.test.ts, cpp-captures.test.ts, cpp-imports.test.ts
- eslint: suppress intentional console.warn in manifest-extractor.ts
- typecheck: restore 'thrift' in ContractType union (was accidentally
  removed) and add thrift case to exhaustive switch in manifest-extractor

* fix(ci): revert unintended group module changes that broke tests

Restore types.ts, config-parser.ts, matching.ts, sync.ts, and
manifest-extractor.ts to upstream/main versions. The original commit
accidentally removed fields (thrift, workspace_deps, exclude_links_paths,
exclude_links_param_only_paths) from DetectConfig/MatchingConfig/ContractType
which are still referenced by matching.test.ts, config-parser.test.ts,
sync.test.ts and other integration tests.

This PR's scope is C++ scope-resolution parity only — group module
type definitions and logic should remain unchanged.

* fix(codeql): address security and quality alerts

- arity-metadata.ts, interpret.ts: replace single-pass template strip
  regex (/<[^>]*>/g) with a while-loop to fully handle nested templates
  like Map<List<int>> — resolves 'Incomplete multi-character sanitization'
- cpp.test.ts: remove unused vitest 'it' import since the file defines
  its own 'it' via createResolverParityIt — resolves 'Assignment to constant'
- include-extractor.test.ts: use fs.mkdtempSync() instead of predictable
  os.tmpdir()+Date.now() paths — resolves 'Insecure temporary file'
- interpret.ts: remove redundant 'name !== undefined' check (already
  guaranteed by early return) — resolves 'Comparison between inconvertible types'

* review: address Claude review findings on PR #1520

- Findings 1-3 (BLOCKERS): restore include-extractor.ts and its test to
  the main baseline. Block-comment fallback regression, suffix-resolve
  false-positive suppression, and the four deleted regression tests
  (#3-#6) are now back. These changes were unrelated to C++ scope
  parity and should not have been in this PR.

- Finding 4 (MAJOR, partial): revert COMPOUND_RECEIVER_MAX_DEPTH 6 to
  4. No C++ test exercises depth > 4 (cpp-chain-call uses a 2-hop
  chain), so the bump risked silent regressions on other migrated
  languages without justification. The wildcard-origin propagation in
  imported-return-types.ts is retained — C++ #include and using
  namespace both emit wildcard-origin bindings (cpp/import-decomposer
  .ts:40,90), so wildcard propagation is causal to C++ parity.

- Finding 6: tighten write-access dedup test with exact per-field
  counts (nameWrites = 2, addrWrites = 1) instead of total-count + sub
  string containment, so a regression in one of the two name writes
  can no longer be masked.

- Finding 8: skipped. Box-drawing characters in cpp/query.ts comments
  match the established convention used in csharp/java/php query
  files.

Finding 5 (int/long normalization tie-breaker) left as documented
follow-up — proper fix requires resolver-level tie-breaker logic and
risks regressing other arity-matching tests.

* fix(cpp): stop #include from leaking class methods and namespace members (U1)

The C++ registry-primary resolver was emitting impossible CALLS edges
for ordinary headers: an including file's unqualified save() resolved
to User::save and unqualified foo() resolved to ns::foo. Two leak
paths converged on localDefs:

1. expandCppWildcardNames (file-local-linkage.ts) iterated the
   flattened localDefs and exported every simple tail, including
   class-owned methods and namespace-contained symbols. Replaced with
   a scope-aware filter: build nodeId -> owning Scope from
   Scope.ownedDefs and skip defs whose owning scope is Namespace or
   Class.

2. The shared global free-call fallback's pickUniqueGlobalCallable
   walks the workspace registry by simple name and would still hit
   class methods / namespace members even with wildcard expansion
   fixed. Plugged the gap via the existing isFileLocalDef hook —
   semantically 'logically invisible cross-file' — by tracking per-
   file non-globally-visible nodeIds (populateCppNonGloballyVisible,
   called from populateOwners) and adding an ownerId !== undefined
   fast-path for class-owned defs.

Side fix in shared finalize-algorithm.ts: when wildcard expansion
resolves to a real target but produces zero propagating names, the
edge was dropped, taking the file-level IMPORTS edge with it.
Preserve the original wildcard edge so #include dependencies survive
even when the header exposes no unqualified bindings.

Tests: cpp-include-no-class-leak, cpp-include-no-namespace-leak, and
cpp-anon-ns-same-file-visible fixtures. Negative tests mode-gated to
REGISTRY_PRIMARY_CPP=1 via the expected-failures registry — legacy
DAG has no scope-aware filtering on the global fallback; backporting
is out of scope. All 2104 resolver integration tests pass under
registry-primary mode.

* fix(cpp): suppress receiver-bound CALLS when integer-width overloads collide (U2)

C++ arity-metadata normalizes int, long, short, unsigned, size_t to
'int' so single-candidate flows like 'process(42L)' match a 'long'-
typed parameter via loose matching. But when both 'process(int)' and
'process(long)' coexist as method overloads, they both end up with
parameterTypes=['int'] in the registry, and pickOverload's narrowing
returns 2 candidates with no way to disambiguate. The previous code
picked candidates[0] arbitrarily, emitting a CALLS edge to the wrong
overload roughly half the time.

Fix:
- Add isOverloadAmbiguousAfterNormalization in overload-narrowing.ts
  that detects >1 candidate sharing identical parameterTypes sequences.
- Have pickOverload return a new OVERLOAD_AMBIGUOUS sentinel when this
  fires.
- In the receiver-bound-calls loop, when pickOverload signals ambiguity,
  suppress the edge AND add the site to handledSites so the late-stage
  emitReferencesViaLookup pass does not re-emit the pre-resolved
  reference. Without the handled-mark, the reference index still
  carries a toDef and emits the same wrong edge.

Graph schema has no ambiguous-target edge model, so emitting two
edges (one per candidate) would require a separate schema change.
Zero-edge is the only safe outcome.

Other languages: the ambiguity check is a precondition gate, not a
behavior change for normal narrowing. Languages whose normalizers do
not collapse distinct types into a single token (verified by grep
over *-arity-metadata.ts) will never produce >1 candidate with
identical parameterTypes from genuinely distinct declarations, so
the branch is effectively C++-only in practice.

Test: cpp-overload-int-long fixture asserts exactly .toBe(0) CALLS
edges. Count=1 = arbitrary pick (the bug); count>1 = unsupported
ambiguous-edge model. Mode-gated to REGISTRY_PRIMARY_CPP=1 — legacy
DAG has no OVERLOAD_AMBIGUOUS wiring; backporting is out of scope.

All 2105 resolver integration tests pass under registry-primary; all
139 cpp tests pass under both modes (3 negative tests skipped in
legacy as documented).

* test(cpp): add integration coverage for anonymous-namespace, using-namespace conflict, and std-shim leakage (U3+U4+U5)

Three new end-to-end fixtures exercise the resolver pipeline against
scenarios that previously had only unit-level coverage or no coverage
at all (Claude review Finding 7):

U3 — cpp-anon-ns-cross-file:
  helper.cpp declares 'namespace { void worker(); }' and calls it
  internally. caller.cpp declares a separate 'void worker()' and calls
  it. Asserts (a) the cross-file CALLS edge from caller's run() does
  not target helper.cpp's anonymous-namespace worker, and (b) the
  same-file edge from helper_entry() to its own worker still resolves
  (positive guard against a 'no edges at all' regression making the
  negative check vacuously pass). Includes a state-isolation guard
  that re-runs the same fixture and asserts identical results,
  proving clearFileLocalNames() is called by the pipeline entry.

U4 — cpp-using-namespace-conflict:
  Two headers each declaring 'namespace a { foo() }' and
  'namespace b { foo() }' respectively, plus a caller doing
  'using namespace a; using namespace b; foo()'. Asserts exactly
  zero CALLS edges. One edge = arbitrary pick (the bug); two edges
  would require an ambiguous-target edge model GitNexus does not
  have. Depends on U1 — without scope-aware filtering, both foo()s
  would already be in the importer's wildcard binding set as simple
  'foo', so the test would pass for the wrong reason.

U5 — cpp-using-namespace-std-smoke:
  Fixture-local 'namespace std { void cout_write(); void println(); }'
  shim rather than real <iostream> — captures the wildcard-leak
  shape deterministically without depending on system-header modeling
  stability (out of scope per plan). Asserts (a) the project-local
  call resolves correctly, (b) no leak to shim STL symbols, and (c)
  no CALLS/ACCESSES edges from the caller into std-shim.h at all.

Negative tests for U2/U4 mode-gated to REGISTRY_PRIMARY_CPP=1 via
the expected-failures registry; legacy DAG lacks the OVERLOAD_AMBIGUOUS
suppression and the namespace-aware filtering, so the leaks persist
there. All 2112 resolver integration tests pass under registry-primary;
all 146 cpp tests pass under both modes (4 negative tests skipped in
legacy as documented).

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(cpp): scope-aware isSuperReceiver classification (U1)

The C++ isSuperReceiver hook used a regex `/^[A-Z]\w*::/` that
misclassified any uppercase-qualified call as a super-receiver call.
Singleton::getInstance(), std::Foo::bar(), and PascalCase namespace
calls all entered the super branch, where the absence of an enclosing
class (or wrong MRO context) dropped the resolution entirely.

Fix:
- New optional ScopeResolver hook isSuperReceiverInContext(text,
  callerScope, scopes). Languages where super classification depends
  on caller context define it; receiver-bound-calls.ts prefers it
  when defined and falls back to the simple isSuperReceiver(text)
  otherwise. Other migrated languages (Python, Java, C#, PHP, Go,
  TypeScript) are unchanged.
- C++ implementation: parse the LHS of '::' from the receiver text,
  resolve via findClassBindingInScope, and return true only when
  the LHS is a class-like def in the caller's enclosing class's MRO.
  Returns false for namespace LHS, unresolved LHS, self-class LHS
  (qualified self-calls aren't super), and any non-'::' form.
- Extended the C++ tree-sitter query to capture the LHS of
  qualified_identifier as @reference.receiver so qualified static
  member calls (Singleton::getInstance()) reach the receiver-bound
  Case 2 (class-name receiver) path. Without the receiver capture,
  qualified calls had no explicit receiver and could not resolve
  through any receiver-bound branch.

Test: cpp-namespace-qualified-not-super fixture. Singleton::getInstance()
from a free function asserts exactly 1 CALLS edge through the
qualified-call path. Passes under both REGISTRY_PRIMARY_CPP=1 and =0.

All 2113 resolver integration tests pass; all 147 cpp tests pass under
both modes.

* fix(cpp): suppress receiver-bound CALLS when default-arg overloads collide (U4)

ISO C++ rejects 's.f(1)' as ambiguous when both 'void f(int)' and
'void f(int, int = 0)' are declared on S. The previous resolver
returned the first viable candidate via pickOverload's fallback.

Extended isOverloadAmbiguousAfterNormalization to take an optional
argCount: when provided, the predicate compares only the first
argCount slots of each candidate's parameterTypes. Candidates whose
declared-prefix matches up to argCount are treated as ambiguous
because default arguments make all of them equally viable for the
call.

Without argCount, behavior is unchanged (the original int/long
normalization-collapse contract, full-length equality required).
pickOverload now passes site.arity so default-arg ambiguity fires.

Test: cpp-overload-default-arg-ambiguous fixture. s.f(1) where S has
f(int) and f(int, int = 0) asserts exactly .toBe(0) CALLS edges.
Passes under both REGISTRY_PRIMARY_CPP=1 and =0.

All 2114 resolver integration tests pass; all 148 cpp tests pass
under both modes.

* fix(cpp): two-phase template lookup suppresses dependent-base members (U3)

ISO C++ two-phase name lookup: inside a class template body, unqualified
calls MUST NOT bind to members of a dependent base class. Only this->name
or Base<T>::name forms make the lookup dependent. GCC and Clang both
reject the unqualified form with 'declaration of f must be available'.

Before this fix, GitNexus's global free-call fallback walked the
workspace registry by simple name and bound unqualified calls inside
template bodies to dependent-base members, producing CALLS edges the
compiler would reject.

Implementation:
- New languages/cpp/two-phase-lookup.ts module: per-pipeline state
  recording (className, dependentBaseName) pairs at capture time and
  resolving them to nodeId sets during populateOwners.
- captures.ts detectCppDependentBases walks the AST once finding every
  template_declaration containing a class/struct definition. For each,
  it collects template-parameter names (typename T, class T, non-type
  int N, template-template parameters) and walks each base in the
  base_class_clause checking whether any inner type_identifier matches
  a template parameter. Conservative bias: typename T::U, decltype,
  and template-template-parameter shapes also classified as dependent.
- Extended scope-resolution contract's isCallableVisibleFromCaller
  hook with optional callerScope and scopes fields. C++ implements
  the hook to consult isCppDependentBaseMember: when the candidate
  is a member of a dependent base of the caller's enclosing class,
  the hook returns false and pickUniqueGlobalCallable skips the
  candidate.
- clearFileLocalNames also clears the dependent-base state per
  pipeline run.

Fixtures:
- cpp-two-phase-dependent-base: Derived<T> deriving from Base<T>,
  unqualified f() and i inside Derived's body. Asserts zero CALLS
  edges and zero ACCESSES edges respectively.
- cpp-two-phase-this-qualified, cpp-two-phase-non-dependent-base,
  cpp-two-phase-namespace-free-call-inside-template: positive
  fixtures left as documented gaps (this-> and qualified-name
  resolution inside template bodies are pre-existing resolver
  weaknesses independent of U3). Tracked separately.

Negative test mode-gated to REGISTRY_PRIMARY_CPP=1 via the expected-
failures registry; legacy DAG has no two-phase lookup.

All 2116 resolver integration tests pass under registry-primary; all
150 cpp tests pass under both modes (5 negative tests skipped in legacy
as documented).

* fix(cpp): implement V1 ADL (Koenig lookup) for free-function calls (U2)

Plan 2026-05-13-001 U2. Adds argument-dependent lookup as a new
candidate-generating tier in `emitFreeCallFallback`: when ordinary
unqualified lookup is empty, ADL surfaces candidates from each
value-class-typed argument's enclosing namespace.

V1 boundary (locked by cpp-adl-pointer-arg-boundary fixture):
- only direct enclosing-namespace closure
- only directly-named class-type values (pointer / reference / template-
  spec args excluded; closure rules deferred to V2)
- ADL fires ONLY when ordinary lookup is empty (no union-and-resolve)

Parenthesized name `(f)(s)` suppresses ADL per ISO C++
[basic.lookup.argdep]/3.1. Multi-candidate ambiguity (e.g. `process(int)`
vs `process(long)` after C++ int-width normalization) returns the
ADL_AMBIGUOUS sentinel — caller suppresses entirely, mirroring the
OVERLOAD_AMBIGUOUS contract from plan 2026-05-12-002 U2.

Implementation:
- `cpp/adl.ts` — new module: per-pipeline argInfoBySite + noAdlSites Maps
  populated at capture time, classToNamespaceQualifiedName Map populated
  during populateOwners; `pickCppAdlCandidates` returns
  SymbolDefinition | ADL_AMBIGUOUS | undefined
- `scope-resolution/contract/scope-resolver.ts` — adds optional
  `resolveAdlCandidates` hook
- `scope-resolution/passes/free-call-fallback.ts` — invokes ADL hook
  between `findCallableBindingInScope` and `pickUniqueGlobalCallable`;
  marks site handled on `'ambiguous'` so emit-references doesn't retry
- `cpp/captures.ts` — detects `parenthesized_expression` function wrap;
  per-arg classification (pointer/reference/value class) preserving the
  shape info the existing arity-narrowing normalizer strips
- `cpp/scope-resolver.ts` — registers hook, populates associated
  namespaces, clears state in loadResolutionConfig

Negative tests (parens, pointer-boundary, ambiguous) gated under
LEGACY_RESOLVER_PARITY_EXPECTED_FAILURES.cpp — legacy DAG has no V1/V2
ADL boundary or ADL_AMBIGUOUS suppression.

154/154 cpp integration tests pass under REGISTRY_PRIMARY_CPP=1;
147 pass + 7 skipped under =0 (legacy parity baseline).

* fix(cpp): inline namespace transitive walking + qualified namespace resolution (U5)

Plan 2026-05-13-001 U5. Two ISO C++ inline-namespace semantics:

1. Unqualified-lookup transitive visibility: inline-namespace members
   reach the enclosing namespace's scope as if declared there. The
   `populateCppNonGloballyVisible` exemption keeps them globally visible
   so cross-file unqualified lookup finds them.

2. Qualified-receiver transitive visibility: `outer::foo()` resolves to
   `outer::v1::foo()` when `v1` is inline (and through arbitrarily-deep
   nesting like `outer::v1::experimental::foo`, matching libc++ `__1` /
   libstdc++ `__cxx11`).

The second behavior required a new resolver case in
`receiver-bound-calls.ts` (Case 1.5: language-specific qualified-receiver
member lookup) because C++ qualified-namespace member calls had no prior
resolution path — receiver-bound Case 1 only handled
`ParsedImport.kind === 'namespace'` (Python/JS-style) and Case 2 handles
class receivers, neither of which fired for `outer::foo()`. The new
hook `resolveQualifiedReceiverMember` is opt-in; languages without
C++-style qualified-name semantics omit it.

Implementation:
- `cpp/inline-namespaces.ts` — new module: per-pipeline
  `inlineNamespaceRangesByFile` + `inlineNamespaceScopeIds` Sets;
  `markCppInlineNamespaceRange` at capture time;
  `populateCppInlineNamespaceScopes` resolves ranges → scope IDs;
  `resolveCppQualifiedNamespaceMember` walks namespace scopes by simple
  name and descends transitively through inline children only.
- `scope-resolution/contract/scope-resolver.ts` — adds optional
  `resolveQualifiedReceiverMember` hook to the contract.
- `scope-resolution/passes/receiver-bound-calls.ts` — Case 1.5 invokes
  the hook between Case 1 (namespace imports) and Case 2 (class-name
  receiver). Returns undefined for non-namespace receivers so Case 2
  still resolves class-qualified calls.
- `cpp/captures.ts` — detects `inline` keyword child on
  `namespace_definition`; records 1-based range to match Scope.range.
- `cpp/file-local-linkage.ts` — `populateCppNonGloballyVisible` exempts
  inline-namespace scopes so cross-file unqualified lookup keeps their
  members visible.
- `cpp/scope-resolver.ts` — wires `populateCppInlineNamespaceScopes`
  into populateOwners (BEFORE `populateCppNonGloballyVisible` so the
  exemption sees populated state); registers
  `resolveQualifiedReceiverMember` hook.

4 fixtures: `cpp-inline-namespace-unqualified`, `-versioned`,
`-nested` (two transitive inline hops, STL `__1` shape), and
`-adl-participation` (composes with U2 — ADL surfaces records declared
inside inline child namespaces). All 4 assert exactly 1 CALLS edge with
correct target file.

Versioned fixture gated under LEGACY_RESOLVER_PARITY_EXPECTED_FAILURES.cpp
— legacy DAG can't disambiguate two same-name foos without inline
awareness. Other 3 coincidentally resolve in legacy.

158/158 cpp integration tests pass under REGISTRY_PRIMARY_CPP=1;
150 pass + 8 skipped under =0 (legacy parity baseline).

* test(cpp): Phase 5 cross-unit composition tests for U1/U2/U3/U5

Plan 2026-05-13-001 Phase 5. Locks in correct behavior at the
intersections between the previously-shipped scope-resolver units.

Enhancement to U1: `isSuperReceiverInContext` strips template-argument
lists (`Base<T>` → `Base`) and namespace prefixes (`outer::v1::Base` →
`Base`) before resolving the receiver in the caller's scope chain. This
makes the super-receiver classification work for template-class
heritage shapes like `Base<T>::method()` and `outer::v1::Base<T>::f()`.

Three fixtures + four tests:

- `cpp-phase5-u1-u3-qualified-base-call`:
  `template<class T> struct Derived : Base<T>` with
  `Base<T>::method()` inside a template body. Asserts NO mis-routing
  (count = 0) — documents the V1 gap that template-class inheritance
  isn't captured as EXTENDS by the legacy DAG, so MRO walks are empty
  and the super branch can't dispatch. The composition still works
  correctly: U1's template-arg-stripping classifies `Base<T>` as a
  super candidate, but the empty-MRO terminates without false edges.

- `cpp-phase5-u2-u3-adl-from-derived`:
  `Derived : Base<T>` where `Base::record` shadows `audit::record`.
  Unqualified `record(e)` inside the template body should resolve via
  ADL to `audit::record` (because U3 + the `isFileLocalDef` class-
  owned filter suppress `Base::record`). Asserts 1 edge to audit.h
  and 0 edges to base.h.

- `cpp-phase5-u3-u5-inline-base`:
  `template<class T> struct Derived : outer::v1::Base<T>` where `v1`
  is inline. Unqualified `f()` inside `Derived<T>::g()` should NOT
  bind to Base::f (dependent-base suppression even across inline
  namespace prefix). Asserts count = 0.

Phase 5 tests asserting no-false-positives are gated under
LEGACY_RESOLVER_PARITY_EXPECTED_FAILURES.cpp — legacy DAG over-
resolves without the template-arg-stripping qualified-receiver path
and without two-phase dependent-base suppression.

162/162 cpp integration tests pass under REGISTRY_PRIMARY_CPP=1;
152 pass + 10 skipped under =0 (legacy parity baseline).

---------

Co-authored-by: HuangWenjie <zhoudeng.hwj@alibaba-inc.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-14 09:30:52 +01:00
e9349ce66a fix(markdown): handle CRLF line endings in section heading parser (#1469)
* fix(markdown): handle CRLF line endings in section heading parser

split('\n') on CRLF content leaves a trailing \r on each line, and the
heading regex /^(#{1,6})\s+(.+)$/ (anchored with $) fails to match
'## Heading\r' because $ matches before end-of-string, not before \r.
Result: Windows-authored markdown silently produces zero Section nodes.

Use split(/\r\n|\r|\n/) to normalize all line-ending conventions.

Pure additive — LF-only files produce identical output. CR-only (Mac OS
Classic) becomes tolerated as a side benefit at zero risk.

Adds integration test markdown-processor-crlf.test.ts covering LF
baseline, CRLF (the regression), CR-only, mixed, and startLine/endLine
correctness.

* test(markdown): strengthen CRLF integration tests + clarify split comment

- Assert section names, levels, line spans, and CONTAINS hierarchy (not only counts)
- Document trailing-newline effect on endLine via exact toEqual expectations
- Reword markdown-processor comment: \$ only at end-of-string vs .+ before \\r

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore: empty commit

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-14 08:58:58 +01:00
a3eef48ce3 fix(cli): make --no-stats actually omit volatile counts (#1477) (#1478)
* fix(cli): make --no-stats actually omit volatile counts (#1477)

Closes #1477.

The `--no-stats` flag on `gitnexus analyze` was advertised as
"Omit volatile file/symbol counts from AGENTS.md and CLAUDE.md"
but had no effect: every reindex still rewrote the markdown with
fresh count phrases, producing chore-commit churn on every run —
the exact problem the flag was added to solve in #704.

Root cause is commander.js negation-flag semantics. `.option(
'--no-stats', ...)` registers the option under the accessor
`stats` (boolean, default `true`; `false` when the flag is passed),
NOT `noStats`. The two action-handler reads in `analyze.ts`
(lines 414 and 500 pre-fix) read `options?.noStats`, which is
always `undefined`, so the `noStats` payload always reached
`runFullAnalysis` / `generateAIContextFiles` as `undefined`/falsy
and the count branch in the template always fired.

Fixed by replacing `options?.noStats` with `options?.stats === false`
at both reads. The strict `=== false` check (rather than
`!options?.stats`) means absent options or absent `.stats` field
fall through as no-stats=false, preserving the documented default-on
behaviour. Also updated the `AnalyzeOptions` interface to declare
`stats?: boolean` (matching commander's actual output) with a
JSDoc explaining the negation, since the prior `noStats?: boolean`
shape was a static-type misrepresentation of what commander
provides at runtime.

Internal call sites that re-pack `{ noStats: ... }` for
downstream consumers (`run-analyze.ts`, `ai-context.ts`) keep
their existing field name — those interfaces are not commander-
shaped, so `noStats` is the correct name there.

## Regression tests

Two new unit tests in `test/unit/ai-context.test.ts`:

* `omits volatile counts when noStats option is set (#1477)` —
  asserts the count parenthetical is absent from both CLAUDE.md
  and AGENTS.md when `noStats: true` is passed.
* `preserves volatile counts when noStats is not set (default)` —
  documents the default-on path so a future refactor can't
  silently flip the default.

Both call `generateAIContextFiles` directly with distinctive numbers
that would unmistakably leak through if the omit branch is broken.

## Manual verification

* `vitest run test/unit/ai-context.test.ts` → 13/13 pass
  (11 prior + 2 new).
* Verified before-fix behaviour by checking out main, running
  `npx gitnexus analyze --no-stats` against an indexed repo, and
  observing the count phrase still present. Re-running on the fix
  branch with the same flag strips the phrase as documented.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(cli): resolve merge conflict markers in analyze.ts (PR #1478)

Remove leftover conflict hunks from main merge; keep commander stats
shape (stats?: boolean), wire noStats: options?.stats === false into
runFullAnalysis and generateAIContextFiles, and retain indexOnly /
skipSkills / skipAgentsMd wiring from main.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(cli): cover analyzeCommand → runFullAnalysis noStats bridge (#1477)

Assert commander-shaped options.stats maps to the internal noStats
payload (including explicit true/false and skipAgentsMd combination)
so the CLI bridge cannot regress without failing tests.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(cli): cover AGENTS.md default stats + skills noStats bridge (#1478)

- Assert volatile stats phrase in both CLAUDE.md and AGENTS.md when noStats is omitted
- Add bridge test for --skills regeneration path with stats:false → generateAIContextFiles noStats
- Note shared noStats expression beside skills-path call; stub process.exit for full analyze path

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-14 08:26:27 +01:00
6229417bd5 feat: gitnexus:keep marker preserves custom context sections (resubmit of #605) (#1508)
* feat: gitnexus:keep marker preserves custom context sections

When <!-- gitnexus:keep --> is present inside the gitnexus block,
analyze only updates the stats line instead of replacing the entire
section with the verbose template. Lets users maintain lean custom
context without it being overwritten on every reindex.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat: improve gitnexus:keep marker to reliably preserve custom sections

The `<!-- gitnexus:keep -->` marker inside a GitNexus block tells
`analyze` to only update the stats line (node/edge/flow counts)
while preserving the user's custom layout. This lets teams trim
the verbose default template to a lean format without having it
overwritten on every reindex.

Changes:
- Broaden stats-line regex to match both "Indexed as" and
  "indexed by GitNexus as" formats
- Improve stats extraction from generated content (prefer
  structured match over greedy parentheses)
- If keep marker is present but no stats line found, preserve
  the section as-is instead of falling through to full replace
- Add tests for keep preservation and no-keep replacement

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address PR #1508 review findings (F1-F5)

Refactor the keep-marker stats-update path and close the test-coverage
gaps surfaced by the production-readiness review.

## Findings 2 + 3 (high) — fragile extraction → silent corruption

Stop re-extracting `newName` (first `**bold**`) and `newStats` (first
`(...)`, with fallback) from generated content. Both are structurally
fragile:

- F2: newName silently picks the wrong value if the template ever
  emits bold text before the project-name line (no current bug; an
  unstated contract with no enforcement)
- F3: newStats fallback `\(([^)]+)\)` matches `({target: "symbolName",
  direction: "upstream"})` from the Always-Do bullet when
  `noStats: true` suppresses the canonical stats line, silently
  corrupting the stats output

Fix: pass `projectName: string` and `stats: RepoStats` directly into
`upsertGitNexusSection`. Build the stats line from those values. Both
callers in `generateAIContextFiles` already have them in scope.

## Finding 1 (high) — misleading return value

When a keep marker is present but no stats line matches the pattern,
the function previously returned `'updated'` without writing,
producing `CLAUDE.md (updated)` in CLI output for a file that was
not touched. Add a distinct `'preserved'` return variant; CLI now
reports `CLAUDE.md (preserved)` honestly.

## Finding 4 (medium) — unanchored stats regex

`/(?:Indexed as|...) \*\*[^*]+\*\* \([^)]+\)/` could match prose
embedded mid-paragraph in user content (e.g. "you'll see it Indexed
as **Foo** (note: ...)"). Anchor with `^...$` plus the `m` flag so
only standalone stats lines match.

## Finding 5 — test coverage gaps

Seven new tests, each cross-referenced to the review finding:

- keep marker OUTSIDE the GitNexus section has no effect
- AGENTS.md keep path preserves custom layout (parity with CLAUDE.md)
- idempotent: second run produces byte-identical output
- CRLF file with keep marker: stats line updates correctly
- noStats + keep marker: not corrupted by Always-Do tuple text (F3 regression guard)
- returns 'preserved' (not 'updated') when no stats line matches (F1 regression guard)
- project name with markdown punctuation (hyphens/slash/dot) lands intact

All 23 ai-context tests pass; typecheck, prettier, eslint clean.

* docs(ai-context): address PR #1508 review findings on keep-marker path

- Clarify that noStats affects generated template only, not keep-section stats updates
- Fix stats-line regex comment to match behavior (no end anchor; trailing suffix kept)
- Assert '. MCP tools.' survives stats replacement in preserve-custom-section test
- Document LF normalization when rewriting CRLF seed in keep-marker CRLF test

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: dp-web4 <dp@web4.ai>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-14 07:40:15 +01:00
dependabot[bot]andGergő Magyar 0566c98b54 chore(deps)(deps): bump @langchain/google-genai in /gitnexus-web (#1554)
Bumps [@langchain/google-genai](https://github.com/langchain-ai/langchainjs) from 2.1.28 to 2.1.30.
- [Release notes](https://github.com/langchain-ai/langchainjs/releases)
- [Commits](https://github.com/langchain-ai/langchainjs/commits)

---
updated-dependencies:
- dependency-name: "@langchain/google-genai"
  dependency-version: 2.1.30
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-14 07:12:36 +01:00
dependabot[bot] 80acaf052f chore(deps): bump sigstore/cosign-installer from 4.1.1 to 4.1.2 (#1557) 2026-05-14 06:45:14 +01:00
dependabot[bot] afa38432a4 chore(deps)(deps-dev): bump vite from 8.0.10 to 8.0.11 in /gitnexus-web (#1555) 2026-05-13 22:17:53 +01:00
Shane Thurston WijayaandGergő Magyar 88d3df77cc feat:(wiki) added --timeout and --retries flags for large module pages to mitigate timeout aborts (#1543)
* feat:(wiki) added --timeout and --retries flags for large module pages to mitigate timeout aborts

* docs(wiki): document --timeout and --retries options

* docs(wiki): document --timeout and --retries in SKILL.md

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-13 18:35:39 +01:00
Léon Simmons 507f84b69a fix(docker): symlink gitnexus binary onto $PATH in runtime image (#1551)
The README documents the Docker workflow as:

    WORKSPACE_DIR=$HOME/code docker compose up -d
    docker compose exec gitnexus-server gitnexus index /workspace/my-repo

…but `gitnexus` is not on $PATH inside the published image:

    $ docker compose exec gitnexus-server which gitnexus
    (empty)
    $ docker compose exec gitnexus-server gitnexus --version
    exec: "gitnexus": executable file not found in $PATH

The package.json `bin` entry (`"gitnexus": "dist/cli/index.js"`) would
normally surface via `node_modules/.bin/gitnexus`, but `npm prune
--omit=dev` in the builder stage strips that directory before the runtime
stage copies it in. The `dist/cli/index.js` itself already has the
`#!/usr/bin/env node` shebang and 755 permissions, so a single symlink
into /usr/local/bin makes the README's literal command work.

Verified locally:

    $ docker build -f Dockerfile.cli -t gitnexus:local-pr-test .
    $ docker run --rm gitnexus:local-pr-test gitnexus --version
    1.6.4
    $ docker run --rm gitnexus:local-pr-test gitnexus --help
    Usage: gitnexus [options] [command]
    …
    $ docker run --rm -d --name t gitnexus:local-pr-test \
      && sleep 4 && docker exec t curl -s localhost:4747/api/health
    {"status":"ok"}

CMD continues to invoke `node gitnexus/dist/cli/index.js serve …`
unchanged, so the change is additive and the server boot path is
untouched.

Refs #1549.
2026-05-13 17:14:52 +01:00
Hugo GuandGergő Magyar 38ff7365e8 fix(docker): install ca-certificates in runtime image for TLS verification (#1545) (#1547)
Close: #1545

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-13 14:45:37 +01:00
dependabot[bot]andGergő Magyar a9d72e2dbf chore(deps): bump urllib3 in /eval in the uv group across 1 directory (#1512)
Bumps the uv group with 1 update in the /eval directory: [urllib3](https://github.com/urllib3/urllib3).


Updates `urllib3` from 2.6.3 to 2.7.0
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.6.3...2.7.0)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.7.0
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-13 13:31:07 +01:00
GoGoLin 4cc4e9c84b fix(build): use platform-aware tsc command for win32 (#1531) 2026-05-13 13:02:58 +01:00
azizur100389andGergő Magyar 48cd55a120 fix(search): guard against undefined bm25Results when FTS unavailable (#1489) (#1540)
* fix(search): guard against undefined bm25Results when FTS unavailable (#1489)

When the FTS extension is unavailable in the MCP process,
searchFTSFromLbug can return an unexpected shape or throw,
leaving bm25Results undefined. The for-loop then crashes with
"bm25Results is not iterable".

- mergeWithRRF: default both inputs via ?? [] so undefined
  never reaches the iteration loops
- hybridSearch: wrap searchFTSFromLbug in try/catch and fall
  back to semantic-only search instead of crashing
- local-backend query handler: guard bm25SearchResult?.results
  and semanticResults with ?? []
- bm25Search: wrap the dynamic import in try/catch for
  sandboxed MCP contexts; guard ftsResponse?.results

Adds 6 regression tests covering undefined inputs and FTS
failure fallback.

Fixes #1489

* fix(search): address review findings on #1489 crash guards

- Guard ftsResponse.results with ?? [] in hybridSearch (Finding 1)
- Add logger.warn on bm25-index.js import failure (Finding 3)
- Add unit test for callTool query FTS throw path (Finding 2)

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-13 12:30:21 +01:00
azizur100389andGergő Magyar e8c8ddec8a fix(wiki): sanitize generated mermaid diagrams (#1539)
* fix(wiki): sanitize generated mermaid diagrams

* fix(wiki): address mermaid sanitizer review

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-13 11:53:09 +01:00
ec4624af87 fix(hooks): cap concurrent augment subprocesses (#1486) (#1510)
* fix(hooks): cap concurrent augment subprocesses to prevent runaway process spawn (#1486)

When Claude Code fires PreToolUse hooks for parallel Grep/Glob/Bash tool
calls, each invocation spawned its own `gitnexus augment` subprocess —
a Node + LadybugDB cold start that holds resources for several seconds.
Under heavy parallel search load (issue #1486: 180+ piled-up processes,
load avg > 100), these accumulated faster than they completed because
nothing capped concurrent in-flight augments.

Add a lockfile-based concurrency guard under `<.gitnexus>/.hook-locks/`:
each running hook claims a `<pid>.lock`, the guard counts live PIDs and
prunes stale entries (>30s mtime or pid no longer alive), and bails
silently when MAX_INFLIGHT (3) is reached. Augment is best-effort
enrichment — missing a few fires under burst load is preferable to
melting the system.

Applied to all three hook variants that spawn augment:
- gitnexus/hooks/claude/gitnexus-hook.cjs (npm-installed Claude hook)
- gitnexus-claude-plugin/hooks/gitnexus-hook.js (plugin Claude hook)
- gitnexus-cursor-integration/hooks/gitnexus-hook.cjs (Cursor hook)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(hooks): make augment concurrency cap a hard cap via atomic slot files

Address Claude's review of #1510. The original count-then-claim guard had
a TOCTOU window: N hooks could each read `active < MAX_INFLIGHT` between
readdirSync and the per-pid `wx` write and all proceed, briefly exceeding
the cap. The PR title's "cap" language overstated this.

Replace with fixed-name `slot-0.lock` ... `slot-N.lock` under `.hook-locks/`.
`O_CREAT|O_EXCL` on a fixed path is OS-atomic — exactly one process wins
each slot, so the cap is hard regardless of burst arrival timing. Each
slot file contains the owning PID so stale-takeover still works when a
hook crashes without releasing.

PID liveness is checked before age (Claude's Finding 3): a slow-but-alive
hook is never wrongly evicted. The 30s age window only kicks in to defend
against PID reuse on a long-abandoned slot, well above the 7s augment
timeout so a healthy run never hits it.

Also adds the missing concurrency-guard tests to cursor-hook.test.ts
(Claude's Finding 2): source-level wiring + dead-PID reclaim + 3-slots-full
bail. Previously only the CJS and Plugin variants had test coverage for
the guard; the Cursor variant was validated only by code inspection.

Tests: 5726 passing, +9 from baseline (1 hard-cap burst test + 4 source
regressions in hooks.test.ts; 3 source + 2 integration in cursor-hook.test.ts).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(hooks): inspect slot mtime + content via single fd (codeql TOCTOU)

CodeQL flagged the stale-takeover path in acquireHookSlot as a potential
filesystem race (js/file-system-race): statSync(slotPath) followed by
readFileSync(slotPath) gives a TOCTOU window where the file could be
swapped between the metadata check and the content read.

Replace the two separate path-based calls with a single openSync + fstatSync
+ readSync + closeSync sequence. Both mtime and owner PID now come from the
same file descriptor, so the operations are atomic on one inode. No
behavioral change beyond closing the race.

Applied to all three hook variants (CJS, Plugin, Cursor).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(hooks): distinguish EPERM from ESRCH in PID liveness check

Cursor Bugbot caught a contradiction with the stated design: the bare
`catch` after `process.kill(owner, 0)` was treating EPERM (process exists
but owned by another user) the same as ESRCH (process gone), which would
evict a live slot whenever the lock dir straddled user boundaries.

Inspect the error code: ESRCH → dead, evict; EPERM → still alive, keep
the slot; anything else → assume alive (be conservative under unexpected
failure rather than over-evict).

Applied to all three hook variants.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(hooks): fail closed when lock dir cannot be created

Previously the mkdirSync catch in acquireHookSlot returned `() => {}`
(a truthy no-op). The caller checks `if (!release) return;` to skip
augment when the guard can't be established — but a truthy no-op
slipped through that check and let augment spawn unguarded. On a
cross-user shared `.gitnexus/` or read-only filesystem, N concurrent
hooks would each take that branch and reintroduce the #1486 fan-out
the guard exists to prevent.

Return `null` instead so the caller's `if (!release) return;` skips
augment cleanly. Augment is best-effort enrichment — skipping it when
the guard fails is strictly safer than running unguarded.

Also clarify the stale-slot comment: PID-liveness wins for slots
younger than HOOK_LOCK_STALE_MS, but age is the final arbiter beyond
30s (PID-reuse defense). The previous wording said "PID-liveness wins
over age" without qualifying it, which contradicted the >30s branch.

Add source-level regression tests in hooks.test.ts and
cursor-hook.test.ts asserting acquireHookSlot returns null (not
() => {}) on lock-dir failure. Note in the Cursor test file that the
10-spawner burst test is not duplicated because the algorithm is
byte-for-byte identical to the CJS hook and already covered there.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(hooks): extract lock guard into helper modules

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/04dd20c5-28fd-433a-83cf-ad83fd03fb32

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
2026-05-13 08:56:27 +01:00
dependabot[bot] aed6cfc7ea chore(deps)(deps): bump mermaid (#1514) 2026-05-13 08:00:14 +01:00
dependabot[bot] 6a23616873 chore(deps)(deps): bump protobufjs from 7.5.5 to 7.5.8 in /gitnexus (#1536) 2026-05-13 06:40:24 +01:00
dependabot[bot] 7637bd1c83 chore(deps)(deps): bump @protobufjs/utf8 in /gitnexus (#1535) 2026-05-12 23:41:05 +01:00
Gergő Magyar 8083c39f6d feat(php): migrate PHP to scope-based resolution model (#938) [supersedes #1124] (#1497) 2026-05-12 16:56:31 +01:00
a2f1b07700 fix: resolve TypeScript ESM .js extension imports to .ts source files (#1525)
* fix: resolve TypeScript ESM .js extension imports to .ts source files

TypeScript ESM requires imports to use .js extensions even when source
files are .ts (moduleResolution: node16/bundler). The import resolver
now strips JS-family extensions (.js/.jsx/.mjs/.cjs) and retries with
TS equivalents (.ts/.tsx/.mts/.cts) when the literal .js file does not
exist. This fallback only applies to TypeScript/JavaScript languages.

Also adds .mts/.cts to the EXTENSIONS list for completeness.

Fixes #1503

* fix: address review findings — normalization, edge-case tests, integration test

- Fix makeCtx to use production normalization (.replace backslash)
  instead of .toLowerCase() (Finding 3)
- Add tests for .mjs/.cjs with competing .ts/.mts siblings (Finding 1)
- Add tests for ./dir.js → dir/index.ts boundary (Finding 2)
- Add integration test verifying full pipeline CALLS edges for ESM
  .js imports (Finding 4)
- Document path alias .js limitation as known follow-up (Finding 5)

* chore(autofix): apply prettier + eslint fixes via /autofix command

* chore: retrigger CI after bot-only tip commit

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-12 16:05:58 +01:00
evolutionandGergő Magyar 0daae93701 fix(lbug): drain checkpoint result before close (#1506)
* fix(lbug): drain checkpoint result before close

* test(lbug): cover checkpoint drain lifecycle

* fix(lbug): close query results after reads

* fix(lbug): close all stream query results

* fix(lbug): harden query result cleanup

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-12 14:03:45 +01:00
4fa40e9881 feat(analyze): incremental indexing (parse cache + DB writeback + scope-res short-circuit) (#1479)
* docs: incremental indexing design spec

Captures the design agreed in brainstorming on 2026-05-10:
- Transitive importer closure with public-surface-change optimization
- Git-only change detection (non-git repos: full rebuild as today)
- New default behavior; --force opts out
- New hydratePhase + loadGraphFromLbug primitive
- Iterative closure expansion with parseCache reuse
- incrementalInProgress dirty flag for crash recovery

Prior art: PR #592 (zenprocess), PR #533 (davidbeesley),
PR #1146 (azeemshaik025) — referenced and credited.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(communities): seed Leiden RNG for deterministic community detection

The vendored Leiden algorithm defaults to Math.random for tie-breaking
and randomized walks, which produces non-deterministic community
assignments and modularity values across runs on the same graph.

Pass a seeded mulberry32 RNG (LEIDEN_SEED=0xC0DE) so:
- The same graph always produces the same partition
- Modularity values are reproducible
- Equivalence tests for incremental indexing can compare community
  assignments byte-for-byte

This is foundational for the upcoming incremental-indexing feature
(see docs/superpowers/specs/2026-05-10-incremental-indexing-design.md)
where the correctness contract is incremental output ≡ full rebuild
output.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(incremental): change-detection, surface signatures, closure expansion

Three new modules supporting the incremental-indexing pipeline:

* core/incremental/git-diff.ts — getChangedFilesSinceCommit() unions
  'git diff lastCommit HEAD' (committed) with 'git status --porcelain'
  (dirty tree). Renames flattened to delete(orig) + add(new). Throws
  LastCommitMissingError when lastCommit is gone (caller falls back to
  full rebuild).

* core/incremental/surface.ts — extractSurfaceSignature() produces a
  stable hash of a file's publicly-visible symbols (functions, classes,
  methods, interfaces, types, heritage). Body-only edits → same hash.
  Signature/heritage changes → different hash. Drives the closure
  scoping optimization.

* core/incremental/closure.ts — computeImporterClosure() iterative
  fixpoint: parse each closure file, extract surface, query DB
  importers, expand. Uses a parseCache so each file is parsed once.
  Generic over TParseResult so closure logic is decoupled from the
  pipeline's parse representation.

32 unit tests across the three modules. Tests cover edge cases:
clean tree, dirty-only, mixed, renames, deletes, multi-hop cascade,
cycle termination, surface invariance, etc.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(lbug): loadGraphFromLbug, queryImporters, deleteAllCommunitiesAndProcesses

Three new primitives in lbug-adapter.ts to support incremental indexing:

* loadGraphFromLbug(graph, unchangedFilePaths) — streams all nodes for
  files in the set across every hydratable node table (excludes
  Community/Process — graph-wide, regenerated downstream). Then loads
  edges where both endpoints belong to loaded nodes, excluding
  MEMBER_OF / STEP_IN_PROCESS edges (also graph-wide).
  FilePaths chunked at 200 per query to keep statement size bounded
  on huge repos. Endpoint-level join filters by source-side filePath
  in the query, target-side checked JS-side via the loadedNodeIds set.

* queryImporters(targetFilePath) — returns DISTINCT a.filePath where
  a -[IMPORTS]-> b and b.filePath = target. Powers closure expansion:
  when a changed file's surface signature changes, all its importers
  must be re-parsed.

* deleteAllCommunitiesAndProcesses() — drops Community/Process nodes
  (and their edges via DETACH DELETE) at the start of each incremental
  run so the communities/processes phases regenerate them from the
  fully-merged graph. Required for the 'Leiden runs on full graph'
  correctness invariant.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(pipeline): hydrate phase + parse-filter for incremental indexing

Wires the incremental-indexing infrastructure into the phase-based
pipeline. Three coordinated changes:

* New hydratePhase (deps: structure) — loads node/edge state for files
  OUTSIDE ctx.options.filesToParse from the existing LadybugDB index.
  Runs before parse so the parse phase can produce a partial graph
  while downstream phases (mro, communities, processes) still see the
  full graph. No-op in full-rebuild mode (filesToParse unset).

* PipelineOptions.filesToParse: optional ReadonlySet<string>. When
  set, parse phase filters scanned files to this set; hydrate fills
  the complement. Set by runFullAnalysis when it detects an eligible
  incremental run; never set by callers directly.

* gitnexus-shared PipelinePhase enum: 'hydrate' added so progress
  callbacks can report the new phase distinctly from 'structure'.

Phase order: scan → structure → hydrate → markdown,cobol → parse
→ routes,tools,orm → crossFile → scopeResolution → mro → communities
→ processes. Communities (Leiden) still runs on the full graph,
satisfying the correctness invariant.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(analyze): incremental orchestrator branch + meta schema

Wires incremental indexing into runFullAnalysis. Highlights:

* RepoMeta schema extended: schemaVersion, surfaceSignatures, and
  incrementalInProgress fields. INCREMENTAL_SCHEMA_VERSION = 1.

* core/incremental/file-hash.ts — v1 surface signature: SHA-256 of file
  content. v2 will switch to a true surface-only signature (defined in
  surface.ts) so body-only edits don't expand the closure. The plumbing
  is signature-agnostic so the swap is local.

* core/incremental/orchestrator.ts — eligibility check, closure
  computation (uses file-hash as the surface signal), dirty-flag
  management, subgraph extraction, signature merge.

* run-analyze.ts adds:
  - hasDirtyTree() check on the existing 'lastCommit==HEAD' early-exit
    so an uncommitted edit triggers re-index (was a coarse equality
    check before).
  - incremental branch: try incremental first; fall through to full
    rebuild on any setup failure or eligibility miss.
  - runIncrementalBranch() — opens existing DB, deletes closure-file
    rows + Community/Process, runs pipeline with filesToParse, writes
    only the changed-subgraph back, refreshes FTS, updates meta with
    new surfaceSignatures and clears the dirty flag.
  - Full-rebuild path now populates surfaceSignatures + schemaVersion
    in meta.json so the next run is eligible for incremental.

Crash recovery: incrementalInProgress is set BEFORE any DB mutation
and cleared on success by overwriting meta.json. A crash anywhere in
between leaves the flag set, and the next analyze run forces a full
rebuild (cheapest path back to a known-good index).

v1 limitation documented: body-only edits trigger 1-hop closure
expansion (content-hash signal). True surface-only optimization is
deferred to v2 — see design doc for the integration path.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(incremental): drop invalid --no-renames=false from git diff

The flag --no-renames=false isn't valid git syntax (it's parsed as a
file path). Git's default rename detection is on; removing the flag
keeps that behavior.

Caught while running an end-to-end smoke test against a small fixture
repo: incremental setup failed with 'Command failed: git diff
--name-status -z --no-renames=false ...'. After the fix, the
incremental path runs cleanly: closure is computed, hydrate phase
loads unchanged-file state from DB, parse phase only re-parses files
in closure, and the writeback updates only changed nodes/edges.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Revert v1 incremental indexing (5 commits)

Reverts the v1 design that parsed only closure files into a fresh
graph and tried to hydrate the rest from DB. Real-repo equivalence
test failed: cross-file resolution operates on partial parse data
(closure files only), so CALLS edges that resolve through unchanged
files silently fall off. Diff against full rebuild on the same
edited state: -50 nodes, -425 edges, -5 communities, -48 processes.

Architecture pivot: switch to PR #533-style content-addressed parse
cache. Pipeline parses every file (cache-served when possible),
giving cross-file resolution full data, with DB writeback then
restricted to changed-file rows.

Reverts:
  d4b9de47 fix(incremental): drop invalid --no-renames=false
  f35f7634 feat(analyze): incremental orchestrator branch + meta schema
  bc039686 feat(pipeline): hydrate phase + parse-filter
  98bb893d feat(lbug): loadGraphFromLbug, queryImporters, ...
  aa8d7ae3 feat(incremental): change-detection, surface signatures, closure

Kept:
  d9e340b0 feat(communities): seed Leiden RNG (foundational)
  8235ca36 docs: incremental indexing design spec (will be revised)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(analyze): incremental DB writeback (Option B)

Equivalence-preserving incremental analyze. The pipeline still parses
every file (correctness invariant: cross-file resolution / scope
resolution / MRO / community detection all need full graph data); the
saving comes from selectively replacing only changed-file rows in
LadybugDB instead of wiping and reloading the whole graph.

How it works:

* On every analyze, we hash all source files (SHA-256 of content) and
  store the map in meta.json.fileHashes alongside schemaVersion.
* The next run loads the prior map and diffs:
  - changed: content hash differs → file's DB rows replaced.
  - added: not in prior map → file's DB rows inserted.
  - deleted: in prior map but not on disk → file's DB rows dropped.
* If the diff is non-empty AND no --force / no schema mismatch / no
  dirty flag, take the incremental path:
  - Set incrementalInProgress dirty flag (BEFORE any DB mutation).
  - Open existing DB (no wipe).
  - deleteNodesForFile() for each changed/added/deleted file.
  - deleteAllCommunitiesAndProcesses() — Leiden regenerates these.
  - extractChangedSubgraph() from the in-memory ctx.graph: nodes whose
    filePath is in the writable set + Community + Process + edges with
    at least one endpoint in the writable set (edges entirely between
    hydrated unchanged nodes are skipped — already in DB).
  - loadGraphToLbug() on the subgraph. Unchanged-file rows in DB
    untouched.
  - Recreate FTS indexes.
  - Update meta with new fileHashes; clear dirty flag.
* Otherwise full-rebuild path runs as before.

Crash recovery: incrementalInProgress is the dirty flag. Set before
destructive ops; cleared on success. Set on next-run startup → forces
full rebuild (cheapest path back to known-good).

Other changes:
* Dirty-tree gate on the existing 'lastCommit==HEAD' early-return:
  uncommitted edits no longer slip through as 'already up to date'.
* deleteAllCommunitiesAndProcesses helper in lbug-adapter.
* Skip the embedding cache+restore cycle when willTryIncremental is
  true — embeddings stay in DB; re-inserting them would PK-conflict.

End-to-end equivalence verified on this repo (993 files, 24K nodes):
incremental run produces byte-identical {nodes, edges, clusters,
flows} to a full rebuild from the same edited state.

Speedup is currently modest (~5% on this repo) because the parse
phase still runs in full. Parse-cache integration is a separate
follow-up that composes cleanly on top of this work.

See docs/superpowers/specs/2026-05-10-incremental-indexing-design.md.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(analyze): chunk-level parse cache for full incremental speedup

Composes with the incremental DB writeback (commit 27f3b49d) to deliver
the major-speedup half of incremental indexing. Previously, the parse
phase ran in full on every analyze; the speedup came purely from
selective DB rewriting. With this commit the parse phase also reuses
prior tree-sitter output for chunks whose contents haven't changed.

How it works:

* Cache layer (gitnexus/src/storage/parse-cache.ts):
  - File: <repo>/.gitnexus/parse-cache.json. Versioned, atomic write.
  - Key: chunk content hash = sha256(sorted(filePath:fileContentHash
    for each file in chunk)).
  - Value: ParseWorkerResult[] (raw worker output for the chunk,
    pre-merge).
  - Granularity: per chunk (~20MB byte-budget). A change to one file
    invalidates only its chunk — typically 1 of ~50 on a 1000-file
    repo (~98% cache hit ratio on a small edit).

* Worker contract (gitnexus/src/core/ingestion/parsing-processor.ts):
  - Extracted the chunk-result merge loop into a public
    mergeChunkResults() so the same logic applies to live worker
    output AND replayed cache entries.
  - processParsingWithWorkers / processParsing accept an optional
    outRawResults out-parameter that captures worker output before
    merging — used by parse-impl to populate the cache after a miss.

* Parse phase wiring (parse-impl.ts):
  - For each chunk, compute its content hash (after reading file
    contents). Cache hit → mergeChunkResults() on cached results,
    skip the worker dispatch entirely. Cache miss → run workers
    normally, capture raw results, store under the chunk hash.
  - Cache mutations happen in-place on the ParseCache passed via
    PipelineOptions.parseCache.

* Lifecycle (run-analyze.ts):
  - loadParseCache() before pipeline runs.
  - Cache passed via runPipelineFromRepo's PipelineOptions.
  - saveParseCache() after the pipeline + DB writeback succeed.

Equivalence verified on this repo (993 files, 24K nodes):

  Cold (no cache, full work):           141.1s
  Warm cache + 1-file edit, incremental: 63.6s  ← 55% speedup
  Warm cache + 1-file edit, --force:     71.6s  ← 49% speedup

All three runs produce byte-identical {nodes, edges, clusters,
flows}. The cache survives --force (content-addressed = always
correct), so even forced rebuilds get the parse-skip benefit.

Why chunk-level rather than per-file: workers process sub-batches and
emit aggregated ParseWorkerResults. Per-file granularity would require
restructuring the worker contract; chunk-level captures most of the
practical speedup with no worker-side changes.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* perf(parse-impl): smaller default chunk budget (20MB→2MB) for cache granularity

The parse cache is keyed at chunk granularity. With the previous 20MB
budget, a typical mid-size repo (e.g. this worktree at 9MB total
parseable source) fits in a single chunk — meaning ANY file change
invalidates the whole chunk and re-parses every file.

2MB default produces ~5x more chunks on the same input, so a one-file
edit invalidates ~1/N of cached chunks instead of the whole thing.
Cold-run overhead from more chunks is <5% (one extra serialization
pass per chunk).

Override via GITNEXUS_CHUNK_BYTE_BUDGET env var for benchmarking.

Measured on this repo (~9MB / 887 parseable files):
  Cold (no cache):                    143s
  Warm cache, no source changes:        2s  (early-return)
  Warm cache + 1-file edit:            81s  (~43% off cold)

Speedup is bounded by the scopeResolution phase (~58s flat regardless
of parse cache) and by GitNexus's own auto-writes during analyze
(AGENTS.md / .claude/skills/ etc. mutate between runs and invalidate
chunks containing them). Both are addressable in follow-ups.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* perf(scope-resolution): reuse worker-produced ParsedFile + stabilize chunk order

Two compounding optimizations that drop warm-cache analyze from
~134s to ~38s on a 1000-file repo (72% faster), and cold rebuild
from ~143s to ~86s (40% faster) by short-circuiting work that was
previously re-done.

1. SCOPE-RESOLUTION: REUSE WORKER PARSEDFILE

Previously, the scope-resolution phase re-parsed every file with
tree-sitter on the main thread (~58s on a 1000-file repo) because
worker-produced tree-sitter Trees can't cross the worker MessageChannel.

But the worker ALSO produces a  artifact via
, which structured-clones fine — and it's exactly
what scope-resolution would re-derive. Threading those ParsedFiles
through the parse phase () into
 ( map) lets scope-
resolution skip its extract loop on a per-file basis.

The fast path is bounded only by  per file (cheap
graph mutation). On this repo: scopeResolution went from 58s → 5s.

2. MAP-PRESERVING PARSE-CACHE SERIALIZATION

 is a
which JSON.stringify collapses to . The first attempt at threading
parsedFiles through the parse cache crashed at runtime with
"importerModule.typeBindings is not iterable" because cached entries
came back as plain objects.

Added a JSON replacer/reviver pair in parse-cache.ts that round-trips
Map and Set instances through tagged plain objects (). Symmetric: save uses replacer, load uses reviver.

3. STABLE CHUNK ORDERING

The byte-budget chunker walked files in filesystem-scan order, which
on Windows isn't guaranteed to be stable across runs. Even with
identical source content, two scans could place files in different
chunks, shifting chunk hashes and causing 100% parse-cache misses.

Added a deterministic alphabetical sort on  before
chunking. Chunk membership is now stable across runs, so a single-file
edit invalidates exactly one chunk, not all of them.

Measured on this repo (993 files, 24K nodes):
  Cold rebuild:                        86s  (was 143s)
  Warm cache, no source changes:        3s  (early-return)
  Warm cache + 1-file edit:            38s  (was 134s)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(incremental): update spec + AGENTS.md + GUARDRAILS.md for shipped design

- Rewrite docs/superpowers/specs/2026-05-10-incremental-indexing-design.md
  to describe the architecture that actually shipped (parse cache +
  incremental DB writeback + scope-resolution short-circuit), with the
  v1 hydrate-phase post-mortem preserved as historical context.
- AGENTS.md "Keeping the Index Fresh" section: note that incremental
  is the new default and --force is the explicit opt-out; mention
  the parse-cache file location and that it's safe to delete.
- GUARDRAILS.md Signs: add an "Index seems corrupt or incremental is
  misbehaving" entry pointing users to --force as the manual escape
  hatch (the dirty flag handles automatic recovery).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(incremental): bugbot review + CI test failures

Bugbot (PR #1479):
- Medium: pruneCache was exported but never called -> cache grew
  unbounded. Wire pruneCache into run-analyze before saveParseCache,
  using a transient usedKeys Set on ParseCache that the parse phase
  populates as it processes chunks.
- Low: willTryIncremental (pre-pipeline) and isIncremental
  (post-pipeline) could desync, silently dropping embeddings on
  mispredicted runs. Removed the prediction; the embedding cache
  now loads unconditionally when shouldLoadCache is true. The
  re-insert step gates on the actual isIncremental value to avoid
  PK-conflicts when the incremental-writeback path keeps DB rows.

CI test failures:
- cli-e2e #1169 + run-analyze.test.ts #1233: my dirty-tree gate on
  the lastCommit==HEAD early-return saw GitNexus's own auto-generated
  outputs (.claude/, .cursor/, AGENTS.md, CLAUDE.md) as dirty,
  perpetually defeating the up-to-date fast path. Extended the
  pathspec exclusion to cover all auto-gen outputs, not just
  .gitnexus/.
- ruby field-type disambig: my chunk-stability sort exposed a
  pre-existing order-dependency in Ruby cross-file resolution
  (`user.address.save -> Address#save` only resolves correctly when
  user.rb parses before address.rb in some configurations). Removed
  the sort. Filesystem ordering is stable enough in practice that
  the parse cache still hits the common case; the pre-existing
  fragility is left for a separate fix.
- pipeline-graph-golden: regenerated. Seeded Leiden RNG produces a
  partition different from the previous Math.random snapshot.
- staleness `parallel calls` was a CI timing flake; passes locally.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(incremental): re-insert cached embeddings on incremental path

Bugbot re-review caught: deleteNodesForFile cascades to the
CodeEmbedding table (DELETE WHERE e.nodeId STARTS WITH ...), so
changed-file embedding rows are wiped along with their nodes. The
previous fix gated re-insert on `!isIncremental`, which silently
dropped those embeddings — a regression versus the full-rebuild path's
"preserve embeddings by default" guarantee.

Remove the `!isIncremental` gate. The per-batch try/catch already
handles the unchanged-file PK-conflict case ("some may fail if node
was removed, that's fine") with the same semantics, so re-inserting
the full cached set on incremental works:

  - changed-file rows: deleted, then re-inserted from cache (preserved)
  - unchanged-file rows: still in DB, re-insert PK-conflicts and is
    silently ignored (existing rows are correct)

Cost: re-inserting ~24K embeddings on incremental when only a few
files changed — most are no-op conflicts. Bounded by batch size of
200; ~3-5s overhead. Worth it for correctness.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(incremental): address Claude+Bugbot review findings + remove design doc

Addresses CHANGES_REQUESTED review on PR #1479:

1. Remove docs/superpowers/specs/2026-05-10-incremental-indexing-design.md
   per maintainer request.

2. BLOCKER (Claude Finding 1, Bugbot Round 3): Stale cross-file edges
   between unchanged files. extractChangedSubgraph excluded edges where
   both endpoints were unchanged-file nodes — when a barrel/re-export
   file changes, cross-file resolution may update CALLS edges between
   two unchanged files that would then be silently lost.

   Fix: 1-hop importer-closure expansion of the writable set in
   run-analyze.ts. Before deleting/rewriting rows, query DB for
   importers of every changed/deleted file and add them to the writable
   set. Their nodes get deleted+rewritten too, so cross-file's refined
   edges land in the DB. Re-added queryImporters to lbug-adapter.ts.

3. BLOCKER (Claude Finding 3): Parse cache key omitted parser version.
   After a GitNexus upgrade, the cache silently replays pre-upgrade
   ParseWorkerResults against the new schema → wrong CALLS/IMPORTS/
   scope edges with no visible signal.

   Fix: PARSE_CACHE_VERSION now embeds the gitnexus npm package
   version (read at module load via createRequire on package.json).
   Format: `${SCHEMA_BUMP}+${PKG_VERSION}` e.g. "1+1.6.4". Any release
   that bumps package.json automatically invalidates the on-disk cache.
   Mismatched versions fall through to an empty cache (next save
   overwrites with the new version baked in).

4. BLOCKER (Claude Finding 2): No automated tests for incremental
   behavior. Added 28 unit tests across 3 files:

     - incremental-file-hash.test.ts (10 tests)
       diffFileHashes classification, computeFileHash determinism,
       computeFileHashes batch / missing-file tolerance, sorted output.

     - incremental-parse-cache.test.ts (12 tests)
       computeChunkHash stability and order-independence, version
       prefix format, pruneCache, load/save round-trip on empty /
       missing / corrupt / version-mismatched files, AND a Map/Set
       round-trip test that pins the JSON replacer/reviver behaviour
       (without it, ParsedFile.scopes[*].typeBindings collapses to
       {} and downstream `.get()` / iteration throws).

     - incremental-subgraph-extract.test.ts (6 tests)
       writable-set node inclusion, Community/Process always kept,
       edge inclusion when at least one endpoint is writable, MEMBER_OF
       edges via graph-wide endpoints, empty subgraph case.

5. Medium (Claude Finding 6): AGENTS.md "Keeping the Index Fresh"
   said "only changed files are re-parsed." Imprecise — the pipeline
   parses every file every run; the cache skips tree-sitter for chunks
   whose contents haven't changed. Reworded to match the design doc.

Test plan still expects:
  [x] Typecheck clean
  [x] All 28 new unit tests pass
  [x] All previously-failing tests still pass on the rebased branch
  [x] Equivalence verified locally (incremental ≡ --force, byte-identical
      stats on this repo)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(incremental): round 3 review feedback — bounded BFS, atomic meta, integration test, docs

Addresses remaining findings on PR #1479 from Claude's re-review of
commit ad7bd31 + verifies the outstanding Bugbot HIGH severity.

1. F1 — Transitive importer expansion (Claude, was Medium-but-noted).
   Previous 1-hop importer expansion missed barrel re-export chains
   (A imports C, C re-exports B; when B changes, only C was pulled in
   — A was left with potentially-stale CALLS edges to refined targets).
   Replaced the single pass with a bounded BFS over the IMPORTS graph
   (depth ≤ 4). Catches nested barrel pyramids without ballooning into
   a near-full rebuild on monorepos with deep re-export trees. `--force`
   remains the escape hatch documented in GUARDRAILS.md for cases that
   exceed the bound.

2. F2 — Integration test for incremental orchestration (Claude, BLOCKER,
   DoD §2.7). The unit tests added in ad7bd31 covered `diffFileHashes`,
   `extractChangedSubgraph`, `computeChunkHash`, `pruneCache`, and the
   Map/Set JSON round-trip — but none of them exercised the real
   `runFullAnalysis` orchestration. Added gitnexus/test/unit/
   incremental-orchestration.test.ts with four end-to-end tests against
   a real git-initialized fixture repo + real LadybugDB:

     a. First run populates fileHashes + schemaVersion and clears
        incrementalInProgress on success.
     b. Second run on unchanged state takes the alreadyUpToDate fast
        path (early-return).
     c. Second run after a source edit takes the incremental path
        (not full rebuild) and rotates fileHashes for the touched file
        while keeping the dirty flag cleared.
     d. A pre-set incrementalInProgress flag forces a full rebuild
        that clears it (crash-recovery wire).

   These would catch any regression that wires `isIncremental` from a
   pre-pipeline prediction (the Bugbot finding from commit 5eb0597) or
   accidentally re-gates the embedding re-insert on `!isIncremental`
   (the Bugbot finding from commit 60c10f1).

3. F3 — GUARDRAILS.md docs accuracy (Claude, Low). Line 33 still said
   "only changed files are re-parsed" — AGENTS.md was already corrected
   in ad7bd31 but GUARDRAILS.md was missed. Reworded to match.

4. F5 — Atomic saveMeta (Claude, Medium; vvladescu-tb fork). The dirty
   flag (`incrementalInProgress`) travels through meta.json. A crash
   mid-write would leave a corrupt meta.json that `loadMeta` would
   silently treat as "no prior index", losing the flag and skipping
   recovery. Switched to tmp-file + rename matching saveParseCache.

5. Bugbot's "Subgraph edges reference nodes absent from subgraph"
   (HIGH severity). Verified as FALSE POSITIVE: `getNodeLabel` in
   lbug-adapter.ts derives labels from the node-ID string (parses
   the table prefix), not from the in-memory graph. The CSV
   generator writes (src_id, dst_id, type) rows without consulting
   node objects; `splitRelCsvByLabelPair` routes by ID-derived label;
   `COPY ... (from=X, to=Y)` resolves both endpoints against the live
   LadybugDB where unchanged-file nodes still exist. No fix needed.

All 213 tests pass locally (including the 4 new integration tests
and the previously-failing CI tests).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(incremental): address Bugbot round-4 findings (added-file shadow seed + dedupe)

Bugbot review on commit e23e4400 surfaced two new findings against the
incremental writeback in run-analyze.ts:

  HIGH — Incremental BFS misses importers of newly added files.
    queryImporters() reads the pre-pipeline DB. For a NEWLY ADDED
    file there are no IMPORTS rows pointing to it yet, so unchanged
    files whose pre-existing import statements now resolve to the
    newcomer keep stale CALLS edges pointing at the OLD resolution
    target.

  LOW — Deleted files double-counted in filesToDelete.
    hashDiff.deleted entries can reappear in writableFiles via the
    BFS expansion (queryImporters can return a now-deleted path),
    so deleteNodesForFile() ran twice for the same file.

Fixes:

  - Add gitnexus/src/core/incremental/shadow-candidates.ts: derive
    the pre-existing file paths whose JS/TS module-resolution claim
    an added file can steal. Pattern catalogue: same-basename/
    different-extension, bare-file-beats-directory-index, and
    directory-index-beats-bare-file. Emit both POSIX and Windows
    separators because the prior fileHashes map may have been
    written from either OS.

  - In run-analyze.ts, seed the BFS frontier with shadow candidates
    that exist in the prior meta.fileHashes. Their importers — found
    via queryImporters — get pulled into the writable set so their
    CALLS edges re-resolve against the new file.

  - Dedupe filesToDelete via Set to avoid the double-call.

Tests: gitnexus/test/unit/incremental-shadow-candidates.test.ts —
8 cases covering each shadow pattern, separator handling, .d.ts as
a single extension token, deduplication, and the no-self-shadow
invariant. All 40 incremental tests (file-hash, parse-cache,
subgraph-extract, shadow-candidates, orchestration) pass locally.

Note on the third Bugbot finding ("Subgraph edges reference nodes
absent from subgraph"): re-anchored from a prior review pass — the
code at subgraph-extract.ts:48 is unchanged. Already verified as a
false positive: getNodeLabel parses labels from ID strings, CSV
write is by ID, and COPY resolves against the live DB.

* chore(autofix): apply prettier + eslint fixes via /autofix command

* test(incremental): exact-equality stats invariant + analyze ≡ analyze --force

Addresses the only remaining Claude production-readiness review finding
on PR #1479 (Low-Medium, test-quality only — Claude itself said it does
NOT block merge, but the central PR claim "incremental ≡ full rebuild"
deserves explicit CI coverage rather than implicit trust).

Changes to gitnexus/test/unit/incremental-orchestration.test.ts:

1) Tighten the existing "comment-only edit takes incremental path" test.
   - Replace toBeGreaterThan(0) bounds assertions on stats.files and
     stats.nodes with exact toBe(firstMeta) per-field equality across
     files / nodes / edges / communities / processes. DoD §2.7 calls
     out bounds-only assertions as masking regressions that drop half
     the graph; this swap closes that gap.
   - Rationale: a comment-only edit must change the file content hash
     (driving the incremental path) without changing any graph data.
     Therefore every stat MUST be identical to the first run. Anything
     else is a regression.

2) New test: incremental output is byte-equivalent to a full rebuild.
   - Run analyze → comment-only edit → analyze (incremental writeback)
     → analyze --force (full rebuild from same on-disk state).
   - Assert files / nodes / edges / communities / processes are exactly
     equal across the incremental and the --force passes.
   - This is the PR's central correctness contract, now proven by a
     test that exercises the real runtime path end-to-end against a
     real on-disk LadybugDB.

All 5 orchestration tests pass locally (52s), including the new
equivalence test — every stat field matches exactly between incremental
and --force on the mini-repo fixture.

tsc --noEmit clean.

* fix(incremental): F1 cross-file edge consistency + F4 stable chunk sort + unit coverage (#1511)

Patch addressing two of the still-open changes-requested findings on PR
#1479, rebased onto the current feat/incremental-indexing head. F3
(parser fingerprint in the cache key), F5 (atomic saveMeta), and F6
(AGENTS.md phrasing) were already handled on the branch, so the
corresponding parts of the original patch were dropped as redundant.

  F1 (Blocker) — Cross-file edges between unchanged files
    Adds `computeEffectiveWriteSet(graph, toWriteSet)` to
    subgraph-extract.ts: a single pass over the new graph's edges that
    pulls the unchanged-side file of every writable-boundary-crossing
    edge into the write set. run-analyze composes it ON TOP of the
    existing importer-BFS expansion and feeds the combined set to BOTH
    `deleteNodesForFile` and `extractChangedSubgraph`, so the delete
    cascade and the writeback subgraph cover identical files (asymmetry
    would leave stale rows or PK-conflict at COPY time). The BFS reads
    IMPORTS from the pre-pipeline DB (catches files that *stopped*
    importing a changed file); the edge walk reads the new graph
    (catches refined CALLS edges the pre-run DB couldn't predict, e.g.
    a barrel re-export shifting a symbol from B to D). `extractChangedSubgraph`
    stays a pure filter — all expansion is the orchestrator's job.

  F4 (Medium) — Restore alphabetical chunk sort
    `parseableScanned` is sorted before chunking. Filesystem-scan order
    isn't stable enough across runs/platforms (notably macOS APFS) to
    keep chunk hashes consistent, so the parse cache thrashes without
    it. The pre-existing Ruby cross-file resolution order-dependency the
    old comment cited is independent — the sort surfaces it but doesn't
    cause it; tracked separately rather than leaving the cache cold.

  Tests — incremental-subgraph-extract.test.ts
    Locks the F1 invariants: `extractChangedSubgraph` is a pure filter
    (includes only the set it's given, plus graph-wide nodes; edges
    fire on one writable endpoint), and `computeEffectiveWriteSet`
    covers the barrel-re-export scenario, the symmetric edge-into-
    changed-file case, the no-boundary-crossed no-op, graph-wide-node
    edges, and input-immutability. Supersedes the prior
    extractChangedSubgraph-only test file on the branch.

Co-authored-by: Val Vladescu <vvladescu-tb@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(call-processor): register properties in pre-pass to fix order-dependent field type disambiguation + regenerate golden snapshot

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2d66666f-861c-432e-a4b0-11f2aefca98a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(call-processor): port worker-path property enrichment into the sequential pre-pass

Copilot's pre-pass in 8184439 fixed the Ruby attr_accessor order-dependence,
but it copied the OLD in-loop registration logic, not the canonical worker
path in parse-worker.ts. That left the sequential and worker paths emitting
non-identical Property nodes/symbols for the same source — silently breaking
the `incremental ≡ --force` invariant the moment a repo crosses the worker
threshold between runs.

Two concrete divergences are closed here:

  * Node id: worker keys Property as `${file}:${className}.${propName}`
    (qualified). Pre-pass was using `${file}:${propName}` (unqualified).
    Same source produced different graph ids depending on which path ran.

  * Field metadata: worker enriches each routed property with
    `provider.fieldExtractor` + `getFieldInfo`, falling back to
    `routedFieldInfo.type` for `declaredType` when the routing payload
    lacks one (e.g. types discovered from `@address = Address.new`
    ctor assignments rather than YARD `@return [Type]`), and propagates
    `visibility` / `isStatic` / `isReadonly`. Pre-pass did none of this,
    so on the sequential path `resolveFieldAccessType` failed to walk
    chains where the type only came from the FieldExtractor.

The pre-pass now mirrors parse-worker.ts:1803-1898 verbatim, with one
deliberate difference: the FieldInfo cache is scoped to a single
`processCalls` invocation rather than module-level (the worker process
is short-lived; the main thread is not, and a module-level cache would
leak state between analyze runs).

Also drops the now-stale "Defer resolution: Ruby attr_accessor properties
are registered during this same loop" comment on `pendingWrites.push` —
the rationale is no longer accurate after Copilot's pre-pass, but the
deferral is still needed so write-access tracking sees inference that
completes during the main loop. Comment updated to reflect that.

Verification:
  * `tsc --noEmit`: 0 errors
  * test/unit (call-processor, call-routing, field-extraction, ruby-self-call): 224 passing
  * test/integration (ruby, ruby-sequential-mixin, pipeline-graph-golden): 137 passing

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(call-processor): key fieldInfoCache by filePath:startIndex, not raw byte offset

Claude's review of 255bdf6 caught a real collision in the FieldInfoCache I
added: keying by `classNode.startIndex` alone is a per-file byte offset, so
two files that both begin with a class at byte 0 — extremely common in Ruby /
Python, where files frequently open with `class Foo`, `module Foo` — collide
on the same cache entry. The second file's `getFieldInfo` then returns the
first file's FieldInfo map, producing wrong `declaredType` / `visibility` /
`isReadonly` on its properties.

Same shape as the bug that already exists in parse-worker.ts:377 (also keyed
by `classNode.startIndex` in a module-level map, persistent across files
processed by the same worker). Fixing the symmetric pre-existing leak in
parse-worker.ts is a separate, scoped follow-up — left out of this commit to
keep the fix minimal and reviewable.

Cache map and key are now both string-typed. Composite key
`${context.filePath}:${classNode.startIndex}` keeps the within-file hit rate
(one FieldExtractor.extract() per class regardless of how many
`attr_accessor` lines it has) while eliminating cross-file aliasing.

Verification on the patched HEAD:
  * `tsc --noEmit`: 0 errors
  * test/unit (call-processor, call-routing, field-extraction, ruby-self-call): 224 passing
  * test/integration (ruby, ruby-sequential-mixin, pipeline-graph-golden): 137 passing

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Val Vladescu <val.vladescu@thirdbridge.com>
Co-authored-by: Val Vladescu <vvladescu-tb@users.noreply.github.com>
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-05-12 13:14:56 +01:00
d4f34905bc feat: migrate Java to scope-based registry resolution (RFC #909 Ring 3) (#1482)
* Initial plan

* feat: implement Java scope-based resolution (RFC #909 Ring 3)

Add scope-resolution pipeline for Java, following the C# pattern:

- query.ts: tree-sitter query for scopes, declarations, imports,
  type bindings, and references against tree-sitter-java grammar
- captures.ts: orchestrator synthesizing import decomposition,
  receiver bindings (this/super), arity metadata, and reference arity
- import-decomposer.ts: decompose import_declaration nodes into
  kind/source/name markers (named, wildcard, static, static-wildcard)
- interpret.ts: convert captures to ParsedImport/ParsedTypeBinding
- receiver-binding.ts: synthesize this/super type-bindings on instance
  methods with superclass support
- arity-metadata.ts: extract parameter count/types using javaMethodConfig
- arity.ts: Java arity compatibility check with varargs support
- merge-bindings.ts: Java shadowing precedence (local > import > wildcard)
- simple-hooks.ts: bindingScopeFor, importOwningScope, receiverBinding
- import-target.ts: package path to file path resolution
- scope-resolver.ts: ScopeResolver implementation registered in registry

Wire scope hooks into javaProvider (java.ts) and register
javaScopeResolver in SCOPE_RESOLVERS registry. Add createResolverParityIt
wrapper to java.test.ts for parity testing.

All 172 existing Java tests pass. Java is NOT added to
MIGRATED_LANGUAGES — the resolver sits idle until the migration flag
is flipped.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix: address review findings 1-4 — varargs arity, static import resolution, importOwningScope, stripGeneric

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/22308da3-59c9-47e6-8e52-738305b1b80a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docs: document registry-primary parity status and CI visibility gap in scope-resolver

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/22308da3-59c9-47e6-8e52-738305b1b80a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: add generic type erasure fallback in stripGeneric + update scope-resolver docs

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/223f77ac-59a7-4487-9316-f2be05eac5d3

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: improve stripGeneric fallback regex — use valid Java identifier chars and handle nested generics

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/223f77ac-59a7-4487-9316-f2be05eac5d3

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: address adversarial review findings 1-6 — flaky test, wildcard import fixture, varargs fixed-prefix test, qualified generic stripping, JSDoc updates

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/172c8a1a-cdf3-4de8-9142-f2c12c14b0a6

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docs: add inline comment explaining stripQualifier/stripGeneric call order

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/172c8a1a-cdf3-4de8-9142-f2c12c14b0a6

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: add varargs 0-arg fixture and strengthen wildcard import assertions

Finding 1: Added `badCall()` method with 0-arg `fmt.format()` call to the
varargs fixture. Test documents that legacy mode still resolves this call
(arity rejection is registry-primary only). The fixture now exercises both
the success path (2-arg, 3-arg) and the undersupplied path (0-arg).

Finding 2: Strengthened wildcard import test to assert `targetFilePath`
on the CALLS edge (`com/example/models/User.java`), confirming the call
resolved through the wildcard-imported type to the correct file.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2b4e5602-9833-485c-ab48-e1d54fdf8465

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-12 09:37:44 +01:00
Abhigyan Patwari f33efa8714 Merge pull request #1523 from magyargergo/fix/claude-review-skip-permissions-ci
ci(claude): fix /review PR comments blocked by Bash approval in Actions
2026-05-12 09:12:46 +01:00
Gergo MagyarandCursor 2bf6d078aa ci(claude): allow Bash in code-review job without interactive approval
Claude Code defaults to prompting for Bash approval. In GitHub Actions there
is no human to approve, so gh pr comment and similar commands fail and the
PR receives no review comment. Pass --dangerously-skip-permissions for the
code-review step only (headless CI; token and checkout are already scoped).

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-12 09:07:47 +01:00
Abhigyan Patwari ba1ac3989d Merge pull request #1522 from abhigyanpatwari/fix/claude-review-pr-comment
fix(ci): make /review reliably post PR comments
2026-05-12 08:29:39 +01:00
Gergo Magyar c5c42a9649 fix(ci): make /review reliably post PR comments
The Run Claude Code Review step passed an invalid PR ref
(owner/repo/pull/N) which gh interprets as a branch name, causing
early gh pr view failures. More importantly, the prompt omitted
--comment, so the code-review plugin only displayed findings in
terminal output and never invoked gh pr comment to post to the PR.

Switch to a full PR URL and add --comment so the plugin posts the
review during the session, which also routes around upstream bugs
anthropics/claude-code-action#1061 and #1087 where the action's
post-step capture can silently drop output on issue_comment triggers.
2026-05-12 08:28:42 +01:00
Abhigyan Patwari a445d38e0b Merge pull request #1521 from abhigyanpatwari/remove_flaky_test
chore(tests): remove flaky regression test for resource exhaustion
2026-05-12 08:00:04 +01:00
Gergo Magyar 0e2c0c77ec chore(tests): remove flaky regression test for resource exhaustion 2026-05-12 07:58:58 +01:00
AntheurusandGergő Magyar fcab1e2e82 fix(augment): add CONTAINS fallback when FTS indexes unavailable (#1476)
* fix(augment): add CONTAINS fallback when FTS indexes unavailable

When the MCP server holds the KuzuDB write lock, the augment CLI opens
the DB read-only. FTS indexes cannot be created in read-only mode, so
searchFTSFromLbug returns ftsAvailable=false and an empty results array.
The existing early-return path silently produced no enrichment.

Add a Cypher name CONTAINS fallback that fires only when ftsAvailable is
false and BM25 produced no symbol matches. This covers the read-only DB
case (concurrent MCP server) and the first-run case (indexes not yet
built). The fallback is wrapped in .catch(() => []) and cannot throw.

When FTS indexes exist, this branch is never reached — behaviour is
unchanged for users without a concurrent MCP server.

* fix(augment): guard against CONTAINS '' and add no-FTS test coverage

Blocker 1 — CONTAINS '' on whitespace-leading patterns:
pattern.split(/\s+/)[0] returns "" when the input has leading whitespace
(e.g. "   ".split(/\s+/) → ["", ""]). In Kuzu, CONTAINS '' matches every
node with a name property, injecting arbitrary graph nodes into LLM context.

Fix: trim() before split, then guard on !firstWord || firstWord.length < 2.
No behaviour change for normal non-empty patterns.

Blocker 2 — zero test coverage on the FTS-unavailable code path:
The new CONTAINS fallback block (engine.ts lines 146-166) was exercised by
no existing test — all existing tests run with FTS indexes built. A second
withTestLbugDB fixture is added with no ftsIndexes, forcing searchFTSFromLbug
to return ftsAvailable: false, and asserts:
1. augment('login', ...) returns non-empty enrichment (fallback works)
2. augment('   ', ...) returns '' (CONTAINS '' guard holds)
3. augment('nxyz_notfound', ...) returns '' (no matching nodes)
4. executeQuery throwing returns '' (.catch(() => []) path)

* fix(augment): extend CONTAINS '' guard to FTS happy path and consolidate

The same split(/\s+/)[0] bug existed at line 125 (BM25 symbol filter,
FTS-available path) — a leading-whitespace pattern produced CONTAINS ''
there too, matching every node in BM25-matched files.

Fix: hoist patternFirstWord computation with trim() and the length guard
to the top of augment(), before any DB interaction. Both CONTAINS sites
(BM25 symbol filter and CONTAINS fallback) now use the single pre-validated
value. No behaviour change for normal patterns; the guard fires once for
all callers instead of being duplicated.

Also tighten the whitespace test in the no-FTS suite from 3 spaces to
4 spaces so it unambiguously exercises the patternFirstWord guard rather
than straddling the outer pattern.length < 3 boundary.

* test(augment): negative-safety test for ftsAvailable=true gate

Asserts the CONTAINS fallback does NOT fire when FTS is available but
BM25 returns zero results. Pins the safety property promised by the PR
description: behavior is unchanged for users without the read-only-DB
condition.

If anyone later loosens the gate to `symbolMatches.length === 0` alone,
this test fails.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-11 16:39:10 +01:00
fdf1effb2a feat(cli): add --skip-skills and --index-only flags to analyze (resubmit of #742) (#1485)
* feat(cli): add --skip-skills and --index-only flags to analyze command

The `installSkills()` call in `generateAIContextFiles()` runs
unconditionally, injecting 6 skill files into `.claude/skills/gitnexus/`
even when `--skip-agents-md` is passed. This is problematic for bulk
indexing operations on read-only mirrors or third-party repos.

Add two new flags:
- `--skip-skills`: suppress standard GitNexus skill file injection
- `--index-only`: pure index mode that suppresses all file injection
  (AGENTS.md, CLAUDE.md, and skills), writing only to `.gitnexus/`

This gives users three levels of control:
- `--skip-agents-md` — suppress only root context files
- `--skip-skills` — suppress only skill injection
- `--index-only` — suppress everything (pure indexing)

Discovery context: while bulk-indexing 176 repos with
`--skip-agents-md`, all 144 indexed repos were contaminated with
`.claude/skills/gitnexus/` files requiring manual cleanup.

* fix(cli): address PR #742 review — gate community skills, drop dangling refs, add tests

Bot review (#742) flagged three issues with the original commit:

1. `--index-only --skills` still wrote community-derived skill files
   to `.claude/skills/generated/`. The `--skills` branch in analyze.ts
   was not gated by `skipAll`, so the "skip all file injection" contract
   was violated. Gate `generateSkillFiles()` with `!skipAll` so
   `--index-only` truly wins over `--skills`.

2. `--skip-skills` without `--skip-agents-md` produced AGENTS.md /
   CLAUDE.md that still referenced `.claude/skills/gitnexus/*/SKILL.md`
   files that were never installed — every agent load incurred 6
   failed reads. Pass `skipSkills` through to `generateGitNexusContent()`
   and omit the standard-skill rows (and the entire `## CLI` heading
   when the table is empty). Community skills, when present via
   `--skills`, are unaffected.

3. No filesystem tests for `skipSkills` / `indexOnly`. Add three
   regression guards to `test/unit/ai-context.test.ts`:
   - `.claude/skills/gitnexus/` is NOT created when skipSkills=true
   - Nothing is written when both skipAgentsMd and skipSkills are true
     (the resolved-flag state from --index-only)
   - AGENTS.md/CLAUDE.md routing table omits standard skill references
     when skipSkills=true, but preserves the load-bearing imperative
     sections (Always Do / Never Do / Resources)

* test(cli): PR 1485 review follow-ups (help text, gate test, --skip-skills docs)

- Assert --skip-skills and --index-only in analyze --help (skip-git-cli.test.ts).

- Export shouldGenerateCommunitySkillFiles; unit-test index-only+skills gate.

- Clarify --skip-skills does not suppress --skills community files; --index-only for full skip.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli): warn when --index-only silently overrides --skills

Address review findings on PR 1485 follow-ups:
- analyze.ts emits a one-line note when both --index-only and --skills
  are set, so users see why a pipeline re-index ran with no skill files
  written.
- index.ts --skills help text now flags the --index-only override.
- shouldGenerateCommunitySkillFiles JSDoc documents the dual role of
  the gate (community skills + AGENTS.md/CLAUDE.md re-generation).
- skip-git-cli.test.ts pins the override-warning surface end-to-end.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-11 15:00:58 +01:00
622f98ade5 feat(embeddings): forward dimensions param in HTTP embedding requests (#1498)
* feat(embeddings): forward GITNEXUS_EMBEDDING_DIMS as dimensions in HTTP request body

When GITNEXUS_EMBEDDING_DIMS is set, include it as the `dimensions` field
in the /v1/embeddings request body. This enables Matryoshka-capable models
(OpenAI text-embedding-3-*, Cohere embed-v3, Voyage) to return truncated
vectors at the requested size.

When the env var is unset, the request body remains `{ input, model }` —
no breaking change for backends that reject unknown fields.

Adds 4 unit tests covering both paths (with/without dimensions) on both
the batch embed and single-query embed code paths.

* fix(embeddings): address review findings — strict parseInt, multi-batch test, comment wording

1. Strict parseInt validation: reject non-numeric strings like '1024abc'
   by checking /^\d+$/ before parseInt (Finding 1).
2. Add multi-batch test asserting dimensions is forwarded in every fetch
   call when inputs exceed batch size (Finding 2).
3. Soften JSDoc comment: backends may ignore or reject the dimensions
   field rather than universally ignoring it (Finding 3).
4. Add test for invalid GITNEXUS_EMBEDDING_DIMS values.

---------

Co-authored-by: henry <zhangwei2017@unipus.cn>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-11 13:56:10 +01:00
juyua9andGergő Magyar 55b7a79beb fix(group): detect httpx async consumers (#1408)
* fix(group): detect httpx async consumers

* test(group): tighten httpx consumer coverage

* test(group): create extractor temp dirs safely

* fix(group): scope httpx async client tracking

* fix(group): tighten httpx module-scope tracking

Prevent module-scope httpx.AsyncClient tracking from matching same-name local variables inside functions.

Also documents the intentionally unsupported direct-import, alias, and typed-assignment forms, and extends the httpx extractor regression fixture to cover module-scope shadowing while keeping module-scope calls detected.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-11 13:13:49 +01:00
uwence a3af0dcce2 fix(docker): include duckdb installer script in runtime image (#1502) 2026-05-11 11:38:59 +01:00
RinandGergő Magyar 6a8947217c fix(server): sanitize repo name to prevent argument injection (#1305)
* fix(server): sanitize repo name to prevent argument injection

Sanitizes the extracted repository name to prevent argument injection during git clone operations and ensures compatibility with various file systems.

1. Strips leading dashes to prevent git command-line argument injection.

2. Replaces unsafe directory characters with underscores.

3. Blocks path traversal segments ('.' and '..') and Windows reserved names.

4. Fixes ReDoS vulnerability in parseRepoNameFromUrl regex.

5. Added unit tests for sanitization and path traversal edge cases.

* fix(server): expand Windows reserved name check to include extensions

- Updated sanitizeRepoName to block Windows reserved names (CON, NUL, etc.) even when they have extensions (e.g., CON.txt).
- Corrected regex and added unit tests for these edge cases to resolve CI failures on Windows.
- Ref: https://github.com/abhigyanpatwari/GitNexus/pull/1305#issuecomment-4407200914

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-11 09:38:07 +01:00
e412d292fe feat: migrate C to scope-based resolution (RFC #909 Ring 3) (#1481)
* Initial plan

* feat: add C scope resolution files for language migration (RFC #909)

Add 11 C language scope resolution files following the Go pattern:
- query.ts: tree-sitter-c query and parser for C constructs
- captures.ts: emit scope captures with arity enrichment
- import-decomposer.ts: decompose #include into structured captures
- arity-metadata.ts: C function declaration/call arity computation
- interpret.ts: interpret C imports and type bindings
- import-target.ts: resolve #include paths via suffix matching
- arity.ts: C arity compatibility (variadic detection)
- merge-bindings.ts: first-wins binding merge by tier
- simple-hooks.ts: null hooks (no receivers/methods in C)
- index.ts: barrel re-exports
- scope-resolver.ts: ScopeResolver implementation for C

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* feat: migrate C to scope-based resolution (RFC #909 Ring 3)

Add C ScopeResolver with:
- tree-sitter-c scope query (structs, unions, enums, functions, macros, variables, includes)
- emitCScopeCaptures with arity enrichment and typedef-struct dedup
- interpretCImport for #include directives (system headers filtered)
- resolveCImportTarget with suffix matching
- cArityCompatibility with variadic detection
- cMergeBindings (first-wins by tier)
- Header file scanning for cross-language #include resolution
- Register in SCOPE_RESOLVERS and MIGRATED_LANGUAGES
- Integration test with 4 passing test cases

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/ddcbc075-2999-492c-a0ac-47cddd401a4b

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: update registry-primary-flag test and add C legacy parity expected failures

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/ddcbc075-2999-492c-a0ac-47cddd401a4b

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor: improve arity-metadata readability per review feedback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/ddcbc075-2999-492c-a0ac-47cddd401a4b

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: address CI failures — unused import, Dirent types, null comparison, lint, formatting

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/f4b6e20d-8d56-4834-8296-db82af27f8e1

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: replace loose comparisons with strict equality, remove optional chaining from childForFieldName

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/20628629-9dc7-45ec-8ae4-f3a14ad29d93

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* Potential fix for pull request finding 'CodeQL / Comparison between inconvertible types'

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

* fix: remove unnecessary optional chaining on non-null decl in findFuncDeclarator

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9b012c35-7494-4180-8c6a-b83da8d8abb9

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: address 5 production readiness review findings

Finding 1: Enforce static functions as file-local via expandsWildcardTo hook
- Add static-linkage.ts tracking module with markStaticName/isStaticName/expandCWildcardNames
- Update captures.ts to detect storage_class_specifier static on functions
- Wire expandsWildcardTo in scope-resolver.ts

Finding 2: Expand test coverage to ≥30 cases (74 unit tests added)
- c-captures.test.ts: 55 tests (scopes, structs, unions, enums, functions, typedef, field, variable, macro, imports, references, type bindings, arity, static)
- c-imports.test.ts: 12 tests (decomposition, interpretation, target resolution, determinism, edge cases)
- c-arity.test.ts: 18 tests (declaration arity, call arity, compatibility)

Finding 3: Deterministic #include resolution on depth ties
- Add lexicographic tiebreak in import-target.ts when candidates tie on path depth

Finding 4: Revert unexplained package-lock.json change
- Restored to pre-PR state (node >=20.0.0)

Finding 5: Planning artifact commit acknowledged (squash on merge)

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/22ee780c-2b44-4e69-b9c4-8843ad6ec1ee

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: address code review feedback — Set-based dedup, SyntaxNode type alias

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/22ee780c-2b44-4e69-b9c4-8843ad6ec1ee

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix: address second review findings 1-4 — isFileLocalDef hook, singleton docs, fn-ptr typedef, static isolation test

Finding 1: Added `isFileLocalDef` hook to ScopeResolver contract + implementation
in free-call-fallback.ts to filter C static functions from global free-call
fallback. Threads caller filePath through pickUniqueGlobalCallable so static
defs in other files are excluded.

Finding 2: Documented single-invocation assumption on staticNames Map. Added
clearStaticNames() call in loadResolutionConfig to prevent cross-repo
contamination in server-mode scenarios.

Finding 3: Added tree-sitter query pattern for function pointer typedef aliases
(typedef void (*callback)(int, int)) in query.ts. Added unit test.

Finding 4: Added c-static-isolation integration fixture (a.c with static helper,
b.c with non-static helper, caller.c) and test asserting no CALLS edge from
caller to a.c's static helper.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/5b948327-9f59-4ca2-9d8c-8c8087feb510

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix: skip static isolation integration test in legacy parity mode

The `caller.c calls b:helper via include, NOT a:static helper` test
requires scope-based wildcard import binding + isFileLocalDef filtering
which is only available in the registry-primary path. The legacy DAG
path does not resolve cross-file calls through #include → prototype
chains. Added to LEGACY_RESOLVER_PARITY_EXPECTED_FAILURES.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/3be48249-b375-4446-973b-657400f530fb

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: address 3 review findings — static leakage in Phase 2, build-dir skip list, same-directory preference

Finding 1: Apply isFileLocalDef filtering in Phase 2 of pickUniqueGlobalCallable
so cross-file static defs cannot leak through the SemanticModel fallback path.

Finding 2: Expand scanHeaderFiles skip list with dist, build, out, target,
_build, .next, cmake-build-* to avoid generated headers shadowing source ones.

Finding 3: Implement same-directory sibling preference in resolveCImportTarget,
matching C compiler #include "…" relative-lookup semantics. Sibling check now
runs before exact match and suffix fallback.

Tests: 11 new header-scan tests, 4 new import-target tests (96 total C tests).

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/5236f2b8-72a0-476d-bf39-cca041781014

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor: skip redundant sibling path computation when targetRaw has no backslashes

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/5236f2b8-72a0-476d-bf39-cca041781014

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix: normalize header-scan paths to forward slashes for Windows compatibility

On Windows, path.relative() produces backslash-separated paths (e.g. src\foo.h)
which caused test failures when comparing against forward-slash expectations.
Normalize output with .replace(/\\/g, '/') matching the convention used
throughout the scope-resolution pipeline.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/88420759-8a8c-4e6d-affa-ecb5458a235a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: address 4 findings — K&R arity, function-pointer docs, prototype docs, _filePath rename

Finding 1: K&R empty parameter list `int foo()` now returns unknown arity `{}`
instead of `{parameterCount:0}`. Distinguishes from explicit `int foo(void)`.
3 unit tests cover K&R definition, prototype, and void comparison.

Finding 2: Added code comment documenting function-pointer-variable call
capture as known architectural trade-off (same as Go resolver).

Finding 3: Added code comment documenting prototype/definition duplication
as graph-quality concern (no false CALLS edges).

Finding 4: Renamed `_filePath` → `filePath` in captures.ts since it is
actively used in markStaticName().

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/ab2164f7-972f-4ea8-81fa-a14ce20d7cce

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-11 09:02:00 +01:00
Hector PratsandGergő Magyar d69eadfb7f fix(windows): 32767-char tree-sitter crash + VECTOR extension SIGSEGV (#1433)
* fix(windows): 32767-char tree-sitter crash + VECTOR extension SIGSEGV

tree-sitter 0.21.x on Windows crashes with SIGSEGV when parsing source
strings longer than 32 767 chars (signed 16-bit integer overflow in the
native binding). Five call sites passed raw file content without any
length guard:

  - captures.ts (C# scope extraction)
  - namespace-siblings.ts (extractFileStructure)
  - parse-worker.ts (worker thread parse path)
  - parsing-processor.ts (sequential parse fallback)

Fix: truncate at the last newline before the limit so the fragment stays
syntactically coherent. Files truncated mid-class produce ERROR roots;
captures.ts returns [] for any ERROR-root tree so the legacy DAG handles
the file silently without orphaned scope errors.

Additional C# scope fixes:
  - scope-tree.ts: Module scopes may share the same range as a top-level
    namespace_declaration (files with no leading `using` directives). The
    rangeStrictlyContains check rejects equal ranges. Added
    rangeNonStrictlyContains for Module parents.
  - scope-extractor.ts: pass1BuildScopes stack-pop used strict containment;
    same Module == Namespace range case caused orphaned scopes. Added
    moduleAwareContains helper.
  - scope-extractor-bridge.ts: empty captures from ERROR-root files still
    called extractScope -> "no Module scope found" warning. Added early
    return for empty/non-array captures.
  - namespace-siblings.ts: three sites pushed onto binding arrays frozen by
    finalize-algorithm. Fixed with spread-copy before mutation.

lbug-adapter.ts: INSTALL VECTOR in loadVectorExtension calls the KuzuDB
native extension installer, which crashes with SIGSEGV on Windows via an
unhandled error path in native code. JS try/catch cannot intercept native
signals. Skip extension loading on win32 — vector/embedding search is
unavailable on Windows but all graph index queries work correctly.

Verified on: Windows 11, Node.js 24, gitnexus 1.6.3, pcf8-game codebase
(61 757 nodes / 111 796 edges / 300 flows after fix).

* fix(windows): skip FTS extension load in pool-adapter on Windows to prevent SIGSEGV

LOAD EXTENSION fts crashes the process with SIGSEGV on Windows when the
FTS extension binary is not installed locally. This is an @ladybugdb/core
native bug — the extension loader hits an unhandled error path that raises
a native signal instead of a JS exception, so try/catch cannot protect here.

Add a process.platform === 'win32' guard in both doInitLbug and
initLbugWithDb. When skipped, bm25-index.js catches the resulting
Kuzu catalog errors (CREATE_FTS_INDEX not defined) and returns empty
BM25 results gracefully. All graph queries (cypher, context, impact)
are unaffected.

This is patch 9 of the Windows fix series for gitnexus on Windows:
patch 8 (same PR) already fixed INSTALL VECTOR SIGSEGV in lbug-adapter.ts.
pool-adapter.ts is the separate MCP-server code path that was not covered.

* fix: address codeql findings on PR #1433

The four `lastIndexOf('\n', ...)` calls were committed with a literal
newline inside the single-quoted string instead of the `\n` escape, so
the files do not parse — `tsc` and CodeQL both flagged them. Replace
the embedded newline with `'\n'`.

Also remove the two helpers that were superseded during review and
became dead code: `rangeNonStrictlyContains` in scope-tree.ts (the
equal-range carve-out is handled by `rangeStrictlyContains` +
`rangesEqual` in `canParentScope`) and `moduleAwareContains` in
scope-extractor.ts (`pass1BuildScopes` calls `canParentScope` directly).

* fix(windows): replace 32767-char truncation with chunked-input parsing

The tree-sitter 0.21.x Node binding crashes (SIGSEGV) on Windows when
parser.parse(string, ...) is handed a JS string longer than 32 767 chars.
The crash is in the bindings V8 string-to-buffer conversion and cannot
be intercepted from JS. Previous mitigation truncated source at the last
newline before that boundary, silently losing the file tail and producing
ERROR-root trees from mid-class cuts.

Switch to the callback (Parser.Input) overload via a new parseSourceSafe
helper. tree-sitter pulls source in 16 KiB chunks via repeated callback
invocations, bypassing the broken conversion path. Files are parsed in
full, no data loss, no platform-specific code path.

Removes the now-unnecessary ERROR-root short-circuit in csharp/captures.ts
and the empty-captures shim in scope-extractor-bridge.ts; both existed only
to swallow truncation-induced parse failures.

* fix(windows): cover all parse sites and correct vector-extension state

Address adversarial review on PR #1433:

1. Extend parseSourceSafe to all remaining parser.parse() call sites that
   handle full file content. The first commit only converted the four
   sites with active truncation hacks; cache-miss paths in
   call-processor (x2), heritage-processor (x2), import-processor, and
   the Go/Python/TypeScript captures + Go range-binding still called
   parser.parse() directly. On Windows those would still SIGSEGV for
   files > 32767 chars.

2. Stop setting vectorExtensionLoaded = true on the win32 short-circuit
   in lbug-adapter.ts. The flag means "successfully loaded" and is
   checked by an early-return at the top of loadVectorExtension; setting
   it on the skip path made the second call return true and let
   QUERY_VECTOR_INDEX run against a DB without the extension.

3. Drop the placeholder issues/... URL in the same comment.

4. Add unit tests for parseSourceSafe at boundary values: 16 KiB
   (direct/callback boundary), the 32 767 Windows crash boundary,
   single-line > chunk size, CRLF near boundary, and large all-Chinese
   source. Confirms the callback path is correct for non-ASCII content,
   which is also exercised by the existing csharp-captures large-file
   test.

Researched the chunking concern: tree-sitter Node binding sets
TSInputEncodingUTF16 and divides byte_index by 2 in ByteCountToJS before
calling the JS callback, so the index argument is a UTF-16 code-unit
offset — matching String.prototype.slice. Splitting tokens across chunks
is safe by API contract; the lexer is chunk-agnostic.

* fix(windows): extend parseSourceSafe to group/embeddings + lint enforcement

Closes the remaining Windows SIGSEGV exposure flagged by the Codex
adversarial review on PR #1433. Six pre-existing parser.parse(content)
call sites bypassed parseSourceSafe and could crash the process on
Windows when a contract IDL, route file, or embedding-target source
exceeded 32 767 chars. Adds a lint rule so the regression vector closes
permanently.

Production code:
- Relocate parseSourceSafe from ingestion/utils/ to core/tree-sitter/
  so group/ and embeddings/ can import without crossing into ingestion
  internals. core/tree-sitter/ already houses parser-loader.ts and is
  the natural shared facade. All 11 existing importers updated; no shim
  left behind in the old location.
- Route through parseSourceSafe in 5 group extractors (grpc, thrift,
  http-route, include, tree-sitter-scanner) and the embeddings
  ensureAndParse helper.
- The seventh direct .parse() call in grpc-patterns/proto.ts:49 is a
  module-load grammar smoke test parsing a 36-char literal. Trivially
  safe by inspection, intentionally direct, filtered out by the lint
  rule via the string-literal-arg skip.

Tests:
- 5 caller-side regression tests with a vi.spyOn assertion on
  parseSourceSafe. The spy is what catches a regression: parser.parse
  on a 40 000-char input succeeds on Linux/macOS, so a "no throw"
  assertion alone would silently pass with the bypass reintroduced.
- The vi.mock boilerplate is centralised in
  gitnexus/test/helpers/parse-source-safe-mock.ts, dynamic-imported
  inside each mock factory so vitest's hoister does not race the
  static import binding.

Lint:
- New custom ESLint rule gitnexus/require-safe-parse, scoped to
  gitnexus/src/core/**, fails on direct <parser>.parse(<non-literal>,
  ...) calls and auto-fixes them to parseSourceSafe(<parser>, ...).
  Skips JSON/URL/marked/Number/Math, string-literal first args
  (smoke tests), test files, and the helper itself. Auto-fix rewrites
  the call site only; the developer adds the import after tsc
  surfaces the missing identifier — same tradeoff as
  unused-imports/no-unused-imports.

Plan: docs/plans/2026-05-10-001-fix-windows-parse-safety-group-and-embeddings-plan.md

* fix(test): use mkdtempSync in http-route-extractor regression test

Address CodeQL js/insecure-temporary-file warning on the new Windows-
SIGSEGV regression test. The test was using path.join(tmpDir, "large-input")
which, when nested inside a Date.now()-based parent tmpDir, lets CodeQL flag
the directory as a predictable-name temp file with race-condition risk.
Switch to fs.mkdtempSync(path.join(tmpDir, "large-input-")) so the suffix
is a secure unique random string.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-10 16:00:36 +01:00
2620b704e0 feat(cursor): upgrade hooks to Cursor 2.4 postToolUse for Read/Grep/Shell coverage (#1467)
* feat(cursor): upgrade hooks to Cursor 2.4 postToolUse for Read/Grep/Shell coverage

Cursor 2.4 (released 2026-01-22) shipped generic preToolUse/postToolUse hooks
matching `Shell|Read|Write|Grep|Delete|Task|MCP:<tool>`, replacing the
2.3-era beforeShellExecution hook that only fired on shell commands. The
existing integration only intercepted the shell path, so Cursor users got
graph augmentation roughly 10% as often as Claude Code users — only when
the agent dropped to rg/grep instead of using its native Read/Grep tools.

This swaps the integration over to postToolUse and ports the bash+jq
hook script to cross-platform Node:

- gitnexus-cursor-integration/hooks/hooks.json: registers a single
  postToolUse hook matching Shell|Read|Grep that invokes the new
  gitnexus-hook.cjs.
- gitnexus-cursor-integration/hooks/gitnexus-hook.cjs: new Node hook
  mirroring the safety patterns from the Claude hook (absolute-cwd
  validation, .gitnexus discovery with linked-worktree fallback,
  npx.cmd on Windows, end-of-options `--` marker, debug truncation,
  graceful failure). Extracts the search pattern per tool kind:
  Grep -> toolInput.query; Read -> file basename stripped to identifier
  chars; Shell -> existing rg/grep arg parser. Emits Cursor-shape
  `{ "additional_context": "..." }` on stdout — no shell, no jq.
- gitnexus-cursor-integration/hooks/augment-shell.sh: removed (Windows
  incompatible, narrower coverage).
- gitnexus/test/unit/cursor-hook.test.ts: 33 regression tests covering
  manifest wiring, source-level invariants (no shell:true, npx.cmd,
  isAbsolute, additional_context output shape, end-of-options marker),
  extractPattern coverage per tool, and behavioral early-exit paths
  (empty/invalid stdin, relative cwd, no .gitnexus, unknown tool name,
  short patterns, non-search shell commands, case-insensitive matching).
- README.md / gitnexus/README.md: editor-support table now lists Cursor
  as Full / hooks=Yes (postToolUse), matching reality.
- gitnexus/src/cli/augment.ts and gitnexus/src/core/augmentation/engine.ts:
  doc-strings updated from `Cursor beforeShellExecution` to
  `Cursor postToolUse`.

Closes #1466.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(cursor): hook timeout is in seconds, not milliseconds

Cursor's `timeout` field in hooks.json is in seconds (per
https://cursor.com/docs/agent/hooks and the original integration's
`"timeout": 5`). I'd written `10000` after blindly copying the issue
body's example — that resolves to ~2.8 hours, not 10 seconds. If the
script ever hangs before reaching its inner spawnSync timeouts (e.g.
during stdin read), Cursor would have waited that long before killing
it.

Drop to `10` (seconds), matching the Claude plugin's hooks.json and
giving plenty of headroom over the inner 7s augment-CLI timeout.

Add a regression-guard assertion in cursor-hook.test.ts so a future
ms/s mixup fails fast.

Reported by Cursor Bugbot on PR #1467.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(cursor): address Claude review findings — payload aliases, debug, install docs

Resolves three findings from Claude reviewer on PR #1467:

1. Cursor payload field-name uncertainty (SIGNIFICANT)
   Claude flagged that the Grep `query` field is an unverified assumption
   per Cursor 2.4 docs (https://cursor.com/docs/agent/hooks). Mitigated:
   - Expanded Grep aliases: query | pattern | regex | q | search | searchQuery
   - Added pickLongestStringValue() last-resort fallback so the hook
     extracts *something* even if Cursor renames every documented field
   - Added GITNEXUS_DEBUG=1 stderr logging of the raw stdin payload so
     users can capture Cursor's actual contract when diagnosing silent
     no-ops, and report it back if aliases drift
   - Added Read alias `filePath` (camelCase variant alongside `file_path`)
   - Inline comment block citing the docs URL and the uncertainty

2. Hook command path resolution + install docs (SIGNIFICANT)
   Claude flagged `node ./hooks/gitnexus-hook.cjs` as relative without
   documented install path. Added gitnexus-cursor-integration/README.md
   with explicit install steps:
   - .cursor/hooks.json + hooks/gitnexus-hook.cjs at project root
   - Confirms Cursor's project-root CWD convention with doc link
   - Verify steps including GITNEXUS_DEBUG capture
   - Pattern-extraction contract table per tool
   - Troubleshooting: not-firing, npx fallback, wrong-pattern diagnosis

3. README "Full" overclaim for Cursor (MODERATE)
   Both README rows now read `Yes (postToolUse, manual install)` linking
   to the new install README, accurately signaling that hooks aren't
   automated by `gitnexus setup` like they are for Claude Code.

4. Shell quoted-pattern parser limitation (MINOR, documented)
   Added inline comment in gitnexus-hook.cjs documenting the known
   `rg "User Service"` -> `User` truncation, plus regression tests in
   cursor-hook.test.ts pinning the behavior so a future change is
   visible.

Test additions (33 -> 41):
- Wide-alias source coverage for Grep (query / pattern / regex / q /
  search / searchQuery) plus pickLongestStringValue fallback
- Read alias coverage including camelCase filePath
- GITNEXUS_DEBUG behavioral test: stderr quiet by default, payload
  echoed when env var set, stdout output contract preserved either way
- Shell quoted-pattern documented behavior tests
- Install README presence + content (.cursor/hooks.json, hooks/, debug
  diagnostics)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-10 13:29:06 +01:00
Gergő Magyar 5d670a530d ci(release): skip rc build on release PRs (#1474)
* ci(release): skip rc build on release PRs

Suppress the auto-fired Release Candidate workflow when:
  1. The HEAD commit subject matches `chore: release vX.Y.Z` (the canonical
     release-PR title), or
  2. The squash-merged PR carries the `release` label.

Either match short-circuits the guard to should_run=false. This prevents the
rc cycle from racing publish.yml on the v-tag (as happened on v1.6.4 where
we had to manually cancel the auto-fired RC run after merging PR #1473).

Adds pull-requests: read to the guard job for the label lookup. A failed
gh API call falls through to the existing dedup logic rather than silently
suppressing rc builds.

* ci(release): address PR #1474 review — anchor regex + sanitise log echo

Two minor follow-ups from Claude's review:

1. End-anchor the release-subject regex. The previous shape
   ^chore: release vX.Y.Z would match noisy variants like
   chore: release v1.0.0 (something unrelated). The new shape
   requires either the bare title or the canonical squash-merge
   (#NNNN) suffix exactly.

2. Sanitise HEAD_SUBJECT before echoing to logs. git %s strips
   newlines so LF injection is impossible, but a hypothetical
   subject containing ::error:: or ::set-output:: could otherwise
   forge GitHub Actions annotation entries. Defence-in-depth.

Both findings flagged minor / does not block merge — applying
anyway since they are trivial.
2026-05-10 09:50:58 +01:00
Gergő Magyar 4848dce9ea test(u8): de-flake regex linearity assertions (#1475)
* test(u8): de-flake regex linearity assertions

The single-trial 2x input + 3x ratio bound was razor-thin: a real macOS
CI run failed at ratio 3.01x with small=7.41ms / large=22.31ms - both
above the 5ms noise floor but close enough that single-shot scheduler
jitter pushed the ratio over.

Replace the methodology with four stacked techniques:
  1. Warmup runs before timing (let the JIT tier up)
  2. Median of 5 trials per measurement (eliminates GC + jitter)
  3. 4x input ratio (was 2x) - linear gives ~4x, O(n^2) gives ~16x
  4. 8x ratio bound with a 20ms noise floor on the LARGE measurement

Headroom: linear is expected at ~4x, bound is 8x = 2x safety margin.
A real O(n^2) regression on a 4x input would clock 16x, well outside.
Catastrophic backtracking is still caught by the absolute <500ms cap.

Verified: 10 consecutive local runs all passed.

* test(u8): address PR #1475 review — tighten floor + rename for accuracy

Two follow-ups from Claude's review:

1. Floor semantics: revert to 'skip when BOTH measurements below floor'
   (AND, not single-check) and lower threshold from 20ms back to 5ms.
   Median-of-5 makes 5ms reliably resolvable above performance.now()'s
   ~10-100us band, so the higher floor was unnecessary defense.
   Closes the gap where an O(n^2) regression on a fast runner could
   stay under 500ms AND below 20ms-large to escape both detectors.

2. Rename assertSubLinearRatio -> assertNearLinearScaling. The bound
   is SIZE_RATIO * 2 = 8x on a 4x input = sub-quadratic with 2x
   headroom over linear, not strict sub-linearity. New name reflects
   the actual semantics.
2026-05-10 09:31:47 +01:00
Gergő Magyar 936a3bed6f chore: release v1.6.4 (#1473)
* chore: release v1.6.4

* chore: expand v1.6.4 CHANGELOG with additional crash-issue closes
2026-05-10 08:33:47 +01:00
dependabot[bot] 874d2aa95b chore(deps)(deps): bump langchain from 1.3.4 to 1.3.5 in /gitnexus-web (#1462) 2026-05-10 06:15:27 +01:00
dependabot[bot] 262ad8b859 chore(deps)(deps-dev): bump @types/dompurify in /gitnexus-web (#1461) 2026-05-10 05:55:44 +01:00
dependabot[bot]andGergő Magyar a2517e050d chore(deps)(deps): bump lucide-react in /gitnexus-web (#1460)
Bumps [lucide-react](https://github.com/lucide-icons/lucide/tree/HEAD/packages/lucide-react) from 1.11.0 to 1.14.0.
- [Release notes](https://github.com/lucide-icons/lucide/releases)
- [Commits](https://github.com/lucide-icons/lucide/commits/1.14.0/packages/lucide-react)

---
updated-dependencies:
- dependency-name: lucide-react
  dependency-version: 1.14.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-09 19:51:12 +01:00
dependabot[bot]andGergő Magyar a5c582f547 chore(deps): bump actions/checkout from 5.0.0 to 6.0.2 (#1459)
Bumps [actions/checkout](https://github.com/actions/checkout) from 5.0.0 to 6.0.2.
- [Release notes](https://github.com/actions/checkout/releases)
- [Changelog](https://github.com/actions/checkout/blob/main/CHANGELOG.md)
- [Commits](https://github.com/actions/checkout/compare/v5...de0fac2e4500dabe0009e67214ff5f5447ce83dd)

---
updated-dependencies:
- dependency-name: actions/checkout
  dependency-version: 6.0.2
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-09 19:26:53 +01:00
CopilotandGergő Magyar 6f1cfffdd7 fix(security): Harden CI permissions (#1454)
* Initial plan

* chore(security): harden workflow permissions and pin Docker base image digests

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2ddc8f2b-7355-48cf-9a0b-c06df66c3f47

* fix(security): restore permissions: {} on publish + release-candidate workflows

These two release-publishing workflows had permissions: {} (the strictest valid form) before PR #1454, which replaced it with permissions: read-all. Every job in both files already declares its own permissions block, so the workflow-level default is only the safety net for future jobs added without one — read-all weakens that net for no benefit. Restore {} and the explanatory comment.

Scorecard's TokenPermissions check accepts both forms, so this preserves U9 compliance.

* fix(security): narrow permissions: read-all to contents: read on 13 workflows

PR #1454 added permissions: read-all to 13 workflows that previously had no top-level permissions block. read-all is Scorecard-compliant but unnecessarily broad — every job in scope only needs contents:read at the workflow level (job-level blocks already grant the writes that any job actually performs).

Snapshot of every job in the 13 workflows confirms contents:read is sufficient:

- ci.yml: quality/tests/scope-parity have explicit contents:read job blocks; save-pr-meta uses upload-artifact only (no token scopes needed); ci-status is pure shell.
- ci-e2e.yml, ci-quality.yml, ci-scope-parity.yml, ci-tests.yml: all jobs do checkout + npm + tsc/vitest/playwright/upload-artifact only; no API token scopes required.
- claude.yml, codeql.yml, dependency-review.yml, docker.yml, gitleaks.yml, pr-labeler.yml, trivy.yml, workflow-lint.yml: all jobs already declare their own job-level blocks (security-events:write, pull-requests:write, packages:write, etc.) so the workflow-level default does not gate them.

zizmor (--min-severity high) is clean on the resulting tree. Pre-existing medium findings (secrets-inherit, artipacked) are in unrelated workflows and untouched by this commit.

scorecard.yml also uses read-all but pre-existed PR #1454 and is deferred to a follow-up PR per the plan's scope boundary.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-09 17:58:22 +01:00
666041d608 fix(security): log-injection, http-to-file-access, client-side-request-forgery (#1456)
* fix(security): U11 log-injection, http-to-file-access, client-side-request-forgery

U11.1: Add validateLLMBaseUrl() in llm-client.ts; called at the top of
callLLM() to reject non-http/https schemes and http:// to non-loopback
hosts before any fetch that writes LLM output to disk.

U11.2: Strip CRLF from groupDir in bridge-db.ts openBridgeDbReadOnly
before logging (defence-in-depth on top of pino's JSON escaping).

U11.3: Replace console.log with logger.debug and sanitize normalizedName
/ job.id in api.ts resolveRepo to close js/log-injection alerts.

U11.4: Add validateBackendUrl() in backend-client.ts; called inside
setBackendUrl() to reject non-http/https schemes before the URL is
stored as a fetch target, closing js/client-side-request-forgery alerts.

U11.5: Tests added:
- wiki-llm-client.test.ts: validateLLMBaseUrl happy/error paths
- server-connection.test.ts: validateBackendUrl and setBackendUrl
  rejection paths

All new tests pass (30/30 wiki-llm-client, 18/18 server-connection,
30/30 bridge-db).

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0452a6ce-711f-4203-9ae6-5dd0b77fb157

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: correct IPv6 loopback check in validateLLMBaseUrl

Node's URL parser preserves brackets in hostname for IPv6 addresses
(e.g. http://[::1]:11434 yields hostname '[::1]'), so strip them
before comparing against '::1'. Add a test to cover this case.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0452a6ce-711f-4203-9ae6-5dd0b77fb157

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: also sanitize error message in bridge-db log call

Sanitize lastErr.message (which may contain a file path from ENOENT
errors) alongside groupDir to prevent CRLF injection from error
message content. Addressed code review feedback.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0452a6ce-711f-4203-9ae6-5dd0b77fb157

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: address security review findings — credential hygiene and test coverage

[LOW] Redact credentials from URL validation error messages:
- validateLLMBaseUrl: malformed URL no longer echoes raw input;
  scheme error shows protocol only; http-non-loopback error uses
  parsed.origin (scheme+host+port) instead of full URL
- validateBackendUrl: same treatment — no raw input in any error path

[INFO] Add state-preservation test for setBackendUrl:
- Proves _backendUrl is unchanged after a rejected call, covering the
  validation-before-assignment ordering.

[INFO] Expand validateLLMBaseUrl adversarial test coverage:
- LOCALHOST uppercase (case-fold path)
- RFC 1918 / IMDS IPs (10.x, 169.254.x)
- Hostname-spoofing (localhost.evil.com, 127.0.0.1.evil.com, localhost.)
- Non-loopback IPv6 (fe80::1, ::ffff:127.0.0.1)
- ftp:// scheme
- Credential-hygiene assertion (sk-secret not in error message)

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/7bb18fa2-3e66-4fe0-949f-6d493fbd351b

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* style: prettier autoformat U11 security fix files

Fixes the failing 'quality / format' check on PR #1456 by running 'prettier --write' over the 6 files touched by the security fix. Formatting only — no logic change.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-05-09 17:26:32 +01:00
e02c56f653 fix(security): Pin Docker Node base images, remove runtime package-manager CVE surface, verify Trivy on PRs, and harden Dependabot policy (#1455)
* fix: pin Docker node base images and remediate bundled npm CVEs

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e0605c79-296e-4b3a-b6c3-4ad375950935

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: run trivy on docker PR changes and remove corepack

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/4d714047-4fc1-4af1-9734-91400a15568f

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore: add docker digest updates and normalize dockerfile comments

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/7d980908-a823-4c28-b074-9134ec672e84

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: add dependabot cooldown policies

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a8531b8d-384b-4c54-84dd-a98b31993c44

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: remove unsupported dependabot cooldown keys

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/166df50e-c2fe-4d7f-ab41-e94c703338f6

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(node): bump CI + engines to Node 22; centralize NPM_VERSION via build ARG

Closes the LOW findings from Claude Final Re-Review on PR #1455:

- Bump engines.node to >=22.0.0 and align all CI workflows (ci-quality,
  pr-autofix, publish, release-candidate) and the composite setup
  actions on Node 22. Node 20 reached EOL on 2026-04-30; the test
  Docker image was already on 22.
- Centralize the bootstrapped npm version in a single ARG NPM_VERSION
  per Dockerfile (cli, web, gitnexus/Dockerfile.test) so a security
  bump only requires updating one default per file with a clear
  cross-reference comment.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-05-09 16:55:31 +01:00
Gergő Magyar 6906be3695 feat(autofix): replace inline reviewdog with /autofix ChatOps button (#1458)
* fix(autofix): verify reviewdog actually posted before claiming "click Apply"

The sticky summary comment was stating "Posted formatting suggestions
inline. Click Apply suggestion on each" even when reviewdog landed zero
inline review comments — typical case: the formatter touched lines
outside the PR's added range, so `-filter-mode=added` (correctly)
filtered everything out. The script unconditionally set `posted=true`
after running reviewdog regardless of whether any comments were
actually created, leaving the user staring at a sticky that promised
buttons that didn't exist.

The publish job now snapshots the count of `github-actions[bot]` review
comments before and after reviewdog. If the delta is zero, surface a
new `diff-no-overlap` UI state that tells the user plainly:

  "Formatter found fixable issues, but they're on lines outside this
  PR's added range — there's nothing to click here. Run locally:
  npm run lint:fix && npm run format."

Plus a matching `gitnexus/autofix` Check Run conclusion (still neutral,
distinct title) so agents reading `gh pr checks` see the same signal.

Three states are now machine-distinguishable in the sticky's
gitnexus-autofix JSON block: suggestions-posted (delta > 0),
diff-no-overlap (delta == 0), skipped-too-large (>3k lines).

* feat(autofix): replace inline reviewdog with /autofix ChatOps button

Pivot the PR autofix UX from per-line reviewdog suggestions to a single
slash-command button. Contributors comment `/autofix` on the PR; a new
trusted workflow downloads the existing autofix patch artifact, applies
it to the PR head, and pushes a commit back.

Why:
- 3K+ diffs hit GitHub's review-comment API 406 limit -> dead end.
- Diffs where the formatter touches lines outside the PR's added range
  ("no-overlap") get filtered by reviewdog's -filter-mode=added -> dead
  end (PR #1457 patched the lying sticky but the underlying UX gap
  remained).
- Per-line click-Apply-suggestion is high-friction for big diffs and
  easy to apply unevenly.
- A single `git apply` + push works at any size and lands fixes
  atomically.

Changes:
- pr-autofix-publish.yml: remove `Install reviewdog` and
  `Post inline suggestions` steps. Collapse three sticky states
  (suggestions-posted, diff-no-overlap, skipped-too-large) into one
  (fixes-available). Bump JSON schema v1 -> v2 with `apply_command`
  field; all v1 fields preserved.
- pr-autofix-apply.yml (new): triggers on issue_comment with body
  `/autofix`, validates body via strict regex, validates commenter
  has write/admin/maintain or is the PR author, locates latest
  successful pr-autofix run for PR head SHA, downloads artifact,
  applies patch, pushes commit. Reacts +1/-1/eyes on triggering
  comment per outcome. Idempotent (`git apply --check --reverse`
  detects already-applied state).
- CONTRIBUTING.md: document v2 schema and the /autofix flow,
  including the maintainer-edit requirement for fork PR pushes.

Trust posture: apply workflow runs from default-branch code only,
under issue_comment trigger. Comment body and author login flow
through env vars and pattern-matched, never interpolated into shell.
Permission gate (write/admin/maintain OR PR author) before any
artifact fetch. Fork PRs require "Allow edits by maintainers"
(GitHub-native; we don't bypass).

Net YAML: -139 lines in publish.yml, +260 in apply.yml. Removes
reviewdog binary pin and the entire review-comment API surface.

* fix(autofix): address Codex adversarial findings on PR #1458

Two findings from the Codex adversarial review of the autofix ChatOps
pivot. Both are localized YAML changes that close trust gaps the pivot
inherited from the original PR #1446 design.

U1 — Cross-verify metadata against workflow_run authority
(.github/workflows/pr-autofix-publish.yml):
  Previously the trusted publisher accepted pr_number, head_sha, and
  head_repo from metadata.json after only an allowlist regex. A
  fork-controlled `npm run lint:fix` could have written a syntactically
  valid metadata.json referencing another PR/SHA, redirecting the
  write-scoped sticky/check-run onto an attacker-chosen target.

  New `Verify metadata against workflow_run authority` step compares
  artifact-claimed identity against:
    - github.event.workflow_run.head_sha
    - github.event.workflow_run.head_repository.full_name
    - workflow_run.pull_requests[].number (within-repo PRs)
    - gh api commits/{sha}/pulls fallback (fork PRs, where
      pull_requests[] is empty)
  Fail closed on mismatch — no sticky, no check-run, no override.

U2 — Lease-protected push in apply workflow
(.github/workflows/pr-autofix-apply.yml):
  Previously the apply step pushed `HEAD:${HEAD_REF}` plain. A force-
  push between resolve (Step 5) and push (Step 9) would silently
  fast-forward an older commit graph over the contributor's newer
  state.

  Push now uses `--force-with-lease=refs/heads/${HEAD_REF}:${HEAD_SHA}`
  against the SHA resolved earlier. Distinct `lease-failed` result code
  + retry-message reply, separated from `push-failed` (fork without
  maintainer-edit) so contributors can diagnose the actual cause.

Plan: docs/plans/2026-05-09-005-fix-autofix-codex-adversarial-findings-plan.md
(local-only per repo convention).

Trust posture preserved: no new permissions, no new workflows, no
contract change. JSON v2 schema unchanged. CodeQL js/server-side-
request-forgery and template-injection posture unchanged — all new
inputs flow via env vars and pattern-matched.

* fix(autofix): close zizmor credential-persistence finding on apply checkout

actions/checkout's default behavior writes the GITHUB_TOKEN into
.git/config as an extraheader. The token then sits on disk in the
checkout directory — an actions/upload-artifact step on that
directory would leak it. We don't upload, but zizmor's
credential-persistence lint correctly flags the latent risk.

Set persist-credentials: false on the Checkout PR head step. Provide
push auth inline via `git -c http.extraheader="Authorization: Basic
<base64-of-x-access-token:TOKEN>"` so the credential never lands on
disk and never appears in process listings (the URL form
https://x-access-token:TOKEN@… is rejected here because it leaks via
ps and git remote -v).

Push lease semantics from U2 unchanged — same --force-with-lease
against the resolved HEAD_SHA, same lease-failed/push-failed/stale
result codes.

* fix(review): apply autofix feedback

ce-code-review surfaced 15 findings on PR #1458; this commit applies
the 7 with concrete fixes (#1, #2, #3, #4, #5, #9, #13). Five P2
findings (#6, #7, #8, #10, #12) are recorded as residual actionable
work for follow-up; two advisory items (#11, #14) skipped.

#1 — applied_run_id schema drift (CONTRIBUTING.md):
  v2 docs claimed `state: applied` enum value and an `applied_run_id`
  field that no code path emits. Trimmed docs to match what the
  workflow actually writes (state: fixes-available; v1 field set as
  superset). Implementing the apply-side sticky upsert that would
  populate `applied_run_id` is deferred — cleaner than carrying a
  contract claim with no code.

#2 — result= unset between idempotency probe and lease push
(pr-autofix-apply.yml):
  After `git apply --check` passed, an early non-zero exit from
  `git config` / `git apply` / `git add` / `git commit` left
  `result=` unset, sending the user to the `*` "unexpected state
  (`unknown`)" arm. Wrapped the apply/commit phase in a single
  if-test that sets `result=apply-failed` on any failure. New
  React-and-reply branch surfaces an actionable message.

#3 — permission lookup conflated transient API failures with denial
(pr-autofix-apply.yml):
  `gh api … 2>/dev/null || echo "none"` swallowed 5xx, 429 secondary
  rate-limit, and network failures, surfacing them as a public 👎
  refusal to legitimate maintainers. Now distinguishes 404
  (genuine non-collaborator) from other API failures via stderr
  match. New `allowed=api-failed` state triggers a 😕 reaction with
  a "transient API failure, retry" reply instead of a misleading
  refusal.

#4 — lease-failure grep missed git's "remote rejected" / branch-
deleted phrasings (pr-autofix-apply.yml):
  Real lease failures got classified as `push-failed` →
  user told to enable maintainer-edit, which won't help. Expanded
  regex to match `remote rejected` and `! [rejected]`.

#5 — broken bullet continuation in CONTRIBUTING.md release-candidate
section: rejoined the split bullet so it renders correctly.

#9 — base64 GITHUB_TOKEN bypassed GitHub's secret-masker
(pr-autofix-apply.yml):
  Added `::add-mask::${auth_header}` immediately after construction
  so any subsequent log line (set -x, GIT_TRACE) gets *** redacted.

#13 — misleading schema-bump comment in pr-autofix-publish.yml:
  Comment claimed all v1 fields preserved exactly, but the `state`
  enum was redefined v1→v2. Updated to make the migration path
  explicit (v1 readers see unfamiliar schema, fall back to prose).

Residual actionable work (deferred to follow-up):
  #6 locate step gh api retry; #7 artifact-expired graceful fallback;
  #8 re-entrancy comment-spam guard; #10 producer-still-running UX;
  #12 gh_retry wrapper for apply.yml.

Validations: yaml.safe_load OK, check-workflow-concurrency.py OK.

* fix(autofix): apply remaining ce-code-review residual findings (#6, #7, #8, #10, #12)

Pulls the deferred items from the previous review pass into this PR so
the workflow ships with full reliability + UX coverage rather than
follow-up debt.

#6 + #12 — gh_retry wrapper on idempotent GETs in apply.yml:
  Permission lookup, PR metadata fetch, and workflow-run lookup are now
  wrapped in the same gh_retry helper publish.yml uses (3 attempts,
  linear backoff). Reaction/comment POSTs remain unwrapped (retrying
  POST would dupe the resource).

#10 — producer-still-running UX:
  The locate step now distinguishes three cases via `found_status`
  output: success (proceed), in-progress / queued / pending / waiting
  (reply ⏳ "wait for autofix run to finish"), not-found (reply 🤔
  "push a commit"), api-failed (reply ⚠️ "transient API failure"). The
  "no successful autofix run" message no longer fires immediately after
  a fresh push while the producer is still mid-run.

#7 — artifact-expired graceful fallback:
  actions/download-artifact gains `continue-on-error: true`. The apply
  step distinguishes patch-file-missing (artifact expired, 1-day
  retention elapsed) from patch-file-zero-bytes (formatter found
  nothing). New `result=artifact-expired` case + ⏳ "push a new commit
  to regenerate" reply.

#8 — re-entrancy loop guard:
  After checkout but before applying, check if HEAD itself is a
  github-actions[bot] `chore(autofix)` commit. If so, refuse to
  re-apply (`result=loop-prevented`) with a 🔁 reply telling the user
  to push a human-authored commit or revert before retrying. Prevents
  formatter-config-drift loops where an automated agent watching the
  sticky could pump arbitrary apply commits.

Net effect: every code path in apply.yml now sets a meaningful `result=`
that maps to a specific user-facing reaction + reply. The `*` "unexpected
state (unknown)" arm becomes truly unreachable in normal operation.

Validations: yaml.safe_load OK, check-workflow-concurrency.py OK.

* fix(autofix): refresh stale reviewdog comments + reject patches touching .github/

Two follow-up findings on PR #1458:

#1 — Stale reviewdog references in workflow header comments:
  pr-autofix-publish.yml's header still described the removed inline-
  suggestion path ("posts inline review-comment suggestions to the PR
  using `reviewdog`", "Reviewdog reporter: github-pr-review reads
  $REVIEWDOG_GITHUB_API_TOKEN…"). The Check Run permissions comment
  enumerated the old outcomes (clean / suggestions-posted /
  skipped-too-large) instead of the current set (clean / fixes-
  available). pr-autofix.yml's header described the trusted job as
  posting "inline review-comment suggestions" and the changed_lines
  comment referenced the dead 3000-line cap. Refreshed all three to
  describe the actual sticky + Check Run + /autofix flow.

#2 — Reject patches touching .github/ (sensitive-paths guard):
  Theoretical supply-chain vector: a malicious PR could ship a custom
  prettier/ESLint config that reformats workflow YAML, dependabot.yml,
  or CODEOWNERS. The producer would capture those edits in
  autofix.patch; a maintainer running `/autofix` would push them under
  `contents: write` without human review. The default GITHUB_TOKEN
  lacks the `workflows` scope so workflow-file pushes would fail at
  the platform layer anyway, but as a generic `push-failed` (which
  misleads users into enabling maintainer-edit). Reject early with
  a specific reason.

  Match runs against the patch with grep on `^(diff --git|---|+++)
  [ab]?/?\.github/`. New `result=sensitive-paths` case + 🛑 reply
  telling the user to apply .github/ formatter changes manually.

  Documented the constraint in CONTRIBUTING.md under the /autofix
  section so contributors aren't surprised when the workflow refuses
  a patch that includes formatter changes to workflow files.

Validations: yaml.safe_load OK, check-workflow-concurrency.py OK.
2026-05-09 16:32:38 +01:00
Gergő Magyar 152a0506c9 feat: shared resilient-fetch (retries + circuit breaker) (#1448)
* feat: shared resilient-fetch (retries + circuit breaker)

Add a small, runtime-agnostic resilience layer in gitnexus-shared and
migrate every backend HTTP outbound call (CLI, MCP, wiki LLM, web → backend)
through it.

Helpers (gitnexus-shared/src/integrations/):

- retry.ts            — withRetry(fn, opts) with caller-supplied
                        retryability classification and full-jitter
                        exponential backoff.
- circuit-breaker.ts  — closed/open/half-open per-process breaker with
                        injectable clock, plus a keyed registry so
                        callers targeting the same endpoint share state.
- resilient-fetch.ts  — composed wrapper: retries 5xx + 429 + retryable
                        network throws, treats AbortSignal.timeout()
                        and 4xx (other than 429) as terminal, honors
                        Retry-After (capped at 30s), throws
                        CircuitOpenError when the breaker opens.

Migrations (no behaviour regression — all existing tests pass):

- gitnexus/src/core/embeddings/http-client.ts (covers analyze + MCP
  query path) — replaces inline linear-backoff retry.
- gitnexus/src/core/wiki/llm-client.ts — preserves Azure content-filter
  branch; resilientFetch handles 5xx/429.
- gitnexus-web/src/services/backend-client.ts (fetchWithTimeout helper)
  — small retry budget (2 attempts, 250–1500 ms) so a dead local
  backend still fails fast for the user.
- gitnexus-web/src/core/llm/settings-service.ts (OpenRouter model list).

Deliberately not migrated:

- gitnexus-web/src/services/backend-client.ts streamJob() — Server-Sent
  Events stream; the existing reconnect-with-Last-Event-ID logic is
  not unary-fetch shaped.
- gitnexus-web/src/components/SettingsPanel.tsx checkOllamaStatus() —
  one-shot health probe; retrying delays the "Ollama not running"
  error rather than improving UX.

41 new helper tests cover backoff math, breaker state transitions,
Retry-After parsing (delta-seconds + HTTP-date), 401/422 terminal
classification, and breaker fail-fast on three exhausted retry batches.

* fix(review): apply autofix feedback

Address Claude's two MEDIUM blocking findings on PR #1448 plus the
CodeQL SSRF false-positive flag.

- backend-client `fetchWithTimeout` now uses `AbortSignal.timeout()`
  merged with the caller's signal via `AbortSignal.any()`. Timer-fired
  aborts surface as `DOMException(name='TimeoutError')` so
  resilientFetch routes them through the terminal-network branch
  (no retry, no breaker hit), instead of incrementing the breaker
  for user-side network slowness.
- Method-aware retry budget in `fetchWithTimeout`: idempotent verbs
  (GET/HEAD/OPTIONS) keep the 2-attempt budget; POST/PATCH/PUT/DELETE
  default to single-attempt so a 5xx on `startAnalyze` cannot start
  a duplicate job. New `forceRetry` parameter for callers that
  know-idempotent mutations (e.g. DELETE of a known-deleted resource).
- `resilient-fetch.ts` carries a documented suppression for CodeQL
  js/server-side-request-forgery on the inner fetch call. Every
  concrete caller passes a hardcoded URL constant or a value from
  configuration (env vars, saved settings); user request input never
  flows into the URL parameter.
- New test file `backend-client-retry.test.ts` covers all three
  paths: GET retries on 503, POST does not retry, timeout does not
  increment the breaker.

* fix(resilient-fetch): address Codex adversarial findings

Closes the three blocking issues from Codex's review on PR #1448.

U1 — Add `recordNeutral()` to CircuitBreaker.
  Third outcome path that's an explicit no-op for state and the
  consecutive-failure counter. Distinct from `recordSuccess` (closes
  the breaker) and `recordFailure` (may open it). Used for outcomes
  that are neither evidence of backend health nor evidence of
  backend failure.

U2 — Route terminal-client / terminal-network through `recordNeutral`.
  Previously a 401 or local timeout called `recordSuccess`, which
  reset `consecutiveFailures` to 0. A 5xx → 401 → 5xx → 401 → 5xx
  sequence would NEVER trip the breaker because each 4xx in between
  erased the running count. Also classify external `AbortError` as
  terminal-network (was retryable-network), so caller-driven
  cancellation no longer retries against an already-aborted signal
  or counts toward breaker failures on exhaustion.

U3 — Per-origin breaker key in web `fetchWithTimeout`.
  Was hardcoded to `'web-backend'` even though `_backendUrl` is
  mutable via `setBackendUrl`. Switching backend URLs after a
  circuit tripped on host-A would strand the user during the full
  cooldown. Key is now `web-backend:<origin>`, so each backend URL
  gets its own breaker state.

Tests: +5 recordNeutral, +4 resilient-fetch (interleaved 4xx/5xx,
external AbortError, prior-state preservation), +1 web switch-backend
regression. All 70 gitnexus integration tests + 15 web tests green.

* fix(resilient-fetch): tolerate header-less fetch mocks on 429

`classifyOutcome` called `resp.headers.get('Retry-After')` directly,
which crashed when a test stubs `fetch` with a plain object like
`{ ok: false, status: 429 }` (no `headers` field). Real `Response`
always has Headers, so this surfaces only in test setups, but the
helper has no business assuming caller-side correctness on this — the
defensive guard is cheap and a missing `Retry-After` falls through to
exponential-backoff retry like any 429 without the header.

Surfaced by `gitnexus/test/unit/http-embedder.test.ts > retries on
rate limit`, which the embeddings migration exercises against a
plain-object 429 stub. Locked in with a new
`classifies 429 from a header-less fetch mock without throwing` case.

* fix(review): apply autofix feedback

Closes findings from the third multi-agent review pass on PR #1448.

#1 (P1) callLLM had no per-attempt timeout
  Wiki LLM calls passed no `signal` to resilientFetch; each of three
  retry attempts could hang indefinitely on a frozen TCP connection.
  Add `signal: AbortSignal.timeout(60_000)` so the per-attempt budget
  matches what http-client.ts and backend-client.ts already provide.

#2 (P2) drop dead `lastRetryableResp` post-loop fallback
  Variable was set in one switch arm but only read in unreachable code
  after the loop. The retry loop always returns/throws on every
  iteration. Keep only the defensive `throw` so TypeScript's
  control-flow analysis still sees `Promise<Response>` as the return.

#5 (P2) gate test-only exports behind a subpath
  `__resetBreakerRegistry__` and `classifyOutcome` were reachable from
  the main `gitnexus-shared` barrel — production code calling
  `__resetBreakerRegistry__` from a tool implementation would silently
  nuke every circuit breaker process-wide. Move to a new
  `gitnexus-shared/test-helpers` subpath export. Production callers
  see the cleaner public API; tests import via the explicit
  `gitnexus-shared/test-helpers` path.

#6 (P2) exhaustiveness guard on Outcome switch
  Add a `default: const _: never = outcome` arm so a future sixth
  `Outcome.kind` won't compile silently — it'll surface at the switch
  site rather than fall through to a retry/no-retry default.

#9 (P3) document cumulative wall-clock budget
  Add a "Cumulative wall-clock budget" paragraph to resilientFetch's
  JSDoc explaining the worst-case total wait (`maxAttempts × (per-attempt
  timeout + capDelayMs)` ≈ 60s with defaults) and pointing callers at
  outer `AbortSignal.timeout()` when they want a tighter bound.

Deferred to follow-up PRs (per review's Auto-resolve recommendation):
  - #3 idempotency knob to shared API (forceRetry into ResilientFetchOptions)
  - #4 publish.ts migration to resilientFetch
  - #7 parseRetryAfter past-HTTP-date / negative-seconds asymmetry
  - #8 recordNeutral counter time-decay (documented breaker semantic)

* fix(circuit-breaker): gate half-open to a single in-flight probe

Closes the Codex adversarial-review finding on PR #1448 that flagged a
recovery-time thundering herd: when cooldown expired, every concurrent
caller transitioned the breaker to half-open and probed the still-
recovering dependency in lockstep, defeating the breaker's "fail fast"
promise.

U1 — probe-permit gate in CircuitBreaker.check()
  Added a `probeInFlight: boolean` field. After cooldown expires, the
  first `check()` admits the probe and consumes the permit; subsequent
  callers throw `CircuitOpenError` with a configurable
  `halfOpenRetryAfterMs` (default 1000ms) until the probe resolves.

  Critical design point: `recordNeutral` now RELEASES the permit but
  does NOT transition state. Without that split, a single `TimeoutError`
  from per-attempt `AbortSignal.timeout` (which routes through neutral
  classification) would permanently park the breaker in half-open. By
  separating permit-release from state-resolution, we keep the
  "neutral doesn't claim health" semantic without creating that wedge.

  Other changes:
  - `halfOpenRetryAfterMs` is now a constructor option for consumers
    with long-running protected ops (LLM streaming, large uploads).
  - `getState()` is documented as a pure read; the implicit
    Open -> Half-Open transition lives in `check()` only, so tests
    that inspect state never inadvertently consume a probe permit.
  - `isProbeInFlight()` test-only accessor for assertion clarity.
  - JSDoc on `check()` records the JS event-loop atomicity dependency
    and the load-bearing `try/finally` pairing invariant.

U2 — End-to-end concurrency regression through resilientFetch
  Three new scenarios in resilient-fetch.test.ts (26 -> 29):
  - 3 concurrent calls + probe gets 200 -> 1 hits fetch, 2 throw
    CircuitOpenError, breaker closes.
  - 3 concurrent calls + probe gets 503 -> ResilientFetchExhaustedError
    on probe; concurrent callers see halfOpenRetryAfterMs (1000ms);
    fresh caller after probe resolves sees the FULL new cooldown
    (10000ms), not the probe-in-flight default.
  - Probe cancelled mid-flight via AbortError -> permit released,
    state stays half-open, next caller becomes the new probe and
    succeeds.

Plus 9 new circuit-breaker unit tests (16 -> 25) covering the permit
gate, recordNeutral-releases-permit semantic, fresh-cooldown distinction,
default vs configurable halfOpenRetryAfterMs, getState() purity, and
the three-probes-via-neutrals chain.

Total integration test count: 70 -> 82. All 106 gitnexus + 15 web
tests pass; both packages typecheck.

Maintainer decisions (deferred per plan 003 Open Questions):
  - Plan 002's deferral judgement was reversed on Codex's argument
    without new measurement / incident data. The reversal is defensible
    on principle (Hystrix / Resilience4j alignment) but lacks workload-
    driven evidence.
  - Probe-blocked callers throw silently (no log / event hook). R4's
    "no new public API" prevents adding observability; loosen if a
    debug log on probe-blocked is wanted.

* refactor(embeddings): replace bespoke HF breaker with shared CircuitBreaker

Deleted the local `HfDownloadCircuitBreaker` class and the manual
retry loop in `withHfDownloadRetry`. Both are now backed by the
shared `gitnexus-shared` primitives:

- `hfDownloadCircuit` is `new CircuitBreaker({ failureThreshold,
  cooldownMs, key: 'hf-download' })` — same state machine as before
  PLUS the single-permit half-open gate that prevents recovery-time
  stampedes when CLI + MCP embedders concurrently re-load the model.
- `withHfDownloadRetry` delegates the loop to `withRetry` from the
  shared package. Per-attempt timeout (`withDownloadTimeout`),
  network-vs-non-network classification, circuit recording, and the
  `onRetry` callback wire through `withRetry`'s `isRetryable`
  callback.

Behaviour preserved:
- Pre-flight `CIRCUIT_OPEN_TAG` rejection when the breaker is open.
- Mid-loop `CIRCUIT_OPEN_TAG` "opened after N consecutive failures"
  when a network error trips the threshold.
- Non-network errors (e.g. CUDA unavailable) bypass retry and go
  through `recordNeutral` instead of resetting the breaker's
  failure-count progress.
- `onRetry(attempt+1, max, err)` fires only when there's a next
  attempt, matching the prior semantic.

Generic CircuitBreaker gained two inspection accessors:
- `getOpenedAt(): number | null`
- `getCooldownMs(): number`
Used by `withHfDownloadRetry` to compute `secsUntilReset` without
consuming a probe permit (which `check()` would do).

Test consolidation: the 7 bespoke `HfDownloadCircuitBreaker`
state-machine tests in hf-env.test.ts were 1:1 duplicates of
existing tests in `circuit-breaker.test.ts` and were deleted.
Remaining 42 hf-env tests all pass; full integration sweep (148
gitnexus + 15 web) green.
2026-05-09 15:18:09 +01:00
CopilotandGergő Magyar 248cb1e634 Add regression coverage for .gitnexusignore behavior with --skip-git (#1450)
* Initial plan

* test: cover --skip-git with .gitnexusignore regression

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/7ca9afd5-10b9-4f39-a260-60e60bde6874

* test: reuse cli path constant in skip-git tests

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/7ca9afd5-10b9-4f39-a260-60e60bde6874

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-09 14:26:06 +01:00
dependabot[bot]andGergő Magyar f26a35b17f chore(deps)(deps): bump @anthropic-ai/sdk (#1442)
Bumps the npm_and_yarn group with 1 update in the /gitnexus-web directory: [@anthropic-ai/sdk](https://github.com/anthropics/anthropic-sdk-typescript).


Updates `@anthropic-ai/sdk` from 0.90.0 to 0.91.1
- [Release notes](https://github.com/anthropics/anthropic-sdk-typescript/releases)
- [Changelog](https://github.com/anthropics/anthropic-sdk-typescript/blob/main/CHANGELOG.md)
- [Commits](https://github.com/anthropics/anthropic-sdk-typescript/compare/sdk-v0.90.0...sdk-v0.91.1)

---
updated-dependencies:
- dependency-name: "@anthropic-ai/sdk"
  dependency-version: 0.91.1
  dependency-type: indirect
  dependency-group: npm_and_yarn
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-09 13:53:07 +01:00
Gergő Magyar b5627f27d8 ci: add fork-safe PR autofix pipeline (#1446)
* ci: add fork-safe PR autofix pipeline

Two-workflow split posts prettier + eslint --fix output as inline
review-comment suggestions on PRs (including fork PRs) without running
fork-controlled ESLint plugins under a privileged token.

- pr-autofix.yml: untrusted, runs lint:fix/format with permissions: {},
  uploads diff artifact. paths-ignore on lockfiles/snapshots/dist to
  avoid reviewdog 406 on >3k-line diffs.
- pr-autofix-publish.yml: trusted workflow_run consumer. Validates every
  metadata.json field with regex allowlists before exporting to
  GITHUB_OUTPUT (closes head_ref newline-injection vector). Concurrency
  keyed on PR number with fork fallback to head-repo+branch. Reviewdog
  pinned to v0.21.0. Sticky comment posts only when patch is non-empty
  (no noise on clean PRs); body carries a fenced gitnexus-autofix JSON
  block under a stable HTML marker for agent parsing. gh API calls go
  through a small retry helper for transient 5xx.

Branch protection should enable merge queue + 'require branches up to
date' to handle PR freshness; chinthakagodawita/autoupdate is dropped
(unmaintained since 2023).

* ci(autofix): close zizmor template-injection findings

Move fork-controlled values (head.ref, head.repo.full_name, head.sha,
pr.number, github.repository) into the step's env: block instead of
interpolating them with `${{ }}` directly into the bash run body. The
job has permissions:{} today so this is defence-in-depth, but a future
scope grant on the untrusted half would otherwise turn a malicious
branch name into shell injection.

Add pr-autofix-publish.yml to the documented dangerous-triggers ignore
list — workflow_run is required to post sticky comments on fork PRs
and the file's structural defences (no fork checkout, allowlist on
metadata.json, base_repo equality check) match the existing
ci-report.yml exemption.

* ci(autofix): close remaining review findings

- Add an actionlint job to workflow-lint.yml. Catches YAML syntax,
  expression typing, shellcheck-inside-run, and deprecated runner
  labels on every .github/** PR — closes the gap that let pr-autofix's
  YAML literal-block bug reach review on this branch.
- pr-autofix-publish.yml emits a `gitnexus/autofix` Check Run on the
  PR head SHA: conclusion `success` for clean, `neutral` (with
  distinct output titles) for suggestions-posted vs.
  skipped-too-large. Stable name lets agents read the outcome via
  `gh pr checks` without parsing the sticky comment.
- Document the autofix signal contract in CONTRIBUTING.md — sticky
  marker, fenced gitnexus-autofix JSON schema, Check Run name. One
  source of truth so the marker / schema fields don't drift across
  the workflow files and consumers.

* ci: fix actionlint/shellcheck findings on PR #1446

Closes the actionlint warnings the new lint job (workflow-lint.yml's
actionlint runner) surfaced once it was wired into CI. Mostly
shellcheck-style cleanups across three workflows.

pr-autofix-publish.yml
  - SC2170: `[ "${{ steps.meta.outputs.changed_lines }}" -gt 3000 ]`
    interpolates a literal string into bash, breaking shellcheck's
    arithmetic-comparison parse. Move `changed_lines` through env: as
    `CHANGED_LINES` and reference as `$CHANGED_LINES` inside bash.

ci-report.yml (Read PR metadata step)
  - SC2002 ×2: `cat file | tr` -> `tr < file`.
  - SC2129: three consecutive `>> "$GITHUB_OUTPUT"` redirects collapsed
    into one `{ ...; } >> "$GITHUB_OUTPUT"` group.

ci-report.yml (Build report step)
  - SC2162 ×2: `read VAR1 VAR2` -> `read -r VAR1 VAR2` so backslashes
    in test-results.json output aren't mangled.
  - SC2034: drop unused `SUITES` aggregate. The per-framework suite
    counts (CLI_SU, WEB_SU) are now read into `_` placeholders since
    the report doesn't surface them anywhere.

release-candidate.yml
  - SC2129 ×2: collapse consecutive `>> "$GITHUB_OUTPUT"` redirects in
    the rc-version computation step and the tag-push step into one
    grouped block each.
2026-05-09 13:35:04 +01:00
KareemandGergő Magyar 3daf8c9984 feat(extractors): strip Unreal Engine reflection macros before C++ parsing (#1439)
* feat(extractors): strip Unreal Engine reflection macros before C++ parsing

Tree-sitter does not expand C preprocessor macros, so Unreal Engine reflection markers (UCLASS, UFUNCTION, UPROPERTY, MODULENAME_API, GENERATED_BODY, ...) are parsed verbatim. The result is mis-parsed UE class/function declarations: in 'class BRAWLUI_API UMyClass : public UObject', tree-sitter-cpp captures BRAWLUI_API as the class name, leaving the actual class without an entry in the graph.

This patch adds an optional 'preprocessSource' hook to LanguageProvider and implements it for C++ via a new 'stripUeMacros' module. The transform is length-preserving (each elided byte becomes a space, newlines preserved) so byte offsets and line/column positions tree-sitter reports remain identical to the original file -- symbol locations in the graph stay accurate.

A cheap detection guard short-circuits files that don't look like UE sources, so non-UE C++ codebases pay no cost (single regex test then bail).

27 unit tests cover the detection guard, length preservation across multiple UE samples, macro removal for UCLASS/UFUNCTION/UPROPERTY/USTRUCT/GENERATED_BODY/MODULE_API/DECLARE_*_DELEGATE/UE_DEPRECATED, false-positive guards (substring matches, balanced parens inside string literals, Qt macros left alone), and class-name extraction sanity. Full unit suite still passes (5337 tests, 0 regressions). Verified end-to-end against an Unreal Engine 5.7 game project (Brawl).

* fix(extractors): address PR review findings on UE macro preprocessor

Resolves three blocking issues raised by automated review:

1. Prettier format: ran prettier --write on call-processor.ts, heritage-processor.ts, import-processor.ts (the three sites where the cache-miss reparse hook insertion landed unformatted).

2. Byte-length contract narrowed: language-provider.ts docblock now states the contract precisely (UTF-16 .length + newline-position preservation, not UTF-8 byte length). Notes that startIndex byte offsets only match the original file when the elided range is pure ASCII -- which is the practical UE case (reflection macros and module-export tokens are ASCII-only).

3. Tree-sitter extraction tests added: new end-to-end tests parse the preprocessed source with tree-sitter-cpp and assert the captured class/struct name is the real UClass identifier (UMyClass, FMyData), never the MODULE_API export macro. Also asserts source positions (startPosition.row) survive the transform.

Plus one moderate fix:

4. _API stripping is now scoped to UE files only. The HAS_UE_HINT guard previously included [A-Z]_API tokens, which would fire on non-UE codebases that use REST_API / HTTP_API / MY_LIB_API as constants or enum values, silently erasing them. The guard now requires a strong UE marker (UCLASS|UFUNCTION|UPROPERTY|USTRUCT|UENUM|UINTERFACE|GENERATED_BODY|UE_DEPRECATED|DECLARE_*_DELEGATE) to be present before any stripping runs. Two new tests confirm REST_API and DECLARE_HANDLER style identifiers in non-UE files are left untouched.

Plus one minor fix:

5. stripUeMacros signature now accepts (source, _filePath?) to match the LanguageProvider.preprocessSource hook contract exactly. The filePath argument is unused; UE detection is purely content-based.

Verification: 34/34 preprocessor tests pass (was 27, +7 new for non-ASCII preservation, REST_API safety, tree-sitter extraction, struct extraction, source position preservation). Full unit suite 5349 pass, 0 regressions. Typecheck clean. Prettier --check clean on all 9 changed files.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-09 12:28:32 +01:00
Alex Macdonald-SmithandGergő Magyar d91428ad9d feat(cli): add gitnexus publish for opt-in understand-quickly registry (#1425)
* feat(cli): add `gitnexus publish` for opt-in understand-quickly registry

Adds a small, opt-in command that fires a single `repository_dispatch`
event at `looptech-ai/understand-quickly` to ask the registry for an
instant resync of the current repo's entry. No graph file is uploaded;
the registry pulls from raw.githubusercontent.com per the protocol at
https://github.com/looptech-ai/understand-quickly/blob/main/docs/integrations/protocol.md.

  - Pure helpers (id parsing, payload construction, validation) live in
    `gitnexus-shared/src/integrations/understand-quickly.ts` so the
    package stays Node-free and the same logic is testable in isolation.
  - The CLI command lives in `gitnexus/src/cli/publish.ts`. Without
    `UNDERSTAND_QUICKLY_TOKEN` it is a no-op (exits 0 with one
    informational line); with the token it POSTs the dispatch and
    surfaces 204 / 401 / 404 / 5xx distinctly.
  - The id defaults to `<owner>/<repo>` parsed from the `origin` remote
    and can be overridden with `--id`.
  - Refuses to publish when no `.gitnexus/` index exists, with a
    `gitnexus analyze` hint.

Tests: a new vitest unit covers the pure helpers (8 + 8 + 2 cases) and
the no-token no-op path with a `fetch` spy that fails the test if the
network is touched. README gets a one-paragraph "Publishing to
understand-quickly" section near the existing CLI docs.

* fix(uq-publish): address review blockers + high-severity items

Addresses CodeQL polynomial-regex (HIGH), token-gate ordering, distinct
401/403/404/422 response branches, fetch timeout, expanded test coverage,
tightened owner/repo validation, and non-GitHub remote rejection.

See response thread on PR #1425 for the per-finding rationale.

Signed-off-by: amacsmith <alex.mac@looptech.ai>

* fix(publish): address Claude review on PR #1425

- AbortError → TimeoutError: AbortSignal.timeout() throws a
  DOMException with name 'TimeoutError', not Error{name:'AbortError'}.
  Match the pattern used in core/embeddings/http-client.ts so the
  user-facing "timed out after 15000ms" message actually fires. Update
  the regression test to throw a real DOMException — the previous fake
  was a false-green.
- isValidOwnerRepo: forbid trailing hyphen in the owner segment.
  GitHub rejects this at account-creation time; allowing it here meant
  hand-typed --id values like 'my-org-/repo' would pass our regex and
  422 from GitHub.
- Add publish-command coverage to cli-index-help.test.ts (asserts on
  --id, --skip-git, the registry name, and the token env var) and
  cli-commands.test.ts (asserts publishCommand is exported as a
  function). Catches accidental command-registration deletion.

---------

Signed-off-by: amacsmith <alex.mac@looptech.ai>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-09 09:52:26 +01:00
32b5c0e3fc feat: add IncludeExtractor for C++ cross-repo include tracking (group) (#1156)
* feat: add IncludeExtractor for C++ cross-repo include tracking (group)

* fix: address CodeQL warnings on include-extractor

- Remove unused HEADER_GLOB constant in include-extractor.ts
- Use fs.mkdtempSync for secure temp dir creation in tests
  (CodeQL: 'Insecure temporary file')

* fix(group): close missing ); in manifest-extractor include branch

The 'include' branch in ManifestExtractor.resolveSymbol was missing
the closing ); for the executor() call, causing a syntax error that
broke ESLint, Prettier, and the full test CI on all platforms.

Reported by Claude PR review on #1156.

* chore: drop test/global-setup.ts + test/vitest.d.ts

Upstream removed these in commit 3f0c74fe (ladybugdb 0.16.0 upgrade).
Commit 3f5d21c5 accidentally restored them during a rebase dance.

* style(group): reformat VALID_CONTRACT_TYPES array to satisfy prettier

Adding 'include' pushed the array over prettier's 100-char limit,
so prettier prefers multi-line. Apply the reformat to unbreak
ci-quality/format job.

* fix(include-extractor): address PR #1156 Claude review findings #3-#7

Claude Deep Review raised 7 findings on the IncludeExtractor. #1/#2
(BLOCKERs) were fixed earlier. This commit closes the remaining five.

#3 HIGH  case-sensitive FS -> provider contract-id collision
  Document the deliberate case-folding trade-off on normalizeIncludePath
  (matches C/C++ convention on Windows/macOS; collapses Foo.h & foo.h on
  Linux). Add a unit test pinning the behavior.

#4 HIGH  suffixResolve short-suffix match silently drops cross-repo include
  When a local file ends with the same basename as an external include
  (e.g. local internal/api.h vs. #include "ext/api.h"), suffixResolve
  returned a bogus local hit and suppressed the cross-repo consumer.
  Replace the suffixResolve lookup inside include-extractor with a
  strict isLocalInclude() that only accepts full-path hits via
  SuffixIndex.get / getInsensitive. Callers of suffixResolve elsewhere
  are unaffected. Add 3 unit tests covering the regression.

#5 MEDIUM regex fallback matched #include inside /* ... */
  Strip block comments before running the fallback regex scan.
  Add a unit test.

#6 MEDIUM meta.source was hard-coded to 'tree_sitter'
  Track the actual extraction path with an extractionSource local and
  write it into meta.source so downstream audits can distinguish
  tree-sitter parses from regex fallbacks. Add 2 unit tests.

#7 MEDIUM missing end-to-end coverage
  Add test/integration/group/include-extractor-sync.test.ts with 3
  cases exercising extractor -> syncGroup -> CrossLink (mocked
  contracts, mixed-case/backslash normalization, real temp repos).

Tests: 21 unit + 3 integration, all green.

* fix(lbug): robust Windows lock acquisition for CI integration tests

LadybugDB's `new Database()` raises `Could not set lock on file` from
local_file_system.cpp synchronously inside the constructor — before any
query is issued, so `withLbugDb`'s query-time retry never sees it. On
Windows CI this surfaces as flaky integration tests due to AV-scanner
holds, libuv handle-release lag, and stale `.wal` sidecars from aborted
prior runs.

This change closes the gap at *open time*:

- `openLbugConnection` now wraps `new lbug.Database()` in a bounded
  busy-retry (5x100ms back-off) inside `lbug-config.ts`. Errors that
  exhaust the budget are tagged via `LBUG_OPEN_RETRY_EXHAUSTED` so
  `withLbugDb`'s outer 3x retry skips re-retrying a freshly-exhausted
  path (eliminates the 3x5=15-attempt / ~6s tail latency).
- For recognized test fixtures only (immediate-parent dir matches a
  known prefix AND resolves under `os.tmpdir()`), one final stale-
  sidecar sweep removes `.wal`/`.lock` and retries once. Production
  paths never enter this branch.
- `safeClose` on Windows runs a bounded `fs.open` probe to absorb
  native handle-release lag; logs a warning if the probe exhausts so
  operators can spot AV interference.
- `isDbBusyError` is now defined in `lbug-config.ts` as the single
  source of truth, re-exported from `lbug-adapter.ts` for compatibility.
- New tests cover open-time retry (happy/retry/exhaust/non-busy/tag),
  stale-sidecar sweep (test-fixture-only, production-rejection,
  preserves-original-error), `isTestFixturePath` direct unit suite
  (accept/reject/traversal/nested/trailing-sep), and
  `waitForWindowsHandleRelease` (openable/ENOENT/no-leak).
- The two new test files are added to vitest's existing serialized
  `lbug-db` project (already `fileParallelism: false`).

Closes the chronic Windows CI flake on lbug-touching integration tests
while preserving the existing single-writable-Database-per-process
LadybugDB contract. No public API surface changed.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(lbug): drop isDbBusyError re-export, import from lbug-config directly

The re-export from lbug-adapter.ts was a transitional convenience — with
the matcher now living in lbug-config.ts, having two import paths for the
same symbol invites future drift. Updated the two real consumers
(lbug-lock-retry.test.ts, lbug-open-retry.test.ts) to import from
lbug-config directly, removed the re-export equality test (now vacuous),
and refreshed the explanatory comment so it no longer references a
re-export pattern that doesn't exist.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(lbug): silence benign LadybugDB v0.16.1 schema-init lock warnings on Windows

doInitLbug logs "⚠️ Schema creation warning: ... Could not set lock on
file" on every CREATE NODE TABLE call after the first init on a given
dbPath, on Windows. The lock is internal to LadybugDB v0.16.1 and is
resolved before the table is created — same tolerance pattern as the
existing "already exists" filter. Genuine cross-process lock contention
still surfaces on the next operation through withLbugDb's retry, so
filtering at the schema-init catch only suppresses noise, not signal.

Also extend the safeClose Windows handle-release probe to cover the
.wal sidecar (the previous Database's WAL handle was the slowest to
release, surfacing as the schema-query lock contention) and switch the
probe back to 'r+' so it actually detects exclusive locks.

Test loop in lbug-close-handle-release.test.ts simplified to 10 plain
iterations now that the underlying noise is filtered upstream.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(lbug): isDbBusyError review fixes

- Drop redundant `could not set lock` term — already subsumed by `lock`.
- Document the intentionally-broad matcher: graph-DB lock-shaped errors
  ("deadlock", "unlock failed", "lock contention", "could not open lock
  file") are all treated as transient. If a non-transient surfaces,
  tighten the matcher rather than raise the retry budget.
- Add positive test cases covering those lock-shaped strings so the
  intent is visible and a future tightening would deliberately break
  these.
- Fix the open-retry back-off comment: max sleep is 100+200+300+400 =
  1000ms (no sleep after the final attempt), not 1.5s.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(group): address PR #1156 follow-up review findings

Addresses two blockers and two mediums from the deep review.

BLOCKER 1: Windows CI ENOTEMPTY in sync.test.ts
  After this PR added writeBridge() to syncGroup, the existing test
  "writes registry to groupDir when skipWrite is false" fails on
  windows-latest. LadybugDB's checkpoint thread briefly outlives
  closeBridgeDb, holding a Win32 lock on bridge.lbug; the test's
  fs.rmSync then fails with ENOTEMPTY. Switched the test cleanup to
  cleanupTempDir from test/helpers/test-db.ts which already tolerates
  EBUSY/EPERM/EACCES/ENOTEMPTY with bounded retries — same pattern
  used elsewhere for LadybugDB-touching tests.

BLOCKER 2: Graph provider absolute-path bug
  extractProvidersGraph queried File.filePath from the LadybugDB graph
  but never stripped the repo root, so provider contract IDs ended up
  as include::/abs/path/foo.h while consumers emitted include::foo.h.
  These never matched through runExactMatch — silently producing 0
  cross-links for any indexed C++ repo (the primary use case).
  Now passes repoPath into extractProvidersGraph and applies
  path.relative(); rows that resolve outside repoPath (stale absolute
  paths from another machine, system headers somehow indexed) are
  dropped instead of polluting the registry.

MEDIUM: `../` relative includes produce spurious noise
  `#include "../foo.h"` is almost always intra-repo, but the suffix
  index can never match a `..`-prefixed path so it became a consumer
  contract no provider could satisfy. Now skipped before matching;
  covers both forward-slash and backslash forms.

MEDIUM: writeBridge error in sync.ts propagates uncaught
  contracts.json is the canonical source of truth and was just written
  successfully when writeBridge runs. A bridge-only failure (disk full,
  schema error, permission denied) shouldn't mask the registry. Wrapped
  writeBridge in try/catch with a logger.warn surfacing the path and
  recovery instructions.

Tests added:
  - extractProvidersGraph repo-relative ID generation (stub Cypher
    executor returns absolute paths)
  - extractProvidersGraph drops rows whose path resolves outside repo
  - `../foo.h` forward-slash skip
  - `..\foo.h` backslash-form skip

Skipped findings:
  - canExtract() removal (#5, low): canExtract is part of the
    ContractExtractor interface; every other extractor implements the
    same `return true` shape. Removing it from IncludeExtractor would
    break the interface contract — keeping for consistency.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(group): close PR #1156 Codex adversarial findings

Two HIGH findings from the Codex adversarial review on
feat/group-include-extractor:

1. Default-on extraction silently changes existing groups (BLOCKER)
   DEFAULT_DETECT.includes was true, so any pre-existing group.yaml
   that omits the new field would gain a wave of include::* contracts
   on the next sync after upgrade. Flipped to false (opt-in). The
   integration test already declares includes: true explicitly so it
   survives unchanged; the unit extractor tests bypass parseGroupConfig
   entirely; the sync test uses extractorOverride. Only config-parser
   needed regression tests covering omitted/explicit/false variants.

2. IncludeExtractor scans outside the indexed file universe (BLOCKER)
   The extractor was running glob('**/*', { ignore: STANDARD_IGNORES })
   twice with a hand-rolled 9-pattern list, no .gitignore/.gitnexusignore
   honoring, and no max-file-size cap. That meant File:<path> contracts
   could appear for files ingestion would never index, producing
   cross-links group impact cannot fan out to (silent false-negatives).
   Refactored to a single discoverIndexableFiles() helper that mirrors
   walkRepositoryPaths exactly: createIgnoreFilter + getMaxFileSizeBytes,
   one discovery pass shared by provider and consumer paths. Dropped
   STANDARD_IGNORES and SOURCE_GLOB entirely.

   third_party and 3rdparty (the C/C++ vendored-deps conventions) were
   in the local ignore list but not in the canonical DEFAULT_IGNORE_LIST
   used by ingestion. Folded both into the canonical set rather than
   keep a parallel list — the whole point of the Codex finding is that
   two file-discovery implementations drift. Single source of truth.

Tests: 5 new regression tests for the discovery alignment (.gitignore,
.gitnexusignore, max-file-size on both provider and consumer paths)
plus 4 for the opt-in default. All 30 include-extractor tests + the
494-test group suite + ignore-service tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(review): apply autofix feedback

ce-code-review surfaced 6 safe_auto findings on commit a9936a9b:

- T1 (testing, P2): the sync.ts:174 gate was untested with includes:false.
  Added a sync-level test mirroring the existing thrift-off pattern at
  sync.test.ts:545, asserting zero include contracts when the gate is
  disabled in a real syncGroup call.

- T3 (testing, P3): third_party and 3rdparty entries in DEFAULT_IGNORE_LIST
  had no regression test. Added both to ignore-service.test.ts's
  dependency-directories it.each block.

- M1 (maintainability, P3): discoverIndexableFiles JSDoc lacked a
  fork-warning relative to walkRepositoryPaths. Added a MAINTENANCE
  note explaining why the duplication is tolerated and the contract
  the two implementations must keep.

- M2 (maintainability, P3): thrift-extractor still hand-rolls its
  ignore array with no signal that DEFAULT_IGNORE_LIST additions
  silently do not apply there. Added TODO(#1156-followup) comments
  above both call sites.

- M3 (maintainability, P3): SOURCE_EXTENSIONS duplicated the four
  HEADER_EXTENSIONS entries with no expressed subset relationship.
  Spread HEADER_EXTENSIONS into SOURCE_EXTENSIONS so future header-
  extension additions propagate.

- C1+T4 (correctness+testing, P3, cross-reviewer corroborated):
  discoverIndexableFiles swallowed all fs.stat errors silently,
  including EACCES/EMFILE/EIO. Narrowed the catch to ENOENT (the
  documented benign glob/stat race) and added a logger.warn for
  any other code so operators can spot permission/resource issues.

All 629 tests pass; typecheck + prettier clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(group): use retryRename in writeContractRegistry to absorb Windows EPERM

`storage.ts:62` used raw `fsp.rename` for the contracts.json atomic swap.
On Windows, AV scanners and concurrent renames briefly hold the
destination handle between rename calls, surfacing as EPERM/EBUSY.
The `insecure-tempfile.test.ts > concurrent writes do not collide`
test was flaking with `EPERM: operation not permitted, rename` on
windows-latest CI.

`bridge-db.ts` already has a battle-tested `retryRename(src, dst, 3)`
helper used at six call sites for exactly this pattern. Reusing it
here keeps the Windows-rename policy single-source-of-truth across
the group package.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(group): drop macro-style #include from consumer contracts

Tree-sitter's `(_) @import.source` wildcard matches the identifier node
of `#include PLATFORM_HEADER`, so the cleaned value `PLATFORM_HEADER`
slipped past the system-header / `..` filters and was emitted as a
permanently orphaned consumer contract (no file is named after a macro
identifier, so no provider can ever match). Add a shape guard that
skips cleaned values lacking both a path separator and an extension
dot, plus regression tests for single and multi-macro files.

Also document `IncludeExtractor.canExtract()` as unused by sync.ts
(gated via `config.detect.includes` instead) and kept solely for
ContractExtractor interface uniformity.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: HuangWenjie <zhoudeng.hwj@alibaba-inc.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 09:31:59 +01:00
30e0c7c726 fix(csharp): include generic typed properties in context and impact (#1399)
* Fix C# context and impact for generic typed properties

* Address C# typed-property review feedback

---------

Co-authored-by: Richard Carmel <rcarmel@seitel.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-09 09:07:24 +01:00
dependabot[bot]andGergő Magyar 98addbd6c4 chore(deps)(deps-dev): bump @types/node in /gitnexus (#1436)
Bumps [@types/node](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/node) from 25.6.1 to 25.6.2.
- [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases)
- [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/node)

---
updated-dependencies:
- dependency-name: "@types/node"
  dependency-version: 25.6.2
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-09 07:30:05 +01:00
dependabot[bot]andGergő Magyar f29147864e chore(deps)(deps): bump fast-uri from 3.1.0 to 3.1.2 in /gitnexus (#1441)
Bumps [fast-uri](https://github.com/fastify/fast-uri) from 3.1.0 to 3.1.2.
- [Release notes](https://github.com/fastify/fast-uri/releases)
- [Commits](https://github.com/fastify/fast-uri/compare/v3.1.0...v3.1.2)

---
updated-dependencies:
- dependency-name: fast-uri
  dependency-version: 3.1.2
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-09 07:09:10 +01:00
dependabot[bot]andGergő Magyar b89ec5b5df chore(deps)(deps): bump hono from 4.12.16 to 4.12.18 in /gitnexus (#1443)
Bumps [hono](https://github.com/honojs/hono) from 4.12.16 to 4.12.18.
- [Release notes](https://github.com/honojs/hono/releases)
- [Commits](https://github.com/honojs/hono/compare/v4.12.16...v4.12.18)

---
updated-dependencies:
- dependency-name: hono
  dependency-version: 4.12.18
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-09 06:49:32 +01:00
dependabot[bot] 5bfe0c5222 chore(deps)(deps): bump onnxruntime-node in /gitnexus (#1435) 2026-05-09 05:53:16 +01:00
azizur100389 5497079ab2 fix(search): surface warning when FTS indexes are missing (#1418) 2026-05-08 17:05:18 +01:00
631 changed files with 58973 additions and 4207 deletions
@@ -17,11 +17,11 @@ npx gitnexus analyze
Run from the project root. This parses all source files, builds the knowledge graph, writes it to `.gitnexus/`, and generates CLAUDE.md / AGENTS.md context files.
| Flag | Effect |
| ------------------- | ------------------------------------------------------------------------------------------------------- |
| `--force` | Force full re-index even if up to date |
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
| Flag | Effect |
| -------------- | ---------------------------------------------------------------- |
| `--force` | Force full re-index even if up to date |
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Claude Code, a PostToolUse hook detects staleness after `git commit` and `git merge` and notifies the agent to run `analyze` — the hook does not run analyze itself, to avoid blocking the agent for up to 120s and risking KuzuDB corruption on timeout.
@@ -1,5 +1,5 @@
name: Setup GitNexus Web
description: Setup Node.js 20.19+ (vite 7 floor), build gitnexus-shared, install web dependencies
description: Setup Node.js 22, build gitnexus-shared, install web dependencies
runs:
using: composite
@@ -7,9 +7,7 @@ runs:
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
with:
# Vite 7 requires Node ^20.19.0 || >=22.12.0 (require(esm) support).
# Pin explicitly so we don't depend on the floating "20" alias resolving
# to a high enough patch version on every runner image.
node-version: '20.19.0'
node-version: 22
cache: npm
cache-dependency-path: gitnexus-web/package-lock.json
+2 -2
View File
@@ -1,5 +1,5 @@
name: Setup GitNexus
description: Setup Node.js 20, install dependencies, and optionally build
description: Setup Node.js 22, install dependencies, and optionally build
inputs:
build:
@@ -12,7 +12,7 @@ runs:
steps:
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
with:
node-version: 20
node-version: 22
cache: npm
cache-dependency-path: gitnexus/package-lock.json
+47
View File
@@ -7,6 +7,8 @@ updates:
directory: /
schedule:
interval: weekly
cooldown:
default-days: 7
open-pull-requests-limit: 5
commit-message:
prefix: chore
@@ -15,6 +17,36 @@ updates:
- dependencies
- ci
# Keep pinned Docker base-image digests current for the root Dockerfiles.
- package-ecosystem: docker
directory: /
schedule:
interval: weekly
cooldown:
default-days: 7
open-pull-requests-limit: 5
commit-message:
prefix: chore(deps)
include: scope
labels:
- dependencies
- ci
# Keep the nested test-image Docker base digest current as well.
- package-ecosystem: docker
directory: /gitnexus
schedule:
interval: weekly
cooldown:
default-days: 7
open-pull-requests-limit: 5
commit-message:
prefix: chore(deps)
include: scope
labels:
- dependencies
- ci
# Gitnexus npm deps — tree-sitter grammars checked daily so we catch
# new releases that unblock the tree-sitter 0.25 upgrade ASAP. Grammars
# are grouped so lockstep bumps produce a single PR. The tree-sitter
@@ -25,6 +57,11 @@ updates:
directory: /gitnexus
schedule:
interval: daily
cooldown:
default-days: 7
semver-major-days: 30
semver-minor-days: 7
semver-patch-days: 3
open-pull-requests-limit: 10
commit-message:
prefix: chore(deps)
@@ -54,6 +91,11 @@ updates:
directory: /gitnexus-web
schedule:
interval: weekly
cooldown:
default-days: 7
semver-major-days: 30
semver-minor-days: 7
semver-patch-days: 3
open-pull-requests-limit: 5
commit-message:
prefix: chore(deps)
@@ -67,6 +109,11 @@ updates:
directory: /gitnexus-shared
schedule:
interval: weekly
cooldown:
default-days: 7
semver-major-days: 30
semver-minor-days: 7
semver-patch-days: 3
open-pull-requests-limit: 5
commit-message:
prefix: chore(deps)
+3
View File
@@ -3,6 +3,9 @@ name: E2E Tests
on:
workflow_call:
permissions:
contents: read
jobs:
check-changes:
name: Check web module changes
+5 -2
View File
@@ -3,6 +3,9 @@ name: Quality Checks
on:
workflow_call:
permissions:
contents: read
jobs:
format:
runs-on: ubuntu-latest
@@ -11,7 +14,7 @@ jobs:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
with:
node-version: 20
node-version: 22
cache: npm
cache-dependency-path: package-lock.json
- run: npm ci
@@ -24,7 +27,7 @@ jobs:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
with:
node-version: 20
node-version: 22
cache: npm
cache-dependency-path: package-lock.json
- run: npm ci
+15 -10
View File
@@ -95,31 +95,33 @@ jobs:
# Validate PR number is a positive integer (artifact comes from
# untrusted fork code, so treat contents defensively).
PR_NUM=$(cat "$DIR/pr_number" | tr -d '[:space:]')
PR_NUM=$(tr -d '[:space:]' < "$DIR/pr_number")
if ! [[ "$PR_NUM" =~ ^[0-9]+$ ]]; then
echo "skip=true" >> "$GITHUB_OUTPUT"
echo "::error::Invalid PR number in artifact: '$PR_NUM'"
exit 0
fi
echo "skip=false" >> "$GITHUB_OUTPUT"
echo "pr_number=$PR_NUM" >> "$GITHUB_OUTPUT"
# Validate job-result strings against known GitHub Actions values.
# Artifact contents come from the PR workflow (potentially untrusted
# fork code), so we whitelist to prevent newline injection into
# GITHUB_OUTPUT.
validate_result() {
local val
val=$(cat "$1" | tr -d '[:space:]')
val=$(tr -d '[:space:]' < "$1")
case "$val" in
success|failure|cancelled|skipped) echo "$val" ;;
*) echo "unknown" ;;
esac
}
echo "quality=$(validate_result "$DIR/quality_result")" >> "$GITHUB_OUTPUT"
echo "tests=$(validate_result "$DIR/tests_result")" >> "$GITHUB_OUTPUT"
echo "e2e=$(validate_result "$DIR/e2e_result")" >> "$GITHUB_OUTPUT"
{
echo "skip=false"
echo "pr_number=$PR_NUM"
echo "quality=$(validate_result "$DIR/quality_result")"
echo "tests=$(validate_result "$DIR/tests_result")"
echo "e2e=$(validate_result "$DIR/e2e_result")"
} >> "$GITHUB_OUTPUT"
- name: Checkout (for vitest config)
if: steps.meta.outputs.skip != 'true'
@@ -279,14 +281,17 @@ jobs:
fi
}
read CLI_T CLI_P CLI_F CLI_S CLI_SU CLI_D <<< "$(sum_results "$RESULTS_FILE")"
read WEB_T WEB_P WEB_F WEB_S WEB_SU WEB_D <<< "$(sum_results "$WEB_RESULTS_FILE")"
# `_` placeholder for the suite-count column — positional
# readability for sum_results' 6-field output, but the value
# isn't surfaced in the report (suites are tracked per-test
# framework, not as a top-line metric).
read -r CLI_T CLI_P CLI_F CLI_S _ CLI_D <<< "$(sum_results "$RESULTS_FILE")"
read -r WEB_T WEB_P WEB_F WEB_S _ WEB_D <<< "$(sum_results "$WEB_RESULTS_FILE")"
TOTAL=$((CLI_T + WEB_T))
PASSED=$((CLI_P + WEB_P))
FAILED=$((CLI_F + WEB_F))
SKIPPED=$((CLI_S + WEB_S))
SUITES=$((CLI_SU + WEB_SU))
DURATION=$((CLI_D > WEB_D ? CLI_D : WEB_D))
# ── Status helpers ──
+3
View File
@@ -28,6 +28,9 @@ name: Scope Resolution Parity
on:
workflow_call:
permissions:
contents: read
jobs:
discover:
name: Discover migrated languages
+105
View File
@@ -3,6 +3,9 @@ name: Tests
on:
workflow_call:
permissions:
contents: read
jobs:
tests:
name: ubuntu / coverage
@@ -72,3 +75,105 @@ jobs:
build: 'true'
- run: npx vitest run
working-directory: gitnexus
# End-to-end smoke test for the #1728 packaging fix: pack the published
# tarball, install it globally into a temp prefix, and assert no junction
# creation (the EPERM root cause) plus working CLI plus vendor cleanliness
# (#836). Runs on windows-latest because that is the platform the fix
# targets; the in-repo `npm ci` job above only exercises the dev-tree path
# and skips the tarball reify step where the historical EPERM occurred.
packaged-install-smoke:
name: packaged install smoke (${{ matrix.os }})
strategy:
fail-fast: false
matrix:
os: [windows-latest, ubuntu-latest]
runs-on: ${{ matrix.os }}
timeout-minutes: 15
steps:
# persist-credentials: false — this job runs npm pack + npm install -g
# from a tarball and never pushes back; the token in .git/config would
# be at risk of leaking through any future artifact-upload step
# (zizmor artipacked audit). Disable upfront.
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
- uses: ./.github/actions/setup-gitnexus
with:
build: 'true'
- name: Pack gitnexus tarball
shell: bash
run: npm pack
working-directory: gitnexus
- name: Install gitnexus tarball into isolated prefix
shell: bash
run: |
set -euo pipefail
PREFIX="$RUNNER_TEMP/gitnexus-smoke"
mkdir -p "$PREFIX"
TARBALL=$(find . -maxdepth 1 -name 'gitnexus-*.tgz' -print -quit)
if [ -z "$TARBALL" ]; then
echo "ERROR: no gitnexus-*.tgz tarball found in $(pwd)" >&2
exit 1
fi
echo "Installing $TARBALL into $PREFIX"
npm install -g --prefix "$PREFIX" "./$TARBALL" --no-audit --no-fund
echo "PREFIX=$PREFIX" >> "$GITHUB_ENV"
working-directory: gitnexus
- name: Assert no junctions or vendor build artifacts
shell: bash
run: |
set -euo pipefail
# Locate the installed gitnexus package across npm prefix layouts
# (lib/node_modules on POSIX, node_modules on Windows).
for candidate in "$PREFIX/lib/node_modules/gitnexus" "$PREFIX/node_modules/gitnexus"; do
if [ -d "$candidate" ]; then
INSTALLED="$candidate"
break
fi
done
if [ -z "${INSTALLED:-}" ]; then
echo "ERROR: installed gitnexus package not found under $PREFIX" >&2
ls -la "$PREFIX" || true
exit 1
fi
echo "Installed package at: $INSTALLED"
# #836 invariant: no node_modules/ or build/ under any vendor/*.
BAD=$(find "$INSTALLED/vendor" \( -name node_modules -o -name build \) -print 2>/dev/null || true)
if [ -n "$BAD" ]; then
echo "ERROR: vendor tree contains forbidden build artifacts (#836):" >&2
echo "$BAD" >&2
exit 1
fi
# #1728 invariant: materialized grammar dirs are real directories,
# not junctions/symlinks (which is what the EPERM regression created).
for name in tree-sitter-dart tree-sitter-proto tree-sitter-swift; do
entry="$INSTALLED/node_modules/$name"
if [ ! -e "$entry" ]; then
echo "WARN: $name not materialized (toolchain/prebuild may be unavailable on $RUNNER_OS)"
continue
fi
if [ -L "$entry" ]; then
echo "ERROR: $entry is a symlink/junction — #1728 regression" >&2
exit 1
fi
if [ ! -d "$entry" ]; then
echo "ERROR: $entry is not a directory" >&2
exit 1
fi
done
- name: Assert gitnexus --version works
shell: bash
run: |
set -euo pipefail
if [ "$RUNNER_OS" = "Windows" ]; then
"$PREFIX/gitnexus.cmd" --version
else
"$PREFIX/bin/gitnexus" --version
fi
+11 -8
View File
@@ -6,16 +6,19 @@ on:
paths-ignore: ['**.md', 'docs/**', 'LICENSE']
workflow_call:
permissions:
contents: read
# Concurrency convention: see CONTRIBUTING.md → "GitHub Actions — Concurrency Convention".
# Hardcoded `CI-` prefix (not `${{ github.workflow }}`) because this workflow is
# invoked as a reusable workflow from publish.yml and release-candidate.yml. In
# called-workflow context `github.workflow` evaluation is ambiguous across GitHub
# Actions versions, and a prefix that could resolve to the caller's name would
# share a concurrency group with the caller → deadlock. A literal prefix is
# immune. Direct `pull_request` invocations use `CI-<ref>`; invocations from a
# reusable-workflow caller fall into a per-run-unique group that never serializes
# with the caller. `push` to main is handled by release-candidate.yml, which
# calls this workflow once before publishing.
# invoked as a reusable workflow from publish.yml. In called-workflow context
# `github.workflow` evaluation is ambiguous across GitHub Actions versions, and a
# prefix that could resolve to the caller's name would share a concurrency group
# with the caller → deadlock. A literal prefix is immune. Direct `pull_request`
# invocations use `CI-<ref>`; invocations from a reusable-workflow caller fall
# into a per-run-unique group that never serializes with the caller. `push` to
# main is handled by publish.yml (RC mode), which calls this workflow once
# before publishing.
concurrency:
group: ${{ github.event_name == 'pull_request' && format('CI-{0}', github.ref) || format('CI-nested-{0}', github.run_id) }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
+6 -1
View File
@@ -19,6 +19,9 @@ on:
pull_request_review:
types: [submitted]
permissions:
contents: read
# Concurrency convention: see CONTRIBUTING.md → "GitHub Actions — Concurrency Convention".
# Serialize per-PR/issue to avoid racing comments.
concurrency:
@@ -155,6 +158,8 @@ jobs:
github_token: ${{ secrets.GITHUB_TOKEN }}
allowed_non_write_users: '*'
show_full_output: true
# Review posts use Bash (`gh`, etc.); default mode asks for approval — impossible in CI.
claude_args: '--dangerously-skip-permissions'
plugin_marketplaces: 'https://github.com/anthropics/claude-code.git'
plugins: 'code-review@claude-code-plugins'
prompt: '/code-review:code-review ${{ github.repository }}/pull/${{ steps.pr.outputs.number }}'
prompt: '/code-review:code-review https://github.com/${{ github.repository }}/pull/${{ steps.pr.outputs.number }} --comment'
+5 -2
View File
@@ -17,6 +17,9 @@ on:
# already-merged code without waiting for the next PR.
- cron: '0 6 * * 1'
permissions:
contents: read
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
@@ -45,7 +48,7 @@ jobs:
persist-credentials: false
- name: Initialize CodeQL
uses: github/codeql-action/init@e46ed2cbd01164d986452f91f178727624ae40d7 # v4.35.3
uses: github/codeql-action/init@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
with:
languages: ${{ matrix.language }}
queries: security-and-quality
@@ -66,6 +69,6 @@ jobs:
- '**/test/fixtures/**'
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@e46ed2cbd01164d986452f91f178727624ae40d7 # v4.35.3
uses: github/codeql-action/analyze@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
with:
category: '/language:${{ matrix.language }}'
+4 -1
View File
@@ -10,6 +10,9 @@ on:
pull_request:
branches: [main]
permissions:
contents: read
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
@@ -30,7 +33,7 @@ jobs:
persist-credentials: false
- name: Dependency Review
uses: actions/dependency-review-action@2031cfc080254a8a887f58cffee85186f0e49e48 # v4.9.0
uses: actions/dependency-review-action@a1d282b36b6f3519aa1f3fc636f609c47dddb294 # v5.0.0
with:
fail-on-severity: high
comment-summary-in-pr: on-failure
+14 -2
View File
@@ -25,6 +25,18 @@ on:
a gitnexus/package.json whose version matches the tag.
required: true
type: string
# Explicit secret contract — callers pass these by name. Replaces the
# blanket `secrets: inherit` pattern (zizmor `secrets-inherit` audit).
# GHCR auth uses the implicit GITHUB_TOKEN; only Docker Hub credentials
# need to be passed through.
secrets:
DOCKERHUB_USERNAME:
required: true
DOCKERHUB_TOKEN:
required: true
permissions:
contents: read
# Concurrency convention: see CONTRIBUTING.md → "GitHub Actions — Concurrency Convention".
# Tag refs are unique per release, so distinct tags run in parallel.
@@ -70,7 +82,7 @@ jobs:
steps:
# Only the workflow_call path requires a non-empty `inputs.tag` — callers
# (e.g. release-candidate.yml) must pass the RC tag explicitly. On direct
# (publish.yml in RC mode) must pass the RC tag explicitly. On direct
# tag pushes the tag comes from `github.ref`, so `inputs.tag` is always
# empty and validating it here would break every real release (#1064).
# The downstream "Verify tag matches gitnexus/package.json version" step
@@ -132,7 +144,7 @@ jobs:
uses: docker/setup-buildx-action@4d04d5d9486b7bd6fa91e7baf45bbb4f8b9deedd # v4.0.0
- name: Install Cosign
uses: sigstore/cosign-installer@cad07c2e89fa2edd6e2d7bab4c1aa38e53f76003 # v4.1.1
uses: sigstore/cosign-installer@6f9f17788090df1f26f669e9d70d6ae9567deba6 # v4.1.2
- name: Log in to GitHub Container Registry
if: ${{ github.event_name != 'pull_request' && !inputs.dry_run }}
+3
View File
@@ -12,6 +12,9 @@ on:
push:
branches: [main]
permissions:
contents: read
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
+595
View File
@@ -0,0 +1,595 @@
name: PR Autofix (apply)
# CHATOPS HALF of the autofix pipeline.
#
# Triggered when a contributor comments `/autofix` on a PR. Validates
# permission, locates the most recent successful `pr-autofix.yml`
# artifact for the PR's current head SHA, applies the patch to the PR
# head, and pushes a commit back to the PR branch.
#
# This workflow runs from the default branch's copy of the file
# regardless of where the comment originates -- that's the trust
# anchor. Comment body and author login are untrusted; both flow
# through env vars and pattern-matched, never interpolated into shell.
#
# Fork PR support: `git push` with the GITHUB_TOKEN succeeds against
# fork branches only when the contributor enabled "Allow edits by
# maintainers" on the PR (the default). When they disabled it, we
# fail loud with a 👎 reaction and an explanation comment.
on:
issue_comment:
types: [created]
concurrency:
# Per-PR scope. issue_comment events expose `github.event.issue.number`
# for both PR and Issue comments; the `pull_request != null` guard on
# the job ensures we only run on PRs, so this number is the PR number.
# cancel-in-progress: false — a second `/autofix` should wait for the
# first to finish (idempotency check on the second invocation handles
# the no-op case).
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.event.issue.number }}
cancel-in-progress: false
permissions: {}
jobs:
apply:
name: apply-autofix
# Pre-filter at the workflow level so non-PR comments and unrelated
# comments don't even spawn a runner. The job-level body re-check
# below (Step 1) is the strict gate.
if: >-
github.event.issue.pull_request != null
&& startsWith(github.event.comment.body, '/autofix')
runs-on: ubuntu-latest
timeout-minutes: 10
permissions:
# React on the triggering comment + post reply comments.
pull-requests: write
# Push the apply commit to the PR head branch.
contents: write
# Required by actions/download-artifact to fetch artifacts produced
# by a different workflow run.
actions: read
steps:
- name: Validate comment body precisely
id: body
env:
BODY: ${{ github.event.comment.body }}
shell: bash
run: |
set -euo pipefail
# Whole-line, case-sensitive match: `^/autofix\s*$`. The
# workflow-level startsWith guard is coarse — `please don't
# /autofix this code` would pass that filter but fail this one.
# We exit silently (no reaction) on body mismatch so quoted
# text in unrelated discussions doesn't get a visible response.
if [[ ! "${BODY}" =~ ^/autofix[[:space:]]*$ ]]; then
echo "Body did not match strict /autofix regex — exiting silently."
echo "match=false" >> "$GITHUB_OUTPUT"
exit 0
fi
echo "match=true" >> "$GITHUB_OUTPUT"
- name: Validate commenter permission
id: perm
if: steps.body.outputs.match == 'true'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
GH_REPO: ${{ github.repository }}
COMMENTER: ${{ github.event.comment.user.login }}
PR_AUTHOR: ${{ github.event.issue.user.login }}
shell: bash
run: |
set -euo pipefail
# Retry wrapper for transient 5xx / 429 / network blips.
# Mirrors the helper in pr-autofix-publish.yml. Used on
# idempotent GETs only; reactions/comment-POSTs are NOT
# wrapped (retrying a POST would dupe the resource).
gh_retry() {
local n=0 max=3
while true; do
if gh "$@"; then return 0; fi
n=$((n+1))
if [ "$n" -ge "$max" ]; then return 1; fi
sleep $((n * 2))
done
}
# Allowlist the commenter login before it flows into a URL.
# GitHub usernames: alphanumeric + dashes, max 39 chars.
if ! [[ "${COMMENTER}" =~ ^[A-Za-z0-9-]{1,39}$ ]]; then
echo "::error::Invalid commenter login format: $(printf '%q' "${COMMENTER}")"
echo "allowed=false" >> "$GITHUB_OUTPUT"
exit 0
fi
# Self-comparison: PR author can always /autofix their own PR.
if [ "${COMMENTER}" = "${PR_AUTHOR}" ]; then
echo "Commenter is PR author — granting access."
echo "allowed=true" >> "$GITHUB_OUTPUT"
exit 0
fi
# Repo permission lookup. admin/write/maintain are sufficient.
# Distinguish API failure (5xx, 429, network) from genuine
# permission denial (404 = not a collaborator). Conflating them
# would silently refuse a legitimate maintainer with a public
# 👎 every time GitHub blips. gh_retry handles transient blips;
# the stderr-grep distinguishes 404 from persistent failure.
perm_stderr=$(mktemp)
if permission=$(gh_retry api "repos/${GH_REPO}/collaborators/${COMMENTER}/permission" \
--jq '.permission' 2>"$perm_stderr"); then
echo "Commenter permission: ${permission}"
case "${permission}" in
admin|write|maintain)
echo "allowed=true" >> "$GITHUB_OUTPUT"
;;
*)
echo "allowed=false" >> "$GITHUB_OUTPUT"
;;
esac
else
err=$(cat "$perm_stderr")
echo "Permission lookup stderr: ${err}" >&2
# 404 (not a collaborator) is a genuine deny.
# Anything else is a transient API/network failure.
if grep -qE "HTTP 404|Not Found" "$perm_stderr"; then
echo "allowed=false" >> "$GITHUB_OUTPUT"
else
echo "::error::Permission lookup failed transiently — refusing to act."
echo "allowed=api-failed" >> "$GITHUB_OUTPUT"
fi
fi
- name: React 😕 on transient permission-API failure
if: steps.body.outputs.match == 'true' && steps.perm.outputs.allowed == 'api-failed'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
GH_REPO: ${{ github.repository }}
COMMENT_ID: ${{ github.event.comment.id }}
PR: ${{ github.event.issue.number }}
RUN_ID: ${{ github.run_id }}
shell: bash
run: |
set -euo pipefail
gh api -X POST "repos/${GH_REPO}/issues/comments/${COMMENT_ID}/reactions" \
-f content="confused" >/dev/null
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" \
-f body="⚠️ Couldn't verify your repo permission (transient GitHub API failure). Please comment \`/autofix\` again. ([apply run](https://github.com/${GH_REPO}/actions/runs/${RUN_ID}))" \
>/dev/null
exit 1
- name: React 👎 on permission denial
if: steps.body.outputs.match == 'true' && steps.perm.outputs.allowed == 'false'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
GH_REPO: ${{ github.repository }}
COMMENT_ID: ${{ github.event.comment.id }}
PR: ${{ github.event.issue.number }}
shell: bash
run: |
set -euo pipefail
gh api -X POST "repos/${GH_REPO}/issues/comments/${COMMENT_ID}/reactions" \
-f content="-1" >/dev/null
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" \
-f body="🚫 \`/autofix\` is restricted to users with write access or the PR author. Comment ignored." \
>/dev/null
# Hard exit so the rest of the job is skipped.
exit 1
- name: React 👀 to acknowledge
if: steps.perm.outputs.allowed == 'true'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
GH_REPO: ${{ github.repository }}
COMMENT_ID: ${{ github.event.comment.id }}
shell: bash
run: |
set -euo pipefail
gh api -X POST "repos/${GH_REPO}/issues/comments/${COMMENT_ID}/reactions" \
-f content="eyes" >/dev/null
- name: Resolve PR head and locate autofix run
id: locate
if: steps.perm.outputs.allowed == 'true'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
GH_REPO: ${{ github.repository }}
PR: ${{ github.event.issue.number }}
shell: bash
run: |
set -euo pipefail
# Same retry wrapper used in the permission step, repeated
# because each YAML `run:` block is a fresh bash session.
gh_retry() {
local n=0 max=3
while true; do
if gh "$@"; then return 0; fi
n=$((n+1))
if [ "$n" -ge "$max" ]; then return 1; fi
sleep $((n * 2))
done
}
# Fetch PR metadata. All fields here are server-controlled API
# output, but we still allowlist before exporting so anything
# weird short-circuits before $GITHUB_OUTPUT. Wrapped in
# gh_retry so transient blips don't surface as "no autofix run
# found" with a wrong remediation.
if ! pr_json=$(gh_retry api "repos/${GH_REPO}/pulls/${PR}"); then
echo "::error::PR metadata fetch failed after retries."
echo "found_status=api-failed" >> "$GITHUB_OUTPUT"
exit 0
fi
head_sha=$(jq -r '.head.sha' <<< "${pr_json}")
head_ref=$(jq -r '.head.ref' <<< "${pr_json}")
head_repo=$(jq -r '.head.repo.full_name' <<< "${pr_json}")
[[ "${head_sha}" =~ ^[0-9a-f]{40}$ ]] || { echo "::error::Bad head_sha"; exit 1; }
[[ "${head_ref}" =~ ^[A-Za-z0-9._/-]+$ ]] || { echo "::error::Bad head_ref"; exit 1; }
[[ "${head_repo}" =~ ^[A-Za-z0-9._-]+/[A-Za-z0-9._-]+$ ]] || { echo "::error::Bad head_repo"; exit 1; }
# Find the latest successful pr-autofix.yml run for this head SHA.
if ! runs_json=$(gh_retry api "repos/${GH_REPO}/actions/workflows/pr-autofix.yml/runs?head_sha=${head_sha}&per_page=10"); then
echo "::error::Workflow run lookup failed after retries."
echo "found_status=api-failed" >> "$GITHUB_OUTPUT"
exit 0
fi
run_id=$(jq -r '[.workflow_runs[] | select(.conclusion == "success")] | .[0].id // empty' <<< "${runs_json}")
if [ -n "${run_id}" ] && [[ "${run_id}" =~ ^[0-9]+$ ]]; then
echo "found_status=success" >> "$GITHUB_OUTPUT"
{
echo "found=true"
echo "head_sha=${head_sha}"
echo "head_ref=${head_ref}"
echo "head_repo=${head_repo}"
echo "run_id=${run_id}"
} >> "$GITHUB_OUTPUT"
exit 0
fi
# No successful run. Distinguish "still running" (producer in
# flight after a recent push) from "never ran / all failed".
# in_progress / queued / pending / waiting cover the GitHub
# workflow-run lifecycle states that precede success/failure.
in_progress=$(jq -r '[.workflow_runs[] | select(.status == "in_progress" or .status == "queued" or .status == "pending" or .status == "waiting")] | length' <<< "${runs_json}")
if [ "${in_progress:-0}" -gt 0 ]; then
echo "::warning::pr-autofix run is still in progress for head ${head_sha}."
echo "found_status=in-progress" >> "$GITHUB_OUTPUT"
else
echo "::warning::No successful pr-autofix run found for head ${head_sha}."
echo "found_status=not-found" >> "$GITHUB_OUTPUT"
fi
# Existing `found` boolean is preserved so downstream gates
# (`steps.locate.outputs.found == 'true'`) still work.
echo "found=false" >> "$GITHUB_OUTPUT"
- name: Reply when locate did not yield a usable run
if: steps.perm.outputs.allowed == 'true' && steps.locate.outputs.found != 'true'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
GH_REPO: ${{ github.repository }}
COMMENT_ID: ${{ github.event.comment.id }}
PR: ${{ github.event.issue.number }}
FOUND_STATUS: ${{ steps.locate.outputs.found_status }}
RUN_ID: ${{ github.run_id }}
shell: bash
run: |
set -euo pipefail
run_url="https://github.com/${GH_REPO}/actions/runs/${RUN_ID}"
case "${FOUND_STATUS}" in
in-progress)
gh api -X POST "repos/${GH_REPO}/issues/comments/${COMMENT_ID}/reactions" \
-f content="confused" >/dev/null
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" \
-f body="⏳ A pr-autofix run is still in progress for this PR's current head SHA. Wait for it to finish, then comment \`/autofix\` again. ([apply run](${run_url}))" \
>/dev/null
;;
api-failed)
gh api -X POST "repos/${GH_REPO}/issues/comments/${COMMENT_ID}/reactions" \
-f content="confused" >/dev/null
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" \
-f body="⚠️ Couldn't reach the GitHub API to look up the autofix run (transient failure after retries). Please comment \`/autofix\` again. ([apply run](${run_url}))" \
>/dev/null
;;
*)
gh api -X POST "repos/${GH_REPO}/issues/comments/${COMMENT_ID}/reactions" \
-f content="-1" >/dev/null
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" \
-f body="🤔 No successful autofix run found for this PR's current head SHA. Push a new commit to trigger one, then comment \`/autofix\` again." \
>/dev/null
;;
esac
exit 1
# Pinned to v8.0.1. Same SHA as pr-autofix-publish.yml.
# `continue-on-error: true` lets the workflow proceed when the
# artifact is expired or pruned (1-day retention). The apply
# step distinguishes "patch file missing entirely" (artifact-
# expired) from "patch file zero bytes" (genuinely empty patch).
- name: Download autofix artifact
if: steps.locate.outputs.found == 'true'
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
continue-on-error: true
with:
name: autofix
run-id: ${{ steps.locate.outputs.run_id }}
github-token: ${{ secrets.GITHUB_TOKEN }}
path: autofix-in
# Pinned to v5.0.4. Verify SHA via:
# gh api repos/actions/checkout/git/refs/tags/v5.0.4
#
# `persist-credentials: false` disables the default behavior where
# actions/checkout writes the GITHUB_TOKEN into `.git/config` as an
# extraheader. That default is convenient (subsequent git commands
# auth automatically) but it means the token is sitting on disk in
# the checkout directory — an `actions/upload-artifact` step on
# this directory would leak the token. We don't upload, but
# zizmor's `credential-persistence` lint flags it defensively.
# Push auth is provided inline at push time via the URL.
- name: Checkout PR head
if: steps.locate.outputs.found == 'true'
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v5.0.4
with:
repository: ${{ steps.locate.outputs.head_repo }}
ref: ${{ steps.locate.outputs.head_sha }}
token: ${{ secrets.GITHUB_TOKEN }}
persist-credentials: false
# Fetch full history so the push doesn't hit shallow-clone errors.
fetch-depth: 0
path: pr-checkout
- name: Apply patch and push
id: apply
if: steps.locate.outputs.found == 'true'
env:
HEAD_REF: ${{ steps.locate.outputs.head_ref }}
HEAD_REPO: ${{ steps.locate.outputs.head_repo }}
# The SHA we resolved earlier in `locate` — this is what the
# remote ref MUST still equal at push time. If the contributor
# force-pushed between resolve and now, the lease fails and
# we surface that distinctly from a fork-without-maintainer
# -edit push failure.
HEAD_SHA: ${{ steps.locate.outputs.head_sha }}
# Auth for the push only — never persisted to disk. Provided
# via env to avoid interpolating into the shell command line.
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
shell: bash
working-directory: pr-checkout
run: |
set -euo pipefail
patch="../autofix-in/autofix.patch"
# Distinguish artifact-expired (file missing entirely, because
# actions/download-artifact ran with continue-on-error and the
# 1-day retention had elapsed) from genuinely empty patch
# (file present, zero bytes, formatter found nothing).
if [ ! -e "$patch" ]; then
echo "::warning::Patch file does not exist — autofix artifact likely expired."
echo "result=artifact-expired" >> "$GITHUB_OUTPUT"
exit 0
fi
if [ ! -s "$patch" ]; then
echo "::warning::Empty patch — nothing to apply."
echo "result=empty-patch" >> "$GITHUB_OUTPUT"
exit 0
fi
# Sensitive-paths guard: refuse to apply patches that touch
# `.github/` — workflow files, action definitions, CODEOWNERS,
# dependabot config, etc. A malicious PR could ship a custom
# prettier/ESLint config that reformats workflow YAML; the
# producer would then capture those edits in autofix.patch,
# and a maintainer running `/autofix` would push them under
# `contents: write`. The default GITHUB_TOKEN lacks `workflows`
# scope so the platform would reject workflow-file pushes
# anyway, but that surfaces as a generic `push-failed` and
# misleads users into enabling maintainer-edit. Reject early
# with a specific reason. CODEOWNERS and dependabot.yml live
# under .github/ but outside .github/workflows/ — the broader
# match is intentional (they all govern trust boundaries).
if grep -qE '^(diff --git|---|\+\+\+) [ab]?/?\.github/' "$patch"; then
echo "::warning::Patch touches .github/ — refusing to apply (sensitive paths)."
echo "result=sensitive-paths" >> "$GITHUB_OUTPUT"
exit 0
fi
# Re-entrancy guard: if HEAD itself is an autofix bot commit,
# refuse to apply again. Without this, lint/formatter config
# drift between runs could pump arbitrary apply commits into
# the same PR if an automated agent watches the sticky and
# re-fires `/autofix` on each new "fixes-available" surface.
# The contributor can still get out by force-pushing a
# human-authored commit to revert the autofix and re-trigger.
head_author=$(git log -1 --format='%ae' HEAD)
head_subject=$(git log -1 --format='%s' HEAD)
if [ "${head_author}" = "41898282+github-actions[bot]@users.noreply.github.com" ] \
&& [[ "${head_subject}" =~ ^chore\(autofix\) ]]; then
echo "::warning::HEAD is an autofix bot commit — refusing to re-apply (loop guard)."
echo "result=loop-prevented" >> "$GITHUB_OUTPUT"
exit 0
fi
# Idempotency probe: does the forward apply work?
if git apply --check "$patch" 2>/dev/null; then
echo "Patch applies cleanly — proceeding."
elif git apply --check --reverse "$patch" 2>/dev/null; then
# Reverse-check passes => the patch is already applied to
# the current tree. Treat as success no-op.
echo "Patch is already applied (reverse-check passed) — no-op."
echo "result=already-applied" >> "$GITHUB_OUTPUT"
exit 0
else
echo "::error::Patch does not apply (stale or conflicting)."
echo "result=stale" >> "$GITHUB_OUTPUT"
exit 0
fi
# Wrap the apply/commit phase so any non-zero exit sets a
# meaningful `result=` instead of leaving it unset (which would
# send the user to the `*` "unexpected state" arm with a
# non-actionable confused-emoji reply).
if ! {
git config user.email "41898282+github-actions[bot]@users.noreply.github.com" &&
git config user.name "github-actions[bot]" &&
git apply "$patch" &&
git add -A &&
git commit -m "chore(autofix): apply prettier + eslint fixes via /autofix command"
}; then
echo "::error::git apply / config / commit failed after idempotency probe passed."
echo "result=apply-failed" >> "$GITHUB_OUTPUT"
exit 0
fi
# Push to the PR head branch with a lease against the resolved
# SHA. The lease ensures the remote ref still points at HEAD_SHA
# when the push lands — if the contributor force-pushed in the
# window between resolve and now, the lease fails and we return
# `lease-failed` (NOT `push-failed`, which would mislead users
# into enabling maintainer-edit). For fork PRs, the push still
# requires "Allow edits by maintainers" to be enabled.
#
# Auth is supplied inline via `-c http.<base>.extraheader` (NOT
# via a `https://x-access-token:TOKEN@…` URL — those leak into
# process listings and `git remote -v` output). The header is
# set per-invocation; it never lands in `.git/config` on disk.
# The token is base64-encoded for the Basic auth header per
# GitHub's documented pattern for this scope.
push_url="https://github.com/${HEAD_REPO}.git"
auth_header="Authorization: Basic $(printf 'x-access-token:%s' "${GITHUB_TOKEN}" | base64 -w0)"
# GitHub's secret-masker only masks the raw token, not its
# base64-encoded form. Mask the encoded value so any subsequent
# log line (set -x, GIT_TRACE, error spew) gets ***-redacted.
echo "::add-mask::${auth_header}"
push_stderr=$(mktemp)
if git -c http.extraheader="${auth_header}" \
push --force-with-lease="refs/heads/${HEAD_REF}:${HEAD_SHA}" \
"${push_url}" "HEAD:${HEAD_REF}" 2>"$push_stderr"; then
echo "result=applied" >> "$GITHUB_OUTPUT"
else
cat "$push_stderr" >&2
# `--force-with-lease` reports "stale info" when the remote
# ref has moved past the expected SHA. Other lease-failure
# phrases git emits include "remote rejected" (server-side
# reject), "non-fast-forward", and the literal flag name. Match
# any of those to distinguish from auth/network/maintainer-
# edit failures.
if grep -qE "stale info|force-with-lease|rejected.*non-fast-forward|remote rejected|! \[rejected\]" "$push_stderr"; then
echo "::error::git push lease failed — branch moved during apply."
echo "result=lease-failed" >> "$GITHUB_OUTPUT"
else
echo "::error::git push failed — likely fork without maintainer-edit enabled."
echo "result=push-failed" >> "$GITHUB_OUTPUT"
fi
exit 0
fi
- name: React and reply on outcome
if: always() && steps.locate.outputs.found == 'true' && steps.apply.outcome != 'skipped'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
GH_REPO: ${{ github.repository }}
COMMENT_ID: ${{ github.event.comment.id }}
PR: ${{ github.event.issue.number }}
RESULT: ${{ steps.apply.outputs.result }}
RUN_ID: ${{ github.run_id }}
shell: bash
run: |
set -euo pipefail
run_url="https://github.com/${GH_REPO}/actions/runs/${RUN_ID}"
case "${RESULT}" in
applied)
gh api -X POST "repos/${GH_REPO}/issues/comments/${COMMENT_ID}/reactions" \
-f content="+1" >/dev/null
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" \
-f body="✅ Applied autofix and pushed a commit. ([apply run](${run_url}))" \
>/dev/null
;;
already-applied)
gh api -X POST "repos/${GH_REPO}/issues/comments/${COMMENT_ID}/reactions" \
-f content="+1" >/dev/null
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" \
-f body="✅ Autofix is already applied — no changes needed." \
>/dev/null
;;
empty-patch)
gh api -X POST "repos/${GH_REPO}/issues/comments/${COMMENT_ID}/reactions" \
-f content="+1" >/dev/null
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" \
-f body="✅ No autofix to apply — formatter found nothing." \
>/dev/null
;;
artifact-expired)
gh api -X POST "repos/${GH_REPO}/issues/comments/${COMMENT_ID}/reactions" \
-f content="confused" >/dev/null
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" \
-f body="⏳ The autofix artifact for this PR's head SHA has expired (1-day retention). Push a new commit to regenerate it, then comment \`/autofix\` again. ([apply run](${run_url}))" \
>/dev/null
exit 1
;;
loop-prevented)
gh api -X POST "repos/${GH_REPO}/issues/comments/${COMMENT_ID}/reactions" \
-f content="confused" >/dev/null
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" \
-f body="🔁 Refusing to re-apply autofix on top of an existing autofix commit. If formatter rules drifted and you genuinely need another pass, push a human-authored commit (or revert the existing autofix commit) before commenting \`/autofix\` again. ([apply run](${run_url}))" \
>/dev/null
exit 1
;;
sensitive-paths)
gh api -X POST "repos/${GH_REPO}/issues/comments/${COMMENT_ID}/reactions" \
-f content="-1" >/dev/null
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" \
-f body="🛑 Refusing to apply: the autofix patch touches files under \`.github/\` (workflow / CODEOWNERS / dependabot config). Apply formatter changes to those files manually in a regular commit so they get human review. ([apply run](${run_url}))" \
>/dev/null
exit 1
;;
stale)
gh api -X POST "repos/${GH_REPO}/issues/comments/${COMMENT_ID}/reactions" \
-f content="-1" >/dev/null
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" \
-f body="⚠️ The autofix patch is stale or conflicts with the current head — push a new commit to regenerate, then comment \`/autofix\` again. ([apply run](${run_url}))" \
>/dev/null
exit 1
;;
apply-failed)
gh api -X POST "repos/${GH_REPO}/issues/comments/${COMMENT_ID}/reactions" \
-f content="-1" >/dev/null
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" \
-f body="⚠️ Autofix applied cleanly in the dry run, but \`git apply\` / \`git commit\` failed when actually landing the patch. This usually means a race with concurrent edits or a corrupt patch. See logs: ${run_url}" \
>/dev/null
exit 1
;;
push-failed)
gh api -X POST "repos/${GH_REPO}/issues/comments/${COMMENT_ID}/reactions" \
-f content="-1" >/dev/null
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" \
-f body="⚠️ Couldn't push the autofix commit. If this is a fork PR, please tick **Allow edits by maintainers** in the PR sidebar, then comment \`/autofix\` again. ([apply run](${run_url}))" \
>/dev/null
exit 1
;;
lease-failed)
gh api -X POST "repos/${GH_REPO}/issues/comments/${COMMENT_ID}/reactions" \
-f content="-1" >/dev/null
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" \
-f body="⚠️ The PR head moved while autofix was applying — a new commit landed in the window between resolve and push. Comment \`/autofix\` again to retry against the latest head. ([apply run](${run_url}))" \
>/dev/null
exit 1
;;
*)
gh api -X POST "repos/${GH_REPO}/issues/comments/${COMMENT_ID}/reactions" \
-f content="confused" >/dev/null
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" \
-f body="❓ Autofix run finished in an unexpected state (\`${RESULT:-unknown}\`). See logs: ${run_url}" \
>/dev/null
exit 1
;;
esac
+316
View File
@@ -0,0 +1,316 @@
name: PR Autofix (publish)
# TRUSTED HALF of the autofix pipeline.
#
# Triggered by `pr-autofix.yml` completing on a PR (including fork PRs).
# Downloads the diff artifact produced by the untrusted job, verifies
# its claimed PR identity against the workflow_run authority, then
# posts (or edits) a single sticky summary comment plus a
# `gitnexus/autofix` Check Run. This job NEVER checks out fork code —
# it only consumes the diff (data) and calls the GitHub API. That
# isolation is what makes it safe to run under `pull-requests: write`
# on fork-triggered events.
#
# The sticky comment is the contributor signal: heading
# "## :sparkles: PR Autofix" in the PR's top-level comments, with a
# fenced `gitnexus-autofix` JSON block carrying machine-readable state
# for AI agents. Contributors apply the patch by commenting `/autofix`
# on the PR — handled by the separate `pr-autofix-apply.yml` workflow.
on:
workflow_run:
workflows: ['PR Autofix']
types: [completed]
concurrency:
# Key on PR identity, NOT workflow_run.id — workflow_run.id is per-run
# unique, which would defeat serialization and let two parallel
# publishes both POST a sticky summary comment. CONTRIBUTING.md
# § GitHub Actions — Concurrency Convention names this anti-pattern
# explicitly. For fork PRs, `pull_requests[]` is empty in the
# workflow_run payload, so we fall back to head-repo + head-branch.
group: ${{ github.workflow }}-${{ github.event.workflow_run.pull_requests[0].number || format('{0}/{1}', github.event.workflow_run.head_repository.full_name, github.event.workflow_run.head_branch) }}
cancel-in-progress: false
permissions: {}
jobs:
publish:
name: publish-autofix
if: >-
github.event.workflow_run.event == 'pull_request'
&& github.event.workflow_run.conclusion == 'success'
runs-on: ubuntu-latest
timeout-minutes: 5
permissions:
pull-requests: write
# Required by actions/download-artifact to fetch artifacts produced
# by a different workflow run.
actions: read
# Required to create the `gitnexus/autofix` Check Run that reports
# the outcome (clean / fixes-available) to the PR's Checks tab.
# Branch protection or agents can grep the conclusion + output
# title without parsing the sticky comment.
checks: write
steps:
# Pinned to v8.0.1. Verify SHA via:
# gh api repos/actions/download-artifact/git/refs/tags/v8.0.1
- name: Download autofix artifact
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
with:
name: autofix
run-id: ${{ github.event.workflow_run.id }}
github-token: ${{ secrets.GITHUB_TOKEN }}
path: autofix-in
- name: Read and validate metadata
id: meta
shell: bash
run: |
set -euo pipefail
test -f autofix-in/metadata.json
jq . autofix-in/metadata.json
# The artifact comes from the untrusted half running fork code.
# Every field is allowlist-validated before it can flow into
# $GITHUB_OUTPUT. A newline in head_ref would otherwise let a
# malicious branch name inject a second `pr_number=N` line and
# redirect this job's reviewdog suggestions / sticky summary
# comment onto a victim PR under github-actions[bot] with
# pull-requests: write.
assert_field() {
local key="$1" pattern="$2" value
value=$(jq -r ".${key} // empty" autofix-in/metadata.json)
if [ -z "$value" ] || ! [[ "$value" =~ $pattern ]]; then
echo "::error::metadata.${key} failed allowlist (got: $(printf '%q' "$value"))"
exit 1
fi
printf '%s' "$value"
}
SCHEMA=$(assert_field schema '^gitnexus\.pr-autofix/v[0-9]+$')
PR_NUMBER=$(assert_field pr_number '^[0-9]+$')
HEAD_SHA=$(assert_field head_sha '^[0-9a-f]{40}$')
HEAD_REF=$(assert_field head_ref '^[A-Za-z0-9._/-]+$')
HEAD_REPO=$(assert_field head_repo '^[A-Za-z0-9._-]+/[A-Za-z0-9._-]+$')
BASE_REPO=$(assert_field base_repo '^[A-Za-z0-9._-]+/[A-Za-z0-9._-]+$')
CHANGED=$(assert_field changed_lines '^[0-9]+$')
# Defence-in-depth: refuse to act if the artifact claims to
# belong to a different repo than the one that triggered us.
if [ "$BASE_REPO" != "${GITHUB_REPOSITORY}" ]; then
echo "::error::Artifact base_repo does not match \$GITHUB_REPOSITORY — refusing to publish."
exit 1
fi
{
echo "schema=${SCHEMA}"
echo "pr_number=${PR_NUMBER}"
echo "head_sha=${HEAD_SHA}"
echo "head_ref=${HEAD_REF}"
echo "head_repo=${HEAD_REPO}"
echo "base_repo=${BASE_REPO}"
echo "changed_lines=${CHANGED}"
} >> "$GITHUB_OUTPUT"
# Cross-verify the artifact's claimed identity against the
# GitHub-controlled workflow_run event. The previous step's
# allowlist only proves the fields are well-formed — not that
# they refer to the PR/SHA that actually triggered this run.
# A fork-controlled `npm run lint:fix` could plausibly mutate
# metadata.json to reference another PR or SHA, redirecting our
# write-scoped sticky/check-run onto an attacker-chosen target.
#
# Authority sources are all server-controlled GitHub event fields:
# - workflow_run.head_sha
# - workflow_run.head_repository.full_name
# - workflow_run.pull_requests[].number (within-repo PRs only;
# empty array on fork PRs — fall back to commits/{sha}/pulls)
#
# Mismatch => fail loud BEFORE any sticky/check-run side effect.
- name: Verify metadata against workflow_run authority
id: verify
if: steps.meta.outputs.changed_lines != '0'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
GH_REPO: ${{ github.repository }}
META_PR_NUMBER: ${{ steps.meta.outputs.pr_number }}
META_HEAD_SHA: ${{ steps.meta.outputs.head_sha }}
META_HEAD_REPO: ${{ steps.meta.outputs.head_repo }}
WF_HEAD_SHA: ${{ github.event.workflow_run.head_sha }}
WF_HEAD_REPO: ${{ github.event.workflow_run.head_repository.full_name }}
WF_PR_NUMBERS: ${{ toJSON(github.event.workflow_run.pull_requests.*.number) }}
shell: bash
run: |
set -euo pipefail
# 1) head_sha must match exactly. workflow_run.head_sha is the
# commit GitHub actually ran the producer against — definitive.
if [ "${META_HEAD_SHA}" != "${WF_HEAD_SHA}" ]; then
echo "::error::Artifact head_sha (${META_HEAD_SHA}) does not match workflow_run.head_sha (${WF_HEAD_SHA}) — refusing to publish."
exit 1
fi
# 2) head_repo must match exactly. Same authority anchor.
if [ "${META_HEAD_REPO}" != "${WF_HEAD_REPO}" ]; then
echo "::error::Artifact head_repo (${META_HEAD_REPO}) does not match workflow_run.head_repository (${WF_HEAD_REPO}) — refusing to publish."
exit 1
fi
# 3) pr_number must reference an open PR with this head SHA.
# Within-repo PRs: workflow_run.pull_requests[] is populated.
# Fork PRs: that array is empty by GitHub design — fall back
# to the REST commit-to-PRs lookup. Fail closed if the lookup
# finds no matching open PR (avoids attacker-forged PR ids).
allowed_numbers=$(jq -c '.' <<< "${WF_PR_NUMBERS}")
if [ "${allowed_numbers}" = "[]" ]; then
echo "workflow_run.pull_requests is empty (fork PR) — falling back to commits/{sha}/pulls."
allowed_numbers=$(gh api "repos/${GH_REPO}/commits/${WF_HEAD_SHA}/pulls" \
--jq '[.[] | select(.state == "open") | .number]' 2>/dev/null || echo "[]")
if [ "${allowed_numbers}" = "[]" ]; then
echo "::error::No open PR found for head ${WF_HEAD_SHA} via commits/{sha}/pulls — refusing to publish."
exit 1
fi
fi
if ! jq -e --argjson n "${META_PR_NUMBER}" 'index($n) != null' <<< "${allowed_numbers}" >/dev/null; then
echo "::error::Artifact pr_number (${META_PR_NUMBER}) is not in the authoritative PR list (${allowed_numbers}) — refusing to publish."
exit 1
fi
echo "Verified: metadata identity matches workflow_run authority (PR=${META_PR_NUMBER}, head_sha=${META_HEAD_SHA}, head_repo=${META_HEAD_REPO})."
- name: Upsert sticky summary comment
# Only post when ci-quality found something fixable (= the
# autofix patch is non-empty). When prettier/eslint are clean
# the patch is zero bytes and the sticky comment is pure noise,
# so we skip it.
if: >-
always()
&& steps.meta.outputs.pr_number != ''
&& steps.meta.outputs.changed_lines != '0'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
GH_REPO: ${{ github.repository }}
PR: ${{ steps.meta.outputs.pr_number }}
CHANGED: ${{ steps.meta.outputs.changed_lines }}
HEAD_SHA: ${{ steps.meta.outputs.head_sha }}
RUN_ID: ${{ github.run_id }}
shell: bash
run: |
set -euo pipefail
# Stable heading + marker — agents grep for these exact strings.
marker="<!-- gitnexus:pr-autofix-summary -->"
heading="## :sparkles: PR Autofix"
# Single state. The /autofix slash command works for any diff
# size — there's no 3K cap and no no-overlap dead-end because
# the apply workflow uses `git apply` + push, not the GitHub
# review-comment API.
ui_state="fixes-available"
prose="Found fixable formatting / unused-import issues across **${CHANGED}** changed lines. **Comment \`/autofix\` on this PR to apply them**, or run \`npm run lint:fix && npm run format\` locally."
# Machine-readable JSON block — agents parse this instead of
# regexing English. Fenced code-block info string is
# `gitnexus-autofix` so agents can locate it without ambiguity.
# Schema bumped from v1 -> v2: adds `apply_command`. The v1
# field set is preserved as a superset, but the `state` enum
# is redefined (v1: suggestions-posted | skipped-too-large |
# diff-no-overlap; v2: fixes-available). v1 readers checking
# `schema == 'gitnexus.pr-autofix/v1'` see an unfamiliar version
# and fall back to prose, which is the intended migration path.
json=$(jq -n -c \
--arg state "${ui_state}" \
--argjson pr_number "${PR}" \
--argjson changed_lines "${CHANGED}" \
--arg head_sha "${HEAD_SHA}" \
--arg run_id "${RUN_ID}" \
'{schema:"gitnexus.pr-autofix/v2", state:$state, pr_number:$pr_number, changed_lines:$changed_lines, head_sha:$head_sha, run_id:$run_id, apply_command:"/autofix"}')
# Multi-line quoted string instead of a column-0 heredoc — YAML's
# `run: |` block ends as soon as a content line dedents below the
# block's first-line indent, which would mis-parse the workflow.
body="${marker}
${heading}
${prose}
\`\`\`gitnexus-autofix
${json}
\`\`\`"
# Strip the leading 10-space indent that the YAML block requires
# so the rendered comment body starts at column 0.
body="$(printf '%s\n' "$body" | sed 's/^ //')"
# Small retry wrapper for transient 5xx / rate-limit responses
# on the GitHub REST API. Three tries with linear backoff. We
# only retry GET (idempotent) and PATCH on a known comment id
# (idempotent). POST is NOT wrapped — retrying a comment-create
# would create duplicates if the first attempt actually landed.
gh_retry() {
local n=0 max=3
while true; do
if gh "$@"; then return 0; fi
n=$((n+1))
if [ "$n" -ge "$max" ]; then return 1; fi
sleep $((n * 2))
done
}
# Find existing bot comment by the marker and edit-in-place; else create.
# CRITICAL: filter by `.user.login == "github-actions[bot]"`. A regular
# user posting a comment containing the marker would otherwise be the
# `head -n1` match; PATCH on someone else's comment 403s, `set -e`
# aborts, and the bot is permanently DoS'd for that PR.
existing=$(gh_retry api "repos/${GH_REPO}/issues/${PR}/comments" \
--paginate --jq ".[] | select(.user.login == \"github-actions[bot]\" and (.body | contains(\"${marker}\"))) | .id" \
| head -n1 || true)
if [ -n "${existing}" ]; then
gh_retry api -X PATCH "repos/${GH_REPO}/issues/comments/${existing}" \
-f body="${body}" >/dev/null
echo "Updated comment ${existing}."
else
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" \
-f body="${body}" >/dev/null
echo "Created summary comment."
fi
- name: Emit gitnexus/autofix Check Run
# Stable check name `gitnexus/autofix` so PR-watching agents can
# `gh pr checks <pr>` and read the conclusion + title without
# parsing the sticky comment. Two outcomes:
# clean → conclusion: success
# fixes-available → conclusion: neutral
# `neutral` does not block branch-protection required-checks but
# is visually distinct from a green pass.
if: always() && steps.meta.outputs.head_sha != ''
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
GH_REPO: ${{ github.repository }}
HEAD_SHA: ${{ steps.meta.outputs.head_sha }}
CHANGED: ${{ steps.meta.outputs.changed_lines }}
shell: bash
run: |
set -euo pipefail
if [ "${CHANGED}" = "0" ]; then
conclusion="success"
title="Formatting clean"
summary="Prettier and ESLint --fix produced no changes."
else
conclusion="neutral"
title="Autofix available — comment /autofix to apply"
summary="Comment \`/autofix\` on this PR to apply formatter + unused-import fixes (works at any diff size). Or run \`npm run lint:fix && npm run format\` locally."
fi
gh api -X POST "repos/${GH_REPO}/check-runs" \
-f name="gitnexus/autofix" \
-f head_sha="${HEAD_SHA}" \
-f status="completed" \
-f conclusion="${conclusion}" \
-f "output[title]=${title}" \
-f "output[summary]=${summary}" \
>/dev/null
echo "Posted check-run gitnexus/autofix=${conclusion} (${title})"
+146
View File
@@ -0,0 +1,146 @@
name: PR Autofix
# UNTRUSTED HALF of the autofix pipeline.
#
# Runs `npm run lint:fix` + `npm run format` against the PR head
# (including fork heads) and uploads the resulting diff as an artifact.
# This job has NO privileged token and CANNOT post to the PR. The trusted
# `pr-autofix-publish.yml` workflow downloads the artifact via
# `workflow_run` and posts a sticky summary comment + Check Run.
# Contributors apply the patch by commenting `/autofix` on the PR —
# handled by the separate `pr-autofix-apply.yml` ChatOps workflow.
#
# Why the split:
# ESLint loads plugins from fork-controlled `node_modules`, so running
# it in a job with `pull-requests: write` would let a malicious fork PR
# ship a poisoned eslint plugin and execute arbitrary code under that
# token. By keeping fork code execution in this job (token: read-only)
# and posting from a separate trusted job that never touches fork
# code, we get the autofix UX for fork PRs without the supply-chain
# hole. (See autofix.ci for the same pattern.)
#
# Removes unused imports via `eslint-plugin-unused-imports`, already in
# devDependencies and wired into the `lint` config.
on:
pull_request:
types: [opened, synchronize, reopened]
# Skip lockfile / generated-file PRs entirely — `action-suggester`
# cannot post on diffs > ~3k lines (GitHub returns 406) and these
# paths produce massive diffs no human wants suggested back inline.
paths-ignore:
- '**/package-lock.json'
- '**/*.snap'
- '**/dist/**'
- '**/node_modules/**'
concurrency:
group: ${{ github.workflow }}-${{ github.event.pull_request.number }}
# Don't cancel in-flight runs; the publish workflow may already be
# downloading the artifact and a cancelled untrusted run produces no
# signal at all (worse DX than waiting).
cancel-in-progress: false
# This workflow runs untrusted fork code. Top-level deny-all and NO
# job-level grants — the job can only read its own checkout.
permissions: {}
jobs:
autofix:
name: autofix
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
# PR head commit (not the synthetic merge ref) — we need the
# exact tree the contributor pushed so suggestions line up.
ref: ${{ github.event.pull_request.head.sha }}
repository: ${{ github.event.pull_request.head.repo.full_name }}
persist-credentials: false
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
with:
node-version: 22
cache: npm
cache-dependency-path: package-lock.json
# `--ignore-scripts` blocks pre/postinstall lifecycle hooks. ESLint
# plugins still load from node_modules (that is the actual escape
# hatch on a typical fork), but this job has no token to abuse —
# which is the whole point of the split.
- run: npm ci --ignore-scripts
- name: ESLint --fix (removes unused imports)
run: npm run lint:fix
# Lint errors that --fix can't auto-resolve must not block the
# diff artifact — partial fixes are still useful as suggestions.
continue-on-error: true
- name: Prettier --write
run: npm run format
continue-on-error: true
- name: Capture diff and metadata
id: capture
# Pass GitHub-context values via env: rather than `${{ }}`
# interpolated directly into the bash body. `head.ref` and
# `head.repo.full_name` are fork-controlled strings; expanding
# them into shell source is the canonical template-injection
# vector zizmor flags. Even though this job has `permissions: {}`,
# routing through env: makes it impossible for a future scope
# grant to turn into RCE. Inside bash, reference as `$HEAD_REF`
# etc. — the values are then plain strings, not code.
env:
PR_NUMBER: ${{ github.event.pull_request.number }}
HEAD_SHA: ${{ github.event.pull_request.head.sha }}
HEAD_REF: ${{ github.event.pull_request.head.ref }}
HEAD_REPO: ${{ github.event.pull_request.head.repo.full_name }}
BASE_REPO: ${{ github.repository }}
shell: bash
run: |
set -euo pipefail
mkdir -p autofix-out
# Produce a unified diff of the working tree vs. the PR head.
# Empty diff => nothing to suggest; the publish job short-circuits.
git diff --no-color > autofix-out/autofix.patch
# NOTE: `changed_lines` is the line-count of the patch file,
# (hunk headers + context lines + added/removed). Surfaced in
# the sticky comment so contributors and AI agents have a
# quick size hint before invoking `/autofix`.
changed_lines=$(wc -l < autofix-out/autofix.patch | tr -d ' ')
echo "changed_lines=${changed_lines}" >> "$GITHUB_OUTPUT"
# Carry PR identity over to the trusted job. workflow_run
# context is base-repo-only, so the publish job needs these
# to call the GitHub PR API on the right resource.
# CONTRACT: keep this schema in sync with pr-autofix-publish.yml's
# `assert_field` validators and the agent-facing JSON block in
# the sticky comment. Bump `schema` when changing field names.
jq -n \
--arg schema 'gitnexus.pr-autofix/v1' \
--argjson pr_number "${PR_NUMBER}" \
--arg head_sha "${HEAD_SHA}" \
--arg head_ref "${HEAD_REF}" \
--arg head_repo "${HEAD_REPO}" \
--arg base_repo "${BASE_REPO}" \
--argjson changed_lines "${changed_lines}" \
'{schema:$schema, pr_number:$pr_number, head_sha:$head_sha, head_ref:$head_ref, head_repo:$head_repo, base_repo:$base_repo, changed_lines:$changed_lines}' \
> autofix-out/metadata.json
echo "--- metadata ---"
cat autofix-out/metadata.json
echo "--- diff (head) ---"
head -c 2000 autofix-out/autofix.patch || true
# Pinned to v7.0.1. Verify SHA via:
# gh api repos/actions/upload-artifact/git/refs/tags/v7.0.1
- name: Upload autofix artifact
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: autofix
path: autofix-out/
retention-days: 1
if-no-files-found: error
+4 -1
View File
@@ -35,6 +35,9 @@ on:
pull_request_target:
types: [opened, edited, reopened]
permissions:
contents: read
# Concurrency convention: see CONTRIBUTING.md → "GitHub Actions — Concurrency Convention".
# Include `github.event_name` so `pull_request` (validate-title) and
# `pull_request_target` (autolabel) runs for the same PR do NOT share a slot
@@ -105,7 +108,7 @@ jobs:
# Pinned to v7.2.0. Verify SHA via:
# gh api repos/release-drafter/release-drafter/git/refs/tags/v7.2.0
# v7 removed `disable-releaser`; use `dry-run: true` to only autolabel.
- uses: release-drafter/release-drafter@563bf132657a13ded0b01fcb723c5a58cdd824e2 # v7.2.1
- uses: release-drafter/release-drafter@c2e2804cc59f45f57076a99af580d0fedb697927 # v7.3.0
with:
config-name: release-drafter.yml
dry-run: true
+831 -34
View File
@@ -1,61 +1,421 @@
name: Publish to npm
name: Publish
# ─────────────────────────────────────────────────────────────────────────────
# Sole publisher for the `gitnexus` npm package, GitHub Releases, and Docker
# images. Replaces the former two-workflow design — see issue #1609 for the
# double-publish race this unification closes.
#
# Two release modes, both routed through this file:
# • Release candidate (rc) — triggered by push to `main` or workflow_dispatch.
# The RC path computes the next rc version, applies it in-CI, pushes a
# detached release commit with v<X.Y.Z>-rc.<N> + rc/<SHA> marker
# atomically, then publishes to npm with --tag rc and creates a GitHub
# prerelease. RC-only docker.yml invocation follows.
# • Stable — triggered by push of a v<X.Y.Z> tag (no -rc.*
# suffix). Verifies package.json matches the tag, publishes to npm with
# --tag latest, creates a stable GitHub Release. No docker (RC-only).
#
# ⚠️ SELF-TRIGGER INVARIANT — DO NOT WEAKEN ⚠️
# The `tags:` filter below uses a negative glob `'!v*-rc.*'` to prevent the
# workflow from re-triggering itself when the RC path pushes its own v-tag.
# Without this exclusion, every RC publish double-fires (the bug fixed by
# #1609). If a NEW prerelease channel is introduced (e.g. `-beta.N`,
# `-alpha.N`, `-next.N`), the negative-glob list MUST be extended in
# lock-step or self-trigger returns. The same invariant applies to the
# `Classify` step further below — its accepted-tag regex must align with
# the trigger filter's exclusion list.
# ─────────────────────────────────────────────────────────────────────────────
on:
push:
branches: [main]
paths-ignore:
- '**.md'
- 'docs/**'
- 'LICENSE'
tags:
# Negative-globbed exclusion of RC tags this workflow itself produces
# (see the SELF-TRIGGER INVARIANT in the header comment).
- 'v*'
- '!v*-rc.*'
workflow_dispatch:
inputs:
bump:
description: >-
Cycle policy. 'auto' (default) continues the active rc cycle on
this branch if there is one, otherwise bumps patch from latest.
Choose 'patch' / 'minor' / 'major' to explicitly start or reset
an rc cycle.
required: false
default: 'auto'
type: choice
options:
- auto
- patch
- minor
- major
force:
description: 'Publish even when HEAD already has an rc marker'
required: false
default: 'false'
type: choice
options:
- 'false'
- 'true'
# Workflow-level deny-all; each job declares the minimum it needs.
permissions: {}
# No workflow-level permissions — scoped per job below.
# Concurrency convention: see CONTRIBUTING.md → "GitHub Actions — Concurrency Convention".
# Tag refs are unique per release, so distinct tags run in parallel. Re-pushes of the
# same tag serialize. cancel-in-progress: false — never cancel a publish mid-flight.
# Distinct refs (refs/heads/main, refs/tags/v*) run in parallel. The
# release-PR-skip in rc-guard is the load-bearing invariant that prevents
# an RC main-push and a stable tag-push colliding on the same release commit.
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: false
jobs:
# ── Phase 1: classify the triggering event into a release mode ─────────────
route:
name: Classify release event
runs-on: ubuntu-latest
timeout-minutes: 2
permissions:
contents: read
outputs:
mode: ${{ steps.classify.outputs.mode }}
head_sha: ${{ steps.classify.outputs.head_sha }}
bump_input: ${{ inputs.bump }}
force_input: ${{ inputs.force }}
steps:
- name: Classify
id: classify
shell: bash
env:
EVENT_NAME: ${{ github.event_name }}
GH_REF: ${{ github.ref }}
GH_REF_NAME: ${{ github.ref_name }}
run: |
set -euo pipefail
HEAD_SHA="${GITHUB_SHA}"
echo "head_sha=${HEAD_SHA}" >> "$GITHUB_OUTPUT"
# Sanitize before logging (annotation-injection defense in depth).
REF_SAFE="${GH_REF//::/__}"
REF_NAME_SAFE="${GH_REF_NAME//::/__}"
echo "event=${EVENT_NAME} ref=${REF_SAFE} ref_name=${REF_NAME_SAFE}"
MODE=""
case "${EVENT_NAME}" in
workflow_dispatch)
# Manual dispatch is only valid on main — that's the only ref
# where a real publish makes sense.
if [ "${GH_REF}" = "refs/heads/main" ]; then
MODE="rc"
else
echo "::error::workflow_dispatch is only permitted on refs/heads/main (got ${REF_SAFE})."
exit 1
fi
;;
push)
case "${GH_REF}" in
refs/heads/main)
MODE="rc"
;;
refs/tags/v*)
# The trigger filter already excluded v*-rc.* tags. Anything
# reaching here is either a stable semver or a malformed v*.
TAG="${GH_REF#refs/tags/}"
if [[ "${TAG}" =~ ^v[0-9]+\.[0-9]+\.[0-9]+$ ]]; then
MODE="stable"
else
echo "::error::malformed v* tag rejected: ${REF_NAME_SAFE}"
echo "::error::stable tags must match ^v[0-9]+\\.[0-9]+\\.[0-9]+\$"
exit 1
fi
;;
*)
echo "::error::unexpected push ref ${REF_SAFE} reached publish workflow."
exit 1
;;
esac
;;
*)
echo "::error::unsupported event ${EVENT_NAME}."
exit 1
;;
esac
echo "mode=${MODE}" >> "$GITHUB_OUTPUT"
echo "Classified as mode=${MODE}"
# ── Phase 2 (RC only): dedup marker + release-PR skip ──────────────────────
rc-guard:
name: RC guard (marker + release-PR skip)
needs: route
if: needs.route.outputs.mode == 'rc'
runs-on: ubuntu-latest
timeout-minutes: 5
permissions:
contents: read
pull-requests: read
outputs:
should_run: ${{ steps.decide.outputs.should_run }}
head_sha: ${{ steps.decide.outputs.head_sha }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
fetch-depth: 0
fetch-tags: true
# rc-guard reads only — no git pushes from this job. Skip the
# default extraheader credential persistence (artipacked audit).
persist-credentials: false
- name: Decide
id: decide
shell: bash
env:
FORCE: ${{ inputs.force }}
BUMP_INPUT: ${{ inputs.bump }}
EVENT_NAME: ${{ github.event_name }}
GH_TOKEN: ${{ github.token }}
REPO: ${{ github.repository }}
run: |
set -euo pipefail
HEAD_SHA=$(git rev-parse HEAD)
echo "head_sha=$HEAD_SHA" >> "$GITHUB_OUTPUT"
if [ "$FORCE" = "true" ]; then
echo "Force flag set — running regardless of marker tag."
echo "should_run=true" >> "$GITHUB_OUTPUT"
exit 0
fi
# Explicit cycle reset on dispatch bypasses dedup.
if [ "$EVENT_NAME" = "workflow_dispatch" ] \
&& [ -n "${BUMP_INPUT:-}" ] \
&& [ "${BUMP_INPUT:-auto}" != "auto" ]; then
echo "Explicit bump=$BUMP_INPUT — bypassing marker dedup."
echo "should_run=true" >> "$GITHUB_OUTPUT"
exit 0
fi
# ── Skip when the merge commit corresponds to a release ───────────
# This skip is load-bearing: it prevents an RC build firing on the
# release-PR commit from racing the imminent stable-tag push on the
# same SHA. Two complementary checks:
# 1. HEAD subject matches `chore: release vX.Y.Z` (the canonical
# release-PR title). Anchored to require the bare title or the
# squash-merge `(#NNNN)` suffix exactly. Case-insensitive so
# `Chore: Release v1.2.3` (IDE auto-capitalization) still
# matches — prior commit-author conventions left the door open.
# 2. Squash-merged PR carries the `release` label.
# Either match suppresses the rc build — stable releases publish on
# the v-tag instead.
HEAD_SUBJECT="$(git log -1 --pretty=%s HEAD)"
# Sanitize GitHub-Actions annotation prefixes before logging — even
# though %s strips newlines, a crafted subject containing `::error::`
# could forge log annotations.
HEAD_SUBJECT_SAFE="${HEAD_SUBJECT//::/__}"
RELEASE_SUBJECT_RE='^chore:[[:space:]]*release[[:space:]]+v[0-9]+\.[0-9]+\.[0-9]+([[:space:]]+\(#[0-9]+\))?$'
shopt -s nocasematch
if [[ "$HEAD_SUBJECT" =~ $RELEASE_SUBJECT_RE ]]; then
shopt -u nocasematch
echo "HEAD commit subject matches a release commit — skipping rc."
echo " subject (sanitised): $HEAD_SUBJECT_SAFE"
echo "should_run=false" >> "$GITHUB_OUTPUT"
exit 0
fi
shopt -u nocasematch
# Squash-merge commits include `(#NNNN)` at the end of the subject.
if [[ "$HEAD_SUBJECT" =~ \(#([0-9]+)\)[[:space:]]*$ ]]; then
PR_NUM="${BASH_REMATCH[1]}"
echo "Detected squash-merge of PR #$PR_NUM — checking labels."
if LABELS_JSON="$(gh pr view "$PR_NUM" --repo "$REPO" --json labels 2>/dev/null)"; then
if printf '%s' "$LABELS_JSON" | jq -e '.labels[] | select(.name == "release")' >/dev/null; then
echo "PR #$PR_NUM has the 'release' label — skipping rc."
echo "should_run=false" >> "$GITHUB_OUTPUT"
exit 0
fi
echo "PR #$PR_NUM has no 'release' label — proceeding."
else
# Lookup failure is not fatal — fall through to dedup check.
echo "::warning::Could not read labels for PR #${PR_NUM} — falling through."
fi
fi
# Dedup: is there already an rc/<HEAD_SHA> marker pointing at HEAD?
MARKER="rc/${HEAD_SHA}"
if git rev-parse "refs/tags/$MARKER" >/dev/null 2>&1; then
echo "HEAD already has marker $MARKER — skipping."
echo "should_run=false" >> "$GITHUB_OUTPUT"
else
echo "No marker on HEAD — proceeding."
echo "should_run=true" >> "$GITHUB_OUTPUT"
fi
# ── Phase 3: reusable CI gate ──────────────────────────────────────────────
# Runs for both rc (when guard says go) and stable. No `secrets:` passed —
# ci.yml and its entire reusable-workflow chain (ci-quality, ci-tests,
# ci-e2e, ci-scope-parity, ci-report) reference zero `secrets.*` values;
# passing any would be unused surface. GITHUB_TOKEN is implicit.
ci:
needs: [route, rc-guard]
if: ${{ always() && (needs.route.outputs.mode == 'stable' || needs.rc-guard.outputs.should_run == 'true') }}
uses: ./.github/workflows/ci.yml
permissions:
contents: read
actions: read
# No pull-requests:write — `ci.yml`'s save-pr-meta job is gated on
# `github.event_name == 'pull_request'`, so it never runs during a
# tag-triggered publish. Least-privilege for release-critical paths.
# ── Phase 4: publish to npm + push refs (RC path) ──────────────────────────
# INVARIANT: `timeout-minutes` MUST stay below the App-token TTL (~60 min
# for actions/create-github-app-token installation tokens). The atomic
# tag-push step relies on the token minted at job start; if the job ever
# runs longer than the TTL, the push fails with an opaque 401. If you
# need to raise the timeout, re-mint the token immediately before the
# `Create and push rc tags` step instead.
publish:
needs: ci
name: Publish to npm
needs: [route, rc-guard, ci]
if: ${{ always() && needs.ci.result == 'success' && (needs.route.outputs.mode == 'stable' || needs.rc-guard.outputs.should_run == 'true') }}
runs-on: ubuntu-latest
timeout-minutes: 15
timeout-minutes: 20
permissions:
# contents: write — RC path needs it for `git push --atomic` (v-tag +
# marker). Stable path runs in the same job and inherits the grant; it
# never invokes `git push`, so the elevated scope is unused there.
# id-token: write — npm provenance attestation.
contents: write
id-token: write
outputs:
# Two distinct step IDs feed this output; exactly one fires per run.
vtag: ${{ steps.rc-tags.outputs.vtag || steps.stable-vtag.outputs.vtag }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
# ── Mint short-lived GitHub App token (RC only) ──────────────────────
# Industry direction (2025-2026): GitHub Apps with
# `actions/create-github-app-token` over long-lived PATs for
# workflow-touching tag pushes. Same fine-grained permission surface,
# ~1h expiry, not tied to a user seat, organizationally auditable.
# Replaces a prior fine-grained PAT.
#
# Required secrets (set in repo Settings → Secrets and variables → Actions):
# secrets.RELEASE_APP_ID — the App's numeric ID
# secrets.RELEASE_APP_PRIVATE_KEY — the App's PEM private key
# (The App ID is technically not sensitive — it's visible on the App's
# settings page — but storing it as a secret is harmless and avoids
# mixing storage classes for the same App.)
# The App must be installed on this repository with:
# - Contents: write (push the v-tag and rc marker)
# - Workflows: write (because the v-tag's tree may touch
# .github/workflows/**, which the default
# GITHUB_TOKEN cannot author)
# - Metadata: read (required for the `gh api /users/<slug>[bot]`
# bot-identity lookup in the tag-push step)
- name: Mint GitHub App token (RC)
if: needs.route.outputs.mode == 'rc'
id: app-token
uses: actions/create-github-app-token@bcd2ba49218906704ab6c1aa796996da409d3eb1 # v3.2.0
with:
# `client-id` is the renamed input that supersedes the deprecated
# `app-id` in v3.x. The action accepts the App's numeric ID or
# its Client ID under this name. We pass the numeric App ID,
# which the action resolves correctly.
client-id: ${{ secrets.RELEASE_APP_ID }}
private-key: ${{ secrets.RELEASE_APP_PRIVATE_KEY }}
# ── Separate checkout steps per mode ─────────────────────────────────
# Conditional `token:` expressions are footguns: empty string passed to
# actions/checkout fails opaquely, and `|| github.token` silently
# degrades a missing token to GITHUB_TOKEN, masking auth failures until
# the eventual `git push`. Two distinct steps make the auth contract
# explicit and fail loudly at checkout when the App token mint failed
# on the RC path.
- name: Checkout (RC)
if: needs.route.outputs.mode == 'rc'
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
fetch-depth: 0
fetch-tags: true
# Short-lived GitHub App installation token. Required because the
# v-tag push lands at a SHA whose tree may touch
# `.github/workflows/**`, which the default GITHUB_TOKEN cannot
# author.
token: ${{ steps.app-token.outputs.token }}
# Do not persist the token in .git/config (artipacked audit). The
# RC tag push uses an inline `http.extraheader` at push time only;
# the credential never lands on disk. See the
# `Create and push rc tags` step below.
persist-credentials: false
- name: Checkout (stable)
if: needs.route.outputs.mode == 'stable'
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
# No `token:` — actions/checkout uses GITHUB_TOKEN by default. Stable
# path performs no git pushes; the default scope is sufficient.
with:
# No git pushes from the stable path either. Skip credential
# persistence (artipacked audit).
persist-credentials: false
- name: Working-tree sanity
# Defense in depth (mirrors the vtag integrity gate, but on the input side):
# if a route-mode regression skipped both checkout `if:` gates, all
# downstream steps would run on a bare runner and produce confusing
# ENOENT errors. Fail loudly and early here instead.
shell: bash
run: |
if [ ! -f gitnexus/package.json ]; then
echo "::error::no working tree at gitnexus/package.json — route classification likely failed silently."
exit 1
fi
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
with:
node-version: 20
registry-url: https://registry.npmjs.org
# Hermetic install for the published artifact — no cache carry-over
# from non-tag contexts. setup-node v5+ caches by default when a
# packageManager field is present in package.json, so the explicit
# opt-out is required to clear the zizmor cache-poisoning audit.
# ~30s slower per release; runs rarely.
# Node 24 ships with npm >= 11.5.x, which is the minimum that
# supports npm Trusted Publishing OIDC. Node 22 ships with npm
# 10.9.x (no OIDC) and `npm install -g npm@latest` to self-upgrade
# is fragile — it can crash the in-flight reify with
# `MODULE_NOT_FOUND` on `promise-retry` etc. Bumping the Node
# version is the clean fix; the package's `engines` field is
# `>=22.0.0` so consumer-side compatibility is unaffected (this
# Node version is only used during publish, not by package users).
node-version: 24
# `registry-url:` is intentionally OMITTED. Under npm Trusted
# Publishing, OIDC only engages when no credential is configured.
# Setting `registry-url:` would make setup-node write
# `//registry.npmjs.org/:_authToken=${NODE_AUTH_TOKEN}` into the
# runner's .npmrc AND export NODE_AUTH_TOKEN from its `token:`
# input (default github.token). `npm publish` would then attempt
# GITHUB_TOKEN as the npm token, get rejected with 404, and OIDC
# would never be tried. See actions/setup-node#1440 and the GitHub
# Community discussion #176761 for the upstream bug and consensus
# workaround.
#
# Hermetic install for published artifacts — opt out of the v5+
# default packageManager-based caching (clears the zizmor
# cache-poisoning audit). ~30s slower per release; runs rarely.
package-manager-cache: false
- name: Build gitnexus-shared
run: npm install && npm run build
working-directory: gitnexus-shared
- run: npm ci
- name: Install gitnexus dependencies
run: npm ci
working-directory: gitnexus
- name: Verify version consistency
# ── Stable-only: verify the tag and package.json agree ───────────────
- name: Verify version consistency (stable)
if: needs.route.outputs.mode == 'stable'
shell: bash
working-directory: gitnexus
run: |
set -euo pipefail
TAG_VERSION="${GITHUB_REF#refs/tags/v}"
if ! [[ "$TAG_VERSION" =~ ^[0-9]+\.[0-9]+\.[0-9]+(-[a-zA-Z0-9.]+)?$ ]]; then
echo "::error::Tag does not follow semver: v$TAG_VERSION"
# Stable mode REJECTS prerelease suffixes — those are filtered at
# trigger by the negative-glob filter, but defend at the bash layer too.
if ! [[ "$TAG_VERSION" =~ ^[0-9]+\.[0-9]+\.[0-9]+$ ]]; then
echo "::error::Stable tag must be ^v[0-9]+.[0-9]+.[0-9]+$ — got v$TAG_VERSION"
exit 1
fi
PKG_VERSION=$(node -p "require('./package.json').version")
@@ -64,24 +424,376 @@ jobs:
exit 1
fi
echo "Version verified: $PKG_VERSION"
working-directory: gitnexus
- name: Build
# ── RC-only: compute the next rc version against the live registry ──
- name: Resolve rc version (rc)
id: rc-version
if: needs.route.outputs.mode == 'rc'
shell: bash
working-directory: gitnexus
env:
BUMP_INPUT: ${{ inputs.bump }}
EVENT_NAME: ${{ github.event_name }}
PKG_NAME: gitnexus
run: |
set -euo pipefail
# 1. Current published `latest` — the floor for any new rc base.
# Only E404 ("never published") falls back to package.json; any
# other error (network, auth, malformed response) fails fast
# (retry-loud policy: never silently substitute on transient errors).
NPM_STDERR_LATEST="$(mktemp)"
if CURRENT_LATEST="$(npm view "$PKG_NAME" version 2>"$NPM_STDERR_LATEST")"; then
:
else
if grep -qiE 'E404|not found' "$NPM_STDERR_LATEST"; then
CURRENT_LATEST="$(node -p "require('./package.json').version")"
echo "Package not on registry (E404) — seeding from package.json: $CURRENT_LATEST"
else
echo "::error::npm registry unreachable for 'view version':" >&2
cat "$NPM_STDERR_LATEST" >&2
rm -f "$NPM_STDERR_LATEST"
exit 1
fi
fi
rm -f "$NPM_STDERR_LATEST"
CURRENT_LATEST_CLEAN="${CURRENT_LATEST%%-*}"
# 2. Full version list — needed for the counter and active-cycle
# inference. Same E404-only fallback.
NPM_STDERR_VERSIONS="$(mktemp)"
if VERSIONS_JSON="$(npm view "$PKG_NAME" versions --json 2>"$NPM_STDERR_VERSIONS")"; then
:
else
if grep -qiE 'E404|not found' "$NPM_STDERR_VERSIONS"; then
VERSIONS_JSON='[]'
echo "No published versions for $PKG_NAME yet (E404)."
else
echo "::error::npm registry unreachable for 'view versions':" >&2
cat "$NPM_STDERR_VERSIONS" >&2
rm -f "$NPM_STDERR_VERSIONS"
exit 1
fi
fi
rm -f "$NPM_STDERR_VERSIONS"
# 3. Base selection.
# - workflow_dispatch + bump != auto → explicit cycle reset.
# - Otherwise (push, or dispatch with bump=auto) → continue the
# highest active rc base > latest if any; else patch from latest.
# Curated wrapper around `npx semver` — bare npx errors are noisy
# and don't distinguish registry-unreachable from invalid-bump-spec.
semver_bump() {
local kind="$1" current="$2" stderr_file out
stderr_file="$(mktemp)"
if out="$(npx --yes -p semver@7 semver -i "$kind" "$current" 2>"$stderr_file")"; then
rm -f "$stderr_file"
printf '%s' "$out"
return 0
fi
echo "::error::semver bump failed (kind=${kind}, current=${current}):" >&2
cat "$stderr_file" >&2
rm -f "$stderr_file"
return 1
}
if [ "$EVENT_NAME" = "workflow_dispatch" ] \
&& [ -n "${BUMP_INPUT:-}" ] \
&& [ "${BUMP_INPUT:-auto}" != "auto" ]; then
BASE="$(semver_bump "$BUMP_INPUT" "$CURRENT_LATEST_CLEAN")"
echo "Explicit bump=$BUMP_INPUT → BASE=$BASE"
else
cat > /tmp/active_base.mjs <<'NODESCRIPT'
const latest = process.env.LATEST;
let v;
try { v = JSON.parse(process.env.VERSIONS_JSON); } catch { v = []; }
if (!Array.isArray(v)) v = [v];
const parse = s => s.split(".").map(n => parseInt(n, 10));
const gt = (a, b) => {
const [A, B] = [parse(a), parse(b)];
for (let i = 0; i < 3; i++) if (A[i] !== B[i]) return A[i] > B[i];
return false;
};
const bases = new Set();
for (const s of v) {
const m = /^(\d+\.\d+\.\d+)-rc\.\d+$/.exec(s);
if (m && gt(m[1], latest)) bases.add(m[1]);
}
if (!bases.size) { process.stdout.write(""); process.exit(0); }
const sorted = [...bases].sort((a, b) => gt(a, b) ? 1 : -1);
process.stdout.write(sorted[sorted.length - 1]);
NODESCRIPT
ACTIVE_BASE="$(LATEST="$CURRENT_LATEST_CLEAN" VERSIONS_JSON="$VERSIONS_JSON" node /tmp/active_base.mjs)"
if [ -n "$ACTIVE_BASE" ]; then
BASE="$ACTIVE_BASE"
echo "Continuing active rc cycle → BASE=$BASE"
else
BASE="$(semver_bump patch "$CURRENT_LATEST_CLEAN")"
echo "No active rc cycle → patch bump from latest → BASE=$BASE"
fi
fi
# 4. Counter: 1 + max existing N for `${BASE}-rc.*`, else 1.
cat > /tmp/next_rc.mjs <<'NODESCRIPT'
const base = process.env.BASE;
const prefix = base + "-rc.";
let v;
try { v = JSON.parse(process.env.VERSIONS_JSON); } catch { v = []; }
if (!Array.isArray(v)) v = [v];
const ns = v
.filter(s => typeof s === "string" && s.startsWith(prefix))
.map(s => parseInt(s.slice(prefix.length), 10))
.filter(n => Number.isInteger(n) && n >= 0);
process.stdout.write(String(ns.length ? Math.max(...ns) + 1 : 1));
NODESCRIPT
NEXT_N="$(BASE="$BASE" VERSIONS_JSON="$VERSIONS_JSON" node /tmp/next_rc.mjs)"
RC_VERSION="${BASE}-rc.${NEXT_N}"
echo "Computed rc: $RC_VERSION"
# 5. Defensive: if the exact version already exists on the registry
# (race with another run), abort before re-publishing.
NPM_STDERR_EXISTS="$(mktemp)"
if npm view "$PKG_NAME@$RC_VERSION" version 2>"$NPM_STDERR_EXISTS" >/dev/null; then
rm -f "$NPM_STDERR_EXISTS"
echo "::error::Version $RC_VERSION already exists on npm — aborting."
exit 1
else
if grep -qiE 'E404|not found' "$NPM_STDERR_EXISTS"; then
rm -f "$NPM_STDERR_EXISTS"
# Version doesn't exist — safe to proceed.
else
echo "::error::npm registry unreachable for existence check:" >&2
cat "$NPM_STDERR_EXISTS" >&2
rm -f "$NPM_STDERR_EXISTS"
exit 1
fi
fi
{
echo "base=$BASE"
echo "rc_n=$NEXT_N"
echo "rc_version=$RC_VERSION"
} >> "$GITHUB_OUTPUT"
- name: Apply rc version in-CI
if: needs.route.outputs.mode == 'rc'
shell: bash
working-directory: gitnexus
run: |
set -euo pipefail
npm version "${{ steps.rc-version.outputs.rc_version }}" \
--no-git-tag-version --allow-same-version
- name: Build gitnexus
run: npm run build
working-directory: gitnexus
- name: Dry-run publish
run: npm publish --dry-run
working-directory: gitnexus
- name: Publish to npm
run: npm publish --provenance --access public
# Cheap verification that the tarball assembles before the real publish.
shell: bash
working-directory: gitnexus
env:
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
NPM_TAG: ${{ needs.route.outputs.mode == 'rc' && 'rc' || 'latest' }}
run: npm publish --dry-run --tag "$NPM_TAG"
- name: Extract release notes from CHANGELOG
# ── Acquire the "rc lock" BEFORE publishing (idempotency anchor) ─────
# We create two refs and push atomically:
# v<RC_VERSION> → annotated tag on a detached release commit whose
# tree contains the rewritten package.json, so the
# tag's source matches the npm tarball.
# rc/<HEAD_SHA> → lightweight tag on HEAD; the guard's dedup key.
# Push fails → nothing published. Push succeeds, npm fails → marker
# blocks retries until manual cleanup (see Rollback Runbook in plan).
- name: Create and push rc tags
id: rc-tags
if: needs.route.outputs.mode == 'rc'
shell: bash
working-directory: gitnexus
env:
RC_VERSION: ${{ steps.rc-version.outputs.rc_version }}
HEAD_SHA: ${{ needs.rc-guard.outputs.head_sha }}
# Short-lived GitHub App token. Auth is supplied inline at push
# time via `http.extraheader` (per GitHub's documented
# x-access-token Basic pattern). It is NOT persisted in
# .git/config (artipacked audit) — checkout above ran with
# `persist-credentials: false`.
PUSH_TOKEN: ${{ steps.app-token.outputs.token }}
# App's slug from create-github-app-token (e.g. `gitnexus-release-bot`).
# Used to attribute the release commit to the App identity rather
# than the generic github-actions[bot]. The bot's numeric user-id
# is resolved at runtime via the GitHub API (the action does not
# expose it directly as of v3.2.0).
APP_SLUG: ${{ steps.app-token.outputs.app-slug }}
GH_TOKEN: ${{ steps.app-token.outputs.token }}
run: |
set -euo pipefail
VTAG="v${RC_VERSION}"
MARKER="rc/${HEAD_SHA}"
# Resolve the App's bot user-id and construct the noreply email
# in the GitHub-canonical `<id>+<slug>[bot]@users.noreply.github.com`
# shape. `[bot]` is part of the actual login on GitHub.
#
# The lookup is wrapped in a bounded retry because the first RC
# after App installation may hit propagation delay (404), and
# transient api.github.com 5xx during heavy org activity is a real
# failure class. Without retry, every transient blip aborts the
# entire release after CI has already succeeded.
BOT_LOGIN="${APP_SLUG}[bot]"
BOT_USER_ID=""
api_stderr="$(mktemp)"
for attempt in 1 2 3; do
if BOT_USER_ID="$(gh api "/users/${BOT_LOGIN}" --jq .id 2>"$api_stderr")" \
&& [[ "${BOT_USER_ID}" =~ ^[0-9]+$ ]]; then
break
fi
BOT_USER_ID=""
if [ "$attempt" -lt 3 ]; then
echo "::warning::bot user-id lookup attempt ${attempt} failed; retrying in $((attempt * 5))s"
sleep $((attempt * 5))
fi
done
if ! [[ "${BOT_USER_ID}" =~ ^[0-9]+$ ]]; then
echo "::error::Could not resolve bot user-id for ${BOT_LOGIN} after 3 attempts."
echo "::error::gh api stderr:"
cat "$api_stderr" >&2 || true
echo "::error::Common causes: (a) newly-installed App — user record still propagating to /users/ (wait ~5min, redispatch with force=true); (b) App lacks Metadata: read permission; (c) transient api.github.com 5xx (redispatch)."
rm -f "$api_stderr"
exit 1
fi
rm -f "$api_stderr"
git config user.name "${BOT_LOGIN}"
git config user.email "${BOT_USER_ID}+${BOT_LOGIN}@users.noreply.github.com"
# Detached release commit with the version bump — main stays
# pristine, but the v-tag's tree matches the published package
# exactly (release-integrity).
git add package.json package-lock.json 2>/dev/null || git add package.json
git commit -m "release: ${VTAG}" --allow-empty
RELEASE_SHA="$(git rev-parse HEAD)"
echo "Detached release commit: $RELEASE_SHA"
git tag -a "$VTAG" "$RELEASE_SHA" -m "$VTAG"
git tag "$MARKER" "$HEAD_SHA"
# Inline auth header. The base64-encoded form is masked as well
# as the raw token, because GitHub's secret-masker only masks the
# raw value — any subsequent `set -x` / GIT_TRACE line would
# otherwise expose the encoded credential.
#
# `set +x` wraps the compute+mask pair so that if an operator
# enables ACTIONS_STEP_DEBUG=true for triage (which turns on
# `set -x` globally), the assignment is NOT traced for the one
# line between compute and mask-registration. Without this wrap,
# debug mode would log `+ auth_header='Authorization: Basic <encoded>'`
# exposing a still-valid (~1h) App token.
{ set +x; } 2>/dev/null
auth_header="Authorization: Basic $(printf 'x-access-token:%s' "${PUSH_TOKEN}" | base64 -w0)"
echo "::add-mask::${auth_header}"
# Re-enable tracing only when explicitly requested via step-debug.
if [ "${ACTIONS_STEP_DEBUG:-false}" = "true" ]; then set -x; fi
# Atomic push of both refs. If either would clobber an existing
# remote ref, the push fails and we stop before npm publish.
git -c http.extraheader="${auth_header}" \
push --atomic origin "refs/tags/$VTAG" "refs/tags/$MARKER"
{
echo "vtag=$VTAG"
echo "marker=$MARKER"
echo "release_sha=$RELEASE_SHA"
} >> "$GITHUB_OUTPUT"
- name: Set vtag (stable)
id: stable-vtag
if: needs.route.outputs.mode == 'stable'
shell: bash
# github.ref_name flows in via env to avoid templating into the
# shell source (template-injection audit). Even though refs are
# constrained by git naming rules, the env-passthrough pattern
# makes injection structurally impossible.
env:
REF_NAME: ${{ github.ref_name }}
run: |
echo "vtag=${REF_NAME}" >> "$GITHUB_OUTPUT"
# ── vtag integrity gate ──────────────────────────────────────────────
# Fail closed before any artifact-producing step (npm publish, Release,
# Docker) runs against an empty or mode-mismatched vtag. Prevents the
# silent "Release named main" / "Docker tagged from ref fallback"
# failure modes that the previous draft was vulnerable to.
- name: vtag integrity gate
id: vtag-gate
shell: bash
env:
MODE: ${{ needs.route.outputs.mode }}
VTAG: ${{ steps.rc-tags.outputs.vtag || steps.stable-vtag.outputs.vtag }}
run: |
set -euo pipefail
if [ -z "$VTAG" ]; then
echo "::error::vtag is empty — refusing to create GitHub Release or trigger Docker."
exit 1
fi
case "$MODE" in
rc)
if ! [[ "$VTAG" =~ ^v[0-9]+\.[0-9]+\.[0-9]+-rc\.[0-9]+$ ]]; then
echo "::error::vtag '${VTAG}' does not match rc shape ^v[0-9]+.[0-9]+.[0-9]+-rc.[0-9]+$"
exit 1
fi
;;
stable)
if ! [[ "$VTAG" =~ ^v[0-9]+\.[0-9]+\.[0-9]+$ ]]; then
echo "::error::vtag '${VTAG}' does not match stable shape ^v[0-9]+.[0-9]+.[0-9]+$"
exit 1
fi
;;
*)
echo "::error::unknown mode '${MODE}' at vtag integrity gate."
exit 1
;;
esac
echo "vtag verified: ${VTAG} (mode=${MODE})"
echo "vtag=${VTAG}" >> "$GITHUB_OUTPUT"
# npm Trusted Publishing (GA'd 2025-07-31). OIDC authentication only
# engages when no npm credential is configured anywhere — the absence
# is the signal. Two upstream behaviors had to be neutralized for
# this to work:
#
# 1. setup-node's `registry-url:` is omitted (see the setup-node
# step above). With it, setup-node writes
# `//registry.npmjs.org/:_authToken=${NODE_AUTH_TOKEN}` into
# .npmrc and exports NODE_AUTH_TOKEN from `token:` (defaulting
# to github.token). npm publish then sends GITHUB_TOKEN as the
# bearer credential and the registry returns 404. OIDC is never
# tried because npm thinks it already has a credential.
# 2. The runner's bundled npm (10.9.x on Node 22) has no OIDC
# support; the upgrade step above pins it to >= 11.5.1.
#
# Provenance is auto-attached by the registry on trusted-publisher
# publishes — no --provenance flag needed.
#
# Prerequisite: register the package as a trusted publisher at
# https://www.npmjs.com/package/gitnexus/access (Publishing access →
# Trusted Publishers → GitHub Actions):
# Owner: abhigyanpatwari
# Repository: GitNexus
# Workflow: publish.yml
# Environment: (none)
- name: Publish to npm
shell: bash
working-directory: gitnexus
env:
NPM_TAG: ${{ needs.route.outputs.mode == 'rc' && 'rc' || 'latest' }}
run: npm publish --access public --tag "$NPM_TAG"
# ── Stable-only: pull CHANGELOG body if present ──────────────────────
- name: Extract release notes from CHANGELOG (stable)
id: changelog
if: needs.route.outputs.mode == 'stable'
shell: bash
run: |
VERSION="${GITHUB_REF#refs/tags/v}"
@@ -97,5 +809,90 @@ jobs:
- name: Create GitHub Release
uses: softprops/action-gh-release@b4309332981a82ec1c5618f44dd2e27cc8bfbfda # v2
with:
body_path: ${{ steps.changelog.outputs.fallback == 'false' && '/tmp/release-notes.md' || '' }}
generate_release_notes: ${{ steps.changelog.outputs.fallback == 'true' }}
tag_name: ${{ steps.vtag-gate.outputs.vtag }}
name: >-
${{ needs.route.outputs.mode == 'rc'
&& format('Release Candidate {0}', steps.vtag-gate.outputs.vtag)
|| steps.vtag-gate.outputs.vtag }}
prerelease: ${{ needs.route.outputs.mode == 'rc' }}
make_latest: ${{ needs.route.outputs.mode == 'stable' && 'true' || 'false' }}
# Stable: prefer CHANGELOG body, fall back to auto-generated.
# RC: always auto-generated + the prerelease body block below.
body_path: >-
${{ needs.route.outputs.mode == 'stable' && steps.changelog.outputs.fallback == 'false'
&& '/tmp/release-notes.md' || '' }}
generate_release_notes: >-
${{ needs.route.outputs.mode == 'rc'
|| steps.changelog.outputs.fallback == 'true' }}
body: >-
${{ needs.route.outputs.mode == 'rc' && format(
'Automated release candidate build from `main`.{0}{0}**npm:** `npm install gitnexus@rc`{0}**Version:** `{1}`{0}**Target base:** `{2}` (rc #{3}){0}**Source commit (main):** {4}{0}**Release commit (versioned tree):** {5}{0}{0}Release candidates are pre-stable builds intended for early testing. Stable releases remain on the `latest` dist-tag.',
'\n',
steps.rc-version.outputs.rc_version,
steps.rc-version.outputs.base,
steps.rc-version.outputs.rc_n,
needs.rc-guard.outputs.head_sha,
steps.rc-tags.outputs.release_sha
) || '' }}
# ── RC partial-failure cleanup ───────────────────────────────────────
# If anything after the atomic tag-push step failed (npm publish
# blew up, GitHub Release call timed out, etc.), the v-tag and
# rc/<SHA> marker are already on origin. External consumers
# (Renovate, Dependabot, Releases RSS) can ingest a phantom tag for
# a version that was never published to npm. This step deletes them
# automatically so the operator's recovery is just "redispatch with
# force=true on the next commit", not a manual ref cleanup.
#
# Scoped strictly to RC + real (non-dry-run) + the rc-tags step
# actually produced a vtag (otherwise nothing to clean up). The
# App token is still valid (~1h TTL, job timeout 20min).
- name: Cleanup pushed tags on partial failure
if: ${{ failure() && needs.route.outputs.mode == 'rc' && steps.rc-tags.outputs.vtag != '' }}
shell: bash
working-directory: gitnexus
env:
VTAG: ${{ steps.rc-tags.outputs.vtag }}
MARKER: ${{ steps.rc-tags.outputs.marker }}
PUSH_TOKEN: ${{ steps.app-token.outputs.token }}
run: |
set -uo pipefail
echo "::warning::Publish step failed after tag push. Cleaning up remote refs to prevent phantom-version ingestion by downstream consumers."
{ set +x; } 2>/dev/null
auth_header="Authorization: Basic $(printf 'x-access-token:%s' "${PUSH_TOKEN}" | base64 -w0)"
echo "::add-mask::${auth_header}"
if [ "${ACTIONS_STEP_DEBUG:-false}" = "true" ]; then set -x; fi
# Delete v-tag and marker. Each delete is best-effort — if one
# is already absent (atomic push partially rejected, or earlier
# cleanup ran), the other still gets attempted.
for ref in "refs/tags/${VTAG}" "refs/tags/${MARKER}"; do
if git -c http.extraheader="${auth_header}" push origin --delete "${ref}" 2>&1; then
echo "deleted origin ${ref}"
else
echo "::warning::could not delete origin ${ref} — may already be absent or protected. Manual cleanup may be required."
fi
done
echo "::notice::Cleanup complete. To retry the release, redispatch the workflow with force=true on the same SHA, or push a new commit to main."
# ── Phase 5 (RC only): Docker images ───────────────────────────────────────
# R6: Docker remains RC-only. Stable Docker builds are explicitly deferred.
# Secrets are passed explicitly (not via `secrets: inherit`) so the
# callee's secret surface is auditable from the caller's source.
docker:
name: Build & Push RC Docker images
needs: [route, publish]
if: ${{ needs.route.outputs.mode == 'rc' && needs.publish.outputs.vtag != '' }}
uses: ./.github/workflows/docker.yml
secrets:
DOCKERHUB_USERNAME: ${{ secrets.DOCKERHUB_USERNAME }}
DOCKERHUB_TOKEN: ${{ secrets.DOCKERHUB_TOKEN }}
permissions:
contents: read
packages: write
id-token: write
attestations: write
with:
tag: ${{ needs.publish.outputs.vtag }}
-409
View File
@@ -1,409 +0,0 @@
name: Release Candidate
on:
# Publish a release-candidate build whenever a merge/commit lands on main.
# Docs/README-only changes are filtered out so prose updates don't
# cut a release.
push:
branches: [main]
paths-ignore:
- '**.md'
- 'docs/**'
- 'LICENSE'
workflow_dispatch:
inputs:
bump:
description: >-
Cycle policy. 'auto' (default) continues the active rc cycle on
this branch if there is one, otherwise bumps patch from latest.
Choose 'patch' / 'minor' / 'major' to explicitly start or reset
an rc cycle.
required: false
default: 'auto'
type: choice
options:
- auto
- patch
- minor
- major
force:
description: 'Publish even when HEAD already has an rc marker'
required: false
default: 'false'
type: choice
options:
- 'false'
- 'true'
# No workflow-level permissions — scoped per job below.
permissions: {}
# Concurrency convention: see CONTRIBUTING.md → "GitHub Actions — Concurrency Convention".
# Serialize all runs on the same ref (push + workflow_dispatch) to prevent two publishes
# racing on the rc counter. cancel-in-progress: false — the earlier merge publishes first.
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: false
jobs:
# ── Skip when HEAD already has an rc marker (retry / duplicate dispatch) ──
# The marker is a lightweight tag `rc/<HEAD_SHA>` pushed *before* `npm
# publish`, so a failed publish leaves the marker in place and the guard
# refuses to re-publish. Recovery path after a partial failure:
# git push --delete origin rc/<HEAD_SHA> v<RC_VERSION>
# then redispatch with force=true.
guard:
name: Check if release candidate should run
runs-on: ubuntu-latest
timeout-minutes: 5
permissions:
contents: read
outputs:
should_run: ${{ steps.decide.outputs.should_run }}
head_sha: ${{ steps.decide.outputs.head_sha }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
fetch-depth: 0
fetch-tags: true
- name: Decide
id: decide
shell: bash
env:
FORCE: ${{ inputs.force }}
BUMP_INPUT: ${{ inputs.bump }}
EVENT_NAME: ${{ github.event_name }}
run: |
set -euo pipefail
HEAD_SHA=$(git rev-parse HEAD)
echo "head_sha=$HEAD_SHA" >> "$GITHUB_OUTPUT"
if [ "$FORCE" = "true" ]; then
echo "Force flag set — running regardless of marker tag."
echo "should_run=true" >> "$GITHUB_OUTPUT"
exit 0
fi
# An explicit cycle reset on dispatch (bump != auto) also bypasses
# the dedup guard — the maintainer is deliberately asking for a
# new rc from the same commit.
if [ "$EVENT_NAME" = "workflow_dispatch" ] \
&& [ -n "${BUMP_INPUT:-}" ] \
&& [ "${BUMP_INPUT:-auto}" != "auto" ]; then
echo "Explicit bump=$BUMP_INPUT — bypassing marker dedup."
echo "should_run=true" >> "$GITHUB_OUTPUT"
exit 0
fi
# Dedup: is there already an rc/<HEAD_SHA> marker pointing at HEAD?
MARKER="rc/${HEAD_SHA}"
if git rev-parse "refs/tags/$MARKER" >/dev/null 2>&1; then
echo "HEAD already has marker $MARKER — skipping."
echo "should_run=false" >> "$GITHUB_OUTPUT"
else
echo "No marker on HEAD — proceeding."
echo "should_run=true" >> "$GITHUB_OUTPUT"
fi
# ── Reuse the stable CI workflow ─────────────────────────────────────
ci:
needs: guard
if: needs.guard.outputs.should_run == 'true'
uses: ./.github/workflows/ci.yml
permissions:
contents: read
secrets: inherit
# ── Publish the rc build to npm + create GitHub prerelease ───────────
publish:
name: Publish release candidate to npm
needs: [guard, ci]
if: needs.guard.outputs.should_run == 'true'
runs-on: ubuntu-latest
timeout-minutes: 20
permissions:
# The default GITHUB_TOKEN cannot be granted `workflows: write`, so
# tag pushes that reach a commit which modified `.github/workflows/**`
# are rejected with: "refusing to allow a GitHub App to create or
# update workflow ... without `workflows` permission". We pass a
# fine-grained PAT (RELEASE_PUSH_TOKEN, scoped to this repo with
# Contents: write + Workflows: write) to `actions/checkout` so that
# the subsequent `git push --atomic` of the v-tag and rc marker
# carries the PAT's identity. Job-level GITHUB_TOKEN keeps its
# scoped permissions for everything else (npm provenance, etc.).
contents: write # push rc tag + marker (via PAT)
id-token: write # npm provenance
outputs:
vtag: ${{ steps.reltag.outputs.vtag }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
fetch-depth: 0
fetch-tags: true
# Use the PAT so `origin` is preauthed for `git push`. Without
# this the default GITHUB_TOKEN is wired into the remote, and a
# workflows-touching tag push is rejected — see the permissions
# block above.
token: ${{ secrets.RELEASE_PUSH_TOKEN }}
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
with:
node-version: 20
registry-url: https://registry.npmjs.org
# Hermetic install — release-candidate produces shipped artifacts.
# setup-node v5+ caches by default when a packageManager field is
# present in package.json; explicit opt-out is required to clear
# the zizmor cache-poisoning audit. See cache-poisoning audit.
package-manager-cache: false
- name: Build gitnexus-shared
run: npm install && npm run build
working-directory: gitnexus-shared
- name: Install gitnexus dependencies
run: npm ci
working-directory: gitnexus
- name: Resolve rc version
id: version
shell: bash
working-directory: gitnexus
env:
BUMP_INPUT: ${{ inputs.bump }}
EVENT_NAME: ${{ github.event_name }}
PKG_NAME: gitnexus
run: |
set -euo pipefail
# 1. Current published `latest` — the floor for any new rc base.
# Only E404 ("never published") falls back to package.json; any
# other error (network, auth, malformed response) fails fast.
NPM_STDERR_LATEST="$(mktemp)"
if CURRENT_LATEST="$(npm view "$PKG_NAME" version 2>"$NPM_STDERR_LATEST")"; then
:
else
if grep -q 'E404' "$NPM_STDERR_LATEST"; then
CURRENT_LATEST="$(node -p "require('./package.json').version")"
echo "Package not on registry (E404) — seeding from package.json: $CURRENT_LATEST"
else
echo "::error::npm registry unreachable for 'view version':" >&2
cat "$NPM_STDERR_LATEST" >&2
rm -f "$NPM_STDERR_LATEST"
exit 1
fi
fi
rm -f "$NPM_STDERR_LATEST"
CURRENT_LATEST_CLEAN="${CURRENT_LATEST%%-*}"
# 2. Full version list — needed for the counter and for active-cycle
# inference. Same E404-only fallback.
NPM_STDERR_VERSIONS="$(mktemp)"
if VERSIONS_JSON="$(npm view "$PKG_NAME" versions --json 2>"$NPM_STDERR_VERSIONS")"; then
:
else
if grep -q 'E404' "$NPM_STDERR_VERSIONS"; then
VERSIONS_JSON='[]'
echo "No published versions for $PKG_NAME yet (E404)."
else
echo "::error::npm registry unreachable for 'view versions':" >&2
cat "$NPM_STDERR_VERSIONS" >&2
rm -f "$NPM_STDERR_VERSIONS"
exit 1
fi
fi
rm -f "$NPM_STDERR_VERSIONS"
# 3. Base selection.
# - workflow_dispatch + bump ∈ {patch,minor,major} → explicit cycle
# reset from latest.
# - Everything else (push, or dispatch with bump=auto) → continue
# the highest active rc base > latest if one exists; else
# default to patch from latest.
if [ "$EVENT_NAME" = "workflow_dispatch" ] \
&& [ -n "${BUMP_INPUT:-}" ] \
&& [ "${BUMP_INPUT:-auto}" != "auto" ]; then
BASE="$(npx --yes -p semver@7 semver -i "$BUMP_INPUT" "$CURRENT_LATEST_CLEAN")"
echo "Explicit bump=$BUMP_INPUT → BASE=$BASE"
else
cat > /tmp/active_base.mjs <<'NODESCRIPT'
const latest = process.env.LATEST;
let v;
try { v = JSON.parse(process.env.VERSIONS_JSON); } catch { v = []; }
if (!Array.isArray(v)) v = [v];
const parse = s => s.split(".").map(n => parseInt(n, 10));
const gt = (a, b) => {
const [A, B] = [parse(a), parse(b)];
for (let i = 0; i < 3; i++) if (A[i] !== B[i]) return A[i] > B[i];
return false;
};
const bases = new Set();
for (const s of v) {
const m = /^(\d+\.\d+\.\d+)-rc\.\d+$/.exec(s);
if (m && gt(m[1], latest)) bases.add(m[1]);
}
if (!bases.size) { process.stdout.write(""); process.exit(0); }
const sorted = [...bases].sort((a, b) => gt(a, b) ? 1 : -1);
process.stdout.write(sorted[sorted.length - 1]);
NODESCRIPT
ACTIVE_BASE="$(LATEST="$CURRENT_LATEST_CLEAN" VERSIONS_JSON="$VERSIONS_JSON" node /tmp/active_base.mjs)"
if [ -n "$ACTIVE_BASE" ]; then
BASE="$ACTIVE_BASE"
echo "Continuing active rc cycle → BASE=$BASE"
else
BASE="$(npx --yes -p semver@7 semver -i patch "$CURRENT_LATEST_CLEAN")"
echo "No active rc cycle → patch bump from latest → BASE=$BASE"
fi
fi
# 4. Counter: 1 + max existing N for `${BASE}-rc.*`, else 1.
cat > /tmp/next_rc.mjs <<'NODESCRIPT'
const base = process.env.BASE;
const prefix = base + "-rc.";
let v;
try { v = JSON.parse(process.env.VERSIONS_JSON); } catch { v = []; }
if (!Array.isArray(v)) v = [v];
const ns = v
.filter(s => typeof s === "string" && s.startsWith(prefix))
.map(s => parseInt(s.slice(prefix.length), 10))
.filter(n => Number.isInteger(n) && n >= 0);
process.stdout.write(String(ns.length ? Math.max(...ns) + 1 : 1));
NODESCRIPT
NEXT_N="$(BASE="$BASE" VERSIONS_JSON="$VERSIONS_JSON" node /tmp/next_rc.mjs)"
RC_VERSION="${BASE}-rc.${NEXT_N}"
echo "Computed rc: $RC_VERSION"
# 5. Defensive: if the exact version already exists on the registry
# (e.g., race with another run), abort before re-publishing.
# Same E404-only pattern used above — a transient network
# failure must fail loudly, not pretend the version is missing.
NPM_STDERR_EXISTS="$(mktemp)"
if npm view "$PKG_NAME@$RC_VERSION" version 2>"$NPM_STDERR_EXISTS" >/dev/null; then
rm -f "$NPM_STDERR_EXISTS"
echo "::error::Version $RC_VERSION already exists on npm — aborting."
exit 1
else
if grep -qiE 'E404|not found' "$NPM_STDERR_EXISTS"; then
rm -f "$NPM_STDERR_EXISTS"
# Version doesn't exist — safe to proceed.
else
echo "::error::npm registry unreachable for existence check:" >&2
cat "$NPM_STDERR_EXISTS" >&2
rm -f "$NPM_STDERR_EXISTS"
exit 1
fi
fi
echo "base=$BASE" >> "$GITHUB_OUTPUT"
echo "rc_n=$NEXT_N" >> "$GITHUB_OUTPUT"
echo "rc_version=$RC_VERSION" >> "$GITHUB_OUTPUT"
- name: Apply rc version in-CI
shell: bash
working-directory: gitnexus
run: |
set -euo pipefail
npm version "${{ steps.version.outputs.rc_version }}" \
--no-git-tag-version --allow-same-version
- name: Build gitnexus
run: npm run build
working-directory: gitnexus
- name: Dry-run publish
run: npm publish --dry-run --tag rc
working-directory: gitnexus
# ── Acquire the "rc lock" BEFORE publishing (fixes idempotency) ─────
# We create two tags and push them atomically:
# v<RC_VERSION> → annotated tag on a detached release commit
# whose tree contains the rewritten package.json
# (so the tag's source matches the npm tarball)
# rc/<HEAD_SHA> → lightweight tag on HEAD; the guard's dedup key
# If this push fails, nothing is published — safe.
# If this push succeeds but npm publish fails, the marker stays on
# the remote and blocks retries until an operator manually cleans up.
- name: Create and push rc tags
id: reltag
shell: bash
working-directory: gitnexus
env:
RC_VERSION: ${{ steps.version.outputs.rc_version }}
HEAD_SHA: ${{ needs.guard.outputs.head_sha }}
run: |
set -euo pipefail
VTAG="v${RC_VERSION}"
MARKER="rc/${HEAD_SHA}"
git config user.name 'github-actions[bot]'
git config user.email '41898282+github-actions[bot]@users.noreply.github.com'
# Detached release commit with the version bump — keeps `main`
# pristine but gives the v-tag a tree that matches the published
# package contents exactly (fixes release-integrity gap).
git add package.json package-lock.json 2>/dev/null || git add package.json
git commit -m "release: ${VTAG}" --allow-empty
RELEASE_SHA="$(git rev-parse HEAD)"
echo "Detached release commit: $RELEASE_SHA"
# Annotated release tag on the release commit.
git tag -a "$VTAG" "$RELEASE_SHA" -m "$VTAG"
# Lightweight marker on the user-visible HEAD for the guard.
git tag "$MARKER" "$HEAD_SHA"
# Atomic push of both refs. If either would clobber an existing
# remote ref, the push fails and we stop before npm publish.
git push --atomic origin "refs/tags/$VTAG" "refs/tags/$MARKER"
echo "vtag=$VTAG" >> "$GITHUB_OUTPUT"
echo "marker=$MARKER" >> "$GITHUB_OUTPUT"
echo "release_sha=$RELEASE_SHA" >> "$GITHUB_OUTPUT"
- name: Publish to npm (rc dist-tag)
run: npm publish --provenance --access public --tag rc
working-directory: gitnexus
env:
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
- name: Create GitHub prerelease
uses: softprops/action-gh-release@b4309332981a82ec1c5618f44dd2e27cc8bfbfda # v2
with:
tag_name: ${{ steps.reltag.outputs.vtag }}
name: Release Candidate ${{ steps.reltag.outputs.vtag }}
prerelease: true
make_latest: 'false'
generate_release_notes: true
body: |
Automated release candidate build from `main`.
**npm:** `npm install gitnexus@rc`
**Version:** `${{ steps.version.outputs.rc_version }}`
**Target base:** `${{ steps.version.outputs.base }}` (rc #${{ steps.version.outputs.rc_n }})
**Source commit (main):** ${{ needs.guard.outputs.head_sha }}
**Release commit (versioned tree):** ${{ steps.reltag.outputs.release_sha }}
Release candidates are pre-stable builds intended for early testing.
Stable releases remain on the `latest` dist-tag.
# ── Build & push RC Docker images ────────────────────────────────────
# Calls docker.yml as a reusable workflow so that the build, signing, and
# attestation logic stays in one place. The publish job exposes `vtag`
# (e.g. `v1.2.3-rc.1`) as an output so we can pass it as the tag input.
# RC images are signed with Cosign keyless signing; the OIDC identity
# will be `docker.yml@refs/heads/main` (the caller's ref) rather than a
# tag ref — see README.md § Docker for the correct verify command for RCs.
docker:
name: Build & Push RC Docker images
needs: [guard, publish]
if: needs.guard.outputs.should_run == 'true' && needs.publish.outputs.vtag != ''
uses: ./.github/workflows/docker.yml
# Reusable workflows do not receive caller secrets unless inherited; without
# this, DOCKERHUB_* / GITHUB_TOKEN are empty in docker.yml → "Username and
# password required" on Docker Hub login (see same pattern on `ci:` above).
secrets: inherit
permissions:
contents: read
packages: write
id-token: write
attestations: write
with:
tag: ${{ needs.publish.outputs.vtag }}
+1 -1
View File
@@ -53,6 +53,6 @@ jobs:
retention-days: 5
- name: Upload to Security tab
uses: github/codeql-action/upload-sarif@e46ed2cbd01164d986452f91f178727624ae40d7 # v4.35.3
uses: github/codeql-action/upload-sarif@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
with:
sarif_file: results.sarif
+14 -5
View File
@@ -1,19 +1,28 @@
name: Trivy Image Scan
# Builds Dockerfile.cli and Dockerfile.web, then scans the resulting images
# for OS-package and language-package CVEs at HIGH/CRITICAL severity.
# for OS-package and language-package CVEs at MEDIUM+ severity.
# Findings upload to the Security tab; record-only (does not block merges).
#
# NOT triggered on PRs — image builds are slow and base-image CVE churn
# shouldn't gate feature delivery.
# Trigger on Dockerfile changes in PRs so base-image/npm-layer remediation can
# be verified before merge without running image scans on every PR.
on:
pull_request:
paths:
- 'Dockerfile.cli'
- 'Dockerfile.web'
- 'gitnexus/Dockerfile.test'
- '.github/workflows/trivy.yml'
push:
branches: [main]
schedule:
- cron: '0 8 * * 1'
workflow_dispatch:
permissions:
contents: read
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: false
@@ -61,13 +70,13 @@ jobs:
image-ref: scan-target:${{ matrix.image.name }}
format: sarif
output: trivy-${{ matrix.image.name }}.sarif
severity: HIGH,CRITICAL
severity: MEDIUM,HIGH,CRITICAL
# Hides CVEs with no available fix in the base image.
ignore-unfixed: true
exit-code: '0'
- name: Upload to Security tab
uses: github/codeql-action/upload-sarif@e46ed2cbd01164d986452f91f178727624ae40d7 # v4.35.3
uses: github/codeql-action/upload-sarif@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
with:
sarif_file: trivy-${{ matrix.image.name }}.sarif
category: trivy-${{ matrix.image.name }}
+33 -6
View File
@@ -1,10 +1,13 @@
name: Workflow Lint (zizmor)
name: Workflow Lint
# Lints .github/workflows/** for known GitHub Actions security misconfigurations:
# unpinned Actions, dangerous ${{ ... }} interpolation in run: blocks,
# missing per-job permissions:, etc.
# Lints .github/workflows/** for both:
# - actionlint: YAML syntax, expression typing, shellcheck inside `run:`
# blocks, unknown contexts, deprecated runner labels.
# - zizmor: security misconfigurations — unpinned actions, dangerous
# `${{ }}` interpolation, missing per-job permissions, etc.
#
# Scoped to PRs that touch .github/** only — keeps off the typical PR critical path.
# Scoped to PRs that touch .github/** only — keeps off the typical PR
# critical path.
on:
pull_request:
@@ -12,11 +15,35 @@ on:
paths:
- '.github/**'
permissions:
contents: read
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
jobs:
actionlint:
name: actionlint
runs-on: ubuntu-latest
timeout-minutes: 5
permissions:
contents: read
steps:
- name: Checkout
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
# Pinned to v2.1.2. Verify SHA via:
# gh api repos/raven-actions/actionlint/git/refs/tags/v2.1.2
# The action wraps the upstream `rhysd/actionlint` binary and emits
# GitHub-annotation-formatted findings on PRs.
- name: Run actionlint
uses: raven-actions/actionlint@205b530c5d9fa8f44ae9ed59f341a0db994aa6f8 # v2.1.2
with:
fail-on-error: true
zizmor:
runs-on: ubuntu-latest
timeout-minutes: 10
@@ -49,7 +76,7 @@ jobs:
continue-on-error: true
- name: Upload SARIF
uses: github/codeql-action/upload-sarif@e46ed2cbd01164d986452f91f178727624ae40d7 # v4.35.3
uses: github/codeql-action/upload-sarif@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
with:
sarif_file: zizmor.sarif
category: zizmor
+14 -4
View File
@@ -14,6 +14,15 @@ rules:
# no checkout of fork code occurs. Header comment in the file documents.
- ci-report.yml
# workflow_run is the trusted half of the autofix pipeline. The
# untrusted half (pr-autofix.yml) runs fork code with permissions:{}
# and produces only a diff artifact (data, not executable code). The
# publish job consumes the artifact, allowlist-validates every field
# of metadata.json before exporting to $GITHUB_OUTPUT, never checks
# out fork code, and never executes anything fork-controlled. Header
# comment in the file documents the split.
- pr-autofix-publish.yml
# pull_request_target needed by claude-code-action to access secrets
# and post review comments on fork PRs. Mitigated by: PR checkouts pin
# the fork's HEAD SHA (not the branch ref) to prevent TOCTOU races,
@@ -28,7 +37,8 @@ rules:
- pr-labeler.yml
# Note: cache-poisoning is NOT exempted. The two prior findings in
# publish.yml and release-candidate.yml were fixed structurally by
# dropping `cache: npm` from those workflows (matches the pattern used
# by PyO3/maturin for the same audit). See the commit that added this
# file for the rationale.
# publish.yml and the former release-candidate.yml were fixed structurally
# by dropping `cache: npm` from those workflows (matches the pattern used
# by PyO3/maturin for the same audit). After the publish-workflow
# unification (issue #1609), only publish.yml remains; the same
# cache-poisoning hardening applies there.
+41 -89
View File
@@ -62,112 +62,64 @@ Commands and gotchas live under **Repo reference** below and in **[CONTRIBUTING.
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
Indexed as **GitNexus** (4325 symbols, 10556 relationships, 300 execution flows). Use MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitNexus** (26675 symbols, 35395 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any tool warns the index is stale, run `npx gitnexus analyze` first.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Do
- **MUST run impact analysis before editing any symbol.** `gitnexus_impact({target: "symbolName", direction: "upstream"})` — report blast radius to the user.
- **MUST run `gitnexus_detect_changes()` before committing** — verify only expected symbols and flows are affected.
- **MUST warn the user** if impact returns HIGH or CRITICAL risk.
- Explore unfamiliar code with `gitnexus_query({query: "concept"})` (process-grouped, ranked) instead of grepping.
- Full context on a symbol: `gitnexus_context({name: "symbolName"})`.
## When Debugging
1. `gitnexus_query({query: "<error or symptom>"})` — find related execution flows
2. `gitnexus_context({name: "<suspect function>"})` — callers, callees, process participation
3. `READ gitnexus://repo/GitNexus/process/{processName}` — trace flow step by step
4. Regressions: `gitnexus_detect_changes({scope: "compare", base_ref: "main"})`
## When Refactoring
- **Rename:** `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` first. Graph edits are safe; text_search edits need manual review.
- **Extract/Split:** `gitnexus_context` (incoming/outgoing refs) then `gitnexus_impact` (upstream callers) before moving code.
- **After any refactor:** `gitnexus_detect_changes({scope: "all"})` to verify scope.
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## Never Do
- Edit a symbol without running `gitnexus_impact` first.
- Ignore HIGH/CRITICAL risk warnings.
- Rename with find-and-replace — use `gitnexus_rename`.
- Commit without `gitnexus_detect_changes()`.
- Add language-specific behavior to shared ingestion code (`gitnexus/src/core/ingestion/`) — use a `LanguageProvider` hook. Seeing `provider.mroStrategy === 'xxx'` or an import from `languages/xxx.ts` in shared code means stop and add a hook.
## Tools Quick Reference
| Tool | When to use | Example |
|------|-------------|---------|
| `list_repos` | Discover indexed repos | `gitnexus_list_repos({})` |
| `query` | Find code by concept | `gitnexus_query({query: "auth validation"})` |
| `context` | 360-degree view of one symbol | `gitnexus_context({name: "validateUser"})` |
| `impact` | Blast radius before editing | `gitnexus_impact({target: "X", direction: "upstream"})` |
| `detect_changes` | Pre-commit scope check | `gitnexus_detect_changes({scope: "staged"})` |
| `rename` | Safe multi-file rename | `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` |
| `cypher` | Custom graph queries | `gitnexus_cypher({query: "MATCH ..."})` |
| `api_impact` | Pre-change API route impact | `gitnexus_api_impact({route: "/api/users", method: "GET"})` |
| `route_map` | Route → handler → consumer map | `gitnexus_route_map({})` |
| `tool_map` | MCP/RPC tool definitions | `gitnexus_tool_map({})` |
| `shape_check` | Response shape vs consumer access | `gitnexus_shape_check({route: "/api/users"})` |
| `group_list` | List repo groups | `gitnexus_group_list({})` |
| `group_sync` | Rebuild group Contract Registry | `gitnexus_group_sync({name: "myGroup"})` |
| `query` (group mode) | Cross-repo search in a group (RRF-merged) | `gitnexus_query({repo: "@myGroup", query: "auth"})` |
| `context` (group mode) | 360° view across all member repos | `gitnexus_context({repo: "@myGroup", name: "validateUser"})` |
| `impact` (group mode) | Cross-repo blast radius via Contract Bridge | `gitnexus_impact({repo: "@myGroup", target: "X", direction: "upstream"})` |
> Group mode: pass `repo: "@<groupName>"` to fan out across all member repos, or `repo: "@<groupName>/<memberPath>"` to target a single member (path keys from `group.yaml`). Optional `service: "<monorepo/path>"` filters by service root. Group-level state (contracts, staleness) lives in the resources table below — there are **no** `group_query` / `group_context` / `group_impact` / `group_contracts` / `group_status` MCP tools.
>
> For a full walkthrough of setting up a group across multiple repos that communicate over gRPC, see [docs/guides/microservices-grpc.md](docs/guides/microservices-grpc.md).
## Impact Risk Levels
| Depth | Meaning | Action |
|-------|---------|--------|
| d=1 | WILL BREAK — direct callers/importers | MUST update |
| d=2 | LIKELY AFFECTED — indirect deps | Should test |
| d=3 | MAY NEED TESTING — transitive | Test if critical path |
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, index freshness |
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
| `gitnexus://group/{name}/contracts` | Group Contract Registry (provider/consumer rows + cross-links) |
| `gitnexus://group/{name}/status` | Per-member index + Contract Registry staleness report |
## Self-Check Before Finishing
## CLI
1. `gitnexus_impact` was run for all modified symbols
2. No HIGH/CRITICAL warnings were ignored
3. `gitnexus_detect_changes()` confirms expected scope
4. All d=1 dependents were updated
## Keeping the Index Fresh
```bash
npx gitnexus analyze # basic refresh; preserves any existing embeddings
npx gitnexus analyze --embeddings # also generate embeddings for new/changed nodes
npx gitnexus analyze --drop-embeddings # explicit opt-in to wipe existing embeddings
```
Check `.gitnexus/meta.json` `stats.embeddings` (0 = none). A plain `analyze` no longer drops existing vectors — pass `--drop-embeddings` to wipe.
> Claude Code: PostToolUse hook detects a stale index after `git commit` and `git merge` and prompts the agent to run `analyze`. The hook does not invoke `analyze` itself.
## CLI Skills
| Task | Skill file |
|------|-----------|
| Architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Debugging / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Refactoring | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools/resources/schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| CLI commands (index, status, clean, wiki) | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
| Work in the Ingestion area (239 symbols) | `.claude/skills/generated/ingestion/SKILL.md` |
| Work in the Extractors area (135 symbols) | `.claude/skills/generated/extractors/SKILL.md` |
| Work in the Components area (112 symbols) | `.claude/skills/generated/components/SKILL.md` |
| Work in the Lbug area (96 symbols) | `.claude/skills/generated/lbug/SKILL.md` |
| Work in the Group area (94 symbols) | `.claude/skills/generated/group/SKILL.md` |
| Work in the Cli area (92 symbols) | `.claude/skills/generated/cli/SKILL.md` |
| Work in the Configs area (92 symbols) | `.claude/skills/generated/configs/SKILL.md` |
| Work in the Type-extractors area (90 symbols) | `.claude/skills/generated/type-extractors/SKILL.md` |
| Work in the Hooks area (88 symbols) | `.claude/skills/generated/hooks/SKILL.md` |
| Work in the Unit area (80 symbols) | `.claude/skills/generated/unit/SKILL.md` |
| Work in the Cpp area (73 symbols) | `.claude/skills/generated/cpp/SKILL.md` |
| Work in the Scope-resolution area (72 symbols) | `.claude/skills/generated/scope-resolution/SKILL.md` |
| Work in the Server area (66 symbols) | `.claude/skills/generated/server/SKILL.md` |
| Work in the Local area (61 symbols) | `.claude/skills/generated/local/SKILL.md` |
| Work in the Wiki area (60 symbols) | `.claude/skills/generated/wiki/SKILL.md` |
| Work in the Workers area (57 symbols) | `.claude/skills/generated/workers/SKILL.md` |
| Work in the Embeddings area (56 symbols) | `.claude/skills/generated/embeddings/SKILL.md` |
| Work in the Typescript area (53 symbols) | `.claude/skills/generated/typescript/SKILL.md` |
| Work in the Storage area (51 symbols) | `.claude/skills/generated/storage/SKILL.md` |
| Work in the Php area (48 symbols) | `.claude/skills/generated/php/SKILL.md` |
<!-- gitnexus:end -->
+64
View File
@@ -52,3 +52,67 @@ If always-on instructions grow, load deep conventions via conditional reads (e.g
## GitNexus rules
See the `<!-- gitnexus:start --> … <!-- gitnexus:end -->` block in **[AGENTS.md](AGENTS.md)** for the canonical MCP tools, impact analysis rules, and index instructions.
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (26675 symbols, 35395 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## Never Do
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
## CLI
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
| Work in the Ingestion area (239 symbols) | `.claude/skills/generated/ingestion/SKILL.md` |
| Work in the Extractors area (135 symbols) | `.claude/skills/generated/extractors/SKILL.md` |
| Work in the Components area (112 symbols) | `.claude/skills/generated/components/SKILL.md` |
| Work in the Lbug area (96 symbols) | `.claude/skills/generated/lbug/SKILL.md` |
| Work in the Group area (94 symbols) | `.claude/skills/generated/group/SKILL.md` |
| Work in the Cli area (92 symbols) | `.claude/skills/generated/cli/SKILL.md` |
| Work in the Configs area (92 symbols) | `.claude/skills/generated/configs/SKILL.md` |
| Work in the Type-extractors area (90 symbols) | `.claude/skills/generated/type-extractors/SKILL.md` |
| Work in the Hooks area (88 symbols) | `.claude/skills/generated/hooks/SKILL.md` |
| Work in the Unit area (80 symbols) | `.claude/skills/generated/unit/SKILL.md` |
| Work in the Cpp area (73 symbols) | `.claude/skills/generated/cpp/SKILL.md` |
| Work in the Scope-resolution area (72 symbols) | `.claude/skills/generated/scope-resolution/SKILL.md` |
| Work in the Server area (66 symbols) | `.claude/skills/generated/server/SKILL.md` |
| Work in the Local area (61 symbols) | `.claude/skills/generated/local/SKILL.md` |
| Work in the Wiki area (60 symbols) | `.claude/skills/generated/wiki/SKILL.md` |
| Work in the Workers area (57 symbols) | `.claude/skills/generated/workers/SKILL.md` |
| Work in the Embeddings area (56 symbols) | `.claude/skills/generated/embeddings/SKILL.md` |
| Work in the Typescript area (53 symbols) | `.claude/skills/generated/typescript/SKILL.md` |
| Work in the Storage area (51 symbols) | `.claude/skills/generated/storage/SKILL.md` |
| Work in the Php area (48 symbols) | `.claude/skills/generated/php/SKILL.md` |
<!-- gitnexus:end -->
+114 -50
View File
@@ -30,17 +30,17 @@ Format: `<type>[(scope)][!]: <subject>`
Allowed types and the release-notes section each one lands in (defined in `.github/release.yml`):
| Type | Label applied | Release-notes section |
|------|---------------|-----------------------|
| `feat` | `enhancement` | 🚀 Features |
| `fix` | `bug` | 🐛 Bug Fixes |
| `perf` | `performance` | 🏎️ Performance |
| `refactor` | `refactor` | 🔄 Refactoring |
| `test` | `test` | 🧪 Tests |
| `ci` | `ci` | 👷 CI/CD |
| `build` / `deps` | `dependencies` | 📦 Dependencies |
| `docs` | `documentation` | (grouped under Other Changes unless a Docs section is added) |
| `chore` / `revert` | `chore` | (excluded from release notes) |
| Type | Label applied | Release-notes section |
| ------------------ | --------------- | ------------------------------------------------------------ |
| `feat` | `enhancement` | 🚀 Features |
| `fix` | `bug` | 🐛 Bug Fixes |
| `perf` | `performance` | 🏎️ Performance |
| `refactor` | `refactor` | 🔄 Refactoring |
| `test` | `test` | 🧪 Tests |
| `ci` | `ci` | 👷 CI/CD |
| `build` / `deps` | `dependencies` | 📦 Dependencies |
| `docs` | `documentation` | (grouped under Other Changes unless a Docs section is added) |
| `chore` / `revert` | `chore` | (excluded from release notes) |
Append `!` to the type (e.g. `feat(api)!: drop /v1 endpoint`) or include `BREAKING CHANGE:` in the PR body to flag a breaking change — the labeler then adds the `breaking` label and the 💥 Breaking Changes section is rendered first.
@@ -81,17 +81,17 @@ Every workflow under `.github/workflows/` MUST declare a top-level `concurrency:
- **Merge queue (`merge_group`)**: when this event is added, use `${{ github.workflow }}-${{ github.event.merge_group.head_ref }}` with `cancel-in-progress: false` (every queue entry is a distinct ref; never cancel).
- **`cancel-in-progress` policy:**
| Event | `cancel-in-progress` | Why |
|-------|----------------------|-----|
| `pull_request` CI run | `true` | New push supersedes old run |
| `push` to `main` | `false` | Every main commit gets validated |
| Tag push (`v*` publish) | `false` | Never cancel mid-publish |
| `push` to `main` for release-candidate | `false` | Never cancel mid-RC publish |
| `workflow_dispatch` (release/publish) | `false` | Manual runs are intentional |
| `workflow_run` (sticky-comment reports) | `false` | Serialize, don't race |
| Per-PR bot workflows (`@claude`, review) | `false` | Serialize comments per PR |
| PR-meta re-checks (pr-description-check) | `true` | Cheap, latest wins |
| Single-slot utilities (triage sweep) | `true` | Latest dispatch supersedes |
| Event | `cancel-in-progress` | Why |
| ---------------------------------------- | -------------------- | -------------------------------- |
| `pull_request` CI run | `true` | New push supersedes old run |
| `push` to `main` | `false` | Every main commit gets validated |
| Tag push (`v*` publish) | `false` | Never cancel mid-publish |
| `push` to `main` for release-candidate | `false` | Never cancel mid-RC publish |
| `workflow_dispatch` (release/publish) | `false` | Manual runs are intentional |
| `workflow_run` (sticky-comment reports) | `false` | Serialize, don't race |
| Per-PR bot workflows (`@claude`, review) | `false` | Serialize comments per PR |
| PR-meta re-checks (pr-description-check) | `true` | Cheap, latest wins |
| Single-slot utilities (triage sweep) | `true` | Latest dispatch supersedes |
- For workflows that serve multiple events at once (e.g. `ci.yml` handles `pull_request`, `push`, and `workflow_call`), make `cancel-in-progress` event-aware:
@@ -103,22 +103,59 @@ Every workflow under `.github/workflows/` MUST declare a top-level `concurrency:
- When adding a new workflow, copy the concurrency block from an existing workflow of the same event shape.
## CI automation contracts
Two workflows produce machine-readable signals on every PR. Coding agents and humans alike can rely on the names and shapes below — change them with intent.
### `gitnexus/autofix`
`pr-autofix.yml` (untrusted) + `pr-autofix-publish.yml` (trusted) run `prettier --write` and `eslint --fix` against the PR head and surface a single ChatOps button on the PR. Three signals are emitted:
| Surface | Where | Notes |
| ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------- |
| Sticky PR comment | Top-level comment with the HTML marker `<!-- gitnexus:pr-autofix-summary -->` and heading `## :sparkles: PR Autofix`. Only posted when there is something to fix; clean PRs stay silent. | Edit-in-place via marker; one comment per PR. |
| Fenced JSON block | Inside the sticky, fenced as `gitnexus-autofix`. Schema `gitnexus.pr-autofix/v2` with fields `state` (`fixes-available`), `pr_number`, `head_sha`, `changed_lines`, `run_id`, and `apply_command` (literal `/autofix`). | Parseable signal — preferred over regexing prose. v1 fields preserved as a superset. |
| Check Run | Stable name `gitnexus/autofix` on the PR head SHA. Conclusion: `success` (clean) or `neutral` (`fixes-available`). The neutral title is `Autofix available — comment /autofix to apply`. | Surfaced under PR Checks; readable via `gh pr checks <pr>`. |
To detect outcome from an agent: `gh pr checks <pr> --json name,conclusion,output | jq '.[] | select(.name == "gitnexus/autofix")'`.
Forks are supported. The untrusted half runs fork code with `permissions: {}` and ships the diff as an artifact; the trusted publish job consumes only the diff (data, not code) and posts the comment + check run.
#### Applying autofix
Comment `/autofix` on the PR (whole-line, no arguments). The `pr-autofix-apply.yml` workflow:
1. Validates the comment body matches `^/autofix\s*$` exactly. Quoted or inline mentions are silently ignored.
2. Validates the commenter has `admin`, `write`, or `maintain` permission on the repo, OR is the PR author. Other commenters get a 👎 reaction and a refusal reply.
3. Locates the most recent successful `pr-autofix.yml` run for the PR's current head SHA, downloads its `autofix` artifact, applies the patch, and pushes a `chore(autofix): ...` commit back to the PR head branch.
4. Reacts ✅ on success, 👎 on stale-patch / push-failure, and posts a short reply with the apply-run URL in either case.
The apply workflow runs from the default branch's copy of the file regardless of where the comment originates — that's the trust anchor. There is no diff-size cap (the apply workflow uses `git apply` + push, not the GitHub review-comment API).
For fork PRs, the push succeeds only when the contributor has **Allow edits by maintainers** enabled on the PR (the default). When they have disabled it, the workflow fails loud with a 👎 reaction and an explanation comment.
Re-invoking `/autofix` after a successful apply is a safe no-op — the workflow detects the already-applied state via `git apply --check --reverse` and reacts ✅ without pushing.
**Sensitive paths.** The apply workflow refuses any patch that touches `.github/` (workflow files, CODEOWNERS, dependabot config). A malicious PR could ship a custom prettier or ESLint config that reformats workflow YAML; if accepted, those edits would be pushed under `contents: write` without human review. Apply formatter changes to files under `.github/` manually in a normal commit so they get the same review every other workflow change gets.
## AI-assisted contributions
If you use coding agents, follow project context files (e.g. `AGENTS.md`, `CLAUDE.md`) and avoid drive-by refactors unrelated to the issue. Prefer incremental, test-backed changes.
## Releases
Two publish workflows ship `gitnexus` to npm:
One workflow ships `gitnexus` to npm — `.github/workflows/publish.yml`. It
routes between two modes based on the triggering event:
- **Stable** (`.github/workflows/publish.yml`) — triggered by pushing any `v*`
tag. Publishes to the `latest` dist-tag with a changelog-backed GitHub
release. Maintainers are expected to tag from `main` as a convention; the
workflow itself does not enforce branch reachability.
- **Release Candidate** (`.github/workflows/release-candidate.yml`) — runs on
every push to `main` (typically a merged PR) plus manual dispatch. Docs-only
changes are skipped via `paths-ignore`. Publishes to the `rc` dist-tag with
version `X.Y.Z-rc.N` and a GitHub prerelease, where:
- **Stable mode** — triggered by pushing any `v<X.Y.Z>` tag (no `-rc.*`
suffix; RC tags are excluded at trigger via a negative glob). Publishes to
the `latest` dist-tag with a changelog-backed GitHub release. Maintainers
are expected to tag from `main` as a convention; the workflow itself does
not enforce branch reachability. No Docker build (RC-only).
- **Release-candidate mode** — runs on every push to `main` (typically a
merged PR) plus manual `workflow_dispatch`. Docs-only changes are skipped
via `paths-ignore`. Publishes to the `rc` dist-tag with version
`X.Y.Z-rc.N` and a GitHub prerelease, where:
- `X.Y.Z` is selected automatically. On push (and on dispatch with
`bump: auto`, the default) the workflow **continues the active rc cycle**:
if the registry already has `X.Y.Z-rc.*` versions with `X.Y.Z` > current
@@ -135,37 +172,64 @@ Two publish workflows ship `gitnexus` to npm:
caller's ref — see README.md § Docker for the verify command).
Idempotency: the workflow pushes an `rc/<HEAD_SHA>` marker tag and a
`v<RC>` release tag **atomically, before** calling `npm publish`. The guard
refuses to re-run once the marker exists, so a post-publish failure will
not mint a duplicate rc for the same commit. The `v<RC>` tag points at a
detached release commit whose `package.json` matches the npm tarball
exactly (traceable releases). Recovery after a partial failure:
`v<RC>` release tag **atomically, before** calling `npm publish`. The
RC guard refuses to re-run once the marker exists, so a post-publish
failure will not mint a duplicate rc for the same commit. The `v<RC>`
tag points at a detached release commit whose `package.json` matches
the npm tarball exactly (traceable releases). The RC tag is excluded
from this workflow's `push: tags:` filter, so it does **not** re-trigger
publishing — preventing the double-publish failure mode tracked in #1609.
Recovery after a partial failure: the workflow's `if: failure()` cleanup
step in the `publish` job auto-deletes the v-tag and marker on most
post-publish failures, so the typical retry is just:
```bash
gh workflow run publish.yml --ref main -f force=true
# or push a new commit to main, which will cut a fresh RC
```
If auto-cleanup didn't run (e.g. the cleanup step itself failed, or the
failure happened in the route/rc-guard phase before the marker was
pushed), manual cleanup is:
```bash
git push --delete origin rc/<HEAD_SHA> v<RC>
# then redispatch the workflow with force: true
# then redispatch with force: true
```
**Release-PR-skip subject pattern.** The rc-guard job recognizes a
squash-merged release commit by matching the commit subject against
`^chore: release vX.Y.Z` (optionally followed by ` (#NNNN)` for the
squash-merge PR-number suffix). Match is case-insensitive — `Chore: Release v1.2.3`
works too. PRs that should suppress the RC build must either use this
subject shape, or carry the `release` label so the label-based fallback
fires. Other release-style subjects (`chore(release): v1.2.3`,
`release: v1.2.3`) will NOT trigger the skip — please name the release
PR exactly `chore: release vX.Y.Z` to keep the dedup deterministic.
**Docker-only partial failure:** if `publish` succeeds (npm tarball + tags
are live) but the `docker` job subsequently fails (e.g. GHCR flakiness),
the npm RC is already published and the `rc/<HEAD_SHA>` marker is in place.
Re-running `release-candidate.yml` with `force: true` will abort at the
"Version already exists on npm" guard. To recover without cutting a new RC:
Recovery without cutting a new RC:
```bash
# 1. Manually trigger only the docker workflow, passing the existing RC tag:
gh workflow run docker.yml --ref main -f tag=v<RC_VERSION>
# (requires a workflow_dispatch trigger on docker.yml — see note below)
# Re-run only the failed docker job from the original workflow run:
gh run rerun <run-id> --failed
```
Because `docker.yml` intentionally has no `workflow_dispatch` (images are
tag-driven by design), the practical recovery options are:
- Wait for the next commit on `main`, which will cut a new RC that includes
the Docker build.
- Manually run `docker build` + `docker push` locally and sign with Cosign
against the same digest.
- Delete `rc/<HEAD_SHA>` and `v<RC>` tags, then redispatch with `force:
true` to re-run the full RC pipeline (cuts a new RC number).
Find the run ID via `gh run list --workflow=publish.yml --branch main`.
`docker.yml` intentionally has no `workflow_dispatch` trigger (images are
tag-driven by design), so the gh-run-rerun path is the supported recovery.
**GitHub Release transient failure** (npm publish succeeded, Release step
failed): the npm artifact is live but no GitHub Release page exists.
Recover by either re-running the failed job (`gh run rerun <run-id> --failed`),
or creating the Release manually:
```bash
gh release create v<RC> --prerelease --generate-notes # RC
gh release create v<X.Y.Z> --notes-file gitnexus/CHANGELOG.md # stable
```
The rc workflow never moves `latest`. To verify after a change, inspect dist-tags:
+31 -9
View File
@@ -1,24 +1,32 @@
ARG BUILDPLATFORM
ARG TARGETPLATFORM
# Pinned npm version used to replace the bundled npm in the upstream Node
# image. Bumping requires a coordinated update in Dockerfile.web and
# gitnexus/Dockerfile.test so all images bootstrap the same npm.
ARG NPM_VERSION=11.14.1
# ── Builder ────────────────────────────────────────────────────────────
# -- Builder -----------------------------------------------------------
# Native modules (tree-sitter-*, onnxruntime-node, node-gyp builds for
# tree-sitter-proto / tree-sitter-swift) require python3 + a C/C++ toolchain.
FROM node:22-trixie-slim AS builder
# node:22-bookworm-slim
FROM node:22-bookworm-slim@sha256:9f6d5975c7dca860947d3915877f85607946403fc55349f39b4bc3688448bb6e AS builder
ARG NPM_VERSION
WORKDIR /app
RUN npx --yes npm@${NPM_VERSION} install -g npm@${NPM_VERSION}
# Toolchain for node-gyp / native builds.
RUN apt-get update && apt-get install -y --no-install-recommends python3 make g++ git && rm -rf /var/lib/apt/lists/*
# Build gitnexus-shared first — gitnexus depends on it as a workspace.
# Build gitnexus-shared first - gitnexus depends on it as a workspace.
COPY gitnexus-shared/package.json gitnexus-shared/package-lock.json ./gitnexus-shared/
RUN npm ci --prefix gitnexus-shared
COPY gitnexus-shared ./gitnexus-shared
RUN rm -f gitnexus-shared/tsconfig.tsbuildinfo
RUN npm run build --prefix gitnexus-shared
# Copy the full gitnexus package before installing — `npm ci` triggers
# Copy the full gitnexus package before installing - `npm ci` triggers
# `postinstall` (patches tree-sitter-swift, builds the vendored
# tree-sitter-proto) and `prepare` (compiles TypeScript via scripts/build.js),
# both of which need the source tree.
@@ -28,11 +36,15 @@ RUN npm ci --prefix gitnexus
# Drop dev dependencies for a smaller runtime layer.
RUN npm prune --omit=dev --prefix gitnexus
# ── Runtime ────────────────────────────────────────────────────────────
FROM node:22-trixie-slim AS runtime
# -- Runtime -----------------------------------------------------------
# node:22-bookworm-slim
FROM node:22-bookworm-slim@sha256:9f6d5975c7dca860947d3915877f85607946403fc55349f39b4bc3688448bb6e AS runtime
# curl for the healthcheck; git so `gitnexus` can clone repos at runtime.
RUN apt-get update && apt-get install -y --no-install-recommends curl git && rm -rf /var/lib/apt/lists/*
# curl for the healthcheck; git for cloning; ca-certificates for TLS verification.
RUN apt-get update && apt-get install -y --no-install-recommends curl git ca-certificates && rm -rf /var/lib/apt/lists/* \
&& rm -rf /usr/local/lib/node_modules/npm \
&& rm -rf /usr/local/lib/node_modules/corepack \
&& rm -f /usr/local/bin/npm /usr/local/bin/npx /usr/local/bin/corepack
WORKDIR /app
@@ -43,11 +55,21 @@ RUN mkdir -p /data/gitnexus && chown -R node:node /data
COPY --from=builder --chown=node:node /app/gitnexus/dist ./gitnexus/dist
COPY --from=builder --chown=node:node /app/gitnexus/node_modules ./gitnexus/node_modules
COPY --from=builder --chown=node:node /app/gitnexus/package.json ./gitnexus/package.json
COPY --from=builder --chown=node:node /app/gitnexus/scripts/install-duckdb-extension.mjs ./gitnexus/scripts/install-duckdb-extension.mjs
COPY --from=builder --chown=node:node /app/gitnexus/vendor ./gitnexus/vendor
# Expose the `gitnexus` binary on PATH so the documented Docker workflow
# (`docker compose exec gitnexus-server gitnexus index /workspace/<repo>`)
# works without users having to invoke `node /app/gitnexus/dist/cli/index.js`.
# `npm prune --omit=dev` in the builder stage strips `node_modules/.bin/`
# entries, so the `gitnexus` bin declared in package.json (`dist/cli/index.js`,
# which already carries `#!/usr/bin/env node` and 755 perms) is otherwise
# unreachable from $PATH.
RUN ln -s /app/gitnexus/dist/cli/index.js /usr/local/bin/gitnexus
USER node
# The web UI defaults to http://localhost:4747 — keep that contract.
# The web UI defaults to http://localhost:4747 - keep that contract.
ENV GITNEXUS_HOME=/data/gitnexus \
NODE_ENV=production \
PORT=4747
+14 -3
View File
@@ -1,10 +1,17 @@
ARG BUILDPLATFORM
ARG TARGETPLATFORM
# Pinned npm version — keep in sync with Dockerfile.cli and
# gitnexus/Dockerfile.test.
ARG NPM_VERSION=11.14.1
FROM --platform=$BUILDPLATFORM node:22-alpine AS builder
# node:22-bookworm-slim
FROM --platform=$BUILDPLATFORM node:22-bookworm-slim@sha256:9f6d5975c7dca860947d3915877f85607946403fc55349f39b4bc3688448bb6e AS builder
ARG NPM_VERSION
WORKDIR /app
RUN npx --yes npm@${NPM_VERSION} install -g npm@${NPM_VERSION}
COPY gitnexus-shared/package.json gitnexus-shared/package-lock.json ./gitnexus-shared/
RUN npm ci --prefix gitnexus-shared
@@ -19,9 +26,13 @@ RUN npm ci --prefix gitnexus-web
COPY gitnexus-web ./gitnexus-web
RUN npm run build --prefix gitnexus-web
FROM node:22-alpine AS runtime
# node:22-bookworm-slim
FROM node:22-bookworm-slim@sha256:9f6d5975c7dca860947d3915877f85607946403fc55349f39b4bc3688448bb6e AS runtime
RUN apk add --no-cache curl
RUN apt-get update && apt-get install -y --no-install-recommends curl && rm -rf /var/lib/apt/lists/* \
&& rm -rf /usr/local/lib/node_modules/npm \
&& rm -rf /usr/local/lib/node_modules/corepack \
&& rm -f /usr/local/bin/npm /usr/local/bin/npx /usr/local/bin/corepack
WORKDIR /app
+7 -1
View File
@@ -30,9 +30,15 @@ Format: **Trigger → Instruction → Reason**. Append new Signs when the same m
### Stale graph after edits
- **Trigger:** MCP warns index is behind `HEAD`, or search doesn't match latest commit.
- **Do:** `npx gitnexus analyze` (plus `--embeddings` if used).
- **Do:** `npx gitnexus analyze` (plus `--embeddings` if used). Runs incrementally by default — the pipeline parses every file every run (cross-file resolution requires it), but tree-sitter dispatch is skipped for unchanged file chunks via the content-addressed cache, and only changed-file rows (plus their importers, transitively) are rewritten in LadybugDB.
- **Why:** Tools query LadybugDB from last analyze; git changes are invisible until re-indexed.
### Index seems corrupt or "incremental" is misbehaving
- **Trigger:** `analyze` produces unexpected results, or `meta.json.incrementalInProgress` is set, or the index is in a half-state after a crash.
- **Do:** `npx gitnexus analyze --force` to rebuild from scratch. The dirty-flag check forces this automatically when a previous incremental run didn't complete cleanly, but `--force` is the manual escape hatch. Safe to delete the `.gitnexus/parse-cache/` directory (and any legacy `.gitnexus/parse-cache.json`) at any time — content-addressed, will be regenerated.
- **Why:** Incremental writeback is selective DB row replacement; if the on-disk state is inconsistent for any reason, a full rebuild is the cheapest path back to a known-good index.
### Embeddings vanished after analyze
- **Trigger:** Semantic search quality drops; `stats.embeddings` in `meta.json` is 0 after refresh.
+114 -81
View File
@@ -1,4 +1,5 @@
# GitNexus
**⚠️ Important Notice:** GitNexus has NO official cryptocurrency, token, or coin. Any token/coin using the GitNexus name on Pump.fun or any other platform is **not affiliated with, endorsed by, or created by** this project or its maintainers. Do not purchase any cryptocurrency claiming association with GitNexus.
<div align="center">
@@ -30,14 +31,9 @@
Indexes any codebase into a knowledge graph — every dependency, call chain, cluster, and execution flow — then exposes it through smart tools so AI agents never miss code.
https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
> *Like DeepWiki, but deeper.* DeepWiki helps you *understand* code. GitNexus lets you *analyze* it — because a knowledge graph tracks every relationship, not just descriptions.
> _Like DeepWiki, but deeper._ DeepWiki helps you _understand_ code. GitNexus lets you _analyze_ it — because a knowledge graph tracks every relationship, not just descriptions.
**TL;DR:** The **Web UI** is a quick way to chat with any repo. The **CLI + MCP** is how you make your AI agent actually reliable — it gives Cursor, Claude Code, Codex, and friends a deep architectural view of your codebase so they stop missing dependencies, breaking call chains, and shipping blind edits. Even smaller models get full architectural clarity, making it compete with Goliath models.
@@ -47,18 +43,17 @@ https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
[![Star History Chart](https://api.star-history.com/svg?repos=abhigyanpatwari/GitNexus&type=date&legend=top-left)](https://www.star-history.com/#abhigyanpatwari/GitNexus&type=date&legend=top-left)
## Two Ways to Use GitNexus
| | **CLI + MCP** | **Web UI** |
| ----------------- | -------------------------------------------------------------- | ------------------------------------------------------------ |
| **What** | Index repos locally, connect AI agents via MCP | Visual graph explorer + AI chat in browser |
| **For** | Daily development with Cursor, Claude Code, Codex, Windsurf, OpenCode | Quick exploration, demos, one-off analysis |
| **Scale** | Full repos, any size | Limited by browser memory (~5k files), or unlimited via backend mode |
| **Install** | `npm install -g gitnexus` | No install — [gitnexus.vercel.app](https://gitnexus.vercel.app) |
| **Storage** | LadybugDB native (fast, persistent) | LadybugDB WASM (in-memory, per session) |
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
| **Privacy** | Everything local, no network | Everything in-browser, no server |
| | **CLI + MCP** | **Web UI** |
| ----------- | --------------------------------------------------------------------- | -------------------------------------------------------------------- |
| **What** | Index repos locally, connect AI agents via MCP | Visual graph explorer + AI chat in browser |
| **For** | Daily development with Cursor, Claude Code, Codex, Windsurf, OpenCode | Quick exploration, demos, one-off analysis |
| **Scale** | Full repos, any size | Limited by browser memory (~5k files), or unlimited via backend mode |
| **Install** | `npm install -g gitnexus` | No install — [gitnexus.vercel.app](https://gitnexus.vercel.app) |
| **Storage** | LadybugDB native (fast, persistent) | LadybugDB WASM (in-memory, per session) |
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
| **Privacy** | Everything local, no network | Everything in-browser, no server |
> **Bridge mode:** `gitnexus serve` connects the two — the web UI auto-detects the local server and can browse all your CLI-indexed repos without re-uploading or re-indexing.
@@ -69,6 +64,7 @@ https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
GitNexus is available as an **enterprise offering** - either as a fully managed **SaaS** or a **self-hosted** deployment. Also available for **commercial use** of the OSS version with proper licensing.
Enterprise includes:
- **PR Review** - automated blast radius analysis on pull requests
- **Auto-updating Code Wiki** - always up-to-date documentation (Code Wiki is also available in OSS)
- **Auto-reindexing** - knowledge graph stays fresh automatically
@@ -77,6 +73,7 @@ Enterprise includes:
- **Priority feature/language support** - request new languages or features
**Upcoming:**
- Auto regression forensics
- End-to-end test generation
@@ -109,7 +106,7 @@ That's it. This indexes the codebase, installs agent skills, registers Claude Co
To configure MCP for your editor, run `npx gitnexus setup` once — or set it up manually below.
> **Faster install (no C++ toolchain needed):** set `GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1` before `npm install -g gitnexus` to skip the native `tree-sitter-dart` and `tree-sitter-proto` builds. Dart/Proto files won't be parsed, but install completes in seconds without `python3`/`make`/`g++`. Strict `=1` only — any other value falls through to the rebuild.
> **Faster install (no C++ toolchain needed):** set `GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1` before `npm install -g gitnexus` to skip vendored grammar materialize/build (`tree-sitter-dart`, `tree-sitter-proto`, `tree-sitter-swift`). Dart/Proto/Swift files won't be parsed, but install completes in seconds without `python3`/`make`/`g++`. Strict `=1` only — any other value falls through to the rebuild.
### MCP Setup
@@ -117,13 +114,13 @@ To configure MCP for your editor, run `npx gitnexus setup` once — or set it up
### Editor Support
| Editor | MCP | Skills | Hooks (auto-augment) | Support |
| --------------------- | --- | ------ | -------------------- | -------------- |
| **Claude Code** | Yes | Yes | Yes (PreToolUse + PostToolUse) | **Full** |
| **Cursor** | Yes | Yes | — | MCP + Skills |
| **Codex** | Yes | Yes | — | MCP + Skills |
| **Windsurf** | Yes | — | — | MCP |
| **OpenCode** | Yes | Yes | — | MCP + Skills |
| Editor | MCP | Skills | Hooks (auto-augment) | Support |
| --------------- | --- | ------ | --------------------------------------------------------------------------------------- | ------------ |
| **Claude Code** | Yes | Yes | Yes (PreToolUse + PostToolUse) | **Full** |
| **Cursor** | Yes | Yes | Yes (postToolUse, [manual install](gitnexus-cursor-integration/README.md#hook-install)) | **Full** |
| **Codex** | Yes | Yes | — | MCP + Skills |
| **Windsurf** | Yes | — | — | MCP |
| **OpenCode** | Yes | Yes | — | MCP + Skills |
> **Claude Code** gets the deepest integration: MCP tools + agent skills + PreToolUse hooks that enrich searches with graph context + PostToolUse hooks that detect a stale index after commits and prompt the agent to reindex.
@@ -131,10 +128,10 @@ To configure MCP for your editor, run `npx gitnexus setup` once — or set it up
Built by the community — not officially maintained, but worth checking out.
| Project | Author | Description |
|---------|--------|-------------|
| [pi-gitnexus](https://github.com/tintinweb/pi-gitnexus) | [@tintinweb](https://github.com/tintinweb) | GitNexus plugin for [pi](https://pi.dev) — `pi install npm:pi-gitnexus` |
| [gitnexus-stable-ops](https://github.com/ShunsukeHayashi/gitnexus-stable-ops) | [@ShunsukeHayashi](https://github.com/ShunsukeHayashi) | Stable ops & deployment workflows (Miyabi ecosystem) |
| Project | Author | Description |
| ----------------------------------------------------------------------------- | ------------------------------------------------------ | ----------------------------------------------------------------------- |
| [pi-gitnexus](https://github.com/tintinweb/pi-gitnexus) | [@tintinweb](https://github.com/tintinweb) | GitNexus plugin for [pi](https://pi.dev) — `pi install npm:pi-gitnexus` |
| [gitnexus-stable-ops](https://github.com/ShunsukeHayashi/gitnexus-stable-ops) | [@ShunsukeHayashi](https://github.com/ShunsukeHayashi) | Stable ops & deployment workflows (Miyabi ecosystem) |
> Have a project built on GitNexus? Open a PR to add it here!
@@ -197,7 +194,8 @@ args = ["-y", "gitnexus@latest", "mcp"]
```bash
gitnexus setup # Configure MCP for your editors (one-time)
gitnexus analyze [path] # Index a repository (or update stale index)
gitnexus analyze --force # Force full re-index
gitnexus analyze --repair-fts # Fast path: rebuild/verify only FTS indexes on existing index data
gitnexus analyze --force # Full rebuild: re-parse + graph rebuild + FTS rebuild
gitnexus analyze --skills # Generate repo-specific skill files from detected communities
gitnexus analyze --skip-embeddings # Skip embedding generation (faster)
gitnexus analyze --skip-agents-md # Preserve custom AGENTS.md/CLAUDE.md gitnexus section edits
@@ -205,6 +203,7 @@ gitnexus analyze --skip-git # Index folders that are not Git repositories
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
gitnexus analyze --worker-timeout 60 # Increase worker idle timeout for slow parses
gitnexus analyze --workers <n> # Parse worker pool size (default: cores-1, capped at 16; 0 = sequential)
gitnexus mcp # Start MCP server (stdio) — serves all indexed repos
gitnexus serve # Start local HTTP server (multi-repo) for web UI connection
gitnexus list # List all indexed repositories
@@ -214,6 +213,7 @@ gitnexus clean --all --force # Delete all indexes
gitnexus wiki [path] # Generate repository wiki from knowledge graph
gitnexus wiki --model <model> # Wiki with custom LLM model (default: gpt-4o-mini)
gitnexus wiki --base-url <url> # Wiki with custom LLM API base URL
gitnexus publish # Notify the understand-quickly registry (opt-in, see below)
# Repository groups (multi-repo / monorepo service tracking)
gitnexus group create <name> # Create a repository group
@@ -228,31 +228,56 @@ gitnexus group status <name> # Check staleness of repos in a group
If `analyze` reports a worker parse timeout on a large or unusual repository, it keeps running and falls back safely. To give slow worker jobs more time, use `gitnexus analyze --worker-timeout 60` or set `GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS=60000`. For very large files, `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` controls the worker job byte budget.
#### Environment variables
Most `analyze` knobs are also CLI flags (`--workers`, `--worker-timeout`, `--max-file-size`, `--verbose`). Use the env-var form when you'd otherwise repeat the same flag every run, or when invoking GitNexus from a long-running host (MCP server, eval-server, CI shell) that already manages its own environment. CLI flags take precedence over env vars; env vars take precedence over built-in defaults.
| Variable | Default | Effect | Tune when… |
| -------------------------------------- | ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- |
| `GITNEXUS_WORKER_POOL_SIZE` | `cores - 1`, capped at 16 | Parse worker pool size. `0` disables the pool (sequential fallback). Equivalent to `--workers <n>`. | Constrained containers (cgroup CPU limits), CI runners with explicit quotas, or debugging a worker-only crash via `0`. |
| `GITNEXUS_PARSE_CHUNK_CONCURRENCY` | `2` | Number of chunks whose file contents may be read into memory in parallel while the pool dispatches the current chunk. Worker dispatch itself stays serial. | Repos large enough to chunk (multi-MB total source) where disk I/O is a measurable fraction of analyze wall-clock. |
| `GITNEXUS_VERBOSE` | unset | When `1`, enables verbose ingestion logs (skipped-file warnings, per-chunk throughput, parse-cache stats). Equivalent to `--verbose`. | Debugging an analyze that "completed" but seems to have missed files; tuning `--workers` / chunk concurrency against observable throughput. |
| `GITNEXUS_MAX_FILE_SIZE` | `512` (KB) | Walker skip threshold in KB. Hard cap is `32768` (tree-sitter buffer ceiling). Equivalent to `--max-file-size <kb>`. | Indexing repos with intentionally-large source files (generated parsers, vendored bundles) that should still be parsed. |
| `GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS` | `30000` | Worker idle timeout in milliseconds before retry/fallback. Equivalent to `--worker-timeout <seconds>` × 1000. | Slow-parsing files (large minified JS, deeply-nested TS types) that legitimately need more than 30s. |
| `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` | `8388608` (8 MB) | Per-job byte budget the pool will send to a worker in one `postMessage`. | Very large individual files; mostly diagnostic — bumping past 8 MB risks structured-clone memory pressure. |
| `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT` | `3` | Max replacement spawns per worker slot before the slot is dropped from the active rotation. Bounds respawn loops on a chronically-crashing slot. | Hosts where a flaky worker should retry more (raise) or fail-fast (lower) before the slot is dropped. |
| `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` | `5 × subBatchTimeoutMs` | Total retry wall-time budget per job before quarantining. Combined with `timeoutBackoffFactor`, prevents exponentially-growing retries from stalling for hours. | Slow files that legitimately need long total retry windows; lower to fail-fast on stalls. |
| `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD`| `max(3, poolSize)` | Per-slot consecutive deaths before the pool's circuit breaker trips. After tripping, every subsequent dispatch rejects until a fresh pool is created. | Hosts where a SIGSEGV-prone native grammar should trip the breaker sooner; CI runners that should fail loudly. |
| `GITNEXUS_CHUNK_BYTE_BUDGET` | `2097152` (2 MB) | Chunk boundary used for cache-key composition and dispatch. Smaller = finer-grained cache hits but more dispatch overhead. | Tuning incremental-analyze cache behavior on monorepos. |
| `GITNEXUS_NO_GITIGNORE` | unset | When set, skips `.gitignore` parsing. `.gitnexusignore` is still honored. | Indexing a repo whose `.gitignore` excludes files you actually want indexed (e.g., generated code committed for cross-repo lookup). |
| `GITNEXUS_SKIP_OPTIONAL_GRAMMARS` | unset | When `=1` strictly, skips vendored grammar materialize/build for `tree-sitter-dart`, `tree-sitter-proto`, and `tree-sitter-swift` at install time. | Installing on a host without a C++ toolchain or where Swift prebuilds don't match; you're willing to skip Dart/Proto/Swift parsing. |
#### Publishing to understand-quickly (opt-in)
[`looptech-ai/understand-quickly`](https://github.com/looptech-ai/understand-quickly) is a public registry of code-knowledge graphs that lists `gitnexus@1` as a first-class format. After registering your repo once (`npx @understand-quickly/cli add` or the [wizard](https://looptech-ai.github.io/understand-quickly/add.html)), `gitnexus publish` fires a single `repository_dispatch` event so the registry resyncs your entry on demand instead of waiting for the nightly job.
It is opt-in and a no-op without `UNDERSTAND_QUICKLY_TOKEN` — a fine-grained GitHub PAT with `Repository dispatches: write` on the registry repo. Nothing else happens; no graph file is uploaded. See the [protocol spec](https://github.com/looptech-ai/understand-quickly/blob/main/docs/integrations/protocol.md) for the full contract.
### What Your AI Agent Gets
**16 tools** exposed via MCP (11 per-repo + 5 group):
| Tool | What It Does | `repo` Param |
| ------------------ | ----------------------------------------------------------------- | -------------- |
| `list_repos` | Discover all indexed repositories | — |
| `query` | Process-grouped hybrid search (BM25 + semantic + RRF) | Optional |
| `context` | 360-degree symbol view — categorized refs, process participation | Optional |
| `impact` | Blast radius analysis with depth grouping and confidence | Optional |
| `detect_changes` | Git-diff impact — maps changed lines to affected processes | Optional |
| `rename` | Multi-file coordinated rename with graph + text search | Optional |
| `cypher` | Raw Cypher graph queries | Optional |
| `group_list` | List configured repository groups | — |
| `group_sync` | Extract contracts and match across repos/services | — |
| `group_contracts`| Inspect extracted contracts and cross-links | — |
| `group_query` | Search execution flows across all repos in a group | — |
| `group_status` | Check staleness of repos in a group | — |
| Tool | What It Does | `repo` Param |
| ----------------- | ---------------------------------------------------------------- | ------------ |
| `list_repos` | Discover all indexed repositories | — |
| `query` | Process-grouped hybrid search (BM25 + semantic + RRF) | Optional |
| `context` | 360-degree symbol view — categorized refs, process participation | Optional |
| `impact` | Blast radius analysis with depth grouping and confidence | Optional |
| `detect_changes` | Git-diff impact — maps changed lines to affected processes | Optional |
| `rename` | Multi-file coordinated rename with graph + text search | Optional |
| `cypher` | Raw Cypher graph queries | Optional |
| `group_list` | List configured repository groups | — |
| `group_sync` | Extract contracts and match across repos/services | — |
| `group_contracts` | Inspect extracted contracts and cross-links | — |
| `group_query` | Search execution flows across all repos in a group | — |
| `group_status` | Check staleness of repos in a group | — |
> When only one repo is indexed, the `repo` parameter is optional. With multiple repos, specify which one: `query({query: "auth", repo: "my-app"})`.
**Resources** for instant context:
| Resource | Purpose |
| ----------------------------------------- | ---------------------------------------------------- |
| Resource | Purpose |
| --------------------------------------- | ---------------------------------------------------- |
| `gitnexus://repos` | List all indexed repositories (read this first) |
| `gitnexus://repo/{name}/context` | Codebase stats, staleness check, and available tools |
| `gitnexus://repo/{name}/clusters` | All functional clusters with cohesion scores |
@@ -263,9 +288,9 @@ If `analyze` reports a worker parse timeout on a large or unusual repository, it
**2 MCP prompts** for guided workflows:
| Prompt | What It Does |
| ----------------- | ------------------------------------------------------------------------- |
| `detect_impact` | Pre-commit change analysis — scope, affected processes, risk level |
| Prompt | What It Does |
| --------------- | ------------------------------------------------------------------------- |
| `detect_impact` | Pre-commit change analysis — scope, affected processes, risk level |
| `generate_map` | Architecture documentation from the knowledge graph with mermaid diagrams |
**4 agent skills** installed to `.claude/skills/` automatically:
@@ -352,10 +377,10 @@ npx gitnexus@latest serve
The official Docker setup ships **two signed images** orchestrated by `docker-compose.yaml`. Each image is published to both **GitHub Container Registry** (GHCR) and **Docker Hub** — same build, same digest, same Cosign signature — so pick whichever registry you prefer:
| Purpose | GHCR (default in `docker-compose.yaml`) | Docker Hub mirror |
| ---------------------------------------------------------------------- | --------------------------------------------- | ------------------------------------------- |
| CLI / `gitnexus serve` backend (HTTP API on port `4747`, MCP, indexer) | `ghcr.io/abhigyanpatwari/gitnexus:latest` | `akonlabs/gitnexus:latest` |
| Static web UI (port `4173`) | `ghcr.io/abhigyanpatwari/gitnexus-web:latest` | `akonlabs/gitnexus-web:latest` |
| Purpose | GHCR (default in `docker-compose.yaml`) | Docker Hub mirror |
| ---------------------------------------------------------------------- | --------------------------------------------- | ------------------------------ |
| CLI / `gitnexus serve` backend (HTTP API on port `4747`, MCP, indexer) | `ghcr.io/abhigyanpatwari/gitnexus:latest` | `akonlabs/gitnexus:latest` |
| Static web UI (port `4173`) | `ghcr.io/abhigyanpatwari/gitnexus-web:latest` | `akonlabs/gitnexus-web:latest` |
> **Heads-up — image rename.** Earlier releases published the web UI under
> `ghcr.io/abhigyanpatwari/gitnexus`. Starting with the introduction of the
@@ -422,7 +447,7 @@ The Docker images are version-locked to the npm package:
Both registries receive the same digest from a single build step, so you can
pull from either and the signature verifies identically.
- Release-candidate images (e.g. `:1.7.0-rc.1`) are published alongside each
RC npm release. They are built by `release-candidate.yml` calling `docker.yml`
RC npm release. They are built by `publish.yml` calling `docker.yml`
as a reusable workflow after the RC tag is created and pushed.
- `:latest` is auto-promoted only from non-prerelease tags by the Docker
metadata action, so it always points at a real, npm-published version.
@@ -455,7 +480,7 @@ registries because both sets of tags were signed at the same digest in one
workflow run.
**Release candidates** — signed from `refs/heads/main` (the caller's ref when
`release-candidate.yml` invokes `docker.yml` as a reusable workflow):
`publish.yml` invokes `docker.yml` as a reusable workflow):
```bash
cosign verify ghcr.io/abhigyanpatwari/gitnexus:1.7.0-rc.1 \
@@ -571,22 +596,22 @@ GitNexus builds a complete knowledge graph of your codebase through a multi-phas
### Supported Languages
| Language | Imports | Named Bindings | Exports | Heritage | Type Annotations | Constructor Inference | Config | Frameworks | Entry Points |
|----------|---------|----------------|---------|----------|-----------------|---------------------|--------|------------|-------------|
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| JavaScript | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Java | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Kotlin | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Go | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Rust | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| PHP | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ |
| Ruby | ✓ | — | ✓ | ✓ | — | ✓ | — | ✓ | ✓ |
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Dart | ✓ | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Language | Imports | Named Bindings | Exports | Heritage | Type Annotations | Constructor Inference | Config | Frameworks | Entry Points |
| ---------- | ------- | -------------- | ------- | -------- | ---------------- | --------------------- | ------ | ---------- | ------------ |
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| JavaScript | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Java | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Kotlin | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Go | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Rust | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| PHP | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ |
| Ruby | ✓ | — | ✓ | ✓ | — | ✓ | — | ✓ | ✓ |
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Dart | ✓ | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
**Imports** — cross-file import resolution · **Named Bindings** — `import { X as Y }` / re-export tracking · **Exports** — public/exported symbol detection · **Heritage** — class inheritance, interfaces, mixins · **Type Annotations** — explicit type extraction for receiver resolution · **Constructor Inference** — infer receiver type from constructor calls (`self`/`this` resolution included for all languages) · **Config** — language toolchain config parsing (tsconfig, go.mod, etc.) · **Frameworks** — AST-based framework pattern detection · **Entry Points** — entry point scoring heuristics
@@ -715,6 +740,14 @@ gitnexus wiki --base-url https://api.anthropic.com/v1
# Force full regeneration
gitnexus wiki --force
# Increase the timeout or retries for large codebase or slow LLM providers
gitnexus wiki --timeout <seconds> # LLM request timeout in seconds (default: disabled)
gitnexus wiki --retries <n> # Max LLM retry attempts per request (default: 3)
# Change the language generation for wiki
gitnexus wiki --lang <lang> # Output language for generated documentation (e.g. english, chinese, spanish, japanese)
```
The wiki generator reads the indexed graph structure, groups files into modules via LLM, generates per-module documentation pages, and creates an overview page — all with cross-references to the knowledge graph.
@@ -723,16 +756,16 @@ The wiki generator reads the indexed graph structure, groups files into modules
## Tech Stack
| Layer | CLI | Web |
| ------------------------- | ------------------------------------- | --------------------------------------- |
| Layer | CLI | Web |
| ------------------- | ------------------------------------- | --------------------------------------- |
| **Runtime** | Node.js (native) | Browser (WASM) |
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
| **Database** | LadybugDB native | LadybugDB WASM |
| **Database** | LadybugDB native | LadybugDB WASM |
| **Embeddings** | HuggingFace transformers.js (GPU/CPU) | transformers.js (WebGPU/WASM) |
| **Search** | BM25 + semantic + RRF | BM25 + semantic + RRF |
| **Agent Interface** | MCP (stdio) | LangChain ReAct agent |
| **Visualization** | — | Sigma.js + Graphology (WebGL) |
| **Frontend** | — | React 18, TypeScript, Vite, Tailwind v4 |
| **Visualization** | — | Sigma.js + Graphology (WebGL) |
| **Frontend** | — | React 18, TypeScript, Vite, Tailwind v4 |
| **Clustering** | Graphology | Graphology |
| **Concurrency** | Worker threads + async | Web Workers + Comlink |
@@ -748,12 +781,12 @@ The wiki generator reads the indexed graph structure, groups files into modules
### Recently Completed
- [X] Constructor-Inferred Type Resolution, `self`/`this` Receiver Mapping
- [X] Wiki Generation, Multi-File Rename, Git-Diff Impact Analysis
- [X] Process-Grouped Search, 360-Degree Context, Claude Code Hooks
- [X] Multi-Repo MCP, Zero-Config Setup, 14 Language Support
- [X] Community Detection, Process Detection, Confidence Scoring
- [X] Hybrid Search, Vector Index
- [x] Constructor-Inferred Type Resolution, `self`/`this` Receiver Mapping
- [x] Wiki Generation, Multi-File Rename, Git-Diff Impact Analysis
- [x] Process-Grouped Search, 360-Degree Context, Claude Code Hooks
- [x] Multi-Repo MCP, Zero-Config Setup, 14 Language Support
- [x] Community Detection, Process Detection, Confidence Scoring
- [x] Hybrid Search, Vector Index
---
+95
View File
@@ -0,0 +1,95 @@
/**
* Custom ESLint rule: require `parseSourceSafe(parser, content, ...)` instead
* of direct `<parser>.parse(<content>, ...)` calls.
*
* Background: tree-sitter's Node.js native binding crashes with SIGSEGV on
* Windows when handed a JS string longer than 32 767 chars. The crash happens
* inside the binding's V8 string-to-buffer conversion and cannot be intercepted
* by JavaScript `try/catch`. `parseSourceSafe` (in
* `gitnexus/src/core/tree-sitter/safe-parse.ts`) routes large inputs through
* the chunked-callback overload of `parser.parse(input, ...)` which bypasses
* the broken conversion path. PR #1433 fixed every direct call site at the
* time; this rule prevents new direct calls from creeping in.
*
* The rule is auto-fixable for the call-site rewrite. It does NOT auto-add the
* import (computing the correct relative path per file is brittle); after the
* call rewrite runs, the consumer file's `tsc` will complain about an
* undefined identifier and the developer adds the import. This is the same
* tradeoff `unused-imports/no-unused-imports` makes in the opposite direction.
*
* False-positive suppression:
* - Skips calls whose receiver is a known non-tree-sitter library (`JSON`,
* `URL`, `marked`, `Number`).
* - Skips calls whose first argument is a string-literal (grammar-load smoke
* tests like `_testParser.parse('service X { rpc Y (R) returns (R); }')`).
* - Skips test files (`.test.ts`/`.test.tsx`/`.spec.ts`).
* - Skips the `safe-parse.ts` helper itself.
*/
const SKIPPED_RECEIVERS = new Set(['JSON', 'URL', 'marked', 'Number', 'Math']);
export default {
meta: {
type: 'problem',
docs: {
description:
'Require parseSourceSafe instead of direct tree-sitter `<parser>.parse(content, ...)` calls (Windows SIGSEGV protection)',
recommended: true,
},
fixable: 'code',
schema: [],
messages: {
useSafeParse:
'Direct `{{receiver}}.parse(...)` can SIGSEGV on Windows for inputs > 32 767 chars (uncatchable from JS). Use `parseSourceSafe({{receiver}}, ...)` from `core/tree-sitter/safe-parse.js`. Auto-fix rewrites the call; add the missing import yourself.',
},
},
create(context) {
const filename = context.filename ?? context.getFilename();
// Don't lint the helper itself or test files.
if (filename.includes('safe-parse')) return {};
if (/[.](?:test|spec)\.tsx?$/.test(filename)) return {};
const sourceCode = context.sourceCode ?? context.getSourceCode();
return {
CallExpression(node) {
const callee = node.callee;
if (callee.type !== 'MemberExpression') return;
if (callee.computed) return;
if (callee.property.type !== 'Identifier') return;
if (callee.property.name !== 'parse') return;
// Skip known non-tree-sitter receivers.
if (callee.object.type === 'Identifier' && SKIPPED_RECEIVERS.has(callee.object.name)) {
return;
}
// Smoke tests pass a string literal directly; those are trivially safe.
const firstArg = node.arguments[0];
if (!firstArg) return;
if (firstArg.type === 'Literal' && typeof firstArg.value === 'string') return;
if (firstArg.type === 'TemplateLiteral' && firstArg.expressions.length === 0) return;
const receiverText = sourceCode.getText(callee.object);
// Receiver-text-shape skip: anything matching well-known JS APIs that
// happen to have a `.parse(<expr>)` shape but aren't tree-sitter.
if (
/^(JSON|URL|marked|Number|Math|Date|globalThis\.JSON)\b/.test(receiverText) ||
/\bjson\.parse\b/i.test(receiverText)
) {
return;
}
context.report({
node,
messageId: 'useSafeParse',
data: { receiver: receiverText },
fix(fixer) {
const argsText = node.arguments.map((arg) => sourceCode.getText(arg)).join(', ');
return fixer.replaceText(node, `parseSourceSafe(${receiverText}, ${argsText})`);
},
});
},
};
},
};
+26
View File
@@ -3,6 +3,15 @@ import tsParser from '@typescript-eslint/parser';
import unusedImports from 'eslint-plugin-unused-imports';
import reactHooks from 'eslint-plugin-react-hooks';
import prettierConfig from 'eslint-config-prettier';
import requireSafeParse from './eslint-rules/require-safe-parse.mjs';
// Local plugin hosting custom rules that enforce GitNexus-specific invariants
// (currently: the Windows-SIGSEGV-safe parser entrypoint).
const gitnexusLocalPlugin = {
rules: {
'require-safe-parse': requireSafeParse,
},
};
// Selectors that protect MCP-reachable code from corrupting the JSON-RPC
// stdio frame stream. The MCP-reachable block below uses these directly;
@@ -135,6 +144,23 @@ export default [
},
},
// Windows SIGSEGV protection: every tree-sitter parse in `core/` must route
// through parseSourceSafe. Direct `<parser>.parse(content, ...)` crashes on
// Windows for inputs > 32 767 chars (V8 string-conversion bug, uncatchable
// from JS). The rule auto-fixes the call site; the developer adds the
// missing import after the fix runs. Out of scope: tests (skipped by the
// rule), the helper itself (`safe-parse.ts`), and the `grpc-patterns/proto.ts`
// grammar-load smoke test (filtered by string-literal-arg skip in the rule).
{
files: ['gitnexus/src/core/**/*.ts'],
plugins: {
gitnexus: gitnexusLocalPlugin,
},
rules: {
'gitnexus/require-safe-parse': 'error',
},
},
// React-specific rules for gitnexus-web
{
files: ['gitnexus-web/src/**/*.{ts,tsx}'],
+57 -2
View File
@@ -162,8 +162,8 @@ Each mode has a `system_{mode}.jinja` + `instance_{mode}.jinja` pair. The agent
```
Agent → bash command → /usr/local/bin/gitnexus-query
→ curl localhost:4848/tool/query (fast path: eval-server, ~100ms)
→ npx gitnexus query (fallback: cold CLI, ~5-10s)
→ curl http://127.0.0.1:4848/tool/query (fast path: eval-server, ~100ms)
→ npx gitnexus query (fallback: cold CLI, ~5-10s)
```
Each tool script in `/usr/local/bin/` is standalone — no sourcing, no env inheritance needed. This is critical because mini-swe-agent runs every command via `subprocess.run` in a fresh subshell.
@@ -176,6 +176,61 @@ The eval-server is a lightweight HTTP daemon that:
- Includes next-step hints to guide tool chaining (query → context → impact → fix)
- Auto-shuts down after idle timeout
**CLI flags:**
| Flag | Default | Purpose |
|------|---------|---------|
| `--port <port>` | `4848` | Port to listen on |
| `--host <host>` | `127.0.0.1` | Bind address — use `0.0.0.0` for cross-container access |
| `--idle-timeout <seconds>` | `0` (disabled) | Auto-shutdown after N seconds of inactivity |
**READY signal:**
When the server is ready, it writes to stdout:
```
# IPv4
GITNEXUS_EVAL_SERVER_READY:127.0.0.1:4848
# IPv6 (bracketed to avoid colon ambiguity)
GITNEXUS_EVAL_SERVER_READY:[::1]:4848
```
Parse the port as the last colon-segment (`split(':').pop()`) — not `split(':')[1]`, which breaks for IPv6 and for non-loopback IPv4 hosts added in this release.
### Custom port and host
`run_eval.py` does not expose `--port` or `--host` as CLI flags. Configure them in your mode YAML under the `environment:` key:
```yaml
# configs/modes/native_augment.yaml (or whichever mode you're running)
environment:
eval_server_port: 4849 # change if 4848 is already in use on the host
eval_server_host: "0.0.0.0" # bind all interfaces — needed for cross-container setups
```
Defaults are `port: 4848` and `host: 127.0.0.1` (loopback only). Use `0.0.0.0` only when the agent container needs to reach the eval-server from a separate network namespace. The health probe and tool scripts connect via the configured bind host (defaulting to `127.0.0.1`), which is reachable for both loopback and all-interface binds.
`"localhost"` is also a valid `eval_server_host` value. The OS resolves it at bind time — typically `127.0.0.1` on dual-stack or IPv4-only systems, and `::1` on IPv6-only systems. The exact result depends on your `/etc/hosts` and `gai.conf`. The READY signal will reflect the actual bound address (e.g. `GITNEXUS_EVAL_SERVER_READY:127.0.0.1:4848` or `GITNEXUS_EVAL_SERVER_READY:[::1]:4848`), not the literal string `localhost`. Use this when you want the server to bind to whichever loopback address the OS prefers rather than forcing IPv4.
**Running eval-server directly in Docker / Docker Compose:**
```bash
# Bind to all interfaces so sibling containers can reach it
gitnexus eval-server --host 0.0.0.0 --port 4848
# Then probe from a sibling container via its service hostname
curl http://eval-container:4848/health
```
If you need a non-default port (e.g. to avoid conflicts), pass `--port <port>` alongside `--host`. The READY signal will reflect both:
```
GITNEXUS_EVAL_SERVER_READY:0.0.0.0:5000
```
Parse the port as the last colon-segment (`split(':').pop()`) — safe for both IPv4 and bracketed IPv6 forms.
### Index caching
SWE-bench repos repeat (Django has 200+ instances at different commits). The harness caches GitNexus indexes per `(repo, commit)` hash in `~/.gitnexus-eval-cache/` to avoid redundant re-indexing.
+18 -5
View File
@@ -39,6 +39,7 @@ logger = logging.getLogger("gitnexus_docker")
DEFAULT_CACHE_DIR = Path.home() / ".gitnexus-eval-cache"
EVAL_SERVER_PORT = 4848
EVAL_SERVER_HOST = "127.0.0.1"
class GitNexusDockerEnvironment(DockerEnvironment):
@@ -62,6 +63,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
skip_embeddings: bool = True,
gitnexus_timeout: int = 120,
eval_server_port: int = EVAL_SERVER_PORT,
eval_server_host: str = EVAL_SERVER_HOST,
**kwargs,
):
super().__init__(**kwargs)
@@ -70,6 +72,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
self.skip_embeddings = skip_embeddings
self.gitnexus_timeout = gitnexus_timeout
self.eval_server_port = eval_server_port
self.eval_server_host = eval_server_host
self.index_time: float = 0.0
self._gitnexus_ready = False
@@ -165,22 +168,29 @@ class GitNexusDockerEnvironment(DockerEnvironment):
def _start_eval_server(self):
"""Start the GitNexus eval-server daemon in the background."""
logger.info(f"Starting eval-server on port {self.eval_server_port}...")
logger.info(
f"Starting eval-server on {self.eval_server_host}:{self.eval_server_port}..."
)
self.execute({
"command": (
f"nohup npx gitnexus eval-server --port {self.eval_server_port} "
f"--host {self.eval_server_host} "
f"--idle-timeout 600 "
f"> /tmp/gitnexus-eval-server.log 2>&1 &"
),
"timeout": 5,
})
# Use 127.0.0.1 for the health probe — reachable whether server binds
# loopback or all interfaces (0.0.0.0), avoiding DNS resolution issues.
health_host = "127.0.0.1"
# Wait for the server to be ready (up to ~15s for KuzuDB init)
for i in range(EVAL_SERVER_HEALTH_RETRIES):
time.sleep(EVAL_SERVER_HEALTH_INTERVAL_SECONDS)
health = self.execute({
"command": f"curl -sf http://127.0.0.1:{self.eval_server_port}/health 2>/dev/null || echo 'NOT_READY'",
"command": f"curl -sf http://{health_host}:{self.eval_server_port}/health 2>/dev/null || echo 'NOT_READY'",
"timeout": EVAL_SERVER_HEALTH_TIMEOUT_SECONDS,
})
output = health.get("output", "").strip()
@@ -201,7 +211,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
)
@staticmethod
def _render_tool_script(spec: ToolScriptSpec, port: str) -> str:
def _render_tool_script(spec: ToolScriptSpec, port: str, host: str = EVAL_SERVER_HOST) -> str:
"""
Render a standalone bash script for a GitNexus tool.
@@ -212,6 +222,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
if spec.endpoint:
lines.append(f'PORT="${{GITNEXUS_EVAL_PORT:-{port}}}"')
lines.append(f'HOST="${{GITNEXUS_EVAL_HOST:-{host}}}"')
if spec.header:
lines.append(spec.header.strip())
@@ -221,7 +232,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
if spec.endpoint:
lines.append(
f'result=$(curl -sf -X POST "http://127.0.0.1:${{PORT}}{spec.endpoint}" '
f'result=$(curl -sf -X POST "http://${{HOST}}:${{PORT}}{spec.endpoint}" '
'-H "Content-Type: application/json" -d "$payload" 2>/dev/null)'
)
lines.append('if [ $? -eq 0 ] && [ -n "$result" ]; then echo "$result"; exit 0; fi')
@@ -244,9 +255,10 @@ class GitNexusDockerEnvironment(DockerEnvironment):
Uses heredocs with quoted delimiter to avoid all quoting/escaping issues.
"""
port = str(self.eval_server_port)
host = self.eval_server_host
for spec in TOOL_SPECS.values():
script_content = self._render_tool_script(spec, port).strip()
script_content = self._render_tool_script(spec, port, host).strip()
# Use heredoc with quoted delimiter — prevents all variable expansion and quoting issues
self.execute({
"command": (
@@ -387,5 +399,6 @@ class GitNexusDockerEnvironment(DockerEnvironment):
"index_time_seconds": round(self.index_time, 2),
"skip_embeddings": self.skip_embeddings,
"eval_server_port": self.eval_server_port,
"eval_server_host": self.eval_server_host,
}
return base
Generated
+6 -6
View File
@@ -760,11 +760,11 @@ wheels = [
[[package]]
name = "idna"
version = "3.11"
version = "3.15"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/6f/6d/0703ccc57f3a7233505399edb88de3cbd678da106337b9fcde432b65ed60/idna-3.11.tar.gz", hash = "sha256:795dafcc9c04ed0c1fb032c2aa73654d8e8c5023a7df64a53f39190ada629902", size = 194582, upload-time = "2025-10-12T14:55:20.501Z" }
sdist = { url = "https://files.pythonhosted.org/packages/82/77/7b3966d0b9d1d31a36ddf1746926a11dface89a83409bf1483f0237aa758/idna-3.15.tar.gz", hash = "sha256:ca962446ea538f7092a95e057da437618e886f4d349216d2b1e294abfdb65fdc", size = 199245, upload-time = "2026-05-12T22:45:57.011Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/0e/61/66938bbb5fc52dbdf84594873d5b51fb1f7c7794e9c0f5bd885f30bc507b/idna-3.11-py3-none-any.whl", hash = "sha256:771a87f49d9defaf64091e6e6fe9c18d4833f140bd19464795bc32d966ca37ea", size = 71008, upload-time = "2025-10-12T14:55:18.883Z" },
{ url = "https://files.pythonhosted.org/packages/d2/23/408243171aa9aaba178d3e2559159c24c1171a641aa83b67bdd3394ead8e/idna-3.15-py3-none-any.whl", hash = "sha256:048adeaf8c2d788c40fee287673ccaa74c24ffd8dcf09ffa555a2fbb59f10ac8", size = 72340, upload-time = "2026-05-12T22:45:55.733Z" },
]
[[package]]
@@ -2278,11 +2278,11 @@ wheels = [
[[package]]
name = "urllib3"
version = "2.6.3"
version = "2.7.0"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/c7/24/5f1b3bdffd70275f6661c76461e25f024d5a38a46f04aaca912426a2b1d3/urllib3-2.6.3.tar.gz", hash = "sha256:1b62b6884944a57dbe321509ab94fd4d3b307075e0c2eae991ac71ee15ad38ed", size = 435556, upload-time = "2026-01-07T16:24:43.925Z" }
sdist = { url = "https://files.pythonhosted.org/packages/53/0c/06f8b233b8fd13b9e5ee11424ef85419ba0d8ba0b3138bf360be2ff56953/urllib3-2.7.0.tar.gz", hash = "sha256:231e0ec3b63ceb14667c67be60f2f2c40a518cb38b03af60abc813da26505f4c", size = 433602, upload-time = "2026-05-07T16:13:18.596Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/39/08/aaaad47bc4e9dc8c725e68f9d04865dbcb2052843ff09c97b08904852d84/urllib3-2.6.3-py3-none-any.whl", hash = "sha256:bf272323e553dfb2e87d9bfd225ca7b0f467b919d7bbd355436d3fd37cb0acd4", size = 131584, upload-time = "2026-01-07T16:24:42.685Z" },
{ url = "https://files.pythonhosted.org/packages/7f/3e/5db95bcf282c52709639744ca2a8b149baccf648e39c8cc87553df9eae0c/urllib3-2.7.0-py3-none-any.whl", hash = "sha256:9fb4c81ebbb1ce9531cce37674bbc6f1360472bc18ca9a553ede278ef7276897", size = 131087, upload-time = "2026-05-07T16:13:17.151Z" },
]
[[package]]
+47 -4
View File
@@ -14,6 +14,8 @@
const fs = require('fs');
const path = require('path');
const { spawnSync } = require('child_process');
const { acquireHookSlot } = require('./hook-lock.js');
const { hasGitNexusDbLockedByGitNexusServer } = require('./hook-db-lock-probe.cjs');
/**
* Read JSON input from stdin synchronously.
@@ -102,6 +104,28 @@ function findGitNexusDir(startDir) {
return null;
}
function hasGitNexusServerOwner(gitNexusDir) {
return hasGitNexusDbLockedByGitNexusServer(path.join(gitNexusDir, 'lbug'), process.pid);
}
function extractAugmentContext(stderr) {
const output = (stderr || '').trim();
const marker = output.indexOf('[GitNexus]');
const debug = process.env.GITNEXUS_DEBUG === '1' || process.env.GITNEXUS_DEBUG === 'true';
if (debug && output.length > 0) {
// Emit the FULL discarded prefix (everything before the marker, or all of
// it when no marker is present) so suppressed diagnostics — LadybugDB lock
// warnings, parser errors, etc. — remain recoverable on the hook's own
// stderr. The untruncated payload lets operators see exactly what was
// filtered out instead of a 180-char JSON-quoted preview.
const discarded = marker === -1 ? output : output.slice(0, marker).trim();
if (discarded.length > 0) {
process.stderr.write(`[GitNexus hook] augment stderr discarded prefix:\n${discarded}\n`);
}
}
return marker === -1 ? '' : output.slice(marker).trim();
}
/**
* Extract search pattern from tool input.
*/
@@ -169,6 +193,15 @@ function extractPattern(toolName, toolInput) {
*/
function runGitNexusCli(args, cwd, timeout) {
const isWin = process.platform === 'win32';
const hookCli = process.env.GITNEXUS_HOOK_CLI_PATH;
if (hookCli !== undefined && String(hookCli).trim() && fs.existsSync(String(hookCli))) {
return spawnSync(process.execPath, [String(hookCli), ...args], {
encoding: 'utf-8',
timeout,
cwd,
stdio: ['pipe', 'pipe', 'pipe'],
});
}
// Detect whether 'gitnexus' is on PATH (cheap check, no execution)
let useDirectBinary = false;
@@ -217,7 +250,8 @@ function sendHookResponse(hookEventName, message) {
function handlePreToolUse(input) {
const cwd = input.cwd || process.cwd();
if (!path.isAbsolute(cwd)) return;
if (!findGitNexusDir(cwd)) return;
const gitNexusDir = findGitNexusDir(cwd);
if (!gitNexusDir) return;
const toolName = input.tool_name || '';
const toolInput = input.tool_input || {};
@@ -226,19 +260,28 @@ function handlePreToolUse(input) {
const pattern = extractPattern(toolName, toolInput);
if (!pattern || pattern.length < 3) return;
if (hasGitNexusServerOwner(gitNexusDir)) {
process.stderr.write('[GitNexus] augment skipped: MCP server owns DB\n');
return;
}
const release = acquireHookSlot(gitNexusDir);
if (!release) return;
let result = '';
try {
const child = runGitNexusCli(['augment', '--', pattern], cwd, 7000);
if (!child.error && child.status === 0) {
result = child.stderr || '';
result = extractAugmentContext(child.stderr || '');
}
} catch {
/* graceful failure */
} finally {
release();
}
if (result && result.trim()) {
sendHookResponse('PreToolUse', result.trim());
if (result) {
sendHookResponse('PreToolUse', result);
}
}
@@ -0,0 +1,238 @@
/**
* Cross-platform best-effort probe: does another process hold dbPath open
* with a command line that looks like a GitNexus MCP/serve server?
*
* Backends (no user-installed Sysinternals):
* - Linux: scan procfs under /proc (per-PID fd entries) via stat(2) (dev+inode); works without lsof;
* optional lsof fallback when proc scan finds nothing.
* - macOS / *BSD / etc.: trusted lsof + ps (absolute paths first).
* - Windows: Restart Manager (rstrtmgr) via bundled PowerShell script +
* Win32_Process for command lines; trusted powershell.exe under %SystemRoot%.
*
* Fail-open on most errors; fail-closed only on lsof ETIMEDOUT (Unix) or
* PowerShell ETIMEDOUT (Windows), matching the hook contract.
*/
const fs = require('fs');
const path = require('path');
const { spawnSync } = require('child_process');
function isGitNexusServerCommand(command) {
const hasServerMode = /(?:^|\s)(mcp|serve)(?:\s|$)/.test(command);
const hasGitNexus =
/(?:^|[/\\\s])gitnexus(?:\.cmd)?(?:\s|$)/.test(command) ||
/node_modules[/\\]gitnexus[/\\]/.test(command);
return hasServerMode && hasGitNexus;
}
function resolveHookBinary(tool) {
const envKey = tool === 'lsof' ? 'GITNEXUS_HOOK_LSOF_PATH' : 'GITNEXUS_HOOK_PS_PATH';
const fromEnv = process.env[envKey];
if (fromEnv && String(fromEnv).trim() && fs.existsSync(String(fromEnv))) {
return String(fromEnv);
}
const candidates =
tool === 'lsof'
? ['/usr/bin/lsof', '/usr/sbin/lsof', '/sbin/lsof', tool]
: ['/bin/ps', '/usr/bin/ps', tool];
for (const candidate of candidates) {
if (candidate === tool) return tool;
try {
if (fs.existsSync(candidate)) return candidate;
} catch {
/* ignore */
}
}
return tool;
}
function resolveWindowsPowerShellPath() {
const fromEnv = process.env.GITNEXUS_HOOK_POWERSHELL_PATH;
if (fromEnv && String(fromEnv).trim() && fs.existsSync(String(fromEnv).trim())) {
return String(fromEnv).trim();
}
const root = process.env.SystemRoot || 'C:\\Windows';
const ps = path.join(root, 'System32', 'WindowsPowerShell', 'v1.0', 'powershell.exe');
if (fs.existsSync(ps)) return ps;
const psWow = path.join(root, 'SysWOW64', 'WindowsPowerShell', 'v1.0', 'powershell.exe');
if (fs.existsSync(psWow)) return psWow;
return 'powershell.exe';
}
// Sentinel:
// undefined = not loaded yet (try the read)
// string = encoded PowerShell command (successful load)
// null = load attempted and failed (do not retry; warning already emitted)
let windowsRmListPsEncodedCommandCache;
let windowsRmListPsLoadFailureWarned = false;
function getWindowsRmListEncodedCommand() {
if (windowsRmListPsEncodedCommandCache !== undefined) {
return windowsRmListPsEncodedCommandCache;
}
try {
const ps1Path = path.join(__dirname, 'win-rm-list-json.ps1');
const src = fs
.readFileSync(ps1Path, 'utf8')
.replace(/^\uFEFF/, '')
.replace(/\r\n/g, '\n');
windowsRmListPsEncodedCommandCache = Buffer.from(src, 'utf16le').toString('base64');
} catch (err) {
windowsRmListPsEncodedCommandCache = null;
if (
!windowsRmListPsLoadFailureWarned &&
(process.env.GITNEXUS_DEBUG === '1' || process.env.GITNEXUS_DEBUG === 'true')
) {
windowsRmListPsLoadFailureWarned = true;
const msg = err && err.message ? String(err.message).slice(0, 200) : 'unknown';
process.stderr.write(`[GitNexus hook] win-rm-list-json.ps1 load failed: ${msg}\n`);
}
}
return windowsRmListPsEncodedCommandCache;
}
function hasGitNexusServerOwnerWindows(dbPathAbs, myPid) {
const encoded = getWindowsRmListEncodedCommand();
if (!encoded) return false;
const psExe = resolveWindowsPowerShellPath();
const r = spawnSync(
psExe,
[
'-NoProfile',
'-NonInteractive',
'-ExecutionPolicy',
'Bypass',
'-STA',
'-EncodedCommand',
encoded,
],
{
encoding: 'utf-8',
timeout: 6000,
stdio: ['ignore', 'pipe', 'ignore'],
env: { ...process.env, GITNEXUS_HOOK_RM_TARGET: dbPathAbs },
},
);
// ETIMEDOUT means the PowerShell probe didn't return in time; treat as 'unresponsive process holds DB' → fail-closed (skip augment).
if (r.error) return r.error.code === 'ETIMEDOUT';
if (r.status !== 0) return false;
let rows;
try {
rows = JSON.parse(String(r.stdout || '').trim() || '[]');
} catch {
return false;
}
if (!Array.isArray(rows)) return false;
for (const row of rows) {
const procId = Number(row.pid);
const cmd = String(row.cmd || '');
if (!Number.isFinite(procId) || procId === myPid) continue;
if (isGitNexusServerCommand(cmd)) return true;
}
return false;
}
function readLinuxCmdline(pidStr) {
try {
return fs.readFileSync(`/proc/${pidStr}/cmdline`, 'utf8').replace(/\0+/g, ' ').trim();
} catch {
return '';
}
}
function linuxProcScanFindGitNexusServer(dbPathAbs, myPid) {
const raw = process.env.GITNEXUS_HOOK_LINUX_PROC_BUDGET_MS;
const budget = Number(raw && String(raw).trim()) ? Number.parseInt(String(raw), 10) : 1200;
const start = Date.now();
let targetStat;
try {
targetStat = fs.statSync(dbPathAbs);
} catch {
return false;
}
let procEntries;
try {
procEntries = fs.readdirSync('/proc', { withFileTypes: true });
} catch {
return false;
}
for (const ent of procEntries) {
if (Date.now() - start > budget) return false;
if (!ent.isDirectory() || !/^\d+$/.test(ent.name)) continue;
const pid = Number.parseInt(ent.name, 10);
if (!Number.isFinite(pid) || pid === myPid) continue;
const fdDir = path.join('/proc', ent.name, 'fd');
let fds;
try {
fds = fs.readdirSync(fdDir);
} catch {
continue;
}
let holds = false;
for (const fd of fds) {
if (Date.now() - start > budget) return false;
try {
const st = fs.statSync(path.join(fdDir, fd));
if (st.dev === targetStat.dev && st.ino === targetStat.ino) {
holds = true;
break;
}
} catch {
/* ignore */
}
}
if (!holds) continue;
if (isGitNexusServerCommand(readLinuxCmdline(ent.name))) return true;
}
return false;
}
function unixLsofPsFindGitNexusServer(dbPathAbs, myPid) {
const lsofPath = resolveHookBinary('lsof');
const lsof = spawnSync(lsofPath, ['-nP', '-t', '--', dbPathAbs], {
encoding: 'utf-8',
timeout: 1000,
stdio: ['ignore', 'pipe', 'ignore'],
});
if (lsof.error) return lsof.error.code === 'ETIMEDOUT';
const pids = (lsof.stdout || '').split(/\s+/).filter(Boolean);
const psPath = resolveHookBinary('ps');
for (const pid of pids) {
if (Number(pid) === myPid) continue;
const ps = spawnSync(psPath, ['-p', pid, '-o', 'command='], {
encoding: 'utf-8',
timeout: 500,
stdio: ['ignore', 'pipe', 'ignore'],
});
if (ps.error) {
if (ps.error.code === 'ETIMEDOUT') return true;
continue;
}
if (isGitNexusServerCommand(ps.stdout || '')) return true;
}
return false;
}
/**
* @param {string} dbPath Absolute or relative path to the DB file (e.g. .../lbug).
* @param {number} myPid Current process PID (hook runner), excluded from matches.
*/
function hasGitNexusDbLockedByGitNexusServer(dbPath, myPid) {
if (!fs.existsSync(dbPath)) return false;
const dbPathAbs = path.resolve(dbPath);
if (process.platform === 'win32') {
return hasGitNexusServerOwnerWindows(dbPathAbs, myPid);
}
if (process.platform === 'linux') {
if (linuxProcScanFindGitNexusServer(dbPathAbs, myPid)) return true;
return unixLsofPsFindGitNexusServer(dbPathAbs, myPid);
}
return unixLsofPsFindGitNexusServer(dbPathAbs, myPid);
}
module.exports = {
hasGitNexusDbLockedByGitNexusServer,
};
+119
View File
@@ -0,0 +1,119 @@
const fs = require('fs');
const path = require('path');
const HOOK_LOCK_SUBDIR = '.hook-locks';
const HOOK_LOCK_MAX_INFLIGHT = 3;
const HOOK_LOCK_STALE_MS = 30000;
function acquireHookSlot(gitNexusDir) {
const lockDir = path.join(gitNexusDir, HOOK_LOCK_SUBDIR);
try {
fs.mkdirSync(lockDir, { recursive: true });
} catch {
// Cannot create lock dir (read-only fs, cross-user perm denial, out of
// inodes, etc.) — fail closed by returning null. Caller skips augment.
// Fail-open here would let N concurrent hooks all proceed unguarded and
// reintroduce the #1486 fan-out the guard exists to prevent.
return null;
}
const myPidStr = String(process.pid);
for (let slot = 0; slot < HOOK_LOCK_MAX_INFLIGHT; slot++) {
const slotPath = path.join(lockDir, `slot-${slot}.lock`);
for (let attempt = 0; attempt < 2; attempt++) {
try {
fs.writeFileSync(slotPath, myPidStr, { flag: 'wx' });
let released = false;
const release = () => {
if (released) return;
released = true;
try {
// Only unlink if we still own the slot. If we appeared stale and
// another hook took over, the file now belongs to it — leave alone.
const content = fs.readFileSync(slotPath, 'utf-8').trim();
if (content === myPidStr) fs.unlinkSync(slotPath);
} catch {
/* already removed or unreadable */
}
};
process.on('exit', release);
return release;
} catch {
// Slot exists. Decide whether to take it over.
// Open once and inspect mtime + content via the same fd so there's
// no TOCTOU between the metadata check and the content read
// (codeql js/file-system-race).
let fd;
try {
fd = fs.openSync(slotPath, 'r');
} catch {
continue; // Vanished between EEXIST and open — retry this slot.
}
let isLive = false;
let mtimeMs = Date.now();
try {
mtimeMs = fs.fstatSync(fd).mtimeMs;
const buf = Buffer.alloc(32);
const n = fs.readSync(fd, buf, 0, 32, 0);
const ownerStr = buf.slice(0, n).toString('utf-8').trim();
if (ownerStr === '') {
// Owner created the file but hasn't written its PID yet. The
// wx open+write window is microseconds; give it the benefit
// of the doubt and treat as live.
isLive = true;
} else {
const owner = Number.parseInt(ownerStr, 10);
if (Number.isFinite(owner) && owner > 0) {
try {
process.kill(owner, 0);
isLive = true;
} catch (e) {
// ESRCH = process gone → treat as dead. EPERM = process exists
// but owned by another user (cross-user lock dir) → still alive,
// keep the slot. Anything else: be conservative, assume alive.
if (e && e.code === 'ESRCH') {
isLive = false;
} else {
isLive = true;
}
}
}
}
} catch {
/* unreadable — treat as dead */
} finally {
try {
fs.closeSync(fd);
} catch {
/* already closed */
}
}
// For slots younger than HOOK_LOCK_STALE_MS, PID-liveness wins —
// a slow-but-alive hook is never wrongly evicted. For older slots,
// age is the final arbiter as a defense against PID reuse on long-
// abandoned slots. 30s >> the 7s augment timeout, so a healthy run
// never crosses this threshold.
if (isLive && Date.now() - mtimeMs > HOOK_LOCK_STALE_MS) {
isLive = false;
}
if (isLive) break; // Try the next slot.
try {
fs.unlinkSync(slotPath);
} catch {
/* another hook beat us to it — retry will hit EEXIST */
}
// Loop and retry this slot.
}
}
}
return null;
}
module.exports = {
HOOK_LOCK_SUBDIR,
HOOK_LOCK_MAX_INFLIGHT,
HOOK_LOCK_STALE_MS,
acquireHookSlot,
};
@@ -0,0 +1,76 @@
$ErrorActionPreference = 'Stop'
$target = $env:GITNEXUS_HOOK_RM_TARGET
if ([string]::IsNullOrWhiteSpace($target)) { Write-Output '[]'; exit 0 }
$target = (Resolve-Path -LiteralPath $target).ProviderPath
if (-not ([Management.Automation.PSTypeName]'GitNexusHookRm.Native').Type) {
Add-Type @'
using System;
using System.Runtime.InteropServices;
namespace GitNexusHookRm {
public static class Native {
public const int ErrorMoreData = 234;
[StructLayout(LayoutKind.Sequential, Pack = 4)]
public struct RM_UNIQUE_PROCESS {
public int dwProcessId;
public long ProcessStartTime;
}
[StructLayout(LayoutKind.Sequential, CharSet = CharSet.Unicode)]
public struct RM_PROCESS_INFO {
public RM_UNIQUE_PROCESS Process;
[MarshalAs(UnmanagedType.ByValTStr, SizeConst = 256)]
public string strAppName;
[MarshalAs(UnmanagedType.ByValTStr, SizeConst = 64)]
public string strServiceShortName;
public uint ApplicationType;
public uint AppStatus;
public uint TSSessionId;
public uint bRestartable;
}
[DllImport("rstrtmgr.dll", CharSet = CharSet.Unicode)]
public static extern int RmStartSession(out uint pSessionHandle, uint dwSessionFlags, string strSessionKey);
[DllImport("rstrtmgr.dll", CharSet = CharSet.Unicode)]
public static extern int RmRegisterResources(uint pSessionHandle, uint nFiles, string[] rgsFileNames, uint nApplications, IntPtr rgApplications, uint nServices, string[] rgsServiceNames);
[DllImport("rstrtmgr.dll")]
public static extern int RmGetList(uint dwSessionHandle, out uint pnProcInfoNeeded, ref uint pnProcInfo, [In, Out] RM_PROCESS_INFO[] rgAffectedApps, ref uint lpdwRebootReasons);
[DllImport("rstrtmgr.dll")]
public static extern int RmEndSession(uint pSessionHandle);
}
}
'@
}
$h = [uint32]0
$key = [guid]::NewGuid().ToString('N')
$rmErr = [GitNexusHookRm.Native]::RmStartSession([ref]$h, 0, $key)
if ($rmErr -ne 0) { Write-Output '[]'; exit 0 }
$files = @($target)
$err = [GitNexusHookRm.Native]::RmRegisterResources($h, 1, $files, 0, [IntPtr]::Zero, 0, $null)
if ($err -ne 0) {
[void][GitNexusHookRm.Native]::RmEndSession($h)
Write-Output '[]'
exit 0
}
$need = [uint32]0
$n = [uint32]0
$reboot = [uint32]0
$err = [GitNexusHookRm.Native]::RmGetList($h, [ref]$need, [ref]$n, $null, [ref]$reboot)
if ($err -ne [GitNexusHookRm.Native]::ErrorMoreData) {
[void][GitNexusHookRm.Native]::RmEndSession($h)
Write-Output '[]'
exit 0
}
$n = $need
$buf = New-Object GitNexusHookRm.Native+RM_PROCESS_INFO[] ([int]$n)
$err = [GitNexusHookRm.Native]::RmGetList($h, [ref]$need, [ref]$n, $buf, [ref]$reboot)
[void][GitNexusHookRm.Native]::RmEndSession($h)
if ($err -ne 0) { Write-Output '[]'; exit 0 }
$out = @()
for ($i = 0; $i -lt [int]$n; $i++) {
$procId = $buf[$i].Process.dwProcessId
$p = Get-CimInstance -ClassName Win32_Process -Filter "ProcessId=$procId" -ErrorAction SilentlyContinue
$cmd = if ($p) { $p.CommandLine } else { '' }
$out += [PSCustomObject]@{ pid = [int]$procId; cmd = $cmd }
}
ConvertTo-Json -InputObject @($out) -Compress
@@ -56,13 +56,15 @@ Generates repository documentation from the knowledge graph using an LLM. Requir
| Flag | Effect |
|------|--------|
| `--force` | Force full regeneration |
| `--force` | Force full regeneration, also required to re-gerenate an existing wiki in a different language |
| `--model <model>` | LLM model (default: minimax/minimax-m2.5) |
| `--base-url <url>` | LLM API base URL |
| `--api-key <key>` | LLM API key |
| `--concurrency <n>` | Parallel LLM calls (default: 3) |
| `--gist` | Publish wiki as a public GitHub Gist |
| `--timeout <seconds>` | LLM request timeout in seconds (default: disabled) |
| `--retries <n>` | Max LLM retry attempts per request (default: 3) |
| `--lang <lang>` | Output language for generated documentation (e.g. english, chinese, spanish, japanese)|
### list — Show all indexed repos
```bash
+91
View File
@@ -0,0 +1,91 @@
# GitNexus — Cursor integration
Static config that adds GitNexus knowledge-graph augmentation and skill files to Cursor.
> **Hooks require Cursor 2.4+.** Earlier versions don't expose `postToolUse` and the hook will silently no-op.
## What you get
| Layer | What it does | How it's installed |
| ------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------- |
| **MCP** | `gitnexus` MCP server with 16 tools (`query`, `context`, `impact`, `detect_changes`, `rename`, …) | `npx gitnexus setup` writes `~/.cursor/mcp.json` automatically. |
| **Skills** | `/gitnexus-exploring`, `/gitnexus-debugging`, `/gitnexus-impact-analysis`, `/gitnexus-refactoring`, `/gitnexus-pr-review` markdown skills | `npx gitnexus setup` copies them to `~/.cursor/skills/gitnexus/`. |
| **Hooks** _(this README)_ | `postToolUse` hook that enriches `Shell` / `Read` / `Grep` tool calls with graph context — same augmentation Claude Code gets | **Manual** — copy the files described below into your project's `.cursor/`. |
## Hook install
Cursor 2.4+ reads `.cursor/hooks.json` from the project root and runs hook commands with the project root as the working directory ([docs](https://cursor.com/docs/agent/hooks)).
From this repo's `gitnexus-cursor-integration/hooks/`, copy the files below into your **project root**:
```text
<your-project>/
├── .cursor/
│ └── hooks.json ← from gitnexus-cursor-integration/hooks/hooks.json
└── hooks/
├── gitnexus-hook.cjs ← from gitnexus-cursor-integration/hooks/gitnexus-hook.cjs
└── hook-lock.cjs ← from gitnexus-cursor-integration/hooks/hook-lock.cjs
```
Equivalent shell commands (run from your project root, with `$GITNEXUS_REPO` pointing at a clone of this repo):
```bash
mkdir -p .cursor hooks
cp "$GITNEXUS_REPO/gitnexus-cursor-integration/hooks/hooks.json" .cursor/hooks.json
cp "$GITNEXUS_REPO/gitnexus-cursor-integration/hooks/gitnexus-hook.cjs" hooks/gitnexus-hook.cjs
cp "$GITNEXUS_REPO/gitnexus-cursor-integration/hooks/hook-lock.cjs" hooks/hook-lock.cjs
```
If you already have a `.cursor/hooks.json`, merge the `hooks.postToolUse` array rather than overwriting.
### Verify
1. Index the project: `npx gitnexus analyze`
2. Reload the Cursor window so it picks up the new hook config.
3. Ask the agent something that triggers `Read` / `Grep` / `Shell rg`. You should see a `[GitNexus]` block appended to the tool result.
4. Diagnose silent no-ops by setting `GITNEXUS_DEBUG=1` in your shell environment — the hook will write Cursor's raw event payload to stderr so you can verify field names.
### What's installed manually vs. automated
| Step | Automated by `gitnexus setup`? |
| -------------------------------------------------------------------- | ------------------------------ |
| `~/.cursor/mcp.json` | ✅ |
| `~/.cursor/skills/gitnexus/*` | ✅ |
| `<project>/.cursor/hooks.json` + `<project>/hooks/gitnexus-hook.cjs` + `<project>/hooks/hook-lock.cjs` | ❌ — copy manually (see above) |
Hook install is per-project (Cursor scopes hooks to a project root); skills and MCP config are global.
## Hook contract
The hook receives a JSON event on stdin matching Cursor 2.4's `postToolUse` shape:
```json
{
"tool_name": "Grep" | "Read" | "Shell",
"tool_input": { /* tool-specific */ },
"tool_output": { /* optional */ },
"cwd": "/absolute/path/to/project"
}
```
It writes augmentation context to stdout as:
```json
{ "additional_context": "[GitNexus] …" }
```
Empty stdout means "no augmentation, continue normally" — the hook never blocks the tool.
### Pattern extraction per tool
| Tool | Pattern source | Notes |
| ------- | ---------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `Grep` | `tool_input.query` (also `pattern`, `regex`, `q`, `search`, `searchQuery`) | Last-resort fallback: longest string value in `tool_input` (≥ 3 chars). |
| `Read` | basename of `tool_input.target_file` (also `file_path`, `filePath`, `path`, `file`), stripped to identifier characters | `auth/handler.ts` → `handler`. |
| `Shell` | First positional argument after `rg` / `grep` in `tool_input.command` | Best-effort tokenizer; quoted multi-word patterns (`rg "User Service"`) extract the first word only. |
## Troubleshooting
- **Nothing happens** — Confirm Cursor is on 2.4+ and the project root has `.cursor/hooks.json` plus both hook files at `hooks/gitnexus-hook.cjs` and `hooks/hook-lock.cjs`. Then `npx gitnexus list` to confirm the project is indexed.
- **`gitnexus` not found** — The hook prefers a locally-resolvable `gitnexus/dist/cli/index.js` and falls back to `npx -y gitnexus`. Install globally with `npm i -g gitnexus` to skip the npx cold-start latency.
- **Wrong pattern extracted** — Set `GITNEXUS_DEBUG=1` and run a tool call. The raw stdin payload is logged to stderr; use it to confirm Cursor's actual `tool_input` field names against the table above. If they differ, file an issue with the captured payload.
@@ -1,50 +0,0 @@
#!/bin/bash
# GitNexus beforeShellExecution hook for Cursor
# Receives JSON on stdin with { command, cwd, timeout }
# Returns JSON on stdout with { permission, agent_message }
#
# Extracts search pattern from grep/rg commands, runs gitnexus augment,
# and injects the enriched context via agent_message.
INPUT=$(cat)
COMMAND=$(echo "$INPUT" | jq -r '.command // empty' 2>/dev/null)
if [ -z "$COMMAND" ]; then
echo '{"permission":"allow"}'
exit 0
fi
# Skip non-search commands
case "$COMMAND" in
cd\ *|npm\ *|yarn\ *|pnpm\ *|git\ commit*|git\ push*|git\ pull*|mkdir\ *|rm\ *|cp\ *|mv\ *|echo\ *|cat\ *)
echo '{"permission":"allow"}'
exit 0
;;
esac
# Extract search pattern from rg/grep commands
PATTERN=""
if echo "$COMMAND" | grep -qE '\brg\b'; then
PATTERN=$(echo "$COMMAND" | sed -n "s/.*\brg\s\+\(--[^ ]*\s\+\)*['\"]\\?\([^'\";\| >]*\\).*/\2/p")
elif echo "$COMMAND" | grep -qE '\bgrep\b'; then
PATTERN=$(echo "$COMMAND" | sed -n "s/.*\bgrep\s\+\(-[^ ]*\s\+\)*['\"]\\?\([^'\";\| >]*\\).*/\2/p")
fi
if [ -z "$PATTERN" ] || [ ${#PATTERN} -lt 3 ]; then
echo '{"permission":"allow"}'
exit 0
fi
# Run gitnexus augment
RESULT=$(npx -y gitnexus augment "$PATTERN" 2>/dev/null)
if [ -n "$RESULT" ]; then
# Escape for JSON
ESCAPED=$(echo "$RESULT" | jq -Rs .)
echo "{\"permission\":\"allow\",\"agent_message\":$ESCAPED}"
else
echo '{"permission":"allow"}'
fi
exit 0
@@ -0,0 +1,266 @@
#!/usr/bin/env node
/**
* GitNexus Cursor postToolUse Hook
*
* Receives a JSON event on stdin describing a finished tool call, derives a
* search pattern (Grep query, Read file basename, or rg/grep arg from a Shell
* command), runs `gitnexus augment <pattern>`, and emits the enriched context
* back as `{ additional_context: "..." }` so the agent sees it alongside the
* tool result.
*
* Replaces the legacy beforeShellExecution / augment-shell.sh pipeline:
* - Cross-platform (no bash, no jq — runs on Windows out of the box)
* - Covers Read and Grep, not just Shell rg/grep
*
* Cursor 2.4+ generic hooks: https://cursor.com/docs/agent/hooks
*/
const fs = require('fs');
const path = require('path');
const { spawnSync } = require('child_process');
const { acquireHookSlot } = require('./hook-lock.cjs');
function readInput() {
try {
const data = fs.readFileSync(0, 'utf-8');
return JSON.parse(data);
} catch {
return {};
}
}
function isGlobalRegistryDir(candidate) {
if (fs.existsSync(path.join(candidate, 'meta.json'))) return false;
return (
fs.existsSync(path.join(candidate, 'registry.json')) ||
fs.existsSync(path.join(candidate, 'repos'))
);
}
function walkForGitNexusDir(startDir) {
let dir = startDir;
for (let i = 0; i < 5; i++) {
const candidate = path.join(dir, '.gitnexus');
if (fs.existsSync(candidate)) {
if (!isGlobalRegistryDir(candidate)) return candidate;
}
const parent = path.dirname(dir);
if (parent === dir) break;
dir = parent;
}
return null;
}
function findCanonicalRepoRoot(cwd) {
try {
const result = spawnSync('git', ['rev-parse', '--path-format=absolute', '--git-common-dir'], {
encoding: 'utf-8',
timeout: 2000,
cwd,
stdio: ['pipe', 'pipe', 'pipe'],
});
if (result.error || result.status !== 0) return null;
const commonDir = (result.stdout || '').trim();
if (!commonDir || !path.isAbsolute(commonDir)) return null;
return path.dirname(commonDir);
} catch {
return null;
}
}
function findGitNexusDir(startDir) {
const cwd = startDir || process.cwd();
const fromCwd = walkForGitNexusDir(cwd);
if (fromCwd) return fromCwd;
const canonicalRoot = findCanonicalRepoRoot(cwd);
if (canonicalRoot && canonicalRoot !== cwd) {
return walkForGitNexusDir(canonicalRoot);
}
return null;
}
function parseRgGrepPattern(cmd) {
const tokens = cmd.split(/\s+/);
let foundCmd = false;
let skipNext = false;
const flagsWithValues = new Set([
'-e',
'-f',
'-m',
'-A',
'-B',
'-C',
'-g',
'--glob',
'-t',
'--type',
'--include',
'--exclude',
]);
for (const token of tokens) {
if (skipNext) {
skipNext = false;
continue;
}
if (!foundCmd) {
if (/\brg$|\bgrep$/.test(token)) foundCmd = true;
continue;
}
if (token.startsWith('-')) {
if (flagsWithValues.has(token)) skipNext = true;
continue;
}
const cleaned = token.replace(/['"]/g, '');
return cleaned.length >= 3 ? cleaned : null;
}
return null;
}
/**
* Extract a search pattern from the tool input. Cursor 2.4 docs at
* https://cursor.com/docs/agent/hooks list the tool *matchers* but do not
* formally specify the per-tool tool_input field names, so we probe a
* generous set of MCP-style aliases. As a last-resort fallback for Grep
* (the highest-frequency search path) we also accept the longest plausible
* string value in tool_input. Set GITNEXUS_DEBUG=1 to log the raw payload
* to stderr if Cursor changes the contract and aliases stop matching.
*/
function pickLongestStringValue(obj) {
let best = null;
if (!obj || typeof obj !== 'object') return null;
for (const v of Object.values(obj)) {
if (typeof v === 'string' && v.length >= 3 && (!best || v.length > best.length)) {
best = v;
}
}
return best;
}
function extractPattern(toolName, toolInput) {
const t = (toolName || '').toLowerCase();
if (t === 'grep') {
const aliases = [
toolInput.query,
toolInput.pattern,
toolInput.regex,
toolInput.q,
toolInput.search,
toolInput.searchQuery,
];
for (const a of aliases) {
if (typeof a === 'string' && a.length >= 3) return a;
}
// Last resort: scan tool_input for any reasonable-looking string value.
return pickLongestStringValue(toolInput);
}
if (t === 'read') {
const filePath =
toolInput.target_file ||
toolInput.file_path ||
toolInput.filePath ||
toolInput.path ||
toolInput.file ||
'';
if (!filePath) return null;
const base = path.basename(String(filePath), path.extname(String(filePath)));
const cleaned = base.replace(/[^a-zA-Z0-9_]/g, '');
return cleaned.length >= 3 ? cleaned : null;
}
if (t === 'shell') {
const cmd = toolInput.command || '';
if (!/\brg\b|\bgrep\b/.test(cmd)) return null;
// NOTE: parseRgGrepPattern uses split(/\s+/) and cannot handle shell
// quoting. `rg "User Service" src/` returns "User" (the first token
// after the rg/grep arg, with surrounding quotes stripped) — the
// multi-word pattern is intentionally not reconstructed since BM25 is
// already token-tolerant. Quoted single tokens (`rg "validateUser"`)
// work fine.
return parseRgGrepPattern(cmd);
}
return null;
}
function resolveCliPath() {
try {
return require.resolve('gitnexus/dist/cli/index.js');
} catch {
return '';
}
}
function runGitNexusCli(cliPath, args, cwd, timeout) {
const isWin = process.platform === 'win32';
if (cliPath) {
return spawnSync(process.execPath, [cliPath, ...args], {
encoding: 'utf-8',
timeout,
cwd,
stdio: ['pipe', 'pipe', 'pipe'],
});
}
return spawnSync(isWin ? 'npx.cmd' : 'npx', ['-y', 'gitnexus', ...args], {
encoding: 'utf-8',
timeout: timeout + 5000,
cwd,
stdio: ['pipe', 'pipe', 'pipe'],
});
}
function main() {
try {
const input = readInput();
if (process.env.GITNEXUS_DEBUG) {
// Echo the payload so users can capture Cursor's actual contract when
// diagnosing why augmentation isn't firing. Stderr only — stdout is
// reserved for the JSON response Cursor consumes.
try {
process.stderr.write(
`GitNexus Cursor hook stdin: ${JSON.stringify(input).slice(0, 500)}\n`,
);
} catch {
/* never let debug logging break the hook */
}
}
const cwd = input.cwd || process.cwd();
if (!path.isAbsolute(cwd)) return;
const gitNexusDir = findGitNexusDir(cwd);
if (!gitNexusDir) return;
const toolName = input.tool_name || '';
const toolInput = input.tool_input || {};
const pattern = extractPattern(toolName, toolInput);
if (!pattern || pattern.length < 3) return;
const release = acquireHookSlot(gitNexusDir);
if (!release) return;
const cliPath = resolveCliPath();
let result = '';
try {
const child = runGitNexusCli(cliPath, ['augment', '--', pattern], cwd, 7000);
if (!child.error && child.status === 0) {
result = child.stderr || '';
}
} catch {
/* graceful failure */
} finally {
release();
}
if (result && result.trim()) {
console.log(JSON.stringify({ additional_context: result.trim() }));
}
} catch (err) {
if (process.env.GITNEXUS_DEBUG) {
console.error('GitNexus Cursor hook error:', (err.message || '').slice(0, 200));
}
}
}
main();
@@ -0,0 +1,119 @@
const fs = require('fs');
const path = require('path');
const HOOK_LOCK_SUBDIR = '.hook-locks';
const HOOK_LOCK_MAX_INFLIGHT = 3;
const HOOK_LOCK_STALE_MS = 30000;
function acquireHookSlot(gitNexusDir) {
const lockDir = path.join(gitNexusDir, HOOK_LOCK_SUBDIR);
try {
fs.mkdirSync(lockDir, { recursive: true });
} catch {
// Cannot create lock dir (read-only fs, cross-user perm denial, out of
// inodes, etc.) — fail closed by returning null. Caller skips augment.
// Fail-open here would let N concurrent hooks all proceed unguarded and
// reintroduce the #1486 fan-out the guard exists to prevent.
return null;
}
const myPidStr = String(process.pid);
for (let slot = 0; slot < HOOK_LOCK_MAX_INFLIGHT; slot++) {
const slotPath = path.join(lockDir, `slot-${slot}.lock`);
for (let attempt = 0; attempt < 2; attempt++) {
try {
fs.writeFileSync(slotPath, myPidStr, { flag: 'wx' });
let released = false;
const release = () => {
if (released) return;
released = true;
try {
// Only unlink if we still own the slot. If we appeared stale and
// another hook took over, the file now belongs to it — leave alone.
const content = fs.readFileSync(slotPath, 'utf-8').trim();
if (content === myPidStr) fs.unlinkSync(slotPath);
} catch {
/* already removed or unreadable */
}
};
process.on('exit', release);
return release;
} catch {
// Slot exists. Decide whether to take it over.
// Open once and inspect mtime + content via the same fd so there's
// no TOCTOU between the metadata check and the content read
// (codeql js/file-system-race).
let fd;
try {
fd = fs.openSync(slotPath, 'r');
} catch {
continue; // Vanished between EEXIST and open — retry this slot.
}
let isLive = false;
let mtimeMs = Date.now();
try {
mtimeMs = fs.fstatSync(fd).mtimeMs;
const buf = Buffer.alloc(32);
const n = fs.readSync(fd, buf, 0, 32, 0);
const ownerStr = buf.slice(0, n).toString('utf-8').trim();
if (ownerStr === '') {
// Owner created the file but hasn't written its PID yet. The
// wx open+write window is microseconds; give it the benefit
// of the doubt and treat as live.
isLive = true;
} else {
const owner = Number.parseInt(ownerStr, 10);
if (Number.isFinite(owner) && owner > 0) {
try {
process.kill(owner, 0);
isLive = true;
} catch (e) {
// ESRCH = process gone → treat as dead. EPERM = process exists
// but owned by another user (cross-user lock dir) → still alive,
// keep the slot. Anything else: be conservative, assume alive.
if (e && e.code === 'ESRCH') {
isLive = false;
} else {
isLive = true;
}
}
}
}
} catch {
/* unreadable — treat as dead */
} finally {
try {
fs.closeSync(fd);
} catch {
/* already closed */
}
}
// For slots younger than HOOK_LOCK_STALE_MS, PID-liveness wins —
// a slow-but-alive hook is never wrongly evicted. For older slots,
// age is the final arbiter as a defense against PID reuse on long-
// abandoned slots. 30s >> the 7s augment timeout, so a healthy run
// never crosses this threshold.
if (isLive && Date.now() - mtimeMs > HOOK_LOCK_STALE_MS) {
isLive = false;
}
if (isLive) break; // Try the next slot.
try {
fs.unlinkSync(slotPath);
} catch {
/* another hook beat us to it — retry will hit EEXIST */
}
// Loop and retry this slot.
}
}
}
return null;
}
module.exports = {
HOOK_LOCK_SUBDIR,
HOOK_LOCK_MAX_INFLIGHT,
HOOK_LOCK_STALE_MS,
acquireHookSlot,
};
+4 -4
View File
@@ -1,11 +1,11 @@
{
"version": 1,
"hooks": {
"beforeShellExecution": [
"postToolUse": [
{
"command": "./hooks/augment-shell.sh",
"timeout": 5,
"matcher": "\\brg\\b|\\bgrep\\b"
"matcher": "Shell|Read|Grep",
"command": "node ./hooks/gitnexus-hook.cjs",
"timeout": 10
}
]
}
+4
View File
@@ -10,6 +10,10 @@
".": {
"types": "./dist/index.d.ts",
"default": "./dist/index.js"
},
"./test-helpers": {
"types": "./dist/test-helpers.d.ts",
"default": "./dist/test-helpers.js"
}
},
"scripts": {
+31 -1
View File
@@ -26,7 +26,7 @@ export type { PipelinePhase, PipelineProgress } from './pipeline.js';
// ─── Scope-based resolution — RFC #909 (Ring 1 #910) ────────────────────────
// Data model (RFC §2)
export type { SymbolDefinition } from './scope-resolution/symbol-definition.js';
export type { ParameterTypeClass, SymbolDefinition } from './scope-resolution/symbol-definition.js';
export type {
ScopeId,
DefId,
@@ -127,8 +127,10 @@ export { CLASS_KINDS, METHOD_KINDS, FIELD_KINDS } from './scope-resolution/regis
export type {
RegistryContext,
RegistryProviders,
OwnedMembersByOwnerLookup,
OwnerScopedContributor,
ArityVerdict,
ConstraintContext,
} from './scope-resolution/registries/context.js';
// Scope tree spine + position lookup (RFC §2.2 + §3.1; Ring 2 SHARED #912)
@@ -143,6 +145,34 @@ export type { ScopeTree } from './scope-resolution/scope-tree.js';
export { buildPositionIndex } from './scope-resolution/position-index.js';
export type { PositionIndex } from './scope-resolution/position-index.js';
// Resilient fetch primitives — bounded retries + per-process circuit breaker.
// Test-only helpers (`__resetBreakerRegistry__`, `classifyOutcome`) are
// reachable via the separate `gitnexus-shared/test-helpers` subpath; do
// NOT add them here. Production consumers must not call them.
export { withRetry, computeBackoffMs } from './integrations/retry.js';
export type { RetryOptions, RetryDecision } from './integrations/retry.js';
export { CircuitBreaker, CircuitOpenError, getBreaker } from './integrations/circuit-breaker.js';
export type { CircuitBreakerOptions } from './integrations/circuit-breaker.js';
export {
resilientFetch,
ResilientFetchExhaustedError,
RETRY_AFTER_CAP_MS,
parseRetryAfter,
} from './integrations/resilient-fetch.js';
export type { ResilientFetchOptions } from './integrations/resilient-fetch.js';
// Understand-Quickly registry integration (opt-in)
export {
UNDERSTAND_QUICKLY_DISPATCH_URL,
UNDERSTAND_QUICKLY_EVENT_TYPE,
UNDERSTAND_QUICKLY_TOKEN_ENV,
buildUqDispatchPayload,
isValidOwnerRepo,
parseOwnerRepoFromRemote,
stripGitSuffix,
} from './integrations/understand-quickly.js';
export type { UqDispatchPayload } from './integrations/understand-quickly.js';
// Shadow-mode diff + aggregation (RFC §6.3; Ring 2 SHARED #918)
export { diffResolutions } from './scope-resolution/shadow/diff.js';
export type {
@@ -0,0 +1,273 @@
/**
* Per-process circuit breaker.
*
* Closed -> Open transition fires after `failureThreshold` consecutive
* failures. While Open, `check` throws `CircuitOpenError` until
* `cooldownMs` has elapsed since the breaker tripped. The first call
* after the cooldown enters Half-Open and consumes the *probe permit*:
* a recorded success returns to Closed; a recorded failure flips back
* to Open with a fresh timestamp.
*
* Half-open admits exactly one in-flight probe at a time. Concurrent
* callers attempting `check()` while a probe is outstanding receive
* `CircuitOpenError` with `retryAfterMs = halfOpenRetryAfterMs` (default
* 1000ms; configurable). This prevents the recovery-time thundering
* herd that defeats the breaker's "fail fast" promise.
*
* Outcome reporting splits permit-release from state-resolution:
* - `recordSuccess` — releases the probe permit, resets the failure
* counter, transitions to Closed. Reserved for true 2xx/3xx outcomes.
* - `recordFailure` — releases the probe permit, increments the
* consecutive-failure counter, transitions to Open with a fresh
* `openedAt` (when called from Half-Open or when the threshold
* trips from Closed).
* - `recordNeutral` — releases the probe permit, BUT leaves state and
* counter untouched. Used for outcomes that are neither evidence of
* backend health nor evidence of backend failure (caller-driven
* cancellation, local timeout, terminal 4xx client errors). Critical
* design point: if `recordNeutral` did not release the permit, a
* single `TimeoutError` from per-attempt `AbortSignal.timeout` would
* route through `recordNeutral` and permanently park the breaker in
* half-open until process restart. Releasing the permit while leaving
* state half-open keeps the "neutral doesn't claim health" semantic
* without creating that wedge.
*
* Pairing invariant: every successful `check()` MUST be paired with
* exactly one `record*()` on every code path including throws. Direct
* consumers should wrap the protected operation in `try/finally`:
*
* breaker.check();
* try {
* const result = await operation();
* breaker.recordSuccess();
* return result;
* } catch (err) {
* // classify err and call recordFailure / recordNeutral / etc.
* throw err;
* }
*
* `resilientFetch`'s catch-all on `fetchImpl` already satisfies this
* for that consumer.
*
* Atomicity model: the half-open gate relies on JavaScript event-loop
* single-threadedness within a synchronous `check()` body. There is no
* `await` inside `check()`; concurrent callers serialize on microtask
* order, and exactly one observes `probeInFlight === false`. Do not
* introduce `await` inside `check()` without revisiting the gate. If
* this code is ever ported to a runtime with shared-memory threads
* (Node `worker_threads` with `SharedArrayBuffer`, Web Workers with
* shared registries), the boolean must become an atomic CAS — Resilience4j
* and Hystrix use atomic permits *because* they run in JVM thread pools.
*
* Runtime-agnostic: depends only on a `now()` clock and standard JS —
* no Node-only imports. Tests inject `now` to advance the clock
* deterministically without `vi.useFakeTimers()`.
*/
export class CircuitOpenError extends Error {
override readonly name = 'CircuitOpenError';
/** Approximate wait time before the breaker may transition to Half-Open
* (or before the in-flight probe is expected to resolve). */
readonly retryAfterMs: number;
constructor(retryAfterMs: number, key?: string) {
super(
key
? `Circuit '${key}' is open; retry in ${Math.ceil(retryAfterMs / 1000)}s`
: `Circuit is open; retry in ${Math.ceil(retryAfterMs / 1000)}s`,
);
this.retryAfterMs = retryAfterMs;
}
}
export interface CircuitBreakerOptions {
/** Consecutive failures required to trip Closed -> Open. */
failureThreshold?: number;
/** Milliseconds Open before the next call may probe (Half-Open). */
cooldownMs?: number;
/**
* Milliseconds to suggest in `CircuitOpenError.retryAfterMs` when the
* breaker is Half-Open with the probe permit consumed. Default 1000ms.
* Consumers with long-running protected ops (LLM streaming, large
* uploads) should raise this — the cooldown clock is no longer the
* right answer because cooldown has elapsed. Returning 0 invites
* retry storms; returning the full cooldown misleads about wait.
*/
halfOpenRetryAfterMs?: number;
/** Optional key for error messages and registry lookups. */
key?: string;
/** Clock override — defaults to `Date.now`. Tests inject deterministic time. */
now?: () => number;
}
type State = 'closed' | 'open' | 'half-open';
export class CircuitBreaker {
private readonly failureThreshold: number;
private readonly cooldownMs: number;
private readonly halfOpenRetryAfterMs: number;
private readonly key: string | undefined;
private readonly now: () => number;
private state: State = 'closed';
private consecutiveFailures = 0;
private openedAt: number | null = null;
/**
* True between a successful `check()` and the next `record*()` call
* during Half-Open. Gates concurrent callers from stampeding a still-
* recovering dependency. Boolean rather than counter — single-permit
* is the conservative end of the Hystrix/Resilience4j spectrum.
*/
private probeInFlight = false;
constructor(opts: CircuitBreakerOptions = {}) {
this.failureThreshold = opts.failureThreshold ?? 3;
this.cooldownMs = opts.cooldownMs ?? 30_000;
this.halfOpenRetryAfterMs = opts.halfOpenRetryAfterMs ?? 1_000;
this.key = opts.key;
this.now = opts.now ?? (() => Date.now());
}
/**
* Throw `CircuitOpenError` if the breaker won't admit this call.
* Otherwise consume the half-open probe permit (if applicable) and
* return so the caller can attempt the protected work.
*
* Three rejection paths:
* 1. Open and still in cooldown → throws with `retryAfterMs` =
* remaining cooldown.
* 2. Open with cooldown elapsed AND a probe is already in flight
* (race: another caller transitioned to half-open and grabbed
* the permit on a microtask before us) → throws with
* `halfOpenRetryAfterMs`.
* 3. Half-Open with probe in flight → throws with `halfOpenRetryAfterMs`.
*
* **Pairing invariant**: every successful return from `check()` MUST
* be paired with exactly one `recordSuccess` / `recordFailure` /
* `recordNeutral` on every code path including thrown exceptions.
* Failing to pair leaves the probe permit consumed forever and
* wedges the breaker. See file-header JSDoc for the canonical
* try/finally pattern.
*/
check(): void {
if (this.state === 'open' && this.openedAt !== null) {
const elapsed = this.now() - this.openedAt;
if (elapsed < this.cooldownMs) {
throw new CircuitOpenError(this.cooldownMs - elapsed, this.key);
}
// Cooldown elapsed — transition to Half-Open. The very next
// `probeInFlight` check below decides whether THIS caller gets
// the permit or hits the gate.
this.state = 'half-open';
}
if (this.state === 'half-open') {
if (this.probeInFlight) {
throw new CircuitOpenError(this.halfOpenRetryAfterMs, this.key);
}
this.probeInFlight = true;
}
// Closed state falls through silently.
}
recordSuccess(): void {
this.probeInFlight = false;
this.consecutiveFailures = 0;
this.state = 'closed';
this.openedAt = null;
}
recordFailure(): void {
this.probeInFlight = false;
this.consecutiveFailures += 1;
if (this.state === 'half-open' || this.consecutiveFailures >= this.failureThreshold) {
this.state = 'open';
this.openedAt = this.now();
}
}
/**
* Releases the probe permit BUT leaves state and counter untouched.
* Use when an attempt produced a response or error that should not
* influence breaker health in either direction — caller-driven aborts,
* local AbortSignal timeouts, terminal 4xx client errors.
*
* Why permit-release-without-state-resolution: if `recordNeutral` did
* not clear `probeInFlight`, a single `TimeoutError` from per-attempt
* `AbortSignal.timeout` (which routes through neutral classification)
* would permanently park the breaker in half-open. Since timeouts are
* an *expected* outcome under flaky-dependency conditions, the cited
* "per-attempt timeout bounds the stuck state" mitigation would itself
* be the trigger for a permanent wedge. Releasing the permit closes
* that loop while keeping the "neutral doesn't claim dependency
* health" semantic.
*
* Calling `recordSuccess` for these would erase legitimate prior
* failure signal; calling `recordFailure` would trip the breaker for
* outcomes the backend isn't responsible for.
*/
recordNeutral(): void {
this.probeInFlight = false;
// State and consecutiveFailures are preserved by design.
}
/**
* Pure read — no state mutation, no permit accounting. Returns the
* *would-be* state at the current instant: 'half-open' if the breaker
* is open with cooldown elapsed (regardless of whether a probe is in
* flight), 'open' if open and still in cooldown, 'closed' otherwise.
*
* Inspection-only; safe to call from tests without consuming a probe
* permit. The implicit Open -> Half-Open transition that mutates
* `state` lives in `check()` only.
*/
getState(): State {
if (this.state === 'open' && this.openedAt !== null) {
const elapsed = this.now() - this.openedAt;
if (elapsed >= this.cooldownMs) return 'half-open';
}
return this.state;
}
getConsecutiveFailures(): number {
return this.consecutiveFailures;
}
/** Inspection-only test accessor for the half-open probe permit. */
isProbeInFlight(): boolean {
return this.probeInFlight;
}
/** Timestamp (ms since epoch) when the breaker last transitioned to Open,
* or `null` if it's currently Closed. Useful for computing remaining
* cooldown without consuming a probe permit via `check()`. */
getOpenedAt(): number | null {
return this.openedAt;
}
/** Configured cooldown duration in milliseconds. */
getCooldownMs(): number {
return this.cooldownMs;
}
}
// ─── Per-process registry ────────────────────────────────────────────
//
// Single shared map keyed on caller-chosen strings. Used by
// `resilient-fetch.ts` so multiple call sites targeting the same logical
// endpoint share breaker state. Per-process only — not persisted.
const registry = new Map<string, CircuitBreaker>();
export function getBreaker(key: string, opts?: CircuitBreakerOptions): CircuitBreaker {
let breaker = registry.get(key);
if (!breaker) {
breaker = new CircuitBreaker({ ...opts, key });
registry.set(key, breaker);
}
return breaker;
}
/**
* Test-only: clear all registered breakers. Tests must call this in
* `beforeEach` to prevent breaker state from leaking across test cases.
*/
export function __resetBreakerRegistry__(): void {
registry.clear();
}
@@ -0,0 +1,279 @@
/**
* `resilientFetch` — fetch wrapped in retry + circuit breaker, with
* GitHub-flavoured retry classification baked in (Retry-After parsing,
* 401/403/404/422 treated as terminal client errors).
*
* Designed for the `gitnexus publish` GitHub `repository_dispatch`
* call, but the classification rules apply to any GitHub REST endpoint.
* Runtime-agnostic — no Node-only imports.
*/
import {
CircuitBreaker,
CircuitOpenError,
getBreaker,
type CircuitBreakerOptions,
} from './circuit-breaker.js';
import { computeBackoffMs, type RetryOptions } from './retry.js';
export { CircuitOpenError };
export interface ResilientFetchOptions {
/** Optional fetch implementation override. Defaults to `globalThis.fetch`. */
fetchImpl?: typeof fetch;
/**
* Logical key for the breaker. Defaults to `<host><pathname>` of the
* request URL — call sites targeting the same endpoint share breaker
* state regardless of query-string differences.
*/
breakerKey?: string;
/** Per-call breaker override. Used for tests and one-off configuration. */
breaker?: CircuitBreaker;
/** Tuning knobs for the breaker registered under `breakerKey`. */
breakerOptions?: CircuitBreakerOptions;
/** Tuning knobs for the retry helper. */
retry?: Partial<Pick<RetryOptions, 'maxAttempts' | 'baseDelayMs' | 'capDelayMs'>> & {
sleep?: RetryOptions['sleep'];
random?: RetryOptions['random'];
};
/** Clock override propagated into Retry-After HTTP-date math and breaker. */
now?: () => number;
}
/** Cap on any single Retry-After wait — protects CLI from a buggy registry. */
export const RETRY_AFTER_CAP_MS = 30_000;
const DEFAULT_RETRY = {
maxAttempts: 3,
baseDelayMs: 500,
capDelayMs: 5_000,
};
/**
* Parse a `Retry-After` header value into milliseconds.
* Accepts either a delta-seconds integer (`"30"`) or an HTTP-date.
* Returns null on parse failure or negative deltas.
*/
export function parseRetryAfter(value: string | null, now: () => number = Date.now): number | null {
if (!value) return null;
const trimmed = value.trim();
if (trimmed === '') return null;
if (/^[0-9]+$/.test(trimmed)) {
const seconds = parseInt(trimmed, 10);
if (Number.isNaN(seconds) || seconds < 0) return null;
return seconds * 1000;
}
const target = Date.parse(trimmed);
if (Number.isNaN(target)) return null;
const delta = target - now();
return delta >= 0 ? delta : 0;
}
/** Internal: outcome classification used by the resilientFetch loop. */
type Outcome =
| { kind: 'success'; resp: Response }
| { kind: 'terminal-client'; resp: Response } // 4xx other than 429: no retry, breaker neutral
| { kind: 'retryable-status'; resp: Response; afterMs: number | undefined } // 5xx, 429
| { kind: 'terminal-network'; err: unknown } // TimeoutError or AbortError: no retry, breaker neutral
| { kind: 'retryable-network'; err: unknown }; // DNS, ECONNRESET, etc.
/** Exported for unit tests. */
export function classifyOutcome(
result: { kind: 'error'; err: unknown } | { kind: 'response'; resp: Response },
now: () => number,
): Outcome {
if (result.kind === 'error') {
// Both timer-fired aborts (`AbortSignal.timeout()` → `TimeoutError`)
// and caller-driven aborts (`AbortController.abort()` → `AbortError`)
// are terminal: retrying against an already-aborted signal would
// fail again immediately, and neither outcome reflects backend
// health. They route through the breaker's neutral path.
if (
result.err instanceof DOMException &&
(result.err.name === 'TimeoutError' || result.err.name === 'AbortError')
) {
return { kind: 'terminal-network', err: result.err };
}
return { kind: 'retryable-network', err: result.err };
}
const resp = result.resp;
if (resp.status >= 200 && resp.status < 400) return { kind: 'success', resp };
if (resp.status === 429) {
// `resp.headers` is always present on a real `Response`, but tests
// sometimes stub `fetch` with a plain `{ ok, status }` object. Be
// defensive — a missing `Retry-After` falls through to exponential
// backoff, which is the correct behaviour anyway.
const retryAfterHeader =
typeof resp.headers?.get === 'function' ? resp.headers.get('Retry-After') : null;
const parsed = parseRetryAfter(retryAfterHeader, now);
return {
kind: 'retryable-status',
resp,
afterMs: parsed !== null ? Math.min(parsed, RETRY_AFTER_CAP_MS) : undefined,
};
}
if (resp.status >= 500) return { kind: 'retryable-status', resp, afterMs: undefined };
return { kind: 'terminal-client', resp };
}
const defaultSleep = (ms: number): Promise<void> =>
new Promise((resolve) => setTimeout(resolve, ms));
function defaultBreakerKey(input: string | URL): string {
try {
const url = typeof input === 'string' ? new URL(input) : input;
return `${url.host}${url.pathname}`;
} catch {
return String(input);
}
}
/** Final error thrown when retries are exhausted on a 5xx / 429. */
export class ResilientFetchExhaustedError extends Error {
override readonly name = 'ResilientFetchExhaustedError';
constructor(public readonly response: Response) {
super(`Request failed after retries (HTTP ${response.status})`);
}
}
/**
* Wrap `fetch` with bounded retries and a per-process circuit breaker.
*
* Semantics:
* - 5xx and 429 responses are retried; 429 honors `Retry-After` (capped).
* - Network throws are retried unless they are `TimeoutError` DOMExceptions.
* - Timeouts and 4xx (other than 429) are returned/thrown without retry
* AND without incrementing the breaker — they reflect caller config
* or local network state, not registry health.
* - Each `fetch` call carries the caller-supplied `signal` (e.g. an
* `AbortSignal.timeout()`) — that timeout bounds each individual
* attempt, not the whole retry sequence.
* - When the breaker is open, throws `CircuitOpenError` synchronously
* without invoking `fetch`.
* - When retries are exhausted on a 5xx / 429, throws
* `ResilientFetchExhaustedError` carrying the last response.
*
* Cumulative wall-clock budget:
* maxAttempts × (per-attempt-timeout + capDelayMs)
* With defaults (3, 500ms base, 5000ms cap) and a typical 15s per-attempt
* timeout from the caller's signal, worst case is ~3 × (15s + 5s) = 60s.
* Callers that want a tighter total bound should reduce `maxAttempts` or
* wrap `resilientFetch` in their own outer `AbortSignal.timeout()`.
*/
export async function resilientFetch(
input: string | URL,
init: RequestInit | undefined,
opts: ResilientFetchOptions = {},
): Promise<Response> {
const fetchImpl = opts.fetchImpl ?? globalThis.fetch;
const now = opts.now ?? (() => Date.now());
const breaker =
opts.breaker ?? getBreaker(opts.breakerKey ?? defaultBreakerKey(input), opts.breakerOptions);
const retryConfig = {
maxAttempts: opts.retry?.maxAttempts ?? DEFAULT_RETRY.maxAttempts,
baseDelayMs: opts.retry?.baseDelayMs ?? DEFAULT_RETRY.baseDelayMs,
capDelayMs: opts.retry?.capDelayMs ?? DEFAULT_RETRY.capDelayMs,
};
const sleep = opts.retry?.sleep ?? defaultSleep;
const random = opts.retry?.random ?? Math.random;
// Fail fast on an open breaker, before invoking fetch.
breaker.check();
for (let attempt = 0; attempt < retryConfig.maxAttempts; attempt++) {
let result: { kind: 'error'; err: unknown } | { kind: 'response'; resp: Response };
try {
// CodeQL js/server-side-request-forgery — flagged because `input`
// is caller-supplied. Suppressed: every concrete caller passes
// either a hardcoded URL constant (UNDERSTAND_QUICKLY_DISPATCH_URL,
// OpenRouter base URL) or a value derived from configuration
// (env vars, saved settings, the local backend URL). User-input
// request fields (e.g. PR title, repo name) never flow into
// `input`. Validating URL shape here would push false-positive
// rejection onto every caller — wrong layer for the check.
// lgtm[js/server-side-request-forgery]
// codeql[js/server-side-request-forgery]
const resp = await fetchImpl(input, init);
result = { kind: 'response', resp };
} catch (err) {
result = { kind: 'error', err };
}
const outcome = classifyOutcome(result, now);
switch (outcome.kind) {
case 'success':
breaker.recordSuccess();
return outcome.resp;
case 'terminal-client':
// 4xx: do not count as breaker failure (the server is healthy
// and rejecting our request — auth, scope, or routing). But
// also do NOT call recordSuccess: a 401 sandwiched between
// 5xx responses would otherwise erase the running outage
// signal. The breaker's neutral path leaves state untouched.
breaker.recordNeutral();
return outcome.resp;
case 'terminal-network':
// Either `AbortSignal.timeout()` fired locally OR an external
// caller cancelled the request via AbortController. The server
// never had a chance to answer; this reflects the user's
// network or an explicit cancel, not registry health. Don't
// punish the breaker AND don't reset its outage signal.
breaker.recordNeutral();
throw outcome.err;
case 'retryable-status':
if (attempt + 1 >= retryConfig.maxAttempts) {
breaker.recordFailure();
throw new ResilientFetchExhaustedError(outcome.resp);
}
await sleep(
computeBackoffMs(
attempt,
retryConfig.baseDelayMs,
retryConfig.capDelayMs,
outcome.afterMs,
random,
),
);
break;
case 'retryable-network':
if (attempt + 1 >= retryConfig.maxAttempts) {
breaker.recordFailure();
throw outcome.err;
}
await sleep(
computeBackoffMs(
attempt,
retryConfig.baseDelayMs,
retryConfig.capDelayMs,
undefined,
random,
),
);
break;
default: {
// Exhaustiveness guard. If a sixth `Outcome` kind is added in
// future, TypeScript will refuse to assign it to `never` and
// this line forces the maintainer to add an explicit arm
// rather than silently fall through to retry/no-retry behaviour.
const _exhaustive: never = outcome;
throw new Error(`resilientFetch: unhandled outcome ${JSON.stringify(_exhaustive)}`);
}
}
}
// Unreachable: every iteration of the loop either returns (success
// / terminal-client) or throws (terminal-network / retry exhaustion).
// The throw is here purely so TypeScript's control-flow analysis sees
// the function never falls off the end without producing `Promise<Response>`.
/* c8 ignore next 2 */
throw new Error('resilientFetch: retry loop terminated unexpectedly');
}
+105
View File
@@ -0,0 +1,105 @@
/**
* Bounded retry helper with full-jitter exponential backoff.
*
* Runtime-agnostic: depends only on `setTimeout`, `Math.random`, and the
* Promise machinery — no Node-only imports. Safe to consume from CLI,
* server, or browser callers.
*
* Pattern reference: gitnexus/src/core/embeddings/http-client.ts. This
* helper is the upgraded form: classification is caller-supplied (so
* 4xx-vs-5xx-vs-timeout decisions live with the protocol that knows
* them), backoff is exponential with full jitter, and an optional
* `afterMs` lets callers honor `Retry-After` headers.
*/
export interface RetryOptions {
/** Initial delay before the first retry attempt, in milliseconds. */
baseDelayMs: number;
/** Upper bound on any single delay, in milliseconds. */
capDelayMs: number;
/** Total attempts including the first call. Must be >= 1. */
maxAttempts: number;
/**
* Decide whether to retry after a thrown error.
* Return `{retry:false}` to terminate immediately and rethrow.
* Return `{retry:true}` to retry with exponential-backoff jitter.
* Return `{retry:true, afterMs}` to wait at least `afterMs` (still
* subject to `capDelayMs`) — used by callers parsing `Retry-After`.
*/
isRetryable: (err: unknown, attempt: number) => RetryDecision;
/** Sleep override — defaults to `setTimeout`. Tests inject fake timers. */
sleep?: (ms: number) => Promise<void>;
/** Random override — defaults to `Math.random`. Tests inject seeded values. */
random?: () => number;
}
export type RetryDecision = { retry: false } | { retry: true; afterMs?: number };
const defaultSleep = (ms: number): Promise<void> =>
new Promise((resolve) => setTimeout(resolve, ms));
/**
* Compute the delay before the next retry attempt.
*
* - When the caller specifies `afterMs` (e.g., from `Retry-After`), use
* `min(afterMs, capDelayMs)` so a misbehaving server can't pin the
* client for an arbitrarily long wait.
* - Otherwise compute full-jitter exponential backoff:
* `random() * min(cap, base * 2^attempt)`. Full jitter (rather than
* "equal jitter") avoids retry-storm thundering herd, per AWS
* guidance on backoff strategies.
*/
export function computeBackoffMs(
attempt: number,
baseDelayMs: number,
capDelayMs: number,
afterMs: number | undefined,
random: () => number,
): number {
if (afterMs !== undefined) {
return Math.min(Math.max(0, afterMs), capDelayMs);
}
const exponential = baseDelayMs * Math.pow(2, attempt);
const upper = Math.min(capDelayMs, exponential);
return Math.floor(random() * upper);
}
/**
* Execute `fn` with bounded retries.
*
* The classification of "retryable" is the caller's responsibility — see
* `resilient-fetch.ts` for the GitHub-dispatch-specific rules. This
* helper is the mechanical retry loop only.
*/
export async function withRetry<T>(
fn: (attempt: number) => Promise<T>,
opts: RetryOptions,
): Promise<T> {
if (opts.maxAttempts < 1) {
throw new Error(`withRetry: maxAttempts must be >= 1, got ${opts.maxAttempts}`);
}
const sleep = opts.sleep ?? defaultSleep;
const random = opts.random ?? Math.random;
let lastError: unknown;
for (let attempt = 0; attempt < opts.maxAttempts; attempt++) {
try {
return await fn(attempt);
} catch (err) {
lastError = err;
const decision = opts.isRetryable(err, attempt);
if (!decision.retry) throw err;
// Don't sleep after the final attempt.
if (attempt + 1 >= opts.maxAttempts) break;
const delayMs = computeBackoffMs(
attempt,
opts.baseDelayMs,
opts.capDelayMs,
decision.afterMs,
random,
);
if (delayMs > 0) await sleep(delayMs);
}
}
throw lastError;
}
@@ -0,0 +1,151 @@
/**
* Understand-Quickly registry integration helpers.
*
* Pure, runtime-agnostic logic for opting in to publishing a GitNexus
* index to the [`looptech-ai/understand-quickly`](https://github.com/looptech-ai/understand-quickly)
* registry. Lives in `gitnexus-shared` so both the Node CLI and any
* future browser-side surface can construct identical dispatch payloads.
*
* Network I/O lives in the CLI command (`gitnexus/src/cli/publish.ts`)
* to keep this module free of Node-only imports — see the comment at
* the top of `gitnexus-shared/src/graph/types.ts`.
*
* The protocol contract (single dispatch event, no graph upload) is
* documented at:
* https://github.com/looptech-ai/understand-quickly/blob/main/docs/integrations/protocol.md
*/
/**
* URL of the registry repo's repository_dispatch endpoint. Hardcoded
* because the registry is the canonical home for this integration —
* users who want a private registry can fork and patch.
*/
export const UNDERSTAND_QUICKLY_DISPATCH_URL =
'https://api.github.com/repos/looptech-ai/understand-quickly/dispatches';
/**
* Event type the registry's sync workflow listens for.
* See `looptech-ai/understand-quickly/.github/workflows/sync.yml`.
*/
export const UNDERSTAND_QUICKLY_EVENT_TYPE = 'sync-entry';
/** Environment variable that gates the dispatch. */
export const UNDERSTAND_QUICKLY_TOKEN_ENV = 'UNDERSTAND_QUICKLY_TOKEN';
export interface UqDispatchPayload {
event_type: typeof UNDERSTAND_QUICKLY_EVENT_TYPE;
client_payload: {
/** `<owner>/<repo>` shape — must match the registered entry. */
id: string;
};
}
/**
* Build the JSON body for the `repository_dispatch` ping. Pure — no
* env reads, no network. Validates that `id` looks like `owner/repo`
* (one slash, no whitespace, both halves non-empty) so a misconfigured
* caller fails loudly before the round-trip.
*/
export function buildUqDispatchPayload(id: string): UqDispatchPayload {
if (!isValidOwnerRepo(id)) {
throw new Error(
`[understand-quickly] expected id of the form "owner/repo", got "${id}". ` +
`The registry uses this string to look up your entry in registry.json — ` +
`it must match the GitHub owner/repo of the source code, not a local path.`,
);
}
return {
event_type: UNDERSTAND_QUICKLY_EVENT_TYPE,
client_payload: { id },
};
}
/**
* `owner/repo` validation. Conservative on purpose: GitHub's actual
* naming rules are looser, but we want to catch local paths
* (`/Users/...`), bare slugs (`my-repo`), and accidental whitespace.
*
* Matches GitHub's published slug rules:
* owner: starts with alnum, then alnum/hyphen only, must end with
* alnum (no trailing hyphen — GitHub rejects this at account
* creation, so a `my-org-/repo` input would otherwise pass us
* and 422 from GitHub). No underscore, no dot. Length cap 39.
* repo: any of alnum/dot/hyphen/underscore. Length cap 100.
*/
export function isValidOwnerRepo(id: string): boolean {
return /^[A-Za-z0-9](?:[A-Za-z0-9-]{0,37}[A-Za-z0-9])?\/[A-Za-z0-9._-]{1,100}$/.test(id);
}
/**
* Strip a single trailing `.git` (case-insensitive) and any trailing
* slashes from a URL-ish string. Bounded linear: each character is
* visited at most twice, no backtracking.
*
* Replaces `s.replace(/\.git\/*$/i, '').replace(/\/+$/, '')` which
* CodeQL's polynomial-regex check (codeql/js/polynomial-redos) flags as
* a worst-case O(n²) on adversarial input like "////.../x".
*/
export function stripGitSuffix(input: string): string {
let end = input.length;
// Trim trailing '/'.
while (end > 0 && input.charCodeAt(end - 1) === 0x2f) end--;
// Drop one trailing '.git' (case-insensitive).
if (end >= 4) {
const tail = input.slice(end - 4, end).toLowerCase();
if (tail === '.git') end -= 4;
}
// Trim trailing '/' that may have sat between '.git' and the rest.
while (end > 0 && input.charCodeAt(end - 1) === 0x2f) end--;
return input.slice(0, end);
}
/**
* Parse `owner/repo` out of a git remote URL. Mirrors the heuristic in
* `gitnexus/src/storage/git.ts:parseRepoNameFromUrl` but keeps both
* halves so we can build a registry id. Returns `null` on shapes we
* don't recognise.
*
* Examples:
* git@github.com:looptech-ai/understand-quickly.git
* https://github.com/looptech-ai/understand-quickly
* ssh://git@github.com/looptech-ai/understand-quickly.git
*/
export function parseOwnerRepoFromRemote(url: string | null | undefined): string | null {
if (!url) return null;
const trimmed = url.trim();
if (!trimmed) return null;
// Strip a trailing `.git` (case-insensitive) and any trailing slashes
// so https://h/o/r and https://h/o/r.git collapse to the same id.
// Bounded-linear helper avoids the polynomial-regex CodeQL alert.
const stripped = stripGitSuffix(trimmed);
// SCP-form SSH (`git@host:owner/repo`). Capture host so we can reject
// non-GitHub remotes — a GitLab origin like
// `https://gitlab.example.com/group/sub/project.git` would otherwise
// silently dispatch the wrong id (LOW 9).
const ssh = stripped.match(/^[^@]+@([^:]+):([^/]+)\/([^/]+)$/);
if (ssh) {
const host = ssh[1].toLowerCase();
if (host !== 'github.com' && host !== 'www.github.com') return null;
return `${ssh[2]}/${ssh[3]}`;
}
// URL forms (https://, ssh://, git://, file://) — last two path segments.
const url2 = stripped.match(/^[a-zA-Z][a-zA-Z0-9+.-]*:\/\/([^/]+)\/(.+)$/);
if (url2) {
// Strip optional `userinfo@` (e.g. `ssh://git@github.com/...`).
const authority = url2[1];
const atIdx = authority.lastIndexOf('@');
const hostAndPort = atIdx >= 0 ? authority.slice(atIdx + 1) : authority;
// Strip `:port` suffix if present.
const colonIdx = hostAndPort.indexOf(':');
const host = (colonIdx >= 0 ? hostAndPort.slice(0, colonIdx) : hostAndPort).toLowerCase();
if (host !== 'github.com' && host !== 'www.github.com') return null;
const segments = url2[2].split('/').filter(Boolean);
if (segments.length >= 2) {
const [owner, repo] = segments.slice(-2);
return `${owner}/${repo}`;
}
}
return null;
}
@@ -833,7 +833,16 @@ function expandWildcard(
if (target === undefined) return [edge];
const names = hooks.expandsWildcardTo(edge.targetModuleScope, workspace);
if (names.length === 0) return [];
if (names.length === 0) {
// Resolved wildcard with zero propagating names is still a real file-
// level dependency (e.g. a C++ header that only declares classes —
// `#include` is a valid IMPORTS edge, but unqualified-binding names
// are correctly empty since class methods require `Class::method`).
// Preserve the original wildcard edge so the file→file IMPORTS edge
// survives; downstream binding materialization sees no propagated
// names because the edge has no `targetExportedName`/`localName`.
return [edge];
}
const expanded: ImportEdge[] = [];
for (const name of names) {
@@ -40,11 +40,27 @@ export interface MethodDispatchIndex {
readonly mroByOwnerDefId: ReadonlyMap<DefId, readonly DefId[]>;
/** Interfaces / traits → classes that implement them. */
readonly implsByInterfaceDefId: ReadonlyMap<DefId, readonly DefId[]>;
/**
* Optional parallel MRO view that EXCLUDES mixin-like augmentation
* (e.g., PHP traits). Populated only when the input supplies
* `computeExtendsOnlyMro`. Used by the super-branch dispatch in
* `receiver-bound-calls` so that `parent::method()` walks the
* inheritance chain only, not the trait-augmented one. Undefined for
* languages without mixin-like semantics — callers should fall back
* to `mroFor` when this is missing.
*/
readonly extendsOnlyMroByOwnerDefId?: ReadonlyMap<DefId, readonly DefId[]>;
/** `mroByOwnerDefId.get`, with an empty frozen array on miss. */
mroFor(ownerDefId: DefId): readonly DefId[];
/** `implsByInterfaceDefId.get`, with an empty frozen array on miss. */
implementorsOf(interfaceDefId: DefId): readonly DefId[];
/**
* `extendsOnlyMroByOwnerDefId.get`, with an empty frozen array on miss.
* Undefined when `extendsOnlyMroByOwnerDefId` was not populated; callers
* should treat this as equivalent to `mroFor` for non-mixin languages.
*/
readonly extendsOnlyMroFor?: (ownerDefId: DefId) => readonly DefId[];
}
export interface MethodDispatchInput {
@@ -81,12 +97,25 @@ export interface MethodDispatchInput {
* write-wins policy and fires at most once per unique owner.
*/
readonly implementsOf: (ownerDefId: DefId) => readonly DefId[];
/**
* Optional: return the EXTENDS-only ancestor chain for `ownerDefId`,
* excluding the owner itself AND any mixin-like augmentation (e.g.,
* PHP traits). Languages without mixin semantics leave this undefined
* and the index's `extendsOnlyMroByOwnerDefId` stays unpopulated.
*
* Same contract as `computeMro`: pure, deterministic, `[]` on no parents.
* Called at most once per unique owner (first-write-wins).
*/
readonly computeExtendsOnlyMro?: (ownerDefId: DefId) => readonly DefId[];
}
// ─── Builder ────────────────────────────────────────────────────────────────
export function buildMethodDispatchIndex(input: MethodDispatchInput): MethodDispatchIndex {
const mroByOwnerDefId = new Map<DefId, readonly DefId[]>();
const extendsOnlyByOwnerDefId = input.computeExtendsOnlyMro
? new Map<DefId, readonly DefId[]>()
: undefined;
const implsBuilding = new Map<DefId, DefId[]>();
const implsSeen = new Map<DefId, Set<DefId>>();
@@ -97,6 +126,14 @@ export function buildMethodDispatchIndex(input: MethodDispatchInput): MethodDisp
const chain = input.computeMro(ownerId);
mroByOwnerDefId.set(ownerId, Object.freeze(chain.slice()));
}
if (
input.computeExtendsOnlyMro !== undefined &&
extendsOnlyByOwnerDefId !== undefined &&
!extendsOnlyByOwnerDefId.has(ownerId)
) {
const extOnly = input.computeExtendsOnlyMro(ownerId);
extendsOnlyByOwnerDefId.set(ownerId, Object.freeze(extOnly.slice()));
}
for (const ifaceId of input.implementsOf(ownerId)) {
let seen = implsSeen.get(ifaceId);
@@ -121,7 +158,7 @@ export function buildMethodDispatchIndex(input: MethodDispatchInput): MethodDisp
implsByInterfaceDefId.set(ifaceId, Object.freeze(owners.slice()));
}
return wrapIndex(mroByOwnerDefId, implsByInterfaceDefId);
return wrapIndex(mroByOwnerDefId, implsByInterfaceDefId, extendsOnlyByOwnerDefId);
}
// ─── Internal ───────────────────────────────────────────────────────────────
@@ -131,8 +168,9 @@ const EMPTY: readonly DefId[] = Object.freeze([]);
function wrapIndex(
mroByOwnerDefId: Map<DefId, readonly DefId[]>,
implsByInterfaceDefId: Map<DefId, readonly DefId[]>,
extendsOnlyMroByOwnerDefId: Map<DefId, readonly DefId[]> | undefined,
): MethodDispatchIndex {
return {
const base: MethodDispatchIndex = {
mroByOwnerDefId,
implsByInterfaceDefId,
mroFor(ownerDefId: DefId): readonly DefId[] {
@@ -142,4 +180,14 @@ function wrapIndex(
return implsByInterfaceDefId.get(interfaceDefId) ?? EMPTY;
},
};
if (extendsOnlyMroByOwnerDefId !== undefined) {
return {
...base,
extendsOnlyMroByOwnerDefId,
extendsOnlyMroFor(ownerDefId: DefId): readonly DefId[] {
return extendsOnlyMroByOwnerDefId.get(ownerDefId) ?? EMPTY;
},
};
}
return base;
}
@@ -21,6 +21,7 @@
* (defined in `./types.ts`).
*/
import type { ParameterTypeClass } from './symbol-definition.js';
import type { Range, ScopeId } from './types.js';
/**
@@ -79,4 +80,11 @@ export interface ReferenceSite {
* (C#: `42` → `'int'`, `"alice"` → `'string'`).
*/
readonly argumentTypes?: readonly string[];
/**
* Optional per-argument type-shape sidecar for languages that need
* cv/ref/pointer distinctions during constraint filtering. This is
* intentionally separate from `argumentTypes`, which stays normalized
* for existing overload narrowing and conversion-rank logic.
*/
readonly argumentTypeClasses?: readonly ParameterTypeClass[];
}
@@ -13,7 +13,7 @@
*/
import type { NodeLabel } from '../../graph/types.js';
import type { SymbolDefinition } from '../symbol-definition.js';
import type { ParameterTypeClass, SymbolDefinition } from '../symbol-definition.js';
import type { Callsite, DefId } from '../types.js';
import type { DefIndex } from '../def-index.js';
import type { QualifiedNameIndex } from '../qualified-name-index.js';
@@ -30,10 +30,50 @@ export interface RegistryProviders {
* when absent, every candidate receives `'unknown'` (neutral signal).
*/
arityCompatibility?(callsite: Callsite, def: SymbolDefinition): ArityVerdict;
/**
* Language-specific constraint compatibility between a callsite and a
* candidate `def`. Mirrors `arityCompatibility` and shares its three-valued
* verdict shape; the third value `'unknown'` MUST keep the candidate
* (monotonicity: adding a predicate can only narrow correctly, never
* produce a wrong edge). Consulted by `narrowOverloadCandidates` after
* arity + type filters when a candidate carries `templateConstraints`.
*
* Optional; when absent the constraint filter is a pass-through. Languages
* with no constrained-overload semantics leave this undefined.
*/
constraintCompatibility?(
callsite: Callsite,
def: SymbolDefinition,
ctx: ConstraintContext,
): ArityVerdict;
}
export type ArityVerdict = 'compatible' | 'unknown' | 'incompatible';
/**
* Context threaded into `constraintCompatibility`. Kept minimal in the
* Tier-A scope (only `argumentTypes`, riding here until a separate
* `Callsite`-widening refactor moves them onto the call site directly).
* Future Tier-B graph-aware predicates (`is_base_of_v`, etc.) will widen
* this interface with `lookupTypeByName` and similar helpers.
*/
export interface ConstraintContext {
/**
* Per-slot argument types at the call site, normalized per the language
* adapter. Empty string means unknown. Same convention as
* `narrowOverloadCandidates`' `argTypes` parameter.
*/
readonly argumentTypes?: readonly string[];
/**
* Optional shape-preserving sidecar aligned with `argumentTypes`.
* Unknown or unsupported slots should be omitted by producers or
* marked with `indirection: 'unknown'`; consumers must preserve the
* monotonic fallback and return 'unknown' instead of guessing.
*/
readonly argumentTypeClasses?: readonly ParameterTypeClass[];
}
// ─── Owner-scoped contributor (concrete shape for `RegistryContributor`) ────
/**
@@ -60,6 +100,19 @@ export interface OwnerScopedContributor {
byName(name: string): readonly SymbolDefinition[];
}
/**
* Required owner-keyed lookup hook for Step 2 receiver/MRO member walks.
* Production callers wire this to the SemanticModel's authoritative
* method/field/nested-type registries so each `(ownerDefId, memberName)`
* probe is O(1). Implementations MUST return `[]` on an indexed miss —
* Step 2 treats `[]` as authoritative and does not consult `defs` for a
* fallback scan.
*/
export type OwnedMembersByOwnerLookup = (
ownerDefId: DefId,
memberName: string,
) => readonly SymbolDefinition[];
// ─── Top-level context threaded through every lookup ───────────────────────
export interface RegistryContext {
@@ -67,6 +120,7 @@ export interface RegistryContext {
readonly defs: DefIndex;
readonly qualifiedNames: QualifiedNameIndex;
readonly moduleScopes: ModuleScopeIndex;
readonly ownedMembersByOwner: OwnedMembersByOwnerLookup;
/**
* Method-dispatch index; required for method/field registries that
* honor `useReceiverTypeBinding`. Omit for class-only lookups.
@@ -27,8 +27,10 @@
* is true, resolve the receiver's type at `startScope` (from
* `scope.typeBindings`), then walk the MRO via
* `MethodDispatchIndex.mroFor(ownerDefId)`. Membership per owner comes
* through `RegistryContext.methodDispatch` + owner lookups into
* `scope.ownedDefs`; each hit records a raw signal with the owner's
* through an optional `RegistryContext.ownedMembersByOwner` hook when
* supplied (`undefined` → fall back to `defs.byId`; `[]` → indexed
* miss), otherwise via the compatibility fallback scan over
* `defs.byId`; each hit records a raw signal with the owner's
* MRO depth.
*
* **Step 3 — Owner-scoped contributor.** When
@@ -263,13 +265,14 @@ function walkReceiverTypeBinding(
// Walk the owner itself at depth 0, then its MRO chain.
const walk: DefId[] = [ownerDefId, ...ctx.methodDispatch.mroFor(ownerDefId)];
for (let mroDepth = 0; mroDepth < walk.length; mroDepth++) {
const currentOwnerId = walk[mroDepth]!;
let mroDepth = 0;
for (const currentOwnerId of walk) {
const members = collectOwnedMembers(currentOwnerId, name, ctx);
for (const def of members) {
if (!acceptedKinds.has(def.type)) continue;
recordTypeBindingHit(perCandidate, def, mroDepth, ownerDefId);
}
mroDepth++;
}
}
@@ -333,23 +336,7 @@ function collectOwnedMembers(
memberName: string,
ctx: RegistryContext,
): readonly SymbolDefinition[] {
// An owner's members are defs whose `ownerId === ownerDefId` and whose
// simple name matches `memberName`. We iterate `defs.byId` — O(D) per
// call today. A future by-owner index would make this O(K); tracked as
// a follow-up optimization before Ring 3 flips go production.
const out: SymbolDefinition[] = [];
for (const def of ctx.defs.byId.values()) {
if (def.ownerId !== ownerDefId) continue;
if (simpleNameOf(def) !== memberName) continue;
out.push(def);
}
return out;
}
function simpleNameOf(def: SymbolDefinition): string | undefined {
if (def.qualifiedName === undefined || def.qualifiedName.length === 0) return undefined;
const dot = def.qualifiedName.lastIndexOf('.');
return dot === -1 ? def.qualifiedName : def.qualifiedName.slice(dot + 1);
return ctx.ownedMembersByOwner(ownerDefId, memberName);
}
function recordTypeBindingHit(
@@ -423,13 +410,30 @@ function applyArityFilter(
}
let anyCompatible = false;
let anyUnknown = false;
for (const state of perCandidate.values()) {
const verdict = arityFn(callsite, state.def);
state.signals.arityVerdict = verdict;
if (verdict === 'compatible') anyCompatible = true;
else if (verdict === 'unknown') anyUnknown = true;
}
if (!anyCompatible) return;
// When ALL candidates are 'incompatible' (none compatible, none unknown),
// the call is genuinely arity-broken — drop every candidate so the
// registry returns no resolution. This matches the PHP variadic case
// f(int $req, ...$rest) called with zero args: every candidate definitively
// rejects, and emitting an edge to a definitively-rejected callable is
// a false positive. When some candidates are 'unknown' (missing metadata),
// keep the set so downstream evidence can break the tie — that's the
// original safety-fallback behavior.
if (!anyCompatible) {
if (!anyUnknown) {
for (const defId of perCandidate.keys()) {
perCandidate.delete(defId);
}
}
return;
}
// Filter: when at least one compatible candidate exists, drop incompatibles.
for (const [defId, state] of perCandidate) {
@@ -11,6 +11,17 @@
import type { NodeLabel } from '../graph/types.js';
export interface ParameterTypeClass {
/** Normalized base type, matching the coarse `parameterTypes` vocabulary when known. */
base: string;
/** Top-level cv signal preserved from the original C++ parameter spelling. */
cv: 'none' | 'const' | 'volatile' | 'const volatile' | 'unknown';
/** Coarse value/reference/pointer shape. */
indirection: 'value' | 'lvalue-ref' | 'rvalue-ref' | 'pointer' | 'unknown';
/** Number of pointer markers when indirection is `pointer`; otherwise 0. */
pointerDepth: number;
}
export interface SymbolDefinition {
nodeId: string;
filePath: string;
@@ -26,10 +37,22 @@ export interface SymbolDefinition {
/** Per-parameter type names for overload disambiguation (e.g. ['int', 'String']).
* Populated when parameter types are resolvable from AST (any typed language). */
parameterTypes?: string[];
/** Additive per-parameter type shape sidecar for languages that need cv/ref/pointer distinctions.
* Does not participate in graph node identity unless a resolver explicitly opts in. */
parameterTypeClasses?: ParameterTypeClass[];
/** Raw return type text extracted from AST (e.g. 'User', 'Promise<User>') */
returnType?: string;
/** Declared type for non-callable symbols — fields/properties (e.g. 'Address', 'List<User>') */
declaredType?: string;
/** Generic/template specialization arguments for class-like symbols (e.g. ['User'], ['T*']). */
templateArguments?: string[];
/** Per-language constraint payload for template / generic overloads
* (e.g. C++ `enable_if_t<P, T>` predicate trees, C++20 `requires` clauses).
* Opaque to shared code — the producing language adapter owns the shape
* and is the only consumer. Read via the optional
* `ScopeResolver.constraintCompatibility` hook during overload narrowing.
* Absent for symbols that have no constraints (the common case). */
templateConstraints?: unknown;
/** Links Method/Constructor/Property to owning Class/Struct/Trait nodeId */
ownerId?: string;
}
+13
View File
@@ -0,0 +1,13 @@
/**
* Test-only helpers.
*
* Symbols here are reachable from `gitnexus-shared/test-helpers` so test
* suites can reset shared registries or exercise internal classifiers,
* but they are deliberately NOT re-exported from the main `gitnexus-shared`
* barrel. Production consumers should never import this module — calling
* `__resetBreakerRegistry__()` from a tool implementation would silently
* nuke every circuit breaker process-wide.
*/
export { __resetBreakerRegistry__ } from './integrations/circuit-breaker.js';
export { classifyOutcome } from './integrations/resilient-fetch.js';
+292 -402
View File
File diff suppressed because it is too large Load Diff
+11 -11
View File
@@ -19,39 +19,39 @@
},
"dependencies": {
"gitnexus-shared": "file:../gitnexus-shared",
"@langchain/anthropic": "^1.3.28",
"@langchain/anthropic": "^1.3.29",
"@langchain/core": "^1.1.44",
"@langchain/google-genai": "^2.1.28",
"@langchain/google-genai": "^2.1.30",
"@langchain/langgraph": "^1.2.9",
"@langchain/ollama": "^1.2.6",
"@langchain/openai": "^1.4.5",
"@sigma/edge-curve": "^3.1.0",
"@tailwindcss/vite": "^4.2.4",
"@tailwindcss/vite": "^4.3.0",
"axios": "^1.16.0",
"d3": "^7.9.0",
"dompurify": "^3.4.2",
"dompurify": "^3.4.3",
"graphology": "^0.26.0",
"graphology-indices": "^0.17.0",
"graphology-layout-force": "^0.2.4",
"graphology-layout-forceatlas2": "^0.10.1",
"graphology-layout-noverlap": "^0.4.2",
"graphology-utils": "^2.3.0",
"langchain": "^1.3.4",
"langchain": "^1.3.5",
"lru-cache": "^11.2.4",
"lucide-react": "^1.11.0",
"mermaid": "^11.14.0",
"lucide-react": "^1.14.0",
"mermaid": "^11.15.0",
"mnemonist": "^0.39.0",
"pandemonium": "^2.4.0",
"react": "^19.2.5",
"react-dom": "^19.2.6",
"react-markdown": "^10.1.0",
"react-syntax-highlighter": "^16.1.0",
"react-syntax-highlighter": "^16.1.1",
"react-zoom-pan-pinch": "^4.0.3",
"remark-gfm": "^4.0.1",
"sigma": "^3.0.2",
"tailwindcss": "^4.2.4",
"uuid": "^14.0.0",
"zod": "^3.25.76"
"zod": "^4.4.3"
},
"devDependencies": {
"@babel/types": "^7.29.0",
@@ -59,7 +59,7 @@
"@testing-library/jest-dom": "^6.9.1",
"@testing-library/react": "^16.3.2",
"@testing-library/user-event": "^14.6.1",
"@types/dompurify": "^3.0.5",
"@types/dompurify": "^3.2.0",
"@types/node": "^25.6.0",
"@types/react": "^19.2.14",
"@types/react-dom": "^19.2.3",
@@ -70,7 +70,7 @@
"jsdom": "^29.1.1",
"tree-sitter-wasms": "^0.1.13",
"typescript": "^5.4.5",
"vite": "^8.0.10",
"vite": "^8.0.11",
"vitest": "^4.1.5",
"wait-on": "^9.0.5"
}
+96 -4
View File
@@ -9,6 +9,7 @@
import { useState, useRef, useEffect, useId } from 'react';
import {
Github,
Gitlab,
FolderOpen,
Loader2,
Check,
@@ -26,15 +27,20 @@ import { AnalyzeProgress } from './AnalyzeProgress';
// ── Helpers ──────────────────────────────────────────────────────────────────
type InputMode = 'github' | 'local';
type InputMode = 'github' | 'gitlab' | 'local';
const GITHUB_RE = /^https?:\/\/(www\.)?github\.com\/[^/\s]+\/[^/\s]+/i;
const GITLAB_RE = /^https?:\/\/[^/\s]+\/[^/\s]+\/[^/\s]+(\/.*)?$/i;
const IS_WINDOWS = navigator.userAgent.toLowerCase().includes('win');
function isValidGithubUrl(value: string): boolean {
return GITHUB_RE.test(value.trim());
}
function isValidGitlabUrl(value: string): boolean {
return GITLAB_RE.test(value.trim());
}
// ── Mode tabs ────────────────────────────────────────────────────────────────
function ModeTabs({ mode, onChange }: { mode: InputMode; onChange: (m: InputMode) => void }) {
@@ -53,6 +59,19 @@ function ModeTabs({ mode, onChange }: { mode: InputMode; onChange: (m: InputMode
<Github className="h-3 w-3" />
GitHub URL
</button>
<button
role="tab"
aria-selected={mode === 'gitlab'}
onClick={() => onChange('gitlab')}
className={`flex flex-1 cursor-pointer items-center justify-center gap-1.5 rounded-md px-3 py-1.5 text-xs font-medium transition-all duration-150 ${
mode === 'gitlab'
? 'bg-accent text-white shadow-sm'
: 'text-text-muted hover:text-text-secondary'
} `}
>
<Gitlab className="h-3 w-3" />
GitLab URL
</button>
<button
role="tab"
aria-selected={mode === 'local'}
@@ -138,6 +157,7 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
const folderInputRef = useRef<HTMLInputElement>(null);
const [mode, setMode] = useState<InputMode>('github');
const [githubUrl, setGithubUrl] = useState('');
const [gitlabUrl, setGitlabUrl] = useState('');
const [localPath, setLocalPath] = useState('');
const [phase, setPhase] = useState<InternalPhase>('input');
const [validationError, setValidationError] = useState<string | null>(null);
@@ -162,6 +182,7 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
const handleModeChange = (m: InputMode) => {
setMode(m);
setGithubUrl('');
setGitlabUrl('');
setLocalPath('');
setValidationError(null);
};
@@ -175,13 +196,19 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
const canSubmit =
mode === 'github'
? isValidGithubUrl(githubUrl) && (phase === 'input' || phase === 'error')
: localPath.trim().length > 1 && (phase === 'input' || phase === 'error');
: mode === 'gitlab'
? isValidGitlabUrl(gitlabUrl) && (phase === 'input' || phase === 'error')
: localPath.trim().length > 1 && (phase === 'input' || phase === 'error');
const handleAnalyze = async () => {
if (mode === 'github' && !isValidGithubUrl(githubUrl)) {
setValidationError('Please enter a valid GitHub repository URL.');
return;
}
if (mode === 'gitlab' && !isValidGitlabUrl(gitlabUrl)) {
setValidationError('Please enter a valid GitLab repository URL.');
return;
}
if (mode === 'local' && localPath.trim().length < 2) {
setValidationError('Please enter a folder path.');
return;
@@ -191,12 +218,22 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
setPhase('starting');
try {
const request = mode === 'github' ? { url: githubUrl.trim() } : { path: localPath.trim() };
const request =
mode === 'github'
? { url: githubUrl.trim() }
: mode === 'gitlab'
? { url: gitlabUrl.trim() }
: { path: localPath.trim() };
const { jobId } = await startAnalyze(request);
jobIdRef.current = jobId;
setPhase('analyzing');
const nameSource = mode === 'github' ? githubUrl.trim() : localPath.trim();
const nameSource =
mode === 'github'
? githubUrl.trim()
: mode === 'gitlab'
? gitlabUrl.trim()
: localPath.trim();
const controller = streamAnalyzeProgress(
jobId,
(p) => setProgress(p),
@@ -297,6 +334,61 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
</div>
)}
{/* GitLab URL input */}
{showInput && mode === 'gitlab' && (
<div className="space-y-2">
<label
htmlFor={inputId}
className="block text-xs font-medium tracking-wider text-text-secondary uppercase"
>
GitLab Repository URL
</label>
<div
className={`flex items-center gap-3 rounded-xl border bg-void px-4 py-3.5 transition-all duration-200 ${
validationError && phase === 'error'
? 'border-red-500/50'
: isValidGitlabUrl(gitlabUrl)
? 'border-accent/50 shadow-[0_0_0_3px_rgba(124,58,237,0.08)]'
: 'border-border-default focus-within:border-accent/40'
} `}
>
<Gitlab className="h-4 w-4 shrink-0 text-text-muted" />
<input
id={inputId}
type="url"
value={gitlabUrl}
onChange={(e) => {
setGitlabUrl(e.target.value);
if (validationError) setValidationError(null);
}}
onKeyDown={(e) => {
if (e.key === 'Enter' && canSubmit && !isLoading) {
e.preventDefault();
handleAnalyze();
}
}}
disabled={isLoading}
placeholder="https://gitlab.com/owner/repo"
autoComplete="url"
spellCheck={false}
className="flex-1 border-none bg-transparent font-mono text-sm text-text-primary outline-none placeholder:text-text-muted disabled:opacity-50"
/>
{gitlabUrl.length > 10 && (
<div className="shrink-0">
{isValidGitlabUrl(gitlabUrl) ? (
<Check className="h-3.5 w-3.5 text-emerald-400" />
) : (
<AlertCircle className="h-3.5 w-3.5 text-text-muted" />
)}
</div>
)}
</div>
<p className="text-xs text-text-muted">
Supports GitLab.com and self-hosted GitLab instances.
</p>
</div>
)}
{/* Local folder input */}
{showInput && mode === 'local' && (
<div className="space-y-2">
@@ -20,6 +20,7 @@ import {
ProviderConfig,
} from './types';
import { DEFAULT_OPENROUTER_BASE_URL, DEFAULT_OLLAMA_BASE_URL } from '../../config/ui-constants';
import { resilientFetch } from 'gitnexus-shared';
const STORAGE_KEY = 'gitnexus-llm-settings';
@@ -407,7 +408,10 @@ export const getAvailableModels = (provider: LLMProvider): string[] => {
*/
export const fetchOpenRouterModels = async (): Promise<Array<{ id: string; name: string }>> => {
try {
const response = await fetch(`${DEFAULT_OPENROUTER_BASE_URL}/models`);
const response = await resilientFetch(`${DEFAULT_OPENROUTER_BASE_URL}/models`, undefined, {
breakerKey: 'openrouter-models',
retry: { maxAttempts: 2, baseDelayMs: 500, capDelayMs: 2_000 },
});
if (!response.ok) throw new Error('Failed to fetch models');
const data = await response.json();
return data.data.map((model: any) => ({
+61
View File
@@ -123,6 +123,67 @@ export {
* defaults to `currentColor`, so Tailwind `text-*` utilities work the same as
* with any other icon in this module.
*/
/**
* GitLab tanuki mark — SVG path data from simple-icons (CC0-1.0).
*
* GitLab's logo (the tanuki/fox-head) is a registered trademark of GitLab Inc.
* We use it here only to indicate GitLab source-repo integration.
*
* API-compatible with `lucide-react` icons (`LucideProps`).
*/
export const Gitlab = forwardRef<SVGSVGElement, LucideProps>(function Gitlab(
{
size = 24,
color = 'currentColor',
className,
strokeWidth: _strokeWidth,
absoluteStrokeWidth: _absoluteStrokeWidth,
...rest
},
ref,
) {
const numericSize = typeof size === 'string' ? Number.parseFloat(size) : size;
const useSmallVariant = Number.isFinite(numericSize) && (numericSize as number) <= 16;
if (useSmallVariant) {
return (
<svg
ref={ref}
xmlns="http://www.w3.org/2000/svg"
width={size}
height={size}
viewBox="0 0 16 16"
fill={color}
className={className}
{...rest}
>
<path d="M8 15.282l1.855-5.717H6.145L8 15.282z" />
<path d="M8 15.282L6.145 9.565H2.333L8 15.282z" />
<path d="M2.333 9.565l-.944-2.942c-.09-.267.067-.553.333-.553h3.153L2.333 9.565z" />
<path d="M4.875 6.07L6.145 9.565H2.333l2.542-3.495z" />
<path d="M13.667 9.565l.944-2.942c.09-.267-.067-.553-.333-.553h-3.153l2.542 3.495z" />
<path d="M11.125 6.07L9.855 9.565h3.812l-2.542-3.495z" />
<path d="M8 15.282l1.855-5.717H6.145L8 15.282z" />
</svg>
);
}
return (
<svg
ref={ref}
xmlns="http://www.w3.org/2000/svg"
width={size}
height={size}
viewBox="0 0 24 24"
fill={color}
className={className}
{...rest}
>
<path d="m23.6004 9.5927-.0337-.0862L20.3.9814a.851.851 0 0 0-.3362-.405.8748.8748 0 0 0-.9997.0539.8748.8748 0 0 0-.29.4399l-2.2055 6.748H7.5375l-2.2057-6.748a.8573.8573 0 0 0-.29-.4412.8748.8748 0 0 0-.9997-.0537.8585.8585 0 0 0-.3362.4049L.4332 9.5015l-.0325.0862a6.0657 6.0657 0 0 0 2.0119 7.0105l.0113.0087.03.0213 4.976 3.7264 2.462 1.8633 1.4995 1.1321a1.0085 1.0085 0 0 0 1.2197 0l1.4995-1.1321 2.4619-1.8633 5.006-3.7489.0125-.01a6.0682 6.0682 0 0 0 2.0094-7.003z" />
</svg>
);
});
export const Github = forwardRef<SVGSVGElement, LucideProps>(function Github(
{
size = 24,
+98 -13
View File
@@ -7,6 +7,7 @@
*/
import type { GraphNode, GraphRelationship } from 'gitnexus-shared';
import { CircuitOpenError, ResilientFetchExhaustedError, resilientFetch } from 'gitnexus-shared';
// ── Types ──────────────────────────────────────────────────────────────────
@@ -204,8 +205,32 @@ export function streamSSE<T = unknown>(url: string, handlers: SSEHandlers<T>): A
let _backendUrl = 'http://localhost:4747';
/**
* Validate that a backend URL is a safe http:// or https:// origin before
* storing it as the fetch target base (CodeQL js/client-side-request-forgery).
*
* Throws if the URL uses a non-HTTP scheme (e.g. javascript:, data:, file://).
* All other well-formed http/https URLs are accepted — the client intentionally
* supports connecting to remote GitNexus servers, not just localhost.
*/
export function validateBackendUrl(url: string): void {
let parsed: URL;
try {
parsed = new URL(url);
} catch {
// Do not echo raw input — it may contain credentials.
throw new Error('Invalid backend URL: must be a well-formed http:// or https:// URL');
}
if (parsed.protocol !== 'http:' && parsed.protocol !== 'https:') {
// Use parsed.protocol only (scheme), not the full URL, to avoid leaking credentials.
throw new Error(`Backend URL must use http:// or https:// (got ${parsed.protocol})`);
}
}
export const setBackendUrl = (url: string): void => {
_backendUrl = url.replace(/\/$/, '');
const trimmed = url.replace(/\/$/, '');
validateBackendUrl(trimmed);
_backendUrl = trimmed;
};
export const getBackendUrl = (): string => _backendUrl;
@@ -237,29 +262,91 @@ export function normalizeServerUrl(input: string): string {
const DEFAULT_TIMEOUT_MS = 30_000;
const PROBE_TIMEOUT_MS = 2_000;
/** Idempotent HTTP methods. Other verbs (POST, PATCH, PUT, DELETE) get
* a single-attempt retry budget by default to avoid duplicate side
* effects on retry — a POST that 5xx'd may have already executed
* server-side. Callers that have idempotency keys or otherwise know
* their mutation is safe to retry can opt in via `forceRetry`. */
const IDEMPOTENT_METHODS = new Set(['GET', 'HEAD', 'OPTIONS']);
const fetchWithTimeout = async (
url: string,
init: RequestInit = {},
timeoutMs: number = DEFAULT_TIMEOUT_MS,
/**
* Force a retry budget on non-idempotent methods. Default false.
* Pass true only when the endpoint is known-idempotent (e.g. DELETE
* of a known-deleted resource — second call is a 404 / no-op) AND
* the duplicate-side-effect window is acceptable.
*/
forceRetry = false,
): Promise<Response> => {
const controller = new AbortController();
// Merge external signal if provided
// Merge the external caller signal (if any) with an
// `AbortSignal.timeout()` so a timer-fired abort produces a
// `DOMException` with `name === 'TimeoutError'` — which
// `resilientFetch` correctly classifies as terminal-network (no
// retry, no breaker hit). A manual `AbortController.abort()` would
// produce `name === 'AbortError'` and route through the
// retryable-network branch, which mis-penalizes the breaker for
// user-side network slowness.
const timeoutSignal = AbortSignal.timeout(timeoutMs);
const externalSignal = init.signal;
if (externalSignal) {
externalSignal.addEventListener('abort', () => controller.abort());
const signal = externalSignal ? AbortSignal.any([timeoutSignal, externalSignal]) : timeoutSignal;
const method = (init.method ?? 'GET').toUpperCase();
const isIdempotent = IDEMPOTENT_METHODS.has(method);
const maxAttempts = isIdempotent || forceRetry ? 2 : 1;
// Key the breaker by the current backend origin so switching backend
// URLs (e.g. recovering from a flapping local server by pointing at
// a different host) gives the new origin a fresh breaker state. A
// single shared `'web-backend'` key would otherwise leave a user
// locked out for the full cooldown after one bad host trips the
// circuit. The malformed-URL fallback is defensive — `setBackendUrl`
// normalizes input, so this branch shouldn't fire in practice.
let breakerKey: string;
try {
breakerKey = `web-backend:${new URL(_backendUrl).origin}`;
} catch {
breakerKey = 'web-backend:invalid';
}
const timer = setTimeout(() => controller.abort(), timeoutMs);
try {
const response = await fetch(url, { ...init, signal: controller.signal });
// Bounded retries + 5xx/429 handling are delegated to resilientFetch.
// Method-aware budget: idempotent verbs retry once on transient
// backend failures; mutations (POST/PATCH/PUT/DELETE) default to
// single-attempt to avoid duplicate side effects.
const response = await resilientFetch(
url,
{ ...init, signal },
{
breakerKey,
retry: { maxAttempts, baseDelayMs: 250, capDelayMs: 1500 },
},
);
return response;
} catch (error: unknown) {
if (error instanceof DOMException && error.name === 'AbortError') {
if (externalSignal?.aborted) {
throw new BackendError('Request aborted', 0, 'network');
}
if (error instanceof CircuitOpenError) {
throw new BackendError(
`GitNexus backend at ${_backendUrl} is unhealthy; retry in ${Math.ceil(error.retryAfterMs / 1000)}s`,
0,
'network',
);
}
if (error instanceof ResilientFetchExhaustedError) {
// Fall through to caller — surface the raw response so assertOk
// can craft the BackendError with the right code.
return error.response;
}
if (error instanceof DOMException && error.name === 'TimeoutError') {
throw new BackendError(`Request to ${url} timed out after ${timeoutMs}ms`, 0, 'timeout');
}
if (error instanceof DOMException && error.name === 'AbortError') {
// External caller-driven cancellation — `timeoutSignal` would
// have surfaced as TimeoutError above, so this branch covers
// only the externally-aborted case.
throw new BackendError('Request aborted', 0, 'network');
}
if (error instanceof TypeError) {
throw new BackendError(
`Network error reaching GitNexus backend at ${_backendUrl}: ${error.message}`,
@@ -268,8 +355,6 @@ const fetchWithTimeout = async (
);
}
throw error;
} finally {
clearTimeout(timer);
}
};
@@ -0,0 +1,110 @@
/**
* Method-aware retry budget + timeout-as-TimeoutError verification for
* backend-client's `fetchWithTimeout`.
*
* Closes review findings on PR #1448:
* - Non-idempotent POST/DELETE must NOT be retried by default —
* a 5xx on `startAnalyze` could otherwise start a duplicate job.
* - Timer-fired timeout must surface as `DOMException(name='TimeoutError')`,
* not `AbortError`, so resilientFetch routes it through the
* terminal-network branch (no retry, no breaker hit).
*/
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest';
import { getBreaker } from 'gitnexus-shared';
import { __resetBreakerRegistry__ } from 'gitnexus-shared/test-helpers';
import { fetchRepos, setBackendUrl, startAnalyze } from '../../src/services/backend-client';
const BASE = 'http://localhost:4747';
describe('backend-client retry budget (method-aware)', () => {
beforeEach(() => {
__resetBreakerRegistry__();
setBackendUrl(BASE);
});
afterEach(() => {
vi.unstubAllGlobals();
});
it('GET retries once on transient 503 (idempotent verb)', async () => {
let n = 0;
const fetchMock = vi.fn(async () => {
n += 1;
if (n === 1) return new Response('boom', { status: 503 });
return new Response('[]', {
status: 200,
headers: { 'Content-Type': 'application/json' },
});
});
vi.stubGlobal('fetch', fetchMock);
const repos = await fetchRepos();
expect(repos).toEqual([]);
// 1 retry budget on idempotent GET → 2 total fetch calls.
expect(fetchMock).toHaveBeenCalledTimes(2);
});
it('POST does NOT retry on 503 by default (non-idempotent verb)', async () => {
const fetchMock = vi.fn(async () => new Response('boom', { status: 503 }));
vi.stubGlobal('fetch', fetchMock);
await expect(startAnalyze({ path: '/tmp/repo' })).rejects.toBeTruthy();
// Single attempt — never duplicates a job-start POST.
expect(fetchMock).toHaveBeenCalledTimes(1);
});
it('switching backend URL after a circuit opens reaches a fresh breaker (U3)', async () => {
// Pre-open the breaker for host-A by directly recording 3 failures.
setBackendUrl('http://host-a.test:4747');
const aKey = 'web-backend:http://host-a.test:4747';
const breakerA = getBreaker(aKey);
breakerA.recordFailure();
breakerA.recordFailure();
breakerA.recordFailure();
expect(breakerA.getState()).toBe('open');
// Switch to host-B and make a request — must succeed against the
// new origin without tripping the host-A circuit. Under the old
// single-key behaviour the call would throw CircuitOpenError.
setBackendUrl('http://host-b.test:4747');
const fetchMock = vi.fn(
async () =>
new Response('[]', {
status: 200,
headers: { 'Content-Type': 'application/json' },
}),
);
vi.stubGlobal('fetch', fetchMock);
const repos = await fetchRepos();
expect(repos).toEqual([]);
expect(fetchMock).toHaveBeenCalledTimes(1);
// Host-A's breaker is still open in cooldown.
expect(breakerA.getState()).toBe('open');
// Host-B has its own (fresh) breaker.
const bKey = 'web-backend:http://host-b.test:4747';
expect(getBreaker(bKey).getState()).toBe('closed');
expect(getBreaker(bKey).getConsecutiveFailures()).toBe(0);
});
it('breaker not incremented when timeout fires (TimeoutError, not AbortError)', async () => {
// Reject directly with a TimeoutError DOMException, mimicking what
// `fetch` produces when its `AbortSignal.timeout()`-wired signal
// fires. The real-fetch path goes signal.reason → reject(reason);
// we shortcut that here so the test doesn't have to wait the
// 30-second default timeout.
const fetchMock = vi.fn(async () => {
throw new DOMException('aborted by timeout', 'TimeoutError');
});
vi.stubGlobal('fetch', fetchMock);
await expect(fetchRepos()).rejects.toMatchObject({ code: 'timeout' });
// The breaker must not have been penalized for a local timeout.
expect(getBreaker(`web-backend:${BASE}`).getConsecutiveFailures()).toBe(0);
// Timeout is terminal — no retry attempted.
expect(fetchMock).toHaveBeenCalledTimes(1);
});
});
@@ -1,5 +1,11 @@
import { afterEach, describe, expect, it, vi } from 'vitest';
import { fetchGraph, normalizeServerUrl, setBackendUrl } from '../../src/services/backend-client';
import {
fetchGraph,
getBackendUrl,
normalizeServerUrl,
setBackendUrl,
validateBackendUrl,
} from '../../src/services/backend-client';
describe('normalizeServerUrl', () => {
it('adds http:// to localhost', () => {
@@ -165,3 +171,61 @@ describe('fetchGraph', () => {
});
});
});
describe('validateBackendUrl', () => {
it('allows http:// URLs', () => {
expect(() => validateBackendUrl('http://localhost:4747')).not.toThrow();
expect(() => validateBackendUrl('http://127.0.0.1:4747')).not.toThrow();
});
it('allows https:// URLs', () => {
expect(() => validateBackendUrl('https://gitnexus.example.com')).not.toThrow();
expect(() => validateBackendUrl('https://my-server.internal:4747')).not.toThrow();
});
it('rejects non-http schemes', () => {
expect(() => validateBackendUrl('javascript:alert(1)')).toThrow('must use http:// or https://');
expect(() => validateBackendUrl('file:///etc/passwd')).toThrow('must use http:// or https://');
expect(() => validateBackendUrl('data:text/plain,evil')).toThrow(
'must use http:// or https://',
);
});
it('rejects malformed URLs', () => {
expect(() => validateBackendUrl('not-a-url')).toThrow('Invalid backend URL');
});
it('does not include the raw URL in error messages (credential hygiene)', () => {
const urlWithCreds = 'javascript:alert("sk-secret")';
let msg = '';
try {
validateBackendUrl(urlWithCreds);
} catch (e) {
msg = (e as Error).message;
}
expect(msg).not.toContain('sk-secret');
expect(msg).not.toContain(urlWithCreds);
});
});
describe('setBackendUrl', () => {
it('accepts valid http URLs', () => {
expect(() => setBackendUrl('http://localhost:4747')).not.toThrow();
});
it('accepts valid https URLs', () => {
expect(() => setBackendUrl('https://my-server.example.com')).not.toThrow();
});
it('rejects non-http/https schemes', () => {
expect(() => setBackendUrl('javascript:alert(1)')).toThrow('must use http:// or https://');
expect(() => setBackendUrl('file:///etc/passwd')).toThrow('must use http:// or https://');
});
it('does not mutate _backendUrl when validation fails', () => {
setBackendUrl('http://localhost:4747');
expect(() => setBackendUrl('javascript:alert(1)')).toThrow();
// State must be preserved — validation must happen before the assignment
expect(getBackendUrl()).toBe('http://localhost:4747');
});
});
+1 -1
View File
@@ -1,3 +1,3 @@
{
"installCommand": "cd ../gitnexus-shared && npm install && npm run build && cd ../gitnexus-web && npm install"
"installCommand": "cd ../gitnexus-shared && npm install && npm run build && cd ../gitnexus-web && npm ci --include=dev"
}
+15 -2
View File
@@ -1,5 +1,18 @@
{
"permissions": {
"allow": ["mcp__plugin_claude-mem_mcp-search__get_observations"]
}
"allow": [
"mcp__plugin_claude-mem_mcp-search__get_observations",
"Skill(gitnexus-exploring)",
"Bash(npx gitnexus *)",
"mcp__obsidian-memory__search_nodes",
"mcp__obsidian-memory__add_observations",
"WebSearch",
"WebFetch(domain:cppreference.net)",
"Bash(xargs grep -l \"templateArguments\\\\|parameterTypes\")",
"Bash(gh issue *)",
"Bash(gh pr *)"
]
},
"enableAllProjectMcpServers": true,
"enabledMcpjsonServers": ["gitnexus"]
}
+121
View File
@@ -4,6 +4,127 @@ All notable changes to GitNexus will be documented in this file.
## [Unreleased]
## [1.6.5] - 2026-05-16
### Added
- **C++ ADL V2** — Argument-Dependent Lookup overhaul. Class-typed reference args (incl. rvalue refs) contribute associated namespaces (#1595); class-pointer args and template-specialization args (with nested template args) included (#1592, #1596); base-class associated namespaces walked via MRO (#1597); free-function reference args contribute enclosing namespace (#1598); ordinary and ADL free-call candidates merged before overload selection (#1599)
- **C++ standard-conversion-sequence ranking** for overload resolution (#1606)
- **C++ scope-resolution migration** — C++ now runs on the registry-primary RFC #909 path (#938, #1520); template-body `this->` + `using ns::name` calls resolved in the scope resolver (#1590); template specializations disambiguated in class graph IDs and receiver routing (#1587); EXTENDS edges for template and qualified template bases (#1581)
- **PHP scope-resolution migration** — PHP moved to scope-based resolution (#938, #1497, supersedes #1124)
- **Java scope-resolution migration** — RFC #909 Ring 3 (#1482)
- **C scope-resolution migration** — RFC #909 Ring 3 (#1481)
- **Incremental indexing** — `gitnexus analyze` now reuses a parse cache, writes back to DB, and short-circuits scope resolution when nothing changed (#1479)
- **`gitnexus:keep` marker** — preserves custom context sections (#605, #1508)
- **`gitnexus analyze --skip-skills` and `--index-only`** flags (#742, #1485)
- **`gitnexus wiki --timeout` and `--retries` flags** — mitigate timeout aborts on large module pages (#1543)
- **HTTP embedding `dimensions` parameter** — now forwarded to the embedding endpoint (#1498)
- **Cursor 2.4 `postToolUse` hooks** — upgraded for Read/Grep/Shell coverage (#1467)
### Fixed
- **Cross-file type propagation** — resolved a stall on large repos (#1626)
- **C++ inline-namespace ambiguity** — detect same-name ambiguity across inline namespace children (#1564, #1600); workspace-wide dependent-base name resolution for cross-file templates (#1586)
- **Parse cache persistence** — sharded on large repos to avoid corruption (#1580)
- **TypeScript ESM `.js` extension** — fallback applied to tsconfig path-alias resolution (#1530) and `.js` → `.ts` source resolution (#1525)
- **Markdown CRLF line endings** — section heading parser now handles them (#1469)
- **`gitnexus analyze --no-stats`** — actually omits volatile counts (#1477, #1478)
- **`ensureGitNexusIgnored`** — tolerate read-only workspaces (#1549, #1550)
- **Claude augment hook** — skipped when GitNexus server owns the DB (#1493)
- **Docker runtime image** — symlink `gitnexus` binary onto `$PATH` (#1551); install `ca-certificates` for TLS verification (#1545, #1547); include duckdb installer script (#1502)
- **Windows reliability** — fix 32767-char tree-sitter crash and VECTOR-extension SIGSEGV (#1433); platform-aware `tsc` build command for win32 (#1531)
- **Search / FTS** — guard against undefined `bm25Results` when FTS is unavailable (#1489, #1540); CONTAINS fallback in augment when FTS indexes unavailable (#1476)
- **Wiki** — sanitize generated mermaid diagrams (#1539)
- **Hooks** — cap concurrent augment subprocesses to prevent runaway fan-out (#1486, #1510)
- **LadybugDB** — drain checkpoint result before close (#1506); recover `gitnexus analyze` from orphan sidecars when the main DB file is missing (#1622)
- **Group / contracts** — detect `httpx` async consumers (#1408)
- **Server hardening** — sanitize repo name to prevent argument injection on `/api/analyze` (#1305)
### Changed
- **CI release pipeline unified under `publish.yml`** — single source of truth for npm publish, provenance, and GitHub Release creation (#1610)
- **CI: skip RC build on release PRs** — release/* branches no longer cut redundant RCs (#1474)
- **CI (Claude review): make `/review` reliably post PR comments** (#1522); allow Bash in code-review job without interactive approval (#1523)
- **CI publish (post-merge fixes)** — bump publish job to Node 24 for npm OIDC support (#1628); engage npm Trusted Publishing OIDC properly (#1627)
- **Tests** — remove flaky regression test for resource exhaustion (#1521); de-flake regex linearity assertions in U8 (#1475)
### Chore / Dependencies
- `vitest` 4.1.5 → 4.1.6 in /gitnexus (#1605)
- `@langchain/google-genai` bump in /gitnexus-web (#1554)
- `vite` 8.0.10 → 8.0.11 in /gitnexus-web (#1555)
- `mermaid` bump (#1514)
- `protobufjs` 7.5.5 → 7.5.8 + `@protobufjs/utf8` in /gitnexus (#1535, #1536)
- `urllib3` bump in /eval uv group (#1512)
- GitHub Actions: `sigstore/cosign-installer` 4.1.1 → 4.1.2 (#1557)
## [1.6.4] - 2026-05-10
### Added
- **`gitnexus publish`** — opt-in command to push your indexed graph to the understand-quickly registry for shareable browsing (#1425)
- **`IncludeExtractor` for C++** — cross-repo include tracking joins the group contract pipeline (#1156)
- **Unreal Engine C++ support** — strips reflection macros (`UCLASS`, `UFUNCTION`, `UPROPERTY`, etc.) before tree-sitter parses, so UE projects index cleanly (#1439)
- **Thrift contracts extractor** — group-mode contract detection for Apache Thrift IDL (#1234)
- **Workspace extractors for Node, Python, Go, Java, Elixir** — group-mode auto-discovery of cross-package boundaries (#1260)
- **Rust workspace cross-crate contracts** — auto-discovery of `[workspace]` member crates and their cross-crate links (#1256)
- **Go scope-resolution hooks** — Go joins Python / C# / TypeScript on the registry-primary RFC #909 path (#1302)
- **TypeScript registry-primary scope resolution (Ring 3)** — TypeScript fully migrated to scope-based resolution (#1050)
- **Configurable group cross-link path exclusions** — reduces false-positive contract links in vendored / monorepo trees (#1093)
- **MCP tool safety annotations** — every MCP tool advertises read-only / mutating semantics so hosts can prompt appropriately (#1127)
- **`--embeddings <limit>` opt-in cap** — bound the embeddings pass on huge graphs (closes #382, #1375)
- **Pino structured logger** — replaces ad-hoc console output across the core with structured JSON logs (with pretty-print for TTY) (#1336)
- **Shared resilient-fetch helper** — single retries + circuit breaker module reused by HF / Docker / publish flows (#1448)
- **`/autofix` ChatOps button** — fork-safe PR autofix pipeline replaces the inline reviewdog flow (#1446, #1458)
- **Automated security & vulnerability scans** in CI (#1297, #1455)
### Fixed
- **FTS read-only DB cluster** — hook resolves canonical repo root and guards read-only FTS ensure; missing-FTS warning is now surfaced. Closes #1255, #1287, #1170, #1449, #1440, #1216, #1438 (#1226, #1418, #1107, #1123)
- **WAL corruption recovery** — quarantine corrupted `.wal` files instead of failing analyze; CHECKPOINT before close prevents recurrence; `safeClose` consolidates flush. Closes #1402, #1236, #1273, #1361 (#1417, #1314, #1377)
- **Embedding download failures** — actionable HF_ENDPOINT guidance, retries, timeout, and circuit breaker; bridge `HF_ENDPOINT` to transformers.js; iterative DFS; HF cache via `os.homedir()`. Closes #1378, #1437, #1205 (#1419, #1252, #1078)
- **Windows reliability** — pin tree-sitter-c/cpp to fix segfault, prefer `.cmd`/`.bat` from `where` output, robust LadybugDB lock acquisition for CI integration tests, surface silent finalize-skips so analyze cannot exit 0 without persisting. Closes #1242, #1427, #1447, #1468, #1400; partial #1218 (#1243, #1299, #1430, #1237, #1226, #1235)
- **DuckDB / LadybugDB native** — bumped to 0.16.0 then 0.16.1; prevent extension install hangs; CHECKPOINT before close; WAL quarantine on corruption. Closes #1162, #1160, #273 (#1235, #1326, #1129, #1314, #1417)
- **C# scope-resolution "Cannot add property" crashes** — generic typed properties included in context and impact, fixing crashes on Unity ECS partial structs and on properties whose name matches the class name. Closes #1426, #1465 (#1399)
- **C# frozen-bucket regression** + scope-resolution I8 hardening — closes #1066 (#1082, #1085)
- **Scope resolution** — same-range Module-as-parent for top-level scopes (closes #1086) (#1087); avoid variadic reference-site aggregation (#1112); skip empty scope extraction (#1100); classify Python class methods as Method (#1102)
- **Python** — index repos with empty `__init__.py` and >32 KB files (#1163); walk ancestors for multi-segment dotted imports (#1241); deterministic multi-segment suffix fallback (#1253)
- **TypeScript** — capture missed CALLS edges from HOF callbacks and JSX (#1175); name HOC-wrapped const declarations (`forwardRef` / `memo` / `useCallback` / `useMemo` / `observer`) (#1261); pair-with-arrow `@declaration.function` anchored on inner arrow
- **Go** — loose equality for `Array.find()` null checks (#1384)
- **Swift** — switched to the official prebuilt parser runtime (#1130)
- **Server hardening cluster (U2–U8)** — JS path-injection on `/api/file` + docker-server (U2, #1322); git-clone path/CLI-injection / ReDoS hardening (U3, #1325); per-route rate limiting on FS-touching endpoints (U4, #1327); URL/regex/tag-filter sanitization (U7, #1330); ReDoS in cobol-preprocessor + rust-workspace + cross-impact resource exhaustion (U8, #1331); critical type-confusion + validation helper (#1317); rate-limit `/api/analyze` and `/api/embed` (closes #1328, #1339); IPv6 ipKeyGenerator (closes #1360, #1374); IPv4-compatible IPv6 / NAT64 SSRF bypasses in `validateGitUrl` (closes #1148, 95814847); predictable tempfile names → `crypto.randomBytes` (#1387); log-injection / http-to-file-access / client-side request forgery (#1456); pin Docker Node base images + Trivy verification + Dependabot policy (#1455)
- **Group / contracts** — `runExactMatch` honours `.gitnexusignore` via shared `IgnoreService` (closes #1185, #1247); custom manifest links resolved against graph symbols (#1254); `IgnoreService` EACCES test under uid=0 (#1108)
- **MCP** — close MCP server timeout via stdout discipline + cold-start friction (#1383); avoid `git` from non-repo cwd in sibling-cwd match (closes #1138, #1293); start MCP bridge correctly when using `npx` (#1114); project `tool_map` flows from handlers (#1113); parallelize staleness checks in `list_repos` (#1416)
- **Storage / CLI** — derive registry name from canonical repo root, not worktree slug (closes #1259, #1296); `--skip-git` treats cwd as index root (#1245); keep GitNexus ignores inside `.gitnexus/` (#1248); surface silent finalize-skips so `analyze` cannot exit 0 without persisting (closes #1169, #1237); ignore global registry during staleness checks (#1141); use `os.homedir()` instead of `process.env.HOME` for HF cache dir (#1078); correct OpenCode skills install path in status message (#1386)
- **Docker / server** — dedicated health endpoint for container healthcheck (closes #1147, #1355); HEAD probe so SSE heartbeat doesn't time out healthcheck (#1182); flush WAL after `/api/embed` so search sees new embeddings (closes #1149, #1359); platform-aware semantic fallback (#1150); skip vector index query on unsupported platforms (closes #1178, #1181); serve web UI at root path instead of 404 (#1048)
- **Worker pool** — wait for replacement worker online before dispatch (#1324); prevent premature pool resolution in worker split-and-retry path (#1321); recover worker parse stalls (#1121); widened CI flake-tolerant timeouts (#1323, #1347, #1354)
- **Embeddings storage** — CHECKPOINT before closing DB to prevent WAL corruption (#1314)
- **Performance** — replace O(n³) C3 merge loop with O(n²) head-pointer algorithm (#1316)
- **Install** — vendor tree-sitter-dart source (#1125)
- **Git utils** — suppress stderr leak in `getCurrentCommit` and `getGitRoot` (closes #1172, #1341)
- **Search** — load FTS during core DB init (#1123); create FTS indexes during `analyze` (#1107); surface warning when FTS indexes are missing (#1418)
- **Hooks** — clarify `PostToolUse` hook is notification-only, not auto-reindex (#1070)
- **Docs** — README Web UI section corrected (closes #1110, #1159, #2ff3e64f); Goliath capitalisation typo (#1126)
- **CI** — fork-safe PR autofix pipeline (#1446); consolidated Claude review workflow (#1258); fine-grained PAT for RC tag push (#1407); handle expired artifacts in base coverage fetch (#1410, #1412); allow expected legacy parity failures (#1099); avoid duplicate main push checks; isolate native LadybugDB / CLI e2e flakes; seed e2e with a small fixture repo (#1249); configure e2e GitNexus home at runtime; widen rate-limit test window for Windows CI (#1347)
### Changed
- **`gitnexus publish` artefact contract** — universal opt-in publish format introduced (#1425, #1458)
- **Refactor: per-language patterns consolidated into `LanguageProvider`** (#1279)
- **Refactor: `safeClose` helper** consolidates WAL flush across LadybugDB call sites (#1377)
- **Quality: exclude `test/fixtures` from CodeQL, ESLint, and Prettier** (#1313)
- **Regression coverage** for `.gitnexusignore` behaviour with `--skip-git` (#1450)
### Chore / Dependencies
- `@ladybugdb/core` 0.16.0 → 0.16.1 (#1235, #1326)
- `@anthropic-ai/sdk` (#1442), `@langchain/anthropic` (#1389), `@langchain/core` (#1394), `@langchain/openai` (#1215)
- `hono` 4.12.9 → 4.12.18 + `@hono/node-server` (#1310, #1311, #1443)
- `axios` (#1345), `fast-uri` 3.1.0 → 3.1.2 (#1441), `lru-cache` 11.3.5 → 11.3.6 (#1344), `mnemonist` 0.40.3 → 0.40.4 (#1239), `express-rate-limit` (#1343, #1397), `onnxruntime-node` (#1213, #1435), `uuid` 13 → 14 in /gitnexus-web (#1211, after revert #1222 / re-land #1250 + #1208)
- `react`/`@types/react` (#1210), `react-dom` 19.2.5 → 19.2.6 (#1396), `react-zoom-pan-pinch` (#1214), `jsdom` 29.0.2 → 29.1.1 (#1395)
- npm_and_yarn group bump (#1312), uv group bump (#1315), `python-dotenv` (#1320), `@types/node` (#1212, #1421, #1436)
- GitHub Actions: `docker/build-push-action` 6.19.2 → 7.1.0 (#1391), `github/codeql-action` 3.35.3 → 4.35.3 (#1390)
## [1.6.3] - 2026-04-24
### Added
+11 -2
View File
@@ -1,6 +1,15 @@
FROM node:20-bookworm
# Pinned npm version — keep in sync with the root Dockerfile.cli and
# Dockerfile.web.
ARG NPM_VERSION=11.14.1
# node:22-bookworm-slim
FROM node:22-bookworm-slim@sha256:9f6d5975c7dca860947d3915877f85607946403fc55349f39b4bc3688448bb6e
ARG NPM_VERSION
WORKDIR /app
RUN apt-get -o Acquire::Check-Valid-Until=false -o Acquire::Check-Date=false update && apt-get install -y python3 make g++ && rm -rf /var/lib/apt/lists/*
RUN npx --yes npm@${NPM_VERSION} install -g npm@${NPM_VERSION} \
&& apt-get -o Acquire::Check-Valid-Until=false -o Acquire::Check-Date=false update \
&& apt-get install -y python3 make g++ \
&& rm -rf /var/lib/apt/lists/*
COPY . .
RUN npm ci --ignore-scripts \
&& npm rebuild tree-sitter-swift 2>&1 \
+13 -2
View File
@@ -33,7 +33,7 @@ To configure MCP for your editor, run `npx gitnexus setup` once — or set it up
| Editor | MCP | Skills | Hooks (auto-augment) | Support |
|--------|-----|--------|---------------------|---------|
| **Claude Code** | Yes | Yes | Yes (PreToolUse) | **Full** |
| **Cursor** | Yes | Yes | — | MCP + Skills |
| **Cursor** | Yes | Yes | Yes (postToolUse, [manual install](../gitnexus-cursor-integration/README.md#hook-install)) | **Full** |
| **Codex** | Yes | Yes | — | MCP + Skills |
| **Windsurf** | Yes | — | — | MCP |
| **OpenCode** | Yes | Yes | — | MCP + Skills |
@@ -151,7 +151,8 @@ Your AI agent gets these tools automatically:
```bash
gitnexus setup # Configure MCP for your editors (one-time)
gitnexus analyze [path] # Index a repository (or update stale index)
gitnexus analyze --force # Force full re-index
gitnexus analyze --repair-fts # Fast path: rebuild/verify only FTS indexes on existing index data
gitnexus analyze --force # Full rebuild: re-parse + graph rebuild + FTS rebuild
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
gitnexus analyze --skip-agents-md # Preserve custom AGENTS.md/CLAUDE.md gitnexus section edits
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
@@ -358,6 +359,16 @@ npx gitnexus analyze
For repositories with very large source files, `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` controls the worker job byte budget. The default is **8388608 bytes (8 MB)**.
### Worker pool resilience tuning
Three env vars expose the pool's resilience layers (respawn budget, cumulative-timeout cap, circuit breaker). Defaults are tuned for typical repos; bump them when an analyze legitimately needs more retries, or lower them to fail-fast on a known-bad shape.
| Variable | Default | Effect |
| ------------------------------------------------- | ------------------------- | --------------------------------------------------------------------------------------------------------------------------------- |
| `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT` | `3` | Max replacement spawns per slot before the slot is dropped from the active rotation. |
| `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` | `5 × subBatchTimeoutMs` | Total retry wall-time budget per job before quarantining. Bounds exponentially-growing retry waits. |
| `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD` | `max(3, poolSize)` | Per-slot consecutive deaths before the pool's circuit breaker trips. After tripping, dispatches require a fresh pool. |
## Privacy
- All processing happens locally on your machine
+175
View File
@@ -0,0 +1,175 @@
# Parse-throughput benchmark (scaffold)
> **Status: methodology + harness scaffold, no measurement data yet.**
> The Latest measurement table below contains `_TBD_` placeholders.
> This file ships intentionally without numbers — populating it
> requires a dedicated bench-pass against the U6 fixture (and ideally
> a real-world TS-root-scale repo) on consistent hardware, which is
> tracked as future work rather than gated on PR #1693's merge.
> Until the table is populated, the load-bearing perf-regression
> protection lives in `gitnexus/test/integration/parse-impl-large-fixture.test.ts`
> (U6, 30 s wall-clock budget via `Promise.race`).
Tracks `runChunkedParseAndResolve` wall-clock + peak heap on a synthetic
fixture so PR #1693's "analyze no longer hangs on TS-root-shaped loads"
claim is measurable, not just asserted by smoke tests. The harness
recipe below is deliberately small enough to re-run in a few minutes
when the bench-pass is undertaken.
---
## Methodology
### Fixture
Synthetic TypeScript repo, _not_ a clone of microsoft/TypeScript. CI cost
of cloning real-world repos is prohibitive; the synthetic shape exercises
the same pipeline paths (chunking, deferred extraction, cross-chunk
imports + heritage) without the disk-I/O overhead. Larger numbers can be
manually captured against real repos and cross-referenced here, but the
authoritative regression-tracking shape is the synthetic fixture so runs
are reproducible across hardware.
The fixture matches the structure pinned by
`gitnexus/test/integration/parse-impl-large-fixture.test.ts` (U6):
- 15 small modules (`mod0.ts` … `mod14.ts`), one exported function each.
- 1 dense `complex.ts` with 30 functions + 1 class + 1 interface.
- 1 `index.ts` re-exporting every symbol from every module.
`GITNEXUS_CHUNK_BYTE_BUDGET=64` forces multi-chunk parsing on this small
fixture — without that override the whole thing fits in one chunk and
the deferred-extraction path is not exercised end-to-end.
### What to measure
| Metric | How |
| --------------------------- | -------------------------------------------------------------------------------- |
| Wall-clock total | `Date.now()` delta around `runChunkedParseAndResolve` |
| Peak heap | Sample `process.memoryUsage().heapUsed` every 50 ms during the run; keep the max |
| Chunks observed | Count distinct `Parsing chunk X/Y` progress messages |
| `getStats()` final snapshot | Quarantined paths, dropped slots, breaker state |
### Hardware shape (record alongside each measurement)
- OS + version
- CPU model + logical core count
- RAM
- Node version
- gitnexus commit SHA (so the snapshot is anchored to a tree, not "main")
---
## Harness recipe
The U6 test (`test/integration/parse-impl-large-fixture.test.ts`) is the
checked-in mini-benchmark — it exercises the same fixture and bounds the
wall-clock at 30 s via `Promise.race`. To produce a richer snapshot for
this doc, run it under instrumentation:
```bash
# From the gitnexus/ subdir:
cd gitnexus
# Single-threaded baseline (sequential fallback):
npx vitest run test/integration/parse-impl-large-fixture.test.ts --reporter=verbose
# Worker-pool path (requires built dist/ — pre-built by `npm run build`):
npm run build && \
GITNEXUS_WORKER_POOL_SIZE=4 \
GITNEXUS_PARSE_CHUNK_CONCURRENCY=2 \
GITNEXUS_VERBOSE=1 \
npx vitest run test/integration/parse-impl-large-fixture.test.ts --reporter=verbose
```
For peak-heap sampling, wrap the dispatch call in a Node script that
polls `process.memoryUsage()`. A future helper at
`gitnexus/bench/scripts/parse-throughput.ts` would automate this — the
plan's stretch goal. Until that lands, capture peak heap manually via:
```bash
node --inspect=0 \
--require ./scripts/heap-sampler.js \
./node_modules/.bin/vitest run test/integration/parse-impl-large-fixture.test.ts
```
---
## Latest measurement
> _No measurement data has been collected yet — this file is the
> methodology + harness scaffold. The single recorded data point is the
> U6 wall-clock smoke baseline below; the worker-pool rows are
> placeholders for future bench-pass output._
The U6 integration test (`gitnexus/test/integration/parse-impl-large-fixture.test.ts`)
was observed completing the synthetic fixture in **~6 seconds** under
the sequential path (`skipWorkers: true`) on the development machine,
well under the 30 s `Promise.race` wall-clock budget. That number is a
smoke baseline only — recorded here for reference, not as a regression
target.
| Path | files/s | wall-clock | peak heap | chunks | quarantined |
| ------------------------------------------ | ------- | -------------------- | --------- | ------ | ----------- |
| Sequential (`skipWorkers: true`, U6 smoke) | _TBD_ | ~6 s _(observation)_ | _TBD_ | 17 | 0 |
| Worker pool, `--workers 4`, concurrency 2 | _TBD_ | _TBD_ | _TBD_ | _TBD_ | 0 |
| Worker pool, `--workers 1`, concurrency 1 | _TBD_ | _TBD_ | _TBD_ | _TBD_ | 0 |
**Hardware:** _TBD — record OS, CPU, RAM, Node version, gitnexus SHA at
the time of the bench-pass that populates the table above._
---
## Operator-tuning quick reference
Cross-links to the env vars documented in the [README](../../README.md#environment-variables).
Use this section as a starting point when the benchmark numbers above
suggest a tuning opportunity for your hardware shape.
- **CPU-bound, big repo, lots of cores:** raise `GITNEXUS_WORKER_POOL_SIZE`
past the default cap of 16. The 16-worker cap exists because past that
point main-thread merge / extraction dominates; if you've measurably
ruled that out, the env var lifts the cap explicitly. (See
`worker-pool.ts` `DEFAULT_POOL_SIZE_CAP`.)
- **Slow files (large minified JS, deep TS types):** raise
`GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS` past 30 000 ms. The cumulative
budget is 5× this value (U10 pins this) so a 60 s idle timeout permits
300 s of total retry-and-split wall-clock before quarantining the file.
- **Constrained container (cgroup CPU limit):** the pool now uses
`os.availableParallelism()` (U3 H2), which honors cgroup limits — no
manual `GITNEXUS_WORKER_POOL_SIZE` override needed unless the auto-
resolved value is too aggressive for your I/O budget.
- **Long-running host (eval-server, MCP daemon) running back-to-back
analyzes:** `--workers` is now threaded through `AnalyzeOptions`
(U2 B2), so per-invocation sizing is honored without `process.env`
state leaking across calls. `GITNEXUS_VERBOSE` is similarly snapshot/
restore-bracketed.
---
## What this benchmark does NOT measure
- **Real-repo performance.** The synthetic fixture is sized for CI; it
doesn't exercise the cumulative-load shape (50k files, occasional
pathological file) that drove the original PR #1693 hang report. Real-
repo numbers should be captured ad-hoc against the user's target repo
and cross-referenced here only as supplementary evidence.
- **Worker-pool resilience under real crashes.** That's verified by the
`worker-pool.test.ts` integration tests (real `process.exit`, real
`error` events, real protocol violations) and the unit suite. The
benchmark cares about throughput on the happy path.
- **IPC repack throughput.** Phase 3 of the PR #1693 plan introduces a
transferList + binary wire-format IPC repack (U16-U17). Once that
lands, an `IPC repack` row should be added to the "Latest measurement"
table above with before/after numbers on the same hardware.
---
## Related artifacts
- Plan: `docs/plans/2026-05-20-001-feat-pr1693-resilience-hardening-and-ipc-repack-plan.md`
- Integration test (mini-benchmark with wall-clock guard): `gitnexus/test/integration/parse-impl-large-fixture.test.ts` (U6)
- Operator env-var reference: `README.md` → Environment variables
- Resilience layer tests: `gitnexus/test/unit/worker-pool-resilience.test.ts`,
`worker-pool-cumulative-timeout.test.ts`,
`worker-pool-windows-quarantine.test.ts`,
`worker-pool-slot-generation.test.ts`
+42 -4
View File
@@ -14,6 +14,8 @@
const fs = require('fs');
const path = require('path');
const { spawnSync } = require('child_process');
const { acquireHookSlot } = require('./hook-lock.cjs');
const { hasGitNexusDbLockedByGitNexusServer } = require('./hook-db-lock-probe.cjs');
/**
* Read JSON input from stdin synchronously.
@@ -102,6 +104,28 @@ function findGitNexusDir(startDir) {
return null;
}
function hasGitNexusServerOwner(gitNexusDir) {
return hasGitNexusDbLockedByGitNexusServer(path.join(gitNexusDir, 'lbug'), process.pid);
}
function extractAugmentContext(stderr) {
const output = (stderr || '').trim();
const marker = output.indexOf('[GitNexus]');
const debug = process.env.GITNEXUS_DEBUG === '1' || process.env.GITNEXUS_DEBUG === 'true';
if (debug && output.length > 0) {
// Emit the FULL discarded prefix (everything before the marker, or all of
// it when no marker is present) so suppressed diagnostics — KuzuDB lock
// warnings, parser errors, etc. — remain recoverable on the hook's own
// stderr. The untruncated payload lets operators see exactly what was
// filtered out instead of a 180-char JSON-quoted preview.
const discarded = marker === -1 ? output : output.slice(0, marker).trim();
if (discarded.length > 0) {
process.stderr.write(`[GitNexus hook] augment stderr discarded prefix:\n${discarded}\n`);
}
}
return marker === -1 ? '' : output.slice(marker).trim();
}
/**
* Extract search pattern from tool input.
*/
@@ -167,6 +191,10 @@ function extractPattern(toolName, toolInput) {
* 3. Fall back to npx (returns empty string)
*/
function resolveCliPath() {
const fromEnv = process.env.GITNEXUS_HOOK_CLI_PATH;
if (fromEnv !== undefined && String(fromEnv).trim() && fs.existsSync(String(fromEnv))) {
return String(fromEnv);
}
let cliPath = path.resolve(__dirname, '..', '..', 'dist', 'cli', 'index.js');
if (!fs.existsSync(cliPath)) {
try {
@@ -207,7 +235,8 @@ function runGitNexusCli(cliPath, args, cwd, timeout) {
function handlePreToolUse(input) {
const cwd = input.cwd || process.cwd();
if (!path.isAbsolute(cwd)) return;
if (!findGitNexusDir(cwd)) return;
const gitNexusDir = findGitNexusDir(cwd);
if (!gitNexusDir) return;
const toolName = input.tool_name || '';
const toolInput = input.tool_input || {};
@@ -216,20 +245,29 @@ function handlePreToolUse(input) {
const pattern = extractPattern(toolName, toolInput);
if (!pattern || pattern.length < 3) return;
if (hasGitNexusServerOwner(gitNexusDir)) {
process.stderr.write('[GitNexus] augment skipped: MCP server owns DB\n');
return;
}
const release = acquireHookSlot(gitNexusDir);
if (!release) return;
const cliPath = resolveCliPath();
let result = '';
try {
const child = runGitNexusCli(cliPath, ['augment', '--', pattern], cwd, 7000);
if (!child.error && child.status === 0) {
result = child.stderr || '';
result = extractAugmentContext(child.stderr || '');
}
} catch {
/* graceful failure */
} finally {
release();
}
if (result && result.trim()) {
sendHookResponse('PreToolUse', result.trim());
if (result) {
sendHookResponse('PreToolUse', result);
}
}
@@ -0,0 +1,238 @@
/**
* Cross-platform best-effort probe: does another process hold dbPath open
* with a command line that looks like a GitNexus MCP/serve server?
*
* Backends (no user-installed Sysinternals):
* - Linux: scan procfs under /proc (per-PID fd entries) via stat(2) (dev+inode); works without lsof;
* optional lsof fallback when proc scan finds nothing.
* - macOS / *BSD / etc.: trusted lsof + ps (absolute paths first).
* - Windows: Restart Manager (rstrtmgr) via bundled PowerShell script +
* Win32_Process for command lines; trusted powershell.exe under %SystemRoot%.
*
* Fail-open on most errors; fail-closed only on lsof ETIMEDOUT (Unix) or
* PowerShell ETIMEDOUT (Windows), matching the hook contract.
*/
const fs = require('fs');
const path = require('path');
const { spawnSync } = require('child_process');
function isGitNexusServerCommand(command) {
const hasServerMode = /(?:^|\s)(mcp|serve)(?:\s|$)/.test(command);
const hasGitNexus =
/(?:^|[/\\\s])gitnexus(?:\.cmd)?(?:\s|$)/.test(command) ||
/node_modules[/\\]gitnexus[/\\]/.test(command);
return hasServerMode && hasGitNexus;
}
function resolveHookBinary(tool) {
const envKey = tool === 'lsof' ? 'GITNEXUS_HOOK_LSOF_PATH' : 'GITNEXUS_HOOK_PS_PATH';
const fromEnv = process.env[envKey];
if (fromEnv && String(fromEnv).trim() && fs.existsSync(String(fromEnv))) {
return String(fromEnv);
}
const candidates =
tool === 'lsof'
? ['/usr/bin/lsof', '/usr/sbin/lsof', '/sbin/lsof', tool]
: ['/bin/ps', '/usr/bin/ps', tool];
for (const candidate of candidates) {
if (candidate === tool) return tool;
try {
if (fs.existsSync(candidate)) return candidate;
} catch {
/* ignore */
}
}
return tool;
}
function resolveWindowsPowerShellPath() {
const fromEnv = process.env.GITNEXUS_HOOK_POWERSHELL_PATH;
if (fromEnv && String(fromEnv).trim() && fs.existsSync(String(fromEnv).trim())) {
return String(fromEnv).trim();
}
const root = process.env.SystemRoot || 'C:\\Windows';
const ps = path.join(root, 'System32', 'WindowsPowerShell', 'v1.0', 'powershell.exe');
if (fs.existsSync(ps)) return ps;
const psWow = path.join(root, 'SysWOW64', 'WindowsPowerShell', 'v1.0', 'powershell.exe');
if (fs.existsSync(psWow)) return psWow;
return 'powershell.exe';
}
// Sentinel:
// undefined = not loaded yet (try the read)
// string = encoded PowerShell command (successful load)
// null = load attempted and failed (do not retry; warning already emitted)
let windowsRmListPsEncodedCommandCache;
let windowsRmListPsLoadFailureWarned = false;
function getWindowsRmListEncodedCommand() {
if (windowsRmListPsEncodedCommandCache !== undefined) {
return windowsRmListPsEncodedCommandCache;
}
try {
const ps1Path = path.join(__dirname, 'win-rm-list-json.ps1');
const src = fs
.readFileSync(ps1Path, 'utf8')
.replace(/^\uFEFF/, '')
.replace(/\r\n/g, '\n');
windowsRmListPsEncodedCommandCache = Buffer.from(src, 'utf16le').toString('base64');
} catch (err) {
windowsRmListPsEncodedCommandCache = null;
if (
!windowsRmListPsLoadFailureWarned &&
(process.env.GITNEXUS_DEBUG === '1' || process.env.GITNEXUS_DEBUG === 'true')
) {
windowsRmListPsLoadFailureWarned = true;
const msg = err && err.message ? String(err.message).slice(0, 200) : 'unknown';
process.stderr.write(`[GitNexus hook] win-rm-list-json.ps1 load failed: ${msg}\n`);
}
}
return windowsRmListPsEncodedCommandCache;
}
function hasGitNexusServerOwnerWindows(dbPathAbs, myPid) {
const encoded = getWindowsRmListEncodedCommand();
if (!encoded) return false;
const psExe = resolveWindowsPowerShellPath();
const r = spawnSync(
psExe,
[
'-NoProfile',
'-NonInteractive',
'-ExecutionPolicy',
'Bypass',
'-STA',
'-EncodedCommand',
encoded,
],
{
encoding: 'utf-8',
timeout: 6000,
stdio: ['ignore', 'pipe', 'ignore'],
env: { ...process.env, GITNEXUS_HOOK_RM_TARGET: dbPathAbs },
},
);
// ETIMEDOUT means the PowerShell probe didn't return in time; treat as 'unresponsive process holds DB' → fail-closed (skip augment).
if (r.error) return r.error.code === 'ETIMEDOUT';
if (r.status !== 0) return false;
let rows;
try {
rows = JSON.parse(String(r.stdout || '').trim() || '[]');
} catch {
return false;
}
if (!Array.isArray(rows)) return false;
for (const row of rows) {
const procId = Number(row.pid);
const cmd = String(row.cmd || '');
if (!Number.isFinite(procId) || procId === myPid) continue;
if (isGitNexusServerCommand(cmd)) return true;
}
return false;
}
function readLinuxCmdline(pidStr) {
try {
return fs.readFileSync(`/proc/${pidStr}/cmdline`, 'utf8').replace(/\0+/g, ' ').trim();
} catch {
return '';
}
}
function linuxProcScanFindGitNexusServer(dbPathAbs, myPid) {
const raw = process.env.GITNEXUS_HOOK_LINUX_PROC_BUDGET_MS;
const budget = Number(raw && String(raw).trim()) ? Number.parseInt(String(raw), 10) : 1200;
const start = Date.now();
let targetStat;
try {
targetStat = fs.statSync(dbPathAbs);
} catch {
return false;
}
let procEntries;
try {
procEntries = fs.readdirSync('/proc', { withFileTypes: true });
} catch {
return false;
}
for (const ent of procEntries) {
if (Date.now() - start > budget) return false;
if (!ent.isDirectory() || !/^\d+$/.test(ent.name)) continue;
const pid = Number.parseInt(ent.name, 10);
if (!Number.isFinite(pid) || pid === myPid) continue;
const fdDir = path.join('/proc', ent.name, 'fd');
let fds;
try {
fds = fs.readdirSync(fdDir);
} catch {
continue;
}
let holds = false;
for (const fd of fds) {
if (Date.now() - start > budget) return false;
try {
const st = fs.statSync(path.join(fdDir, fd));
if (st.dev === targetStat.dev && st.ino === targetStat.ino) {
holds = true;
break;
}
} catch {
/* ignore */
}
}
if (!holds) continue;
if (isGitNexusServerCommand(readLinuxCmdline(ent.name))) return true;
}
return false;
}
function unixLsofPsFindGitNexusServer(dbPathAbs, myPid) {
const lsofPath = resolveHookBinary('lsof');
const lsof = spawnSync(lsofPath, ['-nP', '-t', '--', dbPathAbs], {
encoding: 'utf-8',
timeout: 1000,
stdio: ['ignore', 'pipe', 'ignore'],
});
if (lsof.error) return lsof.error.code === 'ETIMEDOUT';
const pids = (lsof.stdout || '').split(/\s+/).filter(Boolean);
const psPath = resolveHookBinary('ps');
for (const pid of pids) {
if (Number(pid) === myPid) continue;
const ps = spawnSync(psPath, ['-p', pid, '-o', 'command='], {
encoding: 'utf-8',
timeout: 500,
stdio: ['ignore', 'pipe', 'ignore'],
});
if (ps.error) {
if (ps.error.code === 'ETIMEDOUT') return true;
continue;
}
if (isGitNexusServerCommand(ps.stdout || '')) return true;
}
return false;
}
/**
* @param {string} dbPath Absolute or relative path to the DB file (e.g. .../lbug).
* @param {number} myPid Current process PID (hook runner), excluded from matches.
*/
function hasGitNexusDbLockedByGitNexusServer(dbPath, myPid) {
if (!fs.existsSync(dbPath)) return false;
const dbPathAbs = path.resolve(dbPath);
if (process.platform === 'win32') {
return hasGitNexusServerOwnerWindows(dbPathAbs, myPid);
}
if (process.platform === 'linux') {
if (linuxProcScanFindGitNexusServer(dbPathAbs, myPid)) return true;
return unixLsofPsFindGitNexusServer(dbPathAbs, myPid);
}
return unixLsofPsFindGitNexusServer(dbPathAbs, myPid);
}
module.exports = {
hasGitNexusDbLockedByGitNexusServer,
};
+119
View File
@@ -0,0 +1,119 @@
const fs = require('fs');
const path = require('path');
const HOOK_LOCK_SUBDIR = '.hook-locks';
const HOOK_LOCK_MAX_INFLIGHT = 3;
const HOOK_LOCK_STALE_MS = 30000;
function acquireHookSlot(gitNexusDir) {
const lockDir = path.join(gitNexusDir, HOOK_LOCK_SUBDIR);
try {
fs.mkdirSync(lockDir, { recursive: true });
} catch {
// Cannot create lock dir (read-only fs, cross-user perm denial, out of
// inodes, etc.) — fail closed by returning null. Caller skips augment.
// Fail-open here would let N concurrent hooks all proceed unguarded and
// reintroduce the #1486 fan-out the guard exists to prevent.
return null;
}
const myPidStr = String(process.pid);
for (let slot = 0; slot < HOOK_LOCK_MAX_INFLIGHT; slot++) {
const slotPath = path.join(lockDir, `slot-${slot}.lock`);
for (let attempt = 0; attempt < 2; attempt++) {
try {
fs.writeFileSync(slotPath, myPidStr, { flag: 'wx' });
let released = false;
const release = () => {
if (released) return;
released = true;
try {
// Only unlink if we still own the slot. If we appeared stale and
// another hook took over, the file now belongs to it — leave alone.
const content = fs.readFileSync(slotPath, 'utf-8').trim();
if (content === myPidStr) fs.unlinkSync(slotPath);
} catch {
/* already removed or unreadable */
}
};
process.on('exit', release);
return release;
} catch {
// Slot exists. Decide whether to take it over.
// Open once and inspect mtime + content via the same fd so there's
// no TOCTOU between the metadata check and the content read
// (codeql js/file-system-race).
let fd;
try {
fd = fs.openSync(slotPath, 'r');
} catch {
continue; // Vanished between EEXIST and open — retry this slot.
}
let isLive = false;
let mtimeMs = Date.now();
try {
mtimeMs = fs.fstatSync(fd).mtimeMs;
const buf = Buffer.alloc(32);
const n = fs.readSync(fd, buf, 0, 32, 0);
const ownerStr = buf.slice(0, n).toString('utf-8').trim();
if (ownerStr === '') {
// Owner created the file but hasn't written its PID yet. The
// wx open+write window is microseconds; give it the benefit
// of the doubt and treat as live.
isLive = true;
} else {
const owner = Number.parseInt(ownerStr, 10);
if (Number.isFinite(owner) && owner > 0) {
try {
process.kill(owner, 0);
isLive = true;
} catch (e) {
// ESRCH = process gone → treat as dead. EPERM = process exists
// but owned by another user (cross-user lock dir) → still alive,
// keep the slot. Anything else: be conservative, assume alive.
if (e && e.code === 'ESRCH') {
isLive = false;
} else {
isLive = true;
}
}
}
}
} catch {
/* unreadable — treat as dead */
} finally {
try {
fs.closeSync(fd);
} catch {
/* already closed */
}
}
// For slots younger than HOOK_LOCK_STALE_MS, PID-liveness wins —
// a slow-but-alive hook is never wrongly evicted. For older slots,
// age is the final arbiter as a defense against PID reuse on long-
// abandoned slots. 30s >> the 7s augment timeout, so a healthy run
// never crosses this threshold.
if (isLive && Date.now() - mtimeMs > HOOK_LOCK_STALE_MS) {
isLive = false;
}
if (isLive) break; // Try the next slot.
try {
fs.unlinkSync(slotPath);
} catch {
/* another hook beat us to it — retry will hit EEXIST */
}
// Loop and retry this slot.
}
}
}
return null;
}
module.exports = {
HOOK_LOCK_SUBDIR,
HOOK_LOCK_MAX_INFLIGHT,
HOOK_LOCK_STALE_MS,
acquireHookSlot,
};
@@ -0,0 +1,76 @@
$ErrorActionPreference = 'Stop'
$target = $env:GITNEXUS_HOOK_RM_TARGET
if ([string]::IsNullOrWhiteSpace($target)) { Write-Output '[]'; exit 0 }
$target = (Resolve-Path -LiteralPath $target).ProviderPath
if (-not ([Management.Automation.PSTypeName]'GitNexusHookRm.Native').Type) {
Add-Type @'
using System;
using System.Runtime.InteropServices;
namespace GitNexusHookRm {
public static class Native {
public const int ErrorMoreData = 234;
[StructLayout(LayoutKind.Sequential, Pack = 4)]
public struct RM_UNIQUE_PROCESS {
public int dwProcessId;
public long ProcessStartTime;
}
[StructLayout(LayoutKind.Sequential, CharSet = CharSet.Unicode)]
public struct RM_PROCESS_INFO {
public RM_UNIQUE_PROCESS Process;
[MarshalAs(UnmanagedType.ByValTStr, SizeConst = 256)]
public string strAppName;
[MarshalAs(UnmanagedType.ByValTStr, SizeConst = 64)]
public string strServiceShortName;
public uint ApplicationType;
public uint AppStatus;
public uint TSSessionId;
public uint bRestartable;
}
[DllImport("rstrtmgr.dll", CharSet = CharSet.Unicode)]
public static extern int RmStartSession(out uint pSessionHandle, uint dwSessionFlags, string strSessionKey);
[DllImport("rstrtmgr.dll", CharSet = CharSet.Unicode)]
public static extern int RmRegisterResources(uint pSessionHandle, uint nFiles, string[] rgsFileNames, uint nApplications, IntPtr rgApplications, uint nServices, string[] rgsServiceNames);
[DllImport("rstrtmgr.dll")]
public static extern int RmGetList(uint dwSessionHandle, out uint pnProcInfoNeeded, ref uint pnProcInfo, [In, Out] RM_PROCESS_INFO[] rgAffectedApps, ref uint lpdwRebootReasons);
[DllImport("rstrtmgr.dll")]
public static extern int RmEndSession(uint pSessionHandle);
}
}
'@
}
$h = [uint32]0
$key = [guid]::NewGuid().ToString('N')
$rmErr = [GitNexusHookRm.Native]::RmStartSession([ref]$h, 0, $key)
if ($rmErr -ne 0) { Write-Output '[]'; exit 0 }
$files = @($target)
$err = [GitNexusHookRm.Native]::RmRegisterResources($h, 1, $files, 0, [IntPtr]::Zero, 0, $null)
if ($err -ne 0) {
[void][GitNexusHookRm.Native]::RmEndSession($h)
Write-Output '[]'
exit 0
}
$need = [uint32]0
$n = [uint32]0
$reboot = [uint32]0
$err = [GitNexusHookRm.Native]::RmGetList($h, [ref]$need, [ref]$n, $null, [ref]$reboot)
if ($err -ne [GitNexusHookRm.Native]::ErrorMoreData) {
[void][GitNexusHookRm.Native]::RmEndSession($h)
Write-Output '[]'
exit 0
}
$n = $need
$buf = New-Object GitNexusHookRm.Native+RM_PROCESS_INFO[] ([int]$n)
$err = [GitNexusHookRm.Native]::RmGetList($h, [ref]$need, [ref]$n, $buf, [ref]$reboot)
[void][GitNexusHookRm.Native]::RmEndSession($h)
if ($err -ne 0) { Write-Output '[]'; exit 0 }
$out = @()
for ($i = 0; $i -lt [int]$n; $i++) {
$procId = $buf[$i].Process.dwProcessId
$p = Get-CimInstance -ClassName Win32_Process -Filter "ProcessId=$procId" -ErrorAction SilentlyContinue
$cmd = if ($p) { $p.CommandLine } else { '' }
$out += [PSCustomObject]@{ pid = [int]$procId; cmd = $cmd }
}
ConvertTo-Json -InputObject @($out) -Compress
+386 -836
View File
File diff suppressed because it is too large Load Diff
+6 -9
View File
@@ -1,6 +1,6 @@
{
"name": "gitnexus",
"version": "1.6.3",
"version": "1.6.5",
"description": "Graph-powered code intelligence for AI agents. Index any codebase, query via MCP or CLI.",
"author": "Abhigyan Patwari",
"license": "PolyForm-Noncommercial-1.0.0",
@@ -48,7 +48,7 @@
"test:integration": "vitest run test/integration",
"test:watch": "vitest",
"test:coverage": "vitest run --coverage",
"postinstall": "node scripts/build-tree-sitter-dart.cjs && node scripts/build-tree-sitter-proto.cjs",
"postinstall": "node scripts/materialize-vendor-grammars.cjs && node scripts/build-tree-sitter-dart.cjs && node scripts/build-tree-sitter-proto.cjs && node scripts/build-tree-sitter-swift.cjs",
"prepare": "node scripts/build.js",
"prepack": "node scripts/build.js"
},
@@ -60,7 +60,7 @@
"cli-progress": "^3.12.0",
"commander": "^14.0.3",
"cors": "^2.8.5",
"express": "^4.19.2",
"express": "^5.2.1",
"express-rate-limit": "^8.4.1",
"glob": "^13.0.6",
"graphology": "^0.26.0",
@@ -92,15 +92,12 @@
"optionalDependencies": {
"node-addon-api": "^8.0.0",
"node-gyp-build": "^4.8.0",
"tree-sitter-dart": "file:./vendor/tree-sitter-dart",
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-proto": "file:./vendor/tree-sitter-proto",
"tree-sitter-swift": "file:./vendor/tree-sitter-swift"
"tree-sitter-kotlin": "^0.3.8"
},
"devDependencies": {
"@types/cli-progress": "^3.11.6",
"@types/cors": "^2.8.17",
"@types/express": "^4.17.21",
"@types/express": "^5.0.6",
"@types/js-yaml": "^4.0.9",
"@types/node": "^25.6.0",
"@types/uuid": "^11.0.0",
@@ -116,6 +113,6 @@
}
},
"engines": {
"node": ">=20.0.0"
"node": ">=22.0.0"
}
}
@@ -1,4 +1,8 @@
#!/usr/bin/env node
/**
* Build tree-sitter-dart native binding in node_modules/ after materialize-vendor-grammars.cjs.
* Vendored source lives in vendor/ only; see #836 and #1728.
*/
const fs = require('fs');
const path = require('path');
const { execSync } = require('child_process');
+3 -4
View File
@@ -4,7 +4,7 @@
*
* Why this script exists:
* tree-sitter-proto is vendored under gitnexus/vendor/tree-sitter-proto/
* and declared as a `file:` optionalDependency. Previously, the vendored
* and copied into node_modules/ by materialize-vendor-grammars.cjs. Previously, the vendored
* package had its own `dependencies` and `install` script, which caused
* npm to create `vendor/tree-sitter-proto/node_modules/` and
* `vendor/tree-sitter-proto/build/` during install. Those directories
@@ -20,9 +20,8 @@
* gitnexus's own optionalDependencies, and moved native compilation here.
*
* What this does:
* Runs `npx node-gyp rebuild` inside `node_modules/tree-sitter-proto/`
* (which npm creates as a copy of vendor/tree-sitter-proto/ when
* resolving the file: dep). Build output lands in
* Runs `npx node-gyp rebuild` inside `node_modules/tree-sitter-proto/`.
* Build output lands in
* `node_modules/tree-sitter-proto/build/Release/tree_sitter_proto_binding.node`
* — under npm-managed territory, safe on upgrade.
*
@@ -0,0 +1,39 @@
#!/usr/bin/env node
/**
* Probe tree-sitter-swift prebuild availability at install time.
*
* The vendored package ships platform prebuilds; node-gyp-build selects the
* correct binary at require time. This script calls node-gyp-build once
* against the materialized package so a missing-prebuild failure surfaces
* as an install-time warning (with the rest of the gitnexus install
* succeeding) rather than as a runtime error the first time Swift parsing
* is requested. The result is discarded — it does not copy, register, or
* mutate anything; the runtime require() path in parser-loader does the
* actual load. Running this probe here instead of an npm `install` script
* on the vendored package preserves the #836 hygiene (no scripts.install
* inside vendor/).
*/
const fs = require('fs');
const path = require('path');
if (process.env.GITNEXUS_SKIP_OPTIONAL_GRAMMARS === '1') {
console.warn('[tree-sitter-swift] Skipping prebuild probe (GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1).');
process.exit(0);
}
const swiftDir = path.join(__dirname, '..', 'node_modules', 'tree-sitter-swift');
try {
if (!fs.existsSync(path.join(swiftDir, 'bindings', 'node', 'index.js'))) {
process.exit(0);
}
const nodeGypBuild = require('node-gyp-build');
nodeGypBuild(swiftDir);
} catch (err) {
console.warn('[tree-sitter-swift] Prebuild probe failed:', err.message);
console.warn(
'[tree-sitter-swift] Swift parsing will be unavailable. Non-Swift functionality is unaffected.',
);
process.exit(0);
}
+24 -4
View File
@@ -18,14 +18,34 @@ const ROOT = path.resolve(__dirname, '..');
const SHARED_ROOT = path.resolve(ROOT, '..', 'gitnexus-shared');
const DIST = path.join(ROOT, 'dist');
const SHARED_DEST = path.join(DIST, '_shared');
const DEFAULT_BUILD_TIMEOUT_MS = 300_000;
function getBuildTimeoutMs() {
const raw = process.env.GITNEXUS_BUILD_TIMEOUT_MS;
if (raw === undefined || raw.trim() === '') return DEFAULT_BUILD_TIMEOUT_MS;
const parsed = Number.parseInt(raw, 10);
if (Number.isFinite(parsed) && parsed > 0) return parsed;
console.warn(
`[build] ignoring invalid GITNEXUS_BUILD_TIMEOUT_MS=${JSON.stringify(raw)}; using ${DEFAULT_BUILD_TIMEOUT_MS}ms`,
);
return DEFAULT_BUILD_TIMEOUT_MS;
}
const BUILD_TIMEOUT_MS = getBuildTimeoutMs();
// ── 1. Build gitnexus-shared ───────────────────────────────────────
console.log('[build] compiling gitnexus-shared…');
execSync('npx tsc', { cwd: SHARED_ROOT, stdio: 'inherit', timeout: 120_000 });
const tscCmd =
process.platform === 'win32'
? path.join('node_modules', '.bin', 'tsc.cmd')
: path.join('node_modules', '.bin', 'tsc');
execSync(tscCmd, { cwd: SHARED_ROOT, stdio: 'inherit', timeout: BUILD_TIMEOUT_MS });
// ── 2. Build gitnexus ──────────────────────────────────────────────
console.log('[build] compiling gitnexus…');
execSync('npx tsc', { cwd: ROOT, stdio: 'inherit', timeout: 120_000 });
execSync(tscCmd, { cwd: ROOT, stdio: 'inherit', timeout: BUILD_TIMEOUT_MS });
// ── 3. Copy shared dist ────────────────────────────────────────────
console.log('[build] copying shared module into dist/_shared…');
@@ -78,9 +98,9 @@ if (fs.existsSync(path.join(WEB_ROOT, 'package.json'))) {
console.log('[build] building gitnexus-web…');
if (!fs.existsSync(path.join(WEB_ROOT, 'node_modules'))) {
console.log('[build] installing gitnexus-web dependencies…');
execSync('npm ci', { cwd: WEB_ROOT, stdio: 'inherit', timeout: 120_000 });
execSync('npm ci', { cwd: WEB_ROOT, stdio: 'inherit', timeout: BUILD_TIMEOUT_MS });
}
execSync('npm run build', { cwd: WEB_ROOT, stdio: 'inherit', timeout: 120_000 });
execSync('npm run build', { cwd: WEB_ROOT, stdio: 'inherit', timeout: BUILD_TIMEOUT_MS });
// Copy dist → gitnexus/web/ (shipped in the npm package)
fs.rmSync(WEB_DEST, { recursive: true, force: true });
@@ -0,0 +1,72 @@
#!/usr/bin/env node
/**
* Copy vendored tree-sitter grammars into node_modules/ using real files (fs.cpSync).
*
* Published gitnexus used to declare these as optionalDependencies with
* `file:./vendor/...`, which makes npm symlink/junction vendor → node_modules on
* install. Windows without Developer Mode often fails with EPERM (#1728).
*
* Vendor trees stay read-only in gitnexus/vendor/; build artifacts must only
* land under node_modules/ (see #836).
*/
const fs = require('fs');
const path = require('path');
const ROOT = path.join(__dirname, '..');
const VENDORED_GRAMMARS = ['tree-sitter-dart', 'tree-sitter-proto', 'tree-sitter-swift'];
if (process.env.GITNEXUS_SKIP_OPTIONAL_GRAMMARS === '1') {
console.warn(
'[gitnexus] Skipping vendored grammar materialize (GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1). Dart/Proto/Swift parsing will be unavailable.',
);
process.exit(0);
}
for (const name of VENDORED_GRAMMARS) {
const src = path.join(ROOT, 'vendor', name);
const dest = path.join(ROOT, 'node_modules', name);
if (!fs.existsSync(src)) {
console.warn(`[gitnexus] vendor/${name} missing; skipping materialize.`);
continue;
}
// Sequence: copy src → partial; rename dest → backup; rename partial → dest;
// remove backup. If any step fails, restore from backup so a previously-
// materialized grammar is never lost. Targets the #1728 EPERM scenario plus
// narrower failure modes (Windows AV scanner racing on rename, EBUSY mid-swap).
const partial = `${dest}.materialize-tmp`;
const backup = `${dest}.materialize-bak`;
try {
fs.mkdirSync(path.join(ROOT, 'node_modules'), { recursive: true });
fs.rmSync(partial, { recursive: true, force: true });
fs.rmSync(backup, { recursive: true, force: true });
fs.cpSync(src, partial, { recursive: true, verbatim: true });
if (fs.existsSync(dest)) {
fs.renameSync(dest, backup);
}
try {
fs.renameSync(partial, dest);
} catch (renameErr) {
// Best-effort rollback: restore the previous dest from backup.
if (fs.existsSync(backup)) {
try {
fs.renameSync(backup, dest);
} catch {
// If rollback also fails, the prior backup directory still exists on
// disk — the catch block below surfaces both errors via the warning.
}
}
throw renameErr;
}
fs.rmSync(backup, { recursive: true, force: true });
} catch (err) {
// Fail-soft: a single locked/inaccessible file (common on Windows) must not
// abort the whole gitnexus install. Matches build-tree-sitter-*.cjs pattern.
fs.rmSync(partial, { recursive: true, force: true });
console.warn(`[gitnexus] Could not materialize vendor/${name}: ${err.message}`);
console.warn(
`[gitnexus] ${name} parsing will be unavailable. Other functionality is unaffected.`,
);
}
}
+82 -14
View File
@@ -28,6 +28,7 @@ interface RepoStats {
export interface AIContextOptions {
skipAgentsMd?: boolean;
noStats?: boolean;
skipSkills?: boolean;
}
const GITNEXUS_START_MARKER = '<!-- gitnexus:start -->';
@@ -94,6 +95,7 @@ function generateGitNexusContent(
generatedSkills?: GeneratedSkillInfo[],
groupNames?: string[],
noStats?: boolean,
skipSkills?: boolean,
): string {
const generatedRows =
generatedSkills && generatedSkills.length > 0
@@ -105,14 +107,26 @@ function generateGitNexusContent(
.join('\n')
: '';
const skillsTable = `| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | \`.claude/skills/gitnexus/gitnexus-exploring/SKILL.md\` |
// Standard skill rows reference files installed by installSkills(). When
// --skip-skills suppresses that install, these rows must be omitted — else
// AGENTS.md/CLAUDE.md would direct agents to read files that don't exist.
// Community skills (generatedRows) live in .claude/skills/generated/ and
// are independent of --skip-skills, so they remain when present.
const standardSkillsRows = skipSkills
? ''
: `| Understand architecture / "How does X work?" | \`.claude/skills/gitnexus/gitnexus-exploring/SKILL.md\` |
| Blast radius / "What breaks if I change X?" | \`.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md\` |
| Trace bugs / "Why is X failing?" | \`.claude/skills/gitnexus/gitnexus-debugging/SKILL.md\` |
| Rename / extract / split / refactor | \`.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md\` |
| Tools, resources, schema reference | \`.claude/skills/gitnexus/gitnexus-guide/SKILL.md\` |
| Index, status, clean, wiki CLI commands | \`.claude/skills/gitnexus/gitnexus-cli/SKILL.md\` |${generatedRows ? '\n' + generatedRows : ''}`;
| Index, status, clean, wiki CLI commands | \`.claude/skills/gitnexus/gitnexus-cli/SKILL.md\` |`;
const tableBody = [standardSkillsRows, generatedRows].filter(Boolean).join('\n');
const skillsTable = tableBody
? `| Task | Read this skill file |
|------|---------------------|
${tableBody}`
: '';
return `${GITNEXUS_START_MARKER}
# GitNexus — Code Intelligence
@@ -153,11 +167,15 @@ This repository is listed under GitNexus **group(s): ${groupNames.join(', ')}**
`
: ''
}## CLI
}${
skillsTable
? `## CLI
${skillsTable}
${GITNEXUS_END_MARKER}`;
`
: ''
}${GITNEXUS_END_MARKER}`;
}
/**
@@ -181,7 +199,9 @@ async function fileExists(filePath: string): Promise<boolean> {
async function upsertGitNexusSection(
filePath: string,
content: string,
): Promise<'created' | 'updated' | 'appended'> {
projectName: string,
stats: RepoStats,
): Promise<'created' | 'updated' | 'appended' | 'preserved'> {
const exists = await fileExists(filePath);
if (!exists) {
@@ -205,7 +225,50 @@ async function upsertGitNexusSection(
);
if (startIdx !== -1 && endIdx !== -1 && endIdx > startIdx) {
// Replace existing section
const existingSection = existingContent.substring(
startIdx,
endIdx + GITNEXUS_END_MARKER.length,
);
// If the existing section contains <!-- gitnexus:keep -->, preserve the user's
// custom layout and only update the stats line (node/edge/flow counts).
// This lets teams trim the verbose default template to a lean format without
// having it overwritten on every `gitnexus analyze`.
//
// Note: the keep-marker check operates on `existingSection` (the substring
// between valid section markers identified by findSectionMarkerIndex), so
// a keep marker in user prose OUTSIDE the GitNexus block has no effect.
if (existingSection.includes('<!-- gitnexus:keep -->')) {
// Build the new stats line from the caller-provided values directly.
// We do NOT re-extract from `content` because:
// (a) first-bold extraction is fragile if the template evolves
// (b) the parenthesized-text fallback can match unrelated tuples
// like `({target: "symbolName", direction: "upstream"})`
// when noStats is set
// Passing projectName + stats explicitly makes the contract obvious.
// noStats controls template generation, not keep-section stat updates — the user opted into a stats line by keeping it.
const newStatsInner = `${stats.nodes || 0} symbols, ${stats.edges || 0} relationships, ${stats.processes || 0} execution flows`;
const statsLine = `Indexed as **${projectName}** (${newStatsInner})`;
// Match either canonical phrasing at line start (`^` with `m` flag) so we
// cannot replace prose embedded mid-paragraph. Deliberately no `$`: text
// after the closing `)` on the same line (e.g. ". MCP tools.") stays intact.
const statsPattern = /^(?:Indexed as|indexed by GitNexus as) \*\*[^*]+\*\* \([^)]+\)/m;
if (statsPattern.test(existingSection)) {
const updatedSection = existingSection.replace(statsPattern, statsLine);
const before = existingContent.substring(0, startIdx);
const after = existingContent.substring(endIdx + GITNEXUS_END_MARKER.length);
await fs.writeFile(filePath, (before + updatedSection + after).trim() + '\n', 'utf-8');
return 'updated';
}
// Keep marker present but no stats line matched. Section is preserved
// unchanged on disk; return a distinct status so callers/CLI output
// don't mis-report this as 'updated' (which would imply a write).
return 'preserved';
}
// No keep marker — replace existing section with full verbose content
const before = existingContent.substring(0, startIdx);
const after = existingContent.substring(endIdx + GITNEXUS_END_MARKER.length);
const newContent = before + content + after;
@@ -319,28 +382,33 @@ export async function generateAIContextFiles(
generatedSkills,
groupNames,
options?.noStats,
options?.skipSkills,
);
const createdFiles: string[] = [];
if (!options?.skipAgentsMd) {
// Create AGENTS.md (standard for Cursor, Windsurf, OpenCode, Cline, etc.)
const agentsPath = path.join(repoPath, 'AGENTS.md');
const agentsResult = await upsertGitNexusSection(agentsPath, content);
const agentsResult = await upsertGitNexusSection(agentsPath, content, projectName, stats);
createdFiles.push(`AGENTS.md (${agentsResult})`);
// Create CLAUDE.md (for Claude Code)
const claudePath = path.join(repoPath, 'CLAUDE.md');
const claudeResult = await upsertGitNexusSection(claudePath, content);
const claudeResult = await upsertGitNexusSection(claudePath, content, projectName, stats);
createdFiles.push(`CLAUDE.md (${claudeResult})`);
} else {
createdFiles.push('AGENTS.md (skipped via --skip-agents-md)');
createdFiles.push('CLAUDE.md (skipped via --skip-agents-md)');
}
// Install skills to .claude/skills/gitnexus/
const installedSkills = await installSkills(repoPath);
if (installedSkills.length > 0) {
createdFiles.push(`.claude/skills/gitnexus/ (${installedSkills.length} skills)`);
// Install skills to .claude/skills/gitnexus/ (unless --skip-skills)
if (!options?.skipSkills) {
const installedSkills = await installSkills(repoPath);
if (installedSkills.length > 0) {
createdFiles.push(`.claude/skills/gitnexus/ (${installedSkills.length} skills)`);
}
} else {
createdFiles.push('.claude/skills/gitnexus/ (skipped via --skip-skills)');
}
return { files: createdFiles };
+587 -36
View File
@@ -9,10 +9,11 @@
*/
import path from 'path';
import { execFileSync } from 'child_process';
import { spawn } from 'child_process';
import v8 from 'v8';
import cliProgress from 'cli-progress';
import { closeLbug } from '../core/lbug/lbug-adapter.js';
import { isWalCorruptionError, WAL_RECOVERY_SUGGESTION } from '../core/lbug/lbug-config.js';
import {
getStoragePaths,
getGlobalRegistryPath,
@@ -36,6 +37,7 @@ import { isHfDownloadFailure } from '../core/embeddings/hf-env.js';
// previous behaviour silently swallowed stack traces and made #1169
// indistinguishable from a no-op success on Windows.
const realStderrWrite = process.stderr.write.bind(process.stderr);
const realStdoutWrite = process.stdout.write.bind(process.stdout);
const writeFatalToStderr = (label: string, err: unknown): void => {
const isErr = err instanceof Error;
@@ -67,14 +69,354 @@ const installFatalHandlers = (): void => {
});
};
const HEAP_MB = 8192;
const HEAP_FLAG = `--max-old-space-size=${HEAP_MB}`;
const HEAP_MB = 16384;
const TEST_RESPAWN_HEAP_MB = Number(process.env.GITNEXUS_TEST_RESPAWN_HEAP_MB);
const RESPAWN_HEAP_MB =
Number.isFinite(TEST_RESPAWN_HEAP_MB) && TEST_RESPAWN_HEAP_MB > 0
? Math.floor(TEST_RESPAWN_HEAP_MB)
: HEAP_MB;
const HEAP_FLAG = `--max-old-space-size=${RESPAWN_HEAP_MB}`;
/** Increase default stack size (KB) to prevent stack overflow on deep class hierarchies. */
const STACK_KB = 4096;
const STACK_FLAG = `--stack-size=${STACK_KB}`;
const RESPAWN_OUTPUT_TAIL_CHARS = 1024 * 1024;
const RESPAWN_PROGRESS_ENV = 'GITNEXUS_RESPAWN_PROGRESS_TTY';
/** Re-exec the process with an 8GB heap and larger stack if we're currently below that. */
function ensureHeap(): boolean {
interface CliProgressTerminal {
cursorSave(): void;
cursorRestore(): void;
cursor(enabled: boolean): void;
lineWrapping(enabled: boolean): void;
cursorTo(x?: number | null, y?: number | null): void;
cursorRelative(dx?: number | null, dy?: number | null): void;
cursorRelativeReset(): void;
clearRight(): void;
clearLine(): void;
clearBottom(): void;
newline(): void;
write(s: string, rawWrite?: boolean): void;
isTTY(): boolean;
getWidth(): number;
}
const terminalColumns = (): number => {
const parsed = Number(process.env.COLUMNS);
return Number.isFinite(parsed) && parsed > 0 ? Math.floor(parsed) : 80;
};
const ANSI_ESCAPE_PATTERN =
/\x1B(?:\[[0-?]*[ -/]*[@-~]|\][^\x07]*(?:\x07|\x1B\\)|[PX^_][\s\S]*?\x1B\\|[78]|[@-Z\\-_])/y;
interface IntlSegmenterLike {
segment(input: string): Iterable<{ segment: string }>;
}
type IntlWithOptionalSegmenter = typeof Intl & {
Segmenter?: new (
locales?: string | string[],
options?: { granularity?: 'grapheme' },
) => IntlSegmenterLike;
};
const splitGraphemes = (text: string): string[] => {
const Segmenter = (Intl as IntlWithOptionalSegmenter).Segmenter;
if (Segmenter) {
return Array.from(
new Segmenter(undefined, { granularity: 'grapheme' }).segment(text),
(s) => s.segment,
);
}
return Array.from(text);
};
const isZeroWidthCodePoint = (codePoint: number): boolean =>
codePoint === 0x200d ||
(codePoint >= 0x0300 && codePoint <= 0x036f) ||
(codePoint >= 0x1ab0 && codePoint <= 0x1aff) ||
(codePoint >= 0x1dc0 && codePoint <= 0x1dff) ||
(codePoint >= 0x20d0 && codePoint <= 0x20ff) ||
(codePoint >= 0xfe00 && codePoint <= 0xfe0f) ||
(codePoint >= 0xfe20 && codePoint <= 0xfe2f);
const isWideCodePoint = (codePoint: number): boolean =>
codePoint >= 0x1100 &&
(codePoint <= 0x115f ||
codePoint === 0x2329 ||
codePoint === 0x232a ||
(codePoint >= 0x2e80 && codePoint <= 0xa4cf && codePoint !== 0x303f) ||
(codePoint >= 0xac00 && codePoint <= 0xd7a3) ||
(codePoint >= 0xf900 && codePoint <= 0xfaff) ||
(codePoint >= 0xfe10 && codePoint <= 0xfe19) ||
(codePoint >= 0xfe30 && codePoint <= 0xfe6f) ||
(codePoint >= 0xff00 && codePoint <= 0xff60) ||
(codePoint >= 0xffe0 && codePoint <= 0xffe6) ||
(codePoint >= 0x1f300 && codePoint <= 0x1faff) ||
(codePoint >= 0x20000 && codePoint <= 0x3fffd));
const visibleColumns = (text: string): number => {
let columns = 0;
for (const char of Array.from(text)) {
const codePoint = char.codePointAt(0);
if (codePoint === undefined || isZeroWidthCodePoint(codePoint)) continue;
columns += isWideCodePoint(codePoint) ? 2 : 1;
}
return columns;
};
const readAnsiEscapeAt = (text: string, index: number): string | undefined => {
ANSI_ESCAPE_PATTERN.lastIndex = index;
return ANSI_ESCAPE_PATTERN.exec(text)?.[0];
};
const truncateAnsiToColumns = (text: string, maxColumns: number): string => {
if (!Number.isFinite(maxColumns) || maxColumns <= 0) return '';
let output = '';
let columns = 0;
let index = 0;
while (index < text.length) {
const escape = readAnsiEscapeAt(text, index);
if (escape) {
output += escape;
index += escape.length;
continue;
}
const nextEscapeIndex = text.indexOf('\x1B', index);
const plainEnd = nextEscapeIndex === -1 ? text.length : nextEscapeIndex;
const plainText = text.slice(index, plainEnd);
for (const segment of splitGraphemes(plainText)) {
const width = visibleColumns(segment);
if (width > 0 && columns + width > maxColumns) return output;
output += segment;
columns += width;
}
index = plainEnd;
}
return output;
};
const createAnsiPipeTerminal = (stream: NodeJS.WriteStream): CliProgressTerminal => {
let linewrap = true;
let dy = 0;
const write = (s: string): void => {
stream.write(s);
};
const moveVertical = (delta: number): void => {
if (delta > 0) write(`\x1B[${delta}B`);
else if (delta < 0) write(`\x1B[${Math.abs(delta)}A`);
};
return {
cursorSave: () => write('\x1B7'),
cursorRestore: () => write('\x1B8'),
cursor: (enabled) => write(enabled ? '\x1B[?25h' : '\x1B[?25l'),
lineWrapping: (enabled) => {
linewrap = enabled;
write(enabled ? '\x1B[?7h' : '\x1B[?7l');
},
cursorTo: (x = null, y = null) => {
if (typeof y === 'number' && typeof x === 'number') {
write(`\x1B[${y + 1};${x + 1}H`);
return;
}
if (typeof x === 'number') {
write(x === 0 ? '\r' : `\x1B[${x + 1}G`);
}
},
cursorRelative: (dx = null, nextDy = null) => {
if (typeof dx === 'number' && dx !== 0) {
write(dx > 0 ? `\x1B[${dx}C` : `\x1B[${Math.abs(dx)}D`);
}
if (typeof nextDy === 'number' && nextDy !== 0) {
dy += nextDy;
moveVertical(nextDy);
}
},
cursorRelativeReset: () => {
moveVertical(-dy);
write('\r');
dy = 0;
},
clearRight: () => write('\x1B[0K'),
clearLine: () => write('\x1B[2K'),
clearBottom: () => write('\x1B[0J'),
newline: () => {
write('\n');
dy++;
},
write: (s, rawWrite = false) => {
const width = terminalColumns();
write(linewrap && rawWrite === false ? truncateAnsiToColumns(s, width) : s);
},
isTTY: () => true,
getWidth: terminalColumns,
};
};
const shouldBridgeRespawnProgressTty = (): boolean =>
process.stderr.isTTY === true || process.stdout.isTTY === true;
interface RespawnExit {
status?: number | null;
signal?: NodeJS.Signals | null;
stdout?: string;
stderr?: string;
message?: string;
}
const appendOutputTail = (tail: string, chunk: unknown): string => {
const text = Buffer.isBuffer(chunk)
? chunk.toString('utf8')
: typeof chunk === 'string'
? chunk
: String(chunk ?? '');
if (!text) return tail;
const next = tail + text;
return next.length > RESPAWN_OUTPUT_TAIL_CHARS ? next.slice(-RESPAWN_OUTPUT_TAIL_CHARS) : next;
};
/**
* Run the respawned analyzer while teeing child output through to the parent
* and keeping a bounded tail for crash classification.
*
* `execFileSync(..., { stdio: 'inherit' })` preserved live progress but hid
* stderr/stdout from the parent on abnormal exits. That made every
* SIGABRT/status-134 child look like an output-less V8 heap OOM, even when the
* terminal had already shown a native crash such as
* `libc++abi: ... Napi::Error`. Piped streams plus an explicit tee keeps the UX
* and gives `childProcessLikelyOom` the evidence it needs.
*/
const runRespawnedAnalyze = (
args: readonly string[],
env: NodeJS.ProcessEnv,
): Promise<RespawnExit> =>
new Promise((resolve) => {
let stdout = '';
let stderr = '';
let settled = false;
const finish = (exit: RespawnExit): void => {
if (settled) return;
settled = true;
resolve(exit);
};
const child = spawn(process.execPath, [...args], {
stdio: ['inherit', 'pipe', 'pipe'],
env,
});
child.stdout?.on('data', (chunk) => {
stdout = appendOutputTail(stdout, chunk);
realStdoutWrite(chunk);
});
child.stderr?.on('data', (chunk) => {
stderr = appendOutputTail(stderr, chunk);
realStderrWrite(chunk);
});
child.on('error', (err) => {
finish({
status: 1,
signal: null,
stdout,
stderr,
message: err instanceof Error ? err.message : String(err),
});
});
child.on('close', (status, signal) => {
finish({
status,
signal,
stdout,
stderr,
message: `Command failed: ${process.execPath} ${args.join(' ')}`,
});
});
});
/**
* Heuristic for "child re-exec likely died from V8 OOM".
*
* Platform-independent detection is best-effort: V8/Node usually emit stable
* heap-exhaustion phrases in stderr/message across Linux/macOS/Windows (for
* example "JavaScript heap out of memory" or "Reached heap limit"). When the
* child produced no output at all, we still treat status 134/SIGABRT as likely
* heap OOM. If stderr/stdout contains a native crash diagnostic, the output
* evidence wins and we do not print heap guidance.
*/
const childProcessLikelyOom = (err: unknown): boolean => {
if (!err || typeof err !== 'object') return false;
const e = err as {
status?: unknown;
signal?: unknown;
stderr?: unknown;
stdout?: unknown;
message?: unknown;
};
const hasHeapOomSignature = (v: unknown): boolean => {
const text = (
Buffer.isBuffer(v) ? v.toString('utf8') : typeof v === 'string' ? v : ''
).toLowerCase();
if (!text) return false;
return (
text.includes('javascript heap out of memory') ||
text.includes('reached heap limit') ||
text.includes('allocation failed - javascript heap out of memory') ||
text.includes('fatalprocessoutofmemory')
);
};
const fields = [e.message, e.stderr, e.stdout];
if (fields.some((v) => hasHeapOomSignature(v))) return true;
const hasAnyChildOutput = [e.stderr, e.stdout].some(
(v) => (Buffer.isBuffer(v) && v.length > 0) || (typeof v === 'string' && v.length > 0),
);
if (hasAnyChildOutput) return false;
return e.status === 134 || e.signal === 'SIGABRT';
};
const childProcessLikelyNativeAbort = (err: unknown): boolean => {
if (!err || typeof err !== 'object') return false;
const e = err as {
stderr?: unknown;
stdout?: unknown;
message?: unknown;
};
const hasNativeAbortSignature = (v: unknown): boolean => {
const text = (
Buffer.isBuffer(v) ? v.toString('utf8') : typeof v === 'string' ? v : ''
).toLowerCase();
if (!text) return false;
return (
text.includes('napi::error') ||
text.includes('libc++abi: terminating') ||
text.includes('abort trap') ||
text.includes('native stack') ||
text.includes('native worker') ||
text.includes('native binding')
);
};
return [e.message, e.stderr, e.stdout].some((v) => hasNativeAbortSignature(v));
};
const forceHeapOOMForTestIfEnabled = (): void => {
if (process.env.GITNEXUS_TEST_FORCE_HEAP_OOM !== '1') return;
// Allocate JS strings (not Buffers) so pressure lands on V8 heap itself.
// Buffers can allocate off-heap, which makes OOM triggering less reliable.
const chunks: string[] = [];
for (;;) chunks.push('x'.repeat(1024 * 1024));
};
/** Re-exec the process with a 16GB heap and larger stack if we're currently below that. */
async function ensureHeap(): Promise<boolean> {
const nodeOpts = process.env.NODE_OPTIONS || '';
if (nodeOpts.includes('--max-old-space-size')) return false;
@@ -86,19 +428,79 @@ function ensureHeap(): boolean {
const cliFlags = [HEAP_FLAG];
if (!nodeOpts.includes('--stack-size')) cliFlags.push(STACK_FLAG);
try {
execFileSync(process.execPath, [...cliFlags, ...process.argv.slice(1)], {
stdio: 'inherit',
env: { ...process.env, NODE_OPTIONS: `${nodeOpts} ${HEAP_FLAG}`.trim() },
});
} catch (e: any) {
process.exitCode = e.status ?? 1;
const childArgs = [...cliFlags, ...process.argv.slice(1)];
const childEnv = {
...process.env,
NODE_OPTIONS: `${nodeOpts} ${HEAP_FLAG}`.trim(),
};
if (shouldBridgeRespawnProgressTty()) childEnv[RESPAWN_PROGRESS_ENV] = '1';
const childExit = await runRespawnedAnalyze(childArgs, childEnv);
if (childExit.status !== 0 || childExit.signal) {
if (childProcessLikelyOom(childExit)) {
cliError(
` Analysis likely ran out of memory.\n` +
` Retry with a larger heap if your machine allows it:\n` +
` NODE_OPTIONS="--max-old-space-size=24576" gitnexus analyze [your-args]\n` +
` (Windows: set NODE_OPTIONS=--max-old-space-size=24576 && gitnexus analyze [your-args])\n` +
` If this persists, it may be a native crash unrelated to heap size.\n`,
{ recoveryHint: 'heap-oom-respawn' },
);
} else if (childProcessLikelyNativeAbort(childExit)) {
cliError(
` Analysis aborted in a native worker or native binding path.\n` +
` Try one of these recovery paths:\n` +
` gitnexus analyze --workers 0\n` +
` npm uninstall -g gitnexus && npm install -g gitnexus@latest\n` +
` Use Node 22 LTS if you are on a newer non-LTS runtime.\n`,
{ recoveryHint: 'native-worker-abort' },
);
}
const status =
typeof childExit.status === 'number' && childExit.status !== 0 ? childExit.status : 1;
process.exitCode = status;
}
return true;
}
/**
* GITNEXUS_* env vars that `analyzeCommand` writes for backward-compatible
* downstream consumption. Snapshotted at function entry and restored in the
* finally block so that programmatic callers (tests, long-running hosts)
* don't see leaked state across invocations. `GITNEXUS_WORKER_POOL_SIZE` is
* NOT in this list: that knob is threaded through `runFullAnalysis` options
* (see `workerPoolSize` plumbing) so the CLI never has to mutate `process.env`
* for it in the first place.
*/
const ANALYZE_CLI_ENV_KEYS = [
'GITNEXUS_VERBOSE',
'GITNEXUS_MAX_FILE_SIZE',
'GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS',
'GITNEXUS_EMBEDDING_THREADS',
'GITNEXUS_EMBEDDING_BATCH_SIZE',
'GITNEXUS_EMBEDDING_SUB_BATCH_SIZE',
'GITNEXUS_EMBEDDING_DEVICE',
'GITNEXUS_ANALYZE_PROGRESS_ACTIVE',
] as const;
type AnalyzeEnvSnapshot = Record<(typeof ANALYZE_CLI_ENV_KEYS)[number], string | undefined>;
const snapshotAnalyzeEnv = (): AnalyzeEnvSnapshot => {
const snap = {} as AnalyzeEnvSnapshot;
for (const k of ANALYZE_CLI_ENV_KEYS) snap[k] = process.env[k];
return snap;
};
const restoreAnalyzeEnv = (snap: AnalyzeEnvSnapshot): void => {
for (const k of ANALYZE_CLI_ENV_KEYS) {
const v = snap[k];
if (v === undefined) delete process.env[k];
else process.env[k] = v;
}
};
export interface AnalyzeOptions {
force?: boolean;
repairFts?: boolean;
/**
* Embedding generation toggle. Commander parses `--embeddings [limit]` as:
* - `undefined` when the flag is omitted
@@ -117,8 +519,22 @@ export interface AnalyzeOptions {
verbose?: boolean;
/** Skip AGENTS.md and CLAUDE.md gitnexus block updates. */
skipAgentsMd?: boolean;
/** Omit volatile symbol/relationship counts from AGENTS.md and CLAUDE.md. */
noStats?: boolean;
/**
* Stats inclusion in AGENTS.md and CLAUDE.md.
*
* Commander.js represents `--no-stats` as `stats: boolean` (default
* `true`; `false` when the user passes `--no-stats`), NOT as
* `noStats: boolean`. Reading the negated form would always be
* `undefined` and the flag would silently no-op (#1477). Consumers
* that want "did the user request --no-stats?" should compare with
* `=== false` to distinguish the explicit-off case from the
* default-on case.
*/
stats?: boolean;
/** Skip installing standard GitNexus skill files to .claude/skills/gitnexus/. */
skipSkills?: boolean;
/** Pure index mode: skip all file injection (AGENTS.md, CLAUDE.md, skills). */
indexOnly?: boolean;
/** Index the folder even when no .git directory is present. */
skipGit?: boolean;
/**
@@ -144,20 +560,57 @@ export interface AnalyzeOptions {
maxFileSize?: string;
/** Override worker sub-batch idle timeout in seconds. */
workerTimeout?: string;
/** Parse worker pool size; 0 disables workers (sequential fallback). */
workers?: string;
embeddingThreads?: string;
embeddingBatchSize?: string;
embeddingSubBatchSize?: string;
embeddingDevice?: string;
}
/**
* Whether the post-index skill step should run.
*
* The gated block does two things in sequence: (1) generates the community
* skill files from `--skills`, and (2) re-runs `generateAIContextFiles` so
* AGENTS.md/CLAUDE.md can reference the freshly written skills. Both are
* suppressed together — `--index-only` drops the entire step, not just the
* community-skill write. Name retained for the test contract; see call site
* in `analyzeCommand` for the AGENTS.md/CLAUDE.md re-generation it also gates.
*
* Kept as a pure helper so the `--index-only --skills` contract is unit-tested
* without booting the full analyze pipeline (#742 review).
*/
export const shouldGenerateCommunitySkillFiles = (
options: Pick<AnalyzeOptions, 'skills' | 'indexOnly'> | undefined,
pipelineResult: unknown,
): boolean => Boolean(options?.skills && pipelineResult && !options?.indexOnly);
export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOptions) => {
if (ensureHeap()) return;
if (await ensureHeap()) return;
forceHeapOOMForTestIfEnabled();
// Install fatal handlers immediately after re-exec resolution so any
// async error that escapes the try/catch below (#1169) surfaces with
// a stack trace and a non-zero exit code instead of a silent exit 0.
installFatalHandlers();
// Snapshot the GITNEXUS_* env vars that the impl writes for downstream
// consumption, so they don't leak across `analyzeCommand` invocations in
// programmatic callers (tests, long-running hosts). `process.exit(0)` on
// the success path bypasses `finally` — intentional: when the process is
// exiting, restoration is moot. For early-return paths (validation
// errors) and the alreadyUpToDate fast path the finally restores the
// pre-call values.
const envSnap = snapshotAnalyzeEnv();
try {
await analyzeCommandImpl(inputPath, options);
} finally {
restoreAnalyzeEnv(envSnap);
}
};
const analyzeCommandImpl = async (inputPath?: string, options?: AnalyzeOptions): Promise<void> => {
if (options?.verbose) {
process.env.GITNEXUS_VERBOSE = '1';
}
@@ -178,6 +631,26 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
);
}
// `--workers` is threaded through `runFullAnalysis` options → PipelineOptions
// → createWorkerPool, intentionally bypassing the GITNEXUS_WORKER_POOL_SIZE
// env channel so this CLI surface never mutates `process.env` for pool size.
// Tests can therefore re-invoke analyzeCommand with different --workers
// values back-to-back and observe the value they passed, not whatever the
// previous call leaked.
let workerPoolSize: number | undefined;
if (options?.workers !== undefined) {
const parsedWorkers = Number(options.workers);
if (!Number.isInteger(parsedWorkers) || parsedWorkers < 0) {
cliError(
' --workers must be a non-negative integer. ' +
'Pass 0 to disable the worker pool (sequential fallback).\n',
);
process.exitCode = 1;
return;
}
workerPoolSize = parsedWorkers;
}
// Parse `--embeddings [limit]`: `true` → default cap, string → numeric cap
// (0 disables the cap entirely). Validated up here so failures match the
// sibling-validation pattern (exit before bar.start() — otherwise
@@ -243,8 +716,29 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
process.env.GITNEXUS_EMBEDDING_DEVICE = options.embeddingDevice;
}
if (options?.repairFts && options?.force) {
cliError(
' Cannot combine `--repair-fts` with `--force`. ' +
'Use `--repair-fts` for fast FTS-only repair, or `--force` for a full rebuild.\n',
);
process.exitCode = 1;
return;
}
console.log('\n GitNexus Analyzer\n');
// `--index-only` is the stronger contract — it suppresses every form of file
// injection, including community skill writes that `--skills` would normally
// produce. Surface the override explicitly so users don't wonder why a
// pipeline re-index ran but no skill files appeared. The pipeline still
// re-runs (see `force: options?.force || options?.skills` below); the warning
// is purely about the dropped post-index write step.
if (options?.indexOnly && options?.skills) {
console.log(
' Note: --index-only overrides --skills; community skill files will not be written.\n',
);
}
let repoPath: string;
if (inputPath) {
repoPath = path.resolve(inputPath);
@@ -316,19 +810,25 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
}
// ── CLI progress bar setup ─────────────────────────────────────────
const bar = new cliProgress.SingleBar(
{
format: ' {bar} {percentage}% | {phase}',
barCompleteChar: '\u2588',
barIncompleteChar: '\u2591',
hideCursor: true,
barGlue: '',
autopadding: true,
clearOnComplete: false,
stopOnComplete: false,
},
cliProgress.Presets.shades_grey,
);
const barOptions: cliProgress.Options & { terminal?: CliProgressTerminal } = {
format: ' {bar} {percentage}% | {phase}',
barCompleteChar: '\u2588',
barIncompleteChar: '\u2591',
hideCursor: true,
barGlue: '',
autopadding: true,
clearOnComplete: false,
stopOnComplete: false,
};
if (process.env[RESPAWN_PROGRESS_ENV] === '1' && process.stderr.isTTY !== true) {
// Heap respawn pipes stderr so the parent can classify native/OOM crashes.
// The parent was a real TTY when it opted into this env var, so forward
// ANSI cursor controls through the pipe instead of cli-progress' non-TTY
// newline mode. That keeps one-line redraw UX while retaining stderr tail
// capture for diagnostics.
barOptions.terminal = createAnsiPipeTerminal(process.stderr);
}
const bar = new cliProgress.SingleBar(barOptions, cliProgress.Presets.shades_grey);
bar.start(100, 0, { phase: 'Initializing...' });
@@ -362,7 +862,7 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
// eslint-disable-next-line no-console -- intentional console-routing for progress bar UX
const origError = console.error.bind(console);
let barCurrentValue = 0;
const barLog = (...args: any[]) => {
const barLog = (...args: unknown[]) => {
process.stdout.write('\x1b[2K\r');
origLog(args.map((a) => (typeof a === 'string' ? a : String(a))).join(' '));
bar.update(barCurrentValue);
@@ -372,6 +872,7 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
console.warn = barLog;
// eslint-disable-next-line no-console -- intentional console-routing for progress bar UX
console.error = barLog;
process.env.GITNEXUS_ANALYZE_PROGRESS_ACTIVE = '1';
// Track elapsed time per phase
let lastPhaseLabel = 'Initializing...';
@@ -399,6 +900,9 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
// ── Run shared analysis orchestrator ───────────────────────────────
try {
const skipAll = options?.indexOnly;
const skipAgentsMd = skipAll || options?.skipAgentsMd;
const skipSkills = skipAll || options?.skipSkills;
const result = await runFullAnalysis(
repoPath,
{
@@ -406,18 +910,30 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
// needs a fresh pipelineResult. Has no bearing on the registry
// collision guard (see allowDuplicateName below).
force: options?.force || options?.skills,
repairFts: options?.repairFts,
embeddings: embeddingsEnabled,
embeddingsNodeLimit,
dropEmbeddings: options?.dropEmbeddings,
verbose: options?.verbose,
skipGit: options?.skipGit,
skipAgentsMd: options?.skipAgentsMd,
noStats: options?.noStats,
skipAgentsMd,
skipSkills,
// commander.js `.option('--no-stats', …)` registers the flag as
// `options.stats` (boolean, default true; `false` when the user
// passed --no-stats). Reading `options?.noStats` here returns
// undefined every time, so the flag was a no-op on the markdown
// rewrite path before this fix. See #1477.
noStats: options?.stats === false,
registryName: options?.name,
// Registry-collision bypass — its own CLI flag, intentionally NOT
// overloading --force. A user who hits the collision guard should
// be able to accept the duplicate name without also paying the
// cost of a full pipeline re-index. See #829 review round 2.
allowDuplicateName: options?.allowDuplicateName,
// Worker pool size threaded from --workers, replacing the previous
// GITNEXUS_WORKER_POOL_SIZE env mutation. `undefined` defers to the
// env / auto-formula fallback inside the pipeline.
workerPoolSize,
},
{
onProgress: (_phase, percent, message) => {
@@ -447,6 +963,19 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
return;
}
if (result.ftsRepairedOnly) {
clearInterval(elapsedTimer);
process.removeListener('SIGINT', sigintHandler);
console.log = origLog;
// eslint-disable-next-line no-console -- restoring after intentional progress-bar routing
console.warn = origWarn;
// eslint-disable-next-line no-console -- restoring after intentional progress-bar routing
console.error = origError;
bar.stop();
console.log(' FTS indexes repaired successfully\n');
return;
}
// Post-finalize invariant (#1169): runFullAnalysis nominally writes
// meta.json and registers the repo, but on Windows it has been
// observed to return successfully with neither artifact present
@@ -456,8 +985,10 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
// a healthy index.
await assertAnalysisFinalized(repoPath);
// Skill generation (CLI-only, uses pipeline result from analysis)
if (options?.skills && result.pipelineResult) {
// Skill generation (CLI-only, uses pipeline result from analysis).
// Gated so `--index-only --skills` skips community skill writes too
// (`shouldGenerateCommunitySkillFiles` — see unit test).
if (shouldGenerateCommunitySkillFiles(options, result.pipelineResult)) {
updateBar(99, 'Generating skill files...');
try {
const { generateSkillFiles } = await import('./skill-gen.js');
@@ -497,7 +1028,13 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
processes: s.processes,
},
skillResult.skills,
{ skipAgentsMd: options?.skipAgentsMd, noStats: options?.noStats },
{
skipAgentsMd,
skipSkills,
// Mirror runFullAnalysis `noStats` bridge (#1477) — same expression;
// exercised on the `--skills` path by analyze-no-stats-bridge.test.ts.
noStats: options?.stats === false,
},
);
}
} catch {
@@ -534,7 +1071,7 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
}
console.log('');
} catch (err: any) {
} catch (err: unknown) {
clearInterval(elapsedTimer);
process.removeListener('SIGINT', sigintHandler);
console.log = origLog;
@@ -544,7 +1081,7 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
console.error = origError;
bar.stop();
const msg = err.message || String(err);
const msg = err instanceof Error ? err.message : String(err);
// Registry name-collision from --name (#829) — surface as an
// actionable error rather than a generic stack-trace.
@@ -577,6 +1114,20 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
return;
}
// WAL corruption — the index file is unreadable. Give a clear recovery
// path without a confusing stack trace (the native error message alone
// is enough signal).
if (isWalCorruptionError(err) || msg.includes('LadybugDB WAL corruption')) {
cliError(
` The GitNexus index has a corrupted WAL file.\n` +
` This usually happens when a previous analysis was interrupted mid-write.\n` +
` ${WAL_RECOVERY_SUGGESTION}\n`,
{ recoveryHint: 'wal-corruption' },
);
process.exitCode = 1;
return;
}
// HF download failure — show clean guidance without the raw stack trace.
// Checked before writeFatalToStderr so the user sees one focused message
// rather than a stack-trace dump followed by a second remediation block.
+1 -1
View File
@@ -2,7 +2,7 @@
* Augment CLI Command
*
* Fast-path command for platform hooks.
* Shells out from Claude Code PreToolUse / Cursor beforeShellExecution hooks.
* Shells out from Claude Code PreToolUse / Cursor postToolUse hooks.
*
* Usage: gitnexus augment <pattern>
* Returns enriched text to stdout.
+104 -9
View File
@@ -14,9 +14,14 @@
* Agent bash cmd → curl localhost:PORT/tool/query → eval-server → LocalBackend → format → text
*
* Usage:
* gitnexus eval-server # default port 4848
* gitnexus eval-server --port 4848 # explicit port
* gitnexus eval-server --idle-timeout 300 # auto-shutdown after 300s idle
* gitnexus eval-server # default port 4848, binds 127.0.0.1
* gitnexus eval-server --port 4848 # explicit port
* gitnexus eval-server --host 0.0.0.0 # reachable from other VMs / containers
* gitnexus eval-server --idle-timeout 300 # auto-shutdown after 300s idle
*
* READY signal format: GITNEXUS_EVAL_SERVER_READY:<host>:<port>
* IPv4: GITNEXUS_EVAL_SERVER_READY:127.0.0.1:4848
* IPv6: GITNEXUS_EVAL_SERVER_READY:[::1]:4848
*
* API:
* POST /tool/:name — Call a tool. Body is JSON arguments. Returns formatted text.
@@ -25,16 +30,30 @@
*/
import http from 'http';
import { isIPv4, isIPv6 } from 'node:net';
import { writeSync } from 'node:fs';
import { LocalBackend } from '../mcp/local/local-backend.js';
import { logger } from '../core/logger.js';
import { cliInfo, cliWarn } from './cli-message.js';
import { cliInfo, cliWarn, cliError } from './cli-message.js';
export interface EvalServerOptions {
port?: string;
host?: string;
idleTimeout?: string;
}
/**
* Validate the --host value. Accepts IPv4, IPv6, or "localhost".
* Returns the host string unchanged, or null if invalid.
* "localhost" is passed through so the OS resolves it to the correct loopback
* address (127.0.0.1 or ::1) at bind time rather than forcing IPv4.
*/
export function validateHost(raw: string): string | null {
if (raw === 'localhost') return raw;
if (isIPv4(raw) || isIPv6(raw)) return raw;
return null;
}
// ─── Text Formatters ──────────────────────────────────────────────────
// Convert structured JSON results into compact, LLM-friendly text.
// Design: minimize tokens, maximize actionability.
@@ -330,6 +349,22 @@ export async function evalServerCommand(options?: EvalServerOptions): Promise<vo
const port = parseInt(options?.port || '4848');
const idleTimeoutSec = parseInt(options?.idleTimeout || '0');
const rawHost = options?.host ?? '127.0.0.1';
const host = validateHost(rawHost);
if (!host) {
cliError(
`Invalid --host value "${rawHost}":\n` +
` Must be an IP address or "localhost".\n\n` +
` Examples:\n` +
` gitnexus eval-server --host 127.0.0.1 (loopback only, default)\n` +
` gitnexus eval-server --host 0.0.0.0 (all network interfaces)\n` +
` gitnexus eval-server --host 192.168.1.5 (specific interface)\n` +
` gitnexus eval-server --host localhost (OS-resolved loopback)\n`,
{ flag: '--host', value: rawHost },
);
process.exit(1);
}
const backend = new LocalBackend();
const ok = await backend.init();
@@ -426,12 +461,72 @@ export async function evalServerCommand(options?: EvalServerOptions): Promise<vo
}
});
server.listen(port, '127.0.0.1', () => {
server.on('error', (err: NodeJS.ErrnoException) => {
if (err.code === 'EADDRINUSE') {
cliError(
`\nGitNexus eval-server failed to start:\n` +
` Port ${port} is already in use.\n\n` +
` Either:\n` +
` 1. Stop the process already using port ${port}\n` +
` 2. Use a different port: gitnexus eval-server --port 4849\n`,
{ code: err.code, port, host },
);
} else if (err.code === 'EADDRNOTAVAIL') {
// "localhost" may resolve to ::1 on IPv6-only systems; treat it as
// potentially IPv6 so the user gets the right diagnostic hint.
const isIPv6Host = isIPv6(host) || host === 'localhost';
cliError(
`\nGitNexus eval-server failed to start:\n` +
` Address ${host} is not available on this machine.\n\n` +
(isIPv6Host
? ` Address ${host} resolved but is not reachable — IPv6 may be disabled, or the loopback interface may be unavailable.\n` +
` Docker containers and many CI environments disable IPv6 by default.\n\n`
: ` The --host value must be an IP assigned to a local network interface.\n` +
` Run \`ip addr\` (Linux) or \`ipconfig\` (Windows) to list available addresses.\n\n`) +
` Common fixes:\n` +
` gitnexus eval-server --host 127.0.0.1 (loopback, this machine only)\n` +
` gitnexus eval-server --host 0.0.0.0 (all interfaces, reachable from other VMs)\n`,
{ code: err.code, port, host },
);
} else if (err.code === 'EACCES') {
cliError(
`\nGitNexus eval-server failed to start:\n` +
` Permission denied binding to port ${port}.\n\n` +
` Ports below 1024 require elevated privileges.\n` +
` Use a port above 1024: gitnexus eval-server --port 4848\n`,
{ code: err.code, port, host },
);
} else {
cliError(`\nGitNexus eval-server failed to start:\n ${err.message}\n`, {
code: err.code,
port,
host,
});
}
process.exit(1);
});
server.listen(port, host, () => {
// Plain-text banner for the human watching stderr; structured record
// for log aggregation (split into two so the user sees a real banner
// not `{"level":30,"msg":"...","port":4747,"endpoints":[...]}`).
// Use server.address() so the banner and READY signal reflect what the OS
// actually bound to, not the input host string. This matters when "localhost"
// is passed: the OS may resolve it to ::1 on some systems.
const addr = server.address();
// server.listen callback only fires after a successful TCP bind, so
// server.address() is guaranteed to return an AddressInfo object here.
if (typeof addr !== 'object' || addr === null) {
cliError(
`\nGitNexus eval-server: unexpected server.address() value after bind: ${JSON.stringify(addr)}\n`,
);
process.exit(1);
}
const boundPort = addr.port;
const boundAddress = addr.address;
const displayHost = boundAddress.includes(':') ? `[${boundAddress}]` : boundAddress;
const bannerLines = [
`GitNexus eval-server: listening on http://127.0.0.1:${port}`,
`GitNexus eval-server: listening on http://${displayHost}:${boundPort}`,
` POST /tool/query — search execution flows`,
` POST /tool/context — 360-degree symbol view`,
` POST /tool/impact — blast radius analysis`,
@@ -443,8 +538,8 @@ export async function evalServerCommand(options?: EvalServerOptions): Promise<vo
bannerLines.push(` Auto-shutdown after ${idleTimeoutSec}s idle`);
}
cliInfo(bannerLines.join('\n'), {
port,
host: '127.0.0.1',
port: boundPort,
host,
idleTimeoutSec: idleTimeoutSec > 0 ? idleTimeoutSec : undefined,
endpoints: [
'POST /tool/query',
@@ -457,7 +552,7 @@ export async function evalServerCommand(options?: EvalServerOptions): Promise<vo
});
try {
// Use fd 1 directly — LadybugDB captures process.stdout (#324)
writeSync(1, `GITNEXUS_EVAL_SERVER_READY:${port}\n`);
writeSync(1, `GITNEXUS_EVAL_SERVER_READY:${displayHost}:${boundPort}\n`);
} catch {
// stdout may not be available (e.g., broken pipe)
}
+44 -1
View File
@@ -23,6 +23,7 @@ program
.command('analyze [path]')
.description('Index a repository (full analysis)')
.option('-f, --force', 'Force full re-index even if up to date')
.option('--repair-fts', 'Repair/rebuild search FTS indexes without full re-analysis')
.option(
'--embeddings [limit]',
'Enable embedding generation for semantic search (off by default). ' +
@@ -33,9 +34,20 @@ program
'Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` ' +
'preserves any embeddings already present in the index.',
)
.option('--skills', 'Generate repo-specific skill files from detected communities')
.option(
'--skills',
'Generate repo-specific skill files from detected communities ' +
'(no-op when --index-only is also set).',
)
.option('--skip-agents-md', 'Skip updating the gitnexus section in AGENTS.md and CLAUDE.md')
.option('--no-stats', 'Omit volatile file/symbol counts from AGENTS.md and CLAUDE.md')
.option(
'--skip-skills',
'Skip installing standard GitNexus skill files under .claude/skills/gitnexus/. ' +
'Does not suppress community skills from --skills (those use .claude/skills/generated/). ' +
'Use --index-only to skip all AI-context file injection.',
)
.option('--index-only', 'Pure index mode: skip all file injection (AGENTS.md, CLAUDE.md, skills)')
.option(
'--skip-git',
'Treat the provided path/cwd as the index root and skip parent git-root discovery',
@@ -59,6 +71,10 @@ program
'--worker-timeout <seconds>',
'Worker sub-batch idle timeout before retry/fallback. Default: 30.',
)
.option(
'--workers <n>',
'Parse worker pool size. Default: cores-1 capped at 16. Pass 0 to disable workers (sequential).',
)
.option('--embedding-threads <n>', 'Limit local ONNX embedding CPU threads')
.option('--embedding-batch-size <n>', 'Number of nodes per embedding batch')
.option('--embedding-sub-batch-size <n>', 'Number of chunks per embedding model call')
@@ -70,6 +86,11 @@ program
' GITNEXUS_MAX_FILE_SIZE=N Override large-file skip threshold (KB). Default 512, max 32768.\n' +
' GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS=N Worker idle timeout in milliseconds. Default 30000.\n' +
' GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES=N Worker job byte budget. Default 8388608.\n' +
' GITNEXUS_WORKER_POOL_SIZE=N Parse worker count override. Default cores-1 capped at 16.\n' +
' GITNEXUS_PARSE_CHUNK_CONCURRENCY=N Concurrent in-flight parse chunks. Default 2.\n' +
' GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT=N Max replacement spawns per slot before drop. Default 3.\n' +
' GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS=N Total retry wall-time per job. Default 5x sub-batch timeout.\n' +
' GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD=N Per-slot deaths to trip circuit breaker. Default max(3, poolSize).\n' +
' GITNEXUS_EMBEDDING_THREADS=N Limit local ONNX CPU threads for --embeddings.\n' +
' GITNEXUS_SEMANTIC_EXACT_SCAN_LIMIT=N Max embedding chunks for exact-scan fallback. Default 10000.\n' +
'\nTip: `.gitnexusignore` supports `.gitignore`-style negation. Add e.g.\n' +
@@ -150,9 +171,15 @@ program
)
.option('--no-reasoning-model', 'Disable reasoning model mode (overrides saved config)')
.option('--concurrency <n>', 'Parallel LLM calls (default: 3)', '3')
.option('--timeout <seconds>', 'LLM request timeout in seconds (default: disabled)')
.option('--retries <n>', 'Max LLM retry attempts per request (default: 3)')
.option('--gist', 'Publish wiki as a public GitHub Gist after generation')
.option('-v, --verbose', 'Enable verbose output (show LLM commands and responses)')
.option('--review', 'Stop after grouping to review module structure before generating pages')
.option(
'--lang <lang>',
'Output language for generated documentation (e.g. english, chinese, spanish, japanese)',
)
.action(createLazyAction(() => import('./wiki.js'), 'wikiCommand'));
program
@@ -160,6 +187,18 @@ program
.description('Augment a search pattern with knowledge graph context (used by hooks)')
.action(createLazyAction(() => import('./augment.js'), 'augmentCommand'));
program
.command('publish [path]')
.description(
'Notify the understand-quickly registry that this repo has a fresh GitNexus index. ' +
'Opt-in: requires UNDERSTAND_QUICKLY_TOKEN (fine-grained PAT with ' +
'`Repository dispatches: write` on looptech-ai/understand-quickly). ' +
'No-op without the token. See https://github.com/looptech-ai/understand-quickly.',
)
.option('--id <owner/repo>', 'Override the registry id (defaults to the origin remote)')
.option('--skip-git', 'Treat cwd as the repo root and skip parent git-root discovery')
.action(createLazyAction(() => import('./publish.js'), 'publishCommand'));
// ─── Direct Tool Commands (no MCP overhead) ────────────────────────
// These invoke LocalBackend directly for use in eval, scripts, and CI.
@@ -212,6 +251,10 @@ program
.command('eval-server')
.description('Start lightweight HTTP server for fast tool calls during evaluation')
.option('-p, --port <port>', 'Port number', '4848')
.option(
'--host <host>',
'Bind address (default: 127.0.0.1, use 0.0.0.0 to expose to all interfaces)',
)
.option('--idle-timeout <seconds>', 'Auto-shutdown after N seconds idle (0 = disabled)', '0')
.action(createLazyAction(() => import('./eval-server.js'), 'evalServerCommand'));
+12 -8
View File
@@ -1,15 +1,18 @@
/**
* Optional grammar availability check.
*
* tree-sitter-dart and tree-sitter-proto are optionalDependencies that
* require a `node-gyp rebuild` at install time. The build can be skipped
* via GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1 (postinstall scripts), or it can
* silently soft-fail when the C++ toolchain is missing.
* tree-sitter-dart, tree-sitter-proto, and tree-sitter-swift are vendored
* under vendor/ and materialized into node_modules/ at postinstall. Dart
* and Proto are built from source with node-gyp; Swift ships platform
* prebuilds activated via node-gyp-build. All three can be skipped via
* GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1 (postinstall scripts), or can silently
* soft-fail when the toolchain is missing (Dart/Proto) or no prebuild
* matches the host platform (Swift).
*
* Either path produces the same observable: the .node binding is absent
* at runtime. This helper detects that condition and surfaces a single
* stderr line per missing grammar so users learn why .dart/.proto support
* is unavailable instead of silently getting a degraded index.
* stderr line per missing grammar so users learn why .dart/.proto/.swift
* support is unavailable instead of silently getting a degraded index.
*/
import { createRequire } from 'module';
@@ -29,6 +32,7 @@ interface OptionalGrammar {
const OPTIONAL_GRAMMARS: OptionalGrammar[] = [
{ name: 'tree-sitter-dart', pkg: 'tree-sitter-dart', extensions: ['.dart'] },
{ name: 'tree-sitter-proto', pkg: 'tree-sitter-proto', extensions: ['.proto'] },
{ name: 'tree-sitter-swift', pkg: 'tree-sitter-swift', extensions: ['.swift'] },
];
export interface MissingGrammar {
@@ -40,8 +44,8 @@ export interface MissingGrammar {
* Returns the list of optional grammars whose native binding cannot be
* loaded. Actually `require()`s the package — `require.resolve` would
* locate the entry path even when the `.node` binding is absent (the
* `file:` package directory is installed regardless of postinstall
* outcome), giving false negatives for the exact users we want to warn:
* package directory exists without a working `.node` binding), giving false
* negatives for the exact users we want to warn:
* those who installed with `GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1` or whose
* native rebuild soft-failed for missing toolchain.
*
+232
View File
@@ -0,0 +1,232 @@
/**
* `gitnexus publish` — opt-in ping to the understand-quickly registry.
*
* Fires a single `repository_dispatch` event at
* `looptech-ai/understand-quickly` so the registry knows to refresh its
* entry for the current repo. Does NOT upload anything: per the
* understand-quickly protocol, the registry pulls the graph from a
* raw-GitHub URL the user controls.
*
* https://github.com/looptech-ai/understand-quickly/blob/main/docs/integrations/protocol.md
*
* Defaults:
* - Without `UNDERSTAND_QUICKLY_TOKEN` in the env, this is a no-op
* (prints one informational line, exit 0). Same shape as the
* `--publish` patterns in sibling tools.
* - With the token, fires the dispatch and reports the response code.
*
* The `id` is derived from the repo's `origin` remote unless the caller
* passes `--id <owner/repo>` explicitly. We deliberately do NOT auto-add
* the repo to the registry — registration is one-time and uses the
* `npx @understand-quickly/cli add` path documented in the protocol.
*/
import path from 'path';
import {
UNDERSTAND_QUICKLY_DISPATCH_URL,
UNDERSTAND_QUICKLY_TOKEN_ENV,
buildUqDispatchPayload,
isValidOwnerRepo,
parseOwnerRepoFromRemote,
} from 'gitnexus-shared';
import { getGitRoot, getRemoteOriginUrl, getCurrentCommit } from '../storage/git.js';
import { hasIndex } from '../storage/repo-manager.js';
import { cliInfo, cliError } from './cli-message.js';
export interface PublishOptions {
/** Override the auto-derived `owner/repo` id. */
id?: string;
/** Treat the cwd as the repo root (skip git-root walk). */
skipGit?: boolean;
}
const REGISTER_HINT =
'Register your repo once with: npx @understand-quickly/cli add\n' +
'Or use the wizard: https://looptech-ai.github.io/understand-quickly/add.html';
/**
* Hard cap on the dispatch fetch to keep CI publish steps from stalling
* for the OS TCP timeout (~2 min) when api.github.com is unreachable.
* Matches the pattern used in `src/core/embeddings/http-client.ts`.
*/
const DISPATCH_TIMEOUT_MS = 15_000;
export const publishCommand = async (
inputPath?: string,
options: PublishOptions = {},
): Promise<void> => {
// ── 0. Token gate FIRST — guarantees true no-op without the token. ──
// The README, CLI --help, and PR body all promise "exit 0 without
// UNDERSTAND_QUICKLY_TOKEN". Doing the index/repo-root checks before
// the token gate would make those promises false for users who haven't
// run `gitnexus analyze` yet but want to verify the command is wired.
const token = process.env[UNDERSTAND_QUICKLY_TOKEN_ENV];
if (!token) {
cliInfo(
`[understand-quickly] ${UNDERSTAND_QUICKLY_TOKEN_ENV} is not set — skipping dispatch.\n` +
`Set it to a fine-grained PAT with "Repository dispatches: write" on ` +
`looptech-ai/understand-quickly to enable instant resync.\n` +
`(Without the token, the registry's nightly sync still picks up your entry.)`,
{ skipped: 'no-token' },
);
return;
}
// ── 1. Resolve the repo root (same precedence as `analyze`) ──────────
let repoPath: string;
if (inputPath) {
repoPath = path.resolve(inputPath);
} else if (options.skipGit) {
repoPath = path.resolve(process.cwd());
} else {
const gitRoot = getGitRoot(process.cwd());
if (!gitRoot) {
cliError(
'[understand-quickly] not inside a git repository.\n' +
'Run from a repo, or pass --skip-git to publish from the current directory.',
);
process.exitCode = 1;
return;
}
repoPath = gitRoot;
}
// ── 2. Confirm a GitNexus index exists ───────────────────────────────
// Publishing without an index is almost always a mistake — the
// registry's nightly sync would fetch a stale or missing graph file
// and mark the entry `missing`. Refuse loudly with a fix-it hint.
if (!(await hasIndex(repoPath))) {
cliError(
`[understand-quickly] no GitNexus index found at ${repoPath}/.gitnexus.\n` +
'Run `gitnexus analyze` first, then re-run `gitnexus publish`.',
);
process.exitCode = 1;
return;
}
// ── 3. Derive the registry id ─────────────────────────────────────────
const id =
options.id ?? parseOwnerRepoFromRemote(getRemoteOriginUrl(repoPath) ?? undefined) ?? null;
if (!id || !isValidOwnerRepo(id)) {
cliError(
`[understand-quickly] could not derive a registry id from this repo.\n` +
`Pass --id <owner/repo> explicitly (e.g. --id looptech-ai/${path.basename(repoPath)}).\n` +
REGISTER_HINT,
);
process.exitCode = 1;
return;
}
// ── 4. Fire the dispatch ─────────────────────────────────────────────
const payload = buildUqDispatchPayload(id);
let response: Response;
try {
response = await fetch(UNDERSTAND_QUICKLY_DISPATCH_URL, {
method: 'POST',
headers: {
Accept: 'application/vnd.github+json',
Authorization: `Bearer ${token}`,
'X-GitHub-Api-Version': '2022-11-28',
'Content-Type': 'application/json',
'User-Agent': 'gitnexus-cli',
},
body: JSON.stringify(payload),
signal: AbortSignal.timeout(DISPATCH_TIMEOUT_MS),
});
} catch (err) {
// `AbortSignal.timeout()` throws a `DOMException` with `name ===
// 'TimeoutError'` on Node 18.14+ (and on browsers/Bun). It is NOT
// a plain `AbortError`. Match the pattern used in
// gitnexus/src/core/embeddings/http-client.ts so the user sees the
// targeted "timed out" message instead of a generic "operation
// was aborted".
const isTimeout = err instanceof DOMException && err.name === 'TimeoutError';
if (isTimeout) {
cliError(
`[understand-quickly] dispatch timed out after ${DISPATCH_TIMEOUT_MS}ms. ` +
`Check network access to api.github.com and retry.`,
{ id },
);
} else {
const msg = err instanceof Error ? err.message : String(err);
cliError(`[understand-quickly] dispatch network error: ${msg}`, { id });
}
process.exitCode = 1;
return;
}
// GitHub returns 204 on success. Distinct branches for 401/403/404/422
// so users debug without checking the docs.
if (response.status === 204) {
await response.body?.cancel().catch(() => {});
// `getCurrentCommit` is only meaningful in the success path — moving
// it inside this branch removes a wasted child-process spawn on every
// error response (LOW 7).
const commit = getCurrentCommit(repoPath);
cliInfo(
`[understand-quickly] dispatched sync-entry for ${id}` +
(commit ? ` @ ${commit.slice(0, 7)}` : '') +
'.\n' +
`Note: a 204 only confirms GitHub accepted the dispatch. Whether the ` +
`registry workflow finds an entry for "${id}" is logged at ` +
`https://github.com/looptech-ai/understand-quickly/actions/workflows/sync.yml`,
{ id, commit, status: response.status },
);
return;
}
if (response.status === 401) {
cliError(
`[understand-quickly] dispatch returned 401 — the ${UNDERSTAND_QUICKLY_TOKEN_ENV} value is invalid or expired.\n` +
`Regenerate a fine-grained PAT at https://github.com/settings/personal-access-tokens ` +
`with Repository access scoped to looptech-ai/understand-quickly and the ` +
`"Repository dispatches: write" permission, then retry.`,
{ id, status: response.status },
);
process.exitCode = 1;
return;
}
if (response.status === 403) {
cliError(
`[understand-quickly] dispatch returned 403 — the token authenticated but ` +
`lacks the "Repository dispatches: write" permission on ` +
`looptech-ai/understand-quickly. Edit the PAT scopes and retry.`,
{ id, status: response.status },
);
process.exitCode = 1;
return;
}
if (response.status === 404) {
cliError(
`[understand-quickly] dispatch returned 404 — the token cannot reach ` +
`looptech-ai/understand-quickly. Verify the PAT has Repository access to ` +
`that exact repo (not just your own org).`,
{ id, status: response.status },
);
process.exitCode = 1;
return;
}
if (response.status === 422) {
// Malformed event_type / client_payload — a code bug in this CLI,
// not a user mistake. Surface so we get bug reports.
const body422 = await response.text().catch(() => '');
cliError(
`[understand-quickly] dispatch returned 422 (this is a CLI bug; please report).\n` +
`Body: ${body422 || '(empty)'}`,
{ id, status: response.status },
);
process.exitCode = 1;
return;
}
// 5xx and anything else → bubble the body so the user has something to act on.
const body = await response.text().catch(() => '');
cliError(
`[understand-quickly] dispatch failed with HTTP ${response.status}: ${body || '(empty body)'}`,
{ id, status: response.status },
);
process.exitCode = 1;
};
+8 -1
View File
@@ -1,6 +1,7 @@
import { createServer } from '../server/api.js';
import { logger, flushLoggerSync } from '../core/logger.js';
import { cliError } from './cli-message.js';
import { isWalCorruptionError, WAL_RECOVERY_SUGGESTION } from '../core/lbug/lbug-config.js';
// Catch anything that would cause a silent exit. Pino v10's default
// destination is `sync: false` (SonicBoom buffered) — call
@@ -34,7 +35,13 @@ export const serveCommand = async (options?: { port?: string; host?: string }) =
try {
await createServer(port, host);
} catch (err: any) {
if (err.code === 'EADDRINUSE') {
if (isWalCorruptionError(err)) {
cliError(
`\nGitNexus server could not start: the index has a corrupted WAL file.\n` +
` ${WAL_RECOVERY_SUGGESTION}\n`,
{ recoveryHint: 'wal-corruption' },
);
} else if (err.code === 'EADDRINUSE') {
cliError(
`\nFailed to start GitNexus server:\n` +
` ${err.message || err}\n\n` +
+33 -4
View File
@@ -61,11 +61,13 @@ function resolveGitnexusBin(): string | null {
.filter(Boolean);
if (isWin) {
// On Windows, `where` returns multiple entries (e.g. the POSIX shell
// script AND the .cmd/.bat wrapper). Prefer the wrapper because
// child_process.spawn() cannot execute a shell script directly.
// On Windows, npm global installs can surface multiple launchers for the
// same package (e.g. a POSIX shell shim plus .cmd/.bat wrappers). Claude
// and the other MCP hosts need a directly spawnable command path, so only
// accept the Windows wrapper. If it is missing, fall back to the slower
// npx entry instead of persisting a non-spawnable shim path.
const cmdLine = lines.find((l) => /\.(cmd|bat)$/i.test(l));
return cmdLine || lines[0] || null;
return cmdLine || null;
}
return lines[0] || null;
@@ -364,6 +366,33 @@ async function installClaudeCodeHooks(result: SetupResult): Promise<void> {
// Script not found in source — skip
}
try {
await fs.copyFile(
path.join(pluginHooksPath, 'hook-lock.cjs'),
path.join(destHooksDir, 'hook-lock.cjs'),
);
} catch {
// Helper not found in source — skip
}
try {
await fs.copyFile(
path.join(pluginHooksPath, 'hook-db-lock-probe.cjs'),
path.join(destHooksDir, 'hook-db-lock-probe.cjs'),
);
} catch {
// Helper not found in source — skip
}
try {
await fs.copyFile(
path.join(pluginHooksPath, 'win-rm-list-json.ps1'),
path.join(destHooksDir, 'win-rm-list-json.ps1'),
);
} catch {
// Helper not found in source — skip
}
const hookPath = path.join(destHooksDir, 'gitnexus-hook.cjs').replace(/\\/g, '/');
// Escape backslashes FIRST, then quotes (CodeQL js/incomplete-sanitization).
// The previous shape `replace(/"/g, '\\"')` alone would let `path\with"quote`
+60
View File
@@ -33,6 +33,26 @@ export interface WikiCommandOptions {
provider?: LLMProvider;
verbose?: boolean;
review?: boolean;
timeout?: string;
retries?: string;
lang?: string;
}
function parsePositiveIntegerOption(
value: string | undefined,
flag: string,
multiplier = 1,
): number | undefined {
if (value === undefined) return undefined;
const trimmed = value.trim();
if (!/^[1-9]\d*$/.test(trimmed)) {
throw new Error(`${flag} must be a positive integer`);
}
const parsed = parseInt(trimmed, 10);
if (parsed > Math.floor(Number.MAX_SAFE_INTEGER / multiplier)) {
throw new Error(`${flag} is too large`);
}
return parsed;
}
/**
@@ -87,6 +107,24 @@ function prompt(question: string, hide = false): Promise<string> {
}
export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptions) => {
// Snapshot GITNEXUS_VERBOSE at entry — wikiCommand mutates it (the impl
// below) so cursor-client (process.env-driven) sees the right value during
// this run. Restored in finally so back-to-back wiki calls in long-running
// hosts don't leak verbose state from one invocation to the next. Pairs
// with the same snapshot/restore pattern in `analyzeCommand`.
const originalVerbose = process.env.GITNEXUS_VERBOSE;
try {
await wikiCommandImpl(inputPath, options);
} finally {
if (originalVerbose === undefined) {
delete process.env.GITNEXUS_VERBOSE;
} else {
process.env.GITNEXUS_VERBOSE = originalVerbose;
}
}
};
const wikiCommandImpl = async (inputPath?: string, options?: WikiCommandOptions): Promise<void> => {
// Set verbose mode globally for cursor-client to pick up
if (options?.verbose) {
process.env.GITNEXUS_VERBOSE = '1';
@@ -125,6 +163,17 @@ export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptio
return;
}
let timeoutSeconds: number | undefined;
let retries: number | undefined;
try {
timeoutSeconds = parsePositiveIntegerOption(options?.timeout, '--timeout', 1000);
retries = parsePositiveIntegerOption(options?.retries, '--retries');
} catch (error) {
console.log(` Error: ${(error as Error).message}\n`);
process.exitCode = 1;
return;
}
// ── Resolve LLM config (with interactive fallback) ─────────────────
// Save any CLI overrides immediately
if (
@@ -347,6 +396,14 @@ export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptio
}
}
// ── Apply per-run overrides not saved to config ────────────────────
if (timeoutSeconds !== undefined) {
llmConfig.requestTimeoutMs = timeoutSeconds * 1000;
}
if (retries !== undefined) {
llmConfig.maxAttempts = retries;
}
// ── Setup progress bar with elapsed timer ──────────────────────────
const bar = new cliProgress.SingleBar(
{
@@ -383,6 +440,7 @@ export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptio
force: options?.force,
concurrency: options?.concurrency ? parseInt(options.concurrency, 10) : undefined,
reviewOnly: options?.review,
lang: options?.lang,
};
const generator = new WikiGenerator(
@@ -551,6 +609,8 @@ export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptio
if (err.message?.includes('No source files')) {
console.log(`\n ${err.message}\n`);
} else if (err.message?.includes('LLM request timed out after')) {
console.log(`\n Timeout: ${err.message}\n`);
} else if (err.message?.includes('content filter')) {
// Content filter block — actionable message
console.log(`\n Content Filter: ${err.message}\n`);
+2
View File
@@ -25,6 +25,8 @@ const DEFAULT_IGNORE_LIST = new Set([
'bower_components',
'jspm_packages',
'vendor', // PHP/Go
'third_party', // C/C++ (Google-style vendored dependencies)
'3rdparty', // C/C++ (alternate spelling, also Qt convention)
// 'packages' removed - commonly used for monorepo source code (lerna, pnpm, yarn workspaces)
'venv',
'.venv',
+30 -6
View File
@@ -2,8 +2,8 @@
* Augmentation Engine
*
* Lightweight, fast-path enrichment of search patterns with knowledge graph context.
* Designed to be called from platform hooks (Claude Code PreToolUse, Cursor beforeShellExecution)
* when an agent runs grep/glob/search.
* Designed to be called from platform hooks (Claude Code PreToolUse, Cursor postToolUse)
* when an agent runs grep/glob/read/search.
*
* Performance target: <500ms cold start, <200ms warm.
*
@@ -86,6 +86,9 @@ async function findRepoForCwd(cwd: string): Promise<{
export async function augment(pattern: string, cwd?: string): Promise<string> {
if (!pattern || pattern.length < 3) return '';
const patternFirstWord = pattern.trim().replace(/'/g, "''").split(/\s+/)[0];
if (!patternFirstWord || patternFirstWord.length < 2) return '';
const workDir = cwd || process.cwd();
try {
@@ -104,9 +107,7 @@ export async function augment(pattern: string, cwd?: string): Promise<string> {
}
// Step 1: BM25 search (fast, no embeddings)
const bm25Results = await searchFTSFromLbug(pattern, 10, repoId);
if (bm25Results.length === 0) return '';
const { results: bm25Results, ftsAvailable } = await searchFTSFromLbug(pattern, 10, repoId);
// Step 2: Map BM25 file results to symbols
const symbolMatches: Array<{
@@ -124,7 +125,7 @@ export async function augment(pattern: string, cwd?: string): Promise<string> {
repoId,
`
MATCH (n) WHERE n.filePath = '${escaped}'
AND n.name CONTAINS '${pattern.replace(/'/g, "''").split(/\s+/)[0]}'
AND n.name CONTAINS '${patternFirstWord}'
RETURN n.id AS id, n.name AS name, labels(n)[0] AS type, n.filePath AS filePath
LIMIT 3
`,
@@ -143,6 +144,29 @@ export async function augment(pattern: string, cwd?: string): Promise<string> {
}
}
// When FTS indexes are unavailable (read-only DB, first run before indexes are built),
// fall back to a direct name CONTAINS query so enrichment still works.
if (symbolMatches.length === 0 && !ftsAvailable) {
const fallbackRows = await executeQuery(
repoId,
`
MATCH (n)
WHERE n.name CONTAINS '${patternFirstWord}'
RETURN n.id AS id, n.name AS name, labels(n)[0] AS type, n.filePath AS filePath
LIMIT 5
`,
).catch(() => []);
for (const sym of fallbackRows) {
symbolMatches.push({
nodeId: sym.id || sym[0],
name: sym.name || sym[1],
type: sym.type || sym[2],
filePath: sym.filePath || sym[3],
score: 1.0,
});
}
}
if (symbolMatches.length === 0) return '';
// Step 3: Batch-fetch callers/callees/processes/cohesion for top matches
+2 -1
View File
@@ -10,6 +10,7 @@ import {
isLanguageAvailable,
resolveLanguageKey,
} from '../tree-sitter/parser-loader.js';
import { parseSourceSafe } from '../tree-sitter/safe-parse.js';
const parserCache = new Map<string, any>();
@@ -29,7 +30,7 @@ export const ensureAndParse = async (content: string, filePath: string): Promise
parserCache.set(parserKey, parserInstance);
}
return parserInstance.parse(content);
return parseSourceSafe(parserInstance, content);
};
const FUNCTION_LIKE_TYPES = new Set([
+75 -104
View File
@@ -1,6 +1,8 @@
import os from 'node:os';
import { join } from 'node:path';
import { CircuitBreaker, withRetry } from 'gitnexus-shared';
// ---------------------------------------------------------------------------
// Download resilience defaults
// ---------------------------------------------------------------------------
@@ -108,70 +110,19 @@ export function isNetworkFetchError(message: string): boolean {
/** @internal Used by `withHfDownloadRetry` to mark a circuit-open rejection. */
export const CIRCUIT_OPEN_TAG = 'hf-circuit-open';
/** Circuit-breaker states. */
type CircuitState = 'closed' | 'open' | 'half-open';
/**
* Circuit breaker for HuggingFace model downloads.
*
* After `failureThreshold` consecutive network failures the circuit opens and
* all subsequent calls to `withHfDownloadRetry` fail immediately without
* issuing any network requests. After `resetTimeoutMs` the circuit enters the
* half-open state and the next call is attempted — if it succeeds the circuit
* closes again; if it fails the circuit re-opens.
*
* Exported for unit-testing; production code should use the module-level
* `hfDownloadCircuit` singleton.
* Module-level singleton shared by both embedder entry points
* (`core/embeddings/embedder.ts` + `mcp/core/embedder.ts`). Per-process
* only — not persisted across restarts. Backed by the shared
* `CircuitBreaker` from `gitnexus-shared` (same state machine, same
* semantics, plus the single-permit half-open gate that prevents
* recovery-time stampedes).
*/
export class HfDownloadCircuitBreaker {
private _state: CircuitState = 'closed';
private _failures = 0;
/** Timestamp of the last recorded failure (ms since epoch). */
lastFailureAt = 0;
constructor(
readonly failureThreshold: number = CB_FAILURE_THRESHOLD,
readonly resetTimeoutMs: number = CB_RESET_TIMEOUT_MS,
) {}
/** Effective state, factoring in the reset-timeout transition. */
get state(): CircuitState {
if (this._state === 'open' && Date.now() - this.lastFailureAt > this.resetTimeoutMs) {
this._state = 'half-open';
}
return this._state;
}
/** Returns true when the circuit is open and calls should be rejected. */
isOpen(): boolean {
return this.state === 'open';
}
/** Record a successful call — resets the failure counter and closes the circuit. */
recordSuccess(): void {
this._failures = 0;
this._state = 'closed';
}
/** Record a failed call — increments the counter and opens the circuit when the threshold is reached. */
recordFailure(): void {
this._failures++;
this.lastFailureAt = Date.now();
if (this._failures >= this.failureThreshold) {
this._state = 'open';
}
}
/** @internal Reset to initial state (used in tests). */
reset(): void {
this._failures = 0;
this._state = 'closed';
this.lastFailureAt = 0;
}
}
/** Module-level singleton shared by both embedder entry points. */
export const hfDownloadCircuit = new HfDownloadCircuitBreaker();
export const hfDownloadCircuit = new CircuitBreaker({
failureThreshold: CB_FAILURE_THRESHOLD,
cooldownMs: CB_RESET_TIMEOUT_MS,
key: 'hf-download',
});
// ---------------------------------------------------------------------------
// Retry + timeout wrapper
@@ -219,11 +170,6 @@ export function withDownloadTimeout<T>(fn: () => Promise<T>, timeoutMs: number):
});
}
/** @internal Async sleep (exposed for testing). */
export function sleep(ms: number): Promise<void> {
return new Promise((resolve) => setTimeout(resolve, ms));
}
export interface HfRetryOptions {
/** Maximum total attempts including the initial one (default: `HF_MAX_ATTEMPTS`). */
maxAttempts?: number;
@@ -235,7 +181,7 @@ export interface HfRetryOptions {
* Circuit-breaker instance to use. Defaults to the module-level
* `hfDownloadCircuit` singleton. Pass a fresh instance in tests.
*/
circuit?: HfDownloadCircuitBreaker;
circuit?: CircuitBreaker;
/**
* Optional callback invoked before each retry (not the initial attempt).
* @param attempt - 1-based retry number
@@ -295,49 +241,74 @@ export async function withHfDownloadRetry<T>(
circuit = hfDownloadCircuit,
onRetry,
} = options;
if (circuit.isOpen()) {
const secsUntilReset = Math.ceil(
(circuit.resetTimeoutMs - (Date.now() - circuit.lastFailureAt)) / 1000,
);
if (circuit.getState() === 'open') {
// Compute remaining cooldown without consuming a probe permit.
const openedAt = circuit.getOpenedAt();
const secsUntilReset =
openedAt !== null ? Math.ceil((circuit.getCooldownMs() - (Date.now() - openedAt)) / 1000) : 0;
throw new Error(
`${CIRCUIT_OPEN_TAG}: HuggingFace download circuit is open after repeated network failures` +
(secsUntilReset > 0 ? ` — will reset in ~${secsUntilReset}s` : ''),
);
}
let lastError: Error = new Error('unknown error');
// Retry budget delegated to `withRetry` from gitnexus-shared. The
// HF-specific bits — per-attempt timeout, network-vs-non-network
// classification, circuit-breaker recording, onRetry callback — wire
// through the `isRetryable` callback. `circuitTripped` is the
// sentinel that lets us replace the final thrown error with a
// CIRCUIT_OPEN_TAG message when the breaker tripped mid-loop.
let circuitTripped = false;
for (let attempt = 0; attempt < maxAttempts; attempt++) {
try {
const result = await withDownloadTimeout(fn, timeoutMs);
circuit.recordSuccess();
return result;
} catch (err) {
lastError = err instanceof Error ? err : new Error(String(err));
if (!isNetworkFetchError(lastError.message)) {
// Non-network error (e.g. CUDA unavailable) — propagate without retry
throw lastError;
}
circuit.recordFailure();
if (circuit.isOpen()) {
// Circuit just tripped — fail fast, no more retries
throw new Error(
`${CIRCUIT_OPEN_TAG}: HuggingFace download circuit opened after ${circuit.failureThreshold} consecutive failures`,
);
}
if (attempt < maxAttempts - 1) {
const delay = baseDelayMs * Math.pow(2, attempt);
onRetry?.(attempt + 1, maxAttempts, lastError);
await sleep(delay);
}
try {
return await withRetry(
async () => {
const result = await withDownloadTimeout(fn, timeoutMs);
circuit.recordSuccess();
return result;
},
{
maxAttempts,
baseDelayMs,
// Disable the cap to match the bespoke pure-exponential
// progression. With the default `HF_MAX_ATTEMPTS_CAP = 10` and
// `baseDelayMs = 2000`, the largest possible delay is
// `2000 * 2^9 = ~17 minutes` — bounded enough not to need a cap.
capDelayMs: Number.MAX_SAFE_INTEGER,
isRetryable: (err, attempt) => {
const error = err instanceof Error ? err : new Error(String(err));
if (!isNetworkFetchError(error.message)) {
// Non-network error (e.g. CUDA unavailable) — propagate
// without retry. Use recordNeutral so the breaker's existing
// failure-count progress isn't reset by a non-network failure
// that says nothing about the CDN's health.
circuit.recordNeutral();
return { retry: false };
}
circuit.recordFailure();
if (circuit.getState() === 'open') {
// Circuit just tripped — fail fast, no more retries.
circuitTripped = true;
return { retry: false };
}
// Mirror the bespoke onRetry contract: fire only when there's
// actually a next attempt.
if (attempt + 1 < maxAttempts) {
onRetry?.(attempt + 1, maxAttempts, error);
}
return { retry: true };
},
},
);
} catch (err) {
if (circuitTripped) {
throw new Error(
`${CIRCUIT_OPEN_TAG}: HuggingFace download circuit opened after ${CB_FAILURE_THRESHOLD} consecutive failures`,
);
}
// All retries exhausted — rethrow the last network error so
// isNetworkFetchError patterns in the calling code still match and
// surface HF_ENDPOINT guidance.
throw err;
}
// All retries exhausted — throw the last network error so isNetworkFetchError
// patterns in the calling code still match and surface HF_ENDPOINT guidance.
throw lastError;
}
+75 -29
View File
@@ -3,13 +3,22 @@
*
* Shared fetch+retry logic for OpenAI-compatible /v1/embeddings endpoints.
* Imported by both the core embedder (batch) and MCP embedder (query).
*
* Network resilience is delegated to `resilientFetch` from
* `gitnexus-shared` — bounded retries with exponential-backoff jitter,
* `Retry-After` honored on 429, and an in-process circuit breaker that
* fails fast on a flapping endpoint. Per-attempt timeout is enforced
* via `AbortSignal.timeout` on the underlying fetch.
*/
import { CircuitOpenError, ResilientFetchExhaustedError, resilientFetch } from 'gitnexus-shared';
const HTTP_TIMEOUT_MS = 30_000;
const HTTP_MAX_RETRIES = 2;
const HTTP_RETRY_BACKOFF_MS = 1_000;
const HTTP_BATCH_SIZE = 64;
const DEFAULT_DIMS = 384;
const HTTP_BREAKER_KEY = 'embeddings-http';
interface HttpConfig {
baseUrl: string;
@@ -31,8 +40,11 @@ const readConfig = (): HttpConfig | null => {
const rawDims = process.env.GITNEXUS_EMBEDDING_DIMS;
let dimensions: number | undefined;
if (rawDims !== undefined) {
if (!/^\d+$/.test(rawDims)) {
throw new Error(`GITNEXUS_EMBEDDING_DIMS must be a positive integer, got "${rawDims}"`);
}
const parsed = parseInt(rawDims, 10);
if (Number.isNaN(parsed) || parsed <= 0) {
if (parsed <= 0) {
throw new Error(`GITNEXUS_EMBEDDING_DIMS must be a positive integer, got "${rawDims}"`);
}
dimensions = parsed;
@@ -82,7 +94,13 @@ interface EmbeddingItem {
* @param model - Model name for the request body
* @param apiKey - Bearer token (only used in Authorization header)
* @param batchIndex - Logical batch number (for error context)
* @param attempt - Current retry attempt (internal)
* @param dimensions - Optional output-vector size. When provided, sent as
* the `dimensions` field in the request body. Endpoints that implement
* Matryoshka truncation (OpenAI text-embedding-3-*, Cohere embed-v3,
* Voyage) return a truncated vector at that size; endpoints that do not
* recognise the field may ignore it or return 400. Leave
* `GITNEXUS_EMBEDDING_DIMS` unset for strict backends that reject
* unknown fields.
*/
const httpEmbedBatch = async (
url: string,
@@ -90,46 +108,60 @@ const httpEmbedBatch = async (
model: string,
apiKey: string,
batchIndex = 0,
attempt = 0,
dimensions?: number,
): Promise<EmbeddingItem[]> => {
const requestBody: { input: string[]; model: string; dimensions?: number } = {
input: batch,
model,
};
if (dimensions !== undefined) {
requestBody.dimensions = dimensions;
}
let resp: Response;
try {
resp = await fetch(url, {
method: 'POST',
signal: AbortSignal.timeout(HTTP_TIMEOUT_MS),
headers: {
'Content-Type': 'application/json',
Authorization: `Bearer ${apiKey}`,
resp = await resilientFetch(
url,
{
method: 'POST',
signal: AbortSignal.timeout(HTTP_TIMEOUT_MS),
headers: {
'Content-Type': 'application/json',
Authorization: `Bearer ${apiKey}`,
},
body: JSON.stringify(requestBody),
},
body: JSON.stringify({ input: batch, model }),
});
{
breakerKey: HTTP_BREAKER_KEY,
retry: { maxAttempts: HTTP_MAX_RETRIES + 1, baseDelayMs: HTTP_RETRY_BACKOFF_MS },
},
);
} catch (err) {
// Timeouts should not be retried — the server is unresponsive.
// AbortSignal.timeout() throws DOMException with name 'TimeoutError'.
const isTimeout = err instanceof DOMException && err.name === 'TimeoutError';
if (isTimeout) {
if (err instanceof CircuitOpenError) {
throw new Error(
`Embedding endpoint circuit open (${safeUrl(url)}, batch ${batchIndex}): retry in ${Math.ceil(err.retryAfterMs / 1000)}s`,
);
}
if (err instanceof DOMException && err.name === 'TimeoutError') {
throw new Error(
`Embedding request timed out after ${HTTP_TIMEOUT_MS}ms (${safeUrl(url)}, batch ${batchIndex})`,
);
}
// DNS, connection errors — retry with backoff
if (attempt < HTTP_MAX_RETRIES) {
const delay = HTTP_RETRY_BACKOFF_MS * (attempt + 1);
await new Promise((r) => setTimeout(r, delay));
return httpEmbedBatch(url, batch, model, apiKey, batchIndex, attempt + 1);
if (err instanceof ResilientFetchExhaustedError) {
throw new Error(
`Embedding endpoint returned ${err.response.status} (${safeUrl(url)}, batch ${batchIndex})`,
);
}
const reason = err instanceof Error ? err.message : String(err);
throw new Error(`Embedding request failed (${safeUrl(url)}, batch ${batchIndex}): ${reason}`);
}
if (!resp.ok) {
const status = resp.status;
if ((status === 429 || status >= 500) && attempt < HTTP_MAX_RETRIES) {
const delay = HTTP_RETRY_BACKOFF_MS * (attempt + 1);
await new Promise((r) => setTimeout(r, delay));
return httpEmbedBatch(url, batch, model, apiKey, batchIndex, attempt + 1);
}
throw new Error(`Embedding endpoint returned ${status} (${safeUrl(url)}, batch ${batchIndex})`);
// resilientFetch already retried 5xx/429; any non-OK response here is
// a terminal client error (4xx other than 429).
throw new Error(
`Embedding endpoint returned ${resp.status} (${safeUrl(url)}, batch ${batchIndex})`,
);
}
const data = (await resp.json()) as { data: EmbeddingItem[] };
@@ -155,7 +187,14 @@ export const httpEmbed = async (texts: string[]): Promise<Float32Array[]> => {
for (let i = 0; i < texts.length; i += HTTP_BATCH_SIZE) {
const batch = texts.slice(i, i + HTTP_BATCH_SIZE);
const batchIndex = Math.floor(i / HTTP_BATCH_SIZE);
const items = await httpEmbedBatch(url, batch, config.model, config.apiKey, batchIndex);
const items = await httpEmbedBatch(
url,
batch,
config.model,
config.apiKey,
batchIndex,
config.dimensions,
);
if (items.length !== batch.length) {
throw new Error(
@@ -198,7 +237,14 @@ export const httpEmbedQuery = async (text: string): Promise<number[]> => {
if (!config) throw new Error('HTTP embedding not configured');
const url = `${config.baseUrl}/embeddings`;
const items = await httpEmbedBatch(url, [text], config.model, config.apiKey);
const items = await httpEmbedBatch(
url,
[text],
config.model,
config.apiKey,
0,
config.dimensions,
);
if (!items.length) {
throw new Error(`Embedding endpoint returned empty response (${safeUrl(url)})`);
}

Some files were not shown because too many files have changed in this diff Show More