* fix(mcp): key the empty-ascent note on CALL_SUMMARY data, not language (#2802) `pdg-impact.ts` decided whether to append a "return-value ascent is TypeScript/JavaScript-only" caveat to the `impact(mode:'pdg')` note by looking up the criterion file's language. That put language-specific logic in a layer that must be language-agnostic, and it was a lossy proxy for a fact the graph already holds. Whether the ascent can fire is a property of the persisted CALL_SUMMARY edges. The descent already computes it, so thread the resolved-callee and return-flowing counts out of `interproceduralDescent` and key the note on those instead. Three defects the language proxy carried, all gone: - Wrong for `.mjs`/`.cjs`/`.mts`/`.cts`: the provider registry's extension arrays omit them while the ingestion pipeline parses them as TS/JS, so those files were harvested but the note claimed their ascent was empty. - Silently stale: any language whose harvester started recording formal indices would keep getting the caveat until someone edited the list. - Wrong in reverse: a TS/JS callee with no return-flow got no caveat, so an ascent that found nothing read like one that covered the slice. `pdg-impact.ts` now names no language and imports nothing from the language layer, which also drops the analyze-only provider closure from MCP server startup. Measured on overlayfs against a full build: import mcp/local/local-backend.js before 565-648 ms / 548 modules import mcp/local/local-backend.js after 458-463 ms / 170 modules Tests hold CALL_SUMMARY content fixed while varying the file extension across nine languages and assert the note text is identical, then hold the extension fixed and vary the summary to show the note tracks the data. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(mcp): guard MCP startup against the language-provider closure returning The eager `pdg-impact.ts -> core/ingestion/languages` edge was found and lost once already during #2793 before #2802 re-derived it, so it gets a test rather than a comment. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(lbug): record why csv-generator is not lazy-imported #2802 proposed cutting `csv-generator.js` out of the adapter chain to shorten MCP server startup. Measured on a native filesystem, the marginal cost is small relative to the siblings this module already imports, and `core/search/bm25-index.ts` statically imports `normalizeFtsText` from the same module on a path `local-backend.ts` reaches dynamically for FTS — so deferring would relocate the cost to first query, not remove it. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(pdg): pin chained receiver calls reaching BasicBlock.calleeIds The PDG inter-procedural descent hops through `BasicBlock.calleeIds`, so it can only cross a call boundary the resolver resolved. Chained receiver calls reach `calleeIds` through the receiver-typing pass's own `calleeIdSink` — a separate path from plain calls. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(analyze): drop the stale per-language cross-reference (#2802 review P3-4) `pdgModeMismatch`'s comment told readers to keep "the diagnostic per-language refinement in the impact CONSUMER (see pdg-impact.ts assemblePdgImpactResult)". That refinement is no longer per-language — removing it is the point of #2802, which now keys the empty-ascent note on the persisted CALL_SUMMARY data instead. The comment's real invariant is untouched and still correct: the values in `resolvePdgConfig` must stay scalar, because the comparison below is a shallow `!==` and an object would compare by reference. Only the cross-reference was stale. Comment-only; no executable line changes. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(mcp): probe the real module loader for the startup language closure (#2802 review P1-2) The previous guard hand-rolled a regex walk over TypeScript source to assert `core/ingestion/languages` was not statically reachable from MCP startup. Four bypasses were reproduced against it, any one of which let the exact 226-module regression return while the test stayed green: a. Wrong entry root. It walked from `mcp/local/local-backend.ts`, but the server module is `mcp/server.ts` — which imports LocalBackend as `import type`, so the guard's anchor was not even on server.ts's runtime closure. Ten real startup modules sat outside it. b. A top-level `await import(...)` executes during module evaluation, so it is eager at startup — but the walker skipped every `import(...)` by construction. c. The `import type` strip deleted a 16,445-character window of `pdg-impact.ts`: an `export type X =` matched lazily to the next `from "…"`, which lives inside a string literal. Any import in that window was invisible. d. The comment strip treated a `/*` inside a string literal as a comment opener. Replace the approximation with a real module-load probe: spawn a child node process per entry, import the built `dist/` entry, and report what the loader actually pulled in. Rooted at `dist/mcp/server.js` and `dist/cli/mcp.js` (the real startup entries) plus `dist/mcp/local/local-backend.js`. Syntax cannot fool it. One deviation from the two existing sibling probes is load-bearing: `dist/` is ESM, so a `require.cache` diff alone cannot see the first-party `dist/**` graph — it only catches CJS and native modules, which is why `import-closure.test.ts` gets away with it (it asserts on `@ladybugdb/core`). A pure cache diff here would have reported zero language modules unconditionally, i.e. a new vacuous guard. This probe unions `module.registerHooks({ load })` with the cache diff, and each entry carries a non-vacuity anchor and a module floor so an empty result fails loudly. Verified load-bearing: adding a top-level `await import('../core/ingestion/languages/index.js')` to `src/mcp/resources.ts` and rebuilding turns `dist/mcp/server.js` red with 70+ named offenders, while the `local-backend` and `cli/mcp` cases stay green — which is bypass (a) demonstrated directly. The old guard passed that poisoned tree entirely. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(lbug): drop the unreproducible 9p multiplier from the csv-generator note (#2802 review P3-2) The comment justifying why `csv-generator.js` is NOT lazy-imported carried a hard "~40x" figure for how much a 9p mount inflates per-file ESM resolve. Three independent measurements during review produced ~40x, ~7.3x and ~30x, so the multiplier is not a reproducible quantity and had no business being stated as one in a durable comment. Reworked so the STRUCTURAL argument leads and the numbers only support it. That argument is what actually settles the question and it does not rot: `core/search/bm25-index.ts` statically imports `normalizeFtsText` from `csv-generator.js`, and `local-backend.ts` reaches bm25-index through a dynamic import on the FTS query path — so deferring here relocates the cost to first query rather than removing it. Both verified again at `bm25-index.ts:15` and `local-backend.ts:2756`. Remaining figures are re-measured, attributed to a date and issue, and labelled by filesystem: ~1.6 ms marginal (median of 45 cold imports on local disk) versus ~50 ms for the same import on a network mount, stated as environment-bound rather than as a property of the module. The provider-registry cost is given as "several hundred modules" — the static walk, the runtime hook, and the reviewer's probe each counted it differently (375 / 439 / 407), so no single number was picked to go stale. The old "226 modules" was real but counted only the `languages/` subtree and undercounted the win. Also repoints the trailing reference to the guard's new home at `test/integration/mcp/startup-language-closure.test.ts` (same comment block, inseparable from this rewrite). Comment-only; no executable line changes. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(mcp): stop the empty-ascent note asserting a fact an undecodable summary contradicts (#2802 review P2-2) The note claimed "this is a property of the persisted summaries" whenever the descent resolved callees and none carried a return-flow. But `decodeCallSummary` never throws by design: a version-skewed (`2|r:1`), corrupt (`1|r:zz`), or NULL `reason` yields no entry, which was indistinguishable from a cleanly-decoded empty summary. So the note could assert "no formal parameter is recorded as flowing to its return value" about a callee whose CALL_SUMMARY actually records `p0 -> return`. `meta.pdg.hasCallSummary` is a plain boolean and stores no codec version, so nothing else caught it. `calleesWithReturnFlow` now reports three outcomes instead of two — flowing, decoded-empty, and undecodable — and the undecodable count is threaded through the descent to the note. When it is non-zero the note says so and points at a re-index; when every summary decoded, the persisted-summaries claim is kept and now explicitly conditioned on that. Soundness is unchanged: an undecodable summary still licenses no ascent and never enters the return-flowing set, so the ascent path is byte-identical. Only the note's wording moves. Tests drive all three undecodable forms through the mock and assert the false claim is gone, the remedy is reported, and the ascent is still withheld. A companion assertion pins that the all-decoded case KEEPS the persisted-summaries claim, so the fix cannot degenerate into deleting the sentence. Verified load-bearing: reverting the source alone fails 6 of 34. Impact analysis: `calleesWithReturnFlow` upstream LOW (2 callers, both in this file); `assemblePdgImpactResult` upstream LOW (1 caller). Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(pdg): cover every chained-receiver shape and pin the inference gap (#2802 review P2-1, P3-1) The fixture proved chained receiver calls reach `BasicBlock.calleeIds` using exactly one receiver form — a local `const`. That is the shape that works, so a single-shape fixture implied general support the resolver does not have. This repo has been burned by that before: a drop-count gate blind to fixed shapes. Measuring nine forms against the real pipeline also corrects how the gap was originally characterised. It is NOT local-versus-field. An annotated field resolves fine, including the constructor-assigned variant: private p: Outer = new Outer(); -> both links private p: Outer; this.p = new Outer(); -> both links private p = new Outer(); -> EMPTY CELL private p; this.p = new Outer(); -> EMPTY CELL The discriminator is the type ANNOTATION. When a field's type must be inferred from its initializer the whole `calleeIds` cell empties — so even `Outer.inner`, an ordinary named-receiver call, is lost, and the inter-procedural descent cannot cross the boundary at all. Pre-existing; independent of #2802, which does not touch receiver resolution. The fixture is now table-driven over seven working forms (local const, local in a method, annotated field, ctor-assigned annotated, ctor-param assigned, call-result receiver, three-link chain) plus the two inference-typed forms, each row carrying its expected chain-link ids. Assertions moved from substring to exact id membership, split with the production `splitCalleeIds` reader — so `Inner.compute` can no longer be satisfied by `Inner.computeExtra` or `OtherInner.compute`, which matters because the descent keys on exact ids for span and CALL_SUMMARY lookup. The two known-gap rows are pinned with `it.fails` plus a hard assertion on the exact gap-row set, so a resolver fix turns them red instead of passing silently, and an anti-vacuity guard requires every shape to match exactly one block — without it a drifted fixture matching zero blocks would let `it.fails` pass for the wrong reason. Proven by mutation: relabelling a working row as a known gap fails both pins. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(mcp): qualify the empty-ascent note when the examined callee set is incomplete (#2802 review P2-4) The note asserted "none of the N resolved callees carry a CALL_SUMMARY return-flow", and on the all-decoded path that this is "a property of the persisted summaries". Both are universal claims over the callees the descent actually examined, and two mechanisms can leave that set incomplete without the note saying so: 1. Budget truncation. The descent stops on depth/limit/node-cap, so a callee that DOES carry a return-flow can sit in a hop never reached. A 4-deep chain reported "none of the 3 resolved callees" while link 4 held the only summary. 2. Emit-time capping. When a block's `calleeIds` cell was capped, `splitCalleeIds` strips CALLEES_TRUNCATED_SENTINEL, so the dropped callees are invisible to both the scan and the counters — even though the callgraph bridge in this same file already treats such a block as callee-incomplete. Add `calleeIdsWereTruncated`, the counterpart to the sentinel strip, read from the raw cell before splitting so a block whose entire list was capped away still raises the flag. Thread it through the descent to the note. Case 1 needs no new plumbing — the aggregate `truncated` is already on the input object. Using the aggregate rather than a descent-only flag is deliberate: seed truncation and intra-BFS depth truncation also shrink the initial slice, so their callees are never gathered either. It is a sound superset that never under-hedges. When either mechanism fired, one clause naming the reasons is appended and the whole-slice assertion softens to "every summary examined decoded … a property of those summaries". When the set is complete both branches stay byte-identical to before, so this does not become a blanket hedge. Tests pin truncated, untruncated, emit-capped-alone, both-mechanisms, and undecodable+truncated, asserting the truncation premise rather than assuming it. Verified load-bearing: reverting the source alone fails 6 of 42, and the HEAD note printed in those failures is the bug verbatim. Impact analysis: `assemblePdgImpactResult`, `calleeIdsByBlock`, `interproceduralDescent` all upstream LOW; every caller is in this file and `runImpactPDG`'s exported signature is unchanged. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(mcp): stop the empty-ascent note calling call-site references "resolved callees" (#2802 review P3-7) The note printed "none of the N resolved callees carry a CALL_SUMMARY return-flow (no formal parameter is recorded as flowing to its return value)". N counted the raw `BasicBlock.calleeIds` cell, which carries ids `resolveCalleeSpans` never enters — out-of-repo targets, interface methods, and the `Class:` id a `new X()` emits. On the chained-receiver fixture that inflated N from 1 to 3. Two defects, both in the wording rather than the arithmetic: "resolved" implies a symbol-table lookup that did not happen for those ids, and the parenthetical asserted a FORMALS-level property about symbols never resolved to a body. Reworded rather than re-seeded, deliberately. `calleesWithReturnFlow` scans the RAW id set, so the claim "none of these carries a return-flow" is exactly established for all N — the scan really did check the `Class:` id. Re-seeding N from the resolved spans would make the sentence quantify over a strict SUBSET of what was checked, silently dropping the un-enterable references from a claim that genuinely covers them, and would desync N from `calleesUndecodable`, which is derived from the same scan population. none of the N resolved callees carry ... none of the N call-site callee references carry ... and the formals parenthetical is dropped. The note gets shorter, not longer. `calleesResolved` is renamed `calleeReferences` end-to-end (file-local; nothing outside referenced it), and the descent's return-type doc — which called them "callee symbols the descent resolved" and reinforced the wrong reading — now states that un-enterable ids ride the same cell, are scanned, and are never entered. The `> 0` gate is unchanged, so no slice that previously produced the note stops producing one. A test pins that explicitly: an all-un-enterable cell resolves no span, takes no hop, and emits no ascent sentence despite a non-zero count — so a future re-seeding cannot silently move when the note fires. Tests also pin the quoted number and singular/plural against a mixed cell, with a discriminator asserting `reachableBlocks` is byte-identical while the count moves 1 -> 3. Verified load-bearing: reverting the source alone fails 6 of 7 new tests, printing the finding verbatim. Impact analysis: `assemblePdgImpactResult` and `interproceduralDescent` upstream LOW, sole caller `runImpactPDG` in the same file; exported signature unchanged. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(mcp): pin cross-hop callee accumulation and the mixed return-flow contract (#2802 review P2-5) Every case in this file drove a single hop, so the Set union the descent performs across hops (`calleeReferencesSeen` / `calleesReturnFlowingSeen`) was never proven to accumulate rather than overwrite — a one-hop descent cannot tell the two apart. And although a sibling commit added a three-id cell, none of those ids return-flowed, so the "some callees flow, some do not" boundary was entirely unpinned. Extends the mock with a `secondSummary` knob that drives a genuine second hop: `helper2` is named only in `helper`'s own body block, so the descent must cross a second boundary to reach it. Three mock handlers are made faithful to the parameters they already bind — `calleeIdsByBlock` now routes on the asked `$ids`, and the CALL_SUMMARY scan and span resolve answer per asked id — which is what makes a second callee answerable at all. Existing cases are behavior-identical. Five tests: the union count across two hops; a return-flow on hop 0 surviving a later empty hop; a return-flow found only on hop 1; mixed callees in one examined set going silent rather than partial; and a flowing callee alongside an undecodable sibling staying silent including the decode remedy. The mixed case pins a deliberate contract rather than proposing one. The production condition is `calleesReturnFlowing === 0`, so partial coverage is reported as silence. A reviewer considered and dropped "report partial coverage" as a product change; this makes flipping it a conscious edit instead of an accident. Verified load-bearing against three separate source mutations: accumulating only on hop 0 (2 fail), each hop overwriting instead of unioning (3 fail), and flipping the gate to partial-coverage reporting (4 fail). In all three every PRE-EXISTING test still passed — which is the finding restated as evidence. Test-only; `pdg-impact.ts` is byte-identical to HEAD. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(mcp): consolidate the empty-ascent rationale to one canonical site (#2802 review P3-6) The "keyed on observed CALL_SUMMARY data, never on the criterion's language" rationale was restated in full at four comment sites. It exists because a reviewer asked "why not just look up the language?", so it has to stay findable — but not four times. The canonical explanation now lives in `interproceduralDescent`'s return-type doc, where the counters are actually computed, organised as POPULATION (why the raw `calleeIds` tally is the right set to quantify over) and OBSERVED DATA, NEVER THE CRITERION'S LANGUAGE (the full answer, including the producer-change argument and the no-language-naming rule). The other three sites keep only what is locally load-bearing and point here. Deliberately preserved, because each carries a non-obvious fact: why an undecodable summary licenses no ascent, why the aggregate `truncated` is used rather than a descent-only flag, and the raw-id-tally population argument. Net comment delta -11 lines. The reviewer also flagged the local/field naming asymmetry (`calleeReferencesSeen` vs `calleeReferences`). Keeping the suffix, with a comment recording why so it is not re-raised: the premise that every other local matches its field is true, but those locals are identity-returned, whereas these are `Set<string>` accumulators returned as `.size`. Dropping the suffix would give one identifier two types in one file — a `Set` at the accumulation site and a `number` where the note does arithmetic and pluralisation on it ~900 lines away. The Set-ness is also load-bearing: the dedup is why a callee invoked from two hops is not double-counted, which is what makes the note's count correct. Comment-only. Verified mechanically: every added and removed line in `git diff -U0` matches a comment pattern, so the note's template literals are untouched and its rendered text is byte-identical. 89 tests unchanged. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor(mcp): collapse the ascent plumbing accreted across 13 fix commits Quality cleanup, no behavior change. Four independent review passes converged on the same root cause: thirteen commits each fixed one review finding in isolation, and the ascent facts grew one loose field at a time until 62% of the changed region was comments explaining plumbing. Five changes: - `calleeIdsFromBlocks` deleted. Zero call sites anywhere in src/ or test/ — already dead on main, and this branch had edited it to keep it compiling. Its only reference was a stale `{@link}` in a neighbour's doc, now rewritten to stand alone. - `parseCalleeIdsCell` replaces the two-pass read. `calleeIdsWereTruncated` and `splitCalleeIds` were splitting the same cell on adjacent lines, which measured ~2x the parse cost (0.82 -> 1.59 ms at a realistic hop, 57.7 -> 92.7 ms at the per-statement site cap) and was a second independent encoding of the sentinel format — exactly what `splitCalleeIds` was extracted to prevent. One pass classifies as it walks; `splitCalleeIds` stays as a wrapper so its two external callers are untouched. The single-use `export` is gone. - `AscentCoverage` replaces four fields threaded through three signatures. ~12 declaration sites become 3, and the canonical rationale now lives on the type by construction — which is why the earlier doc-consolidation commit was needed at all. - `calleesReturnFlowing` becomes a boolean. Its only reads were `=== 0`, twice; it cost a Set sized to every callee in the slice plus a per-hop union loop. The flag is set inside the existing `returnFlowing.size > 0` branch — equivalent, since the cross-hop union is non-empty iff some hop's was. - The duplicated empty-ascent note head is collapsed to one gate and one head with per-arm tails. Both arms had been edited in lockstep twice in this branch's own history. The rendered note text is byte-identical. Verified structurally and then empirically: both expressions reconstructed standalone and diffed across the full cross product of references x returnFlowing x undecodable x truncated x listTruncated — 288 combinations, 0 mismatches. Net -53 lines. 102 tests pass unedited; the unused-symbol lint warning is gone. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(mcp): parallelise the startup probes, drop a redundant pin, name the mock knobs Quality cleanup from the same review passes. The set of verified behaviors is unchanged except where noted. **Startup probes run concurrently.** `spawnSync` blocks the event loop and vitest runs a file's tests in order, so the three probes strictly serialised. Launching all three with async `spawn` in `beforeAll` and asserting over the collected outcomes cuts the file from ~12.7 s to ~3.9 s wall (-69%). Every promise is caught before `Promise.all`, so all three children are reaped and failures report per entry rather than surfacing only the first rejection. Preserved and each proven by mutation: the missing-dist error names its entry, a raised module floor fails only its own row, and a bogus anchor still reports the loaded-module count. **The two `it.fails` rows are removed.** They pinned the inference-typed receiver gap that the strict `toEqual` pin beside them already covers — and they were the weaker of the two, because `it.fails` passes when the body throws for ANY reason, including `idsFor`'s own non-vacuity guard. A renamed fixture marker would have kept them green on a rotted premise. The strict pin is self-diffing and was verified load-bearing on its own: pointing a known-gap marker at a resolving shape fails it with the two newly-present ids listed. The file header now carries the gap's durable description. **The ascent-note mock takes options objects.** `descentExec` and `run` had grown to five and seven positional parameters in the order five agents added them, so call sites read `run(FILE, true, null, 3, false, undefined, null)` — several carrying `undefined` purely to reach a later argument. All 34 call sites are converted; nine that used only defaults are now bare `run(file)`. No knob renamed — they are orthogonal and correctly named. Code lines are exactly neutral (353 -> 353); the win is at the call sites. Also refreshes five comments that still described `calleesReturnFlowingSeen` and the two-branch note, both of which the preceding commit replaced. 102 unit and 10 integration tests pass; test count moves 9 -> 7 in the chained-receiver file, exactly the two redundant rows. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(mcp): publish return-value-ascent coverage on the PDG impact result `impact(mode:'pdg')` computed four facts about ascent coverage and used them exactly once — to interpolate an English sentence. They never reached the result object, so an agent consuming this MCP output could only ask "was the ascent complete, and if not why" by regexing prose. The cost was already demonstrated: a pure rewording commit earlier in this branch broke ~30 assertions and would have silently broken any consumer keying on the old phrase. Adds `pdgEvidence.ascent`: referencesScanned how many call-site callee references were scanned returnFlowFound did the ascent fire anywhere in this slice undecodableSummaryCount summaries the codec could not decode examinedComplete was the examined set the whole callee list incompleteReasons 'traversal-truncated' | 'callee-list-capped' callSummaryLayerPresent false => pre-FU-C (v3) index Nested under `pdgEvidence` because that is the established counts-and- classification namespace, and `composeUnifiedPdgImpactResult` already spreads it, so the member survives the unified compose untouched. `incompleteReasons` carries CODES, following the existing `truncatedByReasons: ('depth'|'limit')[]` precedent. The prose clause and the structured field now render from one array computed once, so an agent branching on codes and a human reading the note cannot disagree, and a third reason becomes a rendering decision rather than a contract change. Two shape decisions worth recording. `callSummaryLayerPresent` exists because without it a v3 index publishes `referencesScanned: N, returnFlowFound: false`, which reads as "these callees record no return-flow" when the truth is "the layer that records it is absent" — the note already distinguishes those, and the structured surface must not be less honest than the prose. And the field is ABSENT rather than zeroed when the descent never ran (upstream slices): "nothing was scanned" is a different fact from "we scanned and found nothing". `pdgResultVersion` stays 2. The documented trigger is a BREAKING change to the result shape; this removes nothing, renames nothing, and changes no existing field's meaning. Confirmed mechanically: zero top-level key drift across 2304 cases. The historical v2 bump was for changing an existing field's semantics (startLine 0- to 1-based). The note prose is byte-identical, proven across the same 2304 cases with a negative control — perturbing one character of the phrase table produces 60 drifts, so the harness demonstrably detects what it asserts. 14 new tests cover the structured surface and all 14 fail when the source is reverted, while the 54 prose tests pass unchanged. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(helpers): share one module-load probe, and fix two guards that passed on broken builds Three tests independently spawned a child node process to inspect what a built `dist/` entry loads, duplicating the REPO_ROOT derivation, the probe source, the missing-dist guard, the spawn with NODE_OPTIONS cleared, the status-vs-signal rendering, and the payload parse. The newest copy was also the only correct one, so the next author had 2-in-3 odds of copying a weaker probe. The two older probes diff `require.cache` only, which is structurally blind to the first-party ESM `dist/**` graph. That is not theoretical — both were demonstrated passing on genuinely broken builds: - Severing `dist/cli/mcp.js -> stdio-context.js` (a pure ESM change) leaves the require.cache diff EMPTY, so `import-closure.test.ts`'s two assertions reduce to `[].filter(...) === []`. It reported 2 passed on a severed graph. - Severing `registry -> swift/query.js` leaves 76 unrelated CJS entries, which satisfied `registry-import-closure.test.ts`'s indirect guard. The Swift half of its headline had gone vacuous and it reported 1 passed. Both now fail on those same builds, naming the missing anchor. `test/helpers/module-load-probe.ts` unions the ESM `registerHooks({ load })` channel with the cache diff, probes entries concurrently, and makes non-vacuity STRUCTURAL: `anchor` and `minModules` are required fields and the helper throws when either fails. A vacuous probe is a harness failure, not a silently green test, so it cannot be forgotten. Forbidden patterns and remedy text stay per-test — the harness is the shared part, the policy is not. Also fixes `toRepoRelativePosix` resolving non-absolute specifiers against `process.cwd()`, and dedupes modules a CJS-from-ESM import reported once per channel. Faster despite doing more: the registry file goes 12.4s -> 6.75s, because `spawnSync` burned the parent thread polling while the child loaded native grammars. `import-closure` drops to one spawn from two. The `local-backend.js` entry is kept although its closure is currently a strict subset of `server.js`'s: that is an observation, not an invariant. If `server.js` ever stops eagerly reaching the local backend, the server probe stays green while the module #2802 actually changed goes unobserved — and now that anchors are mandatory, that entry is what pins `pdg-impact.js`. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(lbug): trim the csv-generator note and fix the claim it got wrong Two reviewers split on this comment: one wanted it cut to the structural argument, the other said a comment is the right depth for documenting a rejected change since there is no invariant to guard. Both are right, so it stays a comment and gets shorter — 13 lines to 6. Trimmed because it had already taken two corrections (an unreproducible "~40x" figure, and a pointer to a test file that no longer exists), and its tail had drifted from its own guard: the comment said "several hundred modules, ~150 ms" where `startup-language-closure.test.ts` says "~226 extra modules and ~130 ms". Two numbers for one fact. That tail is documented better in the guard's own header, so deleting it loses nothing. It also stated the load-bearing claim inaccurately. The old text said bm25-index imports `normalizeFtsText` "from here" — but `lbug-adapter.ts` neither exports nor re-exports it; the only occurrence of the identifier in this file WAS the comment. Anyone verifying would have grepped, found nothing, and concluded the note was stale. Now names `csv-generator.js` explicitly, re-verified at `bm25-index.ts:15` (static) and `local-backend.ts:2756` (dynamic, on the FTS query path). Comment-only, proven two ways: every changed line matches a comment pattern, and stripping all `//` lines from HEAD and from the working tree yields byte-identical text. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(helpers): extract the temp-repo lifecycle, collapsing five hand-rolled cleanups into one Four cfg integration tests each hand-rolled a `tmpDirs` array, a mkdtemp-and-register step, and an `afterAll` rmSync. It is actually five registrations across six creation sites — `pipeline-pdg.test.ts` keeps a second pool for its C-family fixtures. Seeding genuinely varies four ways (recursive cpSync, single copyFileSync, inline mkdir+writeFile, and nothing at all), so a fixture-copier helper would have fitted about half the sites and made things worse. Extracted the LIFECYCLE instead — mkdtemp, register, afterAll cleanup — which is byte-identical at all five registrations and is the correctness-critical part. `dir()` returns an empty registered directory for callers that seed themselves; `fromFixture()` covers the common case. That fits 6/6. The duplication had already produced a latent defect: `cFamilyTmpDirs` was cleaned by TWO `afterAll` blocks, harmless only because `rmSync` was called with `force: true`. Now one hook. `createTempDirPool` is a function called from each test file's module scope rather than a top-level hook in the helper, because under ESM caching a module-level `afterAll` would register once, against whichever file imported it first. That hazard is documented in the helper. Raw line count is roughly neutral (-44 across the tests, +62 for the helper, 29 of which are the rationale). The win is that a cleanup invariant went from five copies to one. Cleanup verified empirically, including the failure path: a throwaway suite whose `beforeAll` throws still has its directory removed, and every temp directory created by the four migrated files is gone after a run. 46 tests pass across the four files. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(resolvers): pin the inference-typed field receiver gap at the resolver level The gap was pinned only in a PDG test, asserting on `BasicBlock.calleeIds` behind the full `--pdg` pipeline. But it is a resolver fact: when a class field's type must be inferred from its initializer, chained receiver calls resolve to nothing. Whoever closes it will be working in the resolver suite and would have got a red CFG/PDG test with no resolver-side signal. Asserts CALLS edges directly, alongside `python-constructor-field-receiver.test.ts`. Nine receiver shapes run the identical statement; seven resolve, two do not: const o = new Outer() resolves private p: Outer = new Outer() resolves private p: Outer; this.p = new Outer() resolves private p: Outer; this.p = p (ctor arg) resolves constructor(private p: Outer) {} resolves makeOuter().inner().compute() resolves o.inner().mid().compute() (three links) resolves private p = new Outer() NO EDGES private p; this.p = new Outer() NO EDGES Two things the fixture establishes that the PDG-side pin could not. The discriminator is the type ANNOTATION, not local-versus-field — the parameter-property form resolves fine. And the initializer is NOT invisible to the resolver: `new Outer()` still emits its own constructor CALLS edge, byte-identical to the annotated twin. Only the initializer-to-field-type binding is missing, which narrows where a fix belongs. Assertions key on exact node ids rather than names, because `compute` is ambiguous across two classes and keying on the source name collides with `Object.prototype.constructor`. No `describe.skip` and no `it.fails` — the latter passes when the body throws for ANY reason, so it can go green on a rotted premise. The gap is pinned as its explicit current value, which self-diffs: simulating the fix fails one test showing the two newly-resolved ids, and renaming a fixture symbol fails the non-vacuity guard. Runtime is comparable to the PDG-side pin (~9-11s, both dominated by worker startup), so this is an altitude and scope win, not a speed one. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(mcp): replace the extension sweeps with a stronger language-agnosticism pin Two `it.each` sweeps over nine file extensions asserted that the empty-ascent caveat was present (or absent) for each. They looked like the pin for the property the whole change exists for — `pdg-impact.ts` must name no language and its output must not vary by extension — but they were the weakest available form of it. They asserted substring presence/absence, so a language dependence that ADDS text while leaving the caveat intact passes them. Demonstrated, not assumed: injecting a `.py`-only hedge inside the caveat sentence and replaying the two sweeps verbatim against that source gives 18 passed. The byte-identity test beside them caught it. So the sweeps are deleted and the identity test carries the property alone, hardened in two ways: - Two rows instead of one, covering BOTH sides of the caveat gate. The silent (return-flow present) branch previously had no identity counterpart at all — nine runs proving one fact, with nothing checking that its rendering was extension-invariant. - The fingerprint spans the note AND the reachable blocks, not just the note. Strictly more than the sweeps verified. Entailment is exact: identity across the extension set, plus the two existing single-extension content assertions, gives "every extension gets the caveat" and "no extension gets it". Reducing a sweep to one extension was rejected because it reproduces an assertion already present verbatim. Also converts the incompleteness block from six near-identical bodies to a 3-row premise table crossed with two assertions. Each row now names the exact phrase set its clause must contain, so presence and absence are asserted together — which adds three checks the longhand version lacked (the budget row now also proves the emit-cap phrase is absent). And three tests that re-rendered one fixture to make one assertion each are hoisted to a single render. 97 tests, down from 116: -18 sweep cases, -2 from the hoist, +1 identity row. No assertion was lost; several were added. Verified by injection: a `.py`-only note change fails the identity pin, and a dependence in the shared hop sentence fails BOTH rows, confirming the second row is load-bearing rather than decorative. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * perf(mcp): lazy-import syncGroup so MCP startup skips the group extractor closure `core/group/service.ts` statically imported `./sync.js`, which pulls all six contract extractors, five of which statically import the native `tree-sitter` binding. That put the whole parser stack on every MCP server start, for a server that never syncs. Only `groupSync` needs it. The other seven group tools — `group_list`, `group_impact`, `group_query`, `group_contracts`, `group_status`, `group_trace`, `group_context` — do not, and now never load it. `syncGroup` has a single call site, already inside an `async` method, so this is a lazy `await import(...)` at that call site and nothing else: no signature change, no async ripple, no change to `local-backend.ts`. The pattern is already established on this exact module — `cli/group.ts`'s sync command lazy-imports `sync.js` the same way. `service.ts` was the outlier. Measured on a native filesystem (overlayfs; /workspace is a 9p mount that inflates ESM resolve, so it is not a valid measurement surface), 5 cold runs, medians: dist/mcp/server.js 521 ms -> 133 ms (-75%) dist/mcp/local/local-backend.js 453 ms -> 66 ms (-85%) tree-sitter modules at both entries: 11 -> 0 Same defect class as #2802, which cut the language-provider registry from the same startup path; this is what remained. The cost is moved rather than deleted: the first `group_sync` call now pays the module load. That is the right trade — `group_sync` is already a long-running operation, and sessions that never sync pay nothing. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(mcp): guard MCP startup against the group extractor closure returning Sibling forbidden-pattern case in the #2802 startup guard, reusing the concurrent probes it already collects — no new spawn, no new harness. Asserts that none of `dist/mcp/server.js`, `dist/cli/mcp.js`, or `dist/mcp/local/local-backend.js` loads a `core/group/extractors/` module or the native `tree-sitter` package. The parser is matched by package prefix rather than a bare substring, so a source file that merely mentions the word can neither satisfy nor trip it. Verified load-bearing rather than assumed: restoring the static `import { syncGroup }` in `core/group/service.ts` and rebuilding turns `dist/mcp/server.js` red and names all seven offenders — http-route, grpc, thrift, topic, include, manifest and workspace extractors. Reverted and re-confirmed green. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * perf(mcp): keep the analyze-only CFG closure off MCP server startup (#2802 review) `mcp/local/pdg-impact.ts` imported `CALLEES_TRUNCATED_SENTINEL` and `CALLEE_ID_SEP` from `core/ingestion/cfg/emit.ts`. ESM evaluates a module to import any binding from it, so those two strings dragged the whole analyze-only CFG closure into every MCP server start. Measured against a clean build, per entry point: 8 modules — `emit`, `reaching-defs`, `reaching-defs-graph`, `control-dependence`, `post-dominators`, `synthetic-escape`, `call-site-harvest`, `reaching-def-reason-codec` — present at `dist/mcp/server.js`, `dist/mcp/local/local-backend.js` and `dist/mcp/http-transport.js`. Same defect class as the language-provider closure this branch already removed, and the guard could not see it: `FORBIDDEN_RE` covers `core/ingestion/languages/` and `FORBIDDEN_GROUP_RE` covers `core/group/extractors/|node_modules/tree-sitter`, neither of which matches `core/ingestion/cfg/`. The format constants move to a new LEAF module `cfg/callee-cell-format.ts` that imports nothing; `emit.ts` re-exports both names so every existing importer is untouched, and producer and consumer still resolve to one definition — the drift the shared constant exists to prevent stays impossible. Deleted, not deferred — the same bar #2802 held its own csv-generator proposal to. After: cfg modules at startup 8 -> 2, and both survivors (`callee-cell-format`, `reaching-def-reason-codec`) are leaves that import nothing. Totals: `server.js` 387 -> 380, `local-backend.js` 163 -> 156, `http-transport.js` 523 -> 516. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(mcp): stop pdgEvidence.ascent claiming a completeness it cannot have (#2802 review) `examinedComplete` is the field a consumer reads to decide whether `returnFlowFound: false` is a whole-slice claim. It could be published `true` over a callee set the descent never finished examining — the exact false all-clear the field was added to prevent. Root cause: `bfsReachableBlocks` sets `truncatedByDepth` when its frontier is still non-empty at the budget, but both call sites inside `interproceduralDescent` folded only the row-limit flag and dropped the depth flag. The top-level intra BFS's copy of that same flag was already propagated, so the asymmetry was unintended — one `if`-pair folding limit-but-not-depth, within a merge that already folds the node cap too. Reproduced at `maxDepth: 3`, the shipped default: a criterion calling a helper whose body is a 5-block dependence chain, with the return-flowing callee on the block past the clamp. Result reported `truncated: undefined`, `examinedComplete: true`, `incompleteReasons: []` and an unqualified universal note sentence. Fixed by propagating the dropped flags rather than inventing a parallel channel: `intraDepthBudget` is documented in-file as the SAME clamp the top-level intra BFS applies, and that one's depth truncation is already result-level. So the result's own `truncated`/`truncatedBy` were under-reporting for the same reason, and both surfaces are corrected together. Four further honesty fixes to the same published record: - Blocks reached only by the U-C4 ascent went into `reachable` but never `hopReached`, so their `calleeIds` cells were never scanned, never counted, and could not raise `callee-list-capped`. They are slice blocks; they now enter the hop set and get the same treatment as every other one. - `pdgEvidence.ascent` was absent on the empty-slice early return even though the descent had already run and scanned, contradicting the "present iff the descent ran" contract this branch itself added to `tools.ts`. Both exits now classify through one shared helper so they cannot disagree. - A block carrying call sites in `callees` but no resolved ids in `calleeIds` (the whole-file case where `emit.ts` has no fileMap) silently shrank the population while `examinedComplete` still reported `true`. That now raises a third reason, `callee-ids-unrecorded`. - `referencesScanned` is a distinct-callee tally and both surfaces described it as a call-site count. Field name kept — a rename is breaking at `pdgResultVersion: 2` — and the prose corrected instead. `PdgAscentIncompleteReason` gains a member, which is additive, so `pdgResultVersion` stays 2. Visible output change worth knowing: slices whose callee chain outruns `maxDepth` now report `truncatedBy: 'depth'` where they previously reported none, and a repo with id-less call sites now reports `examinedComplete: false`. Both are strictly more honest. Every behavioural change carries a mutation proof — revert the source, watch the new test go red, restore. One exception is documented inline rather than faked: the ascent-side fold cannot be observed independently, because the re-seed shares the caller's `visited` set and so can only reach past the budget when the traversal that covered that closure was already cut and had already raised a flag. Suite: 49 -> 59 tests. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(mcp): anchor each import-closure policy on the edge it polices (#2802 review) `module-load-probe.ts` makes non-vacuity structural via a required `anchor` — but the anchor was one per ENTRY while `startup-language-closure.test.ts` now runs TWO independent policies. The group-extractor policy added in83e8cf7c5therefore had no anchor of its own, and one of its three rows was already vacuous: `cli/mcp.js` loads four leaf modules and reaches no `core/group/` module at all, so its group assertion could not fail for any policy-related reason while its `dist/mcp/stdio-context.js` anchor stayed green. Proven, not argued. `dist/mcp/local/local-backend.js` is the only static importer of `core/group/service.js` in the whole build; severing that one edge — the exact next lazy-load step — and re-probing: OLD shape (anchor per entry): server 385, http-transport 521, local-backend 161 reaches group/service = false, group offenders 0 -> GREEN on all three NEW shape (anchor per policy): -> RED on all three, each naming the missing dist/core/group/service.js Counts fell only 387->385 and 163->161, so `minModules` was structurally blind to the severance; the anchor is the only thing that catches it. `anchor` accepts `string | readonly string[]` and every listed anchor must load. Existing single-anchor call sites are unchanged. `anchorsOf()` lets the group `it.each` DERIVE its entries by filtering on the group anchor, with a test pinning that derivation, so the policy cannot silently register zero cases. `cli/mcp.js` is dropped from the group policy — it cannot honestly carry that anchor — and the doc-comment now states the invariant: an anchor is per-POLICY, not per-entry. Also: - `mcp/http-transport.js` gets a row. It is the largest startup entry (516 modules) and `src/cli/mcp.ts` imports it directly rather than through `server.js`, so nothing about the server row constrained it. Measured clean today; the gap was coverage, not a broken claim. - The three spawn-based closure tests are registered in `SPAWN_CLI`, so the Windows-safety plumbing this branch wrote for them (POSIX normalisation, `pathToFileURL`, `NODE_OPTIONS` clearing, array-form `spawn`) is finally exercised on the Windows/macOS matrix. Measured cost ~11.7s on Linux; budget ~60s on Windows against a 25-minute job. - `PROBE_TARGET` now wins over `extraEnv`, which was spread last and could have silently redirected a probe while `anchor`/`minModules` stayed keyed on `entry`. - The child's JSON payload is validated through a type predicate instead of a bare `as string[]`, and the spawn timeout escalates SIGTERM to SIGKILL so a child stalled in native code is reaped rather than orphaned. - Recorded baselines re-measured (server 380, local-backend 156, cli/mcp 4) and relabelled a snapshot rather than a contract — they moved twice inside this branch alone. The subset claim was re-verified exactly: 0 of local-backend's 156 modules are absent from server's 380. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(helpers): survive a failing temp-dir removal instead of leaking the rest (#2802 review) `createTempDirPool`'s `afterAll` ran a bare `for (const d of created) fs.rmSync(d, {recursive, force})`. `force` suppresses only `ENOENT` — not the `EBUSY`/`EPERM`/`ENOTEMPTY` class a Windows runner produces when a pipeline test still holds a handle — so the FIRST failure threw out of the loop and leaked every directory registered after it. Pre-existing: all four hand-rolled cleanups this helper consolidated had the same shape. But the blast radius is now shared across four consumers, which is exactly why it is worth fixing at the point of consolidation. Cleanup is now per-directory best-effort via `removeTempDirs`, plus Node's own documented mitigation for that error class (`maxRetries: 3, retryDelay: 50`), which costs nothing on the happy path. Warn rather than swallow or rethrow, and the reasoning is in the doc comment, not just here: rethrowing would fail an otherwise green suite from `afterAll` over housekeeping the OS reclaims anyway, where it reads as a test failure and buries the real result — a Windows EBUSY on a temp dir is not a defect in the code under test. Silence is the opposite hazard: a systematic leak would be invisible with nothing naming the responsible suite. The warning carries the path, and the `mkdtemp` prefix is per-pool, so it names the suite that made it. Failure is injected through a scripted remover keyed by path (a Map lookup, so no `if` in a test body and no dependence on producing a real locked handle). Beyond the three behavioural pins there is a wiring pin — a nested `describe` creates a real pool and a sibling `it` declared after it asserts the dirs are gone — so the tested function cannot drift into "tested helper plus an untested copy of the loop". Mutation proof: restoring the abort-on-first-failure loop turns 3 of the 5 tests red, the throw escaping `removeTempDirs` outright so the third real directory is never attempted. With the fix, `[first, blocked, last].map(existsSync)` is `[false, true, false]` — the injected failure survives and the directory after it is really gone, through the remover that actually ships. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(pdg): point the self-diffing receiver pins at #2807, not at this PR (#2802 review) Both pins named the gap "(#2802 follow-up)". The gap has its own tracking issue — #2807, "Inference-typed field receivers resolve to no CALLS edges at all" (open, labeled bug) — and PR #2810 is already open against it. As written, after merge the gap was discoverable only by reading a KNOWN GAP marker inside a test file, not from the issue tracker. Both describe names now read "(known gap: #2807)" and both KNOWN GAP test names carry the number. #2802 is kept only as provenance: the gap was FOUND during #2802 work but is pre-existing and independent of it. Each header gains an explicit "this pin is self-diffing: it will go red on purpose" section naming #2807 with its exact title, noting #2810 is open against it at the time of writing, and stating that the pin asserts the gap EXISTS — so closing #2807 fails it by design, and the correct response is to update the expected value, not to relax the assertion. The same note is repeated inline above each KNOWN GAP test, where a maintainer editing it will actually see it. No pin is weakened. Both deliberately reject `it.fails` in favour of exact `toEqual` assertions with a non-vacuity probe, and that design is left untouched. Refs #2802, #2807 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(group): cover the lazy syncGroup import that no test reached (#2802 review)9ea9676dcturned `GroupService.groupSync`'s `syncGroup` into `await import('./sync.js')` — this branch's one changed control-flow line in production code — and nothing exercised it. Every existing test stopped short: `service.test.ts` returns at the empty-name guard; `group-service-not-found.test.ts` mocks `loadGroupConfig` to reject and never invokes its `syncGroupMock`; `group-sync.test.ts` imports `syncGroup` directly, bypassing `GroupService`; and the startup guard asserts only the negative, that `sync.js` is absent at startup. `tsc` catches a path typo, but nothing verified the import resolves and hands off correctly — while every production `group_sync` call goes through that line. No production change was needed; the reviewed design was sound. This is the missing coverage. The happy-path test mocks nothing: it points `GITNEXUS_HOME` at a pool temp dir, seeds a real `group.yaml`, and calls `groupSync`, so `loadGroupConfig` resolves, `groupDir` is found, and execution falls through into the REAL `syncGroup`. What makes a real sync reachable with no indexed repo: an empty registry puts both members in `missingRepos`, but one declared manifest link still yields synthetic-UID contracts. It asserts the returned counts AND reads back the `contracts.json` that real `syncGroup` wrote into `groupDir` via the production `readContractRegistry`, which pins the option handoff too. Two further tests use `vi.doMock` to re-evaluate the service against a `sync.js` whose load throws: one asserts the call rejects with the load failure in its `cause` chain — so the caller gets a catchable rejection, not a floating unhandled one — and one asserts both pre-import guards still answer with `sync.js` unloadable, which is also a structural pin that the module has no STATIC import of it (a static one would throw at re-import, before any call). Mutation proofs: pointing the specifier at `./sync-nope.js` turns 2 of 3 red ("Cannot find module .../sync-nope.js ... at GroupService.groupSync service.ts:349"); aliasing a real-but-wrong export turns 1 red. Restored, all 3 green, and `service.ts` verified byte-identical to HEAD. Out of scope, stated rather than glossed: the final `isError: true` MCP envelope is produced above `GroupService` and needs a full `LocalBackend`; the rejection test is the in-scope half of that claim. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor(mcp): close the gaps a cleanup pass found in the #2802 review fixes Quality pass over the review-response series (reuse / simplification / efficiency / altitude). No behaviour change except where noted. The two that mattered: - **The cfg/emit fix had no guard.** `FORBIDDEN_RE` covers `core/ingestion/languages/` and `FORBIDDEN_GROUP_RE` covers `core/group/extractors/|tree-sitter`; neither matches `core/ingestion/cfg/`. Because `emit.ts` re-exports the constants, pointing `pdg-impact.ts` back at `cfg/emit.js` typechecks identically and silently restores all 7 modules. Verified: with the import reverted, `tsc --noEmit` still exits 0 and every test stayed green before this commit; after it, 3 rows go red naming the offenders. Written as an ALLOWLIST of genuine leaves rather than a denylist of the 7 already-suffered modules, because the next regression is a module nobody has thought of yet. - **`FORBIDDEN_GROUP_RE`'s parser matcher was forward-slash only** while both sibling probe regexes spell the separator `[\\/]`. Native bindings arrive via the `require.cache` channel as absolute paths and `toRepoRelativePosix` only normalises paths inside the repo root, so a hoisted `node_modules` renders as `…\node_modules\tree-sitter\…` on Windows and matched nothing. The same series put this file on the Windows matrix, where that half of the assertion would have been vacuous. Reuse — three re-implementations of existing helpers: - `removeTempDirRecursive` re-rolled `fs.rmSync` retries; it now delegates to `cleanupTempDirSync` (`test-db.ts`), the repo's Windows-lock-aware remover. The copy had already drifted on both knobs that matter — 3 retries at 50 ms vs 5 at 100–400 ms, and warn-on-everything vs swallow-lock-codes-rethrow-rest — which is how one half of a suite goes green-with-a-warning on the same `EBUSY` the other half fails on. The per-directory try/warn loop, which is the actual fix, is unchanged. - `errorChainText` re-rolled the cause-chain walk that `causeChain` (`src/lib/utils.ts`) exists to be the single copy of — its own doc asks callers not to. - The SIGKILL escalation (a timer, an `unref`, and two `clearTimeout`s) is `spawn`'s own `killSignal` option, which Node's `timeout` already delivers. Simplification and altitude: - `'callee-ids-unrecorded'` documented ONE of its three producer paths. The unnamed common one is a call site that did not RESOLVE — exactly the receiver gaps this repo pins (#2807) — so on a real index the reason fires broadly, driven by resolution quality rather than a missing `--pdg` layer, and "re-run analyze --pdg" is the wrong remedy for it. Doc now names all three and states the consequence: `examinedComplete: true` is the strong, rare signal. - The derived policy-entry list was re-pinned against a hand-written 3-element literal, reinstating one layer down the list the derivation removes. Now asserts the properties that are actually at risk — non-emptiness (a policy going silent) and `cli/mcp.js` staying excluded (a row that cannot fail). - A test fixture spread `ascentBlockCell: 'idless'` and then overrode it to `'capped'` in both runs, so the id-less shape never reached the mock while reading as though it did. - `idlessCallSites` is sticky, so its per-row string allocation now short-circuits once set. - Dropped an unused `export` on `CleanupWarner`. Refs #2802 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * revert(ci): unregister the module-load closure guards from the Windows matrix Registering the three `dist/` closure guards in `SPAWN_CLI` turned the Windows `platform-sensitive 1/3` shard red at the 20-minute watchdog. Baseline83e8cf7c5was green on all three shards;a4245119c(which added them) failed 1/3;d0b201442failed the same way. It is not the files themselves. On the Windows runner they are among the cheapest in the suite — `registry-import-closure` 448 ms, `import-closure` 53 ms — and both passed. vitest shards this list by file COUNT, not runtime, so adding three files RESHUFFLED the split: shard 1 went to 32 files against 26 and 29, concentrating the heavy CLI e2e suites. It timed out with `cli-e2e`, `group/cross-trace-e2e`, `lbug-orphan-sidecar-recovery` and `server-http-startup` still queued — `cli-e2e` being the ~50-spawn suite whose setup flakiness already needed fixing once (PR #2000). That clustering fragility is pre-existing and this file's own header documents it (#2449: "the heaviest spawn suites can cluster on one shard", busiest Windows shard already at 14m57s against the old watchdog). These three files only tipped it over, and unblocking the PR beats holding it for a CI-infra fix that belongs in its own change. Reverted rather than worked around: raising the shard count would keep the coverage but is a repo-wide CI change made on a 25-minute feedback loop with no guarantee the reshuffle balances, and this PR is about MCP startup. The removed entries are replaced by a comment recording WHY they are absent, what they were measured to cost, and the precondition for re-landing them — so the gap is documented at the point someone would otherwise re-add them blind. Verified: the emitted file list is byte-identical to 83e8cf7c5's, so the shard split returns to the configuration that was green. The Windows-specific bug this series found is unaffected — `FORBIDDEN_GROUP_RE` now spells its separator `[\\/]` like its siblings, which was a real forward-slash-only vacuity, and that fix stays. Refs #2802, #2449 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ci): shard the cross-platform matrix by measured weight, not file count Restores the three `dist/` module-load closure guards to the Windows/macOS matrix, and fixes the reason they could not stay there. They must run on every OS — the shared probe in `test/helpers/module-load-probe.ts` IS the platform-varying code (array-form `process.execPath` spawn, cleared NODE_OPTIONS, `pathToFileURL` because Windows rejects a bare absolute path as an ESM specifier, and a `path.sep`→POSIX normalisation the anchors and offender regexes depend on). Ubuntu-only coverage of a platform guard is no coverage. The earlier attempt turned Windows `platform-sensitive 1/3` red at the 20-minute watchdog, and the reflex fix — unregistering them — treated the symptom. The files are among the cheapest in the suite (measured 448 ms, 53 ms, sub-second, and both that completed passed). The defect is that `run-cross-platform.ts` handed vitest all 84 files plus `--shard=i/n`, and vitest partitions by file COUNT. Runtimes here span three orders of magnitude, so a count-split is blind to the thing that decides the budget, AND re-partitions on every insertion: adding three free files reshuffled the list and happened to co-locate `cli-e2e` (361 s) with `cli-limit-e2e` (75 s) and `analyze-heap-oom-e2e` (23 s) — 32 files against 26 and 29 — which timed out with four still queued. The split now happens in `scripts/cross-platform-shard.ts`, longest-processing- time first over measured Windows runtimes, and only the chosen shard's files are passed to vitest (`--shard` is consumed, never forwarded — forwarding would re-partition the slice a second time and silently drop most of it). Weights are measured, from the last green matrix run plus the timed files of the failing one, and every file also carries an 8 s per-file floor. That floor is calibrated, not guessed: the last green busiest shard ran 736 s of wall clock over ~511 s of attributed file time. Without it the balancer isolates the two monsters and then piles every light file onto the remaining shards — trading a runtime imbalance for a count imbalance that costs the same. Result at TOTAL=3, with the three guards back in: 521 s / 527 s / 519 s across 20 / 33 / 34 files. The previous green configuration's busiest shard was 736 s, so this is better balanced than the state before any of this, and the busiest shard is now bounded by construction rather than by sort-order luck. `test/unit/cross-platform-shard.test.ts` pins the properties, and the load-bearing one is not "the split is even" — it is "adding a cheap file cannot move a heavy one", the property whose absence caused the outage. Two details in that test are themselves load-bearing, and earlier drafts got both wrong and were vacuous: the inserted names must sort EARLY (names sorting last disturb nothing under any scheme) and the count must not be a multiple of the shard total (adding exactly `total` files leaves an equal-weight round-robin in the same rotation). Mutation-proved: replacing `weightOf` with a constant — i.e. count-based sharding — turns that test and the per-file-floor test red; restored, all 8 pass. Refs #2802, #2449 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
GitNexus
Graph-powered code intelligence for AI agents. Index any codebase into a knowledge graph, then query it via MCP or CLI.
Works with Cursor, Claude Code, Antigravity (Google), Codex, Windsurf, Cline, OpenCode, CodeBuddy (Tencent), Qoder (Alibaba), and any MCP-compatible tool.
Why?
AI coding tools don't understand your codebase structure. They edit a function without knowing 47 other functions depend on it. GitNexus fixes this by precomputing every dependency, call chain, and relationship into a queryable graph.
Three commands to give your AI agent full codebase awareness.
Quick Start
# Index your repo (run from repo root)
npx gitnexus analyze
That's it. This indexes the codebase, installs agent skills, registers Claude Code hooks, and creates AGENTS.md / CLAUDE.md context files — all in one command.
On npm 11.x?
npxcan crash during install (Cannot destructure property 'package' of 'node.target'). Use the pnpm form instead:pnpm --allow-build=@ladybugdb/core --allow-build=gitnexus --allow-build=tree-sitter dlx gitnexus@latest analyzeSee Troubleshooting →
npx gitnexuscrashes withnode.target is null(npm 11) for the full matrix (global install, npm downgrade).
To configure MCP for your editor, run npx gitnexus setup once — or set it up manually below.
gitnexus setup auto-detects your editors and writes the correct global MCP config. You only need to run it once. To configure only selected integrations, pass --coding-agent/-c with a comma-separated list or repeat the option, for example gitnexus setup -c cursor,codex.
Editor Support
| Editor | MCP | Skills | Hooks (auto-augment) | Support |
|---|---|---|---|---|
| Claude Code | Yes | Yes | Yes (PreToolUse + PostToolUse) | Full |
| Cursor | Yes | Yes | Yes (postToolUse, manual install) | Full |
| Antigravity (Google) | Yes | Yes | Yes (AfterTool, Gemini CLI hooks schema) | Full |
| Codex | Yes | Yes | Yes (PreToolUse + PostToolUse, Codex hooks) | Full |
| OpenCode | Yes | Yes | — | MCP + Skills |
| CodeBuddy (Tencent) | Yes | Yes | — | MCP + Skills |
| Qoder (Alibaba) | Yes | Yes | — | MCP + Skills |
| Windsurf | Yes | — | — | MCP |
Claude Code and Codex get the deepest integration: MCP tools + agent skills + PreToolUse hooks that automatically enrich grep/glob/bash calls with knowledge graph context + PostToolUse hooks that detect a stale index after commits and prompt the agent to reindex.
Community Integrations
| Agent | Install | Source |
|---|---|---|
| pi | pi install npm:pi-gitnexus |
pi-gitnexus |
MCP Setup (manual)
If you prefer to configure manually instead of using gitnexus setup:
Claude Code (full support — MCP + skills + hooks)
# macOS / Linux
claude mcp add gitnexus -- npx -y gitnexus@latest mcp
# Windows
claude mcp add gitnexus -- cmd /c npx -y gitnexus@latest mcp
Codex (full support — MCP + skills + hooks)
codex mcp add gitnexus -- npx -y gitnexus@latest mcp
Codex hooks (PreToolUse graph enrichment + PostToolUse stale-index detection in ~/.codex/hooks.json, same schema as Claude Code) need the bundled adapter script, so they are installed by gitnexus setup -c codex rather than manually.
Alternatively, install everything as a Codex plugin (MCP + skills + hooks in one step):
codex plugin marketplace add abhigyanpatwari/GitNexus
# then inside Codex: /plugins → install "GitNexus"
Codex notes: SessionStart is intentionally not registered — Codex reads AGENTS.md natively, which already carries the GitNexus context block. Newly installed hooks need a one-time approval in Codex via
/hooksbefore they run. Pick one install route (gitnexus setup -c codexor the plugin): plugin hooks load alongside~/.codex/hooks.json, so installing both can fire duplicate hooks per tool call.
Cursor / Windsurf
Add to ~/.cursor/mcp.json (global — works for all projects):
{
"mcpServers": {
"gitnexus": {
"command": "npx",
"args": ["-y", "gitnexus@latest", "mcp"]
}
}
}
OpenCode
Add to ~/.config/opencode/config.json:
{
"mcp": {
"gitnexus": {
"command": "npx",
"args": ["-y", "gitnexus@latest", "mcp"]
}
}
}
CodeBuddy
CodeBuddy reads only the first existing file in its config priority chain: ~/.codebuddy/.mcp.json (recommended) → ~/.codebuddy/mcp.json (deprecated) → ~/.codebuddy.json (legacy). Edit the first non-empty file that exists — creating a higher-priority file would hide the servers in the ones below it. If none exist, create ~/.codebuddy/.mcp.json:
{
"mcpServers": {
"gitnexus": {
"command": "npx",
"args": ["-y", "gitnexus@latest", "mcp"]
}
}
}
Qoder
Add to ~/.qoder.json:
{
"mcpServers": {
"gitnexus": {
"command": "npx",
"args": ["-y", "gitnexus@latest", "mcp"]
}
}
}
How It Works
GitNexus builds a complete knowledge graph of your codebase through a multi-phase indexing pipeline:
- Structure — Walks the file tree and maps folder/file relationships
- Parsing — Extracts functions, classes, methods, and interfaces using Tree-sitter ASTs
- Resolution — Resolves imports and function calls across files with language-aware logic
- Field & Property Type Resolution — Tracks field types across classes and interfaces for deep chain resolution (e.g.,
user.address.city.getName()) - Return-Type-Aware Variable Binding — Infers variable types from function return types, enabling accurate call-result binding
- Field & Property Type Resolution — Tracks field types across classes and interfaces for deep chain resolution (e.g.,
- Clustering — Groups related symbols into functional communities
- Processes — Traces execution flows from entry points through call chains
- Search — Builds hybrid search indexes for fast retrieval
The result is a LadybugDB graph database stored locally in .gitnexus/ with full-text search and semantic embeddings.
Experimental community detection engine
Experimental — not supported for production indexes. The Icebug engine is a research path for #2337. It carries no stability guarantee, may change or be removed without a major version, and partitions differently from the default, so switching engines changes community IDs and any generated context keyed on them. Reindex with
graphologybefore relying on the output.
Community detection uses the bundled Graphology Leiden implementation by default. To try the #2337 Icebug path without changing default analyze behavior, install the optional native package alongside GitNexus and set the engine:
npm i @ladybugmem/icebug
GITNEXUS_COMMUNITY_ENGINE=icebug npx gitnexus analyze
Supported values are graphology, icebug, and auto. Today auto is behaviorally identical to icebug: both try Icebug and fall back to Graphology, while graphology skips Icebug entirely.
Icebug is not a declared dependency — its prebuilds link against system Arrow 24 (libarrow.so.2400), OpenMP, and glibc ≥ 2.38, none of which GitNexus can assume. Analyze falls back to Graphology and reports the reason in progress output when the module is missing, fails to load, or predates the setNumberOfThreads / setSeed controls that reproducible community IDs require (present at icebug-nodejs HEAD, absent from the published 12.8.0 tarball — so the fallback is what you will see today). The engine is pinned to threads: 1, randomize: false for determinism.
Note that the bundled Graphology path is no longer the slow option it once was: #2337 removed an accidental O(communities × N) copy in the vendored Leiden. On a synthetic 200k-node / 800k-edge benchmark graph it went from exceeding the 60s timeout to finishing in ~15s. Real projections vary with their degree distribution, so treat that as a direction, not a guarantee.
MCP Tools
Your AI agent gets 17 tools (15 per-repo + 2 group) automatically:
| Tool | What It Does |
|---|---|
list_repos |
Discover all indexed repositories (paginated — limit/offset) |
query |
Process-grouped hybrid search (BM25 + semantic + RRF) |
context |
360-degree symbol view — categorized refs, process participation |
impact |
Blast radius analysis with depth grouping and confidence |
trace |
Shortest directed path between two symbols (call + class-member edges) |
detect_changes |
Git-diff impact — maps changed lines to affected processes |
check |
Read-only structural checks against the indexed graph |
rename |
Multi-file coordinated rename with graph + text search |
cypher |
Raw Cypher graph queries |
route_map |
API route map — which components fetch which endpoints, and handlers |
tool_map |
MCP/RPC tool definitions — where they're defined and handled |
shape_check |
Validate API response shapes against consumers' property accesses |
api_impact |
Pre-change impact report for an API route handler |
explain |
Explain persisted taint findings (source→sink flows, --pdg indexes) |
pdg_query |
Query control/data dependence at statement level (--pdg indexes) |
group_list |
List configured repository groups |
group_sync |
Rebuild a group's Contract Registry and cross-repo links |
With one indexed repo, the
repoparam is optional. With multiple, specify which:query({search_query: "auth", repo: "my-app"}). Per-repo tools also take an optionalbranchfor indexes pinned withgitnexus analyze --branch; omitting it queries the workspace index, which follows your checked-out working tree.explainandpdg_queryneed an index built withgitnexus analyze --pdg.
MCP Resources
| Resource | Purpose |
|---|---|
gitnexus://repos |
List all indexed repositories (read first) |
gitnexus://setup |
Setup and usage guidance for agents |
gitnexus://repo/{name}/context |
Codebase stats, staleness check, and available tools |
gitnexus://repo/{name}/clusters |
All functional clusters with cohesion scores |
gitnexus://repo/{name}/cluster/{name} |
Cluster members and details |
gitnexus://repo/{name}/processes |
All execution flows |
gitnexus://repo/{name}/process/{name} |
Full process trace with steps |
gitnexus://repo/{name}/schema |
Graph schema for Cypher queries |
gitnexus://group/{name}/contracts |
A group's extracted contracts and cross-links |
gitnexus://group/{name}/status |
Staleness of repos in a group |
MCP Prompts
| Prompt | What It Does |
|---|---|
detect_impact |
Pre-commit change analysis — scope, affected processes, risk level |
generate_map |
Architecture documentation from the knowledge graph with mermaid diagrams |
CLI Commands
gitnexus setup # Configure MCP for detected editors (one-time; use -c to select)
gitnexus uninstall # Preview removal of GitNexus MCP/skills/hooks (add --force to apply)
gitnexus analyze [path] # Index a repository (or update stale index)
gitnexus analyze --repair-fts # Fast path: rebuild/verify only FTS indexes on existing index data
gitnexus analyze --force # Full rebuild: re-parse + graph rebuild + FTS rebuild
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
gitnexus embeddings install # Fetch the optional local embedding stack on demand (--cuda, --force)
gitnexus analyze --skills # Generate repo-specific skill files from detected communities
gitnexus analyze --skip-agents-md # Preserve custom AGENTS.md/CLAUDE.md gitnexus section edits
gitnexus analyze --skip-skills # Skip installing standard .claude/skills/gitnexus-* skill files
gitnexus analyze --skip-git # Index folders that are not Git repositories
gitnexus analyze --workers <n> # Parse worker pool size (>=1; default: cores-1, capped at 16)
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
gitnexus analyze --max-file-size 1024 # Skip files larger than N KB (default: 512, cap: 32768)
gitnexus analyze --worker-timeout 60 # Increase worker idle timeout for slow parses
gitnexus analyze --wal-checkpoint-threshold 67108864 # 64 MiB. Control LadybugDB WAL auto-checkpoint threshold (default: 67108864 = 64 MiB; -1 keeps Ladybug stock ~16 MiB)
gitnexus mcp # Start MCP server (stdio) — serves all indexed repos
gitnexus serve # Start local HTTP server (multi-repo) for web UI
gitnexus index # Register an existing .gitnexus/ folder into the global registry
gitnexus list # List all indexed repositories
gitnexus status # Show index status for current repo
gitnexus clean # Delete index for current repo
gitnexus clean --all --force # Delete all indexes
gitnexus wiki [path] # Generate LLM-powered docs from knowledge graph
gitnexus wiki --model <model> # Wiki with custom LLM model (default: minimax/minimax-m2.5)
gitnexus wiki --base-url http://llama-box.local:8080/v1 --allow-insecure-connection llama-box.local
# Allow an exact LAN/self-hosted HTTP LLM host; env: GITNEXUS_ALLOW_INSECURE_CONNECTION
gitnexus doctor # Show runtime platform capabilities and embedding configuration
# Direct graph queries — the same tools the MCP server exposes, no MCP daemon needed
gitnexus query "<concept>" # Process-grouped hybrid search
gitnexus context <symbol> [--uid <uid> | --file <path>] # 360° symbol view; flags disambiguate a shared name
gitnexus impact <symbol> [--uid <uid> | --file <path> | --kind <kind>] # Blast radius; flags disambiguate a shared name
gitnexus trace <from> <to> # Shortest directed path between two symbols
gitnexus detect-changes # Map the working-tree diff to affected symbols and execution flows
gitnexus check # Read-only structural checks against the indexed graph
gitnexus cypher "<query>" # Run a raw Cypher query against the knowledge graph
# Repository groups (multi-repo / monorepo service tracking)
gitnexus group create <name> # Create a repository group
gitnexus group add <group> <groupPath> <registryName> # Add a repo to a group. <groupPath> is a hierarchy path (e.g. hr/hiring/backend); <registryName> is the repo's name from the registry (see `gitnexus list`)
gitnexus group remove <group> <groupPath> # Remove a repo from a group by its hierarchy path
gitnexus group list [name] # List groups, or show one group's config
gitnexus group sync <name> # Extract contracts and match across repos/services
gitnexus group contracts <name> # Inspect extracted contracts and cross-links
gitnexus group query <name> <q> # Search execution flows across all repos in a group
gitnexus group status <name> # Check staleness of repos in a group
gitnexus group impact <name> --target <symbol> --repo <groupPath> # Cross-repo blast radius
gitnexus uninstallreversesgitnexus setup— it removes the GitNexus MCP entries, hooks, and skill directories it added to each detected editor. Skill directories are identified by bundled gitnexus skill name (e.g.gitnexus-cli/), so if you customized files inside an installed skill directory, back them up first. It is a dry-run preview by default and prints the exact paths it would remove; pass--forceto apply. Per-repo indexes (gitnexus clean --all) and the global npm package (npm uninstall -g gitnexus) are left for you to remove.
Remote Embeddings
Set these env vars to use a remote OpenAI-compatible /v1/embeddings endpoint instead of the local model:
export GITNEXUS_EMBEDDING_URL=http://your-server:8080/v1
export GITNEXUS_EMBEDDING_MODEL=BAAI/bge-large-en-v1.5
export GITNEXUS_EMBEDDING_DIMS=1024 # optional, default 384
export GITNEXUS_EMBEDDING_REQUEST_DIMS=omit # optional: omit "dimensions", or an integer to override it
export GITNEXUS_EMBEDDING_API_KEY=your-key # optional, default: "unused"
export GITNEXUS_EMBEDDING_MAX_ATTEMPTS=3 # optional, total attempts (1-20)
export GITNEXUS_EMBEDDING_RETRY_CAP_MS=5000 # optional, maximum retry delay
export GITNEXUS_EMBEDDING_MIN_INTERVAL_MS=0 # optional, minimum request spacing
export GITNEXUS_EMBEDDING_HTTP_TIMEOUT_MS=180000 # optional, per-request timeout (max 300000)
gitnexus analyze . --embeddings
GITNEXUS_EMBEDDING_REQUEST_DIMS controls only the dimensions field sent in
the request body, independently of GITNEXUS_EMBEDDING_DIMS (which still
validates the returned vector's length):
omit(ornone,off,false,0) — do not senddimensionsat all, for strict backends that return the right vector size but reject the field.- a positive integer — send that value instead of
GITNEXUS_EMBEDDING_DIMS. - unset — send
GITNEXUS_EMBEDDING_DIMS(the previous behavior).
Works with Infinity, vLLM, TEI, llama.cpp, Ollama, LM Studio, or OpenAI. Retry and pacing settings are provider-neutral; provider-specific limits should be supplied through configuration. When unset, local embeddings are used unchanged.
JVM Package Sibling Injection
Java and Kotlin files in the same package receive implicit sibling class bindings
to resolve same-package references. By default, GitNexus injects at most 200
siblings per module scope, nearest first by path. Set
GITNEXUS_MAX_INJECTED_SIBLINGS=0 to remove that per-file limit; this can
substantially increase indexing work for large packages.
export GITNEXUS_MAX_INJECTED_SIBLINGS=200
gitnexus analyze .
When the limit truncates a file's sibling set, that file is marked
visibility-incomplete: same-package references still resolve through the
injected siblings, but wildcard-import attribution (used by the Spring
bean/DI/config passes) is disabled for it rather than resolved against a
partial view. Analyze logs a sibling injection truncated warning naming how
many files were affected.
Packages with more than 500 files are a separate, fixed limit: they are skipped
entirely (logged as skipping package with N files) and every file in them is
marked visibility-incomplete. GITNEXUS_MAX_INJECTED_SIBLINGS does not lift
that skip — including at 0.
Multi-Repo Support
GitNexus supports indexing multiple repositories. Each gitnexus analyze registers the repo in a global registry (~/.gitnexus/registry.json). The MCP server serves all indexed repos automatically.
Supported Languages
TypeScript, JavaScript, Python, Java, C, C++, C#, Go, Rust, PHP, Kotlin, Swift, Ruby, Dart
Language Feature Matrix
| Language | Imports | Named Bindings | Exports | Heritage | Type Annotations | Constructor Inference | Config | Frameworks | Entry Points |
|---|---|---|---|---|---|---|---|---|---|
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| JavaScript | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Java | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Kotlin | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Go | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Rust | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| PHP | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ |
| Ruby | ✓ | — | ✓ | ✓ | — | ✓ | — | ✓ | ✓ |
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Dart | ✓ | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
Imports — cross-file import resolution · Named Bindings — import { X as Y } / re-export tracking · Exports — public/exported symbol detection · Heritage — class inheritance, interfaces, mixins · Type Annotations — explicit type extraction for receiver resolution · Constructor Inference — infer receiver type from constructor calls (self/this resolution included for all languages) · Config — language toolchain config parsing (tsconfig, go.mod, etc.) · Frameworks — AST-based framework pattern detection · Entry Points — entry point scoring heuristics
Agent Skills
GitNexus ships with skill files that teach AI agents how to use the tools effectively:
- Exploring — Navigate unfamiliar code using the knowledge graph
- Debugging — Trace bugs through call chains
- Impact Analysis — Analyze blast radius before changes
- Refactoring — Plan safe refactors using dependency mapping
- Guide — GitNexus tool/resource/schema reference for the agent
- CLI — Run analyze/status/clean/wiki commands on request
- PDG Query — Statement-level control/data dependence queries (
--pdgindex) - Taint Analysis — Source→sink data-flow findings (
--pdgindex) - Plan / Work / Review / LFG — The engineering family: implementation-ready plans, gated plan execution, graph-backed change review with taint + expert lenses, and the end-to-end pipeline
Installed automatically by both gitnexus analyze (per-repo) and gitnexus setup (global). Run gitnexus analyze --skills to additionally generate each detected functional area as a direct project skill under .claude/skills/gitnexus-area-<name>/.
Requirements
- Node.js >= 22
- Git repository (uses git for commit tracking)
- Linux: glibc 2.34 or newer (Ubuntu 22.04+, RHEL/Rocky/Alma 9+, Debian 12+, Fedora 35+). The
LadybugDB native binary ships as a prebuild against that floor, so on an older host it cannot
load and reinstalling does not help — see
Linux:
GLIBC_2.34' not found. - Windows, for full-text search: the Microsoft Visual C++ 2015-2022 Redistributable (x64) and
OpenSSL 3 (
libssl-3-x64.dll,libcrypto-3-x64.dll) resolvable onPATH— see Windows: full-text search unavailable.
Release candidates
Stable releases publish to the default latest dist-tag. When a pull request
with non-documentation changes merges into main, an automated workflow also
publishes a prerelease build under the rc dist-tag, so early adopters can
try in-flight fixes without waiting for the next stable cut. (Docs-only
merges are skipped.)
# Try the latest release candidate (pre-stable — may change at any time)
npm install -g gitnexus@rc
# — or —
npx gitnexus@rc analyze
Release-candidate versions follow the standard semver prerelease format
X.Y.Z-rc.N, where X.Y.Z is the next stable target (bumped from the
current latest by patch by default; minor or major when kicking off a
bigger cycle) and N increments per published rc. Example sequence:
1.6.2-rc.1, 1.6.2-rc.2, …, then once 1.6.2 ships stable,
1.6.3-rc.1. See the Releases page
for the full list; stable latest is unaffected.
Troubleshooting
Cannot destructure property 'package' of 'node.target' as it is null
This error comes from npm 11.x's arborist while installing gitnexus (often via npx), before gitnexus code runs. It is triggered by platform-filtered optionalDependencies in native packages such as onnxruntime-node / @huggingface/transformers (used when indexing with --embeddings). GitNexus cannot catch it at runtime — use one of these workarounds:
pnpm --allow-build=@ladybugdb/core --allow-build=gitnexus --allow-build=tree-sitter dlx gitnexus@latest analyze # auto-selected when pnpm + npm 11+
npm install -g gitnexus@latest # global install avoids per-run npx reify
gitnexus analyze # if already installed globally
On pnpm 10+, lifecycle scripts are blocked unless explicitly allowed — the resolver adds --allow-build for @ladybugdb/core, gitnexus, and tree-sitter automatically when it picks pnpm dlx.
If you must stay on npm 11.x without pnpm, downgrade npm toolchain-wide (last resort):
npm install -g npm@10.9.0
See #1939 and the original #819 thread. An older variant of this crash (tree-sitter-dart tarball URL) was fixed in gitnexus v1.6.2+ (#820); if you still see install failures after upgrading, clear cache:
npm cache clean --force
npx gitnexus@latest analyze
ERR_DLOPEN_FAILED / lbugjs.node missing (pnpm dlx, pnpx)
GitNexus depends on @ladybugdb/core, whose native database addon
(lbugjs.node) is placed by a postinstall script. pnpm dlx, pnpx, and any
install run with --ignore-scripts skip lifecycle scripts, so the addon is
never put in place and the runtime crashes with ERR_DLOPEN_FAILED:
Error: dlopen(.../@ladybugdb/core/lbugjs.node, ...): tried: '...' (no such file)
code: 'ERR_DLOPEN_FAILED'
Options that run install scripts:
# pnpm dlx with explicit build permission (one-off, no global install required)
pnpm --allow-build=@ladybugdb/core --allow-build=gitnexus --allow-build=tree-sitter \
dlx gitnexus@latest serve
# npm: global install (recommended on npm 11+; bare npx may crash — see section above)
npm install -g gitnexus@latest
gitnexus serve
# npx (npm < 11, or after upgrading npm)
npx gitnexus@latest serve
# pnpm: global install with build scripts allowed (pnpm 10.2+; no approve-builds -g on pnpm 11+)
pnpm add -g --allow-build=@ladybugdb/core --allow-build=gitnexus --allow-build=tree-sitter gitnexus
gitnexus serve
Linux: GLIBC_2.34' not found
LadybugDB native binary (lbugjs.node) exists but failed to load:
/lib64/libc.so.6: version `GLIBC_2.34' not found (required by .../lbugjs.node)
The LadybugDB addon ships as a prebuilt binary compiled against glibc 2.34. If your distribution is older (CentOS/RHEL 8 has 2.28, Ubuntu 20.04 has 2.31, Debian 11 has 2.31), the dynamic loader cannot resolve its symbols.
Reinstalling does not help — every download delivers the same prebuilt binary. The fix is a newer C library:
- Run GitNexus on a distribution with glibc 2.34 or newer — Ubuntu 22.04+, RHEL/Rocky/Alma 9+, Debian 12+, Fedora 35+.
- Or run it in the container image, which bundles a current glibc (see Docker).
gitnexus doctor reports the required and detected glibc versions when this happens
(#2672).
Windows: full-text search unavailable
analyze completes, but keyword search is degraded and doctor shows the FTS extension failing
with Windows error 126 (The specified module could not be found). The extension needs two
runtime dependencies Windows does not ship by default:
- Microsoft Visual C++ 2015-2022 Redistributable (x64) — https://aka.ms/vs/17/release/vc_redist.x64.exe
- OpenSSL 3 —
libssl-3-x64.dllandlibcrypto-3-x64.dll, resolvable onPATH
The redistributable alone is not sufficient. If Git for Windows is installed you already have
the OpenSSL DLLs — run gitnexus from Git Bash, or prepend the directory to PATH in the
shell you use:
$env:PATH = "C:\Program Files\Git\mingw64\bin;$env:PATH"
gitnexus analyze --repair-fts
Without them the index is still built, but without search tables, so query returns empty keyword
results until you re-run gitnexus analyze --repair-fts from a shell where the DLLs resolve
(#2669).
Installation fails with native module errors
Some optional language grammars (Dart, Proto, Swift, Kotlin) require native compilation. If they fail, GitNexus still works — those languages will be skipped. To skip them intentionally (no C++ toolchain needed), set GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1 before installing.
If npm install -g gitnexus fails on native modules:
# Ensure build tools are available (Linux/macOS)
# Ubuntu/Debian: sudo apt install python3 make g++
# macOS: xcode-select --install
# Retry installation
npm install -g gitnexus
Installation fails behind an HTTP proxy (onnxruntime-node postinstall)
onnxruntime-node's postinstall downloads optional CUDA GPU binaries from api.nuget.org — outside the npm registry, so registry mirrors don't cover it, and its proxy layer (global-agent) ignores the standard HTTP_PROXY/HTTPS_PROXY variables and rejects 302 redirects (#2370).
Since the packages are optional dependencies, a failed download no longer breaks npm install -g gitnexus — npm skips the embedding stack and everything else works. The stack then self-heals on demand: the first gitnexus analyze --embeddings (or an explicit gitnexus embeddings install) fetches it through your configured npm registry — mirrors and proxies apply, no NuGet download involved — into ~/.gitnexus/embedding-runtime.
# heal a proxy-degraded install manually (CPU embeddings; registry-only)
gitnexus embeddings install
# reinstall into the prefix even when the stack already resolves
gitnexus embeddings install --force
# CUDA GPU hosts: also fetch GPU binaries (NuGet; set the proxy global-agent reads)
GLOBAL_AGENT_HTTPS_PROXY=<proxy-url> gitnexus embeddings install --cuda
The prefix defaults to ~/.gitnexus/embedding-runtime; set GITNEXUS_EMBEDDING_RUNTIME_DIR to install it elsewhere (e.g. a writable path in a container).
Node requirement for the on-demand prefix: the self-heal loads the prefixed packages via
module.registerHooks, available on Node ≥ 22.15 (on the 22.x line) or ≥ 23.5 (on the 23.x line). On an older Node the packages install but can't be loaded from the prefix — reinstall them into the install itself instead (works on every supported Node):ONNXRUNTIME_NODE_INSTALL=skip npm install -g gitnexus(Windows:set ONNXRUNTIME_NODE_INSTALL=skip && npm install -g gitnexus). Skipping only the CUDA download keeps full CPU embeddings (CPU embeddings don't need it). Check the result any time withgitnexus doctor(Embeddings → Support line).
Analyze warns about unavailable FTS or VECTOR extensions
GitNexus uses optional DuckDB extensions for BM25 and vector search. The gitnexus serve and MCP read paths only ever try to LOAD the extensions — they never block on a network install. The analyze command, by default, attempts one bounded out-of-process install if LOAD fails (a plain INSTALL to download a missing extension, escalating to FORCE INSTALL only when the LOAD error shows the existing file is broken or truncated, so a permanent non-file failure does not re-download on every run) and proceeds even when that install times out, so the index is always written to disk; BM25/vector search degrade gracefully until the extensions become available.
Configure the behavior with these environment variables:
| Variable | Values | Default | Effect |
|---|---|---|---|
GITNEXUS_LBUG_EXTENSION_INSTALL |
auto, load-only, never |
auto |
auto runs one bounded install if LOAD fails — a plain INSTALL, escalating to FORCE INSTALL only when the LOAD error shows the present extension file is broken. load-only only uses already-installed extensions (recommended for offline / firewalled environments). never skips optional extensions entirely. |
GITNEXUS_LBUG_EXTENSION_INSTALL_TIMEOUT_MS |
positive integer | 15000 |
Wall-clock budget for the out-of-process extension-install child before it is killed. |
GITNEXUS_FTS_STEMMER |
supported LadybugDB stemmer | porter |
Stemmer used when rebuilding BM25/FTS indexes. Use none for CJK-heavy repositories, or a language stemmer such as german, french, or spanish when that better matches repository comments and identifiers. Re-run gitnexus analyze --repair-fts after changing it. |
GITNEXUS_FTS_CJK_SEGMENTATION |
none, bigram |
none |
bigram inserts overlapping character-bigram boundaries into Chinese/Japanese Han-ideograph spans in content/description before FTS indexing, so LadybugDB's space-only tokenizer can see sub-phrase word boundaries. Scoped to CJK Unified Ideographs only — Japanese Hiragana/Katakana and Korean Hangul are not currently segmented. Unlike GITNEXUS_FTS_STEMMER, this rewrites stored text — enabling it on an already-indexed repo requires a full gitnexus analyze --force; neither --repair-fts nor a plain incremental analyze applies it to previously-indexed files. Set the same value wherever analyze and search-serving processes (CLI query, MCP server, web server) run. |
GITNEXUS_STREAM_GRAPH_EMIT |
0, 1 |
1 (on) |
On by default on a full rebuild (--force); incremental runs ignore it. Holds structural relationships (CALLS, IMPORTS, ACCESSES, CONTAINS, ...) as CSV-on-disk plus compact in-memory columns instead of as objects in three overlapping indexes, cutting peak in-memory graph heap by ~1.4x at no measurable CPU cost (measured A/B on a synthetic 400k-node / 1.08M-edge graph: 819 MB -> 584 MB, iteration at parity, scaling verified linear from 100k to 800k nodes, with every edge still visible through the graph interface; no end-to-end measurement on a real repository yet). Nothing is traded away — community detection, process extraction, PDG taint summaries and the local-symbol pruner all read a complete relationship set and behave identically. Set to 0 only to bisect a suspected streaming-related fault. |
GITNEXUS_COMMUNITY_ENGINE |
graphology, icebug, auto |
graphology |
Community-detection engine used during analyze. graphology is the supported default. icebug and auto are experimental and currently behave identically: both try the optional @ladybugmem/icebug native Leiden over a CSR export and fall back to Graphology if it is not installed, cannot load, or lacks the deterministic thread/seed controls. Experimental engines partition differently, so community IDs are not comparable across engines. |
GITNEXUS_WAL_CHECKPOINT_THRESHOLD |
integer >= -1 |
67108864 (64 MiB) |
LadybugDB WAL auto-checkpoint threshold during analyze (bytes). Auto-checkpoint remains enabled; -1 keeps Ladybug's stock ~16 MiB. Larger thresholds reduce checkpoint frequency but increase the WAL size at rotation time — choose a smaller value on disk-constrained environments. |
GITNEXUS_LBUG_BUFFER_POOL_SIZE |
integer >= 0 (bytes) |
min(2 GiB, 80% RAM) | LadybugDB buffer-pool ceiling for every GitNexus database (analyze, MCP server, serve, group bridges). Bounded so a long-lived gitnexus mcp process or a large incremental analyze cannot grow toward LadybugDB's native 80%-of-RAM default and OOM the host (#2557). 0 restores that native unbounded default; invalid values warn and fall back to the default. During analyze the pool is right-sized to the graph and, on non-4 KiB-page hosts (Apple Silicon 16 KiB, Ascend/aarch64 64 KiB), scaled by the page-size granule ratio up to min(2 GiB × pageSize/4 KiB, 80% RAM) (#2631); this env var overrides all of that as an absolute value. |
GITNEXUS_LBUG_MAX_DB_SIZE |
positive integer (bytes) | 17179869184 (16 GiB) |
Upper bound for a single LadybugDB database file. This is an mmap/disk-address-space ceiling, not a memory limit — it does not constrain the buffer pool (use GITNEXUS_LBUG_BUFFER_POOL_SIZE for that). Raise it when indexing genuinely huge monorepos; invalid values silently fall back to the default. |
# Offline/airgapped: never reach the network for extensions
GITNEXUS_LBUG_EXTENSION_INSTALL=load-only npx gitnexus analyze
# Slow network: give extension downloads more time
GITNEXUS_LBUG_EXTENSION_INSTALL_TIMEOUT_MS=30000 npx gitnexus analyze
# CJK-heavy codebase: rebuild keyword indexes without English stemming
GITNEXUS_FTS_STEMMER=none npx gitnexus analyze --repair-fts
# CJK-heavy codebase: enable sub-phrase search over Chinese/Japanese Han text.
# On an already-indexed repo, the first run after enabling this MUST be --force —
# --repair-fts and plain incremental `analyze` both leave old files un-segmented.
GITNEXUS_FTS_CJK_SEGMENTATION=bigram npx gitnexus analyze --force
Analysis runs out of memory
Memory management is automatic: analyze sizes its heap to the machine
(always below physical RAM), caps each parse worker, and — rather than
grinding into a GC death spiral or crash — stops early with a message telling
you the one thing to do. Repeated
Replacement worker did not report ready within 5000ms warnings on a large
repository are part of the same picture: memory pressure starving healthy
workers, not a worker bug (#2649).
If analyze says the repository doesn't fit, do what the message says:
- The machine has more memory to give (a
NODE_OPTIONS--max-old-space-sizepin from your environment is holding analyze back): re-run without the pin — no flags needed. - The machine is the ceiling: shrink the scope (exclude generated or vendored directories, below) or use a machine with more RAM.
Escape hatches (GITNEXUS_MEMORY=off to decline the autopilot,
GITNEXUS_WORKER_HEAP_MB to size workers yourself) are listed in the
environment-variable table below —
most users never need them.
For very large repositories:
# Increase Node.js heap size
NODE_OPTIONS="--max-old-space-size=16384" npx gitnexus analyze
# Exclude large directories (this repo only)
echo "vendor/" >> .gitnexusignore
echo "dist/" >> .gitnexusignore
# Exclude a directory across every repo you index, without touching each
# repo's own .gitnexusignore or needing push/commit access to it. GitNexus
# reads the same sources `git` itself does: core.excludesFile (all repos)
# and $GIT_DIR/info/exclude (this repo only, untracked). A repo's own
# .gitignore/.gitnexusignore can still override either with a `!pattern`
# negation. Skip both entirely with GITNEXUS_NO_GLOBAL_IGNORE=1.
git config --global core.excludesFile ~/.gitignore_global # applies to every repo
echo "docs/" >> ~/.gitignore_global
echo "build/" >> .git/info/exclude # this repo only, untracked
Large files are being skipped
By default the walker skips files larger than 512 KB (see log line Skipped N large files (>512KB)). Raise the threshold via either the CLI flag or the environment variable — both accept a value in KB:
# CLI flag (takes precedence over the env var)
npx gitnexus analyze --max-file-size 2048 # skip only files > 2 MB
# Environment variable (persists across commands)
export GITNEXUS_MAX_FILE_SIZE=2048
npx gitnexus analyze
Values above 32768 KB (32 MB) are clamped to the tree-sitter parser ceiling; invalid values fall back to the 512 KB default with a one-time warning. When an override is active, analyze prints the effective threshold in its startup banner (e.g. GITNEXUS_MAX_FILE_SIZE: effective threshold 2048KB (default 512KB)).
Analyze reports a worker timeout
Worker parse timeouts are recoverable. GitNexus retries stalled worker jobs with backoff, splits large jobs to isolate slow files, and quarantines a file that repeatedly crashes its worker (respawning the slot so the pool keeps going). If a large repository needs more time per worker job, use either:
# CLI flag, in seconds
npx gitnexus analyze --worker-timeout 60
# Environment variable, in milliseconds
export GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS=60000
npx gitnexus analyze
For repositories with very large source files, GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES controls the worker job byte budget. The default is 8388608 bytes (8 MB).
Worker pool resilience tuning
Four env vars expose the pool's resilience layers (respawn budget, cumulative-timeout cap, circuit breaker, startup handshake). Defaults are tuned for typical repos; bump them when an analyze legitimately needs more retries, or lower them to fail-fast on a known-bad shape.
| Variable | Default | Effect |
|---|---|---|
GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT |
3 |
Max replacement spawns per slot before the slot is dropped from the active rotation. |
GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS |
5 × subBatchTimeoutMs |
Total retry wall-time budget per job before quarantining. Bounds exponentially-growing retry waits. |
GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD |
max(3, poolSize) |
Per-slot consecutive deaths before the pool's circuit breaker trips. After tripping, dispatches require a fresh pool. |
GITNEXUS_WORKER_SHUTDOWN_DRAIN_MS |
30000 |
Max wait at pool shutdown for a retired worker still inside native code — terminated at its next JS-safe point instead of mid-native-call, which would abort the process (Napi::Error, #2432). |
GITNEXUS_WORKER_READY_TIMEOUT_MS |
5000 |
Startup budget for a parse worker to load its grammar bindings and report {type:'ready'}. Slots that miss it are treated as startup crashes. Raise it on a slow or heavily loaded host where a full pool cold-starting concurrently needs more than 5s. |
GITNEXUS_MEMORY |
off |
unset (autopilot on) |
GITNEXUS_WORKER_HEAP_MB |
clamp(512, RAM/2/poolSize, 4096) |
Per-worker V8 old-generation heap cap (#2649). Bounds pool RSS on large repos; a worker exceeding it dies with a real heap error handled by quarantine/respawn. |
GITNEXUS_SERVER_ANALYZE_HEAP_MB |
min(8192, auto cap) |
Heap for the web/MCP server's forked analyze worker (#2649). Defaults to the historical 8192 MB bounded by the machine/container's RAM-aware auto cap; set an absolute MB value to override. |
GITNEXUS_CPP_CAPTURE_BUDGET_MS |
20000 |
Per-file wall-clock budget for C++ capture extraction; on breach the file keeps partial captures with a warning (#2432). 0 expires immediately. |
Graph cleanup tuning
After scope resolution, analyze prunes inert block-local value symbols (a function-local const/let/var that ends up with only its structural File→DEFINES edge) to keep the graph focused on cross-symbol relationships. Module/file-scope symbols, class members, and any local with a real edge are always kept.
| Variable | Default | Effect |
|---|---|---|
GITNEXUS_KEEP_LOCAL_VALUE_SYMBOLS |
unset | Set to 1/true to keep inert block-local value symbols instead of pruning them. |
Programmatic callers can pass keepLocalValueSymbols: true in PipelineOptions instead of setting the env var.
Scope-resolution property-key dispatch cap
During scope resolution GitNexus synthesizes CALLS edges through property-key
dispatch — call sites like hooks.emitScopeCaptures() where a property key is
registered by multiple definitions across the codebase. To keep this fan-in
bounded, each property key is capped at 32 registrations: a key registered
by more than 32 distinct functions is skipped entirely (no CALLS are synthesized
through it), and the dropped key names are surfaced in the analyze log for
operator visibility. The cap is calibrated at 2× this repo's own provider table
(16 legitimate registrations, one per language provider).
| Variable | Default | Effect |
|---|---|---|
GITNEXUS_MAX_PROPERTY_DISPATCH_FANOUT |
32 |
Per-property-key registration cap in the property-dispatch scope-resolution pass. Set to a positive integer to raise it for repositories whose provider/hook tables exceed the default and lose CALLS coverage on a legitimate key; non-integer or < 1 values fall back to 32. Lowering it tightens the overflow budget. |
# A property key registered by 40 functions overflows the default 32 and drops
# all CALLS through it — raise the cap for that repo and rebuild so scope
# resolution reruns.
export GITNEXUS_MAX_PROPERTY_DISPATCH_FANOUT=64
npx gitnexus analyze --force
Scope-resolution dispatch-target cap
During scope resolution GitNexus resolves calls that flow through callable
values — function/method references bound to variables, passed as arguments,
or stored in maps/tables. To keep that inclusion-based resolution finite, each
callable site is capped at 32 dispatch targets. When a site gathers more
candidates than the cap it is treated as overflowed and all of its call
edges are dropped — a cliff, not a tail, so a repository with a legitimately
wide dispatch table (a single callable site resolving to 33+ targets) loses
that site's whole call chain. In that case analyze logs
callable-value-flow: candidate set exceeded the cap; no partial CALLS emitted
alongside a warning carrying the language, the overflowing context, the
candidate count, and the cap (32).
Raise the cap for such repositories:
| Variable | Default | Effect |
|---|---|---|
GITNEXUS_MAX_CALLABLE_VALUE_TARGETS |
32 |
Per-callable-site dispatch-target cap in the callable-value-flow scope-resolution pass. Set to a positive integer to raise it for repositories whose wide dispatch tables overflow the default and lose a whole call chain; non-integer or < 1 values fall back to 32. Lowering it tightens the overflow budget. |
# A callable site resolving to 48 targets overflows the default 32 and drops
# the chain — raise the cap for that repo and rebuild so scope resolution reruns.
export GITNEXUS_MAX_CALLABLE_VALUE_TARGETS=64
npx gitnexus analyze --force
Hook augmentation and skip diagnostics
The Claude Code / Antigravity hooks keep their stderr silent on normal skip
paths so strict hook runners (e.g. Codex PreToolUse) never see unexpected
diagnostic output.
When a GitNexus process holds the repo DB write lock (the common case — the MCP
server is running, or the DB-lock probe timed out and failed closed), the local
CLI augment can't run (LadybugDB is single-writer). Rather than drop the
augmentation, the hook hands the agent a short, conditional MCP-query hint on
stdout (the sanctioned additionalContext channel) — "if the GitNexus MCP tools
are live in this session, call query …" — so an agent that has the tools can
still fetch graph-ranked context. The hint is throttled to at most once per repo
per window (GITNEXUS_MCP_HINT_THROTTLE_MS, default 10 min; 0 disables), so an
owner-locked session isn't nudged on every search. A stale-index reminder, or an
already-current index, stays silent.
To see why a hook skipped the CLI augment, set GITNEXUS_DEBUG=1 and re-run the
action — the hook writes the reason (e.g. [GitNexus] augment skipped: MCP server owns DB) and the stale-index hint to its stderr:
GITNEXUS_DEBUG=1 <your command> # surfaces hook skip/diagnostic reasons on stderr
Only GITNEXUS_DEBUG=1 and GITNEXUS_DEBUG=true enable diagnostics; every other
value (including 0 and false) is treated as off. Diagnostics go to stderr
only — the hook's structured stdout (the JSON the agent consumes) is unaffected.
Privacy
- All processing happens locally on your machine
- No code is sent to any server
- Index stored in
.gitnexus/inside your repo (gitignored) - Global registry at
~/.gitnexus/stores only paths and metadata
Web UI
GitNexus also has a browser-based UI at gitnexus.vercel.app — 100% client-side, your code never leaves the browser.
Local Backend Mode: Run gitnexus serve and open the web UI locally — it auto-detects the server and shows all your indexed repos, with full AI chat support. No need to re-upload or re-index. The agent's tools (Cypher queries, search, code navigation) route through the backend HTTP API automatically.
License
Free for non-commercial use. Contact for commercial licensing.