Compare commits

...
6 Commits
Author SHA1 Message Date
9eaf2e6c4e perf(mcp): cut the analyze-only language-provider closure out of MCP server startup (#2802) (#2806)
* fix(mcp): key the empty-ascent note on CALL_SUMMARY data, not language (#2802)

`pdg-impact.ts` decided whether to append a "return-value ascent is
TypeScript/JavaScript-only" caveat to the `impact(mode:'pdg')` note by
looking up the criterion file's language. That put language-specific
logic in a layer that must be language-agnostic, and it was a lossy proxy
for a fact the graph already holds.

Whether the ascent can fire is a property of the persisted CALL_SUMMARY
edges. The descent already computes it, so thread the resolved-callee and
return-flowing counts out of `interproceduralDescent` and key the note on
those instead.

Three defects the language proxy carried, all gone:

  - Wrong for `.mjs`/`.cjs`/`.mts`/`.cts`: the provider registry's
    extension arrays omit them while the ingestion pipeline parses them
    as TS/JS, so those files were harvested but the note claimed their
    ascent was empty.
  - Silently stale: any language whose harvester started recording formal
    indices would keep getting the caveat until someone edited the list.
  - Wrong in reverse: a TS/JS callee with no return-flow got no caveat, so
    an ascent that found nothing read like one that covered the slice.

`pdg-impact.ts` now names no language and imports nothing from the
language layer, which also drops the analyze-only provider closure from
MCP server startup. Measured on overlayfs against a full build:

  import mcp/local/local-backend.js  before  565-648 ms / 548 modules
  import mcp/local/local-backend.js  after   458-463 ms / 170 modules

Tests hold CALL_SUMMARY content fixed while varying the file extension
across nine languages and assert the note text is identical, then hold the
extension fixed and vary the summary to show the note tracks the data.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(mcp): guard MCP startup against the language-provider closure returning

The eager `pdg-impact.ts -> core/ingestion/languages` edge was found and
lost once already during #2793 before #2802 re-derived it, so it gets a
test rather than a comment.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(lbug): record why csv-generator is not lazy-imported

#2802 proposed cutting `csv-generator.js` out of the adapter chain to
shorten MCP server startup. Measured on a native filesystem, the marginal
cost is small relative to the siblings this module already imports, and
`core/search/bm25-index.ts` statically imports `normalizeFtsText` from the
same module on a path `local-backend.ts` reaches dynamically for FTS — so
deferring would relocate the cost to first query, not remove it.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(pdg): pin chained receiver calls reaching BasicBlock.calleeIds

The PDG inter-procedural descent hops through `BasicBlock.calleeIds`, so
it can only cross a call boundary the resolver resolved. Chained receiver
calls reach `calleeIds` through the receiver-typing pass's own
`calleeIdSink` — a separate path from plain calls.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(analyze): drop the stale per-language cross-reference (#2802 review P3-4)

`pdgModeMismatch`'s comment told readers to keep "the diagnostic
per-language refinement in the impact CONSUMER (see pdg-impact.ts
assemblePdgImpactResult)". That refinement is no longer per-language —
removing it is the point of #2802, which now keys the empty-ascent note on
the persisted CALL_SUMMARY data instead.

The comment's real invariant is untouched and still correct: the values in
`resolvePdgConfig` must stay scalar, because the comparison below is a
shallow `!==` and an object would compare by reference. Only the
cross-reference was stale.

Comment-only; no executable line changes.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(mcp): probe the real module loader for the startup language closure (#2802 review P1-2)

The previous guard hand-rolled a regex walk over TypeScript source to
assert `core/ingestion/languages` was not statically reachable from MCP
startup. Four bypasses were reproduced against it, any one of which let
the exact 226-module regression return while the test stayed green:

  a. Wrong entry root. It walked from `mcp/local/local-backend.ts`, but the
     server module is `mcp/server.ts` — which imports LocalBackend as
     `import type`, so the guard's anchor was not even on server.ts's
     runtime closure. Ten real startup modules sat outside it.
  b. A top-level `await import(...)` executes during module evaluation, so
     it is eager at startup — but the walker skipped every `import(...)`
     by construction.
  c. The `import type` strip deleted a 16,445-character window of
     `pdg-impact.ts`: an `export type X =` matched lazily to the next
     `from "…"`, which lives inside a string literal. Any import in that
     window was invisible.
  d. The comment strip treated a `/*` inside a string literal as a comment
     opener.

Replace the approximation with a real module-load probe: spawn a child
node process per entry, import the built `dist/` entry, and report what
the loader actually pulled in. Rooted at `dist/mcp/server.js` and
`dist/cli/mcp.js` (the real startup entries) plus
`dist/mcp/local/local-backend.js`. Syntax cannot fool it.

One deviation from the two existing sibling probes is load-bearing:
`dist/` is ESM, so a `require.cache` diff alone cannot see the first-party
`dist/**` graph — it only catches CJS and native modules, which is why
`import-closure.test.ts` gets away with it (it asserts on
`@ladybugdb/core`). A pure cache diff here would have reported zero
language modules unconditionally, i.e. a new vacuous guard. This probe
unions `module.registerHooks({ load })` with the cache diff, and each
entry carries a non-vacuity anchor and a module floor so an empty result
fails loudly.

Verified load-bearing: adding a top-level
`await import('../core/ingestion/languages/index.js')` to
`src/mcp/resources.ts` and rebuilding turns `dist/mcp/server.js` red with
70+ named offenders, while the `local-backend` and `cli/mcp` cases stay
green — which is bypass (a) demonstrated directly. The old guard passed
that poisoned tree entirely.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(lbug): drop the unreproducible 9p multiplier from the csv-generator note (#2802 review P3-2)

The comment justifying why `csv-generator.js` is NOT lazy-imported carried
a hard "~40x" figure for how much a 9p mount inflates per-file ESM
resolve. Three independent measurements during review produced ~40x, ~7.3x
and ~30x, so the multiplier is not a reproducible quantity and had no
business being stated as one in a durable comment.

Reworked so the STRUCTURAL argument leads and the numbers only support it.
That argument is what actually settles the question and it does not rot:
`core/search/bm25-index.ts` statically imports `normalizeFtsText` from
`csv-generator.js`, and `local-backend.ts` reaches bm25-index through a
dynamic import on the FTS query path — so deferring here relocates the
cost to first query rather than removing it. Both verified again at
`bm25-index.ts:15` and `local-backend.ts:2756`.

Remaining figures are re-measured, attributed to a date and issue, and
labelled by filesystem: ~1.6 ms marginal (median of 45 cold imports on
local disk) versus ~50 ms for the same import on a network mount, stated
as environment-bound rather than as a property of the module. The
provider-registry cost is given as "several hundred modules" — the static
walk, the runtime hook, and the reviewer's probe each counted it
differently (375 / 439 / 407), so no single number was picked to go stale.
The old "226 modules" was real but counted only the `languages/` subtree
and undercounted the win.

Also repoints the trailing reference to the guard's new home at
`test/integration/mcp/startup-language-closure.test.ts` (same comment
block, inseparable from this rewrite).

Comment-only; no executable line changes.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(mcp): stop the empty-ascent note asserting a fact an undecodable summary contradicts (#2802 review P2-2)

The note claimed "this is a property of the persisted summaries" whenever
the descent resolved callees and none carried a return-flow. But
`decodeCallSummary` never throws by design: a version-skewed (`2|r:1`),
corrupt (`1|r:zz`), or NULL `reason` yields no entry, which was
indistinguishable from a cleanly-decoded empty summary. So the note could
assert "no formal parameter is recorded as flowing to its return value"
about a callee whose CALL_SUMMARY actually records `p0 -> return`.
`meta.pdg.hasCallSummary` is a plain boolean and stores no codec version,
so nothing else caught it.

`calleesWithReturnFlow` now reports three outcomes instead of two —
flowing, decoded-empty, and undecodable — and the undecodable count is
threaded through the descent to the note. When it is non-zero the note
says so and points at a re-index; when every summary decoded, the
persisted-summaries claim is kept and now explicitly conditioned on that.

Soundness is unchanged: an undecodable summary still licenses no ascent
and never enters the return-flowing set, so the ascent path is
byte-identical. Only the note's wording moves.

Tests drive all three undecodable forms through the mock and assert the
false claim is gone, the remedy is reported, and the ascent is still
withheld. A companion assertion pins that the all-decoded case KEEPS the
persisted-summaries claim, so the fix cannot degenerate into deleting the
sentence. Verified load-bearing: reverting the source alone fails 6 of 34.

Impact analysis: `calleesWithReturnFlow` upstream LOW (2 callers, both in
this file); `assemblePdgImpactResult` upstream LOW (1 caller).

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(pdg): cover every chained-receiver shape and pin the inference gap (#2802 review P2-1, P3-1)

The fixture proved chained receiver calls reach `BasicBlock.calleeIds`
using exactly one receiver form — a local `const`. That is the shape that
works, so a single-shape fixture implied general support the resolver does
not have. This repo has been burned by that before: a drop-count gate
blind to fixed shapes.

Measuring nine forms against the real pipeline also corrects how the gap
was originally characterised. It is NOT local-versus-field. An annotated
field resolves fine, including the constructor-assigned variant:

  private p: Outer = new Outer();          -> both links
  private p: Outer; this.p = new Outer();  -> both links
  private p = new Outer();                 -> EMPTY CELL
  private p; this.p = new Outer();         -> EMPTY CELL

The discriminator is the type ANNOTATION. When a field's type must be
inferred from its initializer the whole `calleeIds` cell empties — so even
`Outer.inner`, an ordinary named-receiver call, is lost, and the
inter-procedural descent cannot cross the boundary at all. Pre-existing;
independent of #2802, which does not touch receiver resolution.

The fixture is now table-driven over seven working forms (local const,
local in a method, annotated field, ctor-assigned annotated, ctor-param
assigned, call-result receiver, three-link chain) plus the two
inference-typed forms, each row carrying its expected chain-link ids.

Assertions moved from substring to exact id membership, split with the
production `splitCalleeIds` reader — so `Inner.compute` can no longer be
satisfied by `Inner.computeExtra` or `OtherInner.compute`, which matters
because the descent keys on exact ids for span and CALL_SUMMARY lookup.

The two known-gap rows are pinned with `it.fails` plus a hard assertion on
the exact gap-row set, so a resolver fix turns them red instead of passing
silently, and an anti-vacuity guard requires every shape to match exactly
one block — without it a drifted fixture matching zero blocks would let
`it.fails` pass for the wrong reason. Proven by mutation: relabelling a
working row as a known gap fails both pins.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(mcp): qualify the empty-ascent note when the examined callee set is incomplete (#2802 review P2-4)

The note asserted "none of the N resolved callees carry a CALL_SUMMARY
return-flow", and on the all-decoded path that this is "a property of the
persisted summaries". Both are universal claims over the callees the
descent actually examined, and two mechanisms can leave that set
incomplete without the note saying so:

  1. Budget truncation. The descent stops on depth/limit/node-cap, so a
     callee that DOES carry a return-flow can sit in a hop never reached.
     A 4-deep chain reported "none of the 3 resolved callees" while link 4
     held the only summary.
  2. Emit-time capping. When a block's `calleeIds` cell was capped,
     `splitCalleeIds` strips CALLEES_TRUNCATED_SENTINEL, so the dropped
     callees are invisible to both the scan and the counters — even though
     the callgraph bridge in this same file already treats such a block as
     callee-incomplete.

Add `calleeIdsWereTruncated`, the counterpart to the sentinel strip, read
from the raw cell before splitting so a block whose entire list was capped
away still raises the flag. Thread it through the descent to the note.
Case 1 needs no new plumbing — the aggregate `truncated` is already on the
input object.

Using the aggregate rather than a descent-only flag is deliberate: seed
truncation and intra-BFS depth truncation also shrink the initial slice, so
their callees are never gathered either. It is a sound superset that never
under-hedges.

When either mechanism fired, one clause naming the reasons is appended and
the whole-slice assertion softens to "every summary examined decoded … a
property of those summaries". When the set is complete both branches stay
byte-identical to before, so this does not become a blanket hedge.

Tests pin truncated, untruncated, emit-capped-alone, both-mechanisms, and
undecodable+truncated, asserting the truncation premise rather than
assuming it. Verified load-bearing: reverting the source alone fails 6 of
42, and the HEAD note printed in those failures is the bug verbatim.

Impact analysis: `assemblePdgImpactResult`, `calleeIdsByBlock`,
`interproceduralDescent` all upstream LOW; every caller is in this file and
`runImpactPDG`'s exported signature is unchanged.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(mcp): stop the empty-ascent note calling call-site references "resolved callees" (#2802 review P3-7)

The note printed "none of the N resolved callees carry a CALL_SUMMARY
return-flow (no formal parameter is recorded as flowing to its return
value)". N counted the raw `BasicBlock.calleeIds` cell, which carries ids
`resolveCalleeSpans` never enters — out-of-repo targets, interface
methods, and the `Class:` id a `new X()` emits. On the chained-receiver
fixture that inflated N from 1 to 3.

Two defects, both in the wording rather than the arithmetic: "resolved"
implies a symbol-table lookup that did not happen for those ids, and the
parenthetical asserted a FORMALS-level property about symbols never
resolved to a body.

Reworded rather than re-seeded, deliberately. `calleesWithReturnFlow`
scans the RAW id set, so the claim "none of these carries a return-flow"
is exactly established for all N — the scan really did check the `Class:`
id. Re-seeding N from the resolved spans would make the sentence quantify
over a strict SUBSET of what was checked, silently dropping the
un-enterable references from a claim that genuinely covers them, and would
desync N from `calleesUndecodable`, which is derived from the same scan
population.

  none of the N resolved callees carry ...
  none of the N call-site callee references carry ...

and the formals parenthetical is dropped. The note gets shorter, not
longer. `calleesResolved` is renamed `calleeReferences` end-to-end
(file-local; nothing outside referenced it), and the descent's return-type
doc — which called them "callee symbols the descent resolved" and
reinforced the wrong reading — now states that un-enterable ids ride the
same cell, are scanned, and are never entered.

The `> 0` gate is unchanged, so no slice that previously produced the note
stops producing one. A test pins that explicitly: an all-un-enterable cell
resolves no span, takes no hop, and emits no ascent sentence despite a
non-zero count — so a future re-seeding cannot silently move when the note
fires.

Tests also pin the quoted number and singular/plural against a mixed cell,
with a discriminator asserting `reachableBlocks` is byte-identical while
the count moves 1 -> 3. Verified load-bearing: reverting the source alone
fails 6 of 7 new tests, printing the finding verbatim.

Impact analysis: `assemblePdgImpactResult` and `interproceduralDescent`
upstream LOW, sole caller `runImpactPDG` in the same file; exported
signature unchanged.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(mcp): pin cross-hop callee accumulation and the mixed return-flow contract (#2802 review P2-5)

Every case in this file drove a single hop, so the Set union the descent
performs across hops (`calleeReferencesSeen` / `calleesReturnFlowingSeen`)
was never proven to accumulate rather than overwrite — a one-hop descent
cannot tell the two apart. And although a sibling commit added a
three-id cell, none of those ids return-flowed, so the
"some callees flow, some do not" boundary was entirely unpinned.

Extends the mock with a `secondSummary` knob that drives a genuine second
hop: `helper2` is named only in `helper`'s own body block, so the descent
must cross a second boundary to reach it. Three mock handlers are made
faithful to the parameters they already bind — `calleeIdsByBlock` now
routes on the asked `$ids`, and the CALL_SUMMARY scan and span resolve
answer per asked id — which is what makes a second callee answerable at
all. Existing cases are behavior-identical.

Five tests: the union count across two hops; a return-flow on hop 0
surviving a later empty hop; a return-flow found only on hop 1; mixed
callees in one examined set going silent rather than partial; and a
flowing callee alongside an undecodable sibling staying silent including
the decode remedy.

The mixed case pins a deliberate contract rather than proposing one. The
production condition is `calleesReturnFlowing === 0`, so partial coverage
is reported as silence. A reviewer considered and dropped "report partial
coverage" as a product change; this makes flipping it a conscious edit
instead of an accident.

Verified load-bearing against three separate source mutations: accumulating
only on hop 0 (2 fail), each hop overwriting instead of unioning (3 fail),
and flipping the gate to partial-coverage reporting (4 fail). In all three
every PRE-EXISTING test still passed — which is the finding restated as
evidence.

Test-only; `pdg-impact.ts` is byte-identical to HEAD.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(mcp): consolidate the empty-ascent rationale to one canonical site (#2802 review P3-6)

The "keyed on observed CALL_SUMMARY data, never on the criterion's
language" rationale was restated in full at four comment sites. It exists
because a reviewer asked "why not just look up the language?", so it has to
stay findable — but not four times.

The canonical explanation now lives in `interproceduralDescent`'s
return-type doc, where the counters are actually computed, organised as
POPULATION (why the raw `calleeIds` tally is the right set to quantify
over) and OBSERVED DATA, NEVER THE CRITERION'S LANGUAGE (the full
answer, including the producer-change argument and the no-language-naming
rule). The other three sites keep only what is locally load-bearing and
point here.

Deliberately preserved, because each carries a non-obvious fact: why an
undecodable summary licenses no ascent, why the aggregate `truncated` is
used rather than a descent-only flag, and the raw-id-tally population
argument. Net comment delta -11 lines.

The reviewer also flagged the local/field naming asymmetry
(`calleeReferencesSeen` vs `calleeReferences`). Keeping the suffix, with a
comment recording why so it is not re-raised: the premise that every other
local matches its field is true, but those locals are identity-returned,
whereas these are `Set<string>` accumulators returned as `.size`. Dropping
the suffix would give one identifier two types in one file — a `Set` at the
accumulation site and a `number` where the note does arithmetic and
pluralisation on it ~900 lines away. The Set-ness is also load-bearing: the
dedup is why a callee invoked from two hops is not double-counted, which
is what makes the note's count correct.

Comment-only. Verified mechanically: every added and removed line in
`git diff -U0` matches a comment pattern, so the note's template literals
are untouched and its rendered text is byte-identical. 89 tests unchanged.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(mcp): collapse the ascent plumbing accreted across 13 fix commits

Quality cleanup, no behavior change. Four independent review passes
converged on the same root cause: thirteen commits each fixed one review
finding in isolation, and the ascent facts grew one loose field at a time
until 62% of the changed region was comments explaining plumbing.

Five changes:

  - `calleeIdsFromBlocks` deleted. Zero call sites anywhere in src/ or
    test/ — already dead on main, and this branch had edited it to keep it
    compiling. Its only reference was a stale `{@link}` in a neighbour's
    doc, now rewritten to stand alone.

  - `parseCalleeIdsCell` replaces the two-pass read. `calleeIdsWereTruncated`
    and `splitCalleeIds` were splitting the same cell on adjacent lines,
    which measured ~2x the parse cost (0.82 -> 1.59 ms at a realistic hop,
    57.7 -> 92.7 ms at the per-statement site cap) and was a second
    independent encoding of the sentinel format — exactly what
    `splitCalleeIds` was extracted to prevent. One pass classifies as it
    walks; `splitCalleeIds` stays as a wrapper so its two external callers
    are untouched. The single-use `export` is gone.

  - `AscentCoverage` replaces four fields threaded through three
    signatures. ~12 declaration sites become 3, and the canonical rationale
    now lives on the type by construction — which is why the earlier
    doc-consolidation commit was needed at all.

  - `calleesReturnFlowing` becomes a boolean. Its only reads were
    `=== 0`, twice; it cost a Set sized to every callee in the slice plus a
    per-hop union loop. The flag is set inside the existing
    `returnFlowing.size > 0` branch — equivalent, since the cross-hop union
    is non-empty iff some hop's was.

  - The duplicated empty-ascent note head is collapsed to one gate and one
    head with per-arm tails. Both arms had been edited in lockstep twice in
    this branch's own history.

The rendered note text is byte-identical. Verified structurally and then
empirically: both expressions reconstructed standalone and diffed across
the full cross product of references x returnFlowing x undecodable x
truncated x listTruncated — 288 combinations, 0 mismatches.

Net -53 lines. 102 tests pass unedited; the unused-symbol lint warning is
gone.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(mcp): parallelise the startup probes, drop a redundant pin, name the mock knobs

Quality cleanup from the same review passes. The set of verified behaviors
is unchanged except where noted.

**Startup probes run concurrently.** `spawnSync` blocks the event loop and
vitest runs a file's tests in order, so the three probes strictly
serialised. Launching all three with async `spawn` in `beforeAll` and
asserting over the collected outcomes cuts the file from ~12.7 s to ~3.9 s
wall (-69%). Every promise is caught before `Promise.all`, so all three
children are reaped and failures report per entry rather than surfacing
only the first rejection. Preserved and each proven by mutation: the
missing-dist error names its entry, a raised module floor fails only its
own row, and a bogus anchor still reports the loaded-module count.

**The two `it.fails` rows are removed.** They pinned the inference-typed
receiver gap that the strict `toEqual` pin beside them already covers —
and they were the weaker of the two, because `it.fails` passes when the
body throws for ANY reason, including `idsFor`'s own non-vacuity guard. A
renamed fixture marker would have kept them green on a rotted premise. The
strict pin is self-diffing and was verified load-bearing on its own:
pointing a known-gap marker at a resolving shape fails it with the two
newly-present ids listed. The file header now carries the gap's durable
description.

**The ascent-note mock takes options objects.** `descentExec` and `run`
had grown to five and seven positional parameters in the order five agents
added them, so call sites read `run(FILE, true, null, 3, false, undefined,
null)` — several carrying `undefined` purely to reach a later argument. All
34 call sites are converted; nine that used only defaults are now bare
`run(file)`. No knob renamed — they are orthogonal and correctly named.
Code lines are exactly neutral (353 -> 353); the win is at the call sites.

Also refreshes five comments that still described `calleesReturnFlowingSeen`
and the two-branch note, both of which the preceding commit replaced.

102 unit and 10 integration tests pass; test count moves 9 -> 7 in the
chained-receiver file, exactly the two redundant rows.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(mcp): publish return-value-ascent coverage on the PDG impact result

`impact(mode:'pdg')` computed four facts about ascent coverage and used
them exactly once — to interpolate an English sentence. They never reached
the result object, so an agent consuming this MCP output could only ask
"was the ascent complete, and if not why" by regexing prose. The cost was
already demonstrated: a pure rewording commit earlier in this branch broke
~30 assertions and would have silently broken any consumer keying on the
old phrase.

Adds `pdgEvidence.ascent`:

    referencesScanned        how many call-site callee references were scanned
    returnFlowFound          did the ascent fire anywhere in this slice
    undecodableSummaryCount  summaries the codec could not decode
    examinedComplete         was the examined set the whole callee list
    incompleteReasons        'traversal-truncated' | 'callee-list-capped'
    callSummaryLayerPresent  false => pre-FU-C (v3) index

Nested under `pdgEvidence` because that is the established counts-and-
classification namespace, and `composeUnifiedPdgImpactResult` already
spreads it, so the member survives the unified compose untouched.

`incompleteReasons` carries CODES, following the existing
`truncatedByReasons: ('depth'|'limit')[]` precedent. The prose clause and
the structured field now render from one array computed once, so an agent
branching on codes and a human reading the note cannot disagree, and a
third reason becomes a rendering decision rather than a contract change.

Two shape decisions worth recording. `callSummaryLayerPresent` exists
because without it a v3 index publishes `referencesScanned: N,
returnFlowFound: false`, which reads as "these callees record no
return-flow" when the truth is "the layer that records it is absent" — the
note already distinguishes those, and the structured surface must not be
less honest than the prose. And the field is ABSENT rather than zeroed when
the descent never ran (upstream slices): "nothing was scanned" is a
different fact from "we scanned and found nothing".

`pdgResultVersion` stays 2. The documented trigger is a BREAKING change to
the result shape; this removes nothing, renames nothing, and changes no
existing field's meaning. Confirmed mechanically: zero top-level key drift
across 2304 cases. The historical v2 bump was for changing an existing
field's semantics (startLine 0- to 1-based).

The note prose is byte-identical, proven across the same 2304 cases with a
negative control — perturbing one character of the phrase table produces 60
drifts, so the harness demonstrably detects what it asserts. 14 new tests
cover the structured surface and all 14 fail when the source is reverted,
while the 54 prose tests pass unchanged.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(helpers): share one module-load probe, and fix two guards that passed on broken builds

Three tests independently spawned a child node process to inspect what a
built `dist/` entry loads, duplicating the REPO_ROOT derivation, the probe
source, the missing-dist guard, the spawn with NODE_OPTIONS cleared, the
status-vs-signal rendering, and the payload parse. The newest copy was also
the only correct one, so the next author had 2-in-3 odds of copying a
weaker probe.

The two older probes diff `require.cache` only, which is structurally blind
to the first-party ESM `dist/**` graph. That is not theoretical — both were
demonstrated passing on genuinely broken builds:

  - Severing `dist/cli/mcp.js -> stdio-context.js` (a pure ESM change)
    leaves the require.cache diff EMPTY, so `import-closure.test.ts`'s two
    assertions reduce to `[].filter(...) === []`. It reported 2 passed on a
    severed graph.
  - Severing `registry -> swift/query.js` leaves 76 unrelated CJS entries,
    which satisfied `registry-import-closure.test.ts`'s indirect guard. The
    Swift half of its headline had gone vacuous and it reported 1 passed.

Both now fail on those same builds, naming the missing anchor.

`test/helpers/module-load-probe.ts` unions the ESM `registerHooks({ load })`
channel with the cache diff, probes entries concurrently, and makes
non-vacuity STRUCTURAL: `anchor` and `minModules` are required fields and
the helper throws when either fails. A vacuous probe is a harness failure,
not a silently green test, so it cannot be forgotten. Forbidden patterns
and remedy text stay per-test — the harness is the shared part, the policy
is not.

Also fixes `toRepoRelativePosix` resolving non-absolute specifiers against
`process.cwd()`, and dedupes modules a CJS-from-ESM import reported once
per channel.

Faster despite doing more: the registry file goes 12.4s -> 6.75s, because
`spawnSync` burned the parent thread polling while the child loaded native
grammars. `import-closure` drops to one spawn from two.

The `local-backend.js` entry is kept although its closure is currently a
strict subset of `server.js`'s: that is an observation, not an invariant.
If `server.js` ever stops eagerly reaching the local backend, the server
probe stays green while the module #2802 actually changed goes unobserved —
and now that anchors are mandatory, that entry is what pins `pdg-impact.js`.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(lbug): trim the csv-generator note and fix the claim it got wrong

Two reviewers split on this comment: one wanted it cut to the structural
argument, the other said a comment is the right depth for documenting a
rejected change since there is no invariant to guard. Both are right, so
it stays a comment and gets shorter — 13 lines to 6.

Trimmed because it had already taken two corrections (an unreproducible
"~40x" figure, and a pointer to a test file that no longer exists), and its
tail had drifted from its own guard: the comment said "several hundred
modules, ~150 ms" where `startup-language-closure.test.ts` says "~226
extra modules and ~130 ms". Two numbers for one fact. That tail is
documented better in the guard's own header, so deleting it loses nothing.

It also stated the load-bearing claim inaccurately. The old text said
bm25-index imports `normalizeFtsText` "from here" — but `lbug-adapter.ts`
neither exports nor re-exports it; the only occurrence of the identifier in
this file WAS the comment. Anyone verifying would have grepped, found
nothing, and concluded the note was stale. Now names `csv-generator.js`
explicitly, re-verified at `bm25-index.ts:15` (static) and
`local-backend.ts:2756` (dynamic, on the FTS query path).

Comment-only, proven two ways: every changed line matches a comment
pattern, and stripping all `//` lines from HEAD and from the working tree
yields byte-identical text.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(helpers): extract the temp-repo lifecycle, collapsing five hand-rolled cleanups into one

Four cfg integration tests each hand-rolled a `tmpDirs` array, a
mkdtemp-and-register step, and an `afterAll` rmSync. It is actually five
registrations across six creation sites — `pipeline-pdg.test.ts` keeps a
second pool for its C-family fixtures.

Seeding genuinely varies four ways (recursive cpSync, single copyFileSync,
inline mkdir+writeFile, and nothing at all), so a fixture-copier helper
would have fitted about half the sites and made things worse. Extracted the
LIFECYCLE instead — mkdtemp, register, afterAll cleanup — which is
byte-identical at all five registrations and is the correctness-critical
part. `dir()` returns an empty registered directory for callers that seed
themselves; `fromFixture()` covers the common case. That fits 6/6.

The duplication had already produced a latent defect: `cFamilyTmpDirs` was
cleaned by TWO `afterAll` blocks, harmless only because `rmSync` was called
with `force: true`. Now one hook.

`createTempDirPool` is a function called from each test file's module scope
rather than a top-level hook in the helper, because under ESM caching a
module-level `afterAll` would register once, against whichever file
imported it first. That hazard is documented in the helper.

Raw line count is roughly neutral (-44 across the tests, +62 for the
helper, 29 of which are the rationale). The win is that a cleanup invariant
went from five copies to one.

Cleanup verified empirically, including the failure path: a throwaway suite
whose `beforeAll` throws still has its directory removed, and every
temp directory created by the four migrated files is gone after a run.
46 tests pass across the four files.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(resolvers): pin the inference-typed field receiver gap at the resolver level

The gap was pinned only in a PDG test, asserting on `BasicBlock.calleeIds`
behind the full `--pdg` pipeline. But it is a resolver fact: when a class
field's type must be inferred from its initializer, chained receiver calls
resolve to nothing. Whoever closes it will be working in the resolver
suite and would have got a red CFG/PDG test with no resolver-side signal.

Asserts CALLS edges directly, alongside `python-constructor-field-receiver.test.ts`.
Nine receiver shapes run the identical statement; seven resolve, two do not:

    const o = new Outer()                     resolves
    private p: Outer = new Outer()            resolves
    private p: Outer;  this.p = new Outer()   resolves
    private p: Outer;  this.p = p  (ctor arg) resolves
    constructor(private p: Outer) {}          resolves
    makeOuter().inner().compute()             resolves
    o.inner().mid().compute()  (three links)  resolves
    private p = new Outer()                   NO EDGES
    private p;  this.p = new Outer()          NO EDGES

Two things the fixture establishes that the PDG-side pin could not. The
discriminator is the type ANNOTATION, not local-versus-field — the
parameter-property form resolves fine. And the initializer is NOT invisible
to the resolver: `new Outer()` still emits its own constructor CALLS edge,
byte-identical to the annotated twin. Only the initializer-to-field-type
binding is missing, which narrows where a fix belongs.

Assertions key on exact node ids rather than names, because `compute` is
ambiguous across two classes and keying on the source name collides with
`Object.prototype.constructor`.

No `describe.skip` and no `it.fails` — the latter passes when the body
throws for ANY reason, so it can go green on a rotted premise. The gap is
pinned as its explicit current value, which self-diffs: simulating the fix
fails one test showing the two newly-resolved ids, and renaming a fixture
symbol fails the non-vacuity guard.

Runtime is comparable to the PDG-side pin (~9-11s, both dominated by
worker startup), so this is an altitude and scope win, not a speed one.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(mcp): replace the extension sweeps with a stronger language-agnosticism pin

Two `it.each` sweeps over nine file extensions asserted that the
empty-ascent caveat was present (or absent) for each. They looked like the
pin for the property the whole change exists for — `pdg-impact.ts` must
name no language and its output must not vary by extension — but they were
the weakest available form of it.

They asserted substring presence/absence, so a language dependence that
ADDS text while leaving the caveat intact passes them. Demonstrated, not
assumed: injecting a `.py`-only hedge inside the caveat sentence and
replaying the two sweeps verbatim against that source gives 18 passed. The
byte-identity test beside them caught it.

So the sweeps are deleted and the identity test carries the property alone,
hardened in two ways:

  - Two rows instead of one, covering BOTH sides of the caveat gate. The
    silent (return-flow present) branch previously had no identity
    counterpart at all — nine runs proving one fact, with nothing checking
    that its rendering was extension-invariant.
  - The fingerprint spans the note AND the reachable blocks, not just the
    note. Strictly more than the sweeps verified.

Entailment is exact: identity across the extension set, plus the two
existing single-extension content assertions, gives "every extension gets
the caveat" and "no extension gets it". Reducing a sweep to one extension
was rejected because it reproduces an assertion already present verbatim.

Also converts the incompleteness block from six near-identical bodies to a
3-row premise table crossed with two assertions. Each row now names the
exact phrase set its clause must contain, so presence and absence are
asserted together — which adds three checks the longhand version lacked
(the budget row now also proves the emit-cap phrase is absent). And three
tests that re-rendered one fixture to make one assertion each are hoisted
to a single render.

97 tests, down from 116: -18 sweep cases, -2 from the hoist, +1 identity
row. No assertion was lost; several were added.

Verified by injection: a `.py`-only note change fails the identity pin,
and a dependence in the shared hop sentence fails BOTH rows, confirming the
second row is load-bearing rather than decorative.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* perf(mcp): lazy-import syncGroup so MCP startup skips the group extractor closure

`core/group/service.ts` statically imported `./sync.js`, which pulls all six
contract extractors, five of which statically import the native `tree-sitter`
binding. That put the whole parser stack on every MCP server start, for a
server that never syncs.

Only `groupSync` needs it. The other seven group tools — `group_list`,
`group_impact`, `group_query`, `group_contracts`, `group_status`,
`group_trace`, `group_context` — do not, and now never load it. `syncGroup`
has a single call site, already inside an `async` method, so this is a lazy
`await import(...)` at that call site and nothing else: no signature change,
no async ripple, no change to `local-backend.ts`.

The pattern is already established on this exact module — `cli/group.ts`'s
sync command lazy-imports `sync.js` the same way. `service.ts` was the
outlier.

Measured on a native filesystem (overlayfs; /workspace is a 9p mount that
inflates ESM resolve, so it is not a valid measurement surface), 5 cold runs,
medians:

  dist/mcp/server.js              521 ms -> 133 ms   (-75%)
  dist/mcp/local/local-backend.js 453 ms -> 66 ms    (-85%)
  tree-sitter modules at both entries: 11 -> 0

Same defect class as #2802, which cut the language-provider registry from the
same startup path; this is what remained.

The cost is moved rather than deleted: the first `group_sync` call now pays
the module load. That is the right trade — `group_sync` is already a
long-running operation, and sessions that never sync pay nothing.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(mcp): guard MCP startup against the group extractor closure returning

Sibling forbidden-pattern case in the #2802 startup guard, reusing the
concurrent probes it already collects — no new spawn, no new harness.

Asserts that none of `dist/mcp/server.js`, `dist/cli/mcp.js`, or
`dist/mcp/local/local-backend.js` loads a `core/group/extractors/` module or
the native `tree-sitter` package. The parser is matched by package prefix
rather than a bare substring, so a source file that merely mentions the word
can neither satisfy nor trip it.

Verified load-bearing rather than assumed: restoring the static
`import { syncGroup }` in `core/group/service.ts` and rebuilding turns
`dist/mcp/server.js` red and names all seven offenders —
http-route, grpc, thrift, topic, include, manifest and workspace extractors.
Reverted and re-confirmed green.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* perf(mcp): keep the analyze-only CFG closure off MCP server startup (#2802 review)

`mcp/local/pdg-impact.ts` imported `CALLEES_TRUNCATED_SENTINEL` and
`CALLEE_ID_SEP` from `core/ingestion/cfg/emit.ts`. ESM evaluates a module to
import any binding from it, so those two strings dragged the whole analyze-only
CFG closure into every MCP server start.

Measured against a clean build, per entry point: 8 modules — `emit`,
`reaching-defs`, `reaching-defs-graph`, `control-dependence`, `post-dominators`,
`synthetic-escape`, `call-site-harvest`, `reaching-def-reason-codec` — present at
`dist/mcp/server.js`, `dist/mcp/local/local-backend.js` and
`dist/mcp/http-transport.js`.

Same defect class as the language-provider closure this branch already removed,
and the guard could not see it: `FORBIDDEN_RE` covers `core/ingestion/languages/`
and `FORBIDDEN_GROUP_RE` covers `core/group/extractors/|node_modules/tree-sitter`,
neither of which matches `core/ingestion/cfg/`.

The format constants move to a new LEAF module `cfg/callee-cell-format.ts` that
imports nothing; `emit.ts` re-exports both names so every existing importer is
untouched, and producer and consumer still resolve to one definition — the drift
the shared constant exists to prevent stays impossible.

Deleted, not deferred — the same bar #2802 held its own csv-generator proposal
to. After: cfg modules at startup 8 -> 2, and both survivors
(`callee-cell-format`, `reaching-def-reason-codec`) are leaves that import
nothing. Totals: `server.js` 387 -> 380, `local-backend.js` 163 -> 156,
`http-transport.js` 523 -> 516.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(mcp): stop pdgEvidence.ascent claiming a completeness it cannot have (#2802 review)

`examinedComplete` is the field a consumer reads to decide whether
`returnFlowFound: false` is a whole-slice claim. It could be published `true`
over a callee set the descent never finished examining — the exact false
all-clear the field was added to prevent.

Root cause: `bfsReachableBlocks` sets `truncatedByDepth` when its frontier is
still non-empty at the budget, but both call sites inside `interproceduralDescent`
folded only the row-limit flag and dropped the depth flag. The top-level intra
BFS's copy of that same flag was already propagated, so the asymmetry was
unintended — one `if`-pair folding limit-but-not-depth, within a merge that
already folds the node cap too.

Reproduced at `maxDepth: 3`, the shipped default: a criterion calling a helper
whose body is a 5-block dependence chain, with the return-flowing callee on the
block past the clamp. Result reported `truncated: undefined`,
`examinedComplete: true`, `incompleteReasons: []` and an unqualified universal
note sentence.

Fixed by propagating the dropped flags rather than inventing a parallel channel:
`intraDepthBudget` is documented in-file as the SAME clamp the top-level intra
BFS applies, and that one's depth truncation is already result-level. So the
result's own `truncated`/`truncatedBy` were under-reporting for the same reason,
and both surfaces are corrected together.

Four further honesty fixes to the same published record:

- Blocks reached only by the U-C4 ascent went into `reachable` but never
  `hopReached`, so their `calleeIds` cells were never scanned, never counted, and
  could not raise `callee-list-capped`. They are slice blocks; they now enter the
  hop set and get the same treatment as every other one.
- `pdgEvidence.ascent` was absent on the empty-slice early return even though the
  descent had already run and scanned, contradicting the "present iff the descent
  ran" contract this branch itself added to `tools.ts`. Both exits now classify
  through one shared helper so they cannot disagree.
- A block carrying call sites in `callees` but no resolved ids in `calleeIds`
  (the whole-file case where `emit.ts` has no fileMap) silently shrank the
  population while `examinedComplete` still reported `true`. That now raises a
  third reason, `callee-ids-unrecorded`.
- `referencesScanned` is a distinct-callee tally and both surfaces described it as
  a call-site count. Field name kept — a rename is breaking at
  `pdgResultVersion: 2` — and the prose corrected instead.

`PdgAscentIncompleteReason` gains a member, which is additive, so
`pdgResultVersion` stays 2. Visible output change worth knowing: slices whose
callee chain outruns `maxDepth` now report `truncatedBy: 'depth'` where they
previously reported none, and a repo with id-less call sites now reports
`examinedComplete: false`. Both are strictly more honest.

Every behavioural change carries a mutation proof — revert the source, watch the
new test go red, restore. One exception is documented inline rather than faked:
the ascent-side fold cannot be observed independently, because the re-seed shares
the caller's `visited` set and so can only reach past the budget when the
traversal that covered that closure was already cut and had already raised a flag.

Suite: 49 -> 59 tests.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(mcp): anchor each import-closure policy on the edge it polices (#2802 review)

`module-load-probe.ts` makes non-vacuity structural via a required `anchor` — but
the anchor was one per ENTRY while `startup-language-closure.test.ts` now runs TWO
independent policies. The group-extractor policy added in 83e8cf7c5 therefore had
no anchor of its own, and one of its three rows was already vacuous: `cli/mcp.js`
loads four leaf modules and reaches no `core/group/` module at all, so its group
assertion could not fail for any policy-related reason while its
`dist/mcp/stdio-context.js` anchor stayed green.

Proven, not argued. `dist/mcp/local/local-backend.js` is the only static importer
of `core/group/service.js` in the whole build; severing that one edge — the exact
next lazy-load step — and re-probing:

  OLD shape (anchor per entry):  server 385, http-transport 521, local-backend 161
                                 reaches group/service = false, group offenders 0
                                 -> GREEN on all three
  NEW shape (anchor per policy): -> RED on all three, each naming the missing
                                    dist/core/group/service.js

Counts fell only 387->385 and 163->161, so `minModules` was structurally blind to
the severance; the anchor is the only thing that catches it.

`anchor` accepts `string | readonly string[]` and every listed anchor must load.
Existing single-anchor call sites are unchanged. `anchorsOf()` lets the group
`it.each` DERIVE its entries by filtering on the group anchor, with a test pinning
that derivation, so the policy cannot silently register zero cases. `cli/mcp.js`
is dropped from the group policy — it cannot honestly carry that anchor — and the
doc-comment now states the invariant: an anchor is per-POLICY, not per-entry.

Also:

- `mcp/http-transport.js` gets a row. It is the largest startup entry (516
  modules) and `src/cli/mcp.ts` imports it directly rather than through
  `server.js`, so nothing about the server row constrained it. Measured clean
  today; the gap was coverage, not a broken claim.
- The three spawn-based closure tests are registered in `SPAWN_CLI`, so the
  Windows-safety plumbing this branch wrote for them (POSIX normalisation,
  `pathToFileURL`, `NODE_OPTIONS` clearing, array-form `spawn`) is finally
  exercised on the Windows/macOS matrix. Measured cost ~11.7s on Linux; budget
  ~60s on Windows against a 25-minute job.
- `PROBE_TARGET` now wins over `extraEnv`, which was spread last and could have
  silently redirected a probe while `anchor`/`minModules` stayed keyed on `entry`.
- The child's JSON payload is validated through a type predicate instead of a bare
  `as string[]`, and the spawn timeout escalates SIGTERM to SIGKILL so a child
  stalled in native code is reaped rather than orphaned.
- Recorded baselines re-measured (server 380, local-backend 156, cli/mcp 4) and
  relabelled a snapshot rather than a contract — they moved twice inside this
  branch alone. The subset claim was re-verified exactly: 0 of local-backend's 156
  modules are absent from server's 380.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(helpers): survive a failing temp-dir removal instead of leaking the rest (#2802 review)

`createTempDirPool`'s `afterAll` ran a bare `for (const d of created) fs.rmSync(d, {recursive, force})`.
`force` suppresses only `ENOENT` — not the `EBUSY`/`EPERM`/`ENOTEMPTY` class a
Windows runner produces when a pipeline test still holds a handle — so the FIRST
failure threw out of the loop and leaked every directory registered after it.

Pre-existing: all four hand-rolled cleanups this helper consolidated had the same
shape. But the blast radius is now shared across four consumers, which is exactly
why it is worth fixing at the point of consolidation.

Cleanup is now per-directory best-effort via `removeTempDirs`, plus Node's own
documented mitigation for that error class (`maxRetries: 3, retryDelay: 50`),
which costs nothing on the happy path.

Warn rather than swallow or rethrow, and the reasoning is in the doc comment, not
just here: rethrowing would fail an otherwise green suite from `afterAll` over
housekeeping the OS reclaims anyway, where it reads as a test failure and buries
the real result — a Windows EBUSY on a temp dir is not a defect in the code under
test. Silence is the opposite hazard: a systematic leak would be invisible with
nothing naming the responsible suite. The warning carries the path, and the
`mkdtemp` prefix is per-pool, so it names the suite that made it.

Failure is injected through a scripted remover keyed by path (a Map lookup, so no
`if` in a test body and no dependence on producing a real locked handle). Beyond
the three behavioural pins there is a wiring pin — a nested `describe` creates a
real pool and a sibling `it` declared after it asserts the dirs are gone — so the
tested function cannot drift into "tested helper plus an untested copy of the
loop".

Mutation proof: restoring the abort-on-first-failure loop turns 3 of the 5 tests
red, the throw escaping `removeTempDirs` outright so the third real directory is
never attempted. With the fix, `[first, blocked, last].map(existsSync)` is
`[false, true, false]` — the injected failure survives and the directory after it
is really gone, through the remover that actually ships.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(pdg): point the self-diffing receiver pins at #2807, not at this PR (#2802 review)

Both pins named the gap "(#2802 follow-up)". The gap has its own tracking issue —
#2807, "Inference-typed field receivers resolve to no CALLS edges at all" (open,
labeled bug) — and PR #2810 is already open against it. As written, after merge
the gap was discoverable only by reading a KNOWN GAP marker inside a test file,
not from the issue tracker.

Both describe names now read "(known gap: #2807)" and both KNOWN GAP test names
carry the number. #2802 is kept only as provenance: the gap was FOUND during
#2802 work but is pre-existing and independent of it.

Each header gains an explicit "this pin is self-diffing: it will go red on
purpose" section naming #2807 with its exact title, noting #2810 is open against
it at the time of writing, and stating that the pin asserts the gap EXISTS — so
closing #2807 fails it by design, and the correct response is to update the
expected value, not to relax the assertion. The same note is repeated inline
above each KNOWN GAP test, where a maintainer editing it will actually see it.

No pin is weakened. Both deliberately reject `it.fails` in favour of exact
`toEqual` assertions with a non-vacuity probe, and that design is left untouched.

Refs #2802, #2807

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(group): cover the lazy syncGroup import that no test reached (#2802 review)

9ea9676dc turned `GroupService.groupSync`'s `syncGroup` into
`await import('./sync.js')` — this branch's one changed control-flow line in
production code — and nothing exercised it. Every existing test stopped short:
`service.test.ts` returns at the empty-name guard; `group-service-not-found.test.ts`
mocks `loadGroupConfig` to reject and never invokes its `syncGroupMock`;
`group-sync.test.ts` imports `syncGroup` directly, bypassing `GroupService`; and
the startup guard asserts only the negative, that `sync.js` is absent at startup.
`tsc` catches a path typo, but nothing verified the import resolves and hands off
correctly — while every production `group_sync` call goes through that line.

No production change was needed; the reviewed design was sound. This is the
missing coverage.

The happy-path test mocks nothing: it points `GITNEXUS_HOME` at a pool temp dir,
seeds a real `group.yaml`, and calls `groupSync`, so `loadGroupConfig` resolves,
`groupDir` is found, and execution falls through into the REAL `syncGroup`. What
makes a real sync reachable with no indexed repo: an empty registry puts both
members in `missingRepos`, but one declared manifest link still yields
synthetic-UID contracts. It asserts the returned counts AND reads back the
`contracts.json` that real `syncGroup` wrote into `groupDir` via the production
`readContractRegistry`, which pins the option handoff too.

Two further tests use `vi.doMock` to re-evaluate the service against a `sync.js`
whose load throws: one asserts the call rejects with the load failure in its
`cause` chain — so the caller gets a catchable rejection, not a floating
unhandled one — and one asserts both pre-import guards still answer with
`sync.js` unloadable, which is also a structural pin that the module has no
STATIC import of it (a static one would throw at re-import, before any call).

Mutation proofs: pointing the specifier at `./sync-nope.js` turns 2 of 3 red
("Cannot find module .../sync-nope.js ... at GroupService.groupSync
service.ts:349"); aliasing a real-but-wrong export turns 1 red. Restored, all 3
green, and `service.ts` verified byte-identical to HEAD.

Out of scope, stated rather than glossed: the final `isError: true` MCP envelope
is produced above `GroupService` and needs a full `LocalBackend`; the rejection
test is the in-scope half of that claim.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(mcp): close the gaps a cleanup pass found in the #2802 review fixes

Quality pass over the review-response series (reuse / simplification /
efficiency / altitude). No behaviour change except where noted.

The two that mattered:

- **The cfg/emit fix had no guard.** `FORBIDDEN_RE` covers
  `core/ingestion/languages/` and `FORBIDDEN_GROUP_RE` covers
  `core/group/extractors/|tree-sitter`; neither matches `core/ingestion/cfg/`.
  Because `emit.ts` re-exports the constants, pointing `pdg-impact.ts` back at
  `cfg/emit.js` typechecks identically and silently restores all 7 modules.
  Verified: with the import reverted, `tsc --noEmit` still exits 0 and every
  test stayed green before this commit; after it, 3 rows go red naming the
  offenders. Written as an ALLOWLIST of genuine leaves rather than a denylist of
  the 7 already-suffered modules, because the next regression is a module nobody
  has thought of yet.

- **`FORBIDDEN_GROUP_RE`'s parser matcher was forward-slash only** while both
  sibling probe regexes spell the separator `[\\/]`. Native bindings arrive via
  the `require.cache` channel as absolute paths and `toRepoRelativePosix` only
  normalises paths inside the repo root, so a hoisted `node_modules` renders as
  `…\node_modules\tree-sitter\…` on Windows and matched nothing. The same series
  put this file on the Windows matrix, where that half of the assertion would
  have been vacuous.

Reuse — three re-implementations of existing helpers:

- `removeTempDirRecursive` re-rolled `fs.rmSync` retries; it now delegates to
  `cleanupTempDirSync` (`test-db.ts`), the repo's Windows-lock-aware remover.
  The copy had already drifted on both knobs that matter — 3 retries at 50 ms
  vs 5 at 100–400 ms, and warn-on-everything vs swallow-lock-codes-rethrow-rest
  — which is how one half of a suite goes green-with-a-warning on the same
  `EBUSY` the other half fails on. The per-directory try/warn loop, which is the
  actual fix, is unchanged.
- `errorChainText` re-rolled the cause-chain walk that `causeChain`
  (`src/lib/utils.ts`) exists to be the single copy of — its own doc asks
  callers not to.
- The SIGKILL escalation (a timer, an `unref`, and two `clearTimeout`s) is
  `spawn`'s own `killSignal` option, which Node's `timeout` already delivers.

Simplification and altitude:

- `'callee-ids-unrecorded'` documented ONE of its three producer paths. The
  unnamed common one is a call site that did not RESOLVE — exactly the
  receiver gaps this repo pins (#2807) — so on a real index the reason fires
  broadly, driven by resolution quality rather than a missing `--pdg` layer,
  and "re-run analyze --pdg" is the wrong remedy for it. Doc now names all
  three and states the consequence: `examinedComplete: true` is the strong,
  rare signal.
- The derived policy-entry list was re-pinned against a hand-written 3-element
  literal, reinstating one layer down the list the derivation removes. Now
  asserts the properties that are actually at risk — non-emptiness (a policy
  going silent) and `cli/mcp.js` staying excluded (a row that cannot fail).
- A test fixture spread `ascentBlockCell: 'idless'` and then overrode it to
  `'capped'` in both runs, so the id-less shape never reached the mock while
  reading as though it did.
- `idlessCallSites` is sticky, so its per-row string allocation now
  short-circuits once set.
- Dropped an unused `export` on `CleanupWarner`.

Refs #2802

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* revert(ci): unregister the module-load closure guards from the Windows matrix

Registering the three `dist/` closure guards in `SPAWN_CLI` turned the Windows
`platform-sensitive 1/3` shard red at the 20-minute watchdog. Baseline
83e8cf7c5 was green on all three shards; a4245119c (which added them) failed
1/3; d0b201442 failed the same way.

It is not the files themselves. On the Windows runner they are among the
cheapest in the suite — `registry-import-closure` 448 ms, `import-closure`
53 ms — and both passed. vitest shards this list by file COUNT, not runtime, so
adding three files RESHUFFLED the split: shard 1 went to 32 files against 26 and
29, concentrating the heavy CLI e2e suites. It timed out with `cli-e2e`,
`group/cross-trace-e2e`, `lbug-orphan-sidecar-recovery` and `server-http-startup`
still queued — `cli-e2e` being the ~50-spawn suite whose setup flakiness already
needed fixing once (PR #2000).

That clustering fragility is pre-existing and this file's own header documents
it (#2449: "the heaviest spawn suites can cluster on one shard", busiest Windows
shard already at 14m57s against the old watchdog). These three files only tipped
it over, and unblocking the PR beats holding it for a CI-infra fix that belongs
in its own change.

Reverted rather than worked around: raising the shard count would keep the
coverage but is a repo-wide CI change made on a 25-minute feedback loop with no
guarantee the reshuffle balances, and this PR is about MCP startup. The removed
entries are replaced by a comment recording WHY they are absent, what they were
measured to cost, and the precondition for re-landing them — so the gap is
documented at the point someone would otherwise re-add them blind.

Verified: the emitted file list is byte-identical to 83e8cf7c5's, so the shard
split returns to the configuration that was green.

The Windows-specific bug this series found is unaffected — `FORBIDDEN_GROUP_RE`
now spells its separator `[\\/]` like its siblings, which was a real
forward-slash-only vacuity, and that fix stays.

Refs #2802, #2449

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ci): shard the cross-platform matrix by measured weight, not file count

Restores the three `dist/` module-load closure guards to the Windows/macOS
matrix, and fixes the reason they could not stay there.

They must run on every OS — the shared probe in `test/helpers/module-load-probe.ts`
IS the platform-varying code (array-form `process.execPath` spawn, cleared
NODE_OPTIONS, `pathToFileURL` because Windows rejects a bare absolute path as an
ESM specifier, and a `path.sep`→POSIX normalisation the anchors and offender
regexes depend on). Ubuntu-only coverage of a platform guard is no coverage.

The earlier attempt turned Windows `platform-sensitive 1/3` red at the 20-minute
watchdog, and the reflex fix — unregistering them — treated the symptom. The
files are among the cheapest in the suite (measured 448 ms, 53 ms, sub-second,
and both that completed passed). The defect is that `run-cross-platform.ts`
handed vitest all 84 files plus `--shard=i/n`, and vitest partitions by file
COUNT. Runtimes here span three orders of magnitude, so a count-split is blind
to the thing that decides the budget, AND re-partitions on every insertion:
adding three free files reshuffled the list and happened to co-locate `cli-e2e`
(361 s) with `cli-limit-e2e` (75 s) and `analyze-heap-oom-e2e` (23 s) — 32 files
against 26 and 29 — which timed out with four still queued.

The split now happens in `scripts/cross-platform-shard.ts`, longest-processing-
time first over measured Windows runtimes, and only the chosen shard's files are
passed to vitest (`--shard` is consumed, never forwarded — forwarding would
re-partition the slice a second time and silently drop most of it).

Weights are measured, from the last green matrix run plus the timed files of the
failing one, and every file also carries an 8 s per-file floor. That floor is
calibrated, not guessed: the last green busiest shard ran 736 s of wall clock
over ~511 s of attributed file time. Without it the balancer isolates the two
monsters and then piles every light file onto the remaining shards — trading a
runtime imbalance for a count imbalance that costs the same.

Result at TOTAL=3, with the three guards back in: 521 s / 527 s / 519 s across
20 / 33 / 34 files. The previous green configuration's busiest shard was 736 s,
so this is better balanced than the state before any of this, and the busiest
shard is now bounded by construction rather than by sort-order luck.

`test/unit/cross-platform-shard.test.ts` pins the properties, and the
load-bearing one is not "the split is even" — it is "adding a cheap file cannot
move a heavy one", the property whose absence caused the outage. Two details in
that test are themselves load-bearing, and earlier drafts got both wrong and were
vacuous: the inserted names must sort EARLY (names sorting last disturb nothing
under any scheme) and the count must not be a multiple of the shard total
(adding exactly `total` files leaves an equal-weight round-robin in the same
rotation). Mutation-proved: replacing `weightOf` with a constant — i.e.
count-based sharding — turns that test and the per-file-floor test red; restored,
all 8 pass.

Refs #2802, #2449

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 21:26:13 +01:00
7468cc915b fix(analyze): replace the hand-incremented schema version with a derived DDL fingerprint (#2798) (#2808)
* feat(schema): derive a fingerprint from the DDL this build creates

`SCHEMA_FINGERPRINT` is a sha256 digest of the node and relation DDL that
`runSchemaCreationQueries` actually executes, in the same shape as the existing
`taintModelVersion` stamp (hex, sliced to 12).

It exists because `INCREMENTAL_SCHEMA_VERSION` is hand-picked and has to
*predict* whether an on-disk database matches this build's DDL. That number has
collided with `main` eight times, twice exactly — and an exact clash is the
quiet one, because the reuse gate is a strict `===`.

`EMBEDDING_SCHEMA` is deliberately excluded: its `FLOAT[N]` width comes from
`GITNEXUS_EMBEDDING_DIMS` at module load, so folding it in would make the
digest a function of the environment rather than of code, and two runs of the
same build under different env would thrash full rebuilds.

Refs #2798

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(storage): record the DDL fingerprint in RepoMeta

`RepoMeta.schemaFingerprint` stores the digest of the DDL an index's tables
were actually created from. It is the derived companion to `schemaVersion`,
not its replacement: both are compared, and both must match.

Absent means mismatch, deliberately. Grandfathering a missing fingerprint
would let an incremental top-up stamp a fresh one onto a database whose DDL
was never verified, permanently certifying exactly the wrong-shaped index the
field exists to catch. The cost is one full rebuild per pre-existing index.

The version ladder gains a note that its "re-check against origin/main before
merge" ritual now only guards *semantic* bumps. v25, v26, v30, v31 and v34 all
changed emitted ids, edges or wire formats while leaving the DDL byte-identical,
and the fingerprint cannot see any of them — but DDL collisions no longer need
renumbering.

Refs #2798

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(analyze): gate index reuse on the DDL fingerprint, not just the version (#2798)

`INCREMENTAL_SCHEMA_VERSION` is a hand-incremented integer that has to predict
a derived fact: whether the on-disk DDL matches the code's DDL. It has collided
with `main` eight times, and twice the collision was *exact*.

An exact clash is the silent one. Two builds stamp the same number over
different DDL, the `===` reuse gate reads the index as current, every
`CREATE ... TABLE` is then skipped as "already exists" (suppressed in
`runSchemaCreationQueries`), and the edges whose endpoint pair the live database
cannot hold are dropped by `fallbackRelationshipInserts`' bare `catch`. The
result is a wrong graph, with no error anywhere.

Reuse now requires the version AND the DDL fingerprint to match, in both the
pre-pipeline force-rebuild guard and the `isIncremental` predicate, and the
fingerprint is stamped alongside the version at the end of a run.

Both conditions are necessary. The fingerprint does not replace the integer:
most entries in the version ladder change emitted ids, edges or wire formats
while the DDL stays byte-identical, and a fingerprint-only gate would stop
forcing rebuilds for all of them. What it does buy is that two branches picking
the same number no longer need renumbering.

The new branch sits above the `alreadyUpToDate` fast path for the same reason
the version guard does — a clean tree at an unchanged commit would otherwise
early-return before either check ran.

Closes #2798

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(analyze): pin the DDL fingerprint gate and its two failure cases

`schema-fingerprint.test.ts` pins the properties the gate rests on: the digest
covers exactly the node and relation DDL that gets executed (recomputed from
the exported lists, so adding a table or a FROM/TO pair without the fingerprint
moving is impossible), it excludes the environment-derived embedding DDL, and
it moves when any covered string moves.

The two `incremental-orchestration` cases exercise the production path rather
than modelling it: an index carrying the *current* version with a foreign
fingerprint, and one with no fingerprint at all. Both were run against the
pre-fix tree first and both failed there with `alreadyUpToDate === true` —
the fast path swallowing the mismatch, which is the #2798 symptom exactly.

`call-summary-schema-version.test.ts` widens its gate model to two equalities.
The second argument defaults to the current fingerprint so all 33 existing
version cases read unchanged, and a new case covers the collision, the legacy
absence, and the semantic bump the fingerprint cannot see.

Refs #2798

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(review-skill): point the schema-constant check at the fingerprint, not the deleted integer

All four `gitnexus-review` SKILL.md mirrors told reviewers to verify
`INCREMENTAL_SCHEMA_VERSION` "was bumped or regenerated". That constant no longer
exists, so the instruction sent every future reviewer looking for something they
could not find — and, worse, past its replacement.

The check for graph DDL is now derived: `SCHEMA_FINGERPRINT` moves on its own, so
the question is whether the diff changed a string in `NODE_SCHEMA_QUERIES` /
`REL_SCHEMA_QUERIES`, and whether a newly added DDL array was folded into the
fingerprint at all — the one way the derived gate can still be bypassed.

What did NOT change is called out explicitly: the parse-store `SCHEMA_BUMP` and
the bench fingerprint sets are still hand-maintained and still need the
re-check-against-base ritual, and semantic changes that leave the DDL untouched
fall outside the fingerprint entirely — those rely on the analyzer runner-identity
receipt.

Found by the review swarm's docs lane. The original plan for #2798 claimed no
documentation mentioned the constant; that sweep covered five root docs and never
looked at `.claude/skills/**` or the three mirrors.

Refs #2798

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(migration): record the one-time rebuild the fingerprint switch costs

Replacing `schemaVersion` with `schemaFingerprint` means every index written by
an earlier GitNexus carries no fingerprint, reads as a mismatch, and is rebuilt
once. That is deliberate — grandfathering absence would stamp a fresh fingerprint
onto a database whose DDL was never verified — but until now it was undocumented,
so a user's first post-upgrade analyze would announce a full re-analyze with
nothing to explain it.

MIGRATION.md already sets the precedent: PR #2363's meta.json → gitnexus.json
rename was equally automatic and equally in need of an entry. This follows that
shape, and is explicit about the parts that are easy to undersell:

- the cost is per INDEX, and branch-scoped slots (#2106) each pay separately;
  on a large repository a full re-analyze is substantial, not a blip;
- rollback is safe — an older binary sees no `schemaVersion` and forces its own
  rebuild, which is a cost, never a stale graph;
- alternating between an old and a new binary rebuilds on every switch, because
  the end-of-run meta is written as a fresh literal so neither field survives the
  other's run.

The retired ladder's per-version rationale is pointed at in git history rather
than reproduced: `git show 561f913a3:.../repo-manager.ts`. That commit is an
ancestor of origin/main, so the pointer survives this branch being squash-merged.

Refs #2798

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(identity): cover workspace-linked packages in the analyzer dependency digest

`dependencyNames` enumerated `dependencies`, `optionalDependencies` and
`peerDependencies` only. `gitnexus-shared` is declared as a devDependency
(`file:../gitnexus-shared`), and in a source-mode run the build root is the
gitnexus package tree, which does not contain that sibling. So a change to
gitnexus-shared moved neither `build.digest` nor `dependencyRuntime.digest`.

That gap matters more since #2798 deleted `INCREMENTAL_SCHEMA_VERSION`. A
DDL-affecting edit there is still caught by `SCHEMA_FINGERPRINT`, but a
SEMANTIC-only edit — a new `REL_TYPES` member, say, where the relation table
carries a bare `type STRING` column so no CREATE statement moves — was covered by
nothing at all. Roughly thirty of the retired ladder's entries were exactly that
change class, and the runner-identity receipt is what now carries them.

Only checkout-local specifiers are added: `file:`, `link:`, `workspace:`,
`portal:` and npm's bare local-path shorthands. Pulling in every devDependency
was rejected — vitest, eslint and typescript would enter the digest and force a
full re-analyze on unrelated tool bumps, which is worse than the hole.

Scanning the linked sibling for the first time exposed a latent throw:
`collectArtifacts` honoured `PRUNED_RUNTIME_DIRECTORIES` only for a real
directory, so a SYMLINKED `node_modules` fell through to the payload branch and
died with "Analyzer identity input is not a file". Worktree-style dev layouts and
pnpm shared stores hit this immediately — verified in this worktree, where
`gitnexus-shared/node_modules` is such a symlink. Pruning it loses nothing:
packages beneath are still reached through `resolveDependencyPackageRoot`.

Verified: a real `analyze` in this worktree succeeds with `packageCount` 259;
editing the linked package's source moves the digest, bumping an installed
registry devDependency does not, and removing the link moves it.
`DEPENDENCY_RUNTIME_CANONICALIZATION` is deliberately not bumped — freshness
compares digests, not the label, and the input-set change already moves them.

Follow-up worth having: no fixture in the suite declares `devDependencies`, so
this has no regression test yet.

Refs #2798

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(analyze)!: delete INCREMENTAL_SCHEMA_VERSION, gate reuse on the DDL fingerprint alone

The integer and its ~180-line version ladder are gone, along with
`RepoMeta.schemaVersion`. Index reuse is now decided solely by
`SCHEMA_FINGERPRINT`; a mismatch — including the absent stamp every pre-existing
index carries — warns and forces a full re-analyze, which wipes and recreates the
database so the tables are built from the current DDL.

Deleting the integer is safe because it was already redundant: the
runner-identity guard deep-compares the whole schema-v4 receipt, including a
digest over the build tree, and forces a rebuild on ANY analyzer delta. Verified
empirically — a comment-only edit to logger.ts, with the fingerprint byte
identical, produced "runner identity changed ... forcing a full rebuild".

The fingerprint is not thereby redundant. It fires where that guard cannot: a
DDL-affecting change in `gitnexus-shared`, which is a workspace-linked
devDependency and so sat outside both digests until the companion commit closed
that gap.

Review findings folded in, each correcting a line this rewrite itself introduced
and never published:

- B1: two assertions matched a log string the rewrite had renamed; both tests
  failed. They now assert what production emits.
- B2: the pre-existing downgrade test perturbed `schemaVersion: 7`, a field this
  change deletes, so the spread carried a valid fingerprint, every guard passed,
  and the run legitimately took the fast path. It perturbs the fingerprint now,
  restoring the only integration coverage of the gate-above-the-fast-path
  ordering invariant.
- N5: duplicate `schemaFingerprint` keys silently collapsed two assertions into
  one (TS1117).
- N6: the absent-stamp message told non-git repositories their index was "built
  by an older GitNexus version" — on every run, about an index this exact build
  had just written. Non-git repos never record a fingerprint, and now the message
  says so.
- N9: the on-disk stamp is shape-checked before being echoed, so a crafted
  gitnexus.json cannot push ANSI escapes through the CLI log.
- N7: a test case that re-computed the same digest expression with its operands
  swapped, mislabelled as a randomness check on a module-level const.
- N10: comments claiming the digest "cannot collide" (it is 48 bits), pointing at
  a vector-column gate that does not exist, and asserting storage/ is free of a
  core/ dependency two lines below a core/ value import.

None of these were caught by `tsc -p tsconfig.json`, which covers src only, nor
by eslint, where no-dupe-keys is off. `tsconfig.test.json` reports all three test
defects and is not currently wired into CI.

Refs #2798

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(schema): pin that the fingerprint covers every DDL statement init executes

`SCHEMA_QUERIES` is the list `runSchemaCreationQueries` iterates — the DDL that
actually runs. The fingerprint hashes only two of its three members, and until
now no test imported `SCHEMA_QUERIES` at all, so nothing tied the two together.

A fourth member appended to that array — the one literally named for what init
executes — would have been invisible to the gate. Every existing test would still
pass, because they all recompute the digest from the same two arrays the
fingerprint already uses. An index whose gate passed would then run `initLbug`
over the old database, where `runSchemaCreationQueries` suppresses "already
exists", so the new table would never be created and its edges would be dropped
by `fallbackRelationshipInserts`' bare catch. A wrong graph, no error — exactly
the failure #2798 exists to end.

The check is a pure predicate over (executed, fingerprinted, documented
exclusions) rather than a positional `toEqual`, so `EMBEDDING_SCHEMA` is named as
an exclusion with its reason — its FLOAT[N] width is environment-derived — rather
than sitting in a list where a future reader might "fix" it by folding it in. It
asserts both directions and is order-insensitive, leaving ordering to the digest
assertion that already pins it.

The negative case is pinned in CI rather than checked by hand once: the same
predicate over a synthetic fourth member must report it. If a refactor ever makes
the predicate vacuous, that case fails even though the positive one would not.

Refs #2798

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(analyze): name the invariant the version deletion now rests on

Deleting `INCREMENTAL_SCHEMA_VERSION` moved a load-bearing guarantee into an
implicit one. Roughly thirty of the retired ladder's entries changed no DDL at
all — node ids, wire formats, resolution tiers — and the fingerprint is
structurally incapable of firing on any of them. Their only remaining cover is
the analyzer runner-identity receipt, and nothing in the suite said so.

This adds a table over the real `analyzerRunnerIdentitiesEqual` with a
well-formed schema-v4 receipt: byte-identical reuses; an entrypoint-only
difference reuses (CLI vs analyze worker); a moved build digest with unchanged
DDL forces — that case IS the invariant, commented as such; and a dependency
change, an ABI change, undefined, null, a schema-v3 legacy receipt, a missing
build section and a non-sha256 digest all fail closed.

The deleted `expect(INCREMENTAL_SCHEMA_VERSION).toBe(35)` pin is also worth
naming: it failed CI on every bump by design, which is what made an author stop
and think. Nothing replaced it. This does not restore that — a digest has no
literal to pin — but it does make the mechanism that took over the job visible to
the next person who reads the file.

Refs #2798

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(spring): pin CLASS_SCHEMA's membership in the fingerprinted DDL set

When `INCREMENTAL_SCHEMA_VERSION` went away, its sibling in
basicblock-callee-ids-schema.test.ts got a replacement assertion tying
BASICBLOCK_SCHEMA to the fingerprint's input set. This file's
`>= 23` floor was deleted with nothing put in its place.

The file still asserts CLASS_SCHEMA's CONTENT — that the `frameworkAnnotations`
column exists — but not that CLASS_SCHEMA is part of what the digest covers, and
the second is what makes an index built before that column carry a different
fingerprint and get rebuilt. Mirrors the sibling so the two read the same way.

Refs #2798

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(identity): stop a symlinked directory from aborting the whole analyze

`collectArtifacts` fused two orthogonal facts into one condition: that four
directory names never carry runtime payload, and that a symlink where a real
directory was assumed falls through to the payload branch, where
`snapshotReadableFile` stats the target, sees a directory, and throws
"Analyzer identity input is not a file".

The second was only fixed for those four names. Every other symlinked directory
in a scanned package root still aborted the run — `dist -> build`, a vendored
grammar link, anything inside a linked sibling checkout. Newly reachable,
because making workspace-linked packages scannable pointed the scanner at a live
checkout instead of an immutable registry tarball for the first time.

Split along the actual seam: prune on the NAME alone, and give symlinks their own
branch in the type dispatch, ahead of the payload branch.

Link text is recorded rather than followed. Following was rejected on three
grounds, each checked in source: the traversal is a stack with no visited set, so
a self-referential link would recurse to `runtimeDepth` — which throws, trading
one hard abort for another; `snapshotDirectory` rejects a symlink outright, so
the directory guard could not accept one without a realpath rewrite of its
canonical-path identity; and a link into an already-scanned tree double-counts
against `runtimeEntries`/`runtimeBytes`, which also throw. The cost is stated in
code: a link out of the package contributes its text, not its target's content.
Links resolving to a regular file keep the existing content digest.

The new `'unfollowed-symlink'` kind is threaded through every consumer, including
the cache validator — which re-probes with `mode: 'link'`, since the readable-file
probe resolves the target and would return null for exactly this kind, silently
failing every warm validation.

No canonicalization or cache-schema bump. Digest content changes only for trees
that previously crashed: a delta scan over all 258 scanned roots of this install
found no regular file bearing a pruned name and no symlink failing to resolve to
a file, so `dependencyRuntime.digest` is byte-identical here.

Six of the eight new tests fail against the unfixed tree with the exact production
error; all eight pass after.

Refs #2798

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(analyze): give the reuse gate a real seam and sanitize logs at the funnel

Cleanup pass over the #2798 branch. Net -183 lines.

The gate had no extracted predicate, so its own test asserted it by regex-matching
run-analyze.ts SOURCE TEXT. That pinned production formatting: one pattern froze
three back-to-back single-name imports from './lbug/schema.js', so merging them —
the obvious tidy-up — failed a test named "still imports the DDL digest itself".

`schemaFingerprintMismatch` and `isSchemaFingerprintShaped` now live in
core/lbug/schema.ts beside the constant. Not in run-analyze.ts next to
`pdgModeMismatch`, because storage/ must stay off the analyze pipeline and
mcp/resources.ts is a plausible second consumer — the same reasoning that puts
`cjkSegmentationModeMismatch` in core/search/. The regex block is gone; the test
calls the predicate. The three imports are merged.

ANSI sanitation moved from one field to the funnel. The per-field guard's own
comment stated the general hazard — gitnexus.json is parsed with no runtime shape
validation and the notice reaches console.log — while two sibling guards twelve
lines away echoed `runnerIdentity.schemaVersion` and `cjkSegmentation` from that
same file raw into the same log. `log()` now strips C0/C1 controls, covering all
seven guard messages and any written later.

Also:
- Deleted a duplicate integration test. After the downgrade test was repointed at
  `schemaFingerprint` it became the same scenario as the new one, differing only
  by an extra log assertion — which is now folded into the survivor. Saves a
  fixture and two full pipeline runs per CI pass.
- Replaced a 3-parameter set-difference helper with one set equality. Its doc was
  false at one call site (arguments semantically swapped) and it needed a fourth
  test purely to prove itself non-vacuous; set equality cannot go vacuous.
- Removed ~115 lines of runner-identity table that duplicated
  analyzer-identity.test.ts. The three genuinely uncovered cases moved there, and
  the #2798 invariant — build digest moved while the DDL did not — now asserts
  against a REAL analyzer-build-tree edit rather than a hand-built literal, which
  is strictly stronger than what it replaces.
- MIGRATION.md quoted a log line the code cannot emit; it was written before the
  placeholder changed.
- Restored the rationale on the `capabilities` docstring, which a previous pass
  replaced with its consequence — leaving a maintainer reading "duplicated by
  hand" as a wart to fix by importing, which is what the original forbade.
- Marked the `isIncremental` conjunct as belt-and-braces: `!options.force`
  short-circuits before it, so it cannot decide anything.

Refs #2798

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(analyze): force a rebuild when the vector column width changes

`CodeEmbedding.embedding` is declared `FLOAT[EMBEDDING_DIMS]`, resolved from
`GITNEXUS_EMBEDDING_DIMS` at module load. Nothing gated it. Flip the variable on
a same-commit clean tree and no guard fired at all: `alreadyUpToDate` returned
over a `FLOAT[384]` table while the process embedded at 768. The only reaction
anywhere discards the embedding CACHE and re-embeds — into a column whose type it
never revisits.

This predates #2798; `INCREMENTAL_SCHEMA_VERSION` never covered dims either. It
surfaced because the fingerprint work had to reason about why `EMBEDDING_SCHEMA`
must stay OUT of the digest: its width is environment-derived, so folding it in
would make the same build disagree with itself and thrash rebuilds. That
exclusion is correct, and it leaves the width needing its own guard.

Modelled on `cjkSegmentation`, the closest sibling: an env-resolved scalar
stamped at write time and compared by a small exported predicate that forces on
mismatch. `embeddingDimsMismatch` sits in core/lbug/schema.ts beside
`EMBEDDING_DIMS`, so the query side can adopt it without importing the analyze
pipeline — mcp/local/local-backend.ts already warns on a cjkSegmentation
disagreement and has the identical claim here, since the query path embeds at the
live width against a table of unknown width with no validation at all today.

ABSENCE IS NOT A MISMATCH, deliberately. Forcing on it would be dead code:
`embeddingDims` and `schemaFingerprint` ship together, and a missing fingerprint
already forces exactly one rebuild — which is where this stamp lands. Absence
also carries no signal here, unlike the fingerprint: a missing fingerprint means
"DDL this build cannot vouch for" and ships WITH a DDL change, whereas a missing
dims stamp means only "written before the field existed", and that run's table
agreed with that run's width. Drift requires the env to change, which absence
says nothing about. The `cjkSegmentation` trick of folding absence into the
default was unavailable — there is no width that is safe to assume for an
existing table — so the stamp is instead written unconditionally, giving absence
exactly one meaning. Malformed values are not grandfathered: null, '384', NaN and
objects all read as a mismatch and err toward a rebuild.

Refs #2798

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(mcp): warn when the served index's vector width differs from the query embedder's

The analyze side now forces a rebuild when the vector column width changes. The
query side had no equivalent: a serving process embeds a query at its own width
and searches a table whose width was fixed when the index was built. Disagree and
the user gets wrong or missing semantic results with nothing explaining why.

Mirrors the cjkSegmentation drift warning immediately above it — same warnings[]
array, same per-query recomputation, agent-visible in the tool response, and it
warns rather than refuses. A width mismatch degrades the semantic lane only;
keyword results are unaffected, so `partial` is deliberately not set.

Compares against `getEmbeddingDims()` — the width the query embedder actually
produces — NOT schema.ts's `EMBEDDING_DIMS`. The two diverge exactly when
GITNEXUS_EMBEDDING_DIMS is set on a server that embeds LOCALLY: the query path
ignores that variable and embeds at 384, so comparing against the env-derived
constant would report drift on a lane that works fine. The recorded width is what
the vector CAST actually binds.

`embeddingDimsMismatch` is imported from core/lbug/schema.js rather than
restated, so "absent is not a mismatch" cannot drift between the analyze and
query sides. That predicate was placed in schema.ts precisely so this consumer
could reach it without importing the analyze pipeline.

Two gates keep it quiet when it would be noise: it fires only for a repo where
this process actually produced a query vector, so an index analyzed without
--embeddings (or a server whose embedder is unavailable) never carries it. An
untrusted recorded value — meta.json is schema-less JSON — is reported as "an
unrecognized width" rather than echoed.

`loadMeta` is hoisted out of the neighbouring try so both diagnostics share one
read and an invalid GITNEXUS_FTS_CJK_SEGMENTATION cannot take this one down with
it.

Refs #2798

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(identity): detect an npm-linked dev dependency the specifier cannot see

`isLocallyLinkedSpecifier` admits a devDependency whose SPECIFIER is
checkout-local. `npm link <pkg>` leaves the specifier a registry range while the
node_modules entry symlinks to a checkout — locally linked, invisible to a
specifier check, so a semantic-only edit there still moves neither digest.

The obvious placement is unaffordable, measured rather than assumed: probing
every dev-only name inside collectRuntimePackages costs 1998 resolutions, not
the ~8 it looks like, because dependencyNames runs for every package in the BFS
and published tarballs retain their devDependencies. Persisted path guards go
2221 -> 11050 (+398%), and every guard is re-probed on each warm validation —
the path `status` takes.

Scoped to the root package instead. The declared-intent half is untouched and
still enumerated everywhere: it alone can emit the `<missing>` edge for a
declared link whose checkout is absent, where resolution returns null and cannot
distinguish that from an uninstalled dev tool. The new resolved-location half
runs only when `parent.root === packageRoot`, resolves through the existing
resolver so its path guards are recorded, and admits a name iff the realpath'd
root carries no node_modules segment.

Bounded against mis-fire by EXPANSION. "Not under node_modules" is a proxy for
"checkout-local"; under a relocated pnpm virtual store every dev dep passes it
and the whole dev tree folds into the receipt — against limits that THROW, so a
legitimate install would abort. Measured here: uncapped, that shape takes
259 -> 347 packages and 2250 -> 3786 guards. The cap admits at most four and
DROPS THE WHOLE CHANNEL on overflow rather than an arbitrary prefix, because the
abort comes from the transitive payload of whichever trees get folded in — four
of a mis-fired thirteen is still unbounded, and a sorted-prefix receipt would be
arbitrary. Overflow falls back to the specifier-only receipt that ships today.

Cost on this install: 259 packages unchanged, 13 dev names resolved, guards
2221 -> 2250 (+29, +1.3%). Verified against the real implementation, not just a
replay: validation guards 16295 -> 16324, packageCount and artifactCount
unchanged, and `dependencyRuntime.digest` byte-identical — so this forces no
re-analysis for anyone.

Each test fails on the defect it targets: disabling the channel kills the
npm-link and cap cases; dropping the root-only scope makes the differential
guard-count case fail at 2.8x guards; removing the specifier half kills the
`<missing>` case.

Refs #2798

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 15:04:30 +01:00
561f913a32 fix(embeddings): retry unparseable 200 responses and survive partial embedding failures (#2790) (#2795)
* fix(embeddings): retry unparseable 200 responses and survive partial embedding failures (#2790)

A long-running embedding job against an OpenAI-compatible endpoint could lose
hours of work to a single transient glitch, then refuse to recover on the next
run. Four defects compounded:

1. An HTTP 200 carrying a truncated or non-JSON body was never retried.
   `classifyOutcome` treats any 2xx as success, and the `resp.json()` parse ran
   after `resilientFetch` had already returned, so the parse failure surfaced as
   a terminal error. Measured: a 503 got 3 attempts, a garbage 200 got 1.

   The parse and the response-shape check now run inside the `fetchImpl`
   callback, so a bad body is classified as a retryable failure and gets the
   same backoff as a 5xx. This also stops a garbage 200 from calling the circuit
   breaker's `recordSuccess()`, which previously erased accumulated failures and
   meant an endpoint alternating 5xx and garbage-200 could never trip it.

2. One failed `embedBatch` sub-batch aborted the entire pipeline. Failures are
   now tolerated: the sub-batch's node ids are collected and all of their
   embedding rows are deleted, so those nodes hold zero rows and are re-embedded
   later. Deleting rather than keeping partial rows is deliberate — chunk arrays
   are flat over a 16-node batch and sliced by 8, so a node's chunks can straddle
   a sub-batch boundary, and surviving rows carry the current content hash. The
   hash maps collapse per-chunk rows last-row-wins, so a partially embedded node
   would read as fresh forever and never regenerate its missing chunks.

   A run that fails 5 sub-batches in a row still aborts, and rethrows the first
   error of the streak rather than the last: after 3 failures the circuit breaker
   opens, so later errors degrade into "circuit open, retry in 30s" while the
   first still names the real defect.

3. The Phase 5 `embeddingCount === 0` fail-fast could not tell "wrote nothing"
   from "could not ask" — the count query's catch was silent. The count is now
   tri-state and only a known zero after real work is fatal. A non-numeric count
   previously bypassed the gate entirely, because `Number()` returns NaN and
   `NaN === 0` is false, and then serialized as `embeddings: null`. An unverified
   count no longer certifies `capabilities.vectorSearch.status`.

4. `saveEmbeddingCheckpoint` wrote a completion-shaped meta: it advanced
   `lastCommit`, wrote the new `fileHashes` and cleared `incrementalInProgress`.
   The first checkpoint window fires before a single embedding exists, and on a
   full rebuild the graph is still in a staging database that a crash discards.
   The next run then diffed against the advanced hashes, saw no changes and
   preserved the old graph — the "skipping wipe" symptom in the report. It now
   re-reads meta and replaces only the checkpoint, matching what the server
   endpoint already did.

A partially failed run keeps its checkpoint with the failed ids in
`pendingNodeIds`, so the next plain `analyze` regenerates them through the
existing resume path. Clearing it would have been silent data loss: a plain run
derives `shouldGenerateEmbeddings: false` once embeddings exist, so the pipeline
would never have run again. The old crash-and-abort self-healed only by accident,
via the checkpoint its crash left behind. `gitnexus status` reports the index
incomplete until the nodes recover, and `--drop-embeddings` still abandons them.

`POST /api/embed` is the pipeline's other caller and was discarding the result,
reporting "Embeddings complete" for a partial run. It now persists the pending
ids and reports the run as failed with the underlying endpoint error.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(embeddings): abort a run whose sub-batch failure ratio is too high (#2790)

The consecutive-failure ceiling only catches a total outage, because any
successful sub-batch resets it. An endpoint under load shedding that alternates
success and failure never trips it, so the run walks the whole corpus, deletes
every failed node's rows and exits 0 having dropped a large fraction of the
index. The retained checkpoint made that visible in `gitnexus status`, but a run
that drops a quarter of the corpus should tell the operator to fix their
endpoint, not leave them to notice a status flag.

Adds a cumulative guard: abort once more than 25% of attempted sub-batches have
failed, evaluated as the run progresses and gated behind a floor of 20 attempted
sub-batches. The shape follows Resilience4j's circuit breaker (failure rate plus
a minimum-sample floor) because it is the only one of the surveyed designs that
answers the small-repo case — a three node repo can fail one sub-batch and never
accumulate enough sample for a ratio to mean anything. The rate sits below a live
traffic breaker's 50% because a batch indexer's job is to index the whole corpus
rather than serve degraded traffic, and above Hadoop's single-digit
`failures.maxpercent` because tolerating transient hiccups is the point of the
change this follows.

The guard reuses the existing break-then-cleanup path, so the failed batch's
DELETE still runs before the rethrow, and it wraps the retained first-error-of-
streak rather than inventing a new one, so the message names both the ratio and
the underlying endpoint failure.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(server): record the embedding count after /api/embed so the next analyze cannot wipe it

`POST /api/embed` generated embeddings and wrote them to the database but never
wrote `stats.embeddings` into meta.json. Its checkpoint writer replaced only
`embeddingCheckpoint`, and the finalize write folded in nothing else.

So a repo embedded purely through the server kept whatever count the last CLI
`analyze` stamped, which is 0 for a repo analyzed without embeddings. The next
CLI run read `existingEmbeddingCount = 0`, `deriveEmbeddingMode` returned
`shouldLoadCache: false`, and `gitnexus analyze --force` wiped the database with
no cache load. Every server generated embedding was silently destroyed, with no
warning — the user just lost semantic search.

The route now measures the live count with the same query the CLI uses and folds
it into both meta writes. The measurement is tri-state and deliberately never
falls back to 0: an unverified count is written as absent rather than as zero,
because a wrong-low value is exactly what arms the wipe. It is taken after
`flushWAL()` and inside `withLbugDb`, so it describes durable rows and the
connection is still open. A partial run records its honest count too, alongside
the retained checkpoint, so the next CLI run preserves the partial index instead
of discarding it.

Found while working #2790; not part of that issue.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(embeddings): retry short 200 bodies and stop laundering body-phase timeouts

Two gaps in the #2790 retry fix, both found by review.

A 200 carrying `{"data": []}` or fewer vectors than inputs passed the
in-`fetchImpl` shape check, because `every(isEmbeddingItem)` is vacuously true
for an empty array. `resilientFetch` then classified it `success` and called
`recordSuccess()`, erasing the outage signal, and the cardinality check in
`httpEmbed` threw terminally one attempt later. That is exactly the pair of
properties #2790 was filed about, still broken for this body shape — and worse
than before the fix, since the pipeline now tolerates the error by deleting
those nodes' rows instead of aborting loudly. The count check moves inside the
retried callback; the outer one stays as a backstop.

The `.json()` catch also swallowed every rejection, not just parse errors.
`AbortSignal.any([caller, timeout])` is wired to the body stream, so a stalled
body rejects with a DOMException — which, wrapped in a plain Error, defeated
`classifyOutcome`'s terminal-network test. Measured: the same TimeoutError got
3 attempts and "unparseable response" when raised during the body read, but 1
attempt and "timed out after 180000ms" when raised by fetch itself, and three
such sub-batches opened the process-global breaker that `recordNeutral()`
exists to protect. Abort-like DOMExceptions are now re-raised unchanged.

The dimension check stays outside the loop deliberately: it validates against
`config.dimensions ?? DEFAULT_DIMS`, not the request-dimensions argument, and a
width mismatch is a configuration error where retrying only triples latency and
books failures against a healthy endpoint.

Adds the negative assertion the review found missing: response body text must
never reach the user-facing error string.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(embeddings): scale the sub-batch failure-ratio floor to the run

The cumulative guard needed 20 attempted sub-batches before a failure rate could
abort anything — roughly 160 chunks, or ~80 embeddable nodes at the default
subBatchSize of 8. A 50-node repo whose endpoint sheds every other sub-batch
fails half of them and still exits 0: the ratio guard is below its floor, and
every intervening success resets the consecutive ceiling.

The floor was a good choice for a first run over a small repo, where one failure
out of one sub-batch is 100% and means nothing. The defect is that every resume
run has that shape by construction — its node set is only the pending ids — so
the guard was structurally off in the one run whose entire purpose is retrying
against the endpoint that already failed.

The floor is now sized to the run: clamp(ceil(totalNodes / 16), 5, 20). The
lower bound keeps the case the flat floor protected; the upper bound preserves
today's behavior above 320 nodes and avoids a proportional-only floor perversely
weakening the guard at scale, where a sixteenth of a 20k-node repo would be 1250
sub-batches of damage before a rate could fire. Resilience4j can use a constant
minimumNumberOfCalls because a breaker sits on an unbounded call stream; a batch
indexer has a finite budget, so a constant can exceed the whole run.

The ratio is still evaluated only inside the catch. That is already its local
maximum — both counters have just incremented — so sampling more often would
only ever observe lower ratios.

Also: a failing cleanup DELETE no longer swallows the abort, which was
discarding the retained first-error-of-the-streak that names the real endpoint
fault; `ceilingError` is renamed `abortError` since it carries the ratio abort
too; and three `{ error }` log keys become `{ err }` (#2114 — an arbitrary key
serializes to `{}`, losing message and stack).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(analyze): one tri-state embedding counter, and stop partial runs wedging later runs

The tri-state count doctrine this branch introduced was applied at two of its
three CLI sites, and the two implementations that were meant to mirror each
other had already drifted.

`measurePersistedEmbeddingCount` moves to `core/embedding-count.ts` — beside
`embedding-mode.ts`, with the same no-native-imports property, and outside
`core/embeddings/` so the lazy-embeddings convention (#2370) still holds. All
three call sites now share it.

  - The mid-run `onCheckpoint` counter ran the query bare. A throw there — DB
    busy, connection closed, read-only, the VECTOR DML lock (#2623) — rejected
    the callback out of `runEmbeddingPipeline` and killed the analyze before
    Phase 5 could apply the tri-state that exists for exactly this case. A
    non-numeric cell wrote `stats.embeddings: null` to disk mid-run.
  - Phase 5 used `?? 0` while the server used `?? Number.NaN`, under a comment
    asserting both measured the field the same way. `Number.isFinite(0)` is
    true, so a no-row answer became a *measured* zero and hard-failed a run
    whose embeddings had all persisted.
  - The unknown-count fallback read `existingMeta`, assigned once at run start,
    so it republished the pre-run figure over the fresher count the terminal
    checkpoint had already written. With a prior count of 0 that armed the wipe
    chain: hasExisting false, shouldLoadCache false, and the next --force
    discards live embeddings. It now re-reads the latest on-disk meta, and an
    unverifiable count retains a recovery marker instead of clearing it.

A completed-but-partial run also planted a landmine. Its checkpoint is stamped
with the run's embedding identity, so a later plain `gitnexus analyze` from a
hook, a CI job, or a shell without GITNEXUS_EMBEDDING_URL resolved provider
'local' and threw before any phase ran — after an exit-0 run, where previously
only a visible crash left that state. `--force` did not help: the resume gate
inspected only `--drop-embeddings`.

`RepoMeta.embeddingCheckpoint` gains `kind` to tell the two situations apart.
An 'interrupted' marker (or one with no kind, so markers already on disk keep
the stricter path) still fails closed — its nodes may be half-written, and
resuming under a foreign model would mix vector spaces. A 'partial' marker
names nodes the pipeline already deleted to zero rows, so nothing is at risk: an
identity mismatch drops the pending set with a warning and continues. `--force`
now discards a checkpoint, and `attempts` bounds the retry at
EMBEDDING_RESUME_MAX_ATTEMPTS (3, matching the HTTP embedder's and the WAL
driver's existing per-operation budgets) so a node the endpoint deterministically
rejects converges instead of keeping the repo incomplete forever.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(server): close the SSE stream on terminal job status, not a progress phase

A tolerated partial run reached SSE clients as a clean success — a regression in
this branch's own claim that /api/embed reports a partial run as failed.

The pipeline emits `phase:'ready'` unconditionally before returning, including
when it dropped nodes. The route mapped that to `'complete'`, and
`mountSSEProgress` treated a terminal-looking *progress phase* as terminal:
write the event, `res.end()`, `unsubscribe()`. The route's own
`updateJob({status:'failed'})` then fired into a stream with no listener, and
the web app had already shown "ready". Before this branch the pipeline threw,
which produced `phase:'error'` and did reach the client. Pollers on
GET /api/embed/:jobId were unaffected, so the two consumers disagreed.

Terminality is a property of the job, so the relay now asks the job. Remapping
`ready` alone would have left the trap armed: the `error -> 'failed'` mapping
has the identical shape and would emit `event: failed` with `error: undefined`
before the catch block fills the message in. `ready` is additionally remapped to
`finalizing` so a poller no longer sees `status:'analyzing'` next to
`progress.phase:'complete'`. The single-terminal-event property (#2264) is
preserved on both the clean and partial paths, and /api/analyze is unaffected —
its terminal progress phase is 'done', never 'complete'.

`AnalyzeJob` gains an optional `partial` payload so a client can tell a partial
run from a total failure without a new status member; it is absent on every
other job, so existing payloads stay byte-identical. Consuming it in
gitnexus-web is left to that app's owner — today it renders both as the same
red retry chip.

`resolveEmbedRunOutcome` moves to `embed-run-outcome.ts` and `mountSSEProgress`
to `sse-progress.ts`, both free of Express/LadybugDB/MCP imports, and the local
count copy is replaced by the shared `core/embedding-count.ts`. Reaching three
pure functions previously meant importing the whole server: measured at ~20s
against a 30s test timeout, with one observed timeout failure. That file is now
1.6s.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: document the partial embedding index and its recovery

A run can now finish exit 0 with a partial embedding index, which neither
operator doc described.

GUARDRAILS' "Embeddings vanished after analyze" Sign keys its trigger on
`stats.embeddings` being 0 and lists "the only ways to end up at zero". A
partial run stamps an honest non-zero count and sets `embeddingCheckpoint`, so
the operator's actual symptom is `incompleteReasons:
["embedding-checkpoint-pending"]` — a state that Sign cannot match. Adds a Sign
for it and drops the exhaustive framing from the existing one.

RUNBOOK gains the recovery path: a plain `gitnexus analyze` is correct and needs
no flag, because a retained checkpoint forces generation for the pending nodes
regardless of flags. Also corrects two stale claims — that `stats.embeddings` is
always freshly measured (it can carry forward when the count query cannot
answer, which is why `capabilities.vectorSearch.status` is the certified read),
and that later analyzes must always pass `--embeddings` or lose their vectors,
which contradicts Non-negotiable 5.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(embeddings): one owner for the checkpoint record and the abort predicate

Cleanup pass over the #2790 review fixes. No behavior change except where
noted; the two exceptions are both cases where the code was lying to the
operator or to the other half of itself.

The previous pass extracted `core/embedding-count.ts` because two hand-copied
bodies of "measure the embedding count" had drifted inside a single change. It
then created a second pair of hand-copied publishers — of
`RepoMeta.embeddingCheckpoint` — and those had drifted too: the CLI armed the
attempt counter only after clearing its identity gate, the server derived it
from the resumed marker alone. Only one of the two READERS implemented `kind`
at all, so a 'partial' marker written by `gitnexus analyze` and resumed through
POST /api/embed still hit the permanent wedge `kind` exists to remove.

`core/embedding-checkpoint.ts` now owns the record: `checkpointKind` (the one
home for absent-means-interrupted), the three minters, `nextAttemptCount`, and
`decideEmbeddingResume`, which both gates route through. Five mint sites and
two resume gates become one implementation each.

`resilient-fetch.ts` exports `isTerminalNetworkError` and `classifyOutcome`
calls it, replacing a caller-side copy of the same DOMException test whose
docstring promised it "mirrors classifyOutcome exactly" — an invariant enforced
by prose, where a divergence silently reverts body-phase timeouts to being
retried three times and charged to the shared breaker.

The ratio-guard floor now divides by the run's actual `subBatchSize` instead of
a constant 16 that assumed the default of 8. At `subBatchSize: 32` the old
formula demanded more sub-batches than the run contains, leaving the guard
structurally off — the exact failure the scaled floor was introduced to fix,
and sub-batch size is tuned mainly for the flaky endpoints it protects.

Two operator-facing corrections:

  - The count-recovery marker was stamped `kind: 'partial'` with an empty
    pending set, so `gitnexus status` reported "N node(s) lost their embeddings"
    where N is zero. It gets its own kind and its own incomplete reason.
  - `decideEmbeddingResume` initially keyed its skip-the-identity-gate branch on
    an empty pending set, assuming that meant the count-recovery marker. It does
    not: `onCheckpoint` mints an 'interrupted' marker with no pending nodes
    after every post-window save. That silently cleared an interrupted marker
    under a foreign provider instead of failing closed. Keyed on `kind` now,
    with a regression test.

Also: `isTerminalJobStatus` adopted at the seven sites that still hand-copied
it, including the one gating the single-terminal-event emit; `mountSSEProgress`
re-export dropped and `server-sse-payload.test.ts` repointed at the extracted
module, which takes it from 24.60s to 0.408s — the test that motivated the
extraction was still paying the cost it was meant to remove; the count-mismatch
message and the SSE test harness deduplicated; per-batch error strings made
lazy (~75k needless `new URL()` per large run); `retryable: true` dropped as a
field that can never be false; ~110 lines of restated rationale reduced to
pointers at their canonical home.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 20:43:59 +00:00
azizur100389andCursor 797e4ef8f6 fix(ai-context): document CLI graph fallbacks (#2803)
* fix(ai-context): document CLI graph fallbacks

Teach generated GitNexus guidance to pair mandatory MCP graph checks with repo-scoped CLI fallbacks so agents can keep working when MCP is unavailable.

Co-authored-by: Cursor <cursoragent@cursor.com>

* style(ai-context): satisfy Prettier

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-02 20:52:44 +01:00
Gergő Magyar 990d79ba8c fix(mcp): make impact/context reproducible — deterministic ordering on every capped query (#2787) (#2796) 2026-08-02 17:03:15 +00:00
010a7d806a fix(schema): declare the full scope-resolution relation cross product (#2792) (#2793)
* fix(schema): declare the full scope-resolution relation cross product (#2792)

`RELATION_SCHEMA` was hand-listed, and every prior fix added only the
FROM/TO pair named in a crash report — `Const→Method` in #2769, the
Swift/Rust member pairs before it. So `analyze` kept aborting at
`assertDeclaredPair` on the next codebase whose edges happened to land on
a different pair; #2792 reports `Class→Variable` on Java.

Audit the surface instead of the symptom. `buildGraphNodeLookup` skips
any node whose label is not in `isLinkableLabel`, so the lookup holds
only linkable-labelled nodes — and both endpoints of every graph-bridge
edge resolve through that lookup. The emittable surface is therefore
exactly:

  FROM  LINKABLE_LABELS + File   (the module-level caller fallback)
  TO    LINKABLE_LABELS + CALL_TARGET_TYPES

`isCallerAnchorLabel` is a strict subset of linkable and contributes
nothing on top. `CALL_TARGET_TYPES` contributes `Delegate`, which
`tryEmitEdgeWithExplicitTargetId` can emit without going through the
lookup at all.

Generate that 14x14 block into the DDL rather than listing it: 223 -> 322
declared pairs, and no future pair from these sets can be missing by
construction. The containment/inheritance/DI/route/cluster/PDG pairs stay
hand-declared — no single predicate describes them.

Both label sets live in the ingestion layer, which `core/lbug` must not
import, so schema.ts carries twin lists. test/unit/schema-pair-coverage.ts
derives the requirement from the originals and fails CI when either set
grows without the pairs landing here — the piecemeal loop this fix ends.

Measured before widening: at 322 pairs the cost is inside noise
(1.09s vs 1.12s per 300 anchored queries on a 32-table DB), but the full
32x32 cross product is ~1.8x on untyped-endpoint anchored queries. The
audited subset is the right scope, not "declare everything".

INCREMENTAL_SCHEMA_VERSION 34 -> 35: LadybugDB fixes endpoint pairs when
the rel table is created, so a pre-v35 database physically cannot store
these edges.

Closes #2792

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(schema): declare the non-bridge structural pairs COBOL and Vue emit

The generated scope-resolution block closed the half of RELATION_SCHEMA a
label predicate can describe. The hand-declared half was still stale: with
#2791's Function->Variable fix applied, `analyze` continued to abort on this
repo's own test/fixtures/lang-resolution with

  Relationship label pair Module→Property is not declared

A full sweep (assertDeclaredPair patched to log-and-skip, run over the whole
fixture corpus) found 13 undeclared pairs over 106 edges. This branch already
covered 3 of them via the cross product; the remaining 10 come from emitters
outside the graph bridge:

  - cobol-processor.ts mints Module / Namespace / Record / Property /
    CodeElement and wires them with CONTAINS, CALLS and ACCESSES (9 pairs)
  - vue-sfc-extractor.ts emits BINDS_EVENT_HANDLER from a handler Function to
    the child component's File, the only edge whose target is a File (1 pair)

CodeElement, Namespace, Record and File are in neither scope-bridge label set,
so neither the generated block nor schema-pair-coverage.test.ts can reach them.

Adds test/integration/structural-pair-coverage.test.ts, which derives the
requirement from a corpus instead of a predicate: it runs the real pipeline
over the non-bridge fixtures and requires every FROM/TO pair they produce to
be declared. Mutation-checked — dropping `FROM Function TO File` fails it with
exactly Function|File.

Verified: cobol-app, vue-basic and php-transitive-traits now index instead of
aborting; the full lang-resolution corpus completes at 10,876 nodes / 18,517
edges; scrypster/muninndb at 0b7a4272 (the #2789 repro) completes at 20,069
nodes / 71,580 edges, matching #2791 exactly, so this supersedes that PR.

* refactor(test): simplify the structural pair coverage guard

Cleanup pass over the previous commit. No behaviour change to the schema.

- reuse `FIXTURES` and `runPipelineFromRepo` from resolvers/helpers.ts instead
  of re-deriving the fixture root and importing pipeline.js directly
- gate on `distWorkerExists()` like every other integration test that passes
  `workerUrlForTest`, so a missing dist skips rather than fails
- run the three fixtures with `it.concurrent.each`; they share nothing and the
  cost is almost all worker spawn plus grammar load, which overlaps well
  (tests phase 21-24s -> 5.6s measured)
- replace the sentinel-in-a-Set filter with a plain `.filter()` chain, matching
  the sibling unit test, and move the declared/table lookups off the per-edge
  path onto the deduped set
- move the pure string pin out of the integration tier into
  schema-pair-coverage.test.ts, where the identical construct already lives, so
  it needs no build and survives fixture deletion
- trim the schema and test prose that restated the code, and correct the
  BINDS_EVENT_HANDLER attribution: it is emitted by
  languages/vue/scope-resolver.ts, not vue-sfc-extractor.ts
- amend the v35 comment to mention the 10 structural pairs it now also stamps

Still mutation-checked: dropping `FROM Function TO File` now fails both the
integration sweep and the unit pin with exactly Function|File. 89 tests green.

* fix(schema): generate the attachment pair surface and close four analyze aborts

Review of the generated scope-bridge cross product found four `analyze`
hard-aborts still live at head, each reproduced end-to-end on the default
user path (`analyze --index-only --skip-git`):

  Method→Annotation   Spring `@Bean` + `@ConditionalOnMissingBean` (Java + Kotlin)
  Method→File         Vue Options-API `methods:` handler bound to a child event
  Namespace→Record    COBOL `DECLARATIVES` / `USE AFTER STANDARD ERROR ON <file>`
  Class→Tool          `@mcp.tool()` applied to a class

All four are pre-existing on main, and both existing guards were structurally
blind to them: the unit guard derives from LINKABLE_LABELS ∪ CALL_TARGET_TYPES
(none of Annotation/Tool/Record/File-as-target is a member) and the corpus
guard ran three fixtures that exercise none of these emitters. All 16 tests
passed while all four crashes were live.

The PR's model — "bridge endpoint × structural endpoint" — does not fit:
Namespace→Record is structural on both sides. The property that does hold is
that the ANCHOR is a lookup result, not a literal at the emit site, so the
emitter cannot constrain its label. That gives a second closed-form rule:

  DEFINITION_ANCHOR_LABELS × ATTACHMENT_TARGET_LABELS

DEFINITION_ANCHOR_LABELS is derived from NODE_TABLES by subtraction, so a new
node table joins automatically. 332 → 450 declared pairs.

Sized against a committed harness (gitnexus/bench/schema-pairs), real
@ladybugdb/core, identical data: 450 costs 0.93–1.05× of 332 on untyped-endpoint
anchored queries — inside noise — versus 1.22–1.43× at 641 and 2.03–2.34× at
1024. The harness reproduces the known #2792 cliff, which is what makes the 450
figure trustworthy.

Also in this change:

- Delete the 161 hand-declared pairs the rules already generate (233 → 72).
  The declared set is byte-identical at 450; those lines were load-bearing
  shadow, because the generator suppresses anything already declared
  structurally, so narrowing a rule later would silently keep pairs alive.
  A new guard fails CI if a hand-declared pair is ever re-added inside a rule.
- Import LINKABLE_LABELS / CALL_TARGET_TYPES instead of hand-copying them.
  The twins' stated justification ("the ingestion layer must not be imported
  here") is false: csv-generator.ts and lbug-adapter.ts, siblings in the same
  directory, already do, and no rule in AGENTS.md / ARCHITECTURE.md /
  CONTRIBUTING.md / GUARDRAILS.md states otherwise.
- Resolve `resolveStreamGraphEmit` after the guards that rebind `options.force`,
  not at function entry. It gates on `force`, and every freshness guard runs
  ~360 lines later, so the v34→v35 bump would have pushed every existing index
  down the non-streamed emit path — losing the #2680 memory streaming added for
  the #2649 kernel-scale OOM, for exactly the population most likely to be
  memory-constrained.
- `UndeclaredRelationPairError` now carries the relationship type, both node ids
  and the source file, with a matching CLI branch. The old message named only
  the abstract label pair, which a user could not act on. Found through the
  cause chain, since pipeline-phases/runner.ts rewraps every phase failure.
- Share one classifier (`relPairKeyFor`) across the router, both emit sinks and
  the corpus guard, which previously hand-mirrored the router's skip rule; one
  cause-chain walker in lib/utils.ts; one exported pair-matching regex.
- Corpus guard: four new fixtures reproducing the aborts, per-fixture sentinel
  pairs so a fixture that stops emitting fails loudly instead of passing
  vacuously on an empty graph.

The per-edge path stays allocation-free: the failure context is passed
positionally and the message is built only inside the throw.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0182jkjQqzACkJKYw4MLDnhX

* test(bench): re-baseline the COBOL capture fingerprint for the new fixture

`bench/scope-capture` globs `lang-resolution/cobol-*`, so the
`cobol-declaratives` fixture added in 81daf370e (to reproduce the
`Namespace→Record` analyze abort) joined that corpus and shifted the
fingerprint — 14 → 15 files.

Verified corpus-only, not a capture change: with that one fixture moved
aside the fingerprint is byte-identical to the prior baseline
(d45bb091…), and 81daf370e touches no COBOL capture code. The new value
reproduces CI's reported hash exactly. Scaling 0.677 < 1.5 budget.

`bench/scope-capture/measure.mjs --check` → PASS (15 languages).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0182jkjQqzACkJKYw4MLDnhX

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 16:26:25 +01:00
130 changed files with 17421 additions and 1984 deletions
@@ -17,22 +17,23 @@ description: "Use when the user wants to know what will break if they change som
## Workflow
```
1. impact({target: "X", direction: "upstream"}) → What depends on this
1. impact({target: "X", direction: "upstream"}) or `node .gitnexus/run.cjs impact "X" --direction upstream --repo .`
2. READ gitnexus://repo/{name}/processes → Check affected execution flows
3. detect_changes() → Map current git changes to affected flows
3. detect_changes({scope: "all"}) or `node .gitnexus/run.cjs detect-changes --scope all --repo .`
4. Assess risk and report to user
```
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
> If `.gitnexus/run.cjs` is missing, replace `node .gitnexus/run.cjs` with `npx gitnexus` in the fallback commands.
## Checklist
```
- [ ] impact({target, direction: "upstream"}) to find dependents
- [ ] impact({target, direction: "upstream"}) or CLI fallback to find dependents
- [ ] Review d=1 items first (these WILL BREAK)
- [ ] Check high-confidence (>0.8) dependencies
- [ ] READ processes to check affected execution flows
- [ ] detect_changes() for pre-commit check
- [ ] detect_changes({scope: "all"}) or CLI fallback for pre-commit check
- [ ] Assess risk level and report to user
```
@@ -55,7 +56,7 @@ description: "Use when the user wants to know what will break if they change som
## Tools
**impact** — the primary tool for symbol blast radius:
**impact** — the primary tool for symbol blast radius. If MCP is unavailable, use `node .gitnexus/run.cjs impact <symbol> --direction upstream --repo .` instead:
```
impact({
@@ -73,10 +74,10 @@ impact({
- authRouter (src/routes/auth.ts:22) [CALLS, 95%]
```
**detect_changes** — git-diff based impact analysis:
**detect_changes** — git-diff based impact analysis. If MCP is unavailable, use `node .gitnexus/run.cjs detect-changes --scope all --repo .` instead:
```
detect_changes({scope: "staged"})
detect_changes({scope: "all"})
→ Changed: 5 symbols in 3 files
→ Affected: LoginFlow, TokenRefresh, APIMiddlewarePipeline
@@ -86,7 +87,7 @@ detect_changes({scope: "staged"})
## Example: "What breaks if I change validateUser?"
```
1. impact({target: "validateUser", direction: "upstream"})
1. impact({target: "validateUser", direction: "upstream"}) or `node .gitnexus/run.cjs impact "validateUser" --direction upstream --repo .`
→ d=1: loginHandler, apiMiddleware (WILL BREAK)
→ d=2: authRouter, sessionManager (LIKELY AFFECTED)
+11 -4
View File
@@ -120,10 +120,17 @@ and do not claim a complete graph-backed review.
review surface: when the diff changes what gets emitted or persisted,
verify every schema/version constant gating caches, incremental
writebacks, and fingerprint baselines was bumped or regenerated — in
GitNexus itself, for example: `INCREMENTAL_SCHEMA_VERSION` (the
incremental write set covers only changed files, so new cross-file edges
never reach an existing index without the bump), the parse-store
`SCHEMA_BUMP`, and both bench fingerprint sets.
GitNexus itself, for example: graph DDL needs no manual bump, because
`SCHEMA_FINGERPRINT` (`gitnexus/src/core/lbug/schema.ts`) is derived
from `NODE_SCHEMA_QUERIES` + `REL_SCHEMA_QUERIES` and moves on its own;
the check there is whether the diff changed any string in those arrays,
and, if it added a new DDL array, whether that array was folded into the
fingerprint. The hand-maintained ritual still applies where no
declarative artifact describes the invalidated set: the parse-store
`SCHEMA_BUMP` and both bench fingerprint sets still need an explicit
bump, re-checked against the base branch right before merge. Semantic
changes that leave the DDL untouched are outside the fingerprint; they
rely on the analyzer runner-identity receipt in the index metadata.
## Expert lenses
+8 -7
View File
@@ -111,30 +111,31 @@ mirror. `gitnexus/test/unit/shipped-skills-sync.test.ts` guards the copies. Toke
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (20319 symbols, 54304 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitNexus** (248612 symbols, 565510 relationships, 918 execution flows). Use GitNexus graph tools to understand code, assess impact, and navigate safely.
> Index stale? Run `node .gitnexus/run.cjs analyze` from the project root — it auto-selects an available runner. No `.gitnexus/run.cjs` yet? `npx gitnexus analyze` (npm 11 crash → `npm i -g gitnexus`; #1939).
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows. For regression review, compare against the default branch: `detect_changes({scope: "compare", base_ref: "main"})`.
- **MUST run impact analysis before editing.** Use `impact({target: "symbolName", direction: "upstream"})` (MCP) or `node .gitnexus/run.cjs impact "symbolName" --direction upstream --repo .` (CLI fallback); report callers, processes, and risk. Never substitute grep for graph analysis. For unified PDG impact, add `mode: "pdg"` with optional `line: <N>` — it returns statement-level `affectedStatements` over CDG + REACHING_DEF and inter-procedural symbols in `interproceduralByDepth`/`byDepth`; no-layer/degraded PDG results are UNKNOWN-risk notes (`--pdg` layer). CLI equivalent: `node .gitnexus/run.cjs impact "symbolName" --direction upstream --mode pdg --line <N> --repo .`.
- **MUST analyze graph changes before committing.** Use `detect_changes({scope: "all"})` (MCP) or `node .gitnexus/run.cjs detect-changes --scope all --repo .` (CLI fallback). For regression review: `detect_changes({scope: "compare", base_ref: "main"})` or `node .gitnexus/run.cjs detect-changes --scope compare --base-ref "main" --repo .`.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `query({search_query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `context({name: "symbolName"})`.
- For security review, `explain({target: "fileOrSymbol"})` lists taint findings (source→sink flows; needs `analyze --pdg`).
- For control/data dependence, `pdg_query({mode: "controls", target: "fileOrSymbol"})` answers "under what condition does X run?" (CDG, incl. guard clauses) and `pdg_query({mode: "flows", target, variable})` traces "where does variable Y flow?" (REACHING_DEF). `--pdg` layer.
## Never Do
- NEVER edit a function, class, or method without first running `impact` on it.
- NEVER edit a function, class, or method before MCP/CLI impact analysis.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `rename` which understands the call graph.
- NEVER commit changes without running `detect_changes()` to check affected scope.
- NEVER commit before MCP/CLI graph change analysis.
## Resources
| Resource | Use for |
|----------|---------|
| --- | --- |
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
@@ -143,7 +144,7 @@ This project is indexed by GitNexus as **GitNexus** (20319 symbols, 54304 relati
## CLI
| Task | Read this skill file |
|------|---------------------|
| --- | --- |
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus-debugging/SKILL.md` |
+8 -7
View File
@@ -62,30 +62,31 @@ See the `<!-- gitnexus:start --> … <!-- gitnexus:end -->` block in **[AGENTS.m
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (20319 symbols, 54304 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitNexus** (248612 symbols, 565510 relationships, 918 execution flows). Use GitNexus graph tools to understand code, assess impact, and navigate safely.
> Index stale? Run `node .gitnexus/run.cjs analyze` from the project root — it auto-selects an available runner. No `.gitnexus/run.cjs` yet? `npx gitnexus analyze` (npm 11 crash → `npm i -g gitnexus`; #1939).
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows. For regression review, compare against the default branch: `detect_changes({scope: "compare", base_ref: "main"})`.
- **MUST run impact analysis before editing.** Use `impact({target: "symbolName", direction: "upstream"})` (MCP) or `node .gitnexus/run.cjs impact "symbolName" --direction upstream --repo .` (CLI fallback); report callers, processes, and risk. Never substitute grep for graph analysis. For unified PDG impact, add `mode: "pdg"` with optional `line: <N>` — it returns statement-level `affectedStatements` over CDG + REACHING_DEF and inter-procedural symbols in `interproceduralByDepth`/`byDepth`; no-layer/degraded PDG results are UNKNOWN-risk notes (`--pdg` layer). CLI equivalent: `node .gitnexus/run.cjs impact "symbolName" --direction upstream --mode pdg --line <N> --repo .`.
- **MUST analyze graph changes before committing.** Use `detect_changes({scope: "all"})` (MCP) or `node .gitnexus/run.cjs detect-changes --scope all --repo .` (CLI fallback). For regression review: `detect_changes({scope: "compare", base_ref: "main"})` or `node .gitnexus/run.cjs detect-changes --scope compare --base-ref "main" --repo .`.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `query({search_query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `context({name: "symbolName"})`.
- For security review, `explain({target: "fileOrSymbol"})` lists taint findings (source→sink flows; needs `analyze --pdg`).
- For control/data dependence, `pdg_query({mode: "controls", target: "fileOrSymbol"})` answers "under what condition does X run?" (CDG, incl. guard clauses) and `pdg_query({mode: "flows", target, variable})` traces "where does variable Y flow?" (REACHING_DEF). `--pdg` layer.
## Never Do
- NEVER edit a function, class, or method without first running `impact` on it.
- NEVER edit a function, class, or method before MCP/CLI impact analysis.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `rename` which understands the call graph.
- NEVER commit changes without running `detect_changes()` to check affected scope.
- NEVER commit before MCP/CLI graph change analysis.
## Resources
| Resource | Use for |
|----------|---------|
| --- | --- |
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
@@ -94,7 +95,7 @@ This project is indexed by GitNexus as **GitNexus** (20319 symbols, 54304 relati
## CLI
| Task | Read this skill file |
|------|---------------------|
| --- | --- |
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus-debugging/SKILL.md` |
+7 -1
View File
@@ -44,7 +44,13 @@ Format: **Trigger → Instruction → Reason**. Append new Signs when the same m
- **Trigger:** Semantic search quality drops; `stats.embeddings` in the index metadata (`gitnexus.json` / legacy `meta.json`) is 0 after refresh.
- **Do:** Re-run `npx gitnexus analyze --embeddings` to regenerate. Check the analyze log for a `Warning: could not load cached embeddings` line — if present, the cache restore failed (corrupt DB / schema mismatch) and the rebuild had nothing to preserve. If you intentionally passed `--drop-embeddings`, this is expected.
- **Why:** Plain `analyze` preserves prior vectors by re-inserting them after the rebuild; the only ways to end up at zero are an explicit `--drop-embeddings`, a cache-load failure (now logged), or a model/dimension change that invalidates the cache. A dirty-recovery run that cannot move the crashed WAL aside now either discards it (logged: forensics lost, embeddings still preserved) or fails fast with a lock error naming the holder — it never silently zeroes embeddings.
- **Why:** Plain `analyze` preserves prior vectors by re-inserting them after the rebuild; ways to end up at zero include an explicit `--drop-embeddings`, a cache-load failure (now logged), or a model/dimension change that invalidates the cache — but zero is no longer the only embedding-loss signature to watch for; see the Sign below for the non-zero, partial-failure case. A dirty-recovery run that cannot move the crashed WAL aside now either discards it (logged: forensics lost, embeddings still preserved) or fails fast with a lock error naming the holder — it never silently zeroes embeddings.
### Analyze finishes but embeddings are incomplete (partial embedding index)
- **Trigger:** `npx gitnexus status` reports `incompleteReasons: ["embedding-checkpoint-pending"]` (or the human-readable "Index incomplete reasons" line); `stats.embeddings` is honest and **non-zero**, and the preceding analyze log showed a `Warning: N node(s) lost their embeddings to embedding-endpoint failures` line (#2790).
- **Do:** Re-run plain `npx gitnexus analyze` — no `--embeddings` flag needed. A retained `embeddingCheckpoint` in the index metadata forces embedding generation for exactly the pending nodes regardless of flags, and clears once they succeed. `--drop-embeddings` abandons the pending nodes instead of retrying them; `--force` also discards the checkpoint (with a warning) and rebuilds without resuming it.
- **Why:** A long analyze run against a flaky HTTP embedding endpoint tolerates bounded sub-batch failures instead of aborting the whole run: it deletes the affected nodes' embedding rows (so they hold zero rows, never a partial set) and records those nodes as pending in `embeddingCheckpoint`. `stats.embeddings` stays an honest, non-zero count of everything that did succeed, so this state never trips the "Embeddings vanished" Sign above — `embedding-checkpoint-pending` is the only reliable signal.
### MCP lists no repos
+119
View File
@@ -117,3 +117,122 @@ repo as never analyzed.
The `meta.json` mirror will remain until a future major version. Removal
will be announced in this file and in the changelog before it happens.
## Ambiguous responses report the true match count (PR #2796, issue #2787)
The MCP symbol resolver returns at most 20 candidate rows. Every ambiguous
response used to take its count from that capped window, so a name with 92
matches (`constructor`, in this repo's own index) reported 20. The same PR
pinned the window with an `ORDER BY`, which turned that undercount from
flaky into stable — and a stable wrong number reads as authoritative.
Three consumer-visible changes follow:
- **`impact`'s `totalCandidates` changed meaning.** It was the length of the
capped 20-row window; it is now the true `COUNT(*)` of matching symbols.
Callers using `totalCandidates === candidates.length` as a "not truncated"
proxy will now see the two diverge. This is a bug fix — the old number was
wrong — but it is still a value change on a published field.
- **`totalCandidates` and `candidatesTruncated` are new on other tools.**
They now also appear on `context`, `trace`, the `explain` / `pdg_query`
block-anchor path, and on `rename` (which returns `context`'s ambiguous
payload verbatim). `candidatesTruncated: true` is present only when
`candidates[]` is shorter than `totalCandidates` — absent otherwise, never
`false`.
- **The `message` template gained a `(showing M)` suffix.** It follows the
total — `Found 92 symbols matching 'constructor' (showing 20). …` — and
appears only when the returned window is smaller than the total. `impact`
uses the longer `(showing M of N)` form.
### Do I need to migrate?
**Only if you read `totalCandidates` or parse `message`.** The last two
changes are purely additive — no field was removed or renamed and
`candidates[]` keeps its shape — so PR #888's "no existing field has changed.
No migration required for `context` callers" still holds for `context`.
- Reading `totalCandidates` on `impact`: it is a true total now. Detect a
shortened window with `candidatesTruncated` (or `totalCandidates >
candidates.length`) rather than by comparing it to an array length.
- Parsing `message` for a count: the total is still the first number, but a
`(showing M)` parenthetical may now follow it. Prefer the structured
`totalCandidates` field over the string.
### What happens on re-index?
Nothing — this is an MCP-surface change only. The graph schema, indexer,
and stored data are untouched.
## `schemaVersion` → `schemaFingerprint` (issue #2798)
The field that decides whether an existing index can be reused changed in
`.gitnexus/gitnexus.json` (and in each `branches/<slug>/gitnexus.json`):
`schemaVersion?: number` has been removed and `schemaFingerprint?: string`
added. The new value is a 12-character digest of the graph DDL this build
creates, so it *describes* the schema an index's tables were actually built
from rather than asserting a number about it.
An absent fingerprint is treated as a mismatch, and that is the whole
backward-compatibility story: every index written by an earlier GitNexus
carries no fingerprint, so it is rebuilt exactly once.
### Do I need to migrate?
**No.** There is nothing to run, edit, or pass. The first `analyze` after
upgrading logs one line —
```
index schema changed (built by an unidentified GitNexus build, this build is <fingerprint>); forcing a full re-analyze so the database is recreated from the current schema.
```
— and then performs that full re-analyze itself. The same run stamps the
fingerprint, and every run after it takes the normal incremental path again.
### What happens on re-index?
One automatic full re-analyze, once per index. Nothing else changes; the
resulting graph is what the current build would have produced anyway.
The scope of that one-time cost is worth knowing before you hit it. It is
per **index**, not per machine or per repository — branch-scoped index slots
(#2106) each keep their own `gitnexus.json`, so every slot pays for itself
the first time it is analyzed after the upgrade. On a very large repository
a full re-analyze is substantial, not a blip; plan the first post-upgrade
run accordingly.
### Why a digest instead of a version number?
`schemaVersion` was hand-incremented, and it had to predict something a
number cannot know: whether the DDL an on-disk database was created from
matches this build's. It collided with `main` eight times, twice *exactly* —
and an exact clash was the quiet failure. Two builds stamp the same number
over different DDL, the strict `===` reuse gate reads the index as current,
the `CREATE … TABLE` statements are skipped as "already exists", and edges
whose endpoint pair the live database cannot persist are dropped. A wrong
graph, with no error anywhere.
A derived digest cannot fail that way: two builds agree exactly when their
DDL agrees, so concurrent branches never need renumbering and a mismatch is
always a real mismatch. The retired ladder's per-version rationale (v2
`BasicBlock.callees` through v35's generated relation cross-product) now
lives only in git history:
`git show 561f913a3:gitnexus/src/storage/repo-manager.ts`.
### What about rollback?
Downgrading to an older GitNexus is safe. The older binary looks for
`schemaVersion`, does not find one, treats the index as pre-versioning, and
forces its own full rebuild — the same one-time cost in the other direction,
never a stale or mismatched graph.
### What if I alternate between an old and a new binary?
Every switch forces a rebuild. The end-of-run metadata is written as a fresh
object literal rather than merged over the previous file, so a new build's
write drops `schemaVersion` and an old build's write drops
`schemaFingerprint` — neither field survives the other's run, and each binary
then finds its own gate unsatisfied. This hits anyone running a pinned
`npx gitnexus@<version>` alongside a local build, or an editor hook still on
an older release. It is a cost, not a correctness problem: each run rebuilds
against its own schema, and the graph it serves is correct for the binary
that produced it. Pin one version per index to avoid the churn.
+9 -1
View File
@@ -56,7 +56,15 @@ npx gitnexus list
npx gitnexus analyze --embeddings
```
**Important:** If you already had embeddings, **always** pass `--embeddings` on later analyzes, or they can be dropped. See `stats.embeddings` in `.gitnexus/gitnexus.json` (or its legacy `meta.json` mirror; 0 means none).
**Important:** If you already had embeddings, a plain `npx gitnexus analyze` **preserves** them (Non-negotiable 5 in [GUARDRAILS.md](GUARDRAILS.md)) — pass `--embeddings` when you also want vectors generated for new or changed nodes, and `--drop-embeddings` only for a deliberate wipe. See `stats.embeddings` in `.gitnexus/gitnexus.json` (or its legacy `meta.json` mirror; 0 means none) — but that figure isn't always freshly measured: if a run's embedding-count query can't answer, it carries the previous run's number forward instead of writing a wrong zero. For a certified read, check `capabilities.vectorSearch.status` instead — it reads `unavailable` (never a stale count) whenever GitNexus can't vouch for the live vector index.
**Partial embedding index (analyze exits 0, but some nodes never got embedded):** A long run against a flaky embedding endpoint can finish successfully while a bounded number of sub-batches still fail. Affected nodes are dropped to zero rows (never left half-written) and recorded as a pending `embeddingCheckpoint`; `npx gitnexus status` then reports `incompleteReasons: ["embedding-checkpoint-pending"]`. Recovery is a plain:
```bash
npx gitnexus analyze
```
No `--embeddings` flag needed — a retained checkpoint forces embedding generation for the pending nodes regardless of flags, and clears once they succeed. `--drop-embeddings` abandons the pending nodes instead of retrying them; `--force` also discards the checkpoint (with a warning) and rebuilds without resuming it.
**Large repos:** Analyze may skip or limit embedding work when node counts are very high; watch CLI output.
@@ -17,22 +17,23 @@ description: "Use when the user wants to know what will break if they change som
## Workflow
```
1. impact({target: "X", direction: "upstream"}) → What depends on this
1. impact({target: "X", direction: "upstream"}) or `node .gitnexus/run.cjs impact "X" --direction upstream --repo .`
2. READ gitnexus://repo/{name}/processes → Check affected execution flows
3. detect_changes() → Map current git changes to affected flows
3. detect_changes({scope: "all"}) or `node .gitnexus/run.cjs detect-changes --scope all --repo .`
4. Assess risk and report to user
```
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
> If `.gitnexus/run.cjs` is missing, replace `node .gitnexus/run.cjs` with `npx gitnexus` in the fallback commands.
## Checklist
```
- [ ] impact({target, direction: "upstream"}) to find dependents
- [ ] impact({target, direction: "upstream"}) or CLI fallback to find dependents
- [ ] Review d=1 items first (these WILL BREAK)
- [ ] Check high-confidence (>0.8) dependencies
- [ ] READ processes to check affected execution flows
- [ ] detect_changes() for pre-commit check
- [ ] detect_changes({scope: "all"}) or CLI fallback for pre-commit check
- [ ] Assess risk level and report to user
```
@@ -55,7 +56,7 @@ description: "Use when the user wants to know what will break if they change som
## Tools
**impact** — the primary tool for symbol blast radius:
**impact** — the primary tool for symbol blast radius. If MCP is unavailable, use `node .gitnexus/run.cjs impact <symbol> --direction upstream --repo .` instead:
```
impact({
@@ -73,10 +74,10 @@ impact({
- authRouter (src/routes/auth.ts:22) [CALLS, 95%]
```
**detect_changes** — git-diff based impact analysis:
**detect_changes** — git-diff based impact analysis. If MCP is unavailable, use `node .gitnexus/run.cjs detect-changes --scope all --repo .` instead:
```
detect_changes({scope: "staged"})
detect_changes({scope: "all"})
→ Changed: 5 symbols in 3 files
→ Affected: LoginFlow, TokenRefresh, APIMiddlewarePipeline
@@ -86,7 +87,7 @@ detect_changes({scope: "staged"})
## Example: "What breaks if I change validateUser?"
```
1. impact({target: "validateUser", direction: "upstream"})
1. impact({target: "validateUser", direction: "upstream"}) or `node .gitnexus/run.cjs impact "validateUser" --direction upstream --repo .`
→ d=1: loginHandler, apiMiddleware (WILL BREAK)
→ d=2: authRouter, sessionManager (LIKELY AFFECTED)
@@ -120,10 +120,17 @@ and do not claim a complete graph-backed review.
review surface: when the diff changes what gets emitted or persisted,
verify every schema/version constant gating caches, incremental
writebacks, and fingerprint baselines was bumped or regenerated — in
GitNexus itself, for example: `INCREMENTAL_SCHEMA_VERSION` (the
incremental write set covers only changed files, so new cross-file edges
never reach an existing index without the bump), the parse-store
`SCHEMA_BUMP`, and both bench fingerprint sets.
GitNexus itself, for example: graph DDL needs no manual bump, because
`SCHEMA_FINGERPRINT` (`gitnexus/src/core/lbug/schema.ts`) is derived
from `NODE_SCHEMA_QUERIES` + `REL_SCHEMA_QUERIES` and moves on its own;
the check there is whether the diff changed any string in those arrays,
and, if it added a new DDL array, whether that array was folded into the
fingerprint. The hand-maintained ritual still applies where no
declarative artifact describes the invalidated set: the parse-store
`SCHEMA_BUMP` and both bench fingerprint sets still need an explicit
bump, re-checked against the base branch right before merge. Semantic
changes that leave the DDL untouched are outside the fingerprint; they
rely on the analyzer runner-identity receipt in the index metadata.
## Expert lenses
@@ -16,22 +16,23 @@ description: Analyze blast radius before making code changes
## Workflow
```
1. impact({target: "X", direction: "upstream"}) → What depends on this
1. impact({target: "X", direction: "upstream"}) or `node .gitnexus/run.cjs impact "X" --direction upstream --repo .`
2. READ gitnexus://repo/{name}/processes → Check affected execution flows
3. detect_changes() → Map current git changes to affected flows
3. detect_changes({scope: "all"}) or `node .gitnexus/run.cjs detect-changes --scope all --repo .`
4. Assess risk and report to user
```
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
> If `.gitnexus/run.cjs` is missing, replace `node .gitnexus/run.cjs` with `npx gitnexus` in the fallback commands.
## Checklist
```
- [ ] impact({target, direction: "upstream"}) to find dependents
- [ ] impact({target, direction: "upstream"}) or CLI fallback to find dependents
- [ ] Review d=1 items first (these WILL BREAK)
- [ ] Check high-confidence (>0.8) dependencies
- [ ] READ processes to check affected execution flows
- [ ] detect_changes() for pre-commit check
- [ ] detect_changes({scope: "all"}) or CLI fallback for pre-commit check
- [ ] Assess risk level and report to user
```
@@ -54,7 +55,7 @@ description: Analyze blast radius before making code changes
## Tools
**impact** — the primary tool for symbol blast radius:
**impact** — the primary tool for symbol blast radius. If MCP is unavailable, use `node .gitnexus/run.cjs impact <symbol> --direction upstream --repo .` instead:
```
impact({
target: "validateUser",
@@ -71,9 +72,9 @@ impact({
- authRouter (src/routes/auth.ts:22) [CALLS, 95%]
```
**detect_changes** — git-diff based impact analysis:
**detect_changes** — git-diff based impact analysis. If MCP is unavailable, use `node .gitnexus/run.cjs detect-changes --scope all --repo .` instead:
```
detect_changes({scope: "staged"})
detect_changes({scope: "all"})
→ Changed: 5 symbols in 3 files
→ Affected: LoginFlow, TokenRefresh, APIMiddlewarePipeline
@@ -83,7 +84,7 @@ detect_changes({scope: "staged"})
## Example: "What breaks if I change validateUser?"
```
1. impact({target: "validateUser", direction: "upstream"})
1. impact({target: "validateUser", direction: "upstream"}) or `node .gitnexus/run.cjs impact "validateUser" --direction upstream --repo .`
→ d=1: loginHandler, apiMiddleware (WILL BREAK)
→ d=2: authRouter, sessionManager (LIKELY AFFECTED)
@@ -120,10 +120,17 @@ and do not claim a complete graph-backed review.
review surface: when the diff changes what gets emitted or persisted,
verify every schema/version constant gating caches, incremental
writebacks, and fingerprint baselines was bumped or regenerated — in
GitNexus itself, for example: `INCREMENTAL_SCHEMA_VERSION` (the
incremental write set covers only changed files, so new cross-file edges
never reach an existing index without the bump), the parse-store
`SCHEMA_BUMP`, and both bench fingerprint sets.
GitNexus itself, for example: graph DDL needs no manual bump, because
`SCHEMA_FINGERPRINT` (`gitnexus/src/core/lbug/schema.ts`) is derived
from `NODE_SCHEMA_QUERIES` + `REL_SCHEMA_QUERIES` and moves on its own;
the check there is whether the diff changed any string in those arrays,
and, if it added a new DDL array, whether that array was folded into the
fingerprint. The hand-maintained ritual still applies where no
declarative artifact describes the invalidated set: the parse-store
`SCHEMA_BUMP` and both bench fingerprint sets still need an explicit
bump, re-checked against the base branch right before merge. Semantic
changes that leave the DDL untouched are outside the fingerprint; they
rely on the analyzer runner-identity receipt in the index metadata.
## Expert lenses
+1
View File
@@ -190,6 +190,7 @@ export {
ResilientFetchExhaustedError,
RETRY_AFTER_CAP_MS,
parseRetryAfter,
isTerminalNetworkError,
} from './integrations/resilient-fetch.js';
export type { ResilientFetchOptions } from './integrations/resilient-fetch.js';
@@ -81,6 +81,25 @@ type Outcome =
| { kind: 'terminal-network'; err: unknown } // TimeoutError or AbortError: no retry, breaker neutral
| { kind: 'retryable-network'; err: unknown }; // DNS, ECONNRESET, etc.
/**
* The network errors `resilientFetch` treats as terminal — never retried, and
* routed through the breaker's neutral path.
*
* Both timer-fired aborts (`AbortSignal.timeout()` → `TimeoutError`) and
* caller-driven aborts (`AbortController.abort()` → `AbortError`) qualify:
* retrying against an already-aborted signal would fail again immediately, and
* neither outcome reflects backend health.
*
* Exported because callers that hook into the retry loop (a `fetchImpl` that
* inspects or re-wraps its own throws) have to agree with {@link
* classifyOutcome} about which errors are terminal. Sharing this predicate is
* what makes that agreement structural instead of a hand-copied condition that
* can drift.
*/
export function isTerminalNetworkError(err: unknown): err is DOMException {
return err instanceof DOMException && (err.name === 'TimeoutError' || err.name === 'AbortError');
}
/** Exported for unit tests. */
export function classifyOutcome(
result: { kind: 'error'; err: unknown } | { kind: 'response'; resp: Response },
@@ -88,15 +107,7 @@ export function classifyOutcome(
retryAfterCapMs = RETRY_AFTER_CAP_MS,
): Outcome {
if (result.kind === 'error') {
// Both timer-fired aborts (`AbortSignal.timeout()` → `TimeoutError`)
// and caller-driven aborts (`AbortController.abort()` → `AbortError`)
// are terminal: retrying against an already-aborted signal would
// fail again immediately, and neither outcome reflects backend
// health. They route through the breaker's neutral path.
if (
result.err instanceof DOMException &&
(result.err.name === 'TimeoutError' || result.err.name === 'AbortError')
) {
if (isTerminalNetworkError(result.err)) {
return { kind: 'terminal-network', err: result.err };
}
return { kind: 'retryable-network', err: result.err };
+100
View File
@@ -0,0 +1,100 @@
# Schema pair-set bench (#2793)
What a bigger `CodeRelation` FROM/TO pair set costs at query time, measured
against a real `@ladybugdb/core` database.
```bash
# from gitnexus/
node --import tsx bench/schema-pairs/measure.mjs # print one JSON line per size + a summary
node --import tsx bench/schema-pairs/measure.mjs --check # gate vs baselines.json
```
## Why it exists
`src/core/lbug/schema.ts` generates its relation pairs from two cross products,
and declines to add a third one **on the strength of a number** — roughly 1.04×
at 450 declared pairs, 1.6× at 786, 2.1× at 1024. That measurement used to live
in a scratch directory, so nobody proposing a third rule could re-run it. This
harness is that measurement, committed — and it reproduces those figures.
Run it before widening a rule, and quote the new ratio in the review.
Observed on the reference box, **four runs** (ratios vs the 332-pair list):
| pairs | untyped | typed (floor) |
| ----- | ---------- | ------------- |
| 332 | 1.00× | 1.00× |
| 450 | 0.93–1.05× | 0.98–1.17× |
| 641 | 1.22–1.43× | 1.11–1.23× |
| 786 | 1.52–1.75× | 1.19–1.31× |
| 1024 | 2.03–2.34× | 1.31–1.57× |
Production's 450 came out _faster_ than 332 on three of the four runs, so at this
size the pair count is inside run-to-run noise. Everything past ~640 is not.
**Quote the range, not a single run** — one run is not evidence here.
## What it measures
For each pair-set size it builds a fresh database with all 32 node tables, a
`CodeRelation` table declaring exactly that many FROM/TO pairs, and **identical
data**, then times two query shapes over 40 anchors × 15 reps (median):
- **`untyped_ms_<size>`** — `MATCH (a {id: $id})-[r:CodeRelation]->(b)`. Neither
endpoint is labelled, so LadybugDB must treat every declared pair as a
candidate. This is the shape `impact`, `context` and `detect_changes` issue
when they walk out from one node id, and the only one whose plan depends on
how many pairs the table declares.
- **`typed_ms_<size>`** — `MATCH (a:Function {…})-[r]->(b:Function)`, the lower
bound. Both endpoints labelled prunes the plan to a single pair, so this was
expected to be flat in the pair count. **It is not** — up to 1.17× at 450 and
1.57× at 1024 — so a declared-but-unused pair costs something even when the
planner never considers it. `typed_ratio_*` is therefore the floor, not a noise
control; the real cost of widening sits between it and `ratio_*`. A run where
`typed_ratio` moves _more_ than `ratio` is noise-dominated and should be
rerun.
- **`ratio_<size>`** — `untyped_ms_<size> / untyped_ms_332`. `ratio_450` is the
figure `schema.ts` quotes.
### Sizes
The pair set is a prefix of a fixed 32×32 (`NODE_TABLES`²) enumeration, so each
size is a strict superset of the smaller ones. The four pairs the synthetic data
uses are pinned to the front, so **the same rows are reachable by the same query
at every size** — the only variable is how many unused pairs are declared. The
harness fails if the row counts ever differ across sizes.
| size | what it is |
| ---- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 332 | the pre-#2792 hand-written list — the reference for every ratio |
| 450 | production today (two cross products + 72 hand-declared pairs) |
| 641 | the third cross product `schema.ts` defers (`DEFINITION_ANCHOR_LABELS × {CodeElement, Section, Typedef, Union, Namespace, Impl, TypeAlias, Static, Template}`), which would leave ~29 hand-declared lines |
| 786 | the size an earlier revision of that comment attributed to the third rule — it is 641; kept as a measured waypoint |
| 1024 | the full cross product, the ceiling |
## Correctness gate
Before timing anything, the harness round-trips the **real** `SCHEMA_QUERIES`
through a real database and asserts that `CALL SHOW_CONNECTION('CodeRelation')`
reports exactly the pairs `parseRelationSchemaPairs` finds in `RELATION_SCHEMA`.
No magic number is baked in: the invariant is that the DDL LadybugDB _accepted_
carries the pair set our own parser believes it declares. The absolute count is
reported as `declared_pairs`. A pair declared twice would not reach this check at
all — LadybugDB rejects the `CREATE REL TABLE` outright, which is why a duplicate
kills every `analyze` rather than one repository's.
## What it does NOT measure
- **Ingest / `COPY` cost.** Pair-set size also multiplies the number of per-pair
CSVs the emitter routes to (`src/core/lbug/rel-pair-routing.ts`); that cost is
covered by `bench/emit-persistence`.
- **At-scale absolute numbers.** Row counts here are small and deliberately
constant. The ratios are the signal; the milliseconds are box-specific.
## Regenerating the baseline
`baselines.json` holds one budget, `ratio_450_budget` — the ceiling on what
production's own pair count may cost relative to the 332-pair hand-list it
replaced. Re-run without `--check` **several times** and copy the top of the
observed `ratio_450` range plus headroom — the spread between runs on this box
is wider than the effect being measured at 450, so a single run cannot set it.
@@ -0,0 +1,4 @@
{
"_comment": "ratio_450_budget — ceiling on what production's 450-pair set may cost on untyped-endpoint anchored queries, relative to the 332-pair hand-list it replaced. Observed 0.94x and 1.05x across two runs on the reference box (i.e. inside run-to-run noise; it came out faster than 332 once). The budget carries headroom for that spread — compare typed_ratio_450 (1.10-1.17x) for this box's floor. Raise it only with a measured range, never a single run.",
"ratio_450_budget": 1.3
}
+357
View File
@@ -0,0 +1,357 @@
/**
* What a bigger `CodeRelation` FROM/TO pair set costs at query time (#2793).
*
* `src/core/lbug/schema.ts` declares its relation pairs from two cross products
* plus a small hand-written remainder, and it justifies NOT adding a third cross
* product with a number: anchored queries cost ~1.04× at 450 declared pairs but
* 1.6× at 786 and 2.1× at 1024. That measurement previously lived in a scratch
* directory, so the claim could not be re-checked when someone proposed
* widening a rule. This is it, committed.
*
* WHAT IT MEASURES. Against a real `@ladybugdb/core` database, with byte-identical
* DATA at every size, it times the query shape whose plan actually depends on the
* declared pair set:
*
* MATCH (a {id: $id})-[r:CodeRelation]->(b) RETURN b.id
*
* Neither endpoint is labelled, so LadybugDB must consider every declared
* FROM/TO pair as a candidate — this is the shape `impact`, `context` and
* `detect_changes` all issue when they walk out from one node id.
*
* A LABEL-typed query (`MATCH (a:Function)-[r]->(b:Function)`) is measured
* alongside it as the LOWER BOUND. Its plan prunes to a single pair, so it was
* expected to be flat in the pair count — it is NOT. Measured here it reaches
* 1.17× at 450 and 1.57× at 1024 against the same 332-pair reference, i.e. a
* declared-but-unused pair costs something even when the planner never
* considers it (per-pair catalog/storage overhead the query pays regardless).
* So `typed_ratio_*` is not a noise control: it is the floor, and the true cost
* of a wider pair set lies between it and `ratio_*`. Treat any run where
* `typed_ratio` moves MORE than `ratio` as noise-dominated.
*
* SIZES. The pair set is a prefix of a fixed 32×32 (`NODE_TABLES`²) enumeration
* so every size is a strict SUPERSET of the smaller ones, and the four pairs the
* data actually uses are pinned first — so the same rows are reachable by the
* same query at every size, and the only variable is how many UNUSED pairs the
* table declares:
* - 332 — the pre-#2792 hand-written list (the historical baseline);
* - 450 — production today (two cross products + 72 hand-declared);
* - 641 — the third cross product schema.ts defers
* (`DEFINITION_ANCHOR_LABELS × {CodeElement, Section, Typedef, Union,
* Namespace, Impl, TypeAlias, Static, Template}`), which would leave
* only ~29 hand-declared lines;
* - 786 — the size an earlier revision of that comment attributed to the
* third rule (it is 641; 786 is kept as a measured waypoint);
* - 1024 — the full cross product, the ceiling.
*
* Ratios are reported against 332, the smallest size — `ratio_450` is the
* number schema.ts quotes.
*
* CORRECTNESS GATE. Before timing anything it round-trips the REAL
* `SCHEMA_QUERIES` through a real database and asserts that
* `CALL SHOW_CONNECTION('CodeRelation')` reports exactly the pairs
* `parseRelationSchemaPairs` finds in `RELATION_SCHEMA`. That is the invariant
* that matters and it needs no magic number: it proves the DDL LadybugDB
* ACCEPTED carries the pair set our own parser believes it declares. (A
* duplicated FROM/TO would not even get this far — LadybugDB rejects the
* `CREATE REL TABLE` outright, which is why that failure kills every `analyze`.)
* The absolute count is reported as `declared_pairs` for the record.
*
* Build-free: imports the `.ts` sources through tsx.
*
* node --import tsx bench/schema-pairs/measure.mjs # print JSON lines
* node --import tsx bench/schema-pairs/measure.mjs --check # gate vs baselines.json
*
* `--check` fails if the correctness gate breaks, or if `ratio_450` exceeds its
* budget — i.e. if production's own pair count starts costing materially more
* than the hand-written list it replaced.
*/
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { fileURLToPath } from 'node:url';
import { NODE_TABLES } from 'gitnexus-shared';
import {
NODE_SCHEMA_QUERIES,
RELATION_SCHEMA,
REL_TABLE_NAME,
} from '../../src/core/lbug/schema.ts';
import { parseRelationSchemaPairs } from '../../src/core/lbug/rel-pair-routing.ts';
const __dirname = path.dirname(fileURLToPath(import.meta.url));
const BASELINE_PATH = path.resolve(__dirname, 'baselines.json');
const lbug = (await import('@ladybugdb/core')).default;
// ---- sizes + the pair enumeration every size is a prefix of ----
const SIZES = [332, 450, 641, 786, 1024];
const REFERENCE_SIZE = 332; // ratios are relative to this
const PRODUCTION_SIZE = 450; // the size schema.ts ships
// The four pairs the synthetic data uses. Pinned to the FRONT of the
// enumeration so they are declared at every size — otherwise a smaller pair set
// would simply carry fewer rows and the comparison would measure data volume,
// not pair-set size.
const DATA_PAIRS = [
['File', 'Function'],
['Function', 'Function'],
['Function', 'Class'],
['Class', 'Method'],
];
const pairKey = ([from, to]) => `${from}|${to}`;
// NODE_TABLES² in declaration order, data pairs first, deduped. 32² = 1024.
const PAIR_UNIVERSE = (() => {
const seen = new Set(DATA_PAIRS.map(pairKey));
const all = [...DATA_PAIRS];
for (const from of NODE_TABLES) {
for (const to of NODE_TABLES) {
const key = `${from}|${to}`;
if (seen.has(key)) continue;
seen.add(key);
all.push([from, to]);
}
}
return all;
})();
if (PAIR_UNIVERSE.length !== NODE_TABLES.length ** 2) {
throw new Error(
`bench: pair universe is ${PAIR_UNIVERSE.length}, expected ${NODE_TABLES.length ** 2} ` +
`(NODE_TABLES changed — update SIZES, the 1024 ceiling is no longer the ceiling)`,
);
}
for (const size of SIZES) {
if (size > PAIR_UNIVERSE.length) {
throw new Error(`bench: size ${size} exceeds the ${PAIR_UNIVERSE.length}-pair universe`);
}
}
const relTableDdlFor = (size) => {
const pairs = PAIR_UNIVERSE.slice(0, size).map(([from, to]) => ` FROM \`${from}\` TO \`${to}\``);
return `CREATE REL TABLE ${REL_TABLE_NAME} (\n${pairs.join(',\n')},\n type STRING,\n confidence DOUBLE,\n reason STRING,\n step INT32\n)`;
};
// ---- synthetic data (identical at every size) ----
const FILES = 20;
const FNS_PER_FILE = 8;
const CLASSES = 40;
const METHODS_PER_CLASS = 4;
const CALLS_PER_FN = 3;
const REPS = 15; // median over reps
const ANCHORS = 40; // distinct anchor ids queried per rep
// Batched with UNWIND rather than one statement per row: per-statement overhead
// dwarfs the insert itself here, and load time is not what this bench measures.
function dataStatements() {
const stmts = [];
const fnIds = [];
const classIds = [];
const methodIds = [];
const fileIds = [];
for (let f = 0; f < FILES; f++) fileIds.push(`file-${f}`);
for (let f = 0; f < FILES; f++) {
for (let i = 0; i < FNS_PER_FILE; i++) fnIds.push(`fn-${f}-${i}`);
}
for (let c = 0; c < CLASSES; c++) {
classIds.push(`cls-${c}`);
for (let m = 0; m < METHODS_PER_CLASS; m++) methodIds.push(`m-${c}-${m}`);
}
const nodeBatch = (label, ids) =>
`UNWIND [${ids.map((id) => `{id: '${id}'}`).join(', ')}] AS r ` +
`CREATE (:\`${label}\` {id: r.id, name: r.id, filePath: 'bench.ts'})`;
stmts.push(nodeBatch('File', fileIds));
stmts.push(nodeBatch('Function', fnIds));
stmts.push(nodeBatch('Class', classIds));
stmts.push(nodeBatch('Method', methodIds));
const relBatch = (fromLabel, toLabel, type, edges) =>
`UNWIND [${edges.map(([f, t]) => `{f: '${f}', t: '${t}'}`).join(', ')}] AS e ` +
`MATCH (a:\`${fromLabel}\` {id: e.f}), (b:\`${toLabel}\` {id: e.t}) ` +
`CREATE (a)-[:${REL_TABLE_NAME} {type: '${type}', confidence: 1.0, reason: 'bench', step: 0}]->(b)`;
const contains = [];
for (let f = 0; f < FILES; f++) {
for (let i = 0; i < FNS_PER_FILE; i++) contains.push([`file-${f}`, `fn-${f}-${i}`]);
}
stmts.push(relBatch('File', 'Function', 'CONTAINS', contains));
// Function→Function calls: each fn calls the next CALLS_PER_FN, wrapping.
const calls = [];
for (let i = 0; i < fnIds.length; i++) {
for (let k = 1; k <= CALLS_PER_FN; k++) calls.push([fnIds[i], fnIds[(i + k) % fnIds.length]]);
}
stmts.push(relBatch('Function', 'Function', 'CALLS', calls));
const uses = fnIds.map((id, i) => [id, classIds[i % classIds.length]]);
stmts.push(relBatch('Function', 'Class', 'USES', uses));
const hasMethod = [];
for (let c = 0; c < CLASSES; c++) {
for (let m = 0; m < METHODS_PER_CLASS; m++) hasMethod.push([`cls-${c}`, `m-${c}-${m}`]);
}
stmts.push(relBatch('Class', 'Method', 'HAS_METHOD', hasMethod));
// Anchors: functions, which have out-edges on two distinct declared pairs.
return { stmts, anchors: fnIds.slice(0, ANCHORS) };
}
const { stmts: DATA_STATEMENTS, anchors: ANCHOR_IDS } = dataStatements();
// ---- timing ----
const median = (xs) => {
const s = [...xs].sort((a, b) => a - b);
const m = Math.floor(s.length / 2);
return s.length % 2 ? s[m] : (s[m - 1] + s[m]) / 2;
};
const withDb = async (fn) => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gitnexus-bench-pairs-'));
const db = new lbug.Database(path.join(dir, 'db'));
const conn = new lbug.Connection(db);
try {
return await fn(conn);
} finally {
await conn.close().catch(() => {});
await db.close?.().catch?.(() => {});
fs.rmSync(dir, { recursive: true, force: true });
}
};
async function runAll(conn, statements) {
for (const s of statements) await conn.query(s);
}
// The measured shape: BOTH endpoints untyped, anchored by id. LadybugDB must
// consider every declared FROM/TO pair as a candidate.
const UNTYPED_QUERY = (id) =>
`MATCH (a {id: '${id}'})-[r:${REL_TABLE_NAME}]->(b) RETURN b.id AS id, r.type AS type`;
// The lower bound: both endpoints labelled, so the planner prunes to one pair.
// Still not flat in the pair count (see the header) — an unused declared pair
// costs something even when the plan never touches it.
const TYPED_QUERY = (id) =>
`MATCH (a:Function {id: '${id}'})-[r:${REL_TABLE_NAME}]->(b:Function) RETURN b.id AS id`;
async function timeQueries(conn, build) {
// Warm: run the whole anchor sweep once uncounted (plan cache + page cache).
for (const id of ANCHOR_IDS) await (await conn.query(build(id))).getAll();
const samples = [];
let rows = 0;
for (let rep = 0; rep < REPS; rep++) {
const start = process.hrtime.bigint();
let n = 0;
for (const id of ANCHOR_IDS) n += (await (await conn.query(build(id))).getAll()).length;
samples.push(Number(process.hrtime.bigint() - start) / 1e6);
rows = n;
}
return { ms: median(samples), rows };
}
async function measureSize(size) {
return withDb(async (conn) => {
for (const q of NODE_SCHEMA_QUERIES) await conn.query(q);
await conn.query(relTableDdlFor(size));
await runAll(conn, DATA_STATEMENTS);
const untyped = await timeQueries(conn, UNTYPED_QUERY);
const typed = await timeQueries(conn, TYPED_QUERY);
return {
pairs: size,
untyped_ms: Number(untyped.ms.toFixed(3)),
untyped_rows: untyped.rows,
typed_ms: Number(typed.ms.toFixed(3)),
typed_rows: typed.rows,
};
});
}
// ---- correctness gate: the REAL schema, round-tripped ----
async function verifyRealSchema() {
return withDb(async (conn) => {
for (const q of NODE_SCHEMA_QUERIES) await conn.query(q);
// If RELATION_SCHEMA declared a pair twice, LadybugDB rejects this outright
// — the failure mode that kills every `analyze`, not just one repo's.
await conn.query(RELATION_SCHEMA);
const res = await conn.query(`CALL SHOW_CONNECTION('${REL_TABLE_NAME}') RETURN *`);
const rows = await res.getAll();
const actual = new Set(
rows.map(
(r) =>
`${r['source table name'] ?? r.source}|${r['destination table name'] ?? r.destination}`,
),
);
const expected = parseRelationSchemaPairs(RELATION_SCHEMA);
const missing = [...expected].filter((p) => !actual.has(p)).sort();
const extra = [...actual].filter((p) => !expected.has(p)).sort();
return { declared_pairs: expected.size, db_pairs: actual.size, missing, extra };
});
}
// ---- run ----
const CHECK = process.argv.includes('--check');
const failures = [];
const verified = await verifyRealSchema();
if (verified.missing.length > 0 || verified.extra.length > 0) {
failures.push(
`RELATION_SCHEMA round-trip mismatch: ${verified.missing.length} pair(s) parsed but absent ` +
`from SHOW_CONNECTION (${verified.missing.slice(0, 5).join(', ')}), ${verified.extra.length} ` +
`present in the DB but unparsed (${verified.extra.slice(0, 5).join(', ')})`,
);
}
const results = [];
for (const size of SIZES) results.push(await measureSize(size));
const reference = results.find((r) => r.pairs === REFERENCE_SIZE);
const summary = {
...verified,
missing: undefined,
extra: undefined,
reference_pairs: REFERENCE_SIZE,
};
for (const r of results) {
summary[`untyped_ms_${r.pairs}`] = r.untyped_ms;
summary[`typed_ms_${r.pairs}`] = r.typed_ms;
summary[`ratio_${r.pairs}`] = Number((r.untyped_ms / reference.untyped_ms).toFixed(3));
summary[`typed_ratio_${r.pairs}`] = Number((r.typed_ms / reference.typed_ms).toFixed(3));
}
// Row counts must be identical at every size — otherwise the sizes are not
// carrying the same data and the ratios mean nothing.
const rowShapes = new Set(results.map((r) => `${r.untyped_rows}/${r.typed_rows}`));
if (rowShapes.size !== 1) {
failures.push(
`row counts differ across pair-set sizes (${[...rowShapes].join(' vs ')}) — the data pins ` +
`in DATA_PAIRS are not holding, so the ratios compare different graphs`,
);
}
if (!CHECK) {
for (const r of results) process.stdout.write(JSON.stringify(r) + '\n');
process.stdout.write(JSON.stringify(summary) + '\n');
} else {
const baselines = JSON.parse(fs.readFileSync(BASELINE_PATH, 'utf8'));
const budget = baselines[`ratio_${PRODUCTION_SIZE}_budget`];
if (budget !== undefined && summary[`ratio_${PRODUCTION_SIZE}`] >= budget) {
failures.push(
`production pair set (${PRODUCTION_SIZE}) costs ${summary[`ratio_${PRODUCTION_SIZE}`]}× vs ` +
`${REFERENCE_SIZE} pairs, >= budget ${budget} (untyped ${reference.untyped_ms}ms -> ` +
`${summary[`untyped_ms_${PRODUCTION_SIZE}`]}ms; typed control ` +
`${summary[`typed_ratio_${PRODUCTION_SIZE}`]}×)`,
);
}
process.stdout.write(JSON.stringify(summary) + '\n');
}
if (failures.length > 0) {
for (const f of failures) process.stderr.write(`[schema-pairs] FAIL: ${f}\n`);
process.exit(1);
}
if (CHECK) process.stderr.write(`[schema-pairs --check] PASS (${results.length} sizes)\n`);
+3 -2
View File
@@ -14,10 +14,11 @@
"_rebaselined_2766_callee_position_marker": "#2766 review fix: a call's callee selector is no longer DROPPED at capture. An earlier commit on this branch dropped it outright, which also deleted the genuine field read on a func-typed struct field (`h.dep.Work()` where `Work func() error`) - callback/hook/mock structs lost their only ACCESSES evidence. The match is now emitted carrying `@reference.callee-position`, and the phantom is suppressed at EMIT by the resolved target's kind instead. Go only: the other 14 languages' fingerprints are byte-identical, which is the check that this is not a cross-language capture change. Prior 7bb524a32a2eed57a15b454e3a33480e92a496c683e6856ef02179693c0e02e3 -> e47302079e17a5e73711bbed5416557b49327cb67e4932008700ec6b8fb468b3; scaling 1.001 < 1.5; fixtures 102 (unchanged), capture_groups_fp 2103."
},
"cobol": {
"fingerprint": "d45bb091b0893d0de4fae2486b31ba21719c9377bf35a0908fd3a36fa1c3bf4e",
"fingerprint": "c8c00b56a7da24e04080eb885714fbbf45e3903324f0cf9df0754f5b5a92e3aa",
"scaling_budget": 1.5,
"_rebaselined_callable_flow_2522_followup": "PR #2522 follow-up: COBOL procedure-pointer callable flow facts; multi-topic extraction now consumes each grouped scope/declaration match once instead of requiring a duplicate declaration-only match. Prior 68ee0e95eb9f86f2d92ca35f730f4c2d4d83abc1b5241ae767ff3437780ec8d1 -> d45bb091b0893d0de4fae2486b31ba21719c9377bf35a0908fd3a36fa1c3bf4e; scaling 0.853 < 1.5.",
"_note": "Updated for F17-F23 fixes (P2: TIMES guard, ADD GIVING, SQL AS alias). See PR #1959."
"_note": "Updated for F17-F23 fixes (P2: TIMES guard, ADD GIVING, SQL AS alias). See PR #1959.",
"_rebaselined_2793_declaratives": "PR #2793: corpus-only re-baseline. `cobol-declaratives` was added to test/fixtures/lang-resolution to reproduce the `Namespace\u2192Record` analyze abort (DECLARATIVES / USE AFTER STANDARD ERROR ON <file>), and this bench globs `lang-resolution/cobol-*`, so the corpus grew 14 -> 15 files. Verified capture-neutral: with that one fixture moved aside the fingerprint is byte-identical to the prior d45bb091b0893d0de4fae2486b31ba21719c9377bf35a0908fd3a36fa1c3bf4e. No COBOL capture code changed in that PR. Scaling 0.677 < 1.5."
},
"c": {
"fingerprint": "3418cded9f7072152f68992f0a426f43ae7d9d553579a47075fc0cab185848a5",
+136
View File
@@ -0,0 +1,136 @@
/**
* Weight-aware partitioning for the cross-platform test matrix.
*
* WHY THIS EXISTS. `run-cross-platform.ts` used to hand vitest the whole file
* list plus `--shard=i/n`, and vitest partitions by file COUNT. Runtime on this
* suite is wildly uneven — measured on the Windows runner, `cli-e2e` is 361 s
* and `worker-pool` 221 s, while most files are under a second — so a
* count-split routinely put several of the heaviest suites on one shard. That
* is #2449, and this file's sibling header has documented the symptom ("the
* heaviest spawn suites can cluster on one shard") since the watchdog was first
* raised from 15 to 20 minutes.
*
* It went from a latent hazard to a red matrix when three CHEAP files (the
* `dist/` module-load closure guards: 448 ms, 53 ms, sub-second) were added to
* `SPAWN_CLI`. They cost nothing to run, but a count-split re-partitions on
* every insertion, and the reshuffle happened to land `cli-e2e` + `cli-limit-e2e`
* + `analyze-heap-oom-e2e` together on shard 1/3 — 32 files against 26 and 29 —
* which blew the 20-minute budget with four files still queued. Nothing about
* the added files caused it; they were simply the perturbation.
*
* So the split is done HERE, by weight, and only the chosen shard's files are
* handed to vitest. Two properties follow, and both are pinned in
* `test/unit/cross-platform-shard.test.ts`:
*
* - the heaviest suites are spread across shards by construction, so the
* busiest shard tracks the ideal rather than the luck of the sort order;
* - adding or removing a CHEAP file cannot move a heavy one, so registering a
* new platform-sensitive test is no longer a CI-stability gamble. That is the
* property whose absence caused this.
*/
/**
* Measured wall-clock on the WINDOWS runner (the slowest platform, so it is the
* one that decides the budget), in seconds, from the last fully-green matrix run
* plus the timed files of the run that failed.
*
* Only files heavy enough to matter are listed; everything else is carried by
* {@link PER_FILE_OVERHEAD_SEC} alone. These are load-balancing hints, NOT
* assertions — no
* test asserts a runtime, and drift only makes the split slightly less even, so
* a stale entry is harmless and refreshing them is optional. Deliberately not
* auto-generated: a committed table is reviewable and works offline, and the
* alternative (timing files at CI runtime to decide the split) would make the
* partition depend on the very machine load it is trying to protect against.
*/
export const WINDOWS_WEIGHTS_SEC: Readonly<Record<string, number>> = {
'test/integration/cli-e2e.test.ts': 361,
'test/integration/worker-pool.test.ts': 222,
'test/unit/incremental-vector-extension-ordering.test.ts': 87,
'test/integration/cli-limit-e2e.test.ts': 75,
'test/unit/hooks.test.ts': 26,
'test/integration/analyze-heap-oom-e2e.test.ts': 23,
'test/unit/git-utils.test.ts': 18,
'test/integration/hooks-e2e.test.ts': 15,
'test/integration/tree-sitter-languages.test.ts': 9,
'test/unit/repo-manager.test.ts': 9,
'test/unit/detect-changes-worktree.test.ts': 9,
'test/integration/antigravity-hook-e2e.test.ts': 7,
'test/unit/index-lock.test.ts': 5,
'test/unit/setup.test.ts': 5,
};
/**
* Fixed cost every file pays regardless of what it asserts: a pool worker start,
* module graph evaluation, and (for most of this list) a native addon load.
*
* Added to EVERY file's weight, not just unmeasured ones, and that is the point.
* Calibrated against the last green Windows matrix: its busiest shard ran 736 s
* of wall clock over ~511 s of measured file time, so roughly 8 s per file is
* unattributed setup. Without this term the balancer treats a light file as
* nearly free and, having isolated the two monsters, piles every remaining file
* onto the other shards — trading a runtime imbalance for a file-count one that
* costs just as much. With it, the split balances runtime AND count together.
*/
const PER_FILE_OVERHEAD_SEC = 8;
/**
* Scheduling weight for `file`: its measured runtime (0 if it was fast enough
* that vitest printed no duration) plus the per-file floor above.
*/
export function weightOf(file: string): number {
return (WINDOWS_WEIGHTS_SEC[file] ?? 0) + PER_FILE_OVERHEAD_SEC;
}
/**
* Partition `files` into `total` shards and return the 1-based `index` one.
*
* Longest-processing-time first: sort by weight descending, then repeatedly give
* the next file to the lightest shard so far. LPT is the standard greedy for
* multiprocessor scheduling and is guaranteed within 4/3 of optimal — far more
* than enough here, where the goal is only "no shard gets two monsters".
*
* Ties break on the file path so the partition is DETERMINISTIC: every shard
* computes the same split independently, on a different machine, with no
* coordination — which is what lets each runner select its own slice.
*
* Returns files in the input list's original order, not weight order, so failure
* output and reruns stay readable.
*/
export function shardFiles(
files: readonly string[],
index: number,
total: number,
): readonly string[] {
if (!Number.isInteger(total) || total < 1) {
throw new Error(`shard total must be a positive integer, got ${total}`);
}
if (!Number.isInteger(index) || index < 1 || index > total) {
throw new Error(`shard index must be in 1..${total}, got ${index}`);
}
if (total === 1) return [...files];
const byWeightDesc = [...files].sort((a, b) => {
const diff = weightOf(b) - weightOf(a);
return diff !== 0 ? diff : a.localeCompare(b);
});
const loads = Array.from({ length: total }, () => 0);
const assigned = Array.from({ length: total }, () => new Set<string>());
for (const file of byWeightDesc) {
let lightest = 0;
for (let i = 1; i < total; i++) {
if (loads[i]! < loads[lightest]!) lightest = i;
}
assigned[lightest]!.add(file);
loads[lightest]! += weightOf(file);
}
const mine = assigned[index - 1]!;
return files.filter((f) => mine.has(f));
}
/** Total weight of a file set — the shard cost this balancer is minimising. */
export function shardWeight(files: readonly string[]): number {
return files.reduce((sum, f) => sum + weightOf(f), 0);
}
+27
View File
@@ -186,6 +186,33 @@ const SPAWN_CLI = [
// exposed a file-backend double-admit race here (#2658 review); the reclaim is
// now judgment-verified so a live holder is never displaced.
'test/integration/analyze-index-lock-concurrency.test.ts',
// The three `dist/` module-load closure guards, all built on the shared
// child-process probe in `test/helpers/module-load-probe.ts`. That probe IS
// the platform-varying part: it spawns `process.execPath` in array form,
// clears NODE_OPTIONS, addresses its target via `pathToFileURL` (Windows needs
// the `file:///C:/...` form — a bare absolute path is not a valid ESM
// specifier there), and renders every result through a `path.sep`→POSIX
// normalisation the anchors and offender regexes depend on. None of that is
// proven anywhere else.
//
// Cheap: measured on the Windows runner at 448 ms, 53 ms and sub-second. An
// earlier attempt to register them still turned the matrix red — not from
// their own cost, but because vitest sharded by file COUNT, so inserting any
// file re-partitioned the list and happened to cluster `cli-e2e` (361 s) with
// `cli-limit-e2e` (75 s) on one shard. The split is weight-aware now
// (`scripts/cross-platform-shard.ts`), so a cheap file can no longer move a
// heavy one.
//
// #2802: MCP startup must not eagerly load the analyze-only language
// provider registry or the group contract extractors.
'test/integration/mcp/startup-language-closure.test.ts',
// PR #1383: `cli/mcp.js`'s static-import closure must stay leaf-only so no
// native binding initialises before the stdout sentinel installs.
'test/integration/mcp/import-closure.test.ts',
// #2091/#2093/#2116: the scope-resolution registry must not load the optional
// tree-sitter grammars at import time. The offender regexes match grammar
// paths with either separator, which only the Windows runner proves.
'test/integration/optional-grammars/registry-import-closure.test.ts',
];
// Worker threads tests — exercise real worker_threads which have
+16 -5
View File
@@ -15,6 +15,7 @@ import path from 'path';
import { fileURLToPath } from 'url';
import { ALL_CROSS_PLATFORM } from './cross-platform-tests.js';
import { parseShardArg } from './shard-arg.js';
import { shardFiles, shardWeight } from './cross-platform-shard.js';
const __dirname = path.dirname(fileURLToPath(import.meta.url));
const ROOT = path.resolve(__dirname, '..');
@@ -29,8 +30,10 @@ if (missing.length > 0) {
}
// Optional sharding (CI): `--shard=<i>/<n>` splits the fixed file list across
// parallel matrix shards so each runner processes ~1/n of it. Passed straight
// through to vitest, which partitions the *given* files deterministically. The
// parallel matrix shards. The split is computed HERE, by measured weight, and
// only this shard's files are handed to vitest — it is NOT passed through,
// because vitest partitions by file COUNT and this suite's runtimes span three
// orders of magnitude (see cross-platform-shard.ts). The
// Windows runner is ~5x slower than macOS/Linux on this spawn-heavy suite (~50
// CLI/worker process spawns), so a single shard was creeping past the watchdog
// below; sharding keeps each runner well under it (see ci-tests.yml matrix).
@@ -63,14 +66,22 @@ const timeoutMs =
? timeoutMinutes * 60 * 1000
: DEFAULT_TIMEOUT_MIN * 60 * 1000;
// Resolve the shard to an explicit file list. `--shard=i/n` is consumed here,
// never forwarded: forwarding it as well would re-partition this slice a second
// time and silently drop most of it.
const shardParts = shardArg?.replace('--shard=', '').split('/');
const shardIndex = shardParts ? Number(shardParts[0]) : 1;
const shardTotal = shardParts ? Number(shardParts[1]) : 1;
const files = shardFiles(ALL_CROSS_PLATFORM, shardIndex, shardTotal);
console.log(
`Running ${ALL_CROSS_PLATFORM.length} platform-sensitive tests` +
`${shardArg ? ` (${shardArg.replace('--shard=', 'shard ')})` : ''}...\n`,
`Running ${files.length} of ${ALL_CROSS_PLATFORM.length} platform-sensitive tests` +
`${shardArg ? ` (shard ${shardIndex}/${shardTotal}, ~${shardWeight(files)}s measured weight)` : ''}...\n`,
);
const startedAt = Date.now();
try {
execFileSync('npx', ['vitest', 'run', ...ALL_CROSS_PLATFORM, ...(shardArg ? [shardArg] : [])], {
execFileSync('npx', ['vitest', 'run', ...files], {
cwd: ROOT,
stdio: 'inherit',
timeout: timeoutMs,
+9 -8
View File
@@ -17,22 +17,23 @@ description: "Use when the user wants to know what will break if they change som
## Workflow
```
1. impact({target: "X", direction: "upstream"}) → What depends on this
1. impact({target: "X", direction: "upstream"}) or `node .gitnexus/run.cjs impact "X" --direction upstream --repo .`
2. READ gitnexus://repo/{name}/processes → Check affected execution flows
3. detect_changes() → Map current git changes to affected flows
3. detect_changes({scope: "all"}) or `node .gitnexus/run.cjs detect-changes --scope all --repo .`
4. Assess risk and report to user
```
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
> If `.gitnexus/run.cjs` is missing, replace `node .gitnexus/run.cjs` with `npx gitnexus` in the fallback commands.
## Checklist
```
- [ ] impact({target, direction: "upstream"}) to find dependents
- [ ] impact({target, direction: "upstream"}) or CLI fallback to find dependents
- [ ] Review d=1 items first (these WILL BREAK)
- [ ] Check high-confidence (>0.8) dependencies
- [ ] READ processes to check affected execution flows
- [ ] detect_changes() for pre-commit check
- [ ] detect_changes({scope: "all"}) or CLI fallback for pre-commit check
- [ ] Assess risk level and report to user
```
@@ -55,7 +56,7 @@ description: "Use when the user wants to know what will break if they change som
## Tools
**impact** — the primary tool for symbol blast radius:
**impact** — the primary tool for symbol blast radius. If MCP is unavailable, use `node .gitnexus/run.cjs impact <symbol> --direction upstream --repo .` instead:
```
impact({
@@ -73,10 +74,10 @@ impact({
- authRouter (src/routes/auth.ts:22) [CALLS, 95%]
```
**detect_changes** — git-diff based impact analysis:
**detect_changes** — git-diff based impact analysis. If MCP is unavailable, use `node .gitnexus/run.cjs detect-changes --scope all --repo .` instead:
```
detect_changes({scope: "staged"})
detect_changes({scope: "all"})
→ Changed: 5 symbols in 3 files
→ Affected: LoginFlow, TokenRefresh, APIMiddlewarePipeline
@@ -86,7 +87,7 @@ detect_changes({scope: "staged"})
## Example: "What breaks if I change validateUser?"
```
1. impact({target: "validateUser", direction: "upstream"})
1. impact({target: "validateUser", direction: "upstream"}) or `node .gitnexus/run.cjs impact "validateUser" --direction upstream --repo .`
→ d=1: loginHandler, apiMiddleware (WILL BREAK)
→ d=2: authRouter, sessionManager (LIKELY AFFECTED)
+11 -4
View File
@@ -120,10 +120,17 @@ and do not claim a complete graph-backed review.
review surface: when the diff changes what gets emitted or persisted,
verify every schema/version constant gating caches, incremental
writebacks, and fingerprint baselines was bumped or regenerated — in
GitNexus itself, for example: `INCREMENTAL_SCHEMA_VERSION` (the
incremental write set covers only changed files, so new cross-file edges
never reach an existing index without the bump), the parse-store
`SCHEMA_BUMP`, and both bench fingerprint sets.
GitNexus itself, for example: graph DDL needs no manual bump, because
`SCHEMA_FINGERPRINT` (`gitnexus/src/core/lbug/schema.ts`) is derived
from `NODE_SCHEMA_QUERIES` + `REL_SCHEMA_QUERIES` and moves on its own;
the check there is whether the diff changed any string in those arrays,
and, if it added a new DDL array, whether that array was folded into the
fingerprint. The hand-maintained ritual still applies where no
declarative artifact describes the invalidated set: the parse-store
`SCHEMA_BUMP` and both bench fingerprint sets still need an explicit
bump, re-checked against the base branch right before merge. Semantic
changes that leave the DDL untouched are outside the fingerprint; they
rely on the analyzer runner-identity receipt in the index metadata.
## Expert lenses
+6 -6
View File
@@ -191,18 +191,18 @@ ${tableBody}`
return `${GITNEXUS_START_MARKER}
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **${projectName}**${noStats ? '' : ` (${stats.nodes || 0} symbols, ${stats.edges || 0} relationships, ${stats.processes || 0} execution flows)`}. Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **${projectName}**${noStats ? '' : ` (${stats.nodes || 0} symbols, ${stats.edges || 0} relationships, ${stats.processes || 0} execution flows)`}. Use GitNexus graph tools to understand code, assess impact, and navigate safely.
> Index stale? Run \`${runner} analyze\` from the project root — it auto-selects an available runner. ${bootstrapNote}
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run \`impact({target: "symbolName", direction: "upstream"})\` and report the blast radius (direct callers, affected processes, risk level) to the user.${
- **MUST run impact analysis before editing.** Use \`impact({target: "symbolName", direction: "upstream"})\` (MCP) or \`${runner} impact "symbolName" --direction upstream --repo .\` (CLI fallback); report callers, processes, and risk. Never substitute grep for graph analysis.${
hasPdg
? ` For unified PDG impact, add \`mode: "pdg"\` with optional \`line: <N>\` — it returns statement-level \`affectedStatements\` over CDG + REACHING_DEF and inter-procedural symbols in \`interproceduralByDepth\`/\`byDepth\`; no-layer/degraded PDG results are UNKNOWN-risk notes (\`--pdg\` layer).`
? ` For unified PDG impact, add \`mode: "pdg"\` with optional \`line: <N>\` — it returns statement-level \`affectedStatements\` over CDG + REACHING_DEF and inter-procedural symbols in \`interproceduralByDepth\`/\`byDepth\`; no-layer/degraded PDG results are UNKNOWN-risk notes (\`--pdg\` layer). CLI equivalent: \`${runner} impact "symbolName" --direction upstream --mode pdg --line <N> --repo .\`.`
: ''
}
- **MUST run \`detect_changes()\` before committing** to verify your changes only affect expected symbols and execution flows. For regression review, compare against the default branch: \`detect_changes({scope: "compare", base_ref: ${JSON.stringify(markdownSafeBranch(defaultBranch))}})\`.
- **MUST analyze graph changes before committing.** Use \`detect_changes({scope: "all"})\` (MCP) or \`${runner} detect-changes --scope all --repo .\` (CLI fallback). For regression review: \`detect_changes({scope: "compare", base_ref: ${JSON.stringify(markdownSafeBranch(defaultBranch))}})\` or \`${runner} detect-changes --scope compare --base-ref ${JSON.stringify(markdownSafeBranch(defaultBranch))} --repo .\`.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use \`query({search_query: "concept"})\` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use \`context({name: "symbolName"})\`.
@@ -214,10 +214,10 @@ This project is indexed by GitNexus as **${projectName}**${noStats ? '' : ` (${s
## Never Do
- NEVER edit a function, class, or method without first running \`impact\` on it.
- NEVER edit a function, class, or method before MCP/CLI impact analysis.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use \`rename\` which understands the call graph.
- NEVER commit changes without running \`detect_changes()\` to check affected scope.
- NEVER commit before MCP/CLI graph change analysis.
## Resources
+38 -8
View File
@@ -15,6 +15,8 @@ import v8 from 'v8';
import cliProgress from 'cli-progress';
import { isLbugReady, LbugWipeError } from '../core/lbug/lbug-adapter.js';
import { boundedCheckpointBeforeExit } from '../core/lbug/shutdown-helpers.js';
import { findUndeclaredRelationPairError } from '../core/lbug/rel-pair-routing.js';
import { causeChain } from '../lib/utils.js';
import {
getOsPageSize,
isLbugCheckpointIoError,
@@ -101,15 +103,13 @@ const writeFatalToStderr = (label: string, err: unknown): void => {
// #2068) is only reachable via `.cause`. Without this the user sees the
// wrapper's main-thread stack and never the real frame. `cause.stack` already
// begins with the cause's message, so we print the stack alone (not message +
// stack) to avoid repeating it. Depth-bounded so a cyclic `cause` can't loop
// (the phase runner wraps one level; the bound leaves headroom for future
// nesting); uses realStderrWrite so the redirected console.error's ANSI
// clear-line wrapping can't erase it (#1169).
const MAX_CAUSE_DEPTH = 5;
let cause: unknown = isErr ? (err as { cause?: unknown }).cause : undefined;
for (let depth = 0; depth < MAX_CAUSE_DEPTH && cause instanceof Error; depth++) {
// stack) to avoid repeating it. `causeChain` owns the traversal and the depth
// bound that stops a cyclic `cause` looping — this used to be one of four
// hand-rolled copies that had already drifted apart on both. Uses
// realStderrWrite so the redirected console.error's ANSI clear-line wrapping
// can't erase it (#1169). The head is skipped: it was just printed above.
for (const cause of causeChain(isErr ? (err as { cause?: unknown }).cause : undefined)) {
realStderrWrite(`\n Caused by: ${cause.stack ?? cause.message}\n`);
cause = (cause as { cause?: unknown }).cause;
}
};
@@ -1717,6 +1717,36 @@ const analyzeCommandImpl = async (
return;
}
// An extracted edge whose FROM→TO label pair is missing from GitNexus's own
// relation DDL (#2789). `assertDeclaredPair` aborts the run rather than let
// the bulk COPY fail late and silently drop the edge, so the user sees a
// mid-run crash inside GitNexus internals with nothing to act on. Name the
// pair, the relationship and the file that produced it, and say plainly that
// a re-run cannot help — this is deterministic for the same input.
// Checked by TYPE (repo norm, #2385) BEFORE the message-text heuristics
// below, and through the `cause` chain because the ingestion phase runner
// rewraps every phase failure as `Phase 'X' failed: …`.
const undeclaredPair = findUndeclaredRelationPairError(err);
if (undeclaredPair !== undefined) {
// Render the error's OWN message indented — same idiom as the
// `LbugWipeError` and page-size branches below. `UndeclaredRelationPairError`
// builds a fully self-contained message (pair, relationship type, both node
// ids, source file, issue URL, `.gitnexusignore` workaround) precisely
// because `gitnexus serve` forwards only `err.message` over worker IPC.
// Re-rendering those fields here would be a second copy of one string, free
// to drift from the first — and the actionable half would reach CLI users
// only. `undeclaredPair.message`, not the outer `msg`: the real error may be
// several `cause` levels below the phase wrapper `msg` came from.
cliError(` ${undeclaredPair.message.replace(/\n/g, '\n ')}\n`, {
recoveryHint: 'undeclared-relation-pair',
labelPair: undeclaredPair.pairKey,
relationType: undeclaredPair.relationType,
sourceFile: undeclaredPair.sourceFile,
});
process.exitCode = 1;
return;
}
// WAL corruption — the index file is unreadable. Give a clear recovery
// path without a confusing stack trace (the native error message alone
// is enough signal).
+2 -1
View File
@@ -60,7 +60,8 @@ export type RecoveryHint =
| 'module-not-found'
| 'gitnexusrc-invalid'
| 'default-branch-invalid'
| 'index-lock-timeout';
| 'index-lock-timeout'
| 'undeclared-relation-pair';
/**
* Common shape for the optional structured-field bag passed to
+25 -1
View File
@@ -250,19 +250,43 @@ export function registerGroupCommands(program: Command): void {
} else {
const summary = (raw as { summary?: Record<string, number> })?.summary;
const risk = (raw as { risk?: string })?.risk;
// A truncated fan-out under-reports risk (mergeRisk only grows with
// traversed crossings), and the default human output used to print a
// bare `risk=` indistinguishable from a complete run — the JSON
// already carried `truncated`, but nobody reading the terminal saw it.
const riskFloor =
(raw as { riskEpistemic?: string })?.riskEpistemic === 'lower-bound' ? '+' : '';
const boundaryOnly =
(
raw as {
cross?: Array<{ fanout_status?: string }>;
}
)?.cross?.filter((entry) => entry.fanout_status === 'not_attempted').length ?? 0;
console.log(`Group impact for "${name}" (${String(opts.repo)}): risk=${risk ?? '?'}`);
console.log(
`Group impact for "${name}" (${String(opts.repo)}): risk=${risk ?? '?'}${riskFloor}`,
);
if (summary) {
const boundaryNote = boundaryOnly > 0 ? ` (${boundaryOnly} boundary-only)` : '';
console.log(
` direct=${summary.direct ?? 0} processes=${summary.processes_affected ?? 0} cross=${summary.cross_repo_hits ?? 0}${boundaryNote}`,
);
}
if (riskFloor) {
// `truncated` has two independent causes that point at different
// subsystems, so the note must name the one that actually fired:
// dropped crossings, or a local walk that never finished (most
// often the impact chunk cap, which any symbol with more than a
// thousand locally-impacted nodes hits on every run). `dropped` is
// deduped to distinct repos before it reaches here, so it counts
// repos — reporting it as crossings understates a fan-out cap the
// same way #2787's totals did.
const dropped = (raw as { truncatedRepos?: string[] })?.truncatedRepos ?? [];
console.log(
dropped.length > 0
? ` risk is a LOWER BOUND — fan-out stopped early; crossings to ${dropped.length} repo(s) not traversed: ${dropped.join(', ')}`
: ' risk is a LOWER BOUND — the local impact walk did not complete (every bridge crossing was traversed)',
);
}
}
} finally {
await backend.dispose().catch(() => {});
+348 -42
View File
@@ -82,6 +82,7 @@ type PackageManifest = {
dependencies?: Record<string, unknown>;
optionalDependencies?: Record<string, unknown>;
peerDependencies?: Record<string, unknown>;
devDependencies?: Record<string, unknown>;
};
type StatState = {
@@ -100,6 +101,16 @@ type ReadableFileState = {
symlinkTarget?: string;
};
/**
* State of a symbolic link recorded by its link text rather than by its
* target's payload. There is deliberately no `target` stat: the whole point of
* this shape is that the link was never resolved (see {@link RuntimeArtifact}).
*/
type SymlinkArtifactState = {
link: StatState;
symlinkTarget: string;
};
type RuntimePackage = {
root: string;
locator: string;
@@ -137,11 +148,28 @@ type RuntimeArtifactScanBudget = {
edges: number;
};
type RuntimeArtifact = {
absolutePath: string;
canonicalPath: string;
kind: 'file' | 'symlink';
};
/**
* One runtime payload input.
*
* `file` and `symlink` contribute their target's CONTENT digest; `symlink` also
* carries its link text, so both a retarget and a byte change move the receipt.
*
* `unfollowed-symlink` is a symbolic link that does not resolve to a regular
* file — a linked directory, a dangling link, a device node. It contributes its
* `readlink` TEXT and nothing else. See {@link collectArtifacts} for why the
* scan records such links instead of following them.
*/
type RuntimeArtifact =
| {
absolutePath: string;
canonicalPath: string;
kind: 'file' | 'symlink';
}
| {
absolutePath: string;
canonicalPath: string;
kind: 'unfollowed-symlink';
};
type BuildEntry = {
absolutePath: string;
@@ -157,13 +185,21 @@ type CachedBuildEntry = {
digest?: string;
};
type CachedArtifactEntry = {
absolutePath: string;
canonicalPath: string;
kind: RuntimeArtifact['kind'];
state: ReadableFileState;
digest: string;
};
type CachedArtifactEntry =
| {
absolutePath: string;
canonicalPath: string;
kind: 'file' | 'symlink';
state: ReadableFileState;
digest: string;
}
| {
absolutePath: string;
canonicalPath: string;
kind: 'unfollowed-symlink';
state: SymlinkArtifactState;
digest: string;
};
type CachedBuildDirectoryGuard = {
relativePath: string;
@@ -410,6 +446,28 @@ function snapshotReadableFile(candidate: string): ReadableFileState {
};
}
/**
* Snapshot a symbolic link without resolving it. Unlike
* {@link snapshotReadableFile} this never stats the target, so it is total over
* linked directories, dangling links, and links to device nodes — the inputs
* that make the readable-file snapshot throw.
*/
function snapshotSymlinkArtifact(candidate: string): SymlinkArtifactState {
const link = lstatSync(candidate, { bigint: true });
if (!link.isSymbolicLink()) {
throw new Error(`Analyzer identity input is not a symbolic link: ${candidate}`);
}
return { link: statState(link), symlinkTarget: readlinkSync(candidate) };
}
function snapshotRuntimeArtifact(
artifact: RuntimeArtifact,
): ReadableFileState | SymlinkArtifactState {
return artifact.kind === 'unfollowed-symlink'
? snapshotSymlinkArtifact(artifact.absolutePath)
: snapshotReadableFile(artifact.absolutePath);
}
function snapshotDirectory(candidate: string): StatState {
const stat = lstatSync(candidate, { bigint: true });
if (!stat.isDirectory() || stat.isSymbolicLink()) {
@@ -931,6 +989,100 @@ function runtimePackageLocator(packageRoot: string, runtimeRoot: string): string
return `relative:${relative}`;
}
/** Protocols that name a checkout-local package instead of a registry tarball. */
const LOCAL_LINK_PROTOCOL_PATTERN = /^(?:file|link|workspace|portal):/;
/** npm's bare local-path shorthands: `./x`, `../x`, `/x`, `~/x`, `C:\x`. */
const LOCAL_LINK_PATH_PATTERN = /^(?:\.\.?[/\\]|~[/\\]|[/\\]|[A-Za-z]:)/;
function isLocallyLinkedSpecifier(specifier: unknown): boolean {
if (typeof specifier !== 'string') return false;
const value = specifier.trim();
return LOCAL_LINK_PROTOCOL_PATTERN.test(value) || LOCAL_LINK_PATH_PATTERN.test(value);
}
/**
* Whether a REALPATH'd package root lives inside some installed dependency
* tree. Used as the resolved-location half of "is this dependency a checkout
* this repository owns?" (see {@link undeclaredLocalDevDependencyNames}).
*
* The input must already be realpath'd: `resolveDependencyPackageRoot` returns
* `realpathSync.native`, so a package reached through a link out of
* `node_modules` reports its checkout location and a package that merely lives
* in `node_modules` reports a path that still carries the segment.
*
* `pathApi` is injectable so the Windows separator handling is unit-testable
* from a POSIX runner, exactly as {@link isInside} does. The separator sets
* differ deliberately: `\` is a legal filename character on POSIX, so only
* win32 may treat it as a boundary.
*/
function hasNodeModulesSegment(candidate: string, pathApi: typeof path = path): boolean {
const segments = pathApi.sep === '\\' ? candidate.split(/[\\/]+/) : candidate.split('/');
return segments.includes('node_modules');
}
/** Test seam for {@link hasNodeModulesSegment} (see {@link _isInsideForTests}). */
export const _hasNodeModulesSegmentForTests = hasNodeModulesSegment;
/**
* How many dev dependencies may be admitted by RESOLVED LOCATION alone before
* the whole resolved-location channel is treated as untrustworthy and disabled.
*
* "Realpath carries no `node_modules` segment" is a proxy for "checkout-local",
* and a layout that materializes packages outside `node_modules` — pnpm with a
* relocated `virtual-store-dir`, a custom linker — makes every dev dependency
* pass it. Folding an entire dev tree into the receipt is not a graceful
* degradation: `runtimePackages`/`runtimeEntries`/`runtimeBytes` THROW, so a
* mis-fired proxy on a legitimate install would abort analyze outright.
*
* The bound is therefore on ADMISSIONS, and overflow admits NONE of them rather
* than an arbitrary prefix. A prefix would not bound the failure — the abort
* comes from the transitive payload of whichever trees get folded in — and it
* would make the receipt depend on an arbitrary slice of a sorted name list.
* Dropping the channel wholesale falls back to the specifier-only receipt,
* which is the behaviour that ships today and is known not to abort, and leaves
* the declared-intent half in {@link dependencyNames} untouched.
*
* Four is measured, not guessed. Monorepo and workspace links are declared
* (`file:`/`link:`/`workspace:`) and travel the uncapped declared half, so this
* channel only ever carries UNDECLARED `npm link <pkg>` — a manual, per-package
* developer action, in practice one or two packages. A mis-fire admits the
* entire dev-only set instead: 13 names in this repository's own install, tens
* in a typical application. The cap sits an order of magnitude below the
* mis-fire population and comfortably above realistic link counts.
*/
const MAX_UNDECLARED_LOCAL_DEV_DEPENDENCIES = 4;
/**
* Dependency names whose resolved packages can contribute analyzer semantics.
*
* The three runtime sections are enumerated wholesale. `devDependencies` are
* deliberately not: a registry dev tool (vitest, eslint, typescript) is never
* loaded by the analyzer, and folding the dev tree into the receipt would churn
* `dependencyRuntime.digest` — and force a full re-analysis — on every unrelated
* devDependency bump.
*
* Locally linked dev dependencies are the exception. A `file:`/`link:`/
* `workspace:` sibling is part of this checkout and ships code the analyzer
* imports at runtime: GitNexus links `gitnexus-shared`, whose schema constants
* feed `RELATION_SCHEMA`/`NODE_SCHEMA_QUERIES`. In `kind: 'source'` runs that
* sibling sits outside `buildRoot`, so leaving it out let a semantic change
* there alter analyzer behaviour while moving neither `build.digest` nor
* `dependencyRuntime.digest` — DDL-affecting edits were still caught by the
* schema fingerprint, semantics-only edits by nothing.
*
* An unresolvable link (a published install, where the sibling checkout does not
* exist) still contributes its `<missing>` edge, so the linked package appearing
* or disappearing remains a receipt change rather than a silent one. That is why
* the specifier check cannot be replaced by resolution: resolution returns
* `null` for an absent linked checkout exactly as it does for an uninstalled
* registry dev tool, and the two must not be conflated.
*
* This function is the DECLARED-INTENT half and is enumerated for every package
* in the dependency BFS, so it must stay a pure function of the manifest. The
* RESOLVED-LOCATION half — `npm link <pkg>`, which leaves the specifier a
* registry range — lives in {@link undeclaredLocalDevDependencyNames} and is
* applied to the root package only.
*/
function dependencyNames(manifest: PackageManifest): string[] {
const names = new Set<string>();
for (const section of [
@@ -941,6 +1093,12 @@ function dependencyNames(manifest: PackageManifest): string[] {
if (!section || typeof section !== 'object') continue;
for (const name of Object.keys(section)) names.add(name);
}
const development = manifest.devDependencies;
if (development && typeof development === 'object') {
for (const [name, specifier] of Object.entries(development)) {
if (isLocallyLinkedSpecifier(specifier)) names.add(name);
}
}
return [...names].sort(compareBytes);
}
@@ -980,6 +1138,53 @@ function resolveDependencyPackageRoot(
}
}
/**
* Dev dependencies that are locally linked by INSTALLED LOCATION rather than by
* declared specifier — the `npm link <pkg>` shape, where the manifest still
* carries a registry range while `node_modules/<pkg>` is a symlink into a
* working checkout. {@link isLocallyLinkedSpecifier} is blind to those, yet the
* linked code is exactly as load-bearing for analyzer semantics as a declared
* `file:` sibling, so a semantic-only edit there would move neither digest.
*
* The resolver already knows: {@link resolveDependencyPackageRoot} returns a
* realpath, so a linked package reports a root outside every `node_modules`
* tree while an ordinary installed package cannot.
*
* Two properties are load-bearing and must not be relaxed:
*
* 1. ROOT ONLY. {@link dependencyNames} runs for every package in the BFS, and
* published tarballs keep their `devDependencies`, so probing dev-only names
* everywhere costs 1998 resolutions rather than the ~13 this manifest
* declares — measured on this install, with 0 true positives. Persisted
* `dependencyPathGuards` grow 2220 → 11049, and every guard is re-probed on
* each warm validation, so the cost is recurring and on the `status` path.
* Root-only costs 13 resolutions and ~29 guards.
* 2. An unresolvable name is NEVER admitted. `null` here means "uninstalled
* registry dev tool" far more often than "broken link", and admitting it
* would emit a `<missing>` edge for every dev tool absent from a published
* install. Declared links keep that edge through {@link dependencyNames};
* undeclared ones have no declaration to honour.
*
* The admission count is bounded by {@link MAX_UNDECLARED_LOCAL_DEV_DEPENDENCIES}.
*/
function undeclaredLocalDevDependencyNames(
rootPackage: RuntimePackage,
pathGuards: Map<string, DependencyPathGuardResult>,
limits: AnalyzerIdentityTraversalLimits,
): string[] {
const development = rootPackage.manifest.devDependencies;
if (!development || typeof development !== 'object') return [];
const admitted: string[] = [];
for (const [name, specifier] of Object.entries(development)) {
// Already carried by the declared half; resolving again would only add
// guards. Its `<missing>` edge is that half's responsibility.
if (isLocallyLinkedSpecifier(specifier)) continue;
const resolved = resolveDependencyPackageRoot(rootPackage.root, name, pathGuards, limits);
if (resolved !== null && !hasNodeModulesSegment(resolved)) admitted.push(name);
}
return admitted.length <= MAX_UNDECLARED_LOCAL_DEV_DEPENDENCIES ? admitted : [];
}
function collectRuntimePackages(
packageRoot: string,
directoryGuards: Map<string, DependencyDirectoryGuard>,
@@ -1010,7 +1215,21 @@ function collectRuntimePackages(
for (let index = 0; index < queue.length; index += 1) {
const parent = queue[index];
for (const dependencyName of dependencyNames(parent.manifest)) {
// The declared half is enumerated for every package; the resolved-location
// half is scoped to the root package, where the 1998-resolution /
// 8829-extra-guard blow-up documented on
// `undeclaredLocalDevDependencyNames` cannot occur. Dropping this scope is
// the expensive regression, so it is pinned by a guard-count test.
const dependencies =
parent.root === packageRoot
? [
...new Set([
...dependencyNames(parent.manifest),
...undeclaredLocalDevDependencyNames(parent, pathGuards, limits),
]),
].sort(compareBytes)
: dependencyNames(parent.manifest);
for (const dependencyName of dependencies) {
budget.edges += 1;
if (budget.edges > limits.runtimeEdges) {
throw new Error(
@@ -1108,16 +1327,24 @@ function collectArtifacts(
);
}
for (const entry of entries) {
const absolutePath = path.join(absoluteDir, entry.name);
const relativePath = path.relative(root, absolutePath).split(path.sep).join('/');
const stat = lstatSync(absolutePath);
// Nested dependencies are collected from their manifests as separate
// packages. Only prune those separately traversed trees and VCS
// metadata; generic cache/model directories can contain loadable code,
// native addons, Wasm modules, or data consumed by the runtime.
if (entry.isDirectory() && PRUNED_RUNTIME_DIRECTORIES.has(entry.name)) {
continue;
}
const absolutePath = path.join(absoluteDir, entry.name);
const relativePath = path.relative(root, absolutePath).split(path.sep).join('/');
const stat = lstatSync(absolutePath);
//
// Pruning is decided by NAME alone. These four names never carry analyzer
// payload in any form: `node_modules` is traversed separately through
// `resolveDependencyPackageRoot` (which follows links and guards each
// hop), and a `.git`/`.hg`/`.svn` entry is VCS metadata whether it is a
// directory, a symbolic link into a shared store, or — inside a submodule
// or linked worktree checkout — a regular file holding a gitdir pointer.
// Hashing that pointer would make analyzer identity depend on where the
// checkout happens to live, which is a false-stale source, not a
// semantic input.
if (PRUNED_RUNTIME_DIRECTORIES.has(entry.name)) continue;
if (stat.isDirectory()) {
if (depth >= limits.runtimeDepth) {
throw new Error(
@@ -1125,6 +1352,39 @@ function collectArtifacts(
);
}
pending.push({ absoluteDir: absolutePath, depth: depth + 1 });
} else if (stat.isSymbolicLink() && !isFile(absolutePath)) {
// A symbolic link that does not resolve to a regular file must never
// reach the payload branch below: `snapshotReadableFile` stats the
// target, and a directory (or a dangling link) makes it throw, aborting
// the entire analyze. Workspace-linked checkouts made this reachable
// for every name, not just the pruned four — `dist -> build`, a
// vendored-grammar link, anything a sibling checkout ships.
//
// Such links are RECORDED by their link text rather than followed.
// Following them would (a) recurse without cycle protection — this
// traversal has none, so `self -> .` would ride the depth limit, which
// THROWS, trading one hard abort for another; (b) re-scan trees already
// reached by their real path, inflating the entry/byte budgets that
// also throw; and (c) need a whole containment/TOCTOU trust boundary
// for targets outside the package. Recording the text is cycle-free,
// costs one `readlink`, and still moves the receipt when the link is
// retargeted. The trade-off is that a link's target contributes no
// content of its own: when it points outside the package, only the
// link text is covered. Links that DO resolve to a regular file keep
// their content digest below, unchanged.
if (shouldHashRuntimePayload(relativePath)) {
budget.artifacts += 1;
if (budget.artifacts > limits.runtimePayloads) {
throw new Error(
`Analyzer runtime payload scan exceeded ${limits.runtimePayloads} payloads: ${root}`,
);
}
artifacts.push({
absolutePath,
canonicalPath: `${canonicalPrefix}/${relativePath}`,
kind: 'unfollowed-symlink',
});
}
} else if (
(stat.isFile() || stat.isSymbolicLink()) &&
shouldHashRuntimePayload(relativePath)
@@ -1322,7 +1582,7 @@ function dependencySnapshot(inputs: DependencyInputs): unknown {
absolutePath: artifact.absolutePath,
canonicalPath: artifact.canonicalPath,
kind: artifact.kind,
state: snapshotReadableFile(artifact.absolutePath),
state: snapshotRuntimeArtifact(artifact),
})),
directories: [...inputs.directoryGuards.entries()]
.map(([absolutePath, guard]) => ({ absolutePath, ...guard }))
@@ -1343,10 +1603,30 @@ function hashRuntimeArtifact(
artifact: RuntimeArtifact,
cache: CachedArtifactEntry | undefined,
options: AnalyzerIdentityResolveOptions,
): { digest: string; state: ReadableFileState } {
): CachedArtifactEntry {
if (artifact.kind === 'unfollowed-symlink') {
// The link text is the entire payload, so there is no file read for a
// cached digest to amortize: recompute it and stay independent of the
// cache's freshness. The distinct frame label keeps a link recording from
// ever colliding with a content digest.
const state = snapshotSymlinkArtifact(artifact.absolutePath);
return {
...artifact,
state,
digest: hashCanonicalFrames([
['runtime-payload-link-v1', artifact.kind, state.symlinkTarget],
]),
};
}
const before = snapshotReadableFile(artifact.absolutePath);
if (cache && SHA256_PATTERN.test(cache.digest) && isDeepStrictEqual(cache.state, before)) {
return { digest: cache.digest, state: before };
if (
cache &&
cache.kind !== 'unfollowed-symlink' &&
SHA256_PATTERN.test(cache.digest) &&
isDeepStrictEqual(cache.state, before)
) {
return { ...artifact, state: before, digest: cache.digest };
}
const stable = hashStableFile(artifact.absolutePath);
@@ -1363,7 +1643,7 @@ function hashRuntimeArtifact(
path: artifact.absolutePath,
bytes: stable.bytes,
});
return { digest, state: stable.state };
return { ...artifact, state: stable.state, digest };
}
function compareEdges(a: RuntimeDependencyEdge, b: RuntimeDependencyEdge): number {
@@ -1445,7 +1725,7 @@ function hashDependencyRuntime(
artifact.kind,
digestBytes(hashed.digest),
]);
nextArtifacts.push({ ...artifact, state: hashed.state, digest: hashed.digest });
nextArtifacts.push(hashed);
}
return {
@@ -1479,6 +1759,19 @@ function isReadableFileState(value: unknown): value is ReadableFileState {
);
}
function isSymlinkArtifactState(value: unknown): value is SymlinkArtifactState {
if (typeof value !== 'object' || value === null) return false;
const record = value as Record<string, unknown>;
// `target === undefined` keeps a readable-file state from masquerading as an
// unresolved link recording, which would otherwise be validated against the
// wrong guard mode on the warm path.
return (
isStatState(record.link) &&
typeof record.symlinkTarget === 'string' &&
record.target === undefined
);
}
function isDependencyPathGuardResult(value: unknown): value is DependencyPathGuardResult {
if (value === null) return true;
if (typeof value !== 'object') return false;
@@ -1649,15 +1942,18 @@ function isIdentityCachePayload(
const artifactEntriesValid = record.artifactEntries.every((entry: unknown) => {
if (typeof entry !== 'object' || entry === null) return false;
const item = entry as Record<string, unknown>;
return (
typeof item.absolutePath === 'string' &&
path.isAbsolute(item.absolutePath) &&
typeof item.canonicalPath === 'string' &&
(item.kind === 'file' || item.kind === 'symlink') &&
isReadableFileState(item.state) &&
typeof item.digest === 'string' &&
SHA256_PATTERN.test(item.digest)
);
if (
typeof item.absolutePath !== 'string' ||
!path.isAbsolute(item.absolutePath) ||
typeof item.canonicalPath !== 'string' ||
typeof item.digest !== 'string' ||
!SHA256_PATTERN.test(item.digest)
) {
return false;
}
return item.kind === 'unfollowed-symlink'
? isSymlinkArtifactState(item.state)
: (item.kind === 'file' || item.kind === 'symlink') && isReadableFileState(item.state);
});
const hasBuildRootGuard = record.buildDirectoryGuards.some(
(entry: unknown) =>
@@ -2146,14 +2442,24 @@ function validateIdentityCache(
}
}
for (const artifact of cache.artifactEntries) {
if (
!add(
{ absolutePath: artifact.absolutePath, mode: 'readable-file' },
{ type: 'readable-file', state: artifact.state },
)
) {
return false;
}
// A recorded link is re-probed as a link, never as a readable file: the
// readable-file probe resolves the target and would report `null` for the
// very inputs this kind exists to describe, failing every warm validation.
const probe: { request: CacheGuardRequest; expected: CacheGuardResult } =
artifact.kind === 'unfollowed-symlink'
? {
request: { absolutePath: artifact.absolutePath, mode: 'link' },
expected: {
type: 'symlink',
state: artifact.state.link,
symlinkTarget: artifact.state.symlinkTarget,
},
}
: {
request: { absolutePath: artifact.absolutePath, mode: 'readable-file' },
expected: { type: 'readable-file', state: artifact.state },
};
if (!add(probe.request, probe.expected)) return false;
}
const entries = [...expected.entries()];
+94 -71
View File
@@ -132,6 +132,7 @@ export async function augment(pattern: string, cwd?: string): Promise<string> {
MATCH (n) WHERE n.filePath = '${escaped}'
AND n.name CONTAINS '${patternFirstWord}'
RETURN n.id AS id, n.name AS name, labels(n)[0] AS type, n.filePath AS filePath
ORDER BY id
LIMIT 3
`,
);
@@ -158,6 +159,7 @@ export async function augment(pattern: string, cwd?: string): Promise<string> {
MATCH (n)
WHERE n.name CONTAINS '${patternFirstWord}'
RETURN n.id AS id, n.name AS name, labels(n)[0] AS type, n.filePath AS filePath
ORDER BY id
LIMIT 5
`,
).catch(() => []);
@@ -184,98 +186,119 @@ export async function augment(pattern: string, cwd?: string): Promise<string> {
const idList = uniqueSymbols.map((s) => `'${escapeCypherString(s.nodeId)}'`).join(', ');
// Batch fetch callers
const callersMap = new Map<string, string[]>();
try {
// Callers/callees are windowed PER SYMBOL, not out of one shared budget.
// A single `n.id IN [...] ... ORDER BY targetId LIMIT 15` sorts by the very
// column the rows are grouped by, which turns the cap into a single-bucket
// prefix: the alphabetically-first target takes every row (the hottest
// symbol in this repo's own index has 2163 callers) and the other four
// render with no callers at all. Raising the cap does not bound that, and
// ordering caller-major only moves the bucket. A per-target window is not
// expressible in one statement either — LadybugDB does not preserve a
// per-branch ORDER BY through UNION ALL — so fan out one bounded query per
// symbol (#2787).
//
// `id STARTS WITH 'File:'` sorts container nodes last: node ids are
// `Label:path:name`, so a bare `ORDER BY id` puts `File:` ahead of
// `Function:`/`Method:` and the "Called by" line renders bare filenames
// instead of the calling symbols the hint exists to name.
const NEIGHBOUR_CAP = 3;
const neighbourNames = async (nodeId: string, incoming: boolean): Promise<string[]> => {
const edge = incoming
? `(other)-[:CodeRelation {type: 'CALLS'}]->(n)`
: `(n)-[:CodeRelation {type: 'CALLS'}]->(other)`;
const rows = await executeQuery(
repoId,
`
MATCH (caller)-[:CodeRelation {type: 'CALLS'}]->(n)
WHERE n.id IN [${idList}]
RETURN n.id AS targetId, caller.name AS name
LIMIT 15
MATCH ${edge}
WHERE n.id = '${escapeCypherString(nodeId)}'
RETURN other.name AS name
ORDER BY other.id STARTS WITH 'File:', other.id
LIMIT ${NEIGHBOUR_CAP}
`,
);
).catch(() => []);
const names: string[] = [];
for (const r of rows) {
const tid = r.targetId || r[0];
const name = r.name || r[1];
if (tid && name) {
if (!callersMap.has(tid)) callersMap.set(tid, []);
callersMap.get(tid)!.push(name);
}
const name = r.name || r[0];
if (name) names.push(name);
}
} catch {
/* skip */
}
return names;
};
// Batch fetch callees
const calleesMap = new Map<string, string[]>();
try {
const rows = await executeQuery(
repoId,
`
MATCH (n)-[:CodeRelation {type: 'CALLS'}]->(callee)
WHERE n.id IN [${idList}]
RETURN n.id AS sourceId, callee.name AS name
LIMIT 15
`,
// Keyed by nodeId, like the process/cohesion enrichments below — positional
// arrays would put two lookup disciplines in one object literal and make
// "stays index-aligned with uniqueSymbols" an unenforced invariant.
const neighbourMap = async (incoming: boolean): Promise<Map<string, string[]>> =>
new Map(
await Promise.all(
uniqueSymbols.map(
async (s) => [s.nodeId, await neighbourNames(s.nodeId, incoming)] as const,
),
),
);
for (const r of rows) {
const sid = r.sourceId || r[0];
const name = r.name || r[1];
if (sid && name) {
if (!calleesMap.has(sid)) calleesMap.set(sid, []);
calleesMap.get(sid)!.push(name);
}
}
} catch {
/* skip */
}
// Batch fetch processes
const processesMap = new Map<string, string[]>();
try {
const rows = await executeQuery(
repoId,
`
const fetchProcesses = async (): Promise<Map<string, string[]>> => {
const byNode = new Map<string, string[]>();
try {
const rows = await executeQuery(
repoId,
`
MATCH (n)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
WHERE n.id IN [${idList}]
RETURN n.id AS nodeId, p.heuristicLabel AS label, r.step AS step, p.stepCount AS stepCount
`,
);
for (const r of rows) {
const nid = r.nodeId || r[0];
const label = r.label || r[1];
const step = r.step || r[2];
const stepCount = r.stepCount || r[3];
if (nid && label) {
if (!processesMap.has(nid)) processesMap.set(nid, []);
processesMap.get(nid)!.push(`${label} (step ${step}/${stepCount})`);
);
for (const r of rows) {
const nid = r.nodeId || r[0];
const label = r.label || r[1];
const step = r.step || r[2];
const stepCount = r.stepCount || r[3];
if (nid && label) {
if (!byNode.has(nid)) byNode.set(nid, []);
byNode.get(nid)!.push(`${label} (step ${step}/${stepCount})`);
}
}
} catch {
/* skip */
}
} catch {
/* skip */
}
return byNode;
};
// Batch fetch cohesion
const cohesionMap = new Map<string, number>();
try {
const rows = await executeQuery(
repoId,
`
const fetchCohesion = async (): Promise<Map<string, number>> => {
const byNode = new Map<string, number>();
try {
const rows = await executeQuery(
repoId,
`
MATCH (n)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
WHERE n.id IN [${idList}]
RETURN n.id AS nodeId, c.cohesion AS cohesion
`,
);
for (const r of rows) {
const nid = r.nodeId || r[0];
const coh = r.cohesion ?? r[1] ?? 0;
if (nid) cohesionMap.set(nid, coh);
);
for (const r of rows) {
const nid = r.nodeId || r[0];
const coh = r.cohesion ?? r[1] ?? 0;
if (nid) byNode.set(nid, coh);
}
} catch {
/* skip */
}
} catch {
/* skip */
}
return byNode;
};
// One wave, not three. `augment` runs from the Claude Code PreToolUse hook
// in a cold process against a <500ms budget, so nothing amortizes — and the
// process/cohesion queries depend only on `idList`, never on the neighbour
// results, so serializing them behind the fan-out bought nothing. Each query
// keeps its own try/catch, so one failing still degrades to an empty map
// instead of taking the others down.
const [callerMap, calleeMap, processesMap, cohesionMap] = await Promise.all([
neighbourMap(true),
neighbourMap(false),
fetchProcesses(),
fetchCohesion(),
]);
// Assemble enriched results
const enriched: Array<{
@@ -291,8 +314,8 @@ export async function augment(pattern: string, cwd?: string): Promise<string> {
enriched.push({
name: sym.name,
filePath: sym.filePath,
callers: (callersMap.get(sym.nodeId) || []).slice(0, 3),
callees: (calleesMap.get(sym.nodeId) || []).slice(0, 3),
callers: callerMap.get(sym.nodeId) || [],
callees: calleeMap.get(sym.nodeId) || [],
processes: processesMap.get(sym.nodeId) || [],
cohesion: cohesionMap.get(sym.nodeId) || 0,
});
+301
View File
@@ -0,0 +1,301 @@
/**
* The single owner of `RepoMeta.embeddingCheckpoint` — how it is minted, and
* how a run decides what to do with one it finds.
*
* This module exists for the reason `embedding-count.ts` next door exists, and
* for a sharper one. A review of #2790 found that two hand-copied bodies of
* "measure the embedding count" had drifted inside a single change; the fix for
* that then created a SECOND pair of hand-copied publishers — of this record —
* and they had already drifted too: the CLI armed the attempt counter only
* after clearing its identity gate, the server derived it from the resumed
* marker alone. Worse, only one of the two READERS implemented `kind` at all,
* so a `'partial'` marker written by `gitnexus analyze` and resumed through
* `POST /api/embed` hit exactly the permanent wedge `kind` was introduced to
* remove. Prose promising two sites stay in step is not a mechanism. This is.
*
* Same tier as {@link ./embedding-count.ts} and {@link ./embedding-mode.ts}:
* type-only imports, no native dependencies, so `run-analyze.ts` can import it
* statically without dragging an embeddings module in (#2370) and `server/`
* can import it too.
*/
import type { RepoMeta } from '../storage/repo-manager.js';
export type EmbeddingCheckpoint = NonNullable<RepoMeta['embeddingCheckpoint']>;
export type EmbeddingCheckpointKind = NonNullable<EmbeddingCheckpoint['kind']>;
export interface EmbeddingRunIdentity {
model: string;
dimensions: number;
provider: string;
}
export interface EmbeddingCheckpointProgress {
nodesProcessed: number;
totalNodes: number;
chunksProcessed: number;
}
/**
* How many consecutive resume attempts may fail to clear a `'partial'` pending
* set before it is abandoned.
*
* Not a fresh guess: it is this repo's existing per-operation retry budget —
* the HTTP embedder spends `HTTP_MAX_RETRIES + 1 === 3` per request and the WAL
* driver `CHECKPOINT_RETRY_ATTEMPTS === 3` per flush. Each attempt here is a
* whole analyze invocation with those budgets nested underneath, so 3 is
* already generous while still converging within a day of ordinary committing.
*/
export const EMBEDDING_RESUME_MAX_ATTEMPTS = 3;
/**
* The one home for the back-compat default. Markers written before #2790's
* follow-up carry no `kind`, and they must keep the stricter behavior — an
* implicit default scattered across read sites is precisely how the two
* readers came to disagree.
*/
export const checkpointKind = (checkpoint: EmbeddingCheckpoint): EmbeddingCheckpointKind =>
checkpoint.kind ?? 'interrupted';
/** Pending-node count, tolerating the field's optionality. */
export const pendingNodeCount = (checkpoint: EmbeddingCheckpoint): number =>
checkpoint.pendingNodeIds?.length ?? 0;
/**
* The attempt-chain rule, which is the subtlest thing here and was the part
* that had already diverged.
*
* The counter advances only when this run resumed a `'partial'` marker AND at
* least one node it was handed failed AGAIN. A resume that cleared its set but
* lost *different* nodes is a FRESH partial, so the budget resets — otherwise a
* steadily-flaky endpoint would exhaust the budget on nodes that were never the
* problem. `undefined` rather than `0` when the streak breaks, so a clean
* marker serializes without the field.
*/
export const nextAttemptCount = (
resumedFrom: EmbeddingCheckpoint | undefined,
failedNodeIds: readonly string[],
): number | undefined => {
if (resumedFrom === undefined || checkpointKind(resumedFrom) !== 'partial') return undefined;
const resumedPending = new Set(resumedFrom.pendingNodeIds ?? []);
const failedAgain = failedNodeIds.some((id) => resumedPending.has(id));
return failedAgain ? (resumedFrom.attempts ?? 0) + 1 : undefined;
};
/**
* Mint the marker written BEFORE a bounded write window opens.
*
* `'interrupted'`: if the process dies inside the window these nodes may hold a
* subset of their chunks, which is what forces resume to delete and regenerate
* them even when a persisted row carries the current content hash — and what
* makes an identity mismatch fail closed rather than mix vector spaces. No
* `attempts`: the retry bound belongs to `'partial'`, whose pending set
* provably holds zero rows and can therefore be abandoned safely.
*/
export const mintInterruptedCheckpoint = (
identity: EmbeddingRunIdentity,
progress: EmbeddingCheckpointProgress,
pendingNodeIds: string[],
at: string = new Date().toISOString(),
): EmbeddingCheckpoint => ({
at,
...progress,
model: identity.model,
dimensions: identity.dimensions,
provider: identity.provider,
kind: 'interrupted',
pendingNodeIds,
});
/**
* Mint the marker written AFTER a run that completed while dropping nodes.
*
* `totalNodes` is reconstructed as complete + dropped, because `nodesProcessed`
* counts only the nodes that finished with a full row set. It feeds the resume
* log line and nothing else.
*/
export const mintPartialCheckpoint = (
identity: EmbeddingRunIdentity,
result: { nodesProcessed: number; chunksProcessed: number; failedNodeIds: string[] },
resumedFrom: EmbeddingCheckpoint | undefined,
at: string = new Date().toISOString(),
): EmbeddingCheckpoint => ({
at,
nodesProcessed: result.nodesProcessed,
totalNodes: result.nodesProcessed + result.failedNodeIds.length,
chunksProcessed: result.chunksProcessed,
model: identity.model,
dimensions: identity.dimensions,
provider: identity.provider,
kind: 'partial',
attempts: nextAttemptCount(resumedFrom, result.failedNodeIds),
pendingNodeIds: result.failedNodeIds,
});
/**
* Mint the marker written when a run finished but its embedding count could not
* be verified.
*
* A third state, and it needs its own `kind` rather than borrowing `'partial'`:
* its pending set is EMPTY. Nothing was dropped and nothing needs re-embedding
* — the marker exists only to defeat the same-commit fast return so the next
* run re-derives a count, because clearing it while `stats.embeddings` still
* reads a stale zero is what arms a later `--force` to wipe live embeddings.
* Stamping it `'partial'` made `gitnexus status` tell the operator that N nodes
* had lost their embeddings, where N is zero.
*/
export const mintUnverifiedCountCheckpoint = (
identity: EmbeddingRunIdentity,
progress: EmbeddingCheckpointProgress,
at: string = new Date().toISOString(),
): EmbeddingCheckpoint => ({
at,
...progress,
model: identity.model,
dimensions: identity.dimensions,
provider: identity.provider,
kind: 'unverified-count',
pendingNodeIds: [],
});
export type EmbeddingResumeDecision =
/** An explicit flag cleared it; nothing to resume. */
| { readonly action: 'discard'; readonly log: string }
/**
* Keep running, but abandon the pending set — either the retry budget is
* spent, or the marker was written by a different embedding identity whose
* nodes provably hold zero rows so nothing is at risk.
*/
| { readonly action: 'abandon'; readonly log: string }
/** Resume: regenerate `pendingNodeIds` under the matching identity. */
| {
readonly action: 'resume';
readonly log: string;
readonly pendingNodeIds: ReadonlySet<string>;
readonly resumedFrom: EmbeddingCheckpoint;
}
/** Fail closed: a half-written window under a foreign identity. */
| { readonly action: 'abort'; readonly error: string };
/**
* The one implementation of "what should this run do with the checkpoint it
* found", shared by the CLI resume gate and `POST /api/embed`.
*
* Callers resolve the embedding identity lazily and pass it only when they need
* a verdict beyond the flag checks — resolving it is not free, and the flag
* paths short-circuit before it is needed. Pass `identity: undefined` to get
* just those; a decision that requires an identity returns `'resume'` only when
* one was supplied and matched.
*/
export const decideEmbeddingResume = (
checkpoint: EmbeddingCheckpoint,
identity: EmbeddingRunIdentity | undefined,
options: {
force?: boolean;
dropEmbeddings?: boolean;
maxAttempts?: number;
} = {},
): EmbeddingResumeDecision => {
const kind = checkpointKind(checkpoint);
const pending = pendingNodeCount(checkpoint);
const maxAttempts = options.maxAttempts ?? EMBEDDING_RESUME_MAX_ATTEMPTS;
if (options.dropEmbeddings) {
return { action: 'discard', log: 'Discarding the embedding checkpoint (--drop-embeddings).' };
}
if (options.force) {
return {
action: 'discard',
log:
'Discarding the embedding checkpoint (--force)' +
`${pending > 0 ? ` and its ${pending} pending node(s)` : ''}; this run rebuilds from ` +
'scratch. Pass --embeddings to regenerate them explicitly.',
};
}
// Only the 'unverified-count' marker may skip the identity gate. Its pending
// set is empty BY DEFINITION and it exists purely to force a fresh count, so
// there is nothing a foreign embedding identity could corrupt.
//
// An empty pending set alone must NOT take this path. `onCheckpoint` mints an
// 'interrupted' marker with `pendingNodeIds: []` after every post-window
// save, and legacy markers may omit the field entirely — so keying on
// emptiness would silently clear an interrupted marker under a foreign
// provider instead of failing closed, which is exactly the vector-space
// mixing the 'interrupted' kind exists to prevent.
if (kind === 'unverified-count') {
return {
action: 'abandon',
log: 'Re-deriving the embedding count left unverified by the previous run.',
};
}
if (kind === 'partial' && (checkpoint.attempts ?? 0) >= maxAttempts) {
return {
action: 'abandon',
log:
`Warning: ${pending} node(s) failed to embed on ${checkpoint.attempts} consecutive ` +
'resume attempts and are being abandoned (#2790). The index is registered without ' +
'them; `gitnexus status` stops reporting it as incomplete. Re-run ' +
'`gitnexus analyze --embeddings --force` once the embedding endpoint accepts them.',
};
}
if (identity === undefined) {
return {
action: 'abort',
error: 'Cannot resume embedding checkpoint: no embedding identity was resolved.',
};
}
const providerDiffers = checkpoint.provider !== identity.provider;
const identityDiffers =
providerDiffers ||
checkpoint.model !== identity.model ||
checkpoint.dimensions !== identity.dimensions;
if (identityDiffers && kind !== 'interrupted') {
// Those nodes hold zero rows, so there is no vector space to mix. Throwing
// here is what turned an exit-0 partial run into a permanent wedge: a hook
// or CI job without the endpoint's env vars resolves `provider: 'local'`,
// mismatches, and dies before any phase runs.
return {
action: 'abandon',
log:
`Warning: dropping ${pending} pending node(s) from a ${kind} embedding checkpoint: it ` +
`was written by a different embedding configuration (${checkpoint.model} at ` +
`${checkpoint.dimensions} dimensions) than this run resolves (${identity.model} at ` +
`${identity.dimensions}). Those nodes hold no embedding rows, so nothing is lost that ` +
'a later `--embeddings` run cannot regenerate.',
};
}
if (providerDiffers) {
return {
action: 'abort',
error:
'Cannot resume embedding checkpoint: the embedding provider configuration differs. ' +
'Restore the matching endpoint configuration or pass --drop-embeddings to rebuild without it.',
};
}
if (identityDiffers) {
return {
action: 'abort',
error:
`Cannot resume embedding checkpoint: it uses ${checkpoint.model} at ` +
`${checkpoint.dimensions} dimensions, but this run resolves ${identity.model} at ` +
`${identity.dimensions}. Restore the matching embedding configuration or pass ` +
'--drop-embeddings to rebuild without it.',
};
}
// Reached with an EMPTY pending set too — an 'interrupted' marker written by
// a post-window save names none. That still resumes rather than clearing:
// the marker's presence is what forces embedding generation on a run that
// would otherwise derive `shouldGenerateEmbeddings: false`, and it has now
// cleared the identity gate, so resuming is safe.
return {
action: 'resume',
log:
`Previous analyze ended at an embedding checkpoint (${checkpoint.nodesProcessed}/` +
`${checkpoint.totalNodes} nodes); resuming from persisted hashes` +
`${pending > 0 ? ` and regenerating ${pending} pending node(s)` : ''}.`,
pendingNodeIds: new Set(checkpoint.pendingNodeIds ?? []),
resumedFrom: checkpoint,
};
};
+90
View File
@@ -0,0 +1,90 @@
/**
* The single implementation of "how many embedding rows are actually persisted".
*
* Lives in its own module — beside {@link ./embedding-mode.ts} and with the same
* no-native-imports property — because the CLI (`run-analyze.ts`) and the server
* (`server/api.ts`) both publish `RepoMeta.stats.embeddings`, and a review of
* #2790 found the two hand-copied bodies had already drifted inside a single
* change: one used `?? Number.NaN`, the other `?? 0`, under a comment asserting
* they measured the field "the same way". Two publishers of one field must not
* be able to disagree, so there is now one function and three call sites.
*
* Keeping it out of `core/embeddings/` is deliberate: `run-analyze.ts` may only
* reach embedding modules through `await import(...)` (#2370 — no embeddings
* module loads unless a run actually needs one), and this counter runs on the
* ordinary finalization path of every analyze.
*
* TRI-STATE, and the asymmetry is the whole point. `unknown` means COULD NOT
* ASK — the query throws for reasons that have nothing to do with how many rows
* were written (table missing, connection closed, DB busy or read-only, the
* VECTOR-extension DML lock #2623) — and a fabricated count is never returned in
* its place. A wrong-LOW value is the dangerous direction: `stats.embeddings` is
* the sole input to the next run's `existingEmbeddingCount` → `deriveEmbeddingMode`
* → `shouldLoadCache`, so a false `0` is exactly what makes a later `--force`
* wipe live embeddings without loading the cache.
*/
import { EMBEDDING_TABLE_NAME } from 'gitnexus-shared';
/** Callback shape both callers already satisfy with their `executeQuery`. */
export type EmbeddingCountQuery = (
cypher: string,
) => Promise<Array<Record<string, unknown>> | undefined>;
export type PersistedEmbeddingCount =
| { readonly kind: 'measured'; readonly count: number }
/** `reason` is operator-facing: each caller phrases its own warning around it. */
| { readonly kind: 'unknown'; readonly reason: string };
export const EMBEDDING_COUNT_CYPHER = `MATCH (e:${EMBEDDING_TABLE_NAME}) RETURN count(e) AS cnt`;
/**
* Count the persisted embedding rows, or report that the answer never arrived.
*
* A MISSING row or cell counts as unknown, NOT as zero: an empty table still
* answers with one row holding `0`, so no row / no cell means the query did not
* really answer. That distinction is what `?? Number.NaN` buys — with `?? 0`,
* `Number.isFinite(0)` is true and the unknown branch becomes unreachable for
* exactly the case it exists to catch.
*/
export const measurePersistedEmbeddingCount = async (
runQuery: EmbeddingCountQuery,
): Promise<PersistedEmbeddingCount> => {
let rows: Array<Record<string, unknown>> | undefined;
try {
rows = await runQuery(EMBEDDING_COUNT_CYPHER);
} catch (err) {
return { kind: 'unknown', reason: err instanceof Error ? err.message : String(err) };
}
const row = rows?.[0];
const parsed = Number(row?.cnt ?? row?.[0] ?? Number.NaN);
if (Number.isFinite(parsed)) return { kind: 'measured', count: parsed };
return {
kind: 'unknown',
reason:
row === undefined
? 'the count query returned no row'
: 'the count query returned a non-numeric result',
};
};
/** `undefined` ≡ unknown, for the call sites that only need the optional number. */
export const persistedEmbeddingCountOrUndefined = (
result: PersistedEmbeddingCount,
): number | undefined => (result.kind === 'measured' ? result.count : undefined);
/**
* Fold a measurement into the `RepoMeta.stats.embeddings` a run publishes.
*
* The measurement was already shared; this fold was not, and it is the half
* that decides what actually lands on disk — the CLI's finalization and the
* server's `withMeasuredEmbeddingCount` each carried their own copy of the same
* two lines. NEVER a fabricated `0`: `unknown` carries `priorCount` forward
* (see this file's header for the chain a false zero arms), and `undefined` in
* means `undefined` out, so a repo that never had a count keeps the field
* absent rather than gaining an invented one.
*/
export const resolvePersistedEmbeddingCount = (
measured: PersistedEmbeddingCount,
priorCount: number | undefined,
): number | undefined => (measured.kind === 'measured' ? measured.count : priorCount);
@@ -39,6 +39,7 @@ import {
resolveEmbeddingConfig,
} from './config.js';
import { rankExactEmbeddingRows, type ExactEmbeddingRow } from './exact-search.js';
import { EMBEDDING_COUNT_CYPHER } from '../embedding-count.js';
import { EMBEDDING_TABLE_NAME, EMBEDDING_INDEX_NAME, STALE_HASH_SENTINEL } from '../lbug/schema.js';
import { loadVectorExtension, createVectorIndex } from '../lbug/lbug-adapter.js';
import { escapeCypherString } from '../lbug/cypher-escape.js';
@@ -175,7 +176,9 @@ const queryEmbeddableNodes = async (
}
} catch (error) {
if (isDev) {
logger.warn({ error }, `Query for ${label} nodes failed:`);
// `err` is pino's standard serializer key — an arbitrary key serializes an
// Error to `{}`, losing the message and stack (#2114).
logger.warn({ err: error }, `Query for ${label} nodes failed:`);
}
}
}
@@ -220,7 +223,8 @@ const queryFallbackFileNodes = async (
);
} catch (error) {
if (isDev) {
logger.warn({ error }, 'Fallback File-node embedding query failed:');
// `err`, not `error` — see the serializer note above (#2114).
logger.warn({ err: error }, 'Fallback File-node embedding query failed:');
}
return [];
}
@@ -305,12 +309,156 @@ export const buildVectorIndex = async (): Promise<boolean> => {
};
export interface EmbeddingPipelineResult {
/** Nodes that finished the run with a COMPLETE set of embedding rows. */
nodesProcessed: number;
chunksProcessed: number;
vectorIndexReady: boolean;
semanticMode: 'vector-index' | 'exact-scan';
/**
* Nodes whose embeddings were dropped this run because at least one of their
* chunks lost its sub-batch to a failing endpoint (#2790). Their rows were
* deleted, so they now hold ZERO rows and the next incremental run re-embeds
* them from scratch. Empty on a clean run.
*/
failedNodeIds: string[];
}
/**
* How many sub-batches may fail back-to-back before the pipeline gives up.
*
* Surviving a transient endpoint hiccup is the whole point of tolerating a
* failed sub-batch (#2790) — but a genuinely dead endpoint would otherwise walk
* every remaining node deleting its rows on the way, which on an incremental run
* wipes surviving embeddings wholesale. Rethrowing once the ceiling is reached
* preserves the pre-#2790 behavior for a dead endpoint. Any successful sub-batch
* resets the counter, so only an unbroken run of failures trips it.
*
* ponytail: flat consecutive-failure ceiling; make it rate-based if flaky
* endpoints prove common.
*/
const MAX_CONSECUTIVE_SUB_BATCH_FAILURES = 5;
/**
* Cumulative guard: the share of sub-batches allowed to fail across the whole
* run, and the minimum sample before that share means anything.
*
* The ceiling above only catches a total outage — any success resets it, so an
* endpoint load-shedding every other sub-batch walks the entire repo dropping
* half of it while `analyze` still exits 0 (#2790). Shaped after Resilience4j's
* `failureRateThreshold` + `minimumNumberOfCalls`: judge a failure RATE, but
* only once enough sub-batches have been attempted for a rate to mean anything
* — a 3-node repo losing its single sub-batch is 100% and must not abort. The
* rate sits below a live-traffic breaker's 50% (a batch indexer's job is to
* index the whole corpus, not to serve degraded traffic) and above Hadoop's
* single-digit `mapreduce.map.failures.maxpercent` (tolerating transient
* hiccups is the entire point of #2790).
*
* The sample floor SCALES with the run; it used to be a flat 20. A flat floor
* can exceed a run's ENTIRE sub-batch budget, and then the guard is not "not yet
* armed" — it is structurally off. Every resume run has exactly that shape by
* construction: its node set is only the pending ids, so it produces few
* sub-batches no matter how large the repo is, and the run whose whole purpose
* is retrying a suspect endpoint was the one run this guard could never fire in.
* Resilience4j can afford a constant `minimumNumberOfCalls` because a breaker
* sits on an unbounded call stream — the sample always arrives eventually; a
* batch indexer has a finite budget, so the minimum has to be expressed relative
* to it. Hadoop's `mapreduce.map.failures.maxpercent` takes the other extreme, a
* pure percentage with no minimum at all, which is why a 1-of-1 failure reads as
* "100% failed" there. The clamp below keeps both ends honest:
*
* floor = clamp(ceil(totalNodes / subBatchSize / 2), 5, 20)
*
* - ceil(totalNodes / subBatchSize / 2): the run's own sub-batch budget, halved
* — so the rate is judged over the back half of a short run: a real sample,
* still early enough to stop the run before it walks the rest of the corpus.
* Derived from the ACTUAL `subBatchSize` rather than assuming the default of
* 8, because `GITNEXUS_EMBEDDING_SUB_BATCH_SIZE` is exactly the knob an
* operator turns for a constrained or flaky endpoint — the population this
* guard protects — and a floor computed against the wrong divisor is the same
* structurally-off guard, just at a different repo size. Nodes that chunk
* into several chunks only produce MORE sub-batches, so the guard arms
* earlier still.
* - min 5: under five attempts a "rate" is one or two coin flips. A 3-node repo
* losing its single sub-batch is 100% and must be tolerated, not aborted —
* the exact case the old flat floor was written to protect, kept intact.
* - max 20 (the previous flat value, reached once a run has 40 sub-batches):
* large runs keep today's behavior byte for byte, and a purely proportional
* floor would perversely WEAKEN the guard at scale — half a 20k-node repo's
* sub-batches is 1250 of them (at subBatchSize 8) eaten before a rate could
* abort anything.
*
* Honest limit: this bounds the failure RATE, never the absolute loss. A run
* that steadily fails just under 25% of its sub-batches never aborts, so a
* single run may drop up to a quarter of the corpus and still exit 0 behind one
* warning line. That is by design — the dropped nodes are left holding ZERO rows
* so the next incremental run re-embeds them, which beats failing the whole
* analyze — but it is the real ceiling on what this guard promises.
*
* ponytail: lifetime ratio, evaluated only when a sub-batch fails. A sliding
* window would spot an endpoint that degrades late in a long run sooner —
* upgrade if partial-corpus reports keep arriving from runs that stayed under
* this bar.
*/
const MAX_SUB_BATCH_FAILURE_RATIO = 0.25;
const MIN_SUB_BATCHES_FLOOR = 5;
const MAX_SUB_BATCHES_FLOOR = 20;
/** Judge the rate over the back half of the run: half its sub-batch budget. */
const RATIO_GUARD_SAMPLE_FRACTION_DIVISOR = 2;
const minSubBatchesBeforeFailureRatioGuard = (totalNodes: number, subBatchSize: number): number =>
Math.min(
MAX_SUB_BATCHES_FLOOR,
Math.max(
MIN_SUB_BATCHES_FLOOR,
Math.ceil(totalNodes / subBatchSize / RATIO_GUARD_SAMPLE_FRACTION_DIVISOR),
),
);
/**
* The cumulative-ratio abort (#2790). Deliberately NOT the raw endpoint error
* the consecutive ceiling rethrows: the operator's next action differs — one bad
* batch is noise, a quarter of the corpus failing is an endpoint to fix before
* re-running. The retained streak error still names the real defect, carried in
* both the message and `cause`.
*/
const failureRatioAbortError = (failures: number, attempted: number, cause: unknown): Error =>
new Error(
`[embed] Aborting: ${failures} of ${attempted} embed sub-batches failed ` +
`(${Math.round((failures / attempted) * 100)}%, limit ` +
`${Math.round(MAX_SUB_BATCH_FAILURE_RATIO * 100)}%) — too much of the corpus is failing to ` +
`embed for this run to produce a usable index. Fix the embedding endpoint and re-run. ` +
`Underlying failure: ${cause instanceof Error ? cause.message : String(cause)}`,
{ cause },
);
/**
* Both failed: the run is aborting AND the cleanup DELETE for the dropped nodes
* threw (busy/read-only DB). The abort error stays primary — it is the one that
* names the real defect and the one the operator must act on — and is kept as
* `cause`; the cleanup failure is appended, because it also means those nodes may
* still hold partial rows carrying the CURRENT contentHash, which a later
* incremental run would read as fresh (the corruption the cleanup exists to
* prevent). Without this, a failing DELETE replaced the endpoint error outright.
*/
const abortWithFailedCleanupError = (abortErr: unknown, cleanupErr: unknown): Error =>
new Error(
`${abortErr instanceof Error ? abortErr.message : String(abortErr)} — additionally, ` +
`cleanup of the dropped nodes' embedding rows FAILED, so some may still hold partial ` +
`rows that a later incremental run would read as fresh: ` +
`${cleanupErr instanceof Error ? cleanupErr.message : String(cleanupErr)}`,
{ cause: abortErr },
);
/**
* Cancellation is an instruction, not endpoint flakiness — it must abort the run
* instead of being absorbed as a tolerable sub-batch failure. Checked on the
* error shape as well as on our own signal, because an embed call handed the
* signal rejects with a `DOMException`/`Error` named `AbortError` rather than
* flipping anything the pipeline owns (e.g. a host-supplied transport signal).
*/
const isCancellationError = (err: unknown): boolean =>
typeof err === 'object' && err !== null && (err as { name?: unknown }).name === 'AbortError';
export interface EmbeddingPipelineCheckpoint {
nodesProcessed: number;
totalNodes: number;
@@ -509,6 +657,7 @@ export const runEmbeddingPipeline = async (
chunksProcessed: 0,
vectorIndexReady,
semanticMode: vectorIndexReady ? 'vector-index' : 'exact-scan',
failedNodeIds: [],
};
}
@@ -520,7 +669,38 @@ export const runEmbeddingPipeline = async (
batchSize,
Math.ceil(checkpointEveryNodes / batchSize) * batchSize,
);
// Nodes that finished with a COMPLETE row set. Distinct from the number of
// nodes traversed (see `traversedNodes` in the loop): a node whose chunks lost
// a sub-batch is walked past but NOT processed, and must not be counted as if
// its embeddings exist.
let processedNodes = 0;
// Run-level: every node whose rows were dropped by the sub-batch failure path.
const failedNodeIds = new Set<string>();
// Run-level, not per-batch: a dead endpoint stays dead across batch boundaries.
let consecutiveSubBatchFailures = 0;
// Run-level tally behind MAX_SUB_BATCH_FAILURE_RATIO (#2790). Lifetime, never
// reset — a steadily-degraded endpoint is exactly the shape the consecutive
// counter's reset-on-success blinds it to.
let subBatchesAttempted = 0;
let subBatchFailures = 0;
// Sized from this run's own node set AND its configured sub-batch size, not
// a flat constant, so a resume run (whose node set is only the pending ids)
// and an operator-tuned `subBatchSize` both still arm the guard — see
// minSubBatchesBeforeFailureRatioGuard.
const minSubBatchesForRatioGuard = minSubBatchesBeforeFailureRatioGuard(
totalNodes,
finalConfig.subBatchSize,
);
// #2790: the FIRST error of the current streak — the one the ceiling rethrows.
// Later failures in a streak degrade into generic noise: a misconfigured
// endpoint trips the HTTP client's circuit breaker after 3 rejections, so
// failures 4 and 5 are "circuit open … retry in 30s" (advice to wait for a
// condition that never changes) while the first still names the real defect
// ("unexpected response shape"). Boxed rather than held as a bare `unknown`
// so a thrown `undefined` can't be misread as "no error". Cleared with the
// counter below, so the rethrown error always belongs to the streak that
// actually tripped the ceiling.
let streakFirstError: { err: unknown } | undefined;
onProgress({
phase: 'embedding',
@@ -631,20 +811,69 @@ export const runEmbeddingPipeline = async (
// Embed chunk texts in sub-batches to control memory
const EMBED_SUB_BATCH = finalConfig.subBatchSize;
// Nodes that lost at least one chunk in this batch. Collected as NODE ids,
// not sub-batch indices: `allTexts`/`allUpdates` are flat over the whole
// batch with no node alignment, so one node's chunks can straddle a
// sub-batch boundary and a failure never maps cleanly onto whole nodes.
const batchFailedNodeIds = new Set<string>();
// Set instead of thrown immediately so the cleanup DELETE below still runs.
// Carries EITHER abort reason — the consecutive-ceiling rethrow or the
// cumulative-ratio abort — hence the reason-neutral name.
let abortError: { err: unknown } | undefined;
for (let si = 0; si < allTexts.length; si += EMBED_SUB_BATCH) {
const subTexts = allTexts.slice(si, si + EMBED_SUB_BATCH);
const subUpdates = allUpdates.slice(si, si + EMBED_SUB_BATCH);
subBatchesAttempted += 1;
let embeddings: Float32Array[];
try {
embeddings = await embedBatch(subTexts, { signal: pipelineOptions.signal });
} catch (embedErr) {
logger.error(
{ embedErr },
`❌ embedBatch failed for ${subTexts.length} texts (first: "${subTexts[0]?.substring(0, 80)}..."):`,
// Never swallow a cancel as a "tolerable" failure — the operator asked
// us to stop, and the run has no business continuing (or deleting rows).
if (pipelineOptions.signal?.aborted || isCancellationError(embedErr)) throw embedErr;
// #2790: one transient hiccup used to discard an entire multi-hour run.
// Warn (not error — the run survives), drop this sub-batch's nodes, and
// keep going; the cleanup DELETE after this loop makes the drop safe.
const droppedNodeIds = new Set(subUpdates.map((u) => u.nodeId));
for (const nodeId of droppedNodeIds) batchFailedNodeIds.add(nodeId);
consecutiveSubBatchFailures += 1;
subBatchFailures += 1;
const streakError = (streakFirstError ??= { err: embedErr });
// Log under `err` so pino's standard serializer keeps the message and
// stack — the previous `{ embedErr }` key serialized to `{}` (#2114).
logger.warn(
{ err: embedErr },
`⚠️ embedBatch failed for ${subTexts.length} texts (first: "${subTexts[0]?.substring(0, 80)}...") — ` +
`dropping ${droppedNodeIds.size} node(s) from this run so the next run re-embeds them ` +
`(${consecutiveSubBatchFailures}/${MAX_CONSECUTIVE_SUB_BATCH_FAILURES} consecutive, ` +
`${subBatchFailures}/${subBatchesAttempted} sub-batches failed this run):`,
);
throw embedErr;
if (consecutiveSubBatchFailures >= MAX_CONSECUTIVE_SUB_BATCH_FAILURES) {
// Bail out — but via `break`, not `throw`, so the cleanup DELETE below
// still runs: sub-batches that succeeded before the endpoint went dark
// left partial rows carrying the CURRENT hash, which is exactly the
// read-as-fresh corruption this path exists to prevent.
abortError = streakError;
break;
}
if (
subBatchesAttempted >= minSubBatchesForRatioGuard &&
subBatchFailures / subBatchesAttempted >= MAX_SUB_BATCH_FAILURE_RATIO
) {
// Same break-then-cleanup path as the ceiling: a run aborting because
// the endpoint is shedding a quarter of the corpus still must not leave
// this batch's touched nodes half embedded (#2790).
abortError = {
err: failureRatioAbortError(subBatchFailures, subBatchesAttempted, streakError.err),
};
break;
}
continue;
}
consecutiveSubBatchFailures = 0;
streakFirstError = undefined;
const dbUpdates = subUpdates.map((u, i) => ({
...u,
@@ -655,10 +884,40 @@ export const runEmbeddingPipeline = async (
throwIfCancelled();
}
processedNodes += batch.length;
totalChunks += allUpdates.length;
// A failed sub-batch leaves the node HALF embedded, and half is worse than
// none: this batch's stale rows were already deleted above, a node's chunks
// can straddle a sub-batch boundary, and every surviving row carries the
// CURRENT contentHash. Both downstream hash-map builders collapse a node's
// rows to one entry (last row wins), so a half-embedded node would read as
// FRESH on every future run and its missing chunks would never come back.
// Deleting ALL of its rows leaves zero, which the incremental filter reads
// as "New node — needs embedding" — it self-heals next run. Kept inside the
// per-batch iteration so the inconsistency window stays one batch wide,
// matching the U6 / KTD7 delete-window rationale above.
if (batchFailedNodeIds.size > 0) {
try {
await deleteStaleEmbeddingRows(executeWithReusedStatement, [...batchFailedNodeIds]);
} catch (cleanupErr) {
// A failing cleanup DELETE must never swallow a pending abort: that
// error names the actual defect (the endpoint), while a busy or
// read-only DB is a second, downstream symptom. Without this the
// `throw` below is unreachable and the run reports only the DELETE.
if (!abortError) throw cleanupErr;
throw abortWithFailedCleanupError(abortError.err, cleanupErr);
}
for (const nodeId of batchFailedNodeIds) failedNodeIds.add(nodeId);
}
if (abortError) throw abortError.err;
const embeddingProgress = 20 + (processedNodes / totalNodes) * 70;
// Nodes walked past in this run so far. Drives progress percent and the
// checkpoint cadence, which must stay monotonic and hit `totalNodes`
// exactly — `processedNodes` can now lag behind and would silently stop the
// window-aligned and terminal checkpoints from ever firing.
const traversedNodes = batchIndex + batch.length;
processedNodes += batch.length - batchFailedNodeIds.size;
totalChunks += allUpdates.filter((u) => !batchFailedNodeIds.has(u.nodeId)).length;
const embeddingProgress = 20 + (traversedNodes / totalNodes) * 70;
onProgress({
phase: 'embedding',
percent: Math.round(embeddingProgress),
@@ -670,7 +929,7 @@ export const runEmbeddingPipeline = async (
if (
pipelineOptions.onCheckpoint &&
(processedNodes % checkpointWindowNodeCount === 0 || processedNodes === totalNodes)
(traversedNodes % checkpointWindowNodeCount === 0 || traversedNodes === totalNodes)
) {
await pipelineOptions.onCheckpoint({
nodesProcessed: processedNodes,
@@ -683,10 +942,12 @@ export const runEmbeddingPipeline = async (
// Phase 4: Create vector index
throwIfCancelled();
// Report the nodes actually embedded, not `totalNodes`: with a flaky endpoint
// those differ, and a caller is about to read this number as truth.
onProgress({
phase: 'indexing',
percent: 90,
nodesProcessed: totalNodes,
nodesProcessed: processedNodes,
totalNodes,
});
@@ -699,26 +960,34 @@ export const runEmbeddingPipeline = async (
onProgress({
phase: 'ready',
percent: 100,
nodesProcessed: totalNodes,
nodesProcessed: processedNodes,
totalNodes,
});
if (failedNodeIds.size > 0) {
logger.warn(
`⚠️ Embedding pipeline finished with ${failedNodeIds.size} node(s) dropped after embed failures; ` +
'their rows were removed so the next run re-embeds them.',
);
}
if (isDev) {
logger.info(
`✅ Embedding pipeline complete! (${totalChunks} chunks from ${totalNodes} nodes)`,
`✅ Embedding pipeline complete! (${totalChunks} chunks from ${processedNodes}/${totalNodes} nodes)`,
);
}
return {
nodesProcessed: totalNodes,
nodesProcessed: processedNodes,
chunksProcessed: totalChunks,
vectorIndexReady,
semanticMode: vectorIndexReady ? 'vector-index' : 'exact-scan',
failedNodeIds: [...failedNodeIds],
};
} catch (error) {
const errorMessage = error instanceof Error ? error.message : 'Unknown error';
if (isDev) {
logger.error({ error }, '❌ Embedding pipeline error:');
// `err`, not `error` — see the serializer note in queryEmbeddableNodes (#2114).
logger.error({ err: error }, '❌ Embedding pipeline error:');
}
onProgress({
@@ -785,9 +1054,10 @@ export const semanticSearch = async (
}
if (bestChunks.size === 0) {
const countRows = await executeQuery(
`MATCH (e:${EMBEDDING_TABLE_NAME}) RETURN count(e) AS cnt`,
);
// The Cypher only. NOT `measurePersistedEmbeddingCount`: its tri-state
// exists so a publisher never writes a fabricated 0, whereas here `?? 0`
// is the right answer — an unknown count simply skips the exact scan.
const countRows = await executeQuery(EMBEDDING_COUNT_CYPHER);
const countRow = countRows[0];
const embeddingCount = Number(countRow?.cnt ?? countRow?.[0] ?? 0);
const exactLimit = getExactScanLimit();
+143 -22
View File
@@ -11,7 +11,12 @@
* via `AbortSignal.timeout` on the underlying fetch.
*/
import { CircuitOpenError, ResilientFetchExhaustedError, resilientFetch } from 'gitnexus-shared';
import {
CircuitOpenError,
ResilientFetchExhaustedError,
isTerminalNetworkError,
resilientFetch,
} from 'gitnexus-shared';
const DEFAULT_HTTP_TIMEOUT_MS = 180_000;
const MAX_HTTP_TIMEOUT_MS = 300_000;
@@ -310,6 +315,46 @@ const isEmbeddingItem = (item: unknown): item is EmbeddingItem =>
item !== null &&
Array.isArray((item as { embedding?: unknown }).embedding);
/**
* Module-private signal that a 2xx response carried a body this client cannot
* use (unparseable, or parseable but wrong-shaped). Deliberately not exported:
* it never escapes {@link httpEmbedBatch}, which converts it into the
* user-facing {@link HttpEmbeddingError} it carries in `terminalMessage`.
*
* Thrown from inside the `fetchImpl` callback so `resilientFetch` classifies it
* as `retryable-network` — a truncated or HTML body is the same class of
* endpoint failure as a 503 and deserves the same backoff loop, and routing it
* through the retry loop also makes the circuit breaker see it as a failure
* instead of erasing the outage signal with `recordSuccess()`. See #2790.
*/
class RetryableEmbeddingBodyError extends Error {
constructor(
readonly terminalMessage: string,
options?: { cause?: unknown },
) {
super(terminalMessage, options?.cause !== undefined ? { cause: options.cause } : undefined);
this.name = 'RetryableEmbeddingBodyError';
}
}
/**
* Build the message for a 2xx body carrying the wrong number of vectors.
*
* Hoisted to module scope because two layers report this same fault — the
* in-loop check inside {@link httpEmbedBatch} and the defensive backstop in
* {@link httpEmbed} — and the operator must see one wording regardless of which
* one catches it. Names both counts: "0 vectors for 64 texts" is actionable,
* "unexpected response shape" is not.
*/
const countMismatchMessage = (
received: number,
expected: number,
safeEndpoint: string,
batchIndex: number,
): string =>
`Embedding endpoint returned ${received} vectors for ${expected} texts ` +
`(${safeEndpoint}, batch ${batchIndex})`;
/**
* Send a single batch of texts to the embedding endpoint with retry.
*
@@ -347,7 +392,21 @@ const httpEmbedBatch = async (
requestBody.dimensions = dimensions;
}
// Built on demand, not up front. Both describe faults, so in a healthy run —
// every call of it — they are garbage, and each `safeUrl` parses a URL. A
// 300k-chunk run at subBatchSize 8 makes ~37.5k `httpEmbedBatch` calls, so
// eager construction spent two `new URL()` per call describing failures that
// never happened. Same shape as `countMismatchMessage` above.
const unparseableMessage = (): string =>
`Embedding endpoint returned an unparseable response (${safeUrl(url)}, batch ${batchIndex})`;
const unexpectedShapeMessage = (): string =>
`Embedding endpoint returned an unexpected response shape (${safeUrl(url)}, batch ${batchIndex})`;
let resp: Response;
// Set by `fetchImpl` on the attempt that produced a usable body. Reads back
// as `undefined` only on a path that must already have thrown — see the
// defensive check after the retry loop.
let parsed: EmbeddingItem[] | undefined;
try {
throwIfAborted(requestOptions.signal);
resp = await resilientFetch(
@@ -368,7 +427,54 @@ const httpEmbedBatch = async (
const signal = requestOptions.signal
? AbortSignal.any([requestOptions.signal, timeoutSignal])
: timeoutSignal;
return globalThis.fetch(input, { ...init, signal });
const attemptResp = await globalThis.fetch(input, { ...init, signal });
// Non-OK bodies are none of our business: hand the response straight
// back so `resilientFetch` keeps classifying 4xx/5xx/429 unchanged.
if (!attemptResp.ok) return attemptResp;
// The body is read *here*, inside the retried callback, rather than
// after `resilientFetch` returns. A reachable-but-wrong endpoint (a
// captive portal, a non-embeddings service, a truncated stream) can
// answer 200 with HTML or half a JSON document; parsing outside the
// loop made that a one-shot terminal failure while a 503 got three
// attempts. Throwing from in here gives a bad body the same backoff
// and the same breaker accounting as any other endpoint fault (#2790).
let payload: { data: EmbeddingItem[] };
try {
payload = (await attemptResp.json()) as { data: EmbeddingItem[] };
} catch (err) {
// Not every `.json()` rejection is a parse error: the per-attempt
// signal (`AbortSignal.any([caller, AbortSignal.timeout(...)])`) is
// wired to the body stream, so a stalled body rejects with the abort
// reason. Re-raise those untouched — `isTerminalNetworkError` is
// `resilientFetch`'s own predicate, so this test agrees with
// `classifyOutcome` by construction. Wrapping one would flip its
// verdict from `terminal-network` (returned without retry AND
// without touching the breaker, via `recordNeutral()`) to
// `retryable-network` (retried, then `breaker.recordFailure()`): the
// same timeout would take 3 attempts instead of 1, count toward the
// process-global `embeddings-http` breaker, and reach the operator as
// "unparseable response" so they never reach for the timeout knob.
if (isTerminalNetworkError(err)) throw err;
throw new RetryableEmbeddingBodyError(unparseableMessage(), { cause: err });
}
if (!Array.isArray(payload?.data) || !payload.data.every(isEmbeddingItem)) {
throw new RetryableEmbeddingBodyError(unexpectedShapeMessage());
}
// Cardinality belongs *inside* the retry loop. `every(isEmbeddingItem)`
// is vacuously true for `[]` and true for any array shorter than the
// request, so a 200 carrying `{"data": []}` — or half the vectors —
// used to be classified `success`, call `breaker.recordSuccess()`
// (erasing the outage signal), and only then fail terminally after a
// single attempt. A short body is a truncated body: same backoff, same
// breaker accounting as any other endpoint fault (#2790).
if (payload.data.length !== batch.length) {
throw new RetryableEmbeddingBodyError(
countMismatchMessage(payload.data.length, batch.length, safeUrl(url), batchIndex),
);
}
parsed = payload.data;
return attemptResp;
},
breakerKey: HTTP_BREAKER_KEY,
retry: {
@@ -390,6 +496,12 @@ const httpEmbedBatch = async (
{ cause: err },
);
}
// Retries are exhausted on a bad 2xx body. Surface the message the sentinel
// carried, keeping the underlying parse error in `cause` only — the body
// text must never reach the `sanitizeReason` fallback and leak to stderr.
if (err instanceof RetryableEmbeddingBodyError) {
throw new HttpEmbeddingError(err.terminalMessage, { cause: err.cause });
}
if (err instanceof CircuitOpenError) {
throw new HttpEmbeddingError(
`Embedding endpoint circuit open (${safeUrl(url)}, batch ${batchIndex}): retry in ${Math.ceil(err.retryAfterMs / 1000)}s`,
@@ -425,25 +537,14 @@ const httpEmbedBatch = async (
);
}
// A reachable-but-wrong endpoint (e.g. a captive portal or a non-embeddings
// service) can answer 200 with an HTML/truncated body. Parse inside the
// typed-error boundary so that lands as an endpoint failure the CLI can
// classify, not a raw SyntaxError/TypeError on the generic stack-dump path.
let data: { data: EmbeddingItem[] };
try {
data = (await resp.json()) as { data: EmbeddingItem[] };
} catch (err) {
throw new HttpEmbeddingError(
`Embedding endpoint returned an unparseable response (${safeUrl(url)}, batch ${batchIndex})`,
{ cause: err },
);
if (parsed === undefined) {
// Defensively unreachable: an OK response either sets `parsed` or throws
// out of `fetchImpl`. Kept so the narrowing holds without a non-null
// assertion, and so a future `resilientFetch` change can't return an
// unvalidated body silently.
throw new HttpEmbeddingError(unparseableMessage());
}
if (!Array.isArray(data?.data) || !data.data.every(isEmbeddingItem)) {
throw new HttpEmbeddingError(
`Embedding endpoint returned an unexpected response shape (${safeUrl(url)}, batch ${batchIndex})`,
);
}
return data.data;
return parsed;
};
/**
@@ -482,10 +583,15 @@ export const httpEmbed = async (
config.timeoutMs,
);
// Defensive backstop, deliberately kept: `httpEmbedBatch` now rejects a
// short body from *inside* the retry loop (through the same
// `countMismatchMessage`), so in practice this branch is unreachable. It
// stays so a future change to the in-loop check can't silently hand a short
// vector list to the caller — that failure would land as a Kuzu/FLOAT[N]
// error far from its cause.
if (items.length !== batch.length) {
throw new HttpEmbeddingError(
`Embedding endpoint returned ${items.length} vectors for ${batch.length} texts ` +
`(${safeUrl(url)}, batch ${batchIndex})`,
countMismatchMessage(items.length, batch.length, safeUrl(url), batchIndex),
);
}
@@ -493,6 +599,18 @@ export const httpEmbed = async (
const vec = new Float32Array(item.embedding);
// Fail fast on dimension mismatch rather than inserting bad vectors
// into the FLOAT[N] column which would cause a cryptic Kuzu error.
//
// Unlike the cardinality check this one stays *outside* the retry loop,
// deliberately. The expected width is `config.dimensions ?? DEFAULT_DIMS`
// (GITNEXUS_EMBEDDING_DIMS), which is NOT the `dimensions` value
// `httpEmbedBatch` receives — that is `config.requestDimensions`, which
// `GITNEXUS_EMBEDDING_REQUEST_DIMS` can set to a different number or to
// `undefined` (`omit`). More importantly a width mismatch is an operator
// *configuration* error, not an endpoint fault: retrying it three times
// can never change the answer, and routing it through the retry loop
// would count a healthy endpoint's responses toward the shared circuit
// breaker. The message is an actionable config hint, so it is terminal
// on the first attempt by design (#2790).
const expected = config.dimensions ?? DEFAULT_DIMS;
if (vec.length !== expected) {
const hint = config.dimensions
@@ -539,6 +657,9 @@ export const httpEmbedQuery = async (
config.minIntervalMs,
config.timeoutMs,
);
// Defensive backstop like the `httpEmbed` one above: an empty `data` array is
// now a cardinality mismatch (0 vectors for 1 text) rejected and retried
// inside `httpEmbedBatch`, so this branch is unreachable in practice.
if (!items.length) {
throw new HttpEmbeddingError(`Embedding endpoint returned empty response (${safeUrl(url)})`);
}
+83 -8
View File
@@ -11,6 +11,7 @@ import type {
CrossRepoImpact,
GroupConfig,
GroupImpactResult,
GroupImpactTruncationReason,
MatchType,
OutOfScopeLink,
} from './types.js';
@@ -29,6 +30,7 @@ import {
readBridgeMeta,
} from './bridge-db.js';
import { BRIDGE_SCHEMA_VERSION } from './bridge-schema.js';
import { compareCodeUnits } from '../../lib/utils.js';
// High limit for the local phase of group impact so collectImpactSymbolUids
// sees (nearly) all symbols. Bypasses the MCP-facing default of 100.
@@ -40,6 +42,26 @@ export const MAX_SUPPORTED_CROSS_DEPTH = 1;
/** Default wall-clock budget for the Phase 1 `impact` leg when callers omit `timeoutMs`. */
export const DEFAULT_LOCAL_IMPACT_TIMEOUT_MS = 30_000;
/**
* Cap on neighbour fan-outs attempted per group-impact request.
*
* The bound used to be the wall clock alone, which made the cutoff a function
* of machine load: an idle host traversed more crossings and `mergeRisk`
* escalated to CRITICAL at three, while a loaded host stopped at two and
* reported HIGH or lower — same graph, same arguments, different verdict
* (#2787). A count is deterministic, and the neighbour list carries a total
* order over the full crossing identity (confidence DESC, then repo, uid,
* contract), so the cap keeps the strongest crossings rather than an arbitrary
* prefix.
*
* 50 is borrowed from `MAX_CROSSINGS_TO_TRY` (cross-trace.ts) on cost, not on
* scope — that one caps ContractLinks per repo pair inside a trace, this one
* caps the total across all neighbour repos in one impact call, and group impact
* had no numeric cap at all before #2787. The wall clock stays as a hang
* backstop below; this is the bound that normally binds.
*/
export const MAX_NEIGHBOR_FANOUT = 50;
const CY_NEIGHBORS_UPSTREAM = `
MATCH (consumer:Contract)-[l:ContractLink]->(provider:Contract)
WHERE provider.repo = $localRepo
@@ -358,6 +380,26 @@ export function mergeRisk(localRisk: string, cross: CrossRepoImpact[]): string {
return localRisk;
}
/**
* Build the truncation fields every `runGroupImpact` return path shares.
*
* `riskEpistemic` must follow `truncated` mechanically: it is the marker that
* tells a caller the `risk` value is a floor rather than a verdict, and
* `mergeRisk` can only under-report once a crossing is dropped. Attaching it at
* each return let two of the four paths set `truncated` without it, so a
* truncated result read as complete — deriving it in one place is what keeps
* the invariant from drifting again (#2787).
*/
function truncationFields(
truncated: boolean,
// Only read on the truncated branch, so the not-truncated call sites omit it
// rather than passing a reason that is thrown away.
reasonIfTruncated: GroupImpactTruncationReason = 'partial',
): Pick<GroupImpactResult, 'truncated' | 'truncationReason' | 'riskEpistemic'> {
if (!truncated) return { truncated: false };
return { truncated: true, truncationReason: reasonIfTruncated, riskEpistemic: 'lower-bound' };
}
function addCrossImpact(cross: CrossRepoImpact[], candidate: CrossRepoImpact): void {
const sameBoundary = (entry: CrossRepoImpact): boolean =>
entry.repo_path === candidate.repo_path && entry.contract.id === candidate.contract.id;
@@ -448,7 +490,18 @@ export async function resolveBridgeNeighbors(
const n = rowToNeighbor(raw);
if (n) neighbors.push(n);
}
neighbors.sort((a, b) => b.confidence - a.confidence);
// Sort on the FULL crossing identity — the same triple the fan-out dedups on
// below (`repo\0uid\0contractId`). Two links that share (confidence, repo,
// uid) but differ in contract both survive that dedup, so leaving contractId
// out of the comparator makes them compare 0 and fall back to raw bridge row
// order, which is what decides who lands past MAX_NEIGHBOR_FANOUT (#2787).
neighbors.sort(
(a, b) =>
b.confidence - a.confidence ||
compareCodeUnits(a.neighborRepo, b.neighborRepo) ||
compareCodeUnits(a.neighborUid, b.neighborUid) ||
compareCodeUnits(a.contractId, b.contractId),
);
return neighbors;
}
@@ -514,7 +567,7 @@ export async function runGroupImpact(
group: name,
cross: [],
outOfScope: [],
truncated: true,
...truncationFields(true, 'timeout'),
truncatedRepos: [],
summary: {
direct: 0,
@@ -524,7 +577,6 @@ export async function runGroupImpact(
},
risk: 'UNKNOWN',
timeoutMs,
truncationReason: 'timeout',
crossDepthWarning,
};
}
@@ -548,7 +600,7 @@ export async function runGroupImpact(
group: name,
cross: [],
outOfScope: [],
truncated: false,
...truncationFields(false),
truncatedRepos: [],
summary: {
direct: 0,
@@ -571,7 +623,7 @@ export async function runGroupImpact(
group: name,
cross: [],
outOfScope: [],
truncated: Boolean((local as { partial?: boolean }).partial),
...truncationFields(Boolean((local as { partial?: boolean }).partial), 'partial'),
truncatedRepos: [],
summary: {
direct: s.direct ?? 0,
@@ -581,7 +633,6 @@ export async function runGroupImpact(
},
risk: String((local as { risk?: string }).risk ?? 'LOW'),
timeoutMs,
truncationReason: (local as { partial?: boolean }).partial ? 'partial' : undefined,
crossDepthWarning,
};
}
@@ -593,6 +644,10 @@ export async function runGroupImpact(
const cross: CrossRepoImpact[] = [];
const outOfScope: OutOfScopeLink[] = [];
const truncatedRepos: string[] = [];
/** Real `impactByUid` fan-outs issued — what MAX_NEIGHBOR_FANOUT bounds. */
let attemptedFanouts = 0;
/** True when the wall-clock backstop, not the count cap, cut the fan-out. */
let fanoutTimedOut = false;
try {
const neighbors = await resolveBridgeNeighbors(handle, {
@@ -626,7 +681,18 @@ export async function runGroupImpact(
if (seen.has(key)) continue;
seen.add(key);
// Deterministic bound first: the count decides the cutoff on every host.
// Manifest-only crossings are exempt because they cost no `impactByUid`
// and dropping them regresses #2784.
if (!manifestOnly && attemptedFanouts >= MAX_NEIGHBOR_FANOUT) {
truncatedRepos.push(n.neighborRepo);
continue;
}
// Wall clock second, as a hang backstop only. It is load-dependent by
// construction, so when it is what fired the response must say `timeout`,
// not the generic `partial` it used to report.
if (!manifestOnly && deadline - Date.now() <= 0) {
fanoutTimedOut = true;
truncatedRepos.push(n.neighborRepo);
continue;
}
@@ -676,6 +742,7 @@ export async function runGroupImpact(
// single hung neighbor would pin the request past the clamped
// timeout, which Codex's adversarial review on PR #1331 flagged
// as the still-open half of CodeQL #184 / js/resource-exhaustion.
attemptedFanouts += 1;
const { value: fan, timedOut: neighborTimedOut } = await safeNeighborImpact(
deps.port,
neighborHandle.id,
@@ -690,6 +757,7 @@ export async function runGroupImpact(
remainingMs,
);
if (neighborTimedOut || fan == null) {
if (neighborTimedOut) fanoutTimedOut = true;
truncatedRepos.push(n.neighborRepo);
continue;
}
@@ -721,7 +789,12 @@ export async function runGroupImpact(
group: name,
cross,
outOfScope,
truncated,
// The risk VALUE is never clamped down. `mergeRisk` is monotone increasing
// in the traversed-crossing count, so truncation can only under-report —
// and under-reporting a blast radius is the unsafe direction (an agent told
// LOW proceeds; told CRITICAL it stops). Marking the floor keeps the
// warning intact while making the incompleteness legible.
...truncationFields(truncated, fanoutTimedOut ? 'timeout' : 'partial'),
truncatedRepos: [...new Set(truncatedRepos)],
summary: {
direct: localSum.direct ?? 0,
@@ -731,10 +804,12 @@ export async function runGroupImpact(
},
risk: mergeRisk(localRisk, cross),
timeoutMs,
truncationReason: truncated ? 'partial' : undefined,
crossDepthWarning,
};
return result;
}
export { normalizeServicePrefix, fileMatchesServicePrefix } from './group-path-utils.js';
// Re-exported, not redefined: the single definition lives in lib/utils.ts, but
// this module's own surface is what the #2787 regression test imports it from.
export { compareCodeUnits };
+31 -16
View File
@@ -26,6 +26,7 @@
import { GroupNotFoundError, loadGroupConfig } from './config-parser.js';
import { getGroupDir } from './storage.js';
import { ensureBridgeReady, MAX_SUPPORTED_CROSS_DEPTH } from './cross-impact.js';
import { compareCodeUnits } from '../../lib/utils.js';
import { closeBridgeDb, queryBridge } from './bridge-db.js';
import type {
GroupPdgFlowHop,
@@ -257,7 +258,7 @@ RETURN consumer.symbolUid AS consumerUid,
l.confidence AS confidence,
l.contractId AS contractId,
consumer.type AS contractType
ORDER BY l.confidence DESC
ORDER BY l.confidence DESC, contractId, consumerUid, providerUid
LIMIT ${MAX_CROSSINGS_TO_TRY + 1}
`;
@@ -279,7 +280,7 @@ RETURN consumer.symbolUid AS consumerUid,
l.confidence AS confidence,
l.contractId AS contractId,
consumer.type AS contractType
ORDER BY l.confidence DESC
ORDER BY l.confidence DESC, contractId, consumerUid, providerUid
LIMIT ${MAX_CROSSINGS_TO_TRY + 1}
`;
@@ -321,6 +322,32 @@ function rowToCrossing(r: Record<string, unknown>): CrossingRow | null {
};
}
/**
* Sort a crossing list on the full crossing identity, then apply
* `MAX_CROSSINGS_TO_TRY`.
*
* Both bridge queries already order by confidence DESC; this re-sorts
* defensively so the cap keeps the strongest candidates even if a tuple-mode
* driver reorders rows. The comparator mirrors BOTH queries' `ORDER BY`
* key-for-key — a shorter comparator reproduces less than the order it is
* defending, leaving ties on raw row order (#2787) — so it lives in one place
* rather than being hand-synced across the two callers and two Cypher clauses.
*/
function capCrossings(all: CrossingRow[]): { crossings: CrossingRow[]; truncated: boolean } {
all.sort(
(a, b) =>
b.confidence - a.confidence ||
compareCodeUnits(a.contractId, b.contractId) ||
compareCodeUnits(a.consumerUid, b.consumerUid) ||
compareCodeUnits(a.providerUid, b.providerUid),
);
const truncated = all.length > MAX_CROSSINGS_TO_TRY;
return {
crossings: truncated ? all.slice(0, MAX_CROSSINGS_TO_TRY) : all,
truncated,
};
}
async function listCrossingsBetween(
handle: BridgeHandle,
fromRepo: string,
@@ -335,14 +362,7 @@ async function listCrossingsBetween(
const c = rowToCrossing(raw);
if (c) all.push(c);
}
// The query already orders by confidence DESC; re-sort defensively so the cap
// keeps the strongest candidates even if a tuple-mode driver reorders rows.
all.sort((a, b) => b.confidence - a.confidence);
const truncated = all.length > MAX_CROSSINGS_TO_TRY;
return {
crossings: truncated ? all.slice(0, MAX_CROSSINGS_TO_TRY) : all,
truncated,
};
return capCrossings(all);
}
function destRowToCrossing(r: Record<string, unknown>): CrossingRow | null {
@@ -375,12 +395,7 @@ async function listCrossingsFrom(
const c = destRowToCrossing(raw);
if (c) all.push(c);
}
all.sort((a, b) => b.confidence - a.confidence);
const truncated = all.length > MAX_CROSSINGS_TO_TRY;
return {
crossings: truncated ? all.slice(0, MAX_CROSSINGS_TO_TRY) : all,
truncated,
};
return capCrossings(all);
}
// ── Cross-member symbol resolution ───────────────────────────────────────
@@ -102,6 +102,10 @@ RETURN sym.id AS uid, sym.name AS name, sym.filePath AS filePath,
// degenerate edge-less node NOR inflates the uniqueness count and masks the real
// handler. `LIMIT 2` bounds materialization: distinguishing unique (1) from
// ambiguous (>=2) never needs more than two rows (the count guard stays exact).
//
// determinism: probe — uniqueness discriminator, not a window. `toResolvedSymbol`
// reads row 0 only when `rows.length === 1`; a 2-row result is discarded whole,
// so WHICH two rows came back can never reach a caller.
export const RESOLVE_BY_NAME_QUERY = `
MATCH (n) WHERE labels(n) IN ['Function','Method','CodeElement']
AND n.name = $name AND n.filePath <> ''
@@ -114,6 +118,10 @@ LIMIT 2`;
// the precise rung — it survives aliases and local same-name collisions that a
// repo-wide name lookup cannot, and only resolves on a unique match within that
// module. `LIMIT 2` keeps the uniqueness count exact (see RESOLVE_BY_NAME_QUERY).
//
// determinism: probe — uniqueness discriminator, not a window. Same consumer as
// RESOLVE_BY_NAME_QUERY: `toResolvedSymbol` reads row 0 only when exactly one row
// came back, and discards a 2-row result whole.
export const RESOLVE_IN_MODULE_QUERY = `
MATCH (n) WHERE labels(n) IN ['Function','Method','CodeElement']
AND n.name = $name AND (n.filePath STARTS WITH $fileDot OR n.filePath STARTS WITH $fileSlash)
@@ -23,7 +23,7 @@ export const CUSTOM_CONTRACT_RESOLVE_QUERY = `MATCH (n)
WHERE labels(n) IN ['Function','Method','Class','Interface','Struct','Enum','Trait','Constructor','TypeAlias','Impl','Macro','Union','Typedef','Property','Record','Delegate','Annotation','Template','Const','Static','CodeElement']
AND n.name = $symbolName
RETURN n.id AS uid, n.name AS name, n.filePath AS filePath
ORDER BY n.filePath ASC
ORDER BY n.filePath ASC, n.id ASC
LIMIT 1`;
/**
@@ -242,7 +242,7 @@ export class ManifestExtractor {
`MATCH (handler)-[r:CodeRelation {type: 'HANDLES_ROUTE'}]->(route:Route)
WHERE route.name = $normalized
RETURN handler.id AS uid, handler.name AS name, handler.filePath AS filePath
ORDER BY handler.filePath ASC
ORDER BY handler.filePath ASC, handler.id ASC
LIMIT 1`,
{ normalized },
);
@@ -255,7 +255,7 @@ export class ManifestExtractor {
rows = await executor(
`MATCH (n) WHERE labels(n) IN ['Function','Method','Class','Interface'] AND n.name = $contract
RETURN n.id AS uid, n.name AS name, n.filePath AS filePath
ORDER BY n.filePath ASC
ORDER BY n.filePath ASC, n.id ASC
LIMIT 1`,
{ contract: link.contract },
);
@@ -279,7 +279,7 @@ export class ManifestExtractor {
rows = await executor(
`MATCH (n) WHERE labels(n) IN ['Function','Method'] AND n.name = $methodName
RETURN n.id AS uid, n.name AS name, n.filePath AS filePath
ORDER BY n.filePath ASC
ORDER BY n.filePath ASC, n.id ASC
LIMIT 1`,
{ methodName },
);
@@ -287,7 +287,7 @@ export class ManifestExtractor {
rows = await executor(
`MATCH (n) WHERE labels(n) IN ['Class','Interface'] AND n.name = $serviceName
RETURN n.id AS uid, n.name AS name, n.filePath AS filePath
ORDER BY n.filePath ASC
ORDER BY n.filePath ASC, n.id ASC
LIMIT 1`,
{ serviceName },
);
@@ -304,7 +304,7 @@ export class ManifestExtractor {
rows = await executor(
`MATCH (n) WHERE labels(n) IN ['Module'] AND n.name = $contract
RETURN n.id AS uid, n.name AS name, n.filePath AS filePath
ORDER BY n.filePath ASC
ORDER BY n.filePath ASC, n.id ASC
LIMIT 1`,
{ contract: link.contract },
);
@@ -312,7 +312,7 @@ export class ManifestExtractor {
rows = await executor(
`MATCH (f:File) WHERE f.filePath = $contract
RETURN f.id AS uid, f.name AS name, f.filePath AS filePath
ORDER BY f.filePath ASC
ORDER BY f.filePath ASC, f.id ASC
LIMIT 1`,
{ contract: link.contract },
);
+10 -1
View File
@@ -14,7 +14,10 @@ import {
repoInSubgroup,
} from './group-path-utils.js';
import { getDefaultGitnexusDir, getGroupDir, listGroups, readContractRegistry } from './storage.js';
import { syncGroup } from './sync.js';
// `./sync.js` is imported LAZILY in `groupSync` — see the comment at its call
// site. It statically pulls the six contract extractors and, through them, the
// native tree-sitter binding; a static import here puts all of that on MCP
// server startup, which never syncs.
import { logger } from '../logger.js';
import type {
ContractRegistry,
@@ -338,6 +341,12 @@ export class GroupService {
return { error: `Group "${name}" not found. Run group_list to see configured groups.` };
throw err;
}
// Lazy: `sync.js` reaches the six contract extractors and the native
// tree-sitter binding. `groupSync` is the ONLY consumer — the other seven
// group tools never need it — so deferring it here keeps that closure off
// MCP server startup entirely and off every non-sync group call. The CLI
// already does exactly this at `cli/group.ts`'s sync command.
const { syncGroup } = await import('./sync.js');
const result = await syncGroup(config, {
groupDir,
exactOnly: Boolean(params.exactOnly),
+12
View File
@@ -134,6 +134,18 @@ export interface GroupImpactResult {
cross_repo_hits: number;
};
risk: string;
/**
* `'lower-bound'` when the fan-out was cut short, so `risk` is a FLOOR, not a
* verdict. Same vocabulary as single-repo `impact`'s `epistemic` field.
*
* `mergeRisk` is monotone increasing in the number of traversed crossings, so
* dropping a crossing can only move the reported risk DOWN — a fully
* truncated fan-out returns bare `localRisk`, making a symbol whose blast
* radius crosses a repo boundary indistinguishable from one with no
* cross-repo consumers. Absent this marker a truncated run emits a
* confident-looking `risk` byte-identical to a complete one.
*/
riskEpistemic?: 'lower-bound';
/**
* Milliseconds budget applied to the **Phase 1 local impact** leg (`safeLocalImpact`).
* If the walk hits this wall first, expect `truncationReason: 'timeout'` and a partial `local` payload.
+15 -1
View File
@@ -1,8 +1,10 @@
import { checkpointKind } from './embedding-checkpoint.js';
import type { RepoMeta } from '../storage/repo-manager.js';
export const INDEX_INCOMPLETE_REASONS = [
'incremental-in-progress',
'embedding-checkpoint-pending',
'embedding-count-unverified',
] as const;
export type IndexIncompleteReason = (typeof INDEX_INCOMPLETE_REASONS)[number];
@@ -13,6 +15,18 @@ export function getIndexIncompleteReasons(
): IndexIncompleteReason[] {
const reasons: IndexIncompleteReason[] = [];
if (meta?.incrementalInProgress) reasons.push('incremental-in-progress');
if (meta?.embeddingCheckpoint) reasons.push('embedding-checkpoint-pending');
if (meta?.embeddingCheckpoint) {
// The three checkpoint kinds are not one operator-facing state. GUARDRAILS
// and the runbook document `embedding-checkpoint-pending` as "N node(s)
// lost their embeddings to endpoint failures" — true for 'interrupted' and
// 'partial', a lie for 'unverified-count', whose pending set is empty and
// whose only defect is a count nobody could measure (see the `kind` doc in
// repo-manager.ts).
reasons.push(
checkpointKind(meta.embeddingCheckpoint) === 'unverified-count'
? 'embedding-count-unverified'
: 'embedding-checkpoint-pending',
);
}
return reasons;
}
@@ -0,0 +1,44 @@
/**
* Wire format of the `BasicBlock.callees` / `BasicBlock.calleeIds` cells.
*
* A LEAF module on purpose: it declares two string constants and imports
* nothing. `cfg/emit.ts` produces those cells and `mcp/local/pdg-impact.ts`
* parses them, but `emit.ts` is analyze-only and drags the whole CFG closure
* (reaching-defs, control-dependence, post-dominators, synthetic-escape,
* call-site-harvest) behind it — 8 modules evaluated at every MCP server start
* just to read two strings, since ESM evaluates a module to import any binding
* from it (#2802 review). Splitting the format constants out deletes that cost
* rather than deferring it, which is the same bar #2802 held its own proposals
* to.
*
* `emit.ts` RE-EXPORTS both names, so every existing importer keeps working and
* the producer/consumer pair still resolves to one definition — the drift this
* shared constant exists to prevent stays impossible.
*/
/**
* Reserved token placed in `BasicBlock.callees` when a statement's call sites
* were truncated at the per-statement site cap: the recorded callee list is then
* INCOMPLETE, so over-cap callees are absent. `*` is not a valid identifier
* leaf, so it cannot collide with a real callee name. The impact bridge treats a
* slice containing this sentinel as "callees unknown" and keeps reach
* callgraph-equal (proven), rather than falsely labeling an absent-but-real
* callee `unproven-bridge`.
*/
export const CALLEES_TRUNCATED_SENTINEL = '*';
/**
* Inner separator for the `BasicBlock.calleeIds` cell (resolved callee symbol
* ids). A TAB is used — NOT a space — because resolved ids embed `filePath` and
* C++ overload shape tags with multi-word primitive types (e.g. `unsigned char`,
* `long double`), so an id can legitimately contain a space; a space-joined cell
* then fragments on read and silently drops inter-procedural reach to that
* callee (#2227 tri-review). A tab cannot appear in a tree-sitter-derived id
* token (paths/identifiers/type tokens are tab-free) and round-trips intact
* through `escapeCSVField` (tab is in its preserved set) and the RFC-4180 COPY
* reader (every cell is quoted). Producer (`calleeIdsOfBlock`) and consumer
* (`splitCalleeIds`) both resolve to this single constant so they cannot drift.
* The sibling `callees` (leaf-name) cell stays space-joined — leaf names are
* bare identifiers and never contain a space.
*/
export const CALLEE_ID_SEP = '\t';
+9 -24
View File
@@ -31,34 +31,19 @@ import { augmentForPostDom } from './synthetic-escape.js';
import { DEFAULT_PDG_MAX_SITES_PER_STATEMENT } from './visitors/call-site-harvest.js';
import { calleeIdPosKey } from '../scope-resolution/graph-bridge/callee-id-sink.js';
import { encodeReachingDefReasonPairs } from './reaching-def-reason-codec.js';
import { CALLEES_TRUNCATED_SENTINEL, CALLEE_ID_SEP } from './callee-cell-format.js';
import type { BasicBlockData, BindingEntry, FunctionCfg } from './types.js';
/**
* Reserved token placed in `BasicBlock.callees` when a statement's call sites
* were truncated at {@link DEFAULT_PDG_MAX_SITES_PER_STATEMENT}: the recorded
* callee list is then INCOMPLETE, so over-cap callees are absent. `*` is not a
* valid identifier leaf, so it cannot collide with a real callee name. The
* impact bridge treats a slice containing this sentinel as "callees unknown" and
* keeps reach callgraph-equal (proven), rather than falsely labeling an
* absent-but-real callee `unproven-bridge`.
* Cell-format constants live in the LEAF module `callee-cell-format.ts` and are
* re-exported here so every existing importer keeps working. The consumer side
* (`mcp/local/pdg-impact.ts`) imports them from the leaf directly: importing any
* binding from THIS module evaluates it, and with it the whole analyze-only CFG
* closure — 8 modules on every MCP server start to read two strings (#2802
* review). Producer and consumer still resolve to one definition, so the drift
* these shared constants exist to prevent stays impossible.
*/
export const CALLEES_TRUNCATED_SENTINEL = '*';
/**
* Inner separator for the `BasicBlock.calleeIds` cell (resolved callee symbol
* ids). A TAB is used — NOT a space — because resolved ids embed `filePath` and
* C++ overload shape tags with multi-word primitive types (e.g. `unsigned char`,
* `long double`), so an id can legitimately contain a space; a space-joined cell
* then fragments on read and silently drops inter-procedural reach to that
* callee (#2227 tri-review). A tab cannot appear in a tree-sitter-derived id
* token (paths/identifiers/type tokens are tab-free) and round-trips intact
* through `escapeCSVField` (tab is in its preserved set) and the RFC-4180 COPY
* reader (every cell is quoted). Producer ({@link calleeIdsOfBlock}) and
* consumer (`splitCalleeIds`) import this single constant so they cannot drift.
* The sibling `callees` (leaf-name) cell stays space-joined — leaf names are
* bare identifiers and never contain a space.
*/
export const CALLEE_ID_SEP = '\t';
export { CALLEES_TRUNCATED_SENTINEL, CALLEE_ID_SEP } from './callee-cell-format.js';
/**
* Default per-function CFG edge cap. A pathological generated function could
@@ -54,17 +54,19 @@ import { definitionIdPosition } from '../utils/definition-id.js';
* restricted to function/class-likes, those calls correctly fall
* through to the File-node fallback at the bottom of the walk.
*/
export const CALLER_ANCHOR_LABELS: ReadonlySet<NodeLabel> = new Set<NodeLabel>([
'Function',
'Method',
'Constructor',
'Module',
'Class',
'Interface',
'Struct',
'Enum',
]);
function isCallerAnchorLabel(label: NodeLabel): boolean {
return (
label === 'Function' ||
label === 'Method' ||
label === 'Constructor' ||
label === 'Module' ||
label === 'Class' ||
label === 'Interface' ||
label === 'Struct' ||
label === 'Enum'
);
return CALLER_ANCHOR_LABELS.has(label);
}
function rangeContainsPoint(
@@ -244,38 +244,53 @@ export function buildGraphNodeLookup(graph: KnowledgeGraph): GraphNodeLookup {
return lookup;
}
/**
* Every label {@link buildGraphNodeLookup} registers — and therefore the ONLY
* labels `resolveDefGraphId` can ever return an id for. Both endpoints of every
* scope-resolution edge come from that lookup (the one exception is the File
* fallback in `resolveCallerGraphId`), so this set defines the whole FROM/TO
* surface those edges can produce.
*
* That makes it load-bearing for the LadybugDB relation DDL: a label added here
* without the matching `FROM x TO y` pairs in `RELATION_SCHEMA` crashes
* `analyze` at `assertDeclaredPair` on whichever codebase first emits the pair
* (#2792). `test/unit/schema-pair-coverage.test.ts` derives the required pairs
* from this set and fails in CI instead.
*/
export const LINKABLE_LABELS: ReadonlySet<NodeLabel> = new Set<NodeLabel>([
'Function',
'Method',
'Constructor',
// Program-like module declarations are provider-gated callable-value
// targets and need the same def→graph bridge.
'Module',
'Class',
'Interface',
'Struct',
'Enum',
// Trait nodes are linkable so MRO builders can bridge PHP/Rust trait
// defs between scope-resolution DefIds and the graph's node ids.
// IMPLEMENTS edges from classes to traits are otherwise invisible to
// the scope-resolution MRO pass.
'Trait',
// Variable / Property are linkable too — receiver-bound write/read
// ACCESSES edges target field nodes (e.g. `user.name = "x"` →
// ACCESSES edge to User's `name` Variable/Property node).
'Variable',
'Property',
// Const is linkable so the value-receiver-owner bridge in
// `receiver-bound-calls.ts` Case 5 can translate the scope-resolution
// `Variable` def for `export const fooService = {...}` to the canonical
// `Const:filePath:name` graph node id, against which object-literal
// method symbols register their `ownerId` (PR #1718 / issue #1358).
'Const',
// Macro nodes are linkable so a macro invocation (`log!(…)`) resolved
// via `MacroRegistry` can bridge its scope-resolution `Macro` def to
// the legacy `@definition.macro` graph node and emit the `USES` edge
// (Rust #1934 F72; also covers C/C++ `#define` macro defs).
'Macro',
]);
export function isLinkableLabel(label: NodeLabel): boolean {
return (
label === 'Function' ||
label === 'Method' ||
label === 'Constructor' ||
// Program-like module declarations are provider-gated callable-value
// targets and need the same def→graph bridge.
label === 'Module' ||
label === 'Class' ||
label === 'Interface' ||
label === 'Struct' ||
label === 'Enum' ||
// Trait nodes are linkable so MRO builders can bridge PHP/Rust trait
// defs between scope-resolution DefIds and the graph's node ids.
// IMPLEMENTS edges from classes to traits are otherwise invisible to
// the scope-resolution MRO pass.
label === 'Trait' ||
// Variable / Property are linkable too — receiver-bound write/read
// ACCESSES edges target field nodes (e.g. `user.name = "x"` →
// ACCESSES edge to User's `name` Variable/Property node).
label === 'Variable' ||
label === 'Property' ||
// Const is linkable so the value-receiver-owner bridge in
// `receiver-bound-calls.ts` Case 5 can translate the scope-resolution
// `Variable` def for `export const fooService = {...}` to the canonical
// `Const:filePath:name` graph node id, against which object-literal
// method symbols register their `ownerId` (PR #1718 / issue #1358).
label === 'Const' ||
// Macro nodes are linkable so a macro invocation (`log!(…)`) resolved
// via `MacroRegistry` can bridge its scope-resolution `Macro` def to
// the legacy `@definition.macro` graph node and emit the `USES` edge
// (Rust #1934 F72; also covers C/C++ `#define` macro defs).
label === 'Macro'
);
return LINKABLE_LABELS.has(label);
}
+4 -4
View File
@@ -17,8 +17,8 @@ import { createWriteStream, WriteStream } from 'fs';
import path from 'path';
import type { GraphNode, GraphRelationship } from 'gitnexus-shared';
import { KnowledgeGraph } from '../graph/types.js';
import { NodeTableName, NODE_TABLES, RELATION_SCHEMA } from './schema.js';
import { parseRelationSchemaPairs, RelPairRouter } from './rel-pair-routing.js';
import { NodeTableName, RELATION_SCHEMA } from './schema.js';
import { VALID_NODE_TABLES, parseRelationSchemaPairs, RelPairRouter } from './rel-pair-routing.js';
import { parseTruthyEnv } from '../ingestion/utils/env.js';
import { SYMBOL_NODE_LABELS } from '../ingestion/utils/symbol-labels.js';
import { applyCjkSegmentationIfEnabled } from '../search/cjk-segmentation.js';
@@ -793,13 +793,13 @@ export const streamAllCSVsToDisk = async (
const relRouter = new RelPairRouter(
csvDir,
REL_CSV_HEADER,
new Set<string>(NODE_TABLES),
VALID_NODE_TABLES,
DECLARED_RELATION_PAIRS,
);
try {
let emitted = 0;
for (const rel of orderedRelationships(graph, sortOutput)) {
const pending = relRouter.route(rel.sourceId, rel.targetId, buildRelRow(rel));
const pending = relRouter.route(rel.sourceId, rel.targetId, buildRelRow(rel), rel.type);
if (pending) await pending;
// Periodically hand the event loop back so the overlapped node COPY and
// write-stream drains run instead of starving behind this synchronous
+25 -14
View File
@@ -85,9 +85,10 @@
* ## Correctness contract
*
* Structural sibling of {@link PdgEmitSink}, and reuses its row builder
* (`buildRelRow`), header (`REL_CSV_HEADER`), label derivation (`getNodeLabel`)
* and `RelPairRouter` validity check, so the streamed row SET equals the
* whole-graph emit's and the bulk COPY loads the same rows. Set-level, not
* (`buildRelRow`), header (`REL_CSV_HEADER`) and pair classification
* (`relPairKeyFor`, which is also what `RelPairRouter` routes and skips by), so
* the streamed row SET equals the whole-graph emit's and the bulk COPY loads
* the same rows. Set-level, not
* byte-level: rows stream in emit order and are not re-sorted under
* `GITNEXUS_SORT_GRAPH_OUTPUT`.
*/
@@ -96,8 +97,12 @@ import path from 'path';
import type { GraphNode, GraphRelationship, RelationshipType } from 'gitnexus-shared';
import type { KnowledgeGraph } from '../graph/types.js';
import { DECLARED_RELATION_PAIRS, REL_CSV_HEADER, buildRelRow } from './csv-generator.js';
import { assertDeclaredPair, getNodeLabel } from './rel-pair-routing.js';
import { NODE_TABLES } from './schema.js';
import {
VALID_NODE_TABLES,
assertDeclaredPair,
relPairKeyFor,
splitRelPairKey,
} from './rel-pair-routing.js';
import { DEFAULT_EMIT_CHUNK_ROWS, SyncCsvWriter } from './sync-csv-writer.js';
/**
@@ -229,7 +234,6 @@ export class StreamedRelationshipRemovalError extends Error {
* {@link finalize} once after the pipeline, before `loadGraphToLbug`.
*/
export class GraphEmitSink implements KnowledgeGraph, GraphEmitControl {
private readonly validTables: Set<string>;
private readonly relWriters = new Map<string, SyncCsvWriter>();
/**
* Ids of relationships already streamed. `KnowledgeGraph.addRelationship`
@@ -303,7 +307,6 @@ export class GraphEmitSink implements KnowledgeGraph, GraphEmitControl {
private readonly csvDir: string,
private readonly chunkRows: number = DEFAULT_EMIT_CHUNK_ROWS,
) {
this.validTables = new Set<string>(NODE_TABLES as readonly string[]);
// Own directory, distinct from the PDG sink's: PdgEmitSink wipes and
// recreates its dir on construction and opens with O_EXCL, so a shared dir
// would destroy the other sink's manifest on a combined --pdg run.
@@ -407,16 +410,24 @@ export class GraphEmitSink implements KnowledgeGraph, GraphEmitControl {
}
// Mirror KnowledgeGraph.addRelationship's first-writer-wins dedup.
const fromLabel = getNodeLabel(relationship.sourceId);
const toLabel = getNodeLabel(relationship.targetId);
// Skip edges whose endpoint labels are not valid node tables — mirrors
// `RelPairRouter` exactly so the streamed set matches the whole-graph set.
if (!this.validTables.has(fromLabel) || !this.validTables.has(toLabel)) return;
// Classify + skip via the SHARED `relPairKeyFor`, not a local copy of its
// three lines, so the streamed set cannot drift from the whole-graph set
// `RelPairRouter` produces. `undefined` = an endpoint label is not a node
// table, so the edge is dropped exactly as the router drops it.
const pairKey = relPairKeyFor(relationship.sourceId, relationship.targetId, VALID_NODE_TABLES);
if (pairKey === undefined) return;
const pairKey = `${fromLabel}|${toLabel}`;
assertDeclaredPair(pairKey, DECLARED_RELATION_PAIRS);
assertDeclaredPair(
pairKey,
DECLARED_RELATION_PAIRS,
relationship.type,
relationship.sourceId,
relationship.targetId,
);
let writer = this.relWriters.get(pairKey);
if (writer === undefined) {
// Cold: once per pair, so decoding the key back into its labels is free.
const [fromLabel, toLabel] = splitRelPairKey(pairKey);
try {
writer = new SyncCsvWriter(
path.join(this.csvDir, `rel_${fromLabel}_${toLabel}.csv`),
+33 -6
View File
@@ -19,6 +19,12 @@ import {
STALE_HASH_SENTINEL,
NodeTableName,
} from './schema.js';
// Analyze-only, but reached from MCP startup via `pool-adapter.js`. #2802
// proposed lazy-importing it; rejected — `core/search/bm25-index.ts` statically
// imports `normalizeFtsText` from `csv-generator.js`, and `local-backend.ts`
// dynamically imports bm25-index on the FTS query path, so deferring here
// relocates the startup cost to first query rather than removing it. The
// measured figures live in #2802; they were environment-bound, this is not.
import { streamAllCSVsToDisk, type StreamedCSVResult } from './csv-generator.js';
import type { GraphEmitManifest } from './graph-emit-sink.js';
import type { PdgEmitManifest } from './pdg-emit-sink.js';
@@ -521,6 +527,9 @@ const queryAndDrain = async (targetConn: lbug.Connection, cypher: string): Promi
return isSharedSingletonConn(targetConn) ? withConnLock(run) : run();
};
// determinism: probe — existence only. Every call site runs this through
// `queryAndDrain`, which drains and discards the rows; the ONLY observable is
// whether the read-only shadow replay throws, so no row identity is read.
const READ_ONLY_SHADOW_REPLAY_PROBE = 'MATCH (n) RETURN n LIMIT 1';
/**
@@ -1829,6 +1838,9 @@ export const loadCachedEmbeddings = async (): Promise<{
// Old schema only had (nodeId, embedding); new schema adds (id, chunkIndex, startLine, endLine, contentHash).
// If the query fails (column missing), we return empty cache to force a full rebuild.
try {
// determinism: probe — schema probe, not a sample. `readQueryRows` drains
// the result and the rows are dropped on the floor; only whether the
// new-schema columns parse decides the branch.
const check = await c.query(
`MATCH (e:${EMBEDDING_TABLE_NAME}) RETURN e.nodeId AS nodeId, e.chunkIndex AS chunkIndex LIMIT 1`,
);
@@ -3174,6 +3186,26 @@ export const classifyFtsQueryError = (message: string): FtsQueryFailureClass =>
return 'other';
};
/**
* Build the `QUERY_FTS_INDEX` statement shared by BOTH FTS read paths —
* `queryFTS` below and `queryFTSViaExecutor` in `core/search/bm25-index.ts`
* (the MCP connection-pool path). The two ran byte-identical cypher from two
* places, so every change had to be applied twice in lockstep — the `, node.id`
* ORDER BY tiebreak for #2787 being the latest. Lives beside
* {@link classifyFtsQueryError}, which was already shared for exactly this call.
*/
export const buildFtsQueryCypher = (
tableName: string,
indexName: string,
limit: number,
conjunctive: boolean = false,
): string => `
CALL QUERY_FTS_INDEX('${tableName}', '${indexName}', $query, conjunctive := ${conjunctive})
RETURN node, score
ORDER BY score DESC, node.id
LIMIT ${limit}
`;
/**
* Query a full-text search index
* @param tableName - The node table name
@@ -3196,12 +3228,7 @@ export const queryFTS = async (
throw new Error('LadybugDB not initialized. Call initLbug first.');
}
const cypher = `
CALL QUERY_FTS_INDEX('${tableName}', '${indexName}', $query, conjunctive := ${conjunctive})
RETURN node, score
ORDER BY score DESC
LIMIT ${limit}
`;
const cypher = buildFtsQueryCypher(tableName, indexName, limit, conjunctive);
try {
const rows = await executePrepared(cypher, { query });
+28 -13
View File
@@ -24,8 +24,8 @@
* `storage/parsedfile-store.ts`.
*
* Byte-identity (issue acceptance): the sink reuses the SAME shared row
* builders (`buildBasicBlockRow`, `buildRelRow`) and label derivation
* (`getNodeLabel`) as `streamAllCSVsToDisk`, so the streamed CSV line SET is
* builders (`buildBasicBlockRow`, `buildRelRow`) and pair classification
* (`relPairKeyFor`) as `streamAllCSVsToDisk`, so the streamed CSV line SET is
* identical to the whole-graph emit's, and the bulk COPY loads the same rows →
* the persisted graph is SET-identical and DB-identical. The guarantee is
* set-level, not byte-level on the CSV file: the sink streams rows in emit
@@ -54,9 +54,14 @@ import {
buildBasicBlockRow,
buildRelRow,
} from './csv-generator.js';
import { assertDeclaredPair, getNodeLabel } from './rel-pair-routing.js';
import {
VALID_NODE_TABLES,
assertDeclaredPair,
relPairKeyFor,
splitRelPairKey,
} from './rel-pair-routing.js';
import { DEFAULT_EMIT_CHUNK_ROWS, SyncCsvWriter } from './sync-csv-writer.js';
import { NODE_TABLES, type NodeTableName } from './schema.js';
import { type NodeTableName } from './schema.js';
/**
* PDG edge types streamed per-file (all intra-block BasicBlock→BasicBlock).
@@ -98,7 +103,6 @@ export interface PdgEmitManifest {
* `--pdg` emit, then {@link finalize} once after the last language.
*/
export class PdgEmitSink implements KnowledgeGraph {
private readonly validTables: Set<string>;
private bbWriter: SyncCsvWriter | undefined;
/** pairKey (`From|To`) → writer. PDG edges are all `BasicBlock|BasicBlock`,
* but the map keeps the sink general and the manifest pair-keyed. */
@@ -129,7 +133,6 @@ export class PdgEmitSink implements KnowledgeGraph {
private readonly pdgCsvDir: string,
private readonly chunkRows: number = DEFAULT_PDG_EMIT_CHUNK_ROWS,
) {
this.validTables = new Set<string>(NODE_TABLES as readonly string[]);
// Clear any streamed CSVs left by a previous (possibly crashed) run so a
// later COPY never picks up stale rows.
fs.rmSync(pdgCsvDir, { recursive: true, force: true });
@@ -160,15 +163,27 @@ export class PdgEmitSink implements KnowledgeGraph {
addRelationship(relationship: GraphRelationship): void {
if (PDG_EDGE_TYPES.has(relationship.type)) {
const fromLabel = getNodeLabel(relationship.sourceId);
const toLabel = getNodeLabel(relationship.targetId);
// Skip edges whose endpoint labels are not valid node tables — mirrors
// `RelPairRouter` exactly so the streamed set matches the whole-graph set.
if (!this.validTables.has(fromLabel) || !this.validTables.has(toLabel)) return;
const pairKey = `${fromLabel}|${toLabel}`;
assertDeclaredPair(pairKey, DECLARED_RELATION_PAIRS);
// Classify + skip via the SHARED `relPairKeyFor`, not a local copy of its
// three lines, so the streamed set cannot drift from the whole-graph set
// `RelPairRouter` produces. `undefined` = an endpoint label is not a node
// table, so the edge is dropped exactly as the router drops it.
const pairKey = relPairKeyFor(
relationship.sourceId,
relationship.targetId,
VALID_NODE_TABLES,
);
if (pairKey === undefined) return;
assertDeclaredPair(
pairKey,
DECLARED_RELATION_PAIRS,
relationship.type,
relationship.sourceId,
relationship.targetId,
);
let writer = this.relWriters.get(pairKey);
if (writer === undefined) {
// Cold: once per pair, so decoding the key back into labels is free.
const [fromLabel, toLabel] = splitRelPairKey(pairKey);
try {
writer = new SyncCsvWriter(
path.join(this.pdgCsvDir, `rel_${fromLabel}_${toLabel}.csv`),
+3
View File
@@ -460,6 +460,9 @@ const WAITER_TIMEOUT_MS = 15_000;
// of the lbug-config retry-budget registry.
const LOCK_RETRY_ATTEMPTS = 3;
const LOCK_RETRY_DELAY_MS = 2000;
// determinism: probe — existence only. `probeDatabaseForShadowReplay` calls
// `getAll()` purely to force the shadow replay and then discards the result;
// the function returns void, so no row ever reaches a caller.
const SHADOW_REPLAY_PROBE_QUERY = 'MATCH (n) RETURN n LIMIT 1';
const poolSidecarLogger = {
+231 -30
View File
@@ -31,10 +31,29 @@ import path from 'path';
import { createWriteStream, type WriteStream } from 'fs';
import { once } from 'events';
import { finished } from 'stream/promises';
import { NODE_TABLES } from 'gitnexus-shared';
import { findInCauseChain } from '../../lib/utils.js';
/** Injectable for tests (backpressure/error simulation), mirroring split. */
export type WriteStreamFactory = (filePath: string) => WriteStream;
/**
* Every label LadybugDB has a node table for — the filter that decides whether
* an edge is routable at all.
*
* ONE shared instance, deliberately. `RelPairRouter`, `GraphEmitSink` and
* `PdgEmitSink` each used to build their own `new Set(NODE_TABLES)`; three
* copies of the same immutable set are three chances to seed one of them from
* a different source. Typed `ReadonlySet` because that — not `Object.freeze`,
* which does not touch a Set's internal slots — is what actually stops a
* consumer mutating the shared instance.
*
* Imported straight from `gitnexus-shared` rather than `./schema.js`: schema.ts
* imports `parseRelationSchemaPairs` from this module, so the reverse import
* would close a cycle.
*/
export const VALID_NODE_TABLES: ReadonlySet<string> = new Set<string>(NODE_TABLES);
/**
* Derive a node's table label from its graph id. Matches the legacy
* `getNodeLabel` that lived inline in `loadGraphToLbug`:
@@ -48,20 +67,92 @@ export const getNodeLabel = (nodeId: string): string => {
return nodeId.split(':')[0];
};
/**
* Classify one edge into its `From|To` pair key, or `undefined` when the edge
* must be SKIPPED because an endpoint's label is not a real node table.
*
* THE single definition of "which pair does this edge belong to, and is it
* routable at all". `RelPairRouter.route`, `GraphEmitSink.addRelationship`,
* `PdgEmitSink.addRelationship` and the `structural-pair-coverage` corpus guard
* each used to inline the same three lines (label both ends → drop if either
* label is not a node table → join with `|`). The corpus guard's docblock said
* it "mirrors `RelPairRouter.route`" — a mirror is a drift marker: change the
* skip rule here and the guard would keep classifying by the old one, report
* green, and let `analyze` abort on a pair it had already declared covered.
*
* HOT PATH — called once per edge (~1M on a large repo). Returns the key
* string (which every caller needs anyway for its own Map lookup) rather than
* a `{ pairKey, fromLabel, toLabel }` object or a tuple, so the success path
* allocates nothing beyond what `getNodeLabel` already did. Callers that need
* the two labels back — only when opening a new pair's CSV, once per pair —
* decode the key with {@link splitRelPairKey}.
*/
export const relPairKeyFor = (
fromId: string,
toId: string,
validTables: ReadonlySet<string>,
): string | undefined => {
const fromLabel = getNodeLabel(fromId);
const toLabel = getNodeLabel(toId);
if (!validTables.has(fromLabel) || !validTables.has(toLabel)) return undefined;
return `${fromLabel}|${toLabel}`;
};
/**
* Decode a `From|To` pair key back into its two labels.
*
* Safe because `|` cannot occur inside a node label: every label is a
* `NODE_TABLES` identifier (`[A-Za-z][A-Za-z0-9_]*`), so the FIRST `|` is
* always the separator. That invariant was documented in one comment and
* enforced nowhere while every consumer re-derived it with a bare
* `key.split('|')`.
*
* DECODE ONLY — there is deliberately no matching `encode` helper. The key is
* built once per edge inside {@link relPairKeyFor} (~1M edges on a large
* repo), where a function call is a real regression risk; every decode site is
* cold by construction (once per pair when its CSV is opened, or on the
* throw path of {@link assertDeclaredPair}).
*/
export const splitRelPairKey = (key: string): readonly [from: string, to: string] => {
const sep = key.indexOf('|');
return sep < 0 ? [key, ''] : [key.slice(0, sep), key.slice(sep + 1)];
};
/**
* Build a fresh matcher for the `FROM <label> TO <label>` clauses of a
* relationship DDL. Capture group 1 is the FROM label, group 2 the TO label;
* backticks quote schema labels and are not part of the graph label.
*
* THE SINGLE SOURCE OF TRUTH for that pattern. `parseRelationSchemaPairs`
* below builds its pair set from it, and `test/unit/schema-pair-coverage.test.ts`
* counts raw `FROM…TO` occurrences with it to catch a pair DUPLICATED in the
* DDL (a duplicate makes LadybugDB reject `CREATE REL TABLE`, killing every
* `analyze` — strictly worse than one missing pair). That guard used to inline
* its own copy of the regex: the two matched identically, so it worked, but any
* widening here (dotted identifiers, `IF NOT EXISTS`, a multi-target
* `FROM x TO y, z` form) would have silently degraded it to the tautology
* `declared.size === declared.size`. Consume this factory instead of
* re-inlining a copy.
*
* A FACTORY, not a shared `RegExp`: a module-level `/g` regex carries
* `lastIndex` between calls, so one consumer's `exec`/`test` would corrupt
* everyone else's next match. Each call returns a private instance.
*/
export const createRelationPairMatcher = (): RegExp =>
/\bFROM\s+`?([A-Za-z][A-Za-z0-9_]*)`?\s+TO\s+`?([A-Za-z][A-Za-z0-9_]*)`?/g;
/**
* Extract the FROM→TO pairs accepted by a relationship DDL.
*
* This belongs at the routing boundary: schema.ts owns the DDL, while the CSV
* router owns the fail-fast check that prevents writing a pair LadybugDB cannot
* COPY. Backticks quote schema labels and are not part of the graph label.
* COPY.
*/
export const parseRelationSchemaPairs = (relationSchema: string): ReadonlySet<string> =>
new Set(
[
...relationSchema.matchAll(
/\bFROM\s+`?([A-Za-z][A-Za-z0-9_]*)`?\s+TO\s+`?([A-Za-z][A-Za-z0-9_]*)`?/g,
),
].map((match) => `${match[1]}|${match[2]}`),
[...relationSchema.matchAll(createRelationPairMatcher())].map(
(match) => `${match[1]}|${match[2]}`,
),
);
export interface RelPairMeta {
@@ -69,6 +160,108 @@ export interface RelPairMeta {
rows: number;
}
/**
* Best-effort source file for a graph node id.
*
* Node ids are `<Label>:<repo-relative path>:…` (`comm_*` / `proc_*` synthetic
* ids carry no file). Repo-relative paths are POSIX-normalized and never carry
* a drive letter, so the segment after the label is the file path. Returns
* `undefined` rather than guessing when the id has no path segment — a wrong
* file in a crash message is worse than none.
*/
const deriveNodeFilePath = (nodeId: string): string | undefined => {
if (nodeId.startsWith('comm_') || nodeId.startsWith('proc_')) return undefined;
const filePath = nodeId.split(':')[1];
return filePath === undefined || filePath === '' ? undefined : filePath;
};
/** Where a user files the missing pair. Part of the message — see below. */
const UNDECLARED_PAIR_ISSUE_URL = 'https://github.com/abhigyanpatwari/GitNexus/issues/new';
/**
* An edge whose endpoint-label pair is absent from the relationship DDL.
*
* Carries the context the emit call site already has — relationship type, both
* node ids, and the source file derived from them — so a user whose `analyze`
* just died mid-run can see WHICH of their files produced the edge and file a
* bug report that names the missing pair. The abstract label pair alone is
* unactionable outside GitNexus's own source (#2789).
*
* Classify by TYPE (`err instanceof UndeclaredRelationPairError`, or
* {@link findUndeclaredRelationPairError} when the error may be wrapped in a
* phase `cause` chain) — the repo norm from #2385 — never by message text.
*
* THE MESSAGE IS THE ONLY RENDERING. It carries the five context fields AND
* the two actionable next steps (report the pair; `.gitnexusignore` the file to
* finish the rest of the index), because `gitnexus serve` forwards nothing but
* `err.message` over worker IPC — anything a consumer re-renders from the
* structured fields instead is invisible to a serve-hosted user. The CLI
* branch in `cli/analyze.ts` therefore prints this message indented and adds
* only the machine-readable `cliError` fields, the same idiom `LbugWipeError`
* uses there. It used to re-render the five fields with its own wording; the
* two copies had already drifted on the pair separator, the no-file text and
* the closing sentence within a single PR, and each had its own pinning test.
*/
export class UndeclaredRelationPairError extends Error {
/** `From|To` label pair, exactly as keyed against the declared-pair set. */
readonly pairKey: string;
/** Relationship type of the edge that could not be routed (e.g. `CALLS`). */
readonly relationType: string;
readonly fromId: string;
readonly toId: string;
/** Source file derived from the node ids; `undefined` for synthetic ids. */
readonly sourceFile: string | undefined;
constructor(pairKey: string, relationType: string, fromId: string, toId: string) {
const sourceFile = deriveNodeFilePath(fromId) ?? deriveNodeFilePath(toId);
const [fromLabel, toLabel] = splitRelPairKey(pairKey);
super(
`GitNexus extracted a relationship its own database schema cannot store.\n` +
`Relationship label pair ${fromLabel} → ${toLabel} is not declared in the ` +
`LadybugDB relation schema.\n` +
` relationship type: ${relationType}\n` +
` from node: ${fromId}\n` +
` to node: ${toId}\n` +
` source file: ${sourceFile ?? '(none — synthetic node id)'}\n` +
`This is a gap in GitNexus's own relation schema, not a problem with the ` +
`analyzed code, and re-running the analysis will fail in exactly the same place.\n` +
`Suggestions:\n` +
` 1. Report the missing pair so it can be declared:\n` +
` ${UNDECLARED_PAIR_ISSUE_URL}\n` +
` Include the label pair, the relationship type, and the source file above.\n` +
` 2. To finish indexing the rest meanwhile, add that file (or its directory)\n` +
` to .gitnexusignore and re-run.`,
);
this.name = 'UndeclaredRelationPairError';
this.pairKey = pairKey;
this.relationType = relationType;
this.fromId = fromId;
this.toId = toId;
this.sourceFile = sourceFile;
}
}
/**
* Find an {@link UndeclaredRelationPairError} in `err` or its `cause` chain.
*
* The guard throws deep inside an ingestion phase, and the phase runner rewraps
* every phase failure as `new Error("Phase 'X' failed: …", { cause })` — so a
* bare `instanceof` at the CLI boundary would miss it and fall through to the
* generic stack dump.
*
* The traversal and its depth bound come from `lib/utils.ts` rather than being
* re-rolled here: this was the fourth hand-written copy in the repo and the
* only one that used `depth <= MAX` (six levels) while claiming to mirror
* `cli/analyze.ts`'s `depth < 5`.
*/
export const findUndeclaredRelationPairError = (
err: unknown,
): UndeclaredRelationPairError | undefined =>
findInCauseChain(
err,
(e): e is UndeclaredRelationPairError => e instanceof UndeclaredRelationPairError,
);
/**
* Fail fast on an endpoint-label pair absent from the relationship DDL, the
* same guard `RelPairRouter.route` applies to the whole-graph emit. Exported
@@ -77,18 +270,27 @@ export interface RelPairMeta {
* the bulk insert, and is silently dropped by the per-edge fallback instead
* of failing loudly like the non-streaming path does.
*
* Takes the already-built `From|To` pairKey rather than the two labels — every
* caller needs that same key immediately after for its own Map/stream lookup,
* and this is on the per-edge hot path, so building it twice would be a
* needless allocation per edge. `|` cannot appear inside a label (node labels
* are `NODE_TABLES` identifiers), so splitting it back apart for the error
* message is safe.
* Takes the already-built `From|To` pairKey (from {@link relPairKeyFor})
* rather than the two labels — every caller needs that same key immediately
* after for its own Map/stream lookup, and this is on the per-edge hot path,
* so building it twice would be a needless allocation per edge. The error
* splits it back apart with {@link splitRelPairKey}, which only the throw path
* reaches. The edge context is passed POSITIONALLY for the same reason: a
* `{ relationType, fromId, toId }` context object would allocate on every
* edge, including the ~1M that never fail.
*
* The success path must stay allocation-free: no object literal, no template
* string, no closure, no `Error` constructed before the failure branch.
*/
export const assertDeclaredPair = (pairKey: string, declaredPairs: ReadonlySet<string>): void => {
export const assertDeclaredPair = (
pairKey: string,
declaredPairs: ReadonlySet<string>,
relationType: string,
fromId: string,
toId: string,
): void => {
if (!declaredPairs.has(pairKey)) {
throw new Error(
`Relationship label pair ${pairKey.replaceAll('|', '→')} is not declared in the LadybugDB relation schema`,
);
throw new UndeclaredRelationPairError(pairKey, relationType, fromId, toId);
}
};
@@ -110,7 +312,7 @@ export class RelPairRouter {
constructor(
private readonly csvDir: string,
private readonly header: string,
private readonly validTables: Set<string>,
private readonly validTables: ReadonlySet<string>,
private readonly declaredPairs: ReadonlySet<string>,
private readonly wsFactory: WriteStreamFactory = (p) => createWriteStream(p, 'utf-8'),
) {}
@@ -135,23 +337,25 @@ export class RelPairRouter {
* Returns `void` on the synchronous hot path; a `Promise<void>` only when a
* stream signals backpressure (or a new pair's header does) — the caller
* awaits the promise before routing the next edge.
*
* `relType` is not used for routing — it is carried purely so an undeclared
* pair can name the offending relationship in its error (the row is already
* CSV-escaped by then, so the type is not recoverable from it).
*/
route(fromId: string, toId: string, row: string): void | Promise<void> {
route(fromId: string, toId: string, row: string, relType: string): void | Promise<void> {
if (this.streamError) throw this.streamError;
const fromLabel = getNodeLabel(fromId);
const toLabel = getNodeLabel(toId);
if (!this.validTables.has(fromLabel) || !this.validTables.has(toLabel)) {
const pairKey = relPairKeyFor(fromId, toId, this.validTables);
if (pairKey === undefined) {
this.skipped++;
return;
}
const pairKey = `${fromLabel}|${toLabel}`;
assertDeclaredPair(pairKey, this.declaredPairs);
assertDeclaredPair(pairKey, this.declaredPairs, relType, fromId, toId);
const ws = this.streams.get(pairKey);
if (ws === undefined) {
// First edge for this pair: open the stream, write header + row.
return this.openAndWrite(pairKey, fromLabel, toLabel, row);
return this.openAndWrite(pairKey, row);
}
this.byPair.get(pairKey)!.rows++;
@@ -161,12 +365,9 @@ export class RelPairRouter {
}
}
private async openAndWrite(
pairKey: string,
fromLabel: string,
toLabel: string,
row: string,
): Promise<void> {
/** Cold: runs once per pair, so decoding the key back is free here. */
private async openAndWrite(pairKey: string, row: string): Promise<void> {
const [fromLabel, toLabel] = splitRelPairKey(pairKey);
const csvPath = path.join(this.csvDir, `rel_${fromLabel}_${toLabel}.csv`);
const ws = this.wsFactory(csvPath);
ws.on('error', this.markError);
+364 -161
View File
@@ -9,8 +9,13 @@
* MATCH (f:Function)-[r:CodeRelation {type: 'CALLS'}]->(g:Function) RETURN f, g
*/
import { createHash } from 'crypto';
// Import from shared package (single source of truth) — used in DDL templates below
import { NODE_TABLES, REL_TABLE_NAME, REL_TYPES, EMBEDDING_TABLE_NAME } from 'gitnexus-shared';
import type { NodeLabel, NodeTableName } from 'gitnexus-shared';
import { parseRelationSchemaPairs } from './rel-pair-routing.js';
import { LINKABLE_LABELS } from '../ingestion/scope-resolution/graph-bridge/node-lookup.js';
import { CALL_TARGET_TYPES } from '../ingestion/model/symbol-table.js';
// Re-export so downstream consumers keep the same import path
export { NODE_TABLES, REL_TABLE_NAME, REL_TYPES, EMBEDDING_TABLE_NAME };
export type { NodeTableName, RelType } from 'gitnexus-shared';
@@ -248,92 +253,214 @@ CREATE NODE TABLE BasicBlock (
// Single table with 'type' property - connects all node tables
// ============================================================================
export const RELATION_SCHEMA = `
CREATE REL TABLE ${REL_TABLE_NAME} (
FROM File TO File,
FROM File TO Folder,
FROM File TO Function,
FROM File TO Class,
FROM File TO Interface,
FROM File TO Method,
/**
* Labels the scope-resolution graph bridge can put on the SOURCE side of a
* CALLS / ACCESSES / USES / EXTENDS edge: everything `buildGraphNodeLookup`
* registers (`LINKABLE_LABELS`), plus the `File` node `resolveCallerGraphId`
* falls back to for a module-level call site.
*
* Imported from the ingestion layer, NOT re-listed here: a hand-copied twin is
* pure drift risk, since a label added to it (or dropped from the original) is
* invisible to every guard. `csv-generator.ts` and `lbug-adapter.ts`, both
* siblings in this directory, already import from `../ingestion/`.
*
* The cost is that importing this module pulls five ingestion modules into the
* runtime closure. `gitnexus-web` does not depend on this package at all (only
* on `gitnexus-shared`), so nothing here reaches a browser bundle. The MCP
* server still pays it, though: `local-backend.ts` no longer imports this
* module directly (its two embedding constants come from `gitnexus-shared`),
* but pool-adapter -> lbug-adapter -> csv-generator reaches it anyway.
*/
const SCOPE_BRIDGE_SOURCE_LABELS: readonly NodeLabel[] = ['File', ...LINKABLE_LABELS];
/**
* Labels the bridge can put on the TARGET side: `LINKABLE_LABELS` again (every
* `resolveDefGraphId` hit), plus `CALL_TARGET_TYPES` —
* `tryEmitEdgeWithExplicitTargetId` bypasses the lookup and emits such a def's
* own node id, and C# `Delegate` is in that set without being linkable.
*
* Both sets are `NodeLabel`-typed rather than `NodeTableName`-typed because
* that is what the originals carry, and `NodeLabel` is the wider union — it
* admits five labels with no node table (`Project`, `Package`, `Decorator`,
* `Import`, `Type`). A label from that gap would emit DDL naming a table that
* does not exist, so `test/unit/schema-pair-coverage.test.ts` asserts every
* declared endpoint against `NODE_TABLES`. The hand-written sets below take the
* narrower `NodeTableName` constraint, where a typo is the actual risk.
*/
const SCOPE_BRIDGE_TARGET_LABELS: ReadonlySet<NodeLabel> = new Set<NodeLabel>([
...LINKABLE_LABELS,
...CALL_TARGET_TYPES,
]);
/**
* Node tables that are NOT definitions.
*
* - `Community` / `Process` are analysis overlays synthesized after ingestion;
* nothing is ever attached to one, they are only attached TO.
* - `Route` / `Tool` are framework overlays. They do source exactly two edges
* — `ENTRY_POINT_OF` to a `Process` (`pipeline-phases/processes.ts`) — but
* that emitter hard-codes both labels as literals in one file rather than
* resolving an anchor through a lookup, so those two pairs stay in
* {@link STRUCTURAL_PAIR_DDL}. Admitting them as anchors would mint twelve
* further pairs (`Route→Annotation`, `Tool→Record`, …) no emitter can reach.
* - `Folder` is a filesystem container (`Folder→Folder` / `Folder→File` only).
* - `BasicBlock` is the PDG substrate (`BasicBlock→BasicBlock` only; measured
* over 300k PDG edges, no other pair is emitted).
*/
const NON_DEFINITION_LABELS: readonly NodeTableName[] = [
'Community',
'Process',
'Route',
'Tool',
'Folder',
'BasicBlock',
];
/**
* Every label a DEFINITION node can carry — derived from `NODE_TABLES` by
* subtraction so a new node table joins this set automatically and only an
* explicit entry above can keep it out.
*/
const DEFINITION_ANCHOR_LABELS: readonly NodeTableName[] = NODE_TABLES.filter(
(label) => !NON_DEFINITION_LABELS.includes(label),
);
/**
* Labels whose nodes are minted OUTSIDE the scope-resolution bridge, by a
* phase or framework emitter, and then hung off whichever definition node that
* emitter happened to resolve. For most of them the anchor is a LOOKUP RESULT,
* so its label is not constrained by the emitter — which is exactly why
* hand-listing these pairs has crashed `analyze` four separate times:
*
* | target | emitter | anchor comes from |
* |--------------|------------------------------------------------|---------------------------------------|
* | `Annotation` | `frameworks/spring/conditionals.ts` CONDITIONAL_ON | `resolveDefGraphId` / `resolveCallerGraphId` |
* | `Community` | `pipeline-phases/communities.ts` MEMBER_OF | Leiden membership, `isCommunitySymbol`-gated |
* | `Process` | `pipeline-phases/processes.ts` STEP_IN_PROCESS | trace step node |
* | `Route` | `pipeline-phases/routes.ts` HANDLES_ROUTE | `generateId('File', handlerPath)` — a literal |
* | `Tool` | `pipeline-phases/tools.ts` HANDLES_TOOL | `handlerNodeId` — whatever definition the decorator sat on |
* | `File` | `languages/vue/scope-resolver.ts` BINDS_EVENT_HANDLER | handler node |
* | `Record` | `cobol-processor.ts` × 8 external-resource sites | `scopedCallerLookup` |
*
* The four reproduced hard-aborts are one cell of this table each:
* `Method→Annotation` (Spring `@Bean` + `@ConditionalOnMissingBean`),
* `Method→File` (Vue Options-API handler), `Namespace→Record` (COBOL
* `DECLARATIVES`), `Class→Tool` (`@mcp.tool()` on a class). Declaring
* {@link DEFINITION_ANCHOR_LABELS} × this set covers all four plus every
* sibling the same emitters can reach.
*
* TWO TARGETS ARE LABEL-GATED TODAY, and the cross product over-declares for
* them ON PURPOSE (~47 of the 182 attachment pairs are unreachable right now):
* - `Community` — `isCommunitySymbol` (`community-processor.ts`) admits only
* `Function` / `Class` / `Method` / `Interface` as members, so the other 22
* anchors cannot source a MEMBER_OF edge until that predicate widens.
* - `Route` — HANDLES_ROUTE sources `generateId('File', handlerPath)`, a
* literal `File`, so every non-`File` anchor is headroom.
*
* Those pairs stay declared because the two sides of the error are not
* symmetric: an UNDECLARED pair makes LadybugDB reject the edge and aborts
* `analyze` outright on a user's repo, while an unused DECLARED pair costs
* almost nothing — `bench/schema-pairs` measures the whole 332→450 growth
* (118 pairs, of which these ~47 are a part) at 0.93–1.05×, i.e. inside
* run-to-run noise.
* Every one of the four aborts above came from re-narrowing a set to what one
* predicate looked like it allowed — so a reading of `isCommunitySymbol` is not
* grounds to shrink this. Widening either predicate is then a no-op here.
*
* `Route` / `Tool` being excluded as ANCHORS (see {@link NON_DEFINITION_LABELS})
* is likewise a SIZE choice, not something derived from a rule: they do source
* `ENTRY_POINT_OF`, and admitting them would mint twelve further pairs no
* emitter can currently reach.
*
* Sized deliberately: this rule brings the DDL to 450 pairs. `bench/schema-pairs`
* measures it against real `@ladybugdb/core` with identical data — untyped-endpoint
* anchored queries (`MATCH (a {id: $id})-[r:CodeRelation]->(b)`, the shape
* `impact` / `context` / `detect_changes` issue), relative to the 332-pair
* hand-list this replaced. Four runs on the same box:
*
* 450 → 0.93–1.05× 641 → 1.22–1.43× 786 → 1.52–1.75× 1024 → 2.03–2.34×
*
* 450 is inside run-to-run noise (it came out FASTER than 332 on three of the
* four runs); everything past ~640 is not. The knee sits just above 450, so the containment half below
* stays hand-declared rather than being folded into a third cross product.
* Re-run that bench and quote the range — not one run — before proposing one.
*/
const ATTACHMENT_TARGET_LABELS: readonly NodeTableName[] = [
'Annotation',
'Community',
'Process',
'Route',
'Tool',
'File',
'Record',
];
/**
* The 72 pairs NEITHER rule above generates — everything left after the two
* cross products are subtracted. Carried by CONTAINMENT, inheritance, imports
* and DI: a container label crossed with a contained label. No predicate
* describes that surface (any container can hold any definition).
*
* What survives here is characteristic, not arbitrary. Almost all of it is a
* TARGET no rule reaches — `CodeElement`, `Impl`, `Namespace`, `Template`,
* `TypeAlias`, `Typedef`, `Union`, `Static`, `Section`, `Folder` are in neither
* `SCOPE_BRIDGE_TARGET_LABELS` nor {@link ATTACHMENT_TARGET_LABELS} — plus the
* `Impl|*` and `Template|*` member rows (Rust `impl`/`trait` bodies, C++
* templates), the two `Route|Process` / `Tool|Process` entry points whose
* emitter names both labels as literals, and `BasicBlock|BasicBlock`, the PDG
* substrate.
*
* NOTHING A RULE ALREADY COVERS BELONGS HERE. `generatedRelationPairs` skips
* any pair present in this block, so a redundant line does not merely duplicate
* — it SUPPRESSES generation, and later narrowing a rule would silently keep
* that pair alive with no test failing. 161 such lines were deleted from this
* block (the DDL's pair set is unchanged: they moved into the generated half);
* `test/unit/schema-pair-coverage.test.ts` now fails if one comes back.
*
* Folding this remainder into a third cross product
* (`DEFINITION_ANCHOR_LABELS × {CodeElement, Section, Typedef, Union,
* Namespace, Impl, TypeAlias, Static, Template}`) would take the table to 641
* pairs and leave only ~29 lines here. `bench/schema-pairs` measures 641 at
* 1.22–1.43× on anchored queries, where production's 450 is inside noise — so
* that trade buys ~43 fewer hand-written lines for a real ~22–43% on the query
* shape `impact` uses, which is why it is deferred rather than taken.
*
* Exported so `test/unit/schema-pair-coverage.test.ts` can subtract it and
* assert the GENERATED region of the DDL for exact equality against the two
* rules, rather than one-directional containment.
* `test/integration/structural-pair-coverage.test.ts` guards this half from a
* corpus — that is the guard, not this comment.
*/
export const STRUCTURAL_PAIR_DDL = ` FROM File TO Folder,
FROM File TO CodeElement,
FROM File TO \`Struct\`,
FROM File TO \`Enum\`,
FROM File TO \`Macro\`,
FROM File TO \`Typedef\`,
FROM File TO \`Union\`,
FROM File TO \`Namespace\`,
FROM File TO \`Trait\`,
FROM File TO \`Impl\`,
FROM File TO \`TypeAlias\`,
FROM File TO \`Const\`,
FROM File TO \`Static\`,
FROM File TO \`Variable\`,
FROM File TO \`Property\`,
FROM File TO \`Record\`,
FROM File TO \`Delegate\`,
FROM File TO \`Annotation\`,
FROM File TO \`Constructor\`,
FROM File TO \`Template\`,
FROM File TO \`Module\`,
FROM File TO Section,
FROM Folder TO Folder,
FROM Folder TO File,
FROM Function TO Function,
FROM Function TO Method,
FROM Function TO Class,
FROM Function TO Community,
FROM Function TO \`Macro\`,
FROM Function TO \`Struct\`,
FROM Function TO \`Template\`,
FROM Function TO \`Enum\`,
FROM Function TO \`Namespace\`,
FROM Function TO \`TypeAlias\`,
FROM Function TO \`Module\`,
FROM Function TO \`Impl\`,
FROM Function TO Interface,
FROM Function TO \`Constructor\`,
FROM Function TO \`Const\`,
FROM Function TO \`Typedef\`,
FROM Function TO \`Union\`,
FROM Function TO \`Property\`,
FROM Function TO CodeElement,
FROM Class TO Method,
FROM Class TO Function,
FROM Class TO Class,
FROM Class TO Interface,
FROM Class TO Community,
FROM Class TO \`Template\`,
FROM Class TO \`TypeAlias\`,
FROM Class TO \`Struct\`,
FROM Class TO \`Enum\`,
FROM Class TO \`Annotation\`,
FROM Class TO \`Constructor\`,
FROM Class TO \`Trait\`,
FROM Class TO \`Macro\`,
FROM Class TO \`Impl\`,
FROM Class TO \`Union\`,
FROM Class TO \`Namespace\`,
FROM Class TO \`Typedef\`,
FROM Class TO \`Property\`,
FROM Class TO CodeElement,
FROM Method TO Function,
FROM Method TO Method,
FROM Method TO Class,
FROM Method TO Community,
FROM Method TO \`Template\`,
FROM Method TO \`Struct\`,
FROM Method TO \`TypeAlias\`,
FROM Method TO \`Enum\`,
FROM Method TO \`Macro\`,
FROM Method TO \`Namespace\`,
FROM Method TO \`Module\`,
FROM Method TO \`Impl\`,
FROM Method TO Interface,
FROM Method TO \`Constructor\`,
FROM Method TO \`Property\`,
FROM Method TO \`Variable\`,
FROM Method TO \`Const\`,
FROM Method TO CodeElement,
FROM \`Template\` TO \`Template\`,
FROM \`Template\` TO Function,
@@ -345,134 +472,90 @@ CREATE REL TABLE ${REL_TABLE_NAME} (
FROM \`Template\` TO \`Macro\`,
FROM \`Template\` TO Interface,
FROM \`Template\` TO \`Constructor\`,
FROM \`Module\` TO \`Module\`,
FROM \`Module\` TO CodeElement,
FROM \`Module\` TO \`Namespace\`,
FROM \`Namespace\` TO Function,
FROM CodeElement TO CodeElement,
FROM CodeElement TO \`Module\`,
FROM CodeElement TO \`Property\`,
FROM Section TO Section,
FROM Section TO File,
FROM File TO Route,
FROM Function TO Route,
FROM Method TO Route,
FROM File TO Tool,
FROM Function TO Tool,
FROM Method TO Tool,
FROM CodeElement TO Community,
FROM Interface TO Community,
FROM Interface TO Function,
FROM Interface TO Method,
FROM Interface TO Class,
FROM Interface TO Interface,
FROM Interface TO CodeElement,
FROM Interface TO \`TypeAlias\`,
FROM Interface TO \`Struct\`,
FROM Interface TO \`Constructor\`,
FROM Interface TO \`Property\`,
FROM \`Struct\` TO Community,
FROM \`Struct\` TO \`Trait\`,
FROM \`Struct\` TO \`Struct\`,
FROM \`Struct\` TO Class,
FROM \`Struct\` TO \`Enum\`,
FROM \`Struct\` TO Function,
FROM \`Struct\` TO Method,
FROM \`Struct\` TO Interface,
FROM \`Struct\` TO \`Constructor\`,
FROM \`Struct\` TO \`Property\`,
FROM \`Enum\` TO \`Enum\`,
FROM \`Enum\` TO Community,
FROM \`Enum\` TO Class,
FROM \`Enum\` TO Interface,
FROM \`Enum\` TO Function,
FROM \`Enum\` TO Method,
FROM \`Enum\` TO \`Struct\`,
FROM \`Enum\` TO \`Constructor\`,
FROM \`Enum\` TO \`Property\`,
FROM \`Enum\` TO \`TypeAlias\`,
FROM \`Macro\` TO Community,
FROM \`Macro\` TO Function,
FROM \`Macro\` TO Method,
FROM \`Module\` TO Function,
FROM \`Module\` TO Method,
FROM \`Typedef\` TO Community,
FROM \`Union\` TO Community,
FROM \`Namespace\` TO Community,
FROM \`Namespace\` TO \`Struct\`,
FROM \`Trait\` TO Method,
FROM \`Trait\` TO Function,
FROM \`Trait\` TO \`Constructor\`,
FROM \`Trait\` TO \`Property\`,
FROM \`Trait\` TO Community,
FROM \`Impl\` TO Method,
FROM \`Impl\` TO Function,
FROM \`Impl\` TO \`Constructor\`,
FROM \`Impl\` TO \`Property\`,
FROM \`Impl\` TO Community,
FROM \`Impl\` TO \`Trait\`,
FROM \`Impl\` TO \`Struct\`,
FROM \`Impl\` TO \`Impl\`,
FROM \`TypeAlias\` TO Community,
FROM \`TypeAlias\` TO \`Trait\`,
FROM \`TypeAlias\` TO Class,
FROM \`Const\` TO Community,
FROM \`Const\` TO Method,
FROM \`Static\` TO Community,
FROM \`Variable\` TO Community,
FROM \`Variable\` TO Method,
FROM \`Property\` TO Community,
FROM \`Property\` TO \`Property\`,
FROM \`Property\` TO Class,
FROM \`Property\` TO \`Enum\`,
FROM \`Property\` TO Function,
FROM \`Property\` TO \`Struct\`,
FROM \`Record\` TO Method,
FROM \`Record\` TO \`Constructor\`,
FROM \`Record\` TO \`Property\`,
FROM \`Record\` TO Community,
FROM \`Delegate\` TO Community,
FROM \`Annotation\` TO Community,
FROM \`Constructor\` TO Community,
FROM \`Constructor\` TO Interface,
FROM \`Constructor\` TO Class,
FROM \`Constructor\` TO Method,
FROM \`Constructor\` TO Function,
FROM \`Constructor\` TO \`Constructor\`,
FROM \`Constructor\` TO \`Struct\`,
FROM \`Constructor\` TO \`Macro\`,
FROM \`Constructor\` TO \`Template\`,
FROM \`Constructor\` TO \`TypeAlias\`,
FROM \`Constructor\` TO \`Enum\`,
FROM \`Constructor\` TO \`Annotation\`,
FROM \`Constructor\` TO \`Impl\`,
FROM \`Constructor\` TO \`Namespace\`,
FROM \`Constructor\` TO \`Module\`,
FROM \`Constructor\` TO \`Property\`,
FROM \`Constructor\` TO \`Typedef\`,
FROM \`Template\` TO Community,
FROM \`Module\` TO Community,
FROM Function TO Process,
FROM Method TO Process,
FROM Class TO Process,
FROM Interface TO Process,
FROM \`Struct\` TO Process,
FROM \`Constructor\` TO Process,
FROM \`Module\` TO Process,
FROM \`Macro\` TO Process,
FROM \`Impl\` TO Process,
FROM \`Typedef\` TO Process,
FROM \`TypeAlias\` TO Process,
FROM \`Enum\` TO Process,
FROM \`Union\` TO Process,
FROM \`Namespace\` TO Process,
FROM \`Trait\` TO Process,
FROM \`Const\` TO Process,
FROM \`Static\` TO Process,
FROM \`Variable\` TO Process,
FROM \`Property\` TO Process,
FROM \`Record\` TO Process,
FROM \`Delegate\` TO Process,
FROM \`Annotation\` TO Process,
FROM \`Template\` TO Process,
FROM CodeElement TO Process,
FROM Route TO Process,
FROM Tool TO Process,
FROM BasicBlock TO BasicBlock,
FROM BasicBlock TO BasicBlock`;
/**
* The generated half of the DDL — one ` FROM \`x\` TO \`y\`` line per pair of
* the two cross products below.
*
* 1. SCOPE BRIDGE — `SCOPE_BRIDGE_SOURCE_LABELS × SCOPE_BRIDGE_TARGET_LABELS`.
* Those sets ARE the bridge's emit surface: `buildGraphNodeLookup` holds
* only `LINKABLE_LABELS`, so every id `resolveDefGraphId` returns wears one
* of those labels (#2792).
* 2. ATTACHMENT — `DEFINITION_ANCHOR_LABELS × ATTACHMENT_TARGET_LABELS`, the
* phase/framework overlays hung off a resolved anchor (#2793).
*
* Generated rather than hand-listed because in both families the endpoint
* labels are LOOKUP RESULTS, not literals at the emit site — so any pair drawn
* from the sets can reach `assertDeclaredPair`, and an undeclared one aborts
* `analyze` outright on whichever codebase happens to produce it. Every
* hand-listed fix so far declared only the pair in the stack trace and left the
* rest of its family missing: `Const→Method` (#2781), `Class→Variable` (#2792),
* `Interface→CodeElement` (#2416), then `Method→Annotation` / `Method→File` /
* `Namespace→Record` / `Class→Tool` (#2793) — four more from three different
* emitters, all live at once.
*
* A `Set` guards the emit against a DUPLICATED pair — the asymmetric failure
* mode `rel-pair-routing.ts` documents at length: a duplicate makes LadybugDB
* reject `CREATE REL TABLE` and kills EVERY `analyze`, where a missing pair only
* kills the codebases that emit it. The two target sets are disjoint TODAY, so
* nothing is deduped in practice; the Set is here so that moving one label
* between the rules can never cause it. Pairs already in
* {@link STRUCTURAL_PAIR_DDL} are skipped for the same reason.
*/
const generatedPairDdl = (): string => {
const structural = parseRelationSchemaPairs(STRUCTURAL_PAIR_DDL);
const seen = new Set<string>();
const lines: string[] = [];
const add = (from: NodeLabel, to: NodeLabel): void => {
const pairKey = `${from}|${to}`;
if (structural.has(pairKey) || seen.has(pairKey)) return;
seen.add(pairKey);
lines.push(` FROM \`${from}\` TO \`${to}\``);
};
for (const from of SCOPE_BRIDGE_SOURCE_LABELS) {
for (const to of SCOPE_BRIDGE_TARGET_LABELS) add(from, to);
}
for (const from of DEFINITION_ANCHOR_LABELS) {
for (const to of ATTACHMENT_TARGET_LABELS) add(from, to);
}
return lines.join(',\n');
};
export const RELATION_SCHEMA = `
CREATE REL TABLE ${REL_TABLE_NAME} (
${STRUCTURAL_PAIR_DDL},
${generatedPairDdl()},
type STRING,
confidence DOUBLE,
reason STRING,
@@ -573,3 +656,123 @@ export const NODE_SCHEMA_QUERIES = [
export const REL_SCHEMA_QUERIES = [RELATION_SCHEMA];
export const SCHEMA_QUERIES = [...NODE_SCHEMA_QUERIES, ...REL_SCHEMA_QUERIES, EMBEDDING_SCHEMA];
/**
* Digest of the graph DDL this build creates — the exact statements
* {@link runSchemaCreationQueries} (lbug-adapter.ts) executes for the node and
* relation tables.
*
* This REPLACED `INCREMENTAL_SCHEMA_VERSION` (#2798), a hand-incremented
* integer in repo-manager.ts that had to PREDICT whether an on-disk database
* was created from this build's DDL. It could not: the number collided with
* `main` eight times, twice EXACTLY, and an exact clash was the quiet failure —
* two builds stamp the same number over different DDL, the strict `===` gate
* reads the index as current, every `CREATE … TABLE` is skipped as "already
* exists" (suppressed in `runSchemaCreationQueries`), and the edges whose
* endpoint pair the live DB cannot persist are dropped by
* `fallbackRelationshipInserts`' bare `catch`. A wrong graph, not an error.
*
* A digest cannot collide BY ACCIDENT at this scale: 12 hex chars is 48 bits,
* so even 1,000 distinct DDL variants over the project's whole life put the
* birthday probability of any pair matching at ≈1.8e-9. Two builds agree
* exactly when their DDL agrees, so concurrent branches never need renumbering.
* Do not shorten the slice: the odds double per bit dropped. On mismatch —
* including the ABSENT stamp every pre-#2798 index carries — run-analyze warns
* and forces a full re-analyze, which wipes the database and recreates the
* tables from the DDL below.
*
* {@link EMBEDDING_SCHEMA} is deliberately EXCLUDED. Its `FLOAT[N]` width comes
* from `GITNEXUS_EMBEDDING_DIMS` at module load, so folding it in would make
* this a function of the ENVIRONMENT rather than of code: two runs of the same
* build under different env would disagree and force alternating full rebuilds.
* Vector-column drift is therefore a SEPARATE gate, not an ungated hazard:
* {@link embeddingDimsMismatch} compares the width stamped in
* `RepoMeta.embeddingDims` against {@link EMBEDDING_DIMS} and run-analyze
* forces a rebuild on drift. Do not merge the two — an env-derived value in a
* code digest makes the same build disagree with itself. (The older reaction in
* run-analyze remains, and is to the CACHE, not the schema: when the cached
* vectors' length differs from `EMBEDDING_DIMS` it discards the cache and
* re-embeds.)
*/
export const SCHEMA_FINGERPRINT: string = createHash('sha256')
.update([...NODE_SCHEMA_QUERIES, ...REL_SCHEMA_QUERIES].join('\n'))
.digest('hex')
.slice(0, 12);
/**
* Whether an index built under `recorded` can be reused by this build.
*
* Lives here rather than in run-analyze so the query side can ask the same
* question without importing the analyze pipeline — the reason
* `cjkSegmentationModeMismatch` sits in `core/search/` rather than beside its
* caller. ABSENT counts as a mismatch: that is the backward-compatibility path
* for every index written before the field existed, and grandfathering it would
* stamp a fresh fingerprint onto a database whose DDL was never verified.
*/
export const schemaFingerprintMismatch = (recorded: string | undefined): boolean =>
recorded !== SCHEMA_FINGERPRINT;
/**
* Whether a stamped value has the shape {@link SCHEMA_FINGERPRINT} produces —
* the lowercase-hex prefix of a sha256 digest. The width is read from the live
* constant, so changing the slice above needs no edit here.
*
* Used to decide whether a stamp is worth NAMING in a diagnostic: an index with
* no fingerprint and one carrying a malformed value are both "not this build",
* but only the first has an explanation worth printing. Not a comparison gate —
* {@link schemaFingerprintMismatch} already rejects every value that is not
* exactly this build's.
*/
export const isSchemaFingerprintShaped = (value: unknown): value is string =>
typeof value === 'string' &&
value.length === SCHEMA_FINGERPRINT.length &&
/^[0-9a-f]+$/.test(value);
/**
* Whether the vector-column width an index's `CodeEmbedding` table was created
* at (as persisted in `RepoMeta.embeddingDims`) differs from the width this
* process would embed at ({@link EMBEDDING_DIMS}). The gate
* {@link SCHEMA_FINGERPRINT} deliberately cannot be: `FLOAT[N]` comes from
* `GITNEXUS_EMBEDDING_DIMS` at module load, so folding it into a digest of the
* DDL would make that digest a function of the ENVIRONMENT. Splitting it out
* here keeps the fingerprint purely code-derived and still gates the width —
* before this, flipping `GITNEXUS_EMBEDDING_DIMS` on a same-commit clean tree
* fired no guard at all: `alreadyUpToDate` returned over a `FLOAT[384]` table
* while the process embedded at 768. (The one pre-existing reaction, in
* run-analyze, discards the embedding CACHE and re-embeds — into a column whose
* width was never revisited.) A single scalar, so plain equality suffices.
*
* ABSENT does NOT count as a mismatch — the opposite of
* {@link schemaFingerprintMismatch}, and deliberately:
*
* - Absence carries no signal about the width. A missing fingerprint means
* "DDL this build cannot vouch for", and the field ships WITH a DDL change,
* so absence is itself evidence of drift. A missing dims stamp means only
* "written before the field existed"; the width was whatever that run's env
* resolved, almost always the 384 default, and it was consistent with the
* table it wrote. Drift needs the env to CHANGE, which absence says nothing
* about.
* - Forcing on absence would buy no safety anyway. Every index that lacks this
* stamp also lacks `schemaFingerprint` (both landed together in #2798), and
* that guard already forces a rebuild for exactly those indexes — after
* which the width is stamped and the hazard is closed for good. A second
* trigger for the same one rebuild is dead weight that would keep firing
* forever on any future path that legitimately omits the stamp.
* - The cost of guessing wrong is asymmetric: a fleet-wide full re-analyze
* (minutes to hours per repo) for a hazard that requires a rare, deliberate
* env change.
*
* Absence is precisely `undefined`. Any other recorded value that is not this
* build's width — including a malformed one, since `meta.json` is a schema-less
* `JSON.parse` of on-disk state — reads as a mismatch and errs toward a
* rebuild, which is the safe direction.
*
* Pure + exported for testing, and takes `current` explicitly rather than
* closing over {@link EMBEDDING_DIMS}: that constant is frozen at module load,
* so a parameter is the only way to exercise both sides of the comparison.
* Lives here rather than in run-analyze for the reason
* `cjkSegmentationModeMismatch` lives in `core/search/` — a caller that only
* needs the comparator should not have to pull in the analyze pipeline.
*/
export const embeddingDimsMismatch = (recorded: number | undefined, current: number): boolean =>
recorded !== undefined && recorded !== current;
+413 -117
View File
@@ -74,6 +74,9 @@ import {
inspectLbugSidecars,
} from './lbug/sidecar-recovery.js';
import type { EmbeddingIdentity } from './embeddings/embedding-identity.js';
// Type-only (erased at compile time), so the lazy-embeddings convention
// (#2370: no embeddings module loads unless a run actually needs one) holds.
import type { EmbeddingPipelineResult } from './embeddings/embedding-pipeline.js';
import {
getStoragePaths,
resolveBranchPlacement,
@@ -88,7 +91,6 @@ import {
reconcileMetadataFiles,
isMissingFilesystemError,
INDEX_METADATA_FILE,
INCREMENTAL_SCHEMA_VERSION,
type AnalyzerRunnerIdentity,
type RepoMeta,
} from '../storage/repo-manager.js';
@@ -138,8 +140,15 @@ import {
import type { CachedEmbedding } from './embeddings/types.js';
import { generateAIContextFiles } from '../cli/ai-context.js';
import { sanitizeDetectedBranch } from '../cli/analyze-config.js';
import { EMBEDDING_TABLE_NAME } from './lbug/schema.js';
import { STALE_HASH_SENTINEL } from './lbug/schema.js';
import {
EMBEDDING_TABLE_NAME,
EMBEDDING_DIMS,
STALE_HASH_SENTINEL,
SCHEMA_FINGERPRINT,
schemaFingerprintMismatch,
isSchemaFingerprintShaped,
embeddingDimsMismatch,
} from './lbug/schema.js';
import { isSpringBeanCandidateSourceFile } from './ingestion/frameworks/spring/bean-catalog.js';
import { isSpringBeanFactoryDeclaration } from './ingestion/frameworks/spring/bean-factories.js';
import {
@@ -158,6 +167,41 @@ import {
finalizeAnalyzerRunnerIdentity,
resolveAnalyzerRunnerIdentity,
} from './analyzer-identity.js';
// Static, and deliberately so: `embedding-count.ts` and `embedding-checkpoint.ts`
// live outside `core/embeddings/` precisely because the lazy-embeddings
// convention (#2370 — no embeddings module loads unless a run needs one) must
// keep holding while this counter and the checkpoint rules run on the ORDINARY
// path of every analyze. Neither imports anything native.
import {
measurePersistedEmbeddingCount,
persistedEmbeddingCountOrUndefined,
resolvePersistedEmbeddingCount,
} from './embedding-count.js';
import {
EMBEDDING_RESUME_MAX_ATTEMPTS,
decideEmbeddingResume,
mintInterruptedCheckpoint,
mintPartialCheckpoint,
mintUnverifiedCountCheckpoint,
} from './embedding-checkpoint.js';
import type { EmbeddingCheckpoint } from './embedding-checkpoint.js';
/**
* Strip C0/C1 control characters from a progress/diagnostic message.
*
* Several guard notices below interpolate values read straight out of
* `.gitnexus/gitnexus.json`, which is parsed with no runtime shape validation
* (`loadMeta` does a bare `JSON.parse(...) as RepoMeta`) — the stamped schema
* fingerprint, the runner-identity schema, the CJK mode. On the CLI path these
* reach `console.log` and therefore the user's terminal, so a crafted value
* carrying ANSI escapes (`\x1b[2J`, `\x1b]0;…`) would be replayed verbatim.
*
* Sanitizing at the funnel rather than per field: every message that ever
* interpolates untrusted metadata is covered, including ones not written yet.
* Newline and tab are preserved — multi-line notices are intentional.
*/
const stripControlCharacters = (msg: string): string =>
msg.replace(/[\x00-\x08\x0b\x0c\x0e-\x1f\x7f-\x9f]/g, '');
const ANALYSIS_FEATURES = [
CLASS_FRAMEWORK_ANNOTATIONS_FEATURE,
@@ -727,8 +771,9 @@ export const pdgModeMismatch = (recorded: RepoMeta['pdg'], options: PdgOptions):
// different runs would always be `!==`, tripping pdgModeMismatch on every
// re-analyze and forcing a needless full writeback. e.g. do NOT change
// `hasCallSummary: true` to a per-language object like `{ ts: true, ... }`; keep
// the diagnostic per-language refinement in the impact CONSUMER (see
// pdg-impact.ts assemblePdgImpactResult), not in this version discriminator.
// any diagnostic refinement in the impact CONSUMER (see pdg-impact.ts
// assemblePdgImpactResult, which reports empty ascent from the persisted
// CALL_SUMMARY data), not in this version discriminator.
for (const key of new Set([...Object.keys(reqRecord), ...Object.keys(recRecord)])) {
if (reqRecord[key] !== recRecord[key]) return true;
}
@@ -830,7 +875,7 @@ export async function runFullAnalysis(
// would otherwise stay saturated on a reused process).
resetDegradedParseCounter();
const log = (msg: string) => callbacks.onLog?.(msg);
const log = (msg: string) => callbacks.onLog?.(stripControlCharacters(msg));
const acquireOpts = {
log,
onWaitStart: () =>
@@ -889,14 +934,10 @@ async function runFullAnalysisInner(
writeTarget: WriteTarget,
runnerIdentityAtBootstrap?: AnalyzerRunnerIdentity,
): Promise<AnalyzeResult> {
const log = (msg: string) => callbacks.onLog?.(msg);
const log = (msg: string) => callbacks.onLog?.(stripControlCharacters(msg));
const progress = (phase: string, percent: number, message: string) =>
callbacks.onProgress(phase, percent, message);
// Streamed structural emit (#2680), resolved once so the pipeline flag and the
// CSV-dir resolution below cannot disagree.
const streamGraphEmitActive = resolveStreamGraphEmit(options);
// FTS-config validation and the degraded-parse counter reset happen in the
// `runFullAnalysis` wrapper (before the lock is taken).
@@ -1105,38 +1146,38 @@ async function runFullAnalysisInner(
let resumeEmbeddingCheckpoint = false;
let pendingEmbeddingNodeIds = new Set<string>();
let embeddingIdentityForRun: EmbeddingIdentity | undefined;
// The marker this run resumed, so Phase 5 can tell "the retry cleared the
// set" from "the retry failed the same way again" and bound the latter — see
// `nextAttemptCount` in embedding-checkpoint.ts (#2790).
let resumedEmbeddingCheckpoint: EmbeddingCheckpoint | undefined;
if (existingMeta?.embeddingCheckpoint) {
if (options.dropEmbeddings) {
log('Discarding the interrupted embedding checkpoint (--drop-embeddings).');
options = { ...options, force: true };
} else {
const checkpoint = existingMeta.embeddingCheckpoint;
// The verdict itself lives in embedding-checkpoint.ts, shared with
// `POST /api/embed` — two readers of one marker must not be able to
// disagree about what it means.
//
// The identity stays LAZY, as it has to: the flag and retry-budget verdicts
// short-circuit before one is needed, and resolving it means importing an
// embeddings module (#2370 — none loads unless a run actually needs one).
// `decideEmbeddingResume` asks for it by aborting on `undefined`, which is
// the only abort it can reach without one.
let decision = decideEmbeddingResume(checkpoint, undefined, options);
if (decision.action === 'abort') {
const { resolveEmbeddingIdentity } = await import('./embeddings/embedding-identity.js');
embeddingIdentityForRun = resolveEmbeddingIdentity();
const checkpoint = existingMeta.embeddingCheckpoint;
if (checkpoint.provider !== embeddingIdentityForRun.provider) {
throw new Error(
'Cannot resume embedding checkpoint: the embedding provider configuration differs. ' +
'Restore the matching endpoint configuration or pass --drop-embeddings to rebuild without it.',
);
}
if (
checkpoint.model !== embeddingIdentityForRun.model ||
checkpoint.dimensions !== embeddingIdentityForRun.dimensions
) {
throw new Error(
`Cannot resume embedding checkpoint: it uses ${checkpoint.model} at ` +
`${checkpoint.dimensions} dimensions, but this run resolves ` +
`${embeddingIdentityForRun.model} at ${embeddingIdentityForRun.dimensions}. ` +
'Restore the matching embedding configuration or pass --drop-embeddings to rebuild without it.',
);
}
decision = decideEmbeddingResume(checkpoint, embeddingIdentityForRun, options);
}
if (decision.action === 'abort') throw new Error(decision.error);
log(decision.log);
if (options.dropEmbeddings) {
// --drop-embeddings has always implied a rebuild here; the decision only
// covers the marker.
options = { ...options, force: true };
}
if (decision.action === 'resume') {
resumeEmbeddingCheckpoint = true;
pendingEmbeddingNodeIds = new Set(checkpoint.pendingNodeIds ?? []);
log(
`Previous analyze ended at an embedding checkpoint ` +
`(${checkpoint.nodesProcessed}/${checkpoint.totalNodes} nodes); resuming from persisted hashes` +
`${pendingEmbeddingNodeIds.size > 0 ? ` and regenerating ${pendingEmbeddingNodeIds.size} pending node(s)` : ''}.`,
);
pendingEmbeddingNodeIds = new Set(decision.pendingNodeIds);
resumedEmbeddingCheckpoint = decision.resumedFrom;
}
}
@@ -1241,37 +1282,56 @@ async function runFullAnalysisInner(
options = { ...options, force: true };
}
// ── schema-version mismatch forces full rebuild (#2289 P1) ────────
// Mirrors the pdg-mode block above: a stamp from an older
// INCREMENTAL_SCHEMA_VERSION (e.g. pre-v5 URL-only Route ids) cannot be
// reconciled by an incremental top-up — same-commit re-analyze would
// strand stale rows next to new-schema writes. MUST sit before the
// ── schema mismatch forces full rebuild (#2289 P1, #2798) ─────────
// Mirrors the pdg-mode block above: an index whose tables were created from
// a different DDL cannot be reconciled by an incremental top-up — a
// same-commit re-analyze would strand stale rows next to new-schema writes,
// and LadybugDB fixes a relation table's endpoint pairs at CREATE time, so
// edges the old shape cannot hold are simply dropped. MUST sit before the
// alreadyUpToDate fast path below: an unchanged-commit clean tree would
// otherwise early-return without ever reaching the `isIncremental` gate
// that consults `schemaVersion`, defeating the bump's whole point.
// otherwise early-return without ever reaching the `isIncremental` gate.
//
// `schemaVersion === undefined` covers two cases that should still trip
// this guard: a non-git repo (which never stamps the field) and very old
// meta from before the field existed. Non-git repos take the
// `currentCommit === ''` rebuild branch below regardless, so the redundant
// force here is harmless; the friendlier `'pre-versioning'` log avoids a
// user-visible "stamped vundefined" line in that edge case.
if (existingMeta && existingMeta.schemaVersion !== INCREMENTAL_SCHEMA_VERSION) {
const stampedVersion = existingMeta.schemaVersion ?? 'pre-versioning';
// Forcing here is what recreates the schema: `force` makes the run a full
// rebuild, which wipes the database file and re-runs the DDL against an
// empty one. Re-running `CREATE … TABLE` over the EXISTING database would
// not help — runSchemaCreationQueries suppresses "already exists", so the
// new shape would never be applied.
//
// ABSENT covers two cases and forces in both: an index from a GitNexus
// older than this field (the backward-compatibility path — one rebuild, then
// it is stamped), and a non-git repo, which never stamps it (see the meta
// literal below) and takes the `currentCommit === ''` rebuild branch below
// regardless.
//
// The two cases must not be told the same story. Blaming "an older GitNexus
// version" is FALSE for a non-git repo — the field is absent by design there,
// so this build would keep saying it about an index this exact build just
// wrote, on every run, forever. A stamp is only named when it has the shape
// SCHEMA_FINGERPRINT produces; anything else degrades to a neutral
// placeholder, and a non-git repo is additionally told WHY it has no stamp.
if (existingMeta && schemaFingerprintMismatch(existingMeta.schemaFingerprint)) {
const stamped = existingMeta.schemaFingerprint;
const origin = isSchemaFingerprintShaped(stamped) ? stamped : 'an unidentified GitNexus build';
const nonGitNote =
stamped === undefined && !repoHasGit
? ' Non-git repositories never record a schema fingerprint, so this run rebuilds regardless.'
: '';
log(
`index schema changed (stamped v${stampedVersion}, this build is v${INCREMENTAL_SCHEMA_VERSION}); ` +
`forcing a full rebuild so persisted rows match the current schema.`,
`index schema changed (built by ${origin}, this build is ${SCHEMA_FINGERPRINT}); forcing a ` +
`full re-analyze so the database is recreated from the current schema.${nonGitNote}`,
);
options = { ...options, force: true };
}
// ── independently-versioned analysis capabilities ────────────────
// `schemaVersion` is reserved for graph-wide incremental invariants. Some
// `schemaFingerprint` is reserved for graph-wide incremental invariants. Some
// persisted semantics apply only to repositories containing relevant source
// files, so they carry exact feature versions instead. This guard must also
// run before alreadyUpToDate: current main and this PR both use schema v8,
// while pre-PR v8 indexes lack the Class frameworkAnnotations column and
// Java/Kotlin Bean evidence.
// run before alreadyUpToDate: a feature can change what is EXTRACTED without
// changing the DDL, so an index whose `schemaFingerprint` matches this build
// can still be missing that feature's evidence (e.g. the Class
// frameworkAnnotations values, or Java/Kotlin Bean evidence) — the fingerprint
// guard above would wave it through.
const persistedFilePaths = Object.keys(existingMeta?.fileHashes ?? {});
const expectedPersistedAnalysisFeatures = resolveAnalysisFeatureVersions(
ANALYSIS_FEATURES,
@@ -1321,6 +1381,46 @@ async function runFullAnalysisInner(
options = { ...options, force: true };
}
// ── embedding width mismatch forces full rebuild (#2798) ──────────
// The half of the schema `SCHEMA_FINGERPRINT` deliberately cannot cover:
// `CodeEmbedding.embedding` is declared `FLOAT[EMBEDDING_DIMS]`, and that
// width comes from `GITNEXUS_EMBEDDING_DIMS` at module load, so folding it
// into a digest of CODE would make the same build disagree with itself under
// two envs. Without this block a dims flip on a same-commit clean tree fired
// NO guard: the fast path below returned over a FLOAT[384] table while this
// process embedded at 768. The one older reaction (in the embedding-restore
// block further down) discards the CACHE and re-embeds — into a column whose
// width it never revisits.
//
// Forcing is again what repairs it, and for the same reason as the
// fingerprint guard: only a full rebuild wipes the database and re-runs the
// DDL, and `runSchemaCreationQueries` suppresses "already exists", so
// re-running CREATE over the existing DB would silently keep the old width.
// Not conditioned on the index actually holding vectors — the table is
// created for every index either way, and nothing but a rebuild can retype it.
//
// ABSENT is NOT a mismatch here (see embeddingDimsMismatch for the argument):
// it means an index predating the field, whose width is unknown but was
// consistent with the env that wrote it, and which the fingerprint guard
// above already rebuilds — that rebuild is where the stamp lands.
if (existingMeta && embeddingDimsMismatch(existingMeta.embeddingDims, EMBEDDING_DIMS)) {
// Only NAME a recorded width that could be one, for the reason the
// fingerprint guard gates its stamp on `isSchemaFingerprintShaped`:
// meta.json is a schema-less JSON.parse of on-disk state, so a value that
// is not a positive integer is not worth quoting back at the user.
const recordedDims = existingMeta.embeddingDims;
const built =
typeof recordedDims === 'number' && Number.isInteger(recordedDims) && recordedDims > 0
? `FLOAT[${recordedDims}]`
: 'an unrecognized width';
log(
`embedding dimensions changed (index built with ${built}, this run embeds at ` +
`${EMBEDDING_DIMS}); forcing a full rebuild so the vector column is recreated at the ` +
`new width. Tip: set GITNEXUS_EMBEDDING_DIMS (or --embedding-dims) to pin it across runs.`,
);
options = { ...options, force: true };
}
// ── Early-return: already up to date ──────────────────────────────
if (
existingMeta &&
@@ -1504,6 +1604,22 @@ async function runFullAnalysisInner(
// in-place (cache hits leave entries unchanged; misses add new ones).
const parseCache = await loadParseCache(storagePath);
// Streamed structural emit (#2680). Resolved ONCE, so the pipeline flag and
// the CSV-dir resolution below cannot disagree — and resolved HERE, not at
// function entry, because the POSITION is load-bearing: the gate is
// `options.force`, and every freshness guard above REBINDS `options` with
// `force: true` (embedding-checkpoint drop, dirty-flag recovery, pdg-mode
// flip, schema-fingerprint change, analysis-feature drift, runner-identity change,
// CJK-mode change). Resolving before them froze the answer at `false` for
// every rebuild they trigger — including the whole-fleet rebuild an
// schema-fingerprint change forces on every existing index at once,
// which is exactly when the #2649 memory relief matters most. So this MUST
// stay below the last guard that can set `force` and above its first use.
// (The post-pipeline analysis-feature re-check can also set `force`, but the
// pipeline has already run by then; that run emits non-streamed, precisely as
// `resolveStreamPdgEmit` — read fresh at the same point — behaves.)
const streamGraphEmitActive = resolveStreamGraphEmit(options);
// ── Phase 1: Full Pipeline (0–60%) ────────────────────────────────
const pipelineResult = await runPipelineFromRepo(
repoPath,
@@ -1591,7 +1707,10 @@ async function runFullAnalysisInner(
const isIncremental =
!options.force &&
!!existingMeta &&
existingMeta.schemaVersion === INCREMENTAL_SCHEMA_VERSION &&
// Belt and braces, not a second gate: the guard above already set `force`
// on exactly this condition, and `!options.force` short-circuits before
// this conjunct is reached. Kept so the eligibility contract reads whole.
!schemaFingerprintMismatch(existingMeta.schemaFingerprint) &&
currentAnalysisFeatureMismatches.length === 0 &&
!!existingMeta.fileHashes &&
Object.keys(existingMeta.fileHashes).length > 0 &&
@@ -2400,6 +2519,14 @@ async function runFullAnalysisInner(
const stats = await getLbugStats();
let embeddingSkipped = true;
let semanticMode: 'vector-index' | 'exact-scan' | undefined;
// Hoisted out of the Phase 4 block so the Phase 5 gate can tell "the
// pipeline attempted work and produced nothing" apart from "the pipeline
// had nothing to attempt" (#2790). `undefined` ≡ the pipeline never ran.
let embeddingResult: EmbeddingPipelineResult | undefined;
// What Phase 5 stamps as `embeddingCheckpoint`. `undefined` ≡ clear it
// (the clean-run contract). Built inside Phase 4 so it carries the identity
// of the run that actually wrote it — see the assignment below (#2790).
let pendingEmbeddingCheckpoint: RepoMeta['embeddingCheckpoint'];
if (shouldGenerateEmbeddings) {
const { skipForCap, capDisabled, nodeLimit } = deriveEmbeddingCap(
@@ -2484,6 +2611,31 @@ async function runFullAnalysisInner(
}
}
// ── A checkpoint save writes ONLY the checkpoint (#2790) ──────────
// This used to write a full, SUCCESS-shaped meta: new `lastCommit`, new
// `fileHashes`, `incrementalInProgress: undefined`. All three are lies at
// this point in the run. The first `onCheckpointWindowStart` fires at
// batchIndex 0 — before a single embedding row exists — and on a full
// rebuild the graph is still in the unpublished staging DB
// (`${lbugPath}.staging.<uuid>`), which the atomic swap only renames into
// place AFTER Phase 5. A Phase 4 crash therefore threw the whole staging
// build away while leaving a meta claiming the new commit and the new
// file hashes: the next run diffed those advanced hashes, got
// changed=0/added=0/deleted=0, took the incremental path and "preserved"
// the OLD graph forever — the exact log line reported in #2790 — with the
// `incrementalInProgress` crash-recovery contract (repo-manager.ts) also
// cleared mid-run, so nothing could force the rebuild that would heal it.
//
// Freshness fields may only advance once the index is published. So:
// re-read the on-disk meta immediately before writing (the shape the
// /api/embed checkpoint writer in server/api.ts already uses, which also
// keeps a concurrent writer's update from being reverted by a stale
// snapshot) and replace ONLY `embeddingCheckpoint` — plus
// `stats.embeddings` when the caller actually MEASURED the live count
// (the post-window `onCheckpoint`). The window-start callback passes
// nothing: restating the previous run's count there both re-published a
// stale number and clobbered the live count a preceding `onCheckpoint`
// had just written.
const saveEmbeddingCheckpoint = async (
checkpoint: {
nodesProcessed: number;
@@ -2491,48 +2643,33 @@ async function runFullAnalysisInner(
chunksProcessed: number;
},
pendingNodeIds: string[],
embeddings: number | undefined,
embeddings?: number,
): Promise<void> => {
const fileHashes: Record<string, string> = {};
for (const [key, value] of newFileHashes) fileHashes[key] = value;
await saveMeta(metaDir, {
...(existingMeta ?? {}),
const latestMeta = (await loadMeta(metaDir)) ?? existingMeta;
// First-ever analyze of this repo: no meta exists on disk yet (the
// pre-wipe dirty stamp only fires when one does). Mint the minimum
// RepoMeta requires, with `lastCommit: ''` — never `currentCommit` —
// so a crash here cannot make the next run mistake the discarded
// staging build for an indexed commit.
const base: RepoMeta = latestMeta ?? {
repoPath,
lastCommit: currentCommit,
lastCommit: '',
indexedAt: new Date().toISOString(),
runnerIdentity,
branch: branchLabel ?? existingMeta?.branch,
remoteUrl: hasGitDir(repoPath) ? getRemoteUrl(repoPath) : undefined,
stats: {
files: pipelineResult.totalFileCount,
nodes: stats.nodes,
edges: stats.edges,
communities: pipelineResult.communityResult?.stats.totalCommunities,
processes: pipelineResult.processResult?.stats.totalProcesses,
embeddings,
},
schemaVersion: hasGitDir(repoPath) ? INCREMENTAL_SCHEMA_VERSION : undefined,
unresolvedReceiverMembers: summarizeUnresolvedReceivers(
pipelineResult.resolutionOutcomes ?? [],
),
analysisFeatures: currentAnalysisFeatures,
cjkSegmentation: getSearchFTSCjkSegmentation(),
fileHashes: hasGitDir(repoPath) ? fileHashes : undefined,
cacheKeys: [...parseCache.usedKeys],
incrementalInProgress: undefined,
embeddingCheckpoint: {
at: new Date().toISOString(),
...checkpoint,
model: embeddingIdentity.model,
dimensions: embeddingIdentity.dimensions,
provider: embeddingIdentity.provider,
};
await saveMeta(metaDir, {
...base,
...(embeddings === undefined ? {} : { stats: { ...base.stats, embeddings } }),
// Written by a run that is still IN FLIGHT — see the `kind` doc in
// repo-manager.ts.
embeddingCheckpoint: mintInterruptedCheckpoint(
embeddingIdentity,
checkpoint,
pendingNodeIds,
},
pdg: resolvePdgConfig(options),
),
});
};
const embeddingResult = await runEmbeddingPipeline(
embeddingResult = await runEmbeddingPipeline(
executeQuery,
executeWithReusedStatement,
(p) => {
@@ -2551,19 +2688,56 @@ async function runFullAnalysisInner(
{
forceReembedNodeIds: pendingEmbeddingNodeIds,
onCheckpointWindowStart: async ({ nodeIds, ...checkpoint }) => {
await saveEmbeddingCheckpoint(checkpoint, nodeIds, existingMeta?.stats?.embeddings);
await saveEmbeddingCheckpoint(checkpoint, nodeIds);
},
// ── The mid-run count is a DIAGNOSTIC, not a gate (#2790) ──────
// This used to run the count query bare. THIS callback's rejection
// propagates out of `runEmbeddingPipeline` and kills the whole
// analyze, so an unavailable count took the run down BEFORE Phase 5
// ran at all — meaning the tri-state Phase 5 added for exactly this
// case could never execute.
//
// The shared counter (embedding-count.ts) answers `unknown` instead,
// and `undefined` is already `saveEmbeddingCheckpoint`'s "do not
// touch stats.embeddings" signal — so the checkpoint still lands,
// with whatever count is already on disk left alone.
onCheckpoint: async (checkpoint) => {
await checkpointOnce();
const countResult = await executeQuery(
`MATCH (e:${EMBEDDING_TABLE_NAME}) RETURN count(e) AS cnt`,
const measured = await measurePersistedEmbeddingCount(executeQuery);
if (measured.kind === 'unknown') {
log(
`Warning: could not measure persisted embeddings at the embedding checkpoint ` +
`(${measured.reason}); the checkpoint is saved with the last known count.`,
);
}
await saveEmbeddingCheckpoint(
checkpoint,
[],
persistedEmbeddingCountOrUndefined(measured),
);
const countRow = countResult?.[0];
const embeddings = Number(countRow?.cnt ?? countRow?.[0] ?? 0);
await saveEmbeddingCheckpoint(checkpoint, [], embeddings);
},
},
);
// ── A partial run must NOT clear the checkpoint (#2790) ───────────
// Dropped nodes hold zero embedding rows, but "zero rows" alone heals
// nothing: a plain `gitnexus analyze` derives shouldGenerateEmbeddings
// = false whenever the index already has embeddings, so the pipeline is
// never called and the nodes stay missing until someone passes
// --embeddings/--force/--drop-embeddings. Retaining the checkpoint is
// what restores the pre-#2790 heal: the resume path above forces
// shouldGenerateEmbeddings regardless of flags and feeds
// `pendingNodeIds` into `forceReembedNodeIds`. Stamped with THIS run's
// identity so a later model/provider change trips the resume mismatch
// error rather than resuming under a foreign identity.
if (embeddingResult.failedNodeIds.length > 0) {
// `'partial'` and its attempt chain — see the `kind` doc in
// repo-manager.ts and `nextAttemptCount` in embedding-checkpoint.ts.
pendingEmbeddingCheckpoint = mintPartialCheckpoint(
embeddingIdentity,
embeddingResult,
resumedEmbeddingCheckpoint,
);
}
if (embeddingResult.semanticMode === 'exact-scan') {
semanticMode = 'exact-scan';
log(
@@ -2578,25 +2752,125 @@ async function runFullAnalysisInner(
// ── Phase 5: Finalize (98–100%) ───────────────────────────────────
progress('done', 98, 'Saving metadata...');
// Count embeddings in the index (cached + newly generated)
let embeddingCount = 0;
try {
const embResult = await executeQuery(
`MATCH (e:${EMBEDDING_TABLE_NAME}) RETURN count(e) AS cnt`,
// Count embeddings in the index (cached + newly generated). Tri-state, and
// measured by the SHARED counter rather than a local copy — see
// embedding-count.ts. What that buys HERE: the old silent `catch {}` left
// "cannot ask" indistinguishable from "wrote nothing", so a diagnostic
// failure crashed the run at the gate below with no clue why.
const measuredEmbeddingCount = await measurePersistedEmbeddingCount(executeQuery);
const embeddingCount = persistedEmbeddingCountOrUndefined(measuredEmbeddingCount);
if (measuredEmbeddingCount.kind === 'unknown') {
// Not silent any more: the operator gets the reason the count is unknown.
log(
`Warning: could not count persisted embeddings ` +
`(${measuredEmbeddingCount.reason}); treating the embedding count as unknown.`,
);
const row = embResult?.[0];
embeddingCount = Number(row?.cnt ?? row?.[0] ?? 0);
} catch {
/* table may not exist if embeddings never ran */
}
if (!embeddingSkipped && stats.nodes > 0 && embeddingCount === 0) {
// ── Phase 5 embedding gate (#2790) ────────────────────────────────
// Four genuinely different states used to collapse into
// `embeddingCount === 0`, and the gate hard-crashed on three of them:
// 1. the pipeline never ran (cap-skipped / not requested) —
// `embeddingSkipped`, still short-circuited;
// 2. the pipeline ran but had NOTHING to embed (totalNodes 0 after the
// incremental filter — e.g. a resume whose pending sweep deleted the
// last rows) over a legitimately empty table;
// 3. the count query failed or answered non-numerically (above) — a
// diagnostic failure, not an indexing failure;
// 4. the pipeline embedded and NOTHING persisted — the real defect.
// Only (4) throws. `attemptedEmbedding` is what separates it from (2):
// `nodesProcessed` is now the REAL completed-node count and
// `failedNodeIds` names the nodes whose rows were dropped, so
// "attempted" ≡ at least one node was walked to a conclusion.
const attemptedEmbedding =
!embeddingSkipped &&
embeddingResult !== undefined &&
(embeddingResult.nodesProcessed > 0 || embeddingResult.failedNodeIds.length > 0);
if (attemptedEmbedding && stats.nodes > 0 && embeddingCount === 0) {
throw new Error(
'Embedding generation completed without persisted embeddings. ' +
'The index was not registered to avoid silently reporting embeddings: 0.',
'The index was not registered to avoid silently reporting embeddings: 0. ' +
'Check the embedding endpoint/model configuration (GITNEXUS_EMBEDDING_URL / ' +
'GITNEXUS_EMBEDDING_MODEL) and re-run `gitnexus analyze --embeddings`; ' +
'the graph itself is unaffected, so `--drop-embeddings` indexes without them.',
);
}
if (embeddingCount === undefined) {
log(
'Warning: registering the index without a verified embedding count — the count query ' +
'did not answer, so stats.embeddings falls back to the last known value. ' +
'Re-run `gitnexus analyze --embeddings` if semantic search comes back empty.',
);
}
// ── An unverifiable count must leave a way back (#2790) ───────────────
// The carry-forward below is a GUESS, and the guess is load-bearing (see
// embedding-count.ts). Clearing the checkpoint on top of it would report
// unqualified success and erase the only record that this index was never
// verified.
//
// Retain an identity-matching recovery marker instead — `'unverified-count'`,
// whose whole job is to force the next run past the same-commit fast return
// so the count can be re-derived (see the `kind` doc in repo-manager.ts).
// Self-limiting: once the count answers, the marker is cleared, and the run
// it forces embeds nothing, so `attemptedEmbedding` is false and nothing is
// re-planted.
if (
attemptedEmbedding &&
embeddingCount === undefined &&
pendingEmbeddingCheckpoint === undefined &&
embeddingIdentityForRun !== undefined
) {
log(
'Retaining an embedding checkpoint so the next `gitnexus analyze` re-derives the count ' +
'instead of publishing an unverified one as final (#2790).',
);
pendingEmbeddingCheckpoint = mintUnverifiedCountCheckpoint(embeddingIdentityForRun, {
nodesProcessed: embeddingResult?.nodesProcessed ?? 0,
totalNodes: embeddingResult?.nodesProcessed ?? 0,
chunksProcessed: embeddingResult?.chunksProcessed ?? 0,
});
}
// A partial index that is honest about itself beats no index — see the
// `kind` doc in repo-manager.ts for why the dropped nodes are safe to ship.
if (embeddingResult !== undefined && embeddingResult.failedNodeIds.length > 0) {
log(
`Warning: ${embeddingResult.failedNodeIds.length} node(s) lost their embeddings to ` +
'embedding-endpoint failures and were dropped from this index (#2790). ' +
'They are recorded as an embedding checkpoint, so the next `gitnexus analyze` run ' +
'resumes from it and re-embeds exactly those nodes — `gitnexus status` reports the ' +
`index as incomplete until it succeeds, and the retry gives up after ` +
`${EMBEDDING_RESUME_MAX_ATTEMPTS} consecutive failures rather than staying incomplete ` +
'forever. Pass --force or --drop-embeddings to abandon them instead.',
);
}
// What we can honestly persist as the embedding count: the measurement when
// there is one, else the LAST KNOWN figure — never a fabricated 0 (see
// embedding-count.ts). Folded by the shared `resolvePersistedEmbeddingCount`
// so the CLI and the server cannot drift apart on the carry-forward the way
// they already had on the measurement.
//
// "Last known" is the LATEST ON-DISK meta, re-read here (#2790). It used to
// read `existingMeta`, which is assigned exactly once — before any
// embedding work — so the fallback republished the pre-run figure and
// OVERWROTE the fresher count this run's own terminal `onCheckpoint` had
// already written to disk: prior meta says 0, a clean run inserts
// embeddings and checkpoints the real count, the final probe is
// unavailable, and finalization carries the stale 0 forward while reporting
// success. `loadMeta` never throws (it returns null), and the checkpoint
// writer already re-reads the same way, so this is the same freshness
// discipline applied to the same field.
const latestMetaForCount =
embeddingCount === undefined ? ((await loadMeta(metaDir)) ?? existingMeta) : undefined;
const persistedEmbeddingCount = resolvePersistedEmbeddingCount(
measuredEmbeddingCount,
latestMetaForCount?.stats?.embeddings,
);
const { getRuntimeCapabilities } = await import('./platform/capabilities.js');
const runtimeCapabilities = getRuntimeCapabilities();
// `semanticMode` is authoritative when set (Phase 4 reported what it
@@ -2655,7 +2929,7 @@ async function runFullAnalysisInner(
edges: stats.edges,
communities: pipelineResult.communityResult?.stats.totalCommunities,
processes: pipelineResult.processResult?.stats.totalProcesses,
embeddings: embeddingCount,
embeddings: persistedEmbeddingCount,
},
capabilities: {
graph: { provider: 'ladybugdb', status: runtimeCapabilities.graph },
@@ -2669,7 +2943,20 @@ async function runFullAnalysisInner(
},
vectorSearch: {
provider: effectiveSemanticMode === 'vector-index' ? 'ladybugdb-vector' : 'exact-scan',
status: embeddingCount > 0 ? effectiveSemanticMode : 'unavailable',
// Reads the MEASURED count, not `persistedEmbeddingCount` (#2790).
// The carry-forward exists so a later `--force` doesn't discard a
// live cache — it is a guess, and a guess must never certify the
// vector lane: `--drop-embeddings` + a failed count probe would
// otherwise stamp 'vector-index' with 5000 embeddings over a table
// holding zero. Unknown reads as 'unavailable' (the status union has
// no unknown member, and adding one would touch every consumer);
// the downgrade is recoverable — 'unavailable' is not carried
// forward as `persistedSemanticMode`, so the next run that can
// count restamps the real mode.
status:
embeddingCount !== undefined && embeddingCount > 0
? effectiveSemanticMode
: 'unavailable',
exactScanLimit: runtimeCapabilities.exactScanLimit,
reason: runtimeCapabilities.reason,
},
@@ -2678,7 +2965,9 @@ async function runFullAnalysisInner(
// analyze run can take the incremental DB-writeback path. Setting
// incrementalInProgress to undefined explicitly clears any prior
// dirty flag (full and incremental success paths converge here).
schemaVersion: hasGitDir(repoPath) ? INCREMENTAL_SCHEMA_VERSION : undefined,
// Derived digest of the DDL this run created the tables from (#2798).
// Git-only: non-git repos never take the incremental path.
schemaFingerprint: hasGitDir(repoPath) ? SCHEMA_FINGERPRINT : undefined,
unresolvedReceiverMembers: summarizeUnresolvedReceivers(
pipelineResult.resolutionOutcomes ?? [],
),
@@ -2687,6 +2976,11 @@ async function runFullAnalysisInner(
// `pdg` below, 'none' is a meaningful value to compare, not an
// absence, so this is never conditionally omitted.
cjkSegmentation: getSearchFTSCjkSegmentation(),
// The FLOAT[N] width this run created the vector column at (#2798).
// Always stamped, like `cjkSegmentation` and unlike `schemaFingerprint`:
// the CodeEmbedding table is created for every index, git or not, so
// absence has exactly one meaning — an index older than the field.
embeddingDims: EMBEDDING_DIMS,
fileHashes: hasGitDir(repoPath) ? newFileHashesRecord : undefined,
// This branch's full live chunk-key set (#2106 R6). `usedKeys` is every
// chunk hash touched in this scan — cache HITS included (see parse-impl
@@ -2694,7 +2988,9 @@ async function runFullAnalysisInner(
// so a sibling branch's prune can union it and not evict our shards.
cacheKeys: [...parseCache.usedKeys],
incrementalInProgress: undefined as RepoMeta['incrementalInProgress'],
embeddingCheckpoint: undefined,
// Cleared on a clean run; otherwise the marker Phase 4/5 minted above
// (see the `kind` doc in repo-manager.ts).
embeddingCheckpoint: pendingEmbeddingCheckpoint,
// The effective pdg config this run's DB rows were built under
// (#2099 F1). `undefined` on pdg-off runs — this meta is a fresh
// literal (no spread of existingMeta), so omission is what CLEARS the
+5 -8
View File
@@ -8,8 +8,10 @@
// tri-review Residual-1: `classifyFtsQueryError` now lives in lbug-adapter.ts
// (see its doc comment) so `queryFTS`'s own catch can share the SAME
// classifier instead of maintaining a second, independently-drifting copy
// for the identical `QUERY_FTS_INDEX` cypher call.
import { queryFTS, classifyFtsQueryError } from '../lbug/lbug-adapter.js';
// for the identical `QUERY_FTS_INDEX` cypher call. `buildFtsQueryCypher` is
// that identical call itself — shared for the same reason, so the two paths
// can no longer drift apart the way they had to be edited in lockstep before.
import { queryFTS, classifyFtsQueryError, buildFtsQueryCypher } from '../lbug/lbug-adapter.js';
import { normalizeFtsText } from '../lbug/csv-generator.js';
import { getExtensionCapabilities } from '../lbug/extension-loader.js';
import { redactPaths } from './fts-indexes.js';
@@ -72,12 +74,7 @@ async function queryFTSViaExecutor(
query: string,
limit: number,
): Promise<FTSQueryOutcome> {
const cypher = `
CALL QUERY_FTS_INDEX('${tableName}', '${indexName}', $query, conjunctive := false)
RETURN node, score
ORDER BY score DESC
LIMIT ${limit}
`;
const cypher = buildFtsQueryCypher(tableName, indexName, limit);
try {
const rows = await executor(cypher, { query });
return {
+11 -2
View File
@@ -177,6 +177,13 @@ export async function getInterModuleCallEdges(filePaths: string[]): Promise<{
const fileList = filePaths.map((f) => `'${escapeCypherString(f)}'`).join(', ');
// The sort leads with the symbol names, not the file paths. Ordering by
// `fromFile` first makes the LIMIT a single-file prefix — on this repo's own
// index the 30 outgoing edges of `core/wiki` all came from 1 of its 7 files,
// so the module page described one file's external surface as the module's.
// The four columns are the whole DISTINCT tuple, so any permutation is a
// total order and equally deterministic (#2787); leading with the names just
// spreads the window across files (1 → 7 of 7 here).
const outRows = await executeQuery(
REPO_ID,
`
@@ -184,6 +191,7 @@ export async function getInterModuleCallEdges(filePaths: string[]): Promise<{
WHERE a.filePath IN [${fileList}] AND NOT b.filePath IN [${fileList}]
RETURN DISTINCT a.filePath AS fromFile, a.name AS fromName,
b.filePath AS toFile, b.name AS toName
ORDER BY fromName, toName, fromFile, toFile
LIMIT 30
`,
);
@@ -195,6 +203,7 @@ export async function getInterModuleCallEdges(filePaths: string[]): Promise<{
WHERE NOT a.filePath IN [${fileList}] AND b.filePath IN [${fileList}]
RETURN DISTINCT a.filePath AS fromFile, a.name AS fromName,
b.filePath AS toFile, b.name AS toName
ORDER BY fromName, toName, fromFile, toFile
LIMIT 30
`,
);
@@ -232,7 +241,7 @@ export async function getProcessesForFiles(filePaths: string[], limit = 5): Prom
WHERE s.filePath IN [${fileList}]
RETURN DISTINCT p.id AS id, p.heuristicLabel AS label,
p.processType AS type, p.stepCount AS stepCount
ORDER BY stepCount DESC
ORDER BY stepCount DESC, id
LIMIT ${limit}
`,
);
@@ -281,7 +290,7 @@ export async function getAllProcesses(limit = 20): Promise<ProcessInfo[]> {
MATCH (p:Process)
RETURN p.id AS id, p.heuristicLabel AS label,
p.processType AS type, p.stepCount AS stepCount
ORDER BY stepCount DESC
ORDER BY stepCount DESC, id
LIMIT ${limit}
`,
);
+69
View File
@@ -2,6 +2,75 @@ export const generateId = (label: string, name: string): string => {
return `${label}:${name}`;
};
/**
* Total order on strings by UTF-16 code unit — deliberately NOT `localeCompare`.
*
* `localeCompare` resolves against the host's default ICU locale, so two machines
* can order the same pair differently and a tie-broken cap keeps a different
* subset on each (#2787). Code-unit order is host-independent and matches the
* binary `ORDER BY` the graph queries use, so a JS re-sort reproduces the DB's
* order instead of fighting it.
*
* Use this for every tiebreak that feeds a `.slice()`, a page boundary, or any
* other cut — `Array.prototype.sort` is stable, so a comparator that returns 0
* for distinct elements silently falls back to input order.
*/
export const compareCodeUnits = (a: string, b: string): number => (a < b ? -1 : a > b ? 1 : 0);
/**
* How many links of an `Error.cause` chain any walker below visits, counting
* the head. THE bound — hand-rolled copies had drifted to three different
* numbers with three different loop conditions (`depth < 5` in two places,
* `depth <= 5` in a third, i.e. six levels).
*
* The bound exists purely so a cyclic chain (`a.cause = b; b.cause = a`)
* cannot loop forever. Real chains are one or two links deep — the ingestion
* phase runner wraps a phase failure once as
* `new Error("Phase 'X' failed: …", { cause })` — so five leaves ample
* headroom for future nesting.
*/
export const CAUSE_CHAIN_MAX_DEPTH = 5;
/**
* Walk `err` and its `cause` chain, head first, yielding each `Error` link.
*
* Stops at the first non-`Error` link (a `cause` may legally be any value, and
* a non-Error carries no further `cause` worth following) and after
* `maxDepth` links. Non-`Error` input yields nothing.
*
* THE single cause-chain traversal. Consume this (or {@link findInCauseChain})
* rather than re-rolling the `for (let depth = 0; …; current = current.cause)`
* loop: every hand-rolled copy has to re-decide the bound and the loop
* condition, and they did not agree.
*/
export function* causeChain(err: unknown, maxDepth = CAUSE_CHAIN_MAX_DEPTH): Generator<Error> {
let current: unknown = err;
for (let depth = 0; depth < maxDepth && current instanceof Error; depth++) {
yield current;
current = (current as { cause?: unknown }).cause;
}
}
/**
* Find the first link of `err`'s cause chain that `match` accepts.
*
* Load-bearing for CLI error classification: the ingestion phase runner
* rewraps every phase failure as `new Error("Phase 'X' failed: …", { cause })`,
* so a bare `instanceof` at the CLI boundary misses the real error entirely
* and falls through to a generic stack dump. Classify by TYPE through this
* helper (the repo norm from #2385), never by message text.
*/
export function findInCauseChain<T>(
err: unknown,
match: (e: unknown) => e is T,
maxDepth: number = CAUSE_CHAIN_MAX_DEPTH,
): T | undefined {
for (const link of causeChain(err, maxDepth)) {
if (match(link)) return link;
}
return undefined;
}
/**
* Drop a Windows extended-length (`\\?\`) prefix from a path (#2667).
*
+7
View File
@@ -23,6 +23,11 @@ export async function queryClassBeanMetadata(
symbolType === 'Method'
? 'MATCH (m:Method {id: $symbolId})-[r:CodeRelation]->(b:CodeElement)'
: 'MATCH (m:Method)-[r:CodeRelation]->(b:CodeElement {id: $symbolId})';
// determinism: probe — PK-anchored singleton. A @Bean factory's declaration
// node id is derived from the declaring method's id
// (`CodeElement:spring-bean:<methodId>`, see bean-factories.ts), so exactly
// one DECLARES edge with this reason prefix exists per anchored endpoint,
// whichever end `$symbolId` pins.
const rows = await executeParameterized(
lbugPath,
`${pattern}
@@ -36,6 +41,8 @@ export async function queryClassBeanMetadata(
return row === undefined ? undefined : decodeSpringBeanFactoryReason(row.reason ?? row[0]);
}
// determinism: probe — PK-anchored singleton. `$symbolId` is a node primary
// key, so at most one Class row can match and the LIMIT never chooses.
const rows = await executeParameterized(
lbugPath,
`
File diff suppressed because it is too large Load Diff
+601 -85
View File
@@ -10,12 +10,18 @@ import path from 'path';
import type { executeParameterized } from '../../core/lbug/pool-adapter.js';
import { loadMeta } from '../../storage/repo-manager.js';
import { IMPACT_MAX_DEPTH, PDG_QUERY_DEFAULT_LIMIT, PDG_QUERY_MAX_LIMIT } from '../tools.js';
import { CALLEES_TRUNCATED_SENTINEL, CALLEE_ID_SEP } from '../../core/ingestion/cfg/emit.js';
// Imported from the LEAF `callee-cell-format.js`, NOT from `cfg/emit.js` which
// re-exports them: ESM evaluates a module to import any binding from it, and
// `emit.ts` drags the analyze-only CFG closure (reaching-defs, control-
// dependence, post-dominators, synthetic-escape, call-site-harvest) with it —
// 8 modules on every MCP server start, to read two strings (#2802 review).
import {
CALLEES_TRUNCATED_SENTINEL,
CALLEE_ID_SEP,
} from '../../core/ingestion/cfg/callee-cell-format.js';
import { toDisplayLine } from './line-display.js';
import { decodeCallSummary } from '../../core/ingestion/taint/call-summary-codec.js';
import { decodeReachingDefReason } from '../../core/ingestion/cfg/reaching-def-reason-codec.js';
import { getProviderForFile } from '../../core/ingestion/languages/index.js';
import { SupportedLanguages } from 'gitnexus-shared';
/**
* Parse the `<fnLine>` segment out of a `BasicBlock` id (1-based function start
@@ -66,30 +72,44 @@ const INTERPROC_DEPTH_BUDGET = 3;
const INTERPROC_NODE_BUDGET = 5000;
/**
* Split a tab-joined ({@link CALLEE_ID_SEP}) `BasicBlock.calleeIds` cell into its resolved callee
* symbol ids, dropping the truncation sentinel (a capped block carries the
* sentinel to mark an incomplete call-site list; it is NOT a resolved symbol id
* and must never enter a `has(realId)` set). Empty/whitespace cells yield no ids.
* Parse a tab-joined ({@link CALLEE_ID_SEP}) `BasicBlock.calleeIds` cell in ONE
* pass into its resolved callee symbol ids plus whether the cell was CAPPED at
* emit. The truncation sentinel is NOT a resolved symbol id and must never enter
* a `has(realId)` set, so it is dropped from `ids` — which is exactly why
* `truncated` has to come back alongside them: without it a capped block is
* indistinguishable from a complete one and the dropped callees are invisible to
* every consumer that only reads `ids` (they reach neither the `CALL_SUMMARY`
* scan nor the ascent counters). Empty/whitespace cells yield no ids.
*
* Extracted here (U1) so the two callers — `LocalBackend.calleeIdsOfBlocks` (the
* statement-precise bridge key) and the inter-procedural descent's
* `calleeIdsFromCalleeRows` — cannot diverge on the split-and-drop-sentinel
* logic. Both consume rows of `BasicBlock.calleeIds`; this is the single source.
* The callgraph bridge already treats a capped block as callee-INCOMPLETE (see
* `classifyPdgBridgeEvidence`); this is the same fact, read at the descent side.
*/
export function splitCalleeIds(raw: unknown): string[] {
const out: string[] = [];
function parseCalleeIdsCell(raw: unknown): { ids: string[]; truncated: boolean } {
const ids: string[] = [];
let truncated = false;
// Split on the SHARED CALLEE_ID_SEP (tab) — ids embed file paths / multi-word
// C++ type tokens that can contain a space, so a space split would fragment
// them. Producer (calleeIdsOfBlock) joins with the same constant.
for (const id of String(raw ?? '').split(CALLEE_ID_SEP)) {
if (id && id !== CALLEES_TRUNCATED_SENTINEL) out.push(id);
if (id === CALLEES_TRUNCATED_SENTINEL) truncated = true;
else if (id) ids.push(id);
}
return out;
return { ids, truncated };
}
/**
* Ids-only view of {@link parseCalleeIdsCell}. Exported (U1) so the two callers —
* `LocalBackend.calleeIdsOfBlocks` (the statement-precise bridge key) and the
* inter-procedural descent — cannot diverge on the split-and-drop-sentinel logic.
* Both consume rows of `BasicBlock.calleeIds`; this is the single source.
*/
export function splitCalleeIds(raw: unknown): string[] {
return parseCalleeIdsCell(raw).ids;
}
/**
* Contract version of the mode:'pdg' impact result shape. A stable discriminator
* for external MCP/agent consumers — distinct from the DB INCREMENTAL_SCHEMA_VERSION.
* for external MCP/agent consumers — distinct from the DB schema fingerprint.
* Bump on any breaking change to the PDG result fields.
* v2: `startLine` in the result is now 1-based display (#2380), matching the
* context/query/impact tools (was 0-based).
@@ -552,6 +572,132 @@ export type PdgImpactEvidence =
| 'unproven-bridge'
| 'degraded';
/**
* WHY the callee set the descent EXAMINED for `CALL_SUMMARY` return-flows is a
* strict PREFIX of the slice's real callee list. A STRUCTURED vocabulary, in the
* spirit of `truncatedByReasons: readonly ('depth'|'limit')[]` — the codes are the
* contract, the English phrasing is a rendering of them
* ({@link ASCENT_INCOMPLETE_PHRASE}). A caller branches on the code; only the note
* reads the phrase, so a rewording can never break a consumer.
* - `'traversal-truncated'` — the traversal stopped at its depth/size budget, so
* a callee that DOES carry a return-flow can sit in a hop never reached (the
* same fact the result's `truncated`/`truncatedByReasons` report, read at the
* ascent's granularity).
* - `'callee-list-capped'` — a slice block's `calleeIds` cell was capped at emit;
* `parseCalleeIdsCell` strips the sentinel, so the dropped ids are invisible to
* BOTH the summary scan and the counters.
* - `'callee-ids-unrecorded'` — a slice block records CALL SITES (a non-empty
* `callees` name cell) but NO resolved callee ids. An empty `calleeIds` cell
* carries no sentinel, so those call sites are invisible to the summary scan
* without raising `'callee-list-capped'`. Distinct from the cap: nothing was
* dropped at emit — the ids were never recorded.
*
* THREE producer paths yield it, and the consumer cannot tell them apart —
* do not read this code as naming any one of them (`cfg/emit.ts`,
* `calleeIdsOfBlock`):
* 1. the file's resolved-id map is absent entirely (`fileMap === undefined`);
* 2. a call site has no position anchor;
* 3. a call site's position IS in the map but did not RESOLVE.
* (3) is the ordinary one — it is exactly the receiver-resolution gaps this
* repo pins (e.g. #2807's inference-typed field receivers, where `calleeIds`
* empties while `calleesOfBlock` still writes the leaf names). So on a real
* index this fires broadly and is driven by resolution quality, NOT by a
* missing `--pdg` layer: "re-run analyze --pdg" is the wrong remedy for it,
* and `examinedComplete: false` here is a statement about how much of the
* call graph resolved, not about the traversal giving up.
*
* Consequence worth knowing before branching on it: because (3) is common,
* `examinedComplete: true` is the strong, rare signal and `false` is close to
* the default on a large repo. Distinguishing the three needs a marker at
* emit time, which would move the persisted cell format — deliberately out of
* scope here, and tracked separately.
*/
export type PdgAscentIncompleteReason =
| 'traversal-truncated'
| 'callee-list-capped'
| 'callee-ids-unrecorded';
/**
* Return-value-ascent coverage, published on {@link PdgImpactEvidenceSummary} so a
* consumer can answer "was the ascent complete, and if not why" WITHOUT parsing
* the result `note`. This is MCP output read by agents: the same four facts are
* also narrated in the note (see {@link AscentCoverage} for the canonical
* rationale, and `assemblePdgImpactResult` for the single site that renders both
* from one computation), and the prose is the human surface, not the contract.
*
* Present iff the inter-procedural descent RAN (a downstream slice that reached
* the assembly path). Absent ⇒ nothing was scanned — deliberately not a zeroed
* object, which would read as "we looked and found nothing".
*/
export interface PdgAscentCoverage {
/**
* Count of DISTINCT callee ids scanned for a `CALL_SUMMARY` — a distinct-callee
* tally, NOT a call-site count: two slice blocks invoking the same callee
* contribute 1, and one block invoking it twice also contributes 1. That is the
* correct population for the claim `returnFlowFound` makes, because a
* `CALL_SUMMARY` is a property of the CALLEE, not of the call site.
*
* Callee granularity, NOT "callees resolved to a body": a cell's ids that
* `resolveCalleeSpans` never matches (out-of-repo target, interface method, the
* `Class:` id a `new` expression contributes) are scanned all the same. See
* {@link AscentCoverage} POPULATION. (The field NAME is historical — the
* published shape is versioned by `pdgResultVersion`, so it is kept while the
* prose on both surfaces says "distinct callees".)
*/
referencesScanned: number;
/**
* Whether ANY scanned callee carried a DECODED non-empty return-flow — i.e.
* whether the ascent FIRED anywhere in this slice. `false` with
* `referencesScanned > 0` is the structural counterpart of the note's
* "no return-value ascent in this slice" sentence.
*/
returnFlowFound: boolean;
/**
* Scanned callees whose `CALL_SUMMARY` the codec could not decode (version skew
* / corruption / NULL reason). Each withholds the ascent exactly like an empty
* summary, so a non-zero count means `returnFlowFound: false` is NOT a statement
* about what the persisted summaries record — remedy: re-run `analyze --pdg`.
*/
undecodableSummaryCount: number;
/**
* Whether {@link referencesScanned} ranges over EVERY callee the descent's slice
* blocks recorded a resolved id for. `false` ⇒ the counters above range over a
* strict subset, so `returnFlowFound: false` is not a whole-slice claim. Reasons
* in {@link incompleteReasons}.
*
* SCOPE — the population is "callees the INDEX recorded resolved ids for on the
* blocks the DESCENT visited", never "every call the source text makes". Two
* gaps are named structurally rather than assumed away: a block whose ids the
* emitter capped raises `'callee-list-capped'`, and a block that records call
* sites but no ids at all raises `'callee-ids-unrecorded'`. What is NOT modelled
* (and cannot be, from the persisted graph) is a call the CFG never materialised
* as a call site — so `true` means "nothing the index recorded was skipped", not
* "the program makes no other calls".
*/
examinedComplete: boolean;
/** Empty iff {@link examinedComplete}; otherwise every mechanism that fired. */
incompleteReasons: readonly PdgAscentIncompleteReason[];
/**
* Whether the index carries the `CALL_SUMMARY` layer at all. `false` ⇒ a PRE-FU-C
* (v3) `--pdg` index, where the scan COULD NOT have found a return-flow — without
* this a consumer would read `returnFlowFound: false` as "these callees record no
* return-flow" when the truth is "the layer that records it does not exist here"
* (the note distinguishes the two in prose; this keeps the structured surface
* from being false-safe). Remedy: re-run `analyze --pdg`.
*
* READING IT WITH THE OTHER FIELDS. `{referencesScanned: N>0, returnFlowFound:
* false, callSummaryLayerPresent: false}` is SELF-CONSISTENT and expected on a
* v3 index, not a contradiction: the scan really did run over N callees and
* really did find nothing, because there was no layer in which a return-flow
* could be recorded. Read this field FIRST — while it is `false`,
* `returnFlowFound` and `undecodableSummaryCount` say nothing about the callees
* themselves and must not be used to conclude "no callee returns a
* slice-dependent value". `examinedComplete` is orthogonal to all of this: it
* reports coverage of the callee POPULATION, never the presence of the layer.
*/
callSummaryLayerPresent: boolean;
}
export interface PdgImpactEvidenceSummary {
statements?: PdgImpactEvidence;
localSymbols?: PdgImpactEvidence;
@@ -560,6 +706,12 @@ export interface PdgImpactEvidenceSummary {
unresolvedBlockCount?: number;
ambiguousProjectionCount?: number;
interproceduralEvidenceCounts?: Partial<Record<PdgImpactEvidence, number>>;
/**
* Return-value-ascent coverage — the same "counts + classification" kind as the
* three counters above, scoped to the U-C4 ascent. Optional because the descent
* does not run for every slice; see {@link PdgAscentCoverage}.
*/
ascent?: PdgAscentCoverage;
}
export interface PdgInterproceduralImpact {
@@ -723,6 +875,130 @@ export function makePdgLayerDegradedResult(input: {
};
}
/**
* What the inter-procedural descent OBSERVED about return-value ascent, threaded
* from `interproceduralDescent` through `runImpactPDG` to `assemblePdgImpactResult`.
* When references were scanned but none carried a decodable return-flow the note
* reports that the ascent was structurally empty for this slice. This is the
* DESCENT-SIDE record; `assemblePdgImpactResult` renders it into TWO surfaces —
* the result `note` (prose, for humans) and `pdgEvidence.ascent`
* ({@link PdgAscentCoverage}, structured, for agents) — from ONE computation, so
* the two can never disagree. CANONICAL rationale for all four members; the sites
* that thread it point here.
*
* POPULATION. {@link references} is the DISTINCT-ID tally of the slice's
* `BasicBlock.calleeIds` cells — distinct CALLEES, deliberately NOT call sites
* and NOT "callees the descent resolved to a body". Two distinctions, both
* load-bearing:
* - not call SITES: the accumulator is a `Set` of callee ids, so two blocks
* invoking the same callee are tallied once. That is the right population,
* because a `CALL_SUMMARY` is a property of the CALLEE — scanning the same id
* twice could not change the answer, and quoting a site count would over-state
* the size of the set the universal claim ranges over;
* - not "resolved to a body": `resolveCalleeSpans` matches only
* `Function`/`Method`/`Constructor`, so an out-of-repo target, an interface
* method, and a node kind with no CFG body (e.g. the `Class:` id a `new`
* expression contributes) yield no span and are never descended into — yet
* they ride the same cell and ARE scanned for a `CALL_SUMMARY`. Narrowing to
* the descended set would under-state what was checked, and "resolved" would
* assert a symbol-table lookup that did not happen for part of the set.
* Both surfaces must word it at DISTINCT-CALLEE granularity.
*
* OBSERVED DATA, NEVER THE CRITERION'S LANGUAGE (#2802 — a reviewer asked why
* the note does not just look the language up). Whether a callee's return value
* can be ascended is a property of its persisted `CALL_SUMMARY`, not of a
* language name, so asking the graph is both correct for every language and
* correct as producers change: a harvester that starts recording formal indices
* needs no edit at the note site, and a callee that genuinely has no return-flow
* is never described as if the ascent had covered it. This module must not name
* languages — nor must the shared `core/ingestion` pipeline.
*
* INCOMPLETENESS. {@link undecodable} keeps the note honest: without it an
* unreadable summary would be reported as one that records no return-flow. Three
* mechanisms can make the examined set a strict SUBSET of the slice's real callee
* list — the traversal's own truncation flags, {@link listTruncated}, and
* {@link idlessCallSites} — which is what stops the note quantifying universally
* ("none of the N callees …") over a set it knows is incomplete. Those three are
* what {@link PdgAscentIncompleteReason} names structurally for the published
* surface. The traversal flags are DELIBERATELY the result-level ones: a callee's
* own intra BFS exhausting the depth budget hides deeper call sites exactly the
* way the top-level intra BFS does, so both fold into the same signal.
*/
interface AscentCoverage {
/** Count of DISTINCT callee ids scanned for a `CALL_SUMMARY`. */
readonly references: number;
/** Whether ANY scanned callee carried a DECODED non-empty return-flow. */
readonly anyReturnFlow: boolean;
/** Scanned callees whose `CALL_SUMMARY` the codec could not decode. */
readonly undecodable: number;
/** Whether any gathered block's `calleeIds` cell was CAPPED at emit. */
readonly listTruncated: boolean;
/**
* Whether any gathered block recorded CALL SITES (a non-empty `callees` name
* cell) but NO resolved callee ids — an empty `calleeIds` cell, which carries no
* cap sentinel and so is invisible to {@link listTruncated}. Those call sites
* are silently outside {@link references}.
*/
readonly idlessCallSites: boolean;
}
/**
* Render table for {@link PdgAscentIncompleteReason} — the ONLY place a code
* becomes English. Keeping the mapping here (rather than building sentences at
* the point the mechanism is detected) is what lets the published vocabulary and
* the note's wording move independently: a reworded phrase is invisible to every
* consumer branching on the code, and a new code cannot silently change the
* joiner the existing sentence uses.
*/
const ASCENT_INCOMPLETE_PHRASE: Readonly<Record<PdgAscentIncompleteReason, string>> = {
'traversal-truncated': 'the traversal stopped at its depth/size budget',
'callee-list-capped': "a slice block's call-site list was capped at emit",
'callee-ids-unrecorded': 'a slice block records call sites but no resolved callee ids',
};
/**
* The ONE place {@link AscentCoverage} plus the result-level truncation flag
* become the published {@link PdgAscentIncompleteReason} codes. Extracted so the
* two exits that publish coverage — `assemblePdgImpactResult` (the slice result,
* which also renders the codes into the note) and `runImpactPDG`'s empty-slice
* return — cannot classify the same descent differently.
*
* Emission order is the array order below and is part of what the note renders,
* so a new code appends rather than inserts.
*/
function ascentIncompleteReasonsOf(input: {
truncated: boolean;
coverage: AscentCoverage | undefined;
}): PdgAscentIncompleteReason[] {
const reasons: PdgAscentIncompleteReason[] = [];
if (input.truncated) reasons.push('traversal-truncated');
if (input.coverage?.listTruncated === true) reasons.push('callee-list-capped');
if (input.coverage?.idlessCallSites === true) reasons.push('callee-ids-unrecorded');
return reasons;
}
/**
* Project the descent's {@link AscentCoverage} onto the published
* {@link PdgAscentCoverage}. Shared by both exits that publish `pdgEvidence.ascent`
* so the contract sentence "present iff the inter-procedural descent ran" holds on
* BOTH — a descent that ran and scanned callees must not go unreported merely
* because the slice happened to reach no DISTINCT downstream block.
*/
function publishedAscentCoverage(input: {
coverage: AscentCoverage;
incompleteReasons: readonly PdgAscentIncompleteReason[];
callSummaryAvailable: boolean;
}): PdgAscentCoverage {
return {
referencesScanned: input.coverage.references,
returnFlowFound: input.coverage.anyReturnFlow,
undecodableSummaryCount: input.coverage.undecodable,
examinedComplete: input.incompleteReasons.length === 0,
incompleteReasons: input.incompleteReasons,
callSummaryLayerPresent: input.callSummaryAvailable,
};
}
/**
* Assemble the consumer-safe PDG impact result (U4 / KTD8 parity matrix).
*
@@ -795,6 +1071,8 @@ function assemblePdgImpactResult(input: {
* and steers to a re-index. `true` ⇒ ascent active (no extra note).
*/
callSummaryAvailable?: boolean;
/** Observed ascent inputs from the descent — rationale on {@link AscentCoverage}. */
ascentCoverage?: AscentCoverage;
}): PdgImpactSuccessResult {
const { target, direction, reachableBlocks, projection } = input;
const { symbols, unresolvedCount, ambiguousCount } = projection;
@@ -831,6 +1109,21 @@ function assemblePdgImpactResult(input: {
const byDepth: Record<number, unknown[]> = items.length > 0 ? { 1: items } : {};
const byDepthCounts: Record<number, number> = { 1: items.length };
// ── Ascent coverage: ONE computation, TWO surfaces ─────────────────────────
// The empty-ascent sentence quantifies UNIVERSALLY over the callees the descent
// actually EXAMINED, and three mechanisms can make that set a strict
// subset of the slice's real callee list (rationale: AscentCoverage
// INCOMPLETENESS, vocabulary: PdgAscentIncompleteReason). Classified ONCE by the
// shared `ascentIncompleteReasonsOf` and consumed twice — by `pdgEvidence.ascent`
// (structured, the contract) and by the note's qualifier clause (prose, rendered
// through ASCENT_INCOMPLETE_PHRASE). Deriving both from one array is what stops
// an agent branching on the codes and a human reading the note from ever
// disagreeing.
const ascentIncompleteReasons = ascentIncompleteReasonsOf({
truncated: input.truncated,
coverage: input.ascentCoverage,
});
const noteParts: string[] = statementMode
? [
`mode:'pdg' — intra-procedural slice from line ${input.criterionLine} of ` +
@@ -876,21 +1169,58 @@ function assemblePdgImpactResult(input: {
`CALL_SUMMARY edges and enable it.`,
);
} else if (input.callSummaryAvailable === true) {
// The CALL_SUMMARY layer is present, but return-value ascent is populated
// ONLY for TypeScript/JavaScript today (the formal-index it needs is set
// solely by the TS/JS harvester). For a criterion in any other language the
// ascent is structurally empty, so say so rather than letting the omission
// read as "ascent ran and found nothing". Sound — never claims ascent fired.
// Language is derived HERE in mcp/local, which may name languages; the
// shared core/ingestion pipeline must not.
const lang = getProviderForFile(target.filePath)?.id;
const ascentLanguage =
lang === SupportedLanguages.TypeScript || lang === SupportedLanguages.JavaScript;
if (!ascentLanguage) {
// The CALL_SUMMARY layer is present, but that only means the index CAN
// carry return-flow summaries — not that the callees in THIS slice have
// one. When none of them does, the ascent is structurally empty and the
// note says so, rather than letting the omission read as "ascent ran and
// found nothing". Sound — never claims the ascent fired. Keyed on the
// OBSERVED summaries, never on the criterion's language (#2802) — see
// {@link AscentCoverage} for why, and for the population the sentence below
// quantifies over: DISTINCT CALLEES (a `Set` of callee ids — two call sites
// to the same callee count once), NOT callees resolved to a body. Hence the
// "distinct callee(s)" wording; a call-SITE count would over-state the set.
const coverage = input.ascentCoverage;
const references = coverage?.references ?? 0;
const undecodable = coverage?.undecodable ?? 0;
// When the examined set is a strict prefix (the codes computed once above)
// the claim is qualified: the note may describe what was examined, never
// assert a property of the whole slice the traversal did not establish. The
// codes are mapped to phrases HERE — the note is a rendering of the same
// vocabulary `pdgEvidence.ascent.incompleteReasons` publishes.
const examinedIncomplete = ascentIncompleteReasons.length > 0;
const incompleteClause = examinedIncomplete
? ` (${ascentIncompleteReasons
.map((reason) => ASCENT_INCOMPLETE_PHRASE[reason])
.join(' and ')}, so callees past the examined set were not checked)`
: '';
if (references > 0 && coverage?.anyReturnFlow !== true) {
// ONE head for both arms — a shared gate and a shared opening sentence, so
// the two cannot drift on wording or pluralization. When at least one
// summary could not be DECODED the note must not assert what the persisted
// summaries record: an undecodable `reason` may well encode a return-flow
// this reader cannot unpack (`decodeCallSummary` never throws, so a
// version-skewed / corrupt / NULL reason is otherwise indistinguishable
// from a cleanly-decoded empty one). The ascent is withheld either way;
// only the tail that explains it changes.
noteParts.push(
`return-value ascent is currently TypeScript/JavaScript-only (only the TS/JS harvester ` +
`records the formal-index it needs), so a caller statement depending on a non-TS/JS ` +
`callee's RETURN value is not in the slice. Descent and the intra slice are unaffected.`,
`no return-value ascent in this slice: none of the ${references} distinct ` +
`${references === 1 ? 'callee carries' : 'callees carry'} a ` +
`${undecodable > 0 ? 'decodable ' : ''}CALL_SUMMARY return-flow${incompleteClause}` +
(undecodable > 0
? `, and ${undecodable} callee ` +
`${undecodable === 1 ? 'summary' : 'summaries'} could not be decoded (version ` +
`skew or corruption) — re-run gitnexus analyze --pdg to rebuild them. A caller ` +
`statement depending on a callee's RETURN value is not in the slice; descent and ` +
`the intra slice are unaffected.`
: `. So a caller statement depending on a callee's RETURN value is ` +
`not in the slice. ` +
(examinedIncomplete
? `Every summary examined decoded, so this is a property of those summaries, not `
: `Every summary in this slice decoded, so this is a property of the persisted ` +
`summaries, not `) +
`of the criterion's language — a callee whose producer records no formal index and ` +
`one with genuinely no return-flow are indistinguishable here. Descent and the ` +
`intra slice are unaffected.`),
);
}
}
@@ -930,6 +1260,28 @@ function assemblePdgImpactResult(input: {
localSymbolCount: impactedCount,
unresolvedBlockCount: unresolvedCount,
ambiguousProjectionCount: ambiguousCount,
// Structured ascent coverage — the note's facts, published so a caller never
// has to regex prose to learn whether the ascent was complete. Emitted iff
// the DESCENT RAN (`ascentCoverage` present), which is exactly the contract
// sentence on {@link PdgAscentCoverage}: absent ⇒ nothing was scanned because
// the descent never ran (an upstream slice), present ⇒ it ran and these are
// its counts.
//
// A zeroed-but-PRESENT record is therefore a real, honest reading — "the
// descent ran and the slice's blocks recorded no callee ids to scan" — not a
// placeholder. What used to make that reading unsafe was a block carrying
// call sites the index left id-less, which vanished from the population with
// no signal; that case now raises `'callee-ids-unrecorded'`, so a zero here
// with `examinedComplete: true` really does mean there was nothing to scan.
...(input.ascentCoverage
? {
ascent: publishedAscentCoverage({
coverage: input.ascentCoverage,
incompleteReasons: ascentIncompleteReasons,
callSummaryAvailable: input.callSummaryAvailable === true,
}),
}
: {}),
},
// Statement-level slice: the dependent source statements (line + text) the
// change reaches. This is the primary useful output of statement mode; the
@@ -1166,6 +1518,9 @@ export async function pdgLayerStatus(deps: {
let edgesVisible = false;
let probeError: string | undefined;
try {
// determinism: probe — layer existence. Only `rows.length > 0` is read; the
// projected `r.type` is never consumed, so which of the two edge types the
// one row happens to carry cannot change the returned state or note.
const rows = await deps.executeParameterized(
deps.lbugPath,
`MATCH (:BasicBlock)-[r:CodeRelation]->(:BasicBlock) WHERE r.type IN ['CDG', 'REACHING_DEF'] RETURN r.type AS type LIMIT 1`,
@@ -1239,6 +1594,16 @@ function blockAnchorForResolvedSymbol(sym: {
return { anchorClause: 'a.id STARTS WITH $idPrefix', queryParams: { idPrefix } };
}
/**
* The seed-block query both anchor builders feed — the top-level target seed and
* the per-callee span seeds. It was duplicated byte-for-byte at those two call
* sites, so the `ORDER BY a.startLine, id` tiebreak (#2787) had to be added
* twice; sharing it keeps the anchor and its query together, the way
* {@link blockAnchorForResolvedSymbol} already is.
*/
const seedBlockQuery = (anchorClause: string, probeLimit: number): string =>
`MATCH (a:BasicBlock) WHERE ${anchorClause} RETURN a.id AS id ORDER BY a.startLine, id LIMIT ${probeLimit}`;
/**
* Build a STATEMENT seed anchor: the BasicBlock(s) starting at a specific
* 1-based source `line` WITHIN the resolved symbol. This is what makes
@@ -1323,6 +1688,7 @@ async function bfsReachableBlocks(input: {
`MATCH (a:BasicBlock)-[r:CodeRelation]->(b:BasicBlock)
WHERE r.type IN ['CDG', 'REACHING_DEF'] AND ${matchEndpoint}.id IN $frontier
RETURN DISTINCT ${collectEndpoint}.id AS id
ORDER BY id
LIMIT ${probeLimit}`,
{ frontier },
);
@@ -1356,25 +1722,6 @@ interface CalleeSpan {
endLine: number;
}
/**
* Gather the resolved callee symbol ids invoked across a set of slice blocks
* (`BasicBlock.calleeIds`). Reuses the SHARED `splitCalleeIds` so the descent
* cannot diverge from `LocalBackend.calleeIdsOfBlocks` on the split/drop-sentinel
* logic. A pre-namespace-v4 index (no `calleeIds` column → empty cells) yields no
* ids, so the descent degrades cleanly to intra-only (no inter-procedural hop).
*/
async function calleeIdsFromBlocks(
lbugPath: string,
blockIds: string[],
exec: typeof executeParameterized,
): Promise<Set<string>> {
const ids = new Set<string>();
for (const { calleeIds } of await calleeIdsByBlock(lbugPath, blockIds, exec)) {
for (const id of calleeIds) ids.add(id);
}
return ids;
}
/** One slice block paired with the resolved callee ids it invokes. */
interface BlockCallees {
blockId: string;
@@ -1382,60 +1729,123 @@ interface BlockCallees {
}
/**
* Per-block variant of {@link calleeIdsFromBlocks}: keep the CALL block → its
* `calleeIds` mapping rather than flattening it. The return-value ascent (U-C4)
* needs this association — it re-seeds the caller's intra closure FROM the
* specific call block whose callee's `CALL_SUMMARY` licenses the ascent, so the
* flattened id-only set is insufficient. Reuses the SHARED `splitCalleeIds` so
* the split/drop-sentinel logic cannot diverge from the flattening caller. A
* block with no callee ids (empty/whitespace cell, or a pre-v4 index with no
* `calleeIds` column) yields an empty `calleeIds` — skipped by the consumer.
* Gather the resolved callee ids (`BasicBlock.calleeIds`) invoked across a set of
* slice blocks, keeping the CALL block → callees association rather than
* flattening it: the return-value ascent (U-C4) re-seeds the caller's intra
* closure FROM the specific call block whose callee's `CALL_SUMMARY` licenses the
* ascent, so a flat id set is insufficient. Reuses the SHARED
* {@link parseCalleeIdsCell} so the split/drop-sentinel logic cannot diverge from
* `LocalBackend.calleeIdsOfBlocks`. A block with no callee ids (empty/whitespace
* cell, or a pre-namespace-v4 index with no `calleeIds` column) yields an empty
* `calleeIds` — skipped by the consumer, so such an index degrades cleanly to
* intra-only (no inter-procedural hop).
*
* `calleeListTruncated` reports whether ANY of the queried blocks carried the
* emit-time cap sentinel. It is read from the RAW cell, so a block whose entire
* list was capped away (sentinel only ⇒ no ids ⇒ not emitted as a `BlockCallees`
* row) still raises it.
*
* `idlessCallSites` is the OTHER way a block's call sites leave the population
* unannounced: `calleeIdsOfBlock` emits an EMPTY `calleeIds` cell for a whole file
* whose resolved-id map is absent, and an empty cell carries no sentinel, so the
* cap flag cannot see it. The sibling `callees` (leaf NAMES) cell is read purely
* to tell that case apart from a block that genuinely calls nothing — names
* present + ids absent means the index recorded call sites it could not resolve.
* The name cell is never used for the descent itself (the resolved id is the sound
* key); it only keeps the coverage claim honest.
*/
async function calleeIdsByBlock(
lbugPath: string,
blockIds: string[],
exec: typeof executeParameterized,
): Promise<BlockCallees[]> {
if (blockIds.length === 0) return [];
): Promise<{ blocks: BlockCallees[]; calleeListTruncated: boolean; idlessCallSites: boolean }> {
if (blockIds.length === 0)
return { blocks: [], calleeListTruncated: false, idlessCallSites: false };
const rows = await exec(
lbugPath,
`MATCH (b:BasicBlock) WHERE b.id IN $ids RETURN b.id AS id, b.calleeIds AS calleeIds`,
`MATCH (b:BasicBlock) WHERE b.id IN $ids
RETURN b.id AS id, b.calleeIds AS calleeIds, b.callees AS callees`,
{ ids: blockIds },
);
const out: BlockCallees[] = [];
let calleeListTruncated = false;
let idlessCallSites = false;
// Narrow the awaited rows ONCE at the boundary to a typed record shape; read
// the aliased cells via bracket access — no per-field `as any`.
for (const r of rows as Array<Record<string, unknown>>) {
const blockId = String(r['id'] ?? '');
if (!blockId) continue;
const calleeIds = splitCalleeIds(r['calleeIds']);
// ONE pass over the cell classifies BOTH facts — a second full split just to
// re-test the sentinel doubled the per-row parse cost.
const { ids: calleeIds, truncated } = parseCalleeIdsCell(r['calleeIds']);
if (truncated) calleeListTruncated = true;
// Ids absent while NAMES are present ⇒ recorded call sites with no resolved
// id. Gated on `!truncated` so a capped-to-nothing cell keeps reporting the
// cap (the more specific mechanism) rather than both.
// `!idlessCallSites` first: the flag is sticky, so once it is set the string
// allocation below is pure waste on every remaining row of every later hop.
if (
!idlessCallSites &&
calleeIds.length === 0 &&
!truncated &&
String(r['callees'] ?? '').trim().length > 0
) {
idlessCallSites = true;
}
if (calleeIds.length > 0) out.push({ blockId, calleeIds });
}
return out;
return { blocks: out, calleeListTruncated, idlessCallSites };
}
/**
* Of a set of resolved callee symbol ids, which ones have a persisted
* `CALL_SUMMARY` self-loop edge recording a NON-EMPTY return-value ascent
* (≥1 formal parameter flows to the callee's return). This is the FU-C consumer
* side of the producer's per-callee summary (see `call-summary-codec.ts`).
* The THREE outcomes a persisted `CALL_SUMMARY` can have for one callee. The
* ascent itself only ever consults {@link returnFlowing}; {@link undecodable} is
* carried so the result `note` can tell "the summaries say there is no return-
* flow" apart from "we could not read the summaries" — two facts a single
* has-flow/has-no-flow boolean conflates (`decodeCallSummary` never throws, so a
* version-skewed / corrupt / NULL `reason` is otherwise indistinguishable from a
* cleanly-decoded EMPTY summary).
*/
interface CalleeReturnFlowScan {
/**
* Callees whose summary DECODED and records ≥1 formal parameter flowing to the
* return value — the only ones that license a return-value ascent.
*/
returnFlowing: Set<string>;
/**
* Callees that HAVE a `CALL_SUMMARY` row whose `reason` did not decode
* (unsupported version prefix, malformed segment, invalid hex payload, or a
* NULL/non-string reason). Never licenses an ascent — a decode failure means
* "no usable ascent fact", the codec's documented sound default.
*/
undecodable: Set<string>;
}
/**
* Scan the persisted `CALL_SUMMARY` self-loops of a set of resolved callee
* symbol ids. This is the FU-C consumer side of the producer's per-callee
* summary (see `call-summary-codec.ts`).
*
* The summary is a self-loop on the Function/Method/Constructor node:
* `(c)-[r:CodeRelation {type:'CALL_SUMMARY'}]->(c) WHERE c.id IN $ids`. The
* `reason` carries the param→return bitset; `decodeCallSummary` unpacks it and
* NEVER throws — a malformed / absent / empty (`r:0`) summary yields NO entry
* (the sound default: never claim a false return-flow). A PRE-FU-C (v3) `--pdg`
* index has NO `CALL_SUMMARY` edges, so this returns the empty set and the
* ascent is a clean no-op (the intra slice is unchanged — the documented
* "re-index for CALL_SUMMARY" degradation).
* NEVER throws. Three outcomes, per {@link CalleeReturnFlowScan}: a non-empty
* decoded return-flow, a cleanly-decoded EMPTY (`r:0`) summary, and an
* UNDECODABLE reason. Only the first licenses an ascent — the other two yield no
* ascent (the sound default: never claim a false return-flow) but are reported
* separately so the note never states a fact about summaries it could not read.
* A PRE-FU-C (v3) `--pdg` index has NO `CALL_SUMMARY` edges, so both sets come
* back empty and the ascent is a clean no-op (the intra slice is unchanged — the
* documented "re-index for CALL_SUMMARY" degradation).
*/
async function calleesWithReturnFlow(
lbugPath: string,
calleeIds: string[],
exec: typeof executeParameterized,
): Promise<Set<string>> {
const out = new Set<string>();
if (calleeIds.length === 0) return out;
): Promise<CalleeReturnFlowScan> {
const returnFlowing = new Set<string>();
const undecodable = new Set<string>();
if (calleeIds.length === 0) return { returnFlowing, undecodable };
const rows = await exec(
lbugPath,
`MATCH (c)-[r:CodeRelation]->(c)
@@ -1447,6 +1857,13 @@ async function calleesWithReturnFlow(
const id = String(r['id'] ?? '');
if (!id) continue;
const decoded = decodeCallSummary(r['reason']);
// A typed decode failure is NOT an empty summary — record it separately and
// withhold the ascent all the same (the codec's contract: a decode failure
// means "no usable ascent fact"). Only the note's wording depends on this.
if (!decoded.ok) {
undecodable.add(id);
continue;
}
// ARG→FORMAL trace precision: the conservative-but-sound default — ascend if
// ANY formal is return-flowing (the call site's argument is, by construction
// of the descent, in the slice: the call block is itself a slice block). A
@@ -1455,9 +1872,9 @@ async function calleesWithReturnFlow(
// per-arg list), so this never drops a real ascent; it may over-include
// (bounded — the result still flows to a slice statement). See the descent
// doc-comment + the result `note` caveat.
if (decoded.ok && decoded.returnFlowParams.length > 0) out.add(id);
if (decoded.returnFlowParams.length > 0) returnFlowing.add(id);
}
return out;
return { returnFlowing, undecodable };
}
/**
@@ -1574,6 +1991,12 @@ async function interproceduralDescent(input: {
* WHICH blocks got the ascent so the statement projection can expand them.
*/
ascentBlocks: Set<string>;
/**
* What the descent observed about return-value ascent across all hops, for the
* result note. Rationale — population, why it keys on observed data, and why
* the incompleteness flags matter — on {@link AscentCoverage}.
*/
ascentCoverage: AscentCoverage;
}> {
const {
lbugPath,
@@ -1599,13 +2022,37 @@ async function interproceduralDescent(input: {
// U-C4 return-value ascent: CALL blocks whose callee has a non-empty
// CALL_SUMMARY return-flow → the call's result depends on the slice.
const ascentBlocks = new Set<string>();
// Ascent-coverage accumulators (rationale: {@link AscentCoverage}). Sets so a
// callee invoked from two hops is tallied once; the `Seen` suffix marks them as
// accumulators whose `.size` — not the set — is what gets returned.
const calleeReferencesSeen = new Set<string>();
const calleesUndecodableSeen = new Set<string>();
// Sticky across hops. `anyReturnFlow` is the cross-hop union being non-empty,
// which holds iff SOME hop's return-flowing set was — so the flag is set inside
// the hop's existing non-empty branch rather than accumulating another Set.
let anyReturnFlow = false;
let calleeListTruncated = false;
let idlessCallSites = false;
hopLoop: for (let hop = 0; hop < depthBudget; hop++) {
if (sliceBlocks.length === 0) break;
// Blocks this hop newly reached — the NEXT hop's slice, and therefore the set
// whose `calleeIds` cells the next hop gathers. Declared BEFORE the U-C4
// ascent below so the ascent's own newly-reached blocks land in it: they are
// slice blocks (they are unioned into `reachable` and published in
// `reachableBlocks`), so their call sites must reach the CALL_SUMMARY scan and
// the coverage counters exactly like a descent-reached block's.
const hopReached = new Set<string>();
// Keep the CALL block → callee association (U-C4 needs it to re-seed the
// caller's intra closure FROM the specific call block the ascent licenses);
// the flattened id set still drives the descent's fresh-callee bookkeeping.
const blockCallees = await calleeIdsByBlock(lbugPath, sliceBlocks, exec);
const {
blocks: blockCallees,
calleeListTruncated: hopCellCapped,
idlessCallSites: hopIdless,
} = await calleeIdsByBlock(lbugPath, sliceBlocks, exec);
if (hopCellCapped) calleeListTruncated = true;
if (hopIdless) idlessCallSites = true;
const calleeIds = new Set<string>();
for (const { calleeIds: ids } of blockCallees) for (const id of ids) calleeIds.add(id);
@@ -1617,8 +2064,15 @@ async function interproceduralDescent(input: {
// that consumes the result is captured. Monotone: only ADDS to `reachable`,
// reusing the shared `visited` set, so it stays bounded + terminating. A
// pre-v4 index (no CALL_SUMMARY) yields no return-flowing callees → no-op.
const returnFlowing = await calleesWithReturnFlow(lbugPath, [...calleeIds], exec);
const summaryScan = await calleesWithReturnFlow(lbugPath, [...calleeIds], exec);
const returnFlowing = summaryScan.returnFlowing;
for (const id of calleeIds) calleeReferencesSeen.add(id);
// An undecodable summary withholds the ascent exactly like an empty one; it
// is tracked only so the note reports "could not read" rather than "records
// no return-flow".
for (const id of summaryScan.undecodable) calleesUndecodableSeen.add(id);
if (returnFlowing.size > 0) {
anyReturnFlow = true;
for (const { blockId, calleeIds: ids } of blockCallees) {
// Bound the ascent re-seeds the same way the descent bounds its per-span
// BFS (line ~1496): a wide fan-out of return-flowing call blocks must not
@@ -1647,8 +2101,24 @@ async function interproceduralDescent(input: {
stepLimit,
probeLimit,
});
// BOTH budgets, not just the row budget: the re-seed runs the SAME BFS
// under the SAME depth clamp as the top-level intra pass, whose depth
// exhaustion is result-level truncation.
//
// The depth fold here is a CONSISTENCY guard with no independent
// observable, and deliberately so — do not go hunting for the test that
// pins it. The re-seed shares the caller's `visited` set, so it can only
// discover new ground past the depth budget when the traversal that
// already covered this closure (the top-level intra BFS, or the callee's
// own BFS at a later hop) was ITSELF cut short — which has already raised
// one of these flags. Keeping it is what stops that reasoning from
// silently becoming load-bearing if the sharing of `visited` ever changes.
if (ascent.truncatedByLimit) truncatedByLimit = true;
for (const id of ascent.reachable) reachable.add(id);
if (ascent.truncatedByDepth) truncatedByDepth = true;
for (const id of ascent.reachable) {
reachable.add(id);
hopReached.add(id);
}
}
}
@@ -1672,7 +2142,7 @@ async function interproceduralDescent(input: {
const { anchorClause, queryParams } = blockAnchorForResolvedSymbol(span);
const rawSeedRows = await exec(
lbugPath,
`MATCH (a:BasicBlock) WHERE ${anchorClause} RETURN a.id AS id LIMIT ${probeLimit}`,
seedBlockQuery(anchorClause, probeLimit),
queryParams,
);
const exceeded = rawSeedRows.length > stepLimit;
@@ -1714,7 +2184,6 @@ async function interproceduralDescent(input: {
),
);
const hopReached = new Set<string>();
for (let si = 0; si < spans.length; si++) {
// Node budget is checked INSIDE the per-span MERGE (in span order) so the
// mid-hop short-circuit stays byte-identical: the cumulative reachable size
@@ -1740,6 +2209,13 @@ async function interproceduralDescent(input: {
const bfs = spanBfs[si];
if (bfs === null) continue;
if (bfs.truncatedByLimit) truncatedByLimit = true;
// A callee whose own dependence chain outruns `intraDepthBudget` is the SAME
// kind of incompleteness the top-level intra BFS reports through this flag
// (the budget is deliberately the same clamp — see `intraDepthBudget`), so
// it folds into the same result-level signal. Without this the slice could
// stop mid-callee while `truncated` stayed false and `examinedComplete`
// published a false all-clear over the callees past the frontier.
if (bfs.truncatedByDepth) truncatedByDepth = true;
// The per-callee BFS ran against a clone, so fold its discovered blocks
// into the shared `visited`/`reachable` here (the sequential path did this
// inside the BFS); Sets dedup, so order across siblings is irrelevant.
@@ -1756,9 +2232,12 @@ async function interproceduralDescent(input: {
}
sliceBlocks = [...hopReached];
}
// Frontier of callees still expandable after the hop budget ⇒ depth truncation.
// (Conservative: if the last hop reached blocks AND we used the full budget,
// deeper callees may exist.)
// Frontier of callees still expandable after the FUNCTION-hop budget ⇒ depth
// truncation. (Conservative: if the last hop reached blocks AND we used the full
// budget, deeper callees may exist.) This is the hop-level source; the per-callee
// and ascent BFS passes above fold their own block-hop depth exhaustion into the
// same flag, so `truncatedByDepth` means "some dependence frontier was cut by a
// depth budget", at either granularity.
if (hopsReached >= depthBudget && sliceBlocks.length > 0) truncatedByDepth = true;
return {
@@ -1768,6 +2247,13 @@ async function interproceduralDescent(input: {
truncatedByLimit,
truncatedByNodeCap,
ascentBlocks,
ascentCoverage: {
references: calleeReferencesSeen.size,
anyReturnFlow,
undecodable: calleesUndecodableSeen.size,
listTruncated: calleeListTruncated,
idlessCallSites,
},
};
}
@@ -1841,7 +2327,7 @@ export async function runImpactPDG(deps: RunPdgImpactDeps): Promise<PdgImpactRes
const probeLimit = stepLimit + 1;
const rawSeedRows = await exec(
repo.lbugPath,
`MATCH (a:BasicBlock) WHERE ${anchorClause} RETURN a.id AS id LIMIT ${probeLimit}`,
seedBlockQuery(anchorClause, probeLimit),
queryParams,
);
const seedRows = rawSeedRows.slice(0, stepLimit) as Array<Record<string, unknown>>;
@@ -1966,6 +2452,13 @@ export async function runImpactPDG(deps: RunPdgImpactDeps): Promise<PdgImpactRes
// call lines (a coalesced call block spans several statements whose results
// chain through it — the statement-granularity realisation of the ascent).
let ascentBlocks = new Set<string>();
// Observed ascent inputs, plumbed to the note AND to `pdgEvidence.ascent`
// (rationale: AscentCoverage). Left UNDEFINED when the descent never ran (an
// upstream slice): "nothing was scanned" is a different fact from "we scanned
// and found nothing", and a zeroed record would publish the second. The note's
// ascent branch is gated on `interproceduralHops > 0`, which only a descent can
// produce, so the prose is unaffected either way.
let ascentCoverage: AscentCoverage | undefined;
if (direction === 'downstream') {
const interproc = await interproceduralDescent({
lbugPath: repo.lbugPath,
@@ -1992,6 +2485,7 @@ export async function runImpactPDG(deps: RunPdgImpactDeps): Promise<PdgImpactRes
});
interproceduralHops = interproc.hopsReached;
ascentBlocks = interproc.ascentBlocks;
ascentCoverage = interproc.ascentCoverage;
for (const id of interproc.reachable) reachable.add(id);
if (interproc.truncatedByDepth) truncatedByDepth = true;
if (interproc.truncatedByLimit) truncatedByLimit = true;
@@ -2128,6 +2622,27 @@ export async function runImpactPDG(deps: RunPdgImpactDeps): Promise<PdgImpactRes
...(truncated ? { truncated: true } : {}),
...(truncatedBy ? { truncatedBy } : {}),
...(truncatedByReasons ? { truncatedByReasons } : {}),
// The descent may well have RUN and scanned callees before the slice turned
// out to reach no DISTINCT downstream block (a seed line whose only callee
// is invoked directly on it). `pdgEvidence.ascent` is contracted as "present
// iff the inter-procedural descent ran", so it must be published here too —
// omitting it made the documented "absent ⇒ nothing was scanned" reading
// false on exactly this exit. Classified through the SAME shared helpers the
// assembled slice result uses, so the two exits cannot disagree.
...(ascentCoverage
? {
pdgEvidence: {
ascent: publishedAscentCoverage({
coverage: ascentCoverage,
incompleteReasons: ascentIncompleteReasonsOf({
truncated,
coverage: ascentCoverage,
}),
callSummaryAvailable,
}),
},
}
: {}),
...emptyPdgParityFields(),
};
}
@@ -2159,6 +2674,7 @@ export async function runImpactPDG(deps: RunPdgImpactDeps): Promise<PdgImpactRes
truncatedByReasons,
interproceduralHops,
callSummaryAvailable,
ascentCoverage,
});
}
+11 -7
View File
@@ -280,7 +280,7 @@ Shows categorized incoming/outgoing references (calls, imports, extends, impleme
WHEN TO USE: After query() to understand a specific symbol in depth. When you need to know all callers, callees, and what execution flows a symbol participates in.
AFTER THIS: Use impact() if planning changes, or READ gitnexus://repo/{name}/process/{processName} for full execution trace.
Handles disambiguation: if multiple symbols share the same name, returns ranked candidates (each with a relevance score) for you to pick from. Use uid for zero-ambiguity lookup, or narrow the search with file_path and/or kind hints.
Handles disambiguation: if multiple symbols share the same name, returns ranked candidates (each with a relevance score) for you to pick from. Use uid for zero-ambiguity lookup, or narrow the search with file_path and/or kind hints. The ambiguous response carries totalCandidates — the TRUE match count, not candidates[].length — plus candidatesTruncated:true and a "(showing M)" suffix on message when candidates[] is the shorter window.
NOTE: ACCESSES edges (field read/write tracking) are included in context results with reason 'read' or 'write'. CALLS edges resolve through field access chains and method-call chains (e.g., user.address.getCity().save() produces CALLS edges at each step).
@@ -413,7 +413,9 @@ AFTER THIS: Run detect_changes() to verify no unexpected side effects.
Each edit is tagged with confidence:
- "graph": found via knowledge graph relationships (high confidence, safe to accept)
- "text_search": found via regex text search (lower confidence, review carefully)`,
- "text_search": found via regex text search (lower confidence, review carefully)
Handles disambiguation via context()'s payload verbatim: an ambiguous symbol_name returns status "ambiguous" with ranked candidates and totalCandidates — the TRUE match count, not candidates[].length — plus candidatesTruncated:true and a "(showing M)" suffix on message when candidates[] is the shorter window. Re-call with symbol_uid.`,
annotations: DESTRUCTIVE_TOOL_ANNOTATIONS,
inputSchema: {
type: 'object',
@@ -447,7 +449,7 @@ MODE (opt-in): "callgraph" (default) walks symbol→symbol edges (CALLS/IMPORTS/
STATEMENT-ANCHORED PDG SLICE: with mode:'pdg', pass "line" (1-based source line within the target symbol) to seed the dependence slice on the statement at that line and return what depends on it in affectedStatements (line + text). Inter-procedural symbols are still reported through interproceduralByDepth/pdgInterprocedural and the compatibility byDepth bucket. Without "line", pdg returns whole-symbol inter-procedural reach plus local whole-symbol PDG diagnostics.
PDG OUTPUT CONTRACT: every mode:'pdg' result (success, empty, degraded, or error) carries pdgResultVersion:2 — a stable discriminator for external consumers that bumps on any breaking change to the PDG result shape (distinct from the DB schema version). Successful PDG results include mode:'pdg', a full target envelope (id/name/type/filePath), affectedStatements, affectedStatementCount, interproceduralByDepth/pdgInterprocedural for cross-function reach, compatibility byDepth/byDepthCounts, risk:'UNKNOWN', and a note describing the unified contract. Degraded PDG results (no-layer, sub-layer-missing, unknown) keep mode:'pdg', pdgResultVersion:2, target metadata when the target resolves, risk:'UNKNOWN', note/remediation, and empty byDepth parity fields — never a false-safe zero. If depth and limit both bound the slice, truncatedByReasons reports both causes while truncatedBy remains scalar.
PDG OUTPUT CONTRACT: every mode:'pdg' result (success, empty, degraded, or error) carries pdgResultVersion:2 — a stable discriminator for external consumers that bumps on any breaking change to the PDG result shape (distinct from the DB schema version). Successful PDG results include mode:'pdg', a full target envelope (id/name/type/filePath), affectedStatements, affectedStatementCount, interproceduralByDepth/pdgInterprocedural for cross-function reach, compatibility byDepth/byDepthCounts, risk:'UNKNOWN', and a note describing the unified contract. Degraded PDG results (no-layer, sub-layer-missing, unknown) keep mode:'pdg', pdgResultVersion:2, target metadata when the target resolves, risk:'UNKNOWN', note/remediation, and empty byDepth parity fields — never a false-safe zero. If depth and limit both bound the slice, truncatedByReasons reports both causes while truncatedBy remains scalar. Return-value-ascent coverage is published structurally at pdgEvidence.ascent — present iff the inter-procedural descent ran, including on an empty slice — with referencesScanned (DISTINCT callees scanned for a CALL_SUMMARY: a distinct-id tally, not a call-site count — two call sites to the same callee count once), returnFlowFound (whether the ascent fired anywhere in the slice), undecodableSummaryCount, examinedComplete (whether that scan covered every callee the index recorded a resolved id for on the visited blocks), incompleteReasons ('traversal-truncated' | 'callee-list-capped' | 'callee-ids-unrecorded'), and callSummaryLayerPresent. Read callSummaryLayerPresent FIRST: false ⇒ a pre-CALL_SUMMARY index, so {referencesScanned:N>0, returnFlowFound:false} is self-consistent and says nothing about the callees — the scan ran, but no layer existed in which a return-flow could be recorded (remedy: re-run gitnexus analyze --pdg). Branch on those fields; the note narrates the same facts in prose for humans and is not a stable contract.
WHEN TO USE: Before making code changes — especially refactoring, renaming, or modifying shared code. Shows what would break.
AFTER THIS: Review d=1 items (WILL BREAK). Use context() on high-risk symbols.
@@ -476,12 +478,12 @@ TIP: For hub symbols (base error classes, shared utilities) with many direct cal
TIP: Default traversal uses CALLS/IMPORTS/EXTENDS/IMPLEMENTS. For class members, include HAS_METHOD and HAS_PROPERTY in relationTypes. For field access analysis, include ACCESSES in relationTypes.
Handles disambiguation: when multiple symbols share the target name, returns ranked candidates (each with a relevance score) instead of silently picking one. Use target_uid for zero-ambiguity lookup, or narrow with file_path and/or kind hints.
Handles disambiguation: when multiple symbols share the target name, returns ranked candidates (each with a relevance score) instead of silently picking one. Use target_uid for zero-ambiguity lookup, or narrow with file_path and/or kind hints. totalCandidates is the TRUE match count — it reported the capped resolver window before #2787, so it can now exceed candidates.length; candidatesTruncated:true and a "(showing M of N)" suffix on message mark the shorter window.
EdgeType: CALLS, IMPORTS, EXTENDS, IMPLEMENTS, HAS_METHOD, HAS_PROPERTY, METHOD_OVERRIDES, METHOD_IMPLEMENTS, ACCESSES
Confidence: 1.0 = certain, <0.8 = fuzzy match
GROUP MODE: set "repo" to "@<groupName>" for cross-repo impact anchored at the default member (lexicographically first key in group.yaml "repos"), or "@<groupName>/<groupRepoPath>" to choose the member (same path keys as in group.yaml). Phase-1 walk runs in that member; cross-boundary fan-out uses the group bridge. A cross entry with fanout_status:"not_attempted" proves the declared repository boundary, but its far endpoint has no graph symbol; do not interpret empty by_depth or affected_processes on that entry as a completed zero-impact walk.
GROUP MODE: set "repo" to "@<groupName>" for cross-repo impact anchored at the default member (lexicographically first key in group.yaml "repos"), or "@<groupName>/<groupRepoPath>" to choose the member (same path keys as in group.yaml). Phase-1 walk runs in that member; cross-boundary fan-out uses the group bridge. A cross entry with fanout_status:"not_attempted" proves the declared repository boundary, but its far endpoint has no graph symbol; do not interpret empty by_depth or affected_processes on that entry as a completed zero-impact walk. The fan-out attempts at most 50 neighbour crossings, strongest-confidence first; when it stops early the response carries truncated:true, truncatedRepos, and riskEpistemic:"lower-bound" — dropping a crossing can only move risk DOWN, so treat that risk as a floor, never as a verdict.
SERVICE: optional monorepo path prefix (case-sensitive path segments). When "repo" starts with "@", scopes the local impact walk and cross-repo symbol paths to files under that prefix; ignored for a normal indexed repo name.`,
annotations: READ_ONLY_TOOL_ANNOTATIONS,
@@ -632,7 +634,7 @@ Each finding carries the sink category (command-injection, code-injection, path-
WHEN TO USE: Security review — "what taint findings exist in this repo / file / function?". Requires the repo to be indexed with \`gitnexus analyze --pdg\`; without that layer the tool returns a clear "no taint layer" note, not an error.
ANCHORLESS (no "target"): enumerates all persisted findings for the repo — bounded ("limit", deterministic order), with "totalFindings" and a "truncated" flag.
ANCHORED ("target" = file path or symbol/function name): full hop detail for that anchor. A file-ish target (contains "/" or an extension) filters by file; a symbol name resolves like context() — ambiguous names return ranked candidates, unknown names return not-found. Symbol anchoring is line-range granular for intra-procedural findings; cross-function findings match when the symbol is the source OR sink function.
ANCHORED ("target" = file path or symbol/function name): full hop detail for that anchor. A file-ish target (contains "/" or an extension) filters by file; a symbol name resolves like context() — ambiguous names return ranked candidates plus totalCandidates (the TRUE match count, not candidates[].length), candidatesTruncated:true and a "(showing M)" suffix on message when candidates[] is the shorter window; unknown names return not-found. Symbol anchoring is line-range granular for intra-procedural findings; cross-function findings match when the symbol is the source OR sink function.
CONTRACT CAVEATS (absent flows are NOT proof of safety):
- Cross-function flows ARE modeled (#2084 M4): a source flowing through helper functions into a sink is found, via summary composition over the call graph (context-insensitive — return/call-site merging is accepted).
@@ -678,7 +680,7 @@ MODES:
WHEN TO USE: comprehension ("what guards this statement?"), data-flow tracing within a function, guard-clause discovery. Requires \`gitnexus analyze --pdg\`; without that layer the tool returns a clear "no PDG layer" note, not an error.
ANCHORING (required): \`target\` is a file path or a symbol/function name (resolved like context()). PDG queries are ALWAYS anchored — there is no whole-repo enumeration (an unanchored basic-block path scan is unbounded; LadybugDB has no rel-property index). A symbol target is line-range granular; an ambiguous name returns ranked candidates, unknown returns not-found.
ANCHORING (required): \`target\` is a file path or a symbol/function name (resolved like context()). PDG queries are ALWAYS anchored — there is no whole-repo enumeration (an unanchored basic-block path scan is unbounded; LadybugDB has no rel-property index). A symbol target is line-range granular; an ambiguous name returns ranked candidates plus totalCandidates (the TRUE match count, not candidates[].length), candidatesTruncated:true and a "(showing M)" suffix on message when candidates[] is the shorter window; unknown returns not-found.
CONTRACT CAVEATS:
- CDG labels are binary 'T'/'F' in M5/M6; per-case \`switch\` arm conditions are not yet distinguished (every case dispatch is 'T').
@@ -855,6 +857,8 @@ Traverses CALLS edges plus HAS_METHOD (class → member) edges, so a trace can d
Returns: ordered hops with file:line, and an aligned edges[] of edge type + confidence. When no path exists, reports the furthest reachable node so you know where the chain breaks (and truncated: true if a traversal cap was hit first).
Handles disambiguation: an ambiguous from/to name returns status "ambiguous" with role ("from" or "to"), ranked candidates and totalCandidates — the TRUE match count, not candidates[].length — plus candidatesTruncated:true and a "(showing M)" suffix on message when candidates[] is the shorter window. Re-call with from_uid/to_uid.
CROSS-REPO (experimental): pass repo as "@groupName" to trace across repositories in a group. When from/to live in different member repos, the trace stitches the two repo-local segments across a single ContractLink boundary (e.g. an HTTP consumer→provider link), clamped to one crossing. The result adds crossings[] (the bridged contract with matchType/confidence), tags each hop with its member repo, and a notes[] channel for degraded states. The boundary hop is reported with edge type CONTRACT_LINK. Pass pdg:true to also attach the intra-procedural data-flow (REACHING_DEF) for boundary-adjacent segments when those repos were indexed with --pdg; absent a PDG layer it degrades to call-level hops with a note.
DESTINATION TRACE (cross-repo): for an "@groupName" trace, OMIT to/to_uid/to_file to trace 'from' to wherever its outgoing HTTP call lands. The result ends at the provider endpoint (reported by route + file even when the handler is an anonymous function with no nameable symbol). This is the way to follow a client call to a backend handler you cannot name.`,
+52 -4
View File
@@ -19,14 +19,59 @@ export interface AnalyzeJobProgress {
message: string;
}
export type AnalyzeJobStatus =
| 'queued'
| 'cloning'
| 'analyzing'
| 'loading'
| 'complete'
| 'failed';
/**
* A job's outcome is settled — `complete` and `failed` are the only two states
* from which nothing more is emitted (see `updateJob`'s immutability guard).
*
* Exported because terminality is a property of the JOB, and every consumer
* that has to decide "is this stream over?" must ask the same question of the
* same field. The SSE relay used to decide from the PHASE STRING of a progress
* event instead, which let a mid-run `phase: 'complete'` close the stream
* before the route had decided the actual outcome (#2790).
*/
export const isTerminalJobStatus = (status: AnalyzeJobStatus): boolean =>
status === 'complete' || status === 'failed';
/**
* Structured detail for a job that ended `failed` while its work PARTIALLY
* succeeded — an embedding run that persisted most nodes and dropped a few to
* endpoint failures is not the same event as one that produced nothing.
*
* `AnalyzeJob.status` deliberately gains no `partial` member: the status union
* is consumed by the web app, the CLI and every poller, and a new member would
* be an unhandled value in each of them. This is additive and optional instead
* — absent on every job that is not a partial embedding run, so `JSON.stringify`
* omits it and existing payloads stay byte-identical — while giving a client
* that wants to distinguish "retry these N nodes" from "nothing worked" enough
* to do it (#2790 review). A UI that ignores it still sees the honest `failed`.
*/
export interface AnalyzeJobPartialOutcome {
/** Which kind of partial result this is; only embedding runs produce one today. */
kind: 'embedding-partial';
/** Nodes whose rows were dropped and are checkpointed for retry. */
pendingNodeCount: number;
/** Nodes that completed; their rows are durable. */
nodesProcessed: number;
}
export interface AnalyzeJob {
id: string;
status: 'queued' | 'cloning' | 'analyzing' | 'loading' | 'complete' | 'failed';
status: AnalyzeJobStatus;
repoUrl?: string;
repoPath?: string;
repoName?: string;
progress: AnalyzeJobProgress;
error?: string;
/** Set only when a terminal `failed` job still persisted usable work. */
partial?: AnalyzeJobPartialOutcome;
startedAt: number;
completedAt?: number;
/** Number of times the worker has been retried after a crash. */
@@ -96,7 +141,10 @@ export class JobManager {
updateJob(
id: string,
update: Partial<
Pick<AnalyzeJob, 'status' | 'progress' | 'error' | 'repoPath' | 'repoName' | 'completedAt'>
Pick<
AnalyzeJob,
'status' | 'progress' | 'error' | 'partial' | 'repoPath' | 'repoName' | 'completedAt'
>
>,
) {
const job = this.jobs.get(id);
@@ -116,7 +164,7 @@ export class JobManager {
}
// Emit exactly one event per updateJob call to prevent SSE double-write
if (update.status === 'complete' || update.status === 'failed') {
if (update.status !== undefined && isTerminalJobStatus(update.status)) {
// Terminal event takes precedence — don't also emit the progress event
this.emitter.emit(`progress:${id}`, {
phase: update.status,
@@ -209,7 +257,7 @@ export class JobManager {
}
private isTerminal(status: AnalyzeJob['status']): boolean {
return status === 'complete' || status === 'failed';
return isTerminalJobStatus(status);
}
private cleanup() {
+4 -4
View File
@@ -24,7 +24,7 @@ import {
} from '../storage/repo-manager.js';
import { logger } from '../core/logger.js';
import { autoHeapCapMb } from '../core/ingestion/utils/effective-ram.js';
import type { JobManager } from './analyze-job.js';
import { isTerminalJobStatus, type JobManager } from './analyze-job.js';
import type { WorkerMessage } from './analyze-worker.js';
const _require = createRequire(import.meta.url);
@@ -162,7 +162,7 @@ export function createLaunchAnalysisWorker(deps: LaunchDeps) {
const forkWorker = () => {
const currentJob = jobManager.getJob(job.id);
if (!currentJob || currentJob.status === 'complete' || currentJob.status === 'failed') return;
if (!currentJob || isTerminalJobStatus(currentJob.status)) return;
const child = fork(workerPath, [], {
execArgv: [...tsxHookArgs, `--max-old-space-size=${workerHeapMb}`],
@@ -182,7 +182,7 @@ export function createLaunchAnalysisWorker(deps: LaunchDeps) {
// re-release the repo lock or flip the reported status. Mirrors the `exit`
// handler guard below; pairs with the worker's terminal-claim (#2264 P3).
const current = jobManager.getJob(job.id);
if (!current || current.status === 'complete' || current.status === 'failed') return;
if (!current || isTerminalJobStatus(current.status)) return;
if (msg.type === 'progress') {
jobManager.updateJob(job.id, {
@@ -230,7 +230,7 @@ export function createLaunchAnalysisWorker(deps: LaunchDeps) {
child.on('exit', (code) => {
const j = jobManager.getJob(job.id);
if (!j || j.status === 'complete' || j.status === 'failed') return;
if (!j || isTerminalJobStatus(j.status)) return;
// Worker crashed — attempt retry if under the limit
if (j.retryCount < MAX_WORKER_RETRIES) {
+146 -127
View File
@@ -41,7 +41,19 @@ import { ftsDegradedWarning } from '../core/search/fts-indexes.js';
import { LocalBackend } from '../mcp/local/local-backend.js';
import { mountMCPEndpoints } from './mcp-http.js';
import { fileURLToPath } from 'url';
import { JobManager } from './analyze-job.js';
import { isTerminalJobStatus, JobManager, type AnalyzeJobPartialOutcome } from './analyze-job.js';
import { mountSSEProgress } from './sse-progress.js';
import {
resolveEmbedRunOutcome,
withMeasuredEmbeddingCount,
type EmbedRunFinalizeContext,
} from './embed-run-outcome.js';
import { decideEmbeddingResume, mintInterruptedCheckpoint } from '../core/embedding-checkpoint.js';
import {
measurePersistedEmbeddingCount,
persistedEmbeddingCountOrUndefined,
type PersistedEmbeddingCount,
} from '../core/embedding-count.js';
import { assertString, escapeRegExp, BadRequestError, createRouteLimiter } from './validation.js';
import {
extractRepoName,
@@ -464,98 +476,6 @@ export const streamGraphNdjson = async (
});
};
/**
* Mount an SSE progress endpoint for a JobManager.
* Handles: initial state, terminal events, heartbeat, event IDs, client disconnect.
*
* Terminal payloads carry `repoPath` (the analyzed path) alongside the display
* `repoName` so clients can reconnect by path identity — with duplicate
* basenames, a name-only reconnect resolves to the first same-named sibling.
* Exported for unit tests that lock the wire payload shape.
*/
export const mountSSEProgress = (app: express.Express, routePath: string, jm: JobManager) => {
app.get(routePath, (req, res) => {
let jobId: string;
try {
jobId = assertString(req.params.jobId, 'jobId');
} catch (err: any) {
res.status(err.status ?? 400).json({ error: err.message });
return;
}
const job = jm.getJob(jobId);
if (!job) {
res.status(404).json({ error: 'Job not found' });
return;
}
let eventId = 0;
res.writeHead(200, {
'Content-Type': 'text/event-stream',
'Cache-Control': 'no-cache',
Connection: 'keep-alive',
'X-Accel-Buffering': 'no',
});
// Send current state immediately
eventId++;
res.write(`id: ${eventId}\ndata: ${JSON.stringify(job.progress)}\n\n`);
// If already terminal, send event and close
if (job.status === 'complete' || job.status === 'failed') {
eventId++;
res.write(
`id: ${eventId}\nevent: ${job.status}\ndata: ${JSON.stringify({
repoName: job.repoName,
repoPath: job.repoPath,
error: job.error,
})}\n\n`,
);
res.end();
return;
}
// Heartbeat to detect zombie connections
const heartbeat = setInterval(() => {
try {
res.write(':heartbeat\n\n');
} catch {
clearInterval(heartbeat);
unsubscribe();
}
}, 30_000);
// Subscribe to progress updates
const unsubscribe = jm.onProgress(job.id, (progress) => {
try {
eventId++;
if (progress.phase === 'complete' || progress.phase === 'failed') {
const eventJob = jm.getJob(jobId);
res.write(
`id: ${eventId}\nevent: ${progress.phase}\ndata: ${JSON.stringify({
repoName: eventJob?.repoName,
repoPath: eventJob?.repoPath,
error: eventJob?.error,
})}\n\n`,
);
clearInterval(heartbeat);
res.end();
unsubscribe();
} else {
res.write(`id: ${eventId}\ndata: ${JSON.stringify(progress)}\n\n`);
}
} catch {
clearInterval(heartbeat);
unsubscribe();
}
});
req.on('close', () => {
clearInterval(heartbeat);
unsubscribe();
});
});
};
const statusFromError = (err: any): number => {
// Validation helpers throw BadRequestError / ForbiddenError with a typed
// .status field — honor it before falling back to message-string matching.
@@ -1317,6 +1237,9 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
// Label is validated against NODE_TABLES (compile-time safe identifiers);
// nodeId uses $nid parameter binding to prevent injection
const [connRes, clusterRes, procRes] = await Promise.all([
// determinism: probe — aggregate singleton. Both projections are
// `collect(...)` with no grouping key, which yields exactly one
// row, and `n` is PK-anchored on `$nid`; the LIMIT never chooses.
executePrepared(
`
MATCH (n:${nodeLabel} {id: $nid})
@@ -1334,6 +1257,7 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
MATCH (n:${nodeLabel} {id: $nid})
MATCH (n)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
RETURN c.label AS label, c.description AS description
ORDER BY c.id
LIMIT 1
`,
{ nid: nodeId },
@@ -1742,7 +1666,7 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
res.status(404).json({ error: 'Job not found' });
return;
}
if (job.status === 'complete' || job.status === 'failed') {
if (isTerminalJobStatus(job.status)) {
res.status(400).json({ error: `Job already ${job.status}` });
return;
}
@@ -1788,13 +1712,16 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
const EMBED_TIMEOUT_MS = 30 * 60 * 1000;
const embedTimeout = setTimeout(() => {
const current = embedJobManager.getJob(job.id);
if (current && current.status !== 'complete' && current.status !== 'failed') {
if (current && !isTerminalJobStatus(current.status)) {
embedJobManager.cancelJob(job.id, 'Embedding timed out (30 minute limit)');
}
}, EMBED_TIMEOUT_MS);
// Run embedding pipeline asynchronously
(async () => {
// Set inside withLbugDb, read after it closes (#2790).
let partialRunError: string | undefined;
let partialRunDetail: AnalyzeJobPartialOutcome | undefined;
try {
const lbugPath = path.join(entry.storagePath, 'lbug');
await withLbugDb(lbugPath, async () => {
@@ -1808,23 +1735,25 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
throw new Error('Repository metadata is missing; run gitnexus analyze first');
}
const priorCheckpoint = embeddingMeta.embeddingCheckpoint;
if (priorCheckpoint && priorCheckpoint.provider !== embeddingIdentity.provider) {
throw new Error(
'Cannot resume embedding checkpoint: the embedding provider configuration differs.',
);
// The SAME decision the CLI's resume gate makes
// (core/embedding-checkpoint.ts). This route used to hard-throw on
// any identity mismatch and ignore `attempts` entirely, so a
// `'partial'` marker written by `gitnexus analyze` and resumed
// here hit exactly the permanent wedge `kind` exists to remove:
// two readers of one record disagreeing about the rule it encodes.
// No `force`/`--drop-embeddings` equivalent exists on this route,
// so the flag options go unset and `'discard'` is unreachable —
// it is folded into the abandon arm rather than given an invented
// flag. `maxAttempts` is left to the shared default.
const resume = priorCheckpoint
? decideEmbeddingResume(priorCheckpoint, embeddingIdentity)
: undefined;
if (resume?.action === 'abort') throw new Error(resume.error);
if (resume?.action === 'abandon' || resume?.action === 'discard') {
logger.warn({ repo: entry.name }, resume.log);
}
if (
priorCheckpoint &&
(priorCheckpoint.model !== embeddingIdentity.model ||
priorCheckpoint.dimensions !== embeddingIdentity.dimensions)
) {
throw new Error(
`Cannot resume embedding checkpoint: it uses ${priorCheckpoint.model} at ` +
`${priorCheckpoint.dimensions} dimensions, but this run resolves ` +
`${embeddingIdentity.model} at ${embeddingIdentity.dimensions}.`,
);
}
const forceReembedNodeIds = new Set(priorCheckpoint?.pendingNodeIds ?? []);
const forceReembedNodeIds: ReadonlySet<string> =
resume?.action === 'resume' ? resume.pendingNodeIds : new Set<string>();
const saveEmbeddingCheckpoint = async (
checkpoint: {
nodesProcessed: number;
@@ -1832,6 +1761,7 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
chunksProcessed: number;
},
pendingNodeIds: string[],
embeddings?: PersistedEmbeddingCount,
): Promise<void> => {
// tri-review NEW-2: re-read immediately before writing (mirrors
// the pattern in run-analyze.ts's --repair-fts stamp) instead of
@@ -1841,17 +1771,42 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
// --repair-fts capability stamp) would be silently reverted on
// every checkpoint save for the job's whole lifetime.
const latestMeta = (await loadMeta(entry.storagePath)) ?? embeddingMeta;
embeddingMeta = {
...latestMeta,
embeddingCheckpoint: {
at: new Date().toISOString(),
...checkpoint,
...embeddingIdentity,
pendingNodeIds,
// `stats.embeddings` only moves when the caller MEASURED the
// live count (the post-flush `onCheckpoint`). The window-start
// callback measures nothing and passes nothing: restating the
// old count there would re-publish a stale number and clobber
// what a preceding `onCheckpoint` just wrote (same split as
// run-analyze.ts's checkpoint writer).
embeddingMeta = withMeasuredEmbeddingCount(
{
...latestMeta,
// In flight ⇒ `kind: 'interrupted'` (embedding-checkpoint.ts).
embeddingCheckpoint: mintInterruptedCheckpoint(
embeddingIdentity,
checkpoint,
pendingNodeIds,
),
},
};
embeddings,
);
await saveMeta(entry.storagePath, embeddingMeta);
};
/**
* Count the persisted rows, or report the answer never arrived.
* The TRI-STATE is carried to the fold rather than collapsed here:
* `unknown` is not 0, and only the fold knows what to carry
* forward instead (core/embedding-count.ts).
*/
const countPersistedEmbeddings = async (): Promise<PersistedEmbeddingCount> => {
const counted = await measurePersistedEmbeddingCount(executeQuery);
if (counted.kind === 'unknown') {
logger.warn(
{ reason: counted.reason },
'[embed] could not count persisted embeddings; leaving stats.embeddings untouched',
);
}
return counted;
};
// Fetch existing content hashes for incremental embedding.
// Delegated to lbug-adapter which owns the DB query logic and legacy-fallback handling.
const { fetchExistingEmbeddingHashes } = await import('../core/lbug/lbug-adapter.js');
@@ -1861,14 +1816,25 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
`[embed] ${existingEmbeddings.size} nodes already embedded — incremental run with content-hash comparison`,
);
}
await runEmbeddingPipeline(
const pipelineResult = await runEmbeddingPipeline(
executeQuery,
executeWithReusedStatement,
(p) => {
embedJobManager.updateJob(job.id, {
progress: {
// `ready` maps to 'finalizing', NOT 'complete' (#2790).
// The pipeline emits `ready`/100% unconditionally before
// returning — including when it dropped nodes to endpoint
// failures — and the route has not measured the index or
// decided the outcome yet, so 'complete' here would make
// the job record contradict itself (`status: 'analyzing'`,
// `progress.phase: 'complete'`).
phase:
p.phase === 'ready' ? 'complete' : p.phase === 'error' ? 'failed' : p.phase,
p.phase === 'ready'
? 'finalizing'
: p.phase === 'error'
? 'failed'
: p.phase,
percent: p.percent,
message:
p.phase === 'loading-model'
@@ -1878,7 +1844,7 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
: p.phase === 'indexing'
? 'Creating vector index...'
: p.phase === 'ready'
? 'Embeddings complete'
? 'Finalizing embeddings...'
: `${p.phase} (${p.percent}%)`,
},
});
@@ -1893,8 +1859,10 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
await saveEmbeddingCheckpoint(checkpoint, nodeIds);
},
onCheckpoint: async (checkpoint) => {
// Count AFTER the flush, so the number describes rows that
// are durable rather than rows still pending in the WAL.
await flushWAL();
await saveEmbeddingCheckpoint(checkpoint, []);
await saveEmbeddingCheckpoint(checkpoint, [], await countPersistedEmbeddings());
},
},
);
@@ -1904,16 +1872,64 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
// handles this during process exit, but the server keeps the
// connection open for other routes — a CHECKPOINT is enough.
await flushWAL();
// Same re-read-before-write reasoning as saveEmbeddingCheckpoint above.
// Measure inside withLbugDb, after the flush and while the
// connection is still open — this is the route's only chance to
// stamp `stats.embeddings` (embed-run-outcome.ts). A partial run
// gets the same stamp: an honest count of a partial index is what
// makes it survivable.
const measuredEmbeddings = await countPersistedEmbeddings();
// Same re-read-before-write reasoning as saveEmbeddingCheckpoint
// above — and the outcome decision reads it too: its
// `embeddingCheckpoint` is the marker this run's own mid-run
// writer saved, which is the only record of the work when the
// count query could not answer.
const finalMeta = (await loadMeta(entry.storagePath)) ?? embeddingMeta;
embeddingMeta = { ...finalMeta, embeddingCheckpoint: undefined };
const finalizeContext: EmbedRunFinalizeContext = {
measuredEmbeddings: persistedEmbeddingCountOrUndefined(measuredEmbeddings),
onDisk: finalMeta,
// The marker the job STARTED from — `finalMeta`'s has since been
// overwritten by the in-flight writer, so only this one carries
// the `'partial'` attempt chain.
resumedFrom: priorCheckpoint,
};
const outcome = resolveEmbedRunOutcome(
embeddingIdentity,
pipelineResult,
finalizeContext,
);
partialRunError = outcome.error;
partialRunDetail = outcome.partial;
embeddingMeta = withMeasuredEmbeddingCount(
{ ...finalMeta, embeddingCheckpoint: outcome.checkpoint },
measuredEmbeddings,
);
await saveMeta(entry.storagePath, embeddingMeta);
});
// Don't overwrite 'failed' if the job was cancelled while the pipeline was running
const current = embedJobManager.getJob(job.id);
if (!current || current.status !== 'failed') {
embedJobManager.updateJob(job.id, { status: 'complete' });
// The ONLY terminal event for the job, on both branches — nothing
// earlier maps to a terminal status, and `updateJob` synthesizes
// exactly one event per terminal status (#2264 P3). The explicit
// `progress` keeps the record self-consistent for a poller reading
// `progress.phase` (#2790).
embedJobManager.updateJob(
job.id,
partialRunError === undefined
? {
status: 'complete',
progress: { phase: 'complete', percent: 100, message: 'Embeddings complete' },
}
: {
status: 'failed',
error: partialRunError,
// Lets a client separate "retry these N nodes" from "the
// run produced nothing" without a new status member.
partial: partialRunDetail,
progress: { phase: 'failed', percent: 100, message: partialRunError },
},
);
}
} catch (err: any) {
const current = embedJobManager.getJob(job.id);
@@ -1953,6 +1969,9 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
repoName: job.repoName,
progress: job.progress,
error: job.error,
// Absent unless the run was a partial one — omitted by JSON.stringify, so
// the response shape is unchanged for every other outcome (#2790).
partial: job.partial,
startedAt: job.startedAt,
completedAt: job.completedAt,
});
@@ -1969,7 +1988,7 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
res.status(404).json({ error: 'Job not found' });
return;
}
if (job.status === 'complete' || job.status === 'failed') {
if (isTerminalJobStatus(job.status)) {
res.status(400).json({ error: `Job already ${job.status}` });
return;
}
+143
View File
@@ -0,0 +1,143 @@
/**
* What POST /api/embed persists and reports once the embedding pipeline returns.
*
* Extracted from api.ts (#2790 review, finding 9): these are pure functions, and
* reaching them through `await import('../../src/server/api.js')` pulled in
* Express, cors, the LadybugDB native adapter and the whole MCP wiring — one
* measured run of the test file TIMED OUT at 30s and the two that passed took
* ~20s and ~22s. Nothing here imports the database, MCP or Express.
*
* This module ROUTES that decision onto the route's meta write and job status;
* it does not author the checkpoint record. Minting and the resume rules live in
* `core/embedding-checkpoint.ts`, shared with the CLI's Phase 5, so the two
* writers of one field cannot drift apart again.
*/
import type { RepoMeta } from '../storage/repo-manager.js';
import { mintPartialCheckpoint, type EmbeddingRunIdentity } from '../core/embedding-checkpoint.js';
import {
resolvePersistedEmbeddingCount,
type PersistedEmbeddingCount,
} from '../core/embedding-count.js';
import type { AnalyzeJobPartialOutcome } from './analyze-job.js';
export type { EmbeddingRunIdentity };
/** The subset of `EmbeddingPipelineResult` the outcome decision reads. */
export interface EmbeddingRunResult {
nodesProcessed: number;
chunksProcessed: number;
failedNodeIds: string[];
}
/** Everything the decision needs beyond the run's own result. */
export interface EmbedRunFinalizeContext {
/**
* Embedding rows counted after the final flush; `undefined` ≡ COULD NOT ASK
* (see `core/embedding-count.ts`). Never a fabricated 0.
*/
measuredEmbeddings?: number;
/**
* The meta currently on disk, re-read immediately before the finalize write.
* Its `embeddingCheckpoint` is the marker the run's own mid-run checkpoint
* writer last saved — the recovery evidence the clean-run branch may keep.
*/
onDisk?: Pick<RepoMeta, 'stats' | 'embeddingCheckpoint'>;
/**
* The marker this run RESUMED from, read at job start — not the same object
* as `onDisk.embeddingCheckpoint`, which the mid-run writer has since
* overwritten with an `'interrupted'` marker. Carries the attempt chain.
*/
resumedFrom?: RepoMeta['embeddingCheckpoint'];
}
export interface EmbedRunOutcome {
/** What to write as `embeddingCheckpoint`; `undefined` ≡ clear it. */
checkpoint: RepoMeta['embeddingCheckpoint'];
/** Set ≡ the job must be reported `failed`. */
error?: string;
/** Set ≡ the failure is a PARTIAL one; relayed on the job and the SSE payload. */
partial?: AnalyzeJobPartialOutcome;
}
/**
* Decide the checkpoint + reported outcome for a finished /api/embed run.
*
* PARTIAL RUN: the pipeline no longer throws when a sub-batch loses its
* endpoint — it deletes the affected nodes' rows and names them in
* `failedNodeIds`. Dropping that receipt made the route clear
* `embeddingCheckpoint` and mark the job 'complete', and the dropped nodes were
* then never retried: a plain `analyze` derives `shouldGenerateEmbeddings:
* false` once any embeddings exist, so nothing would ever call the pipeline
* again. So the checkpoint is RETAINED, carrying the dropped ids as
* `pendingNodeIds` (the next run's `forceReembedNodeIds`), and the run is
* reported failed — the `AnalyzeJob` status union has no partial member and this
* is what the pre-#2790 throw produced, so a poller keeps seeing "not a clean
* success". `partial` carries the distinction a client needs without a new
* status member.
*
* CLEAN RUN: clear the checkpoint — with one exception, below.
*/
export const resolveEmbedRunOutcome = (
identity: EmbeddingRunIdentity,
result: EmbeddingRunResult,
context: EmbedRunFinalizeContext = {},
): EmbedRunOutcome => {
if (result.failedNodeIds.length === 0) {
// ── Clearing the marker requires PROOF the index is accounted for ──────
// When the count query did not answer, the route cannot stamp
// `stats.embeddings` (a fabricated value is worse than none — see
// `core/embedding-count.ts`), so a repo whose recorded count is still 0
// would end this run with NOTHING on disk saying embeddings exist. Keeping
// the mid-run marker leaves one durable record. It costs a re-embed of its
// pending set on the next run and clears itself as soon as any run can
// measure the table.
const countIsKnown = context.measuredEmbeddings !== undefined;
const metaAlreadyRecordsEmbeddings = (context.onDisk?.stats?.embeddings ?? 0) > 0;
if (countIsKnown || metaAlreadyRecordsEmbeddings) return { checkpoint: undefined };
return { checkpoint: context.onDisk?.embeddingCheckpoint };
}
return {
// `'partial'` + the attempt-chain rule are the minter's (#2790).
checkpoint: mintPartialCheckpoint(identity, result, context.resumedFrom),
error:
`Embedding generation finished partially: ${result.failedNodeIds.length} node(s) lost ` +
'their embeddings to endpoint failures and were dropped. They are checkpointed as ' +
'pending — run embedding generation again to retry exactly those nodes.',
partial: {
kind: 'embedding-partial',
pendingNodeCount: result.failedNodeIds.length,
nodesProcessed: result.nodesProcessed,
},
};
};
/** A write that did not even ask — folded exactly like a failed measurement. */
const UNMEASURED: PersistedEmbeddingCount = {
kind: 'unknown',
reason: 'this write measured nothing',
};
/**
* Fold a measurement into the meta /api/embed is about to save.
*
* The /api/embed count omission: this route generated embeddings and wrote
* `embeddingCheckpoint`, but never `stats.embeddings` — so a repo embedded
* purely through the server kept whatever count the last CLI `analyze` stamped
* (0 for a repo analyzed without embeddings), and the next `analyze --force`
* wiped every server-generated embedding.
*
* The carry-forward itself is `resolvePersistedEmbeddingCount`'s, not this
* module's: the CLI publishes the same field, and "unknown ⇒ carry forward,
* never fabricate 0" only holds if both publishers ask the same function
* (core/embedding-count.ts).
*/
export const withMeasuredEmbeddingCount = (
meta: RepoMeta,
/** `undefined` ≡ this write measured nothing (e.g. the window-start save). */
measured: PersistedEmbeddingCount | undefined,
): RepoMeta => {
const embeddings = resolvePersistedEmbeddingCount(measured ?? UNMEASURED, meta.stats?.embeddings);
return embeddings === undefined ? meta : { ...meta, stats: { ...meta.stats, embeddings } };
};
+137
View File
@@ -0,0 +1,137 @@
/**
* SSE progress relay for JobManager-backed jobs (`/api/analyze` and `/api/embed`).
*
* Extracted from api.ts so the relay can be driven by a test without booting
* Express + the LadybugDB native adapter + the whole MCP wiring: the bug this
* module exists to prevent is only observable END TO END (a client watching the
* stream), and a contract that can only be checked by regex over api.ts's source
* is not a contract. Nothing here imports the database or MCP; `express` is a
* TYPE-only import, so the runtime cost of loading this module is validation.ts
* plus analyze-job.ts. Import it from HERE — routing through api.ts re-imports
* everything the extraction was meant to avoid.
*/
import type express from 'express';
import { assertString, BadRequestError } from './validation.js';
import { isTerminalJobStatus, type AnalyzeJob, type JobManager } from './analyze-job.js';
/**
* The wire payload of a terminal (`event: complete` / `event: failed`) frame.
*
* Read from the JOB at emit time rather than from the progress event, so it
* always carries whatever the terminal `updateJob` call just committed.
* `undefined` fields are dropped by `JSON.stringify`, which keeps a clean run's
* payload byte-identical to the pre-`partial` shape.
*/
const terminalPayload = (job: AnalyzeJob | undefined) => ({
repoName: job?.repoName,
repoPath: job?.repoPath,
error: job?.error,
// Lets a client tell a partial embedding run ("retry these N nodes") from a
// total failure ("nothing worked") without a new `status` member (#2790).
partial: job?.partial,
});
/**
* Mount an SSE progress endpoint for a JobManager.
* Handles: initial state, terminal events, heartbeat, event IDs, client disconnect.
*
* Terminal payloads carry `repoPath` (the analyzed path) alongside the display
* `repoName` so clients can reconnect by path identity — with duplicate
* basenames, a name-only reconnect resolves to the first same-named sibling.
* Exported for unit tests that lock the wire payload shape.
*
* ── The stream closes on the JOB'S STATUS, never on a phase string (#2790) ──
* This used to end the response as soon as a progress event carried
* `phase: 'complete'` or `'failed'`. /api/embed maps the pipeline's `ready`
* phase, which `runEmbeddingPipeline` emits UNCONDITIONALLY — including when it
* dropped nodes to endpoint failures — and the route only decides the real
* outcome after the pipeline returns. So the relay wrote `event: complete` with
* `error: undefined`, called `res.end()` and unsubscribed; the route's later
* `updateJob({status:'failed'})` was emitted into a stream nobody was listening
* to. An SSE client saw a tolerated partial run as a clean success while a
* `GET /api/embed/:jobId` poller saw `failed` — the two consumers of one job
* disagreeing about whether the data is complete.
*
* `updateJob` already synthesizes exactly one terminal progress event (and
* refuses every update once the job is terminal, #2264 P3), and it assigns the
* status BEFORE emitting — so asking the job is both sufficient and the only
* source that cannot be spoofed by an intermediate phase label.
*/
export const mountSSEProgress = (app: express.Express, routePath: string, jm: JobManager) => {
app.get(routePath, (req, res) => {
let jobId: string;
try {
jobId = assertString(req.params.jobId, 'jobId');
} catch (err) {
const status = err instanceof BadRequestError ? err.status : 400;
res.status(status).json({ error: err instanceof Error ? err.message : 'Invalid jobId' });
return;
}
const job = jm.getJob(jobId);
if (!job) {
res.status(404).json({ error: 'Job not found' });
return;
}
let eventId = 0;
res.writeHead(200, {
'Content-Type': 'text/event-stream',
'Cache-Control': 'no-cache',
Connection: 'keep-alive',
'X-Accel-Buffering': 'no',
});
// Send current state immediately
eventId++;
res.write(`id: ${eventId}\ndata: ${JSON.stringify(job.progress)}\n\n`);
// If already terminal, send event and close
if (isTerminalJobStatus(job.status)) {
eventId++;
res.write(
`id: ${eventId}\nevent: ${job.status}\ndata: ${JSON.stringify(terminalPayload(job))}\n\n`,
);
res.end();
return;
}
// Heartbeat to detect zombie connections
const heartbeat = setInterval(() => {
try {
res.write(':heartbeat\n\n');
} catch {
clearInterval(heartbeat);
unsubscribe();
}
}, 30_000);
// Subscribe to progress updates
const unsubscribe = jm.onProgress(job.id, (progress) => {
try {
eventId++;
const eventJob = jm.getJob(jobId);
if (eventJob !== undefined && isTerminalJobStatus(eventJob.status)) {
res.write(
`id: ${eventId}\nevent: ${eventJob.status}\ndata: ${JSON.stringify(
terminalPayload(eventJob),
)}\n\n`,
);
clearInterval(heartbeat);
res.end();
unsubscribe();
} else {
res.write(`id: ${eventId}\ndata: ${JSON.stringify(progress)}\n\n`);
}
} catch {
clearInterval(heartbeat);
unsubscribe();
}
});
req.on('close', () => {
clearInterval(heartbeat);
unsubscribe();
});
});
};
+3 -1
View File
@@ -128,7 +128,9 @@ import type { ParseWorkerResult } from '../core/ingestion/workers/parse-worker.j
// both genuinely shipped as 21 — renumbering would misstate history. Read this
// as the reason to re-check SCHEMA_BUMP against origin/main immediately before
// merging, not just when the branch is cut; the same collision hit
// INCREMENTAL_SCHEMA_VERSION in #2653/#2654.
// the DB schema version in #2653/#2654 (that constant is gone — the DB side is
// a derived fingerprint now, see SCHEMA_FINGERPRINT; SCHEMA_BUMP below is still
// hand-maintained because no declarative artifact describes a capture set).
// v27: generator EXPRESSIONS bound to a name emit a callable definition capture,
// and nested-callable caller attribution appends the localIdentity suffix the
// definition phase already used. Both are parse-time, so a warm cache would
+98 -284
View File
@@ -195,8 +195,9 @@ export interface RepoMeta {
* that `--repair-fts` changed FTS availability (`doctor` still prints
* platform-derived capabilities separately; `graph`/`vectorSearch` remain
* forensic-only). The status unions mirror `CapabilityStatus` /
* `SemanticSearchMode` in core/platform/capabilities.ts; inlined to keep
* storage/ free of a core/ type dependency.
* `SemanticSearchMode` in core/platform/capabilities.ts; inlined so storage/
* takes no core/ import for a pair of string unions, at the cost of keeping
* the two in sync by hand.
*/
capabilities?: {
graph: { provider: string; status: 'available' | 'degraded' | 'unavailable' };
@@ -209,15 +210,33 @@ export interface RepoMeta {
};
};
/**
* Bumped whenever incremental-indexing invariants change in an
* incompatible way (delete-and-rewrite logic, subgraph extraction,
* graph-wide node handling). On mismatch, runFullAnalysis forces a
* full rebuild rather than risk an inconsistent incremental update.
* Digest of the graph DDL this index's tables were actually created from
* (`SCHEMA_FINGERPRINT`, core/lbug/schema.ts). On mismatch, runFullAnalysis
* warns and forces a full rebuild, which wipes and recreates the database so
* the tables are built from the current DDL (#2798).
*
* This REPLACED `schemaVersion`, a hand-incremented integer that had to
* predict the same fact and could not: it collided with `main` eight times,
* twice exactly, and an exact clash passed the `===` gate silently. The
* digest is derived, so it cannot collide by accident at this scale (48
* bits; see SCHEMA_FINGERPRINT) — two builds agree exactly when their DDL
* agrees.
*
* ABSENT ≡ mismatch, deliberately. That is the backward-compatibility path:
* every index built by an older GitNexus carries no fingerprint, gets the
* warning, and is rebuilt once against the current schema. Grandfathering
* absence would instead stamp a fresh fingerprint onto a database whose DDL
* was never verified.
*
* Stamped only for git repos — non-git repos never take the incremental path.
* Declared as a plain string rather than importing the constant: that would
* be a RUNTIME value import of core/lbug/schema.ts, pulling the whole DDL and
* its `gitnexus-shared` module graph into every storage/ consumer.
*/
schemaVersion?: number;
schemaFingerprint?: string;
/**
* Exact versions of independently-gated analysis capabilities produced by
* the successful run. Unlike schemaVersion, these may apply only to repos
* the successful run. Unlike schemaFingerprint, these may apply only to repos
* containing relevant source files.
*/
analysisFeatures?: Record<string, number>;
@@ -231,6 +250,34 @@ export interface RepoMeta {
* compare, not an absence.
*/
cjkSegmentation?: string;
/**
* The `FLOAT[N]` width this index's `CodeEmbedding` vector column was
* actually created at — `EMBEDDING_DIMS` (core/lbug/schema.ts), resolved from
* `GITNEXUS_EMBEDDING_DIMS` at module load (#2798). On mismatch with the live
* process's width, runFullAnalysis forces a full rebuild, which wipes the
* database and recreates the table at the new width; an incremental run never
* revisits a column's type, so nothing else can.
*
* Sits beside `schemaFingerprint` rather than inside it on purpose: the
* fingerprint is a digest of CODE, and this width comes from the
* ENVIRONMENT, so folding it in would make the same build disagree with
* itself across two runs and thrash rebuilds.
*
* ABSENT means an index written before this field existed — NOT a mismatch,
* unlike `schemaFingerprint` above. Absence says nothing about the width
* (that run used whatever its env resolved, almost always the 384 default,
* and the table it wrote agreed with it), and every such index also predates
* `schemaFingerprint`, so the guard above already rebuilds it once and this
* stamp lands then. See `embeddingDimsMismatch` for the full argument.
*
* Always stamped, like `cjkSegmentation` and unlike `schemaFingerprint`: the
* column is created for every index, git or not, so there is no case where
* omitting it is correct — which keeps absence meaning exactly one thing.
* A plain number rather than an import of the constant, for the same reason
* `schemaFingerprint` is a plain string: storage/ takes no runtime import of
* core/lbug/schema.ts.
*/
embeddingDims?: number;
/**
* Member names whose call sites were DROPPED because the receiver's type
* could not be established (#2744, the second half of #2708). Read by
@@ -293,11 +340,15 @@ export interface RepoMeta {
droppedImporterChunks?: number;
};
/**
* Durable embedding-resume marker. Before a bounded write window begins,
* `pendingNodeIds` records every node that could become partially persisted;
* after the LadybugDB checkpoint it is cleared while progress is retained.
* A matching runtime resumes from persisted hashes and regenerates pending
* nodes; a model or dimension mismatch fails before mutation.
* Durable embedding-resume marker, written in two distinct situations that
* `kind` tells apart — see below. A matching runtime resumes from persisted
* hashes and regenerates the pending nodes.
*
* Cleared by a clean run. NOT cleared by a run that completed while dropping
* nodes to endpoint failures (#2790): retaining it is what makes those nodes
* come back, because a plain `analyze` derives `shouldGenerateEmbeddings:
* false` once any embeddings exist, so nothing would ever call the pipeline
* again.
*/
embeddingCheckpoint?: {
at: string;
@@ -309,10 +360,37 @@ export interface RepoMeta {
/** `local` or a secret-free SHA-256 fingerprint of the HTTP endpoint identity. */
provider: string;
/**
* Nodes in the current checkpoint window. Any of these may have only a
* subset of their chunks persisted after an abrupt process termination,
* so resume must delete and regenerate them even when a persisted row has
* the current content hash.
* Which situation wrote this marker. Absent ≡ `'interrupted'`, so markers
* written by older versions keep the stricter behavior.
*
* - `'interrupted'` — written BEFORE a bounded write window. Its
* `pendingNodeIds` may be half-persisted if the process died mid-window,
* so resume must delete and regenerate them even when a persisted row
* carries the current content hash, and an identity mismatch must fail
* closed: resuming under a foreign model would mix vector spaces.
* - `'partial'` — written AFTER a run that completed but dropped nodes to
* endpoint failures. The pipeline already deleted every row of those
* nodes, so they provably hold ZERO rows. Nothing is at risk from a
* different embedding identity, so an identity mismatch may drop the
* pending set with a warning instead of aborting the run.
* - `'unverified-count'` — written after a run whose embedding count could
* not be measured. `pendingNodeIds` is EMPTY: nothing was dropped and
* nothing needs re-embedding. It exists only to defeat the same-commit
* fast return so the next run re-derives a count, because clearing it
* while `stats.embeddings` still reads a stale zero is what arms a later
* `--force` to wipe live embeddings.
*/
kind?: 'interrupted' | 'partial' | 'unverified-count';
/**
* Consecutive resume attempts that have failed to clear `pendingNodeIds`
* (`'partial'` only). Bounds the retry so a node the endpoint rejects
* deterministically — an oversized chunk, content it refuses — cannot keep
* a repo permanently incomplete. See EMBEDDING_RESUME_MAX_ATTEMPTS.
*/
attempts?: number;
/**
* Nodes to regenerate on resume. For `'interrupted'` these may hold a
* subset of their chunks; for `'partial'` they hold none.
*/
pendingNodeIds?: string[];
};
@@ -342,9 +420,9 @@ export interface RepoMeta {
* compares this against the requested options and forces a full
* writeback on any mismatch — the incremental path only persists
* changed-file nodes and would otherwise silently drop (or strand) the
* CFG layer on a mode flip. Additive/optional, no
* INCREMENTAL_SCHEMA_VERSION bump (a bump would force a one-time full
* rebuild for every user). NOTE the removal mechanism is load-bearing:
* CFG layer on a mode flip. Additive/optional: it is metadata, not DDL, so
* it does not move `schemaFingerprint` and costs no rebuild for anyone whose
* pdg mode is unchanged. NOTE the removal mechanism is load-bearing:
* the end-of-run meta is a fresh object literal, NOT a spread of the
* prior meta, so omitting this field on a pdg-off run is what clears
* the stamp after an on→off flip.
@@ -426,270 +504,6 @@ export interface RepoMeta {
};
}
/**
* Bumped whenever incremental-indexing invariants change incompatibly.
* v2: `BasicBlock.callees` column added (statement-precise inter-procedural
* reach substrate) — an index built before this lacks the column, so a full
* re-analyze is required rather than an incremental top-up.
* v3: `BasicBlock.calleeIds` column added (sound resolved-callee-id parallel
* to `callees`, #2227) — same contract: an index built before this lacks the
* column, so a full re-analyze is forced rather than an incremental top-up.
* v4: `CALL_SUMMARY` relation type added (per-callee RETURN-VALUE ASCENT
* summary edges, PDG FU-C). A pre-v4 `--pdg` index has NO CALL_SUMMARY edges,
* so the engine would silently UNDER-REPORT return-value ascent on an
* incremental top-up; force a full re-analyze instead (same contract as v2/v3).
* This single bump covers the whole FU-C re-index window (and the later FU-B-2).
* v5: `Route` node identity changed to `(method, url)` (#2289 — a same-URL
* GET/POST pair is now two distinct Route nodes). Every declarative-route node
* id moved from `Route:/x` to `Route:GET /x` (filesystem routes keep their
* URL-only id). The incremental writeback preserves unchanged-file rows, so a
* top-up against a pre-v5 index would strand old url-keyed Route nodes alongside
* new composite-keyed ones — force a full re-analyze instead.
* v6: line-number storage flipped to uniform 0-based for the last 1-based
* GraphNode emitters — COBOL/JCL/markdown/scope (#2377/#2379/#2380). Incremental
* writeback preserves unchanged-file rows, so a top-up against a pre-v6 index
* would MIX old 1-based rows with new 0-based ones — and the 1-based MCP display
* would render the stale rows one line too high — so force a full re-analyze.
* v7: callable-value-flow CALLS/USES edges added (#2437/#2522) — new edges can
* connect two files whose content did not change, but the incremental write set
* only covers changed files (`computeEffectiveWriteSet`), so a top-up against a
* pre-v7 index would silently omit the new edges for every unchanged file pair;
* force a full re-analyze instead (same contract as v2–v6).
* v8: Java anonymous class bodies became first-class Class nodes (#2550):
* `new Runnable() { run(){} }` now emits `Class:...:Worker$1` and its methods
* re-keyed from `Worker.run` to `Worker$1.run`. Node identities move on
* unchanged files — a top-up against a pre-v8 index would strand the old
* `Worker.run`-keyed Method nodes alongside the new ones (the v5 Route
* precedent); force a full re-analyze instead.
* v9: Java enum constant bodies joined the instance model and anonymous
* naming switched to JLS 13.1 immediately-enclosing-type chains (#2555): `enum E { A {
* hook(){} } }` now emits `Class:...:E$1` with methods re-keyed from
* `E.hook` to `E$1.hook`, and nested-host anonymous names re-key
* (`EnumWrap$1` → `EnumWrap$Mode$1`). Same contract as v8: identities move
* on unchanged files; force a full re-analyze.
* v10: Java `record_declaration` now emits a first-class `Record` graph node
* (#2564): a record's container node was previously never created (JAVA_QUERIES
* had no capture for it), so its methods existed as ownerless Method nodes
* with no `HAS_METHOD` edge. The incremental write set only covers changed
* files — a top-up against a pre-v10 index would keep silently omitting the
* `Record` node and its `HAS_METHOD` edges for every unchanged record file
* (same v7 contract: new nodes/edges the incremental path would otherwise
* never backfill); force a full re-analyze instead.
* v11: Rust abstract trait methods (`fn foo(&self) -> T;`, no body) now get a
* scope + declaration capture (#2604): RUST_SCOPE_QUERY had no
* `function_signature_item` pattern, so a `&dyn Trait` receiver could never
* dispatch a CALLS edge to the trait's own method. Same v7/v10 contract: the
* incremental write set only covers changed files, so a top-up against a
* pre-v11 index would keep silently missing these CALLS edges for every
* unchanged Rust trait file; force a full re-analyze instead.
* v12: Rust range-binding stopped restoring ambiguous duplicate type names
* (#2514): a function/struct name defined three or more times used to
* re-resolve to the last-scanned file (a presence toggle), so odd duplicate
* counts emitted a wrong cross-file CALLS edge. Same v7/v11 contract: the
* incremental write set only covers changed files, so a top-up against a
* pre-v12 index would keep these spurious CALLS edges on every unchanged Rust
* file. v12 also changes edges in the other direction: range-binding now
* RESOLVES import-disambiguated duplicate names (`for item in make()` /
* `let Struct { f } = ..` where a `use` or `use x::*` import pins one of several
* same-named definitions) to the imported definition's type. Both the removed
* spurious edges and these new resolved edges are cross-file, so a pre-v12
* top-up would leave unchanged Rust files stale either way; force a full
* re-analyze instead.
* v13: Java local classes, enums, records, and interfaces use
* source-type-relative JLS 13.1 identities (`Outer$1Local`). Number allocation
* matches javac: one sequence per (enclosing type, local simple name), with a
* separate sequence for anonymous types. Existing type/member ids, lexical
* bindings, and ownership edges must not be mixed with newly named unchanged
* Java files; force a full re-analyze.
* v14: C# and Kotlin free-call fallback now rejects same-file methods whose
* instance owner is outside the caller's enclosing class/MRO (#2563). The
* incremental write set would otherwise retain those stale CALLS edges on
* every unchanged C# and Kotlin file; force a full re-analyze instead.
* v15: `const X = <arrow | function-expression>` no longer emits an edgeless
* `Const:<file>:X` twin beside its `Function` node (#2687). The incremental
* write set only covers changed files, so every unchanged TS/JS file would
* keep its twin and `impact`/`context` would stay ambiguous on those names;
* force a full re-analyze instead.
* v16: calls through a closure-valued binding (`val f = { }; f()`) now resolve
* in Kotlin, Swift, Dart, Ruby, Java, C# and PHP (#2693). These are NEW `CALLS`
* edges, and those languages also gain callable graph nodes for closure
* bindings that previously carried a value label or no node at all (including
* JS/TS `var f = () => {}`). The incremental write set only covers changed
* files, so unchanged files would keep reporting a zero blast radius for those
* symbols; force a full re-analyze instead.
* v17: `this` inside a JS/TS ordinary `function` no longer resolves to the
* lexically enclosing class (#2701). This REMOVES `CALLS`/`ACCESSES` edges —
* including ones that are correct at runtime via `.bind(this)`, `.call`, or a
* `forEach` thisArg, which the graph does not model. The incremental write set
* only covers changed files, so every unchanged TS/JS file would keep its
* fabricated `this` edges; force a full re-analyze instead.
* v18: function-local callables carry their enclosing-callable chain plus their
* own position, so a local closure no longer shares a node id with a same-named
* file-level function (#2699) — `Function:f.ts:save` ->
* `Function:f.ts:run.save@2:2`. JavaScript/TypeScript also gain block scopes
* (`statement_block`), without which two `const` of one name in sibling blocks
* stay indistinguishable to the resolver and each call resolves to BOTH. This
* CHANGES PERSISTED NODE IDS for every function-local callable and changes
* which node a local call resolves to. An incremental top-up would leave
* unchanged files pointing at the old ids while changed files emit the new
* ones, splitting each symbol in two; force a full re-analyze instead.
* v19: the enclosing-callable walk now stops at class BODIES and anonymous-class
* construction sites, not only at class DECLARATIONS (#2699 follow-up). v18 shipped
* with `CLASS_CONTAINER_TYPES` as the only boundary, which lists no node for a Java
* anonymous class (`object_creation_expression > class_body`), so the walk reached the
* enclosing method and re-keyed `Worker$1.run` as `Worker.makeHandler.run@7:12` —
* destroying the javac-compatible JLS identity of #2550/#2555/#2562. An index stamped
* v18 therefore holds WRONG Java ids, and without this bump it passes the reuse gate
* and keeps them on every unchanged file; force a full re-analyze instead.
* v20: a NAMED explicit receiver no longer resolves its member through the lexical
* scope chain (#2699 follow-up). `options.baseUrl` used to bind to an unrelated
* function-local `const baseUrl`; measured on a 762-file corpus this removes 709
* such edges and adds none. `this`/`self` are exempt, so the 2 genuine self-alias
* reads it also covered are kept. A v19 index holds those false CALLS/ACCESSES on
* every unchanged file and would keep serving them through the reuse gate; force a
* full re-analyze instead.
* v21: a closure bound to a name is a call SOURCE in every language, not only a
* TARGET (#2699 part B). PHP/Rust/Kotlin/Ruby/Dart closure bindings gained the
* declaration rule, Rust gained the graph NODE it never emitted, and Dart locals
* gained the enclosing-callable + position identity that made two same-named
* closures collapse onto one node — which had them asserting a CALLS edge
* present nowhere in the source. All of that changes emitted node ids AND edges
* on files that did not themselves change, so a v20 index topped up
* incrementally keeps serving the old attribution; force a full re-analyze.
*
* v22: CommonJS export forms are indexed (#2723) — `exports.X`/`module.exports.X`,
* aliased receivers, module-level `this`, re-export forwarding, `module.exports = fn`,
* plus prototype/`this` members as Methods with owner edges; and the #2729 review
* fixes that stopped a text-only exports receiver inventing exports inside UMD
* factories and stopped the shadow guard deleting or fabricating call edges.
* These change what is emitted for source whose CONTENT has not changed, so a v21
* index would keep serving the pre-fix graph for every unchanged CommonJS file —
* the exact "Target not found" symptom #2723 reported. Force a full re-analyze.
* v23: Rust module-qualified calls resolve against the module tree (#2730).
* RUST_SCOPE_QUERY gained `@declaration.namespace` on `mod_item` and
* `@reference.qualified-name` on scoped call sites, and a new resolution tier
* binds `tools::dispatch(..)` to the module the path names instead of the
* lexically nearest same-named fn. Same v11/v12 contract: the incremental
* write set only covers CHANGED files, so a top-up against a pre-v23 index
* would keep the wrong self-loop — and keep reporting the callee as unreached
* — for every unchanged Rust file, which is exactly the symptom #2730
* reported. Force a full re-analyze.
*
* v28: structural receiver typing is active for ALL 14 languages, and the fold no
* longer types a bare identifier that merely SHADOWS a class name as that class.
* v27 landed with TypeScript-only emission and with the permissive base lookup, so
* an index stamped 27 by an intermediate build carries both pre-rollout edges for 13
* languages AND the fabricated edges the shadowing bug produced. The reuse gate is a
* strict `===`, so such an index would be treated as current. Re-bumped here so the
* version tracks the final edge semantics. Force a full re-analyze.
*
* v26: receiver expressions are typed from captured structure rather than from
* their source text. `svc?.getUser().save()`, `svc!.getUser().save()` and
* `svc.getTyped<User>().save()` previously emitted NO `CALLS` edge — the text
* cascade split the receiver on punctuation it could not parse — and two of the
* three recorded no drop either, because a later case marked the site handled,
* which suppresses the drop record. So the caller was missing from
* `impact(direction: "upstream")` and `context()` AND the count still claimed
* `epistemic: 'exact'`. Same v11/v12 contract: the incremental write set covers
* only CHANGED files, so a top-up against a pre-v26 index keeps serving the
* pre-fix graph — and the pre-fix confident count — for every unchanged file.
* Worse than merely incomplete: the drop summary is a whole-repo recompute while
* the edges are a changed-files write, so the two would disagree. Force a full
* re-analyze.
*
* v24: inline constructor receivers resolve — `Service(db).do_work()` (Python),
* `new Service(db).doWork()` (JS/TS, C#), `Service.new.do_work` (Ruby), plus the
* generic, qualified, chain-head and keyword-trivia spellings of the same shape
* (#2708). These calls previously emitted NO `CALLS` edge, so the caller was
* missing from `impact(direction: "upstream")` and `context()`. The Ruby
* selector fix also moves an edge: `factory.new.run`, where the class defines an
* instance method named `new`, now resolves through that method again instead of
* being read as construction. All of it changes what is emitted for source whose
* CONTENT has not changed, so a v22 index topped up incrementally — or served by
* the same-commit "already up to date" fast path — keeps returning the pre-fix
* graph for every unchanged file, which is exactly the missing-caller symptom
* #2708 reported. Force a full re-analyze.
*
* v29: Spring @Bean declarations are CodeElement providers and INJECTS may run
* from a consumer Class or factory Method to that CodeElement (#2413). The
* relation DDL gained Class→CodeElement; a pre-v29 database cannot persist that
* label pair, so force a one-time rebuild against the expanded schema.
*
* (This shipped as v25 on its own branch; `main` took 25 through 28 first, so it
* is renumbered at merge time. Re-check both constants against origin/main
* immediately before merging — this is the fifth time that collision has bitten.)
*
* v26: unresolved-receiver member names are persisted
* (`unresolvedReceiverMembers`) so `impact()`/`context()` can report
* `epistemic: 'lower-bound'` instead of a confident `'exact'` when a call site
* was dropped for want of a receiver type (#2744). A pre-v26 index carries no
* such summary, and an absent summary is indistinguishable from "nothing was
* dropped" — so topping one up incrementally would keep reporting `exact` for
* exactly the symbols the signal exists to flag. Force a full re-analyze.
*
* (This shipped as v25 on its own branch; `main` took 25 for #2742 first, so it
* is renumbered here. Re-check both constants against origin/main immediately
* before merging — this is the fourth time that collision has bitten.)
*
* v25: Rust items are qualified by their enclosing `mod` chain (#2742), so
* `mod inner { fn dispatch }` and a crate-root `fn dispatch` in one file are
* finally DISTINCT nodes instead of collapsing onto `Function:<file>:dispatch`
* first-wins. Node IDS CHANGE for every Rust item inside any `mod` block —
* `#[cfg(test)] mod tests` makes that close to every Rust repo — so a pre-v25
* index holds ids an incremental top-up cannot reconcile and would simply
* strand. Force a full re-analyze.
*
* v30: bound-callable graph `startLine` follows the initializer (#2735), so a
* multi-line closure binding joins the scope channel and emits its CALLS edge.
* Pre-v30 indexes keep the wrapper line on unchanged files and would keep
* failing closed (no edge) through the reuse gate. Force a full re-analyze.
*
* v31: Python named imports that resolve to concrete submodules are finalized
* as namespace edges (#2746), enabling qualified constructor and method CALLS
* edges. Pre-v31 indexes retain the old package-target/missing-edge graph for
* unchanged files through the reuse gate. Force a full re-analyze.
*
* v32: the relation DDL (the single shared `CodeRelation` REL TABLE) gains
* sixteen FROM/TO pairs carried by `HAS_METHOD`/`HAS_PROPERTY` and
* scope-resolution edges: Enum→{Function, Method, Struct, Constructor,
* Property, TypeAlias}, Property→{Class, Enum, Function, Struct},
* Method→{Variable, Const}, Trait→Function, Impl→Function, Const→Method and
* Variable→Method. The Enum/Property set was observed on Swift (enums carry
* computed properties, methods, initializers and nested types) and is also
* reached by Java/PHP enum members; Trait/Impl→Function covers a Rust
* `impl`/`trait` method, which is minted as a `Function` node, not `Method`;
* Const/Variable→Method and its sibling Method→Const cover a JS/TS object
* literal's shorthand methods, whose owner is labelled `Const`/`Variable`. A
* pre-v32 database physically lacks these from-to pairs — see
* `assertDeclaredPair` (rel-pair-routing.ts) for why an incremental top-up
* fails loudly on one path and silently on the other. Force a full re-analyze.
*
* (This shipped as v31 on its own branch; `main` took 31 for #2746 first, so
* it is renumbered here. Re-check both constants against origin/main
* immediately before merging — this is the sixth time that collision has
* bitten. If this change is ever reverted, do not free 32 for reuse — the
* reuse gate is exact equality, so an index already stamped 32 would satisfy
* it against a differently-shaped reverted DB. Start the next allocation at
* 33 instead.)
*
* v33: Spring AOP evidence adds the Interface→CodeElement relation pair
* (#2416). LadybugDB fixes allowed endpoint pairs when the relation table is
* created, so an older index cannot persist these edges through incremental
* writeback. Force a full re-analyze.
*
* v34: receiver-chain wire format v2 (name-free `await` / `index` steps). Every
* persisted `ReferenceSite.receiverChain` string changed prefix, and a v2
* decoder refuses a v1 payload by design, so a pre-v34 index carries chains this
* build cannot read. Resolution would silently fall back to the text cascade for
* every chain-carrying site — no error, just quietly worse edges. Force a full
* re-analyze.
*
* Numbered 34, not 33: `main` took 33 for Spring AOP (#2416) mid-flight, landing
* on exactly this branch's number — the seventh collision in this series and the
* first exact clash. Re-check against origin/main before merge.
*/
export const INCREMENTAL_SCHEMA_VERSION = 34;
export interface IndexedRepo {
repoPath: string;
storagePath: string;
@@ -0,0 +1,48 @@
*****************************************************************
* DECLARATIVES + USE AFTER STANDARD ERROR.
*
* `cobol-processor.ts` turns each DECLARATIVES handler into an
* ACCESSES edge from the handler SECTION's `Namespace` node to a
* synthesized `Record` node for the file the USE clause names --
* the `Namespace -> Record` FROM/TO pair. Both endpoints are
* structural (neither label is in `LINKABLE_LABELS`), so no
* scope-bridge cross product reaches it.
*****************************************************************
IDENTIFICATION DIVISION.
PROGRAM-ID. ERRDEMO.
ENVIRONMENT DIVISION.
INPUT-OUTPUT SECTION.
FILE-CONTROL.
SELECT CUSTOMER-FILE ASSIGN TO "CUST.DAT"
ORGANIZATION IS SEQUENTIAL.
SELECT AUDIT-FILE ASSIGN TO "AUDIT.DAT"
ORGANIZATION IS SEQUENTIAL.
DATA DIVISION.
FILE SECTION.
FD CUSTOMER-FILE.
01 CUSTOMER-REC.
05 CUST-ID PIC X(10).
05 CUST-BALANCE PIC 9(7)V99.
FD AUDIT-FILE.
01 AUDIT-REC.
05 AUDIT-TEXT PIC X(60).
WORKING-STORAGE SECTION.
01 WS-EOF PIC X VALUE "N".
PROCEDURE DIVISION.
DECLARATIVES.
CUSTOMER-ERR-HANDLER SECTION.
USE AFTER STANDARD ERROR ON CUSTOMER-FILE.
CUSTOMER-ERR-PARA.
DISPLAY "CUSTOMER IO ERROR".
AUDIT-ERR-HANDLER SECTION.
USE AFTER STANDARD ERROR ON AUDIT-FILE.
AUDIT-ERR-PARA.
DISPLAY "AUDIT IO ERROR".
END DECLARATIVES.
MAIN-SECTION SECTION.
MAIN-PARA.
OPEN INPUT CUSTOMER-FILE
OPEN OUTPUT AUDIT-FILE
CLOSE CUSTOMER-FILE
CLOSE AUDIT-FILE
STOP RUN.
@@ -0,0 +1,22 @@
"""`@mcp.tool()` applied to a CLASS rather than a function.
`pipeline-phases/tools.ts` hangs the HANDLES_TOOL edge off whatever node the
tool decorator sat on (`handlerNodeId`), so a class-decorated tool produces a
`Class -> Tool` edge. Every other tool fixture decorates a function or falls
back to the file, so `Function -> Tool` / `File -> Tool` are the only pairs
they exercise.
"""
from mcp import tool
def _render(payload: dict) -> str:
return str(payload)
@mcp.tool()
class WeatherTool:
"""Class-based MCP tool."""
def run(self, payload: dict) -> str:
return _render(payload)
@@ -0,0 +1,30 @@
package com.example;
import org.springframework.boot.autoconfigure.condition.ConditionalOnMissingBean;
import org.springframework.boot.autoconfigure.condition.ConditionalOnProperty;
import org.springframework.context.annotation.Bean;
import org.springframework.context.annotation.Configuration;
/**
* `frameworks/spring/conditionals.ts` mints one `Annotation` node per Spring
* condition and links its OWNER to it with a CONDITIONAL_ON edge. When the
* owner is an `@Bean` factory METHOD (rather than the `@Configuration` class),
* the edge is `Method -> Annotation` — the source label comes from the
* scope-resolution bridge, the target is a structural `Annotation` node that
* is in neither scope-bridge label set.
*/
@Configuration
public class AppAutoConfiguration {
@Bean
@ConditionalOnMissingBean
public PaymentService paymentService() {
return new PaymentService();
}
@Bean
@ConditionalOnProperty(prefix = "billing", name = "enabled", havingValue = "true")
public PaymentService fallbackPaymentService() {
return new PaymentService();
}
}
@@ -0,0 +1,5 @@
package com.example;
public class PaymentService {
public void pay() {}
}
@@ -0,0 +1,20 @@
package com.example.kotlin
import org.springframework.boot.autoconfigure.condition.ConditionalOnMissingBean
import org.springframework.context.annotation.Bean
import org.springframework.context.annotation.Configuration
class ReportService
/**
* Kotlin twin of the Java `@Bean` + `@ConditionalOnMissingBean` case — the same
* `Method -> Annotation` pair, reached through the Kotlin conditional metadata
* adapter instead of the Java one.
*/
@Configuration
class KotlinAutoConfiguration {
@Bean
@ConditionalOnMissingBean
fun reportService(): ReportService = ReportService()
}
@@ -0,0 +1,29 @@
<template>
<div class="options-host">
<Button variant="secondary" @click="onOptionsClick">
Options click
</Button>
</div>
</template>
<script lang="ts">
import { defineComponent } from 'vue';
import Button from './components/Button.vue';
/**
* Options-API host: the `@click` handler lives in `methods:`, so the graph node
* for it is a `Method`, not a `Function`. The BINDS_EVENT_HANDLER edge Vue's
* scope resolver emits for a component-element binding therefore has a `Method`
* source and a `File` target — the `Method→File` pair. The `<script setup>`
* sibling (App.vue) only ever produces `Function→File`.
*/
export default defineComponent({
name: 'OptionsHost',
components: { Button },
methods: {
onOptionsClick() {
return 'options-clicked';
},
},
});
</script>
+416
View File
@@ -0,0 +1,416 @@
/**
* Shared child-process module-load probe for the `dist/` import-closure tests.
*
* Extracted at its THIRD consumer (`mini-repo.ts` set the precedent at its
* second). `test/integration/mcp/import-closure.test.ts`,
* `test/integration/optional-grammars/registry-import-closure.test.ts` and
* `test/integration/mcp/startup-language-closure.test.ts` had each grown their
* own copy of the same machinery: the `REPO_ROOT` derivation, the probe source,
* the "dist missing — run `npm run build`" guard, the spawn with `NODE_OPTIONS`
* cleared, the status-vs-signal error rendering, and the JSON payload parse.
*
* The copies were not equal, which is what made the duplication actively
* harmful rather than merely verbose. Two of the three diffed `require.cache`
* only — structurally BLIND to the first-party ESM `dist/**` graph they walk
* (`dist/` is `"type": "module"`), so they could not see most of what they
* traversed — and one of those had no non-vacuity guard at all, meaning a
* severed entry passed it green. The next author had 2-in-3 odds of copying a
* broken probe.
*
* So the HARNESS is shared and the POLICY is not: which modules are forbidden,
* and what the remedy is, stays in each test, because that advice is specific
* to the regression that test exists to prevent.
*
* Two load channels, unioned:
* - `module.registerHooks({ load })` sees every module the ESM loader
* resolves, including the first-party `dist/**` graph. (Added in Node
* 22.15; the package `engines` floor is `^22.18.0 || >=24.11.0`.)
* - a `require.cache` diff catches CJS/native modules, which is how a
* tree-sitter grammar binding or a `.node` addon surfaces.
*
* Non-vacuity is STRUCTURAL here, not a convention a caller can forget: every
* request MUST declare an `anchor` module and a `minModules` floor, and the
* probe throws unless both hold. "The probe loaded nothing" is the one failure
* mode that turns every one of these tests green while asserting nothing, so it
* is not left to the test author to remember.
*
* An anchor is PER-POLICY, not per-entry. One probe is routinely asserted over
* by several INDEPENDENT policies ("loads no language provider" AND "loads no
* group extractor"), and each policy is only non-vacuous while the chain IT
* polices is still walked. A single anchor on one of those chains, plus the
* module-count floor, both stay green when a DIFFERENT chain is severed — and
* the policy that rode on it silently stops being able to fail. So `anchor`
* takes a list: name one module per policy, e.g.
* `anchor: ['dist/mcp/resources.js', 'dist/core/group/service.js']`. A bare
* string is the single-policy shorthand.
*
* Lazy `await import(...)` inside a function body remains the sanctioned escape
* hatch throughout: it does not run at module evaluation, so the probe does not
* see it. A TOP-LEVEL `await import(...)` does run, and the probe reports it —
* which is the point.
*/
import { spawn } from 'node:child_process';
import fs from 'node:fs';
import path from 'node:path';
import { fileURLToPath, pathToFileURL } from 'node:url';
/** `test/helpers/module-load-probe.ts` → the gitnexus package root is two levels up. */
const REPO_ROOT = path.resolve(path.dirname(fileURLToPath(import.meta.url)), '..', '..');
/**
* Payload delimiters. The child writes its JSON between them so a stray
* `console.log` from an imported module cannot corrupt the payload — several of
* the entries probed here print banners on load.
*/
const BEGIN = '<<<GITNEXUS_PROBE>>>';
const END = '<<<END_GITNEXUS_PROBE>>>';
/**
* Hard bound on one child. Generous: the heaviest entry probed today (the
* scope-resolution registry, which eagerly loads ~40 tree-sitter bindings)
* takes ~9 s, and CI machines are slower. This is a wedge-breaker, not a
* performance assertion — nothing here asserts on elapsed time.
*/
const PROBE_TIMEOUT_MS = 60_000;
const PROBE_SOURCE = `
import { createRequire, registerHooks } from 'node:module';
const req = createRequire(import.meta.url);
const loaded = new Set();
registerHooks({
load(url, context, nextLoad) {
loaded.add(url);
return nextLoad(url, context);
},
});
const beforeCjs = new Set(Object.keys(req.cache));
await import(process.env.PROBE_TARGET);
for (const key of Object.keys(req.cache)) {
if (!beforeCjs.has(key)) loaded.add(key);
}
process.stdout.write('${BEGIN}' + JSON.stringify([...loaded]) + '${END}');
`;
/**
* `path.relative(from, to)` rendered with POSIX separators, so paths compare
* and print identically on Windows. Inline copies of this three-step dance had
* accumulated across the test tree; new ones should call this.
*/
export function relativePosix(from: string, to: string): string {
return path.relative(from, to).split(path.sep).join('/');
}
/**
* Render one probe entry — a `file:` URL from the ESM hook, or an absolute path
* from `require.cache` — as a repo-relative POSIX path.
*
* Anything that is not an absolute path inside the repo (a `node:` builtin, a
* globally-linked dependency) is returned VERBATIM, so failure output still
* names it recognizably and so the rendering never depends on `process.cwd()`.
*/
export function toRepoRelativePosix(entry: string): string {
const asPath = entry.startsWith('file:') ? fileURLToPath(entry) : entry;
if (!path.isAbsolute(asPath)) return entry;
const relative = relativePosix(REPO_ROOT, asPath);
return relative.startsWith('..') || relative === '' ? entry : relative;
}
/** What to probe, and what proves the probe actually reached the target's graph. */
export interface ModuleLoadRequest {
/**
* Path under `dist/`, POSIX separators: `'mcp/local/local-backend.js'`.
* Also the probe's label in every message it emits.
*/
readonly entry: string;
/**
* Module(s) the entry genuinely loads, as `toRepoRelativePosix` renders them
* (e.g. `'dist/mcp/resources.js'`). Every one must be present or the probe
* throws.
*
* Non-vacuity guard, and REQUIRED: if a refactor severs the entry from its
* real graph, the probe fails loudly instead of letting "none of the
* forbidden modules loaded" pass green over a graph nothing walked. Each
* anchor must sit on the same edge a policy asserted over this probe
* polices — one whose disappearance would make that policy's assertion
* meaningless. ONE PER POLICY: a file running two independent policies over
* one probe needs two anchors (see the module doc), because an anchor on
* policy A's chain says nothing about whether policy B's chain is still
* walked.
*/
readonly anchor: string | readonly string[];
/**
* Floor on the number of distinct modules loaded — a coarser second
* non-vacuity guard, and REQUIRED. Set it well below the observed count so
* ordinary dependency churn does not trip it.
*/
readonly minModules: number;
/**
* Extra child environment, merged last. Use it to neutralise env that would
* change what the child loads (the caller's own env is inherited, with
* `NODE_OPTIONS` already cleared).
*/
readonly env?: Readonly<Record<string, string>>;
}
/** What one entry actually loaded. */
export interface ModuleLoadProbe {
/** `dist/mcp/server.js` — the entry, as it appears in failure messages. */
readonly label: string;
/**
* Every distinct module the child loaded, in load order, rendered by
* `toRepoRelativePosix`. Includes the ESM `dist/**` graph and the CJS/native
* modules under it.
*/
readonly modules: readonly string[];
/** The loaded modules matching `pattern` — the offender list for a policy assertion. */
matching(pattern: RegExp): readonly string[];
}
/** The probes recorded by one concurrent `probeModuleLoads` call. */
export interface ModuleLoadProbes {
/** The probe for `entry`. Throws when it was never requested — never returns empty. */
get(entry: string): ModuleLoadProbe;
}
interface ProbeProcessResult {
readonly status: number | null;
readonly signal: NodeJS.Signals | null;
readonly stdout: string;
readonly stderr: string;
}
/** A settled probe: the outcome, or the failure that stopped it — always labelled. */
type SettledProbe =
| { readonly label: string; readonly probe: ModuleLoadProbe }
| { readonly label: string; readonly error: Error };
/** The label a request is reported and looked up under. */
function labelOf(entry: string): string {
return `dist/${entry}`;
}
/**
* Normalise {@link ModuleLoadRequest.anchor} to a list. Exported so a test can
* derive "which entries does policy X apply to?" from the anchors themselves
* rather than from a hand-maintained second list that can drift out of step
* with them.
*/
export function anchorsOf(anchor: string | readonly string[]): readonly string[] {
return typeof anchor === 'string' ? [anchor] : anchor;
}
/**
* Run the probe against `targetUrl` in a fresh child process.
*
* Async `spawn` rather than `spawnSync` so several entries can be probed
* CONCURRENTLY: `spawnSync` blocks the event loop, and vitest runs a file's
* tests sequentially, so a per-test sync probe serialises N full Node starts
* that share nothing.
*/
function spawnProbe(targetUrl: string, extraEnv: Readonly<Record<string, string>>) {
return new Promise<ProbeProcessResult>((resolve, reject) => {
const child = spawn(process.execPath, ['--input-type=module', '-e', PROBE_SOURCE], {
cwd: REPO_ROOT,
// NODE_OPTIONS is cleared so a session-pinned --max-old-space-size (or a
// loader flag) can't perturb which modules the child evaluates.
//
// PROBE_TARGET is spread AFTER `extraEnv` deliberately: the harness's own
// target must always win. A request whose `env` set PROBE_TARGET would
// otherwise redirect the probe at a different module while `anchor` and
// `minModules` stayed keyed on `entry` — the exact vacuity this file
// exists to make impossible.
env: { ...process.env, NODE_OPTIONS: '', ...extraEnv, PROBE_TARGET: targetUrl },
timeout: PROBE_TIMEOUT_MS,
// SIGKILL rather than the default SIGTERM, which is catchable and
// ignorable — a child wedged in synchronous native code (the failure this
// guards) would survive it. Same reasoning as `lbug-config.ts`'s spawn.
// Node's own `timeout` delivers this, so no second timer to keep in step.
killSignal: 'SIGKILL',
});
let stdout = '';
let stderr = '';
child.stdout.setEncoding('utf8');
child.stderr.setEncoding('utf8');
child.stdout.on('data', (chunk: string) => {
stdout += chunk;
});
child.stderr.on('data', (chunk: string) => {
stderr += chunk;
});
child.on('error', (error) => {
reject(error);
});
child.on('close', (status, signal) => {
resolve({ status, signal, stdout, stderr });
});
});
}
function runProbe(request: ModuleLoadRequest, label: string): Promise<ModuleLoadProbe> {
const target = path.join(REPO_ROOT, 'dist', ...request.entry.split('/'));
if (!fs.existsSync(target)) {
return Promise.reject(
new Error(
`${target} missing — run \`npm run build\` first (or \`npm run test:integration\`, ` +
`which builds via pretest:integration).`,
),
);
}
return spawnProbe(pathToFileURL(target).href, request.env ?? {}).then((result) => {
if (result.status !== 0) {
// `status` is null when the child died to a signal (e.g. a native addon
// SIGSEGV) — report the signal so that reads differently from a plain
// non-zero exit.
const exit =
result.status !== null ? `status ${result.status}` : `signal ${result.signal ?? 'unknown'}`;
throw new Error(
`probing ${label} failed (${exit}):\n` +
`stderr:\n${result.stderr}\nstdout:\n${result.stdout}`,
);
}
const begin = result.stdout.indexOf(BEGIN);
const end = result.stdout.indexOf(END);
if (begin < 0 || end < begin) {
throw new Error(
`probe output for ${label} had no payload markers.\n` +
`stdout:\n${result.stdout}\nstderr:\n${result.stderr}`,
);
}
const raw = parsePayload(result.stdout.slice(begin + BEGIN.length, end), label);
// Deduplicate AFTER rendering: a CJS module imported from ESM is reported
// once per channel (a `file:` URL and an absolute path) and would otherwise
// appear twice in every offender list. Load order is preserved — it is the
// most useful thing in a failure dump.
const modules = [...new Set(raw.map(toRepoRelativePosix))];
assertNonVacuous(request, label, modules);
return {
label,
modules,
matching: (pattern: RegExp): readonly string[] => modules.filter((m) => pattern.test(m)),
};
});
}
/**
* Parse the child's payload back into a module list.
*
* Validated rather than cast: this crosses a process boundary, so the shape is
* an assumption about another process's output, not a fact the type system
* knows. A malformed payload must read as a HARNESS failure with the raw text
* attached, never as `undefined` flowing into the offender lists.
*/
function parsePayload(payload: string, label: string): readonly string[] {
const parsed: unknown = JSON.parse(payload);
if (!isModuleList(parsed)) {
throw new Error(
`probe payload for ${label} was not an array of module strings — the child's ` +
`protocol changed, or a module wrote between the payload markers. Payload:\n${payload}`,
);
}
return parsed;
}
function isModuleList(value: unknown): value is readonly string[] {
return Array.isArray(value) && value.every((entry: unknown) => typeof entry === 'string');
}
/**
* Fail unless the probe demonstrably walked the entry's real graph. Thrown, not
* asserted, because a vacuous probe is a harness failure: every policy
* assertion built on it is meaningless, so no test should get the chance to
* evaluate one.
*
* EVERY anchor must be present, not merely one: they are one-per-policy, so a
* surviving anchor cannot vouch for a severed sibling's chain.
*/
function assertNonVacuous(
request: ModuleLoadRequest,
label: string,
modules: readonly string[],
): void {
const missing = anchorsOf(request.anchor).filter((anchor) => !modules.includes(anchor));
if (missing.length > 0) {
throw new Error(
`${label} did not load its anchor(s) ${missing.join(', ')}. If that edge moved, repoint ` +
`the anchor — otherwise this probe is reporting over an unexercised graph and every ` +
`assertion on it is vacuous. Loaded (${modules.length}):\n${modules.join('\n')}`,
);
}
if (modules.length < request.minModules) {
throw new Error(
`${label} loaded ${modules.length} modules, below its floor of ${request.minModules} — ` +
`the probe did not reach the entry's real graph. Loaded:\n${modules.join('\n')}`,
);
}
}
/**
* Probe every request CONCURRENTLY and return the results by entry.
*
* Each request is an independent child process paying a full Node start, so
* running them in parallel is worth roughly a 60% wall-clock cut on a
* three-entry file. Call this once from `beforeAll` and keep the `it` bodies
* pure assertions over what it recorded.
*
* Every failure is reported WITH its entry — a shared hook must not collapse N
* distinct probes into one anonymous "beforeAll failed". Rejections are caught
* per request rather than raced, which also guarantees every child is reaped
* before this resolves.
*/
export async function probeModuleLoads(
requests: readonly ModuleLoadRequest[],
): Promise<ModuleLoadProbes> {
const settled = await Promise.all(
requests.map((request): Promise<SettledProbe> => {
const label = labelOf(request.entry);
return runProbe(request, label).then(
(probe) => ({ label, probe }),
(error: unknown) => ({
label,
error: error instanceof Error ? error : new Error(String(error)),
}),
);
}),
);
const failures = settled.flatMap((r) => ('error' in r ? [`${r.label}: ${r.error.message}`] : []));
if (failures.length > 0) {
throw new Error(
`${failures.length} of ${requests.length} module-load probes failed:\n\n` +
failures.join('\n\n'),
);
}
const byLabel = new Map(
settled.flatMap((r) => ('probe' in r ? ([[r.label, r.probe]] as const) : [])),
);
return {
get(entry: string): ModuleLoadProbe {
const label = labelOf(entry);
const probe = byLabel.get(label);
if (probe === undefined) {
throw new Error(
`no module-load probe recorded for ${label} — probed: ` +
`${[...byLabel.keys()].join(', ')}. (A typo here would otherwise read as a pass.)`,
);
}
return probe;
},
};
}
/** Probe a single entry. Same guarantees as {@link probeModuleLoads}. */
export async function probeModuleLoad(request: ModuleLoadRequest): Promise<ModuleLoadProbe> {
return (await probeModuleLoads([request])).get(request.entry);
}
+68
View File
@@ -0,0 +1,68 @@
/**
* Shared harness for the suites that drive the REAL `mountSSEProgress` relay
* over a REAL HTTP server (server-sse-payload.test.ts and analyze-api.test.ts).
*
* Both files had grown their own listen/port/close dance plus their own SSE
* frame parser, and the parsers had already diverged in what they do with a
* missing frame. The bug these suites exist to catch is only observable END TO
* END, so the harness stays real — one Express app, one ephemeral port, one
* JobManager — and only the boilerplate is shared (the embedding-seed.ts
* precedent: a helper module has no describe-registration problem, unlike
* importing a sibling test file).
*/
import express from 'express';
import http from 'node:http';
import { JobManager } from '../../src/server/analyze-job.js';
import { mountSSEProgress } from '../../src/server/sse-progress.js';
export interface SSEHarness {
/** The JobManager the mounted relay reads. Drive the test through this. */
readonly manager: JobManager;
/** `http://127.0.0.1:<ephemeral port>` — prefix for the mounted route. */
readonly baseUrl: string;
/** Disposes the JobManager and closes the server. Call from `afterEach`. */
close(): Promise<void>;
}
/**
* Mount the relay at `routePath` on a fresh Express app bound to an ephemeral
* localhost port. `routePath` mirrors whichever production mount in
* `createServer()` the suite is standing in for.
*/
export const startSSEHarness = async (routePath: string): Promise<SSEHarness> => {
const manager = new JobManager();
const app = express();
mountSSEProgress(app, routePath, manager);
const server = await new Promise<http.Server>((resolve) => {
const s = app.listen(0, '127.0.0.1', () => resolve(s));
});
const addr = server.address();
const port = typeof addr === 'object' && addr ? addr.port : 0;
return {
manager,
baseUrl: `http://127.0.0.1:${port}`,
close: () => {
manager.dispose();
return new Promise<void>((resolve, reject) => {
server.close((err) => (err ? reject(err) : resolve()));
});
},
};
};
/**
* The parsed JSON payload of the named terminal SSE frame, or `undefined` when
* the stream carried no such frame — "the client was never told" is an outcome
* these suites assert on, so an absent frame must not throw.
*/
export const terminalFrame = (body: string, event: 'complete' | 'failed'): unknown => {
const frame = body.split('\n\n').find((f) => f.includes(`event: ${event}`));
const dataLine = frame?.split('\n').find((line) => line.startsWith('data: '));
return dataLine === undefined ? undefined : JSON.parse(dataLine.slice('data: '.length));
};
/** How many terminal frames of any kind the client received. */
export const terminalFrameCount = (body: string): number =>
body.split('\n\n').filter((f) => /^event: (complete|failed)$/m.test(f)).length;
+130
View File
@@ -0,0 +1,130 @@
/**
* Temp-directory lifecycle for pipeline-level integration tests.
*
* The pipeline mutates the repo it is handed (parse caches, `.gitnexus/`), so
* these tests each run against a throwaway copy of a fixture. Every consumer
* had hand-rolled the SAME three parts — a `string[]` of created dirs, a
* `mkdtempSync` that pushes onto it, and an `afterAll` that `rmSync`s the lot.
* Extracted at the fourth consumer (`pipeline-pdg`, `pipeline-pdg-streaming`,
* `interproc-taint`, `pdg-chained-receiver-callees`); the copies had already
* drifted — `pipeline-pdg` registered two cleanup hooks over one array.
*
* Only the LIFECYCLE is shared, deliberately: seeding differs per test (a
* recursive fixture copy, a single file, an inline-written source, or nothing
* at all), so `dir()` hands back an empty registered directory and the caller
* fills it however it likes. `fromFixture()` is the common case.
*
* `createTempDirPool` calls `afterAll` itself, so it must be called from a
* test file's module scope (not from this module's top level — ESM caching
* would register the hook once, for whichever file imported it first).
* Directories are registered at creation, before any seeding runs, so a
* fixture copy or a pipeline run that throws still leaves them cleaned up.
*
* Cleanup is best-effort PER DIRECTORY — see `removeTempDirs`. Every one of the
* hand-rolled copies looped bare `rmSync` calls, so the first failure aborted
* the removal of every directory after it; consolidating them made that one
* loop the single point of failure for four suites.
*/
import fs from 'fs';
import os from 'os';
import path from 'path';
import { afterAll } from 'vitest';
import { cleanupTempDirSync } from './test-db.js';
export interface TempDirPool {
/** A fresh empty temp dir, registered for cleanup. Seed it yourself. */
dir(): string;
/** A fresh temp dir seeded with a recursive copy of `fixture`. */
fromFixture(fixture: string): string;
}
/** Removes one registered directory. A seam, so the failure path is testable. */
export type TempDirRemover = (dir: string) => void;
/** Reports one cleanup failure. A seam, for the same reason. */
type CleanupWarner = (message: string) => void;
/**
* Delegates to `cleanupTempDirSync`, the repo's existing Windows-lock-aware
* remover — do NOT re-roll `fs.rmSync` here. It already encodes the whole
* problem this pool hit: `force` suppresses only `ENOENT`, while a handle a
* pipeline test left open surfaces as `EBUSY`/`EPERM`, so it retries 5× with a
* 100–400 ms backoff and then swallows exactly the Windows lock codes and
* `ENOTEMPTY` — rethrowing anything else, so a genuine bug still surfaces
* through `removeTempDirs`' per-directory catch below.
*
* A second copy here had already drifted from it on both knobs that matter
* (3 retries at 50 ms, and warn-on-everything), which is how one half of the
* suite ends up green-with-a-warning on the same `EBUSY` the other half fails on.
*
* Exported so the cleanup pin can inject a failure for ONE directory while the
* others still go through the removal that actually ships — a proof against a
* stand-in `fs.rmSync` call in the test would not be one.
*/
export const removeTempDirRecursive: TempDirRemover = (dir) => {
cleanupTempDirSync(dir);
};
// Node's console methods are bound, so this can be the default directly.
const warnToConsole: CleanupWarner = console.warn;
/**
* Remove every registered directory, best-effort: one failure must not abort
* the removal of the directories after it.
*
* WARN — not throw, not swallow. Throwing would fail an otherwise green suite
* over housekeeping the OS reclaims anyway, and it would do so from `afterAll`,
* where it reads as a test failure and buries the real result. Swallowing is
* its own hazard: a systematic leak (a runner whose tmpdir keeps filling) would
* then be invisible, with nothing naming the suite responsible. A warning
* carrying the path costs nothing on the happy path, and the path's `mkdtemp`
* prefix is per-pool, so it names the suite that made it.
*
* Exported so the failure path can be pinned by injecting a throwing `remove`:
* a real `EBUSY` is not reproducible on demand, and a test that waited for one
* would be non-deterministic. The `afterAll` below calls exactly this function,
* so that pin is over the loop that actually ships.
*/
export function removeTempDirs(
dirs: readonly string[],
remove: TempDirRemover = removeTempDirRecursive,
warn: CleanupWarner = warnToConsole,
): void {
for (const dir of dirs) {
try {
remove(dir);
} catch (error) {
const reason = error instanceof Error ? error.message : String(error);
warn(`[temp-dir-pool] could not remove ${dir}: ${reason}`);
}
}
}
/**
* Create a pool of temp directories that are removed after the calling test
* file finishes. `prefix` is the `mkdtemp` prefix (e.g. `'gn-pdg-'`), kept
* per-pool so a leaked directory still names the suite that made it.
*/
export function createTempDirPool(prefix: string): TempDirPool {
const created: string[] = [];
const dir = (): string => {
const made = fs.mkdtempSync(path.join(os.tmpdir(), prefix));
created.push(made);
return made;
};
afterAll(() => {
removeTempDirs(created);
});
return {
dir,
fromFixture(fixture: string): string {
const made = dir();
fs.cpSync(fixture, made, { recursive: true });
return made;
},
};
}
@@ -12,33 +12,23 @@
* the parse worker, and a stale dist is a spurious red.
*/
import { describe, it, expect, afterAll } from 'vitest';
import fs from 'fs';
import os from 'os';
import { describe, it, expect } from 'vitest';
import path from 'path';
import { runPipelineFromRepo } from '../../../src/core/ingestion/pipeline.js';
import type { PipelineResult } from '../../../src/types/pipeline.js';
import { decodeTaintPath } from '../../../src/core/ingestion/taint/path-codec.js';
import { createTempDirPool } from '../../helpers/temp-dir-pool.js';
const FIXTURE = path.join(__dirname, 'fixtures', 'interproc-repo');
const tmpDirs: string[] = [];
function freshRepo(): string {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gn-interproc-'));
fs.cpSync(FIXTURE, dir, { recursive: true });
tmpDirs.push(dir);
return dir;
}
const repos = createTempDirPool('gn-interproc-');
const freshRepo = (): string => repos.fromFixture(FIXTURE);
function taintPaths(result: PipelineResult) {
return [...result.graph.iterRelationships()].filter((r) => r.type === 'TAINT_PATH');
}
describe('U9 — end-to-end interprocedural taint (--pdg)', () => {
afterAll(() => {
for (const d of tmpDirs) fs.rmSync(d, { recursive: true, force: true });
});
it('with --pdg: composes a cross-file source→sink into a TAINT_PATH edge', async () => {
const result = await runPipelineFromRepo(freshRepo(), () => {}, { pdg: true });
const paths = taintPaths(result);
@@ -0,0 +1,305 @@
/**
* The PDG inter-procedural descent hops through `BasicBlock.calleeIds`, so it
* can only cross a call boundary that the RESOLVER managed to resolve. Chained
* receiver calls (`out.inner().compute(x)`) are resolved by the receiver-typing
* pass, whose resolved ids reach `calleeIds` through a separate sink from the
* plain-call path — which means the chain could regress there without any
* plain-call test noticing.
*
* This pins the resolver -> PDG seam for a chain: the block holding the chained
* statement must carry the id of EVERY link, not just the first. The descent's
* behaviour once the ids are present is covered by impact-pdg-interproc and
* impact-pdg-fullchain-e2e; what those cannot catch is a chain link silently
* missing from the column they both read.
*
* ── WHERE THE SUPPORT ACTUALLY STOPS ──────────────────────────────────────────
*
* Chained resolution is NOT general. Measured against this fixture (one repo per
* shape and all shapes in one repo agree, so the rows do not contaminate each
* other), the discriminator is whether the receiver's type is DECLARED, not
* whether it is a local or a field:
*
* receiver form calleeIds cell
* --------------------------------------------------- ------------------------
* local `const o = new Outer()` Outer.inner + Inner.compute
* field `private p: Outer = new Outer()` Outer.inner + Inner.compute
* field `private p: Outer;` + ctor `this.p = new ...` Outer.inner + Inner.compute
* receiver is a call result `makeOuter().inner()...` makeOuter + both links
* three links `o.inner().mid().compute()` all three links
* field `private p = new Outer()` (INFERRED) EMPTY <- known gap
* field `private p;` + ctor `this.p = new Outer()` EMPTY <- known gap
*
* An inference-typed receiver does not merely lose the CHAINED link — it empties
* the whole cell, so the descent cannot cross into `Outer.inner` either, even
* though that call has a perfectly ordinary named receiver. Those two rows are
* pinned by `KNOWN GAP: an inference-typed receiver empties the WHOLE calleeIds
* cell` below, which asserts the current empty value EXACTLY.
*
* ── THIS PIN IS SELF-DIFFING: IT WILL GO RED ON PURPOSE ───────────────────────
*
* The gap is tracked as issue #2807 ("Inference-typed field receivers resolve to
* no CALLS edges at all"); PR #2810 is open against it at the time of writing.
* The `KNOWN GAP` test asserts that the gap EXISTS — the empty cell, exactly —
* so it is not a regression guard, it is a record. Whoever closes #2807 will see
* it fail with the newly resolved ids in the diff; that is the intended signal,
* and the fix is to update this file (the table above, the rows' `resolution`,
* and the pin's expected value), not to relax the assertion. A single-shape
* fixture, or an `it.fails` row, would instead have kept quietly implying that
* chained receivers work in general.
*
* The gap was FOUND during #2802 work but is PRE-EXISTING and independent of it:
* nothing on that branch touches receiver typing. Track the gap itself at #2807.
*
* Self-contained fixture rather than an addition to `fixtures/pdg-repo` — that
* fixture is shared by eight suites including a snapshot test, so growing it to
* cover one seam churns unrelated expectations.
*/
import { describe, it, expect, beforeAll } from 'vitest';
import fs from 'fs';
import path from 'path';
import { runPipelineFromRepo } from '../../../src/core/ingestion/pipeline.js';
import { createTempDirPool } from '../../helpers/temp-dir-pool.js';
// The PRODUCTION reader of the cell: splits on `CALLEE_ID_SEP`
// (src/core/ingestion/cfg/emit.ts) and drops the truncation sentinel. Both the
// statement-precise bridge and the inter-procedural descent go through it, so
// asserting on its output is asserting on exactly the ids the descent sees —
// and it yields whole ids, which a substring match over the raw cell would not.
import { splitCalleeIds } from '../../../src/mcp/local/pdg-impact.js';
const FIXTURE_PATH = 'src/app.ts';
// Every caller below chains `.compute()` onto the RESULT of `.inner()`; the
// second call has no named receiver, so it resolves only if the receiver's type
// is carried through the chain. Only the receiver FORM varies between rows.
const CHAINED_SOURCE = `export class Mid {
compute(v: number): number {
return v * 3;
}
}
export class Inner {
compute(v: number): number {
return v * 2;
}
mid(): Mid {
return new Mid();
}
}
export class Outer {
inner(): Inner {
return new Inner();
}
}
export function makeOuter(): Outer {
return new Outer();
}
export function runLocalConst(x: number): number {
const localConst = new Outer();
const r = localConst.inner().compute(x);
return r;
}
export function runCallResultReceiver(x: number): number {
const r = makeOuter().inner().compute(x);
return r;
}
export function runThreeLink(x: number): number {
const threeLink = new Outer();
const r = threeLink.inner().mid().compute(x);
return r;
}
export class AnnotatedFieldCaller {
private annotated: Outer = new Outer();
run(x: number): number {
const r = this.annotated.inner().compute(x);
return r;
}
}
export class InferredFieldCaller {
private inferred = new Outer();
run(x: number): number {
const r = this.inferred.inner().compute(x);
return r;
}
}
export class CtorAssignedAnnotatedCaller {
private ctorTyped: Outer;
constructor() {
this.ctorTyped = new Outer();
}
run(x: number): number {
const r = this.ctorTyped.inner().compute(x);
return r;
}
}
export class CtorAssignedInferredCaller {
private ctorUntyped;
constructor() {
this.ctorUntyped = new Outer();
}
run(x: number): number {
const r = this.ctorUntyped.inner().compute(x);
return r;
}
}
`;
// EXACT resolved ids — never substrings. `Inner.compute` as a substring is also
// satisfied by `Inner.computeExtra` and by `OtherInner.compute`, while the
// descent keys on the whole id for its span and CALL_SUMMARY lookups. The `#N`
// suffix is the arity disambiguator the resolver mints.
const OUTER_INNER = `Method:${FIXTURE_PATH}:Outer.inner#0`;
const INNER_COMPUTE = `Method:${FIXTURE_PATH}:Inner.compute#1`;
const INNER_MID = `Method:${FIXTURE_PATH}:Inner.mid#0`;
const MID_COMPUTE = `Method:${FIXTURE_PATH}:Mid.compute#1`;
const MAKE_OUTER = `Function:${FIXTURE_PATH}:makeOuter`;
/**
* `reaches-pdg` — every link's id lands in the cell today.
* `known-gap-empty-cell` — the resolver cannot type the receiver, so the cell is
* emitted EMPTY and the descent cannot cross ANY link of the chain.
*/
type ChainResolution = 'reaches-pdg' | 'known-gap-empty-cell';
interface ReceiverShape {
/** Row name; also the key of the known-gap pin below. */
readonly name: string;
/** Unique fragment of the chained statement, used to find its block. */
readonly marker: string;
/** Every link of the chain, as an exact resolved id. */
readonly links: readonly string[];
readonly resolution: ChainResolution;
}
const RECEIVER_SHAPES: readonly ReceiverShape[] = [
{
name: 'local-const',
marker: 'localConst.inner().compute(',
links: [OUTER_INNER, INNER_COMPUTE],
resolution: 'reaches-pdg',
},
{
name: 'annotated-field',
marker: 'this.annotated.inner().compute(',
links: [OUTER_INNER, INNER_COMPUTE],
resolution: 'reaches-pdg',
},
{
name: 'ctor-assigned-annotated',
marker: 'this.ctorTyped.inner().compute(',
links: [OUTER_INNER, INNER_COMPUTE],
resolution: 'reaches-pdg',
},
{
name: 'call-result-receiver',
marker: 'makeOuter().inner().compute(',
links: [MAKE_OUTER, OUTER_INNER, INNER_COMPUTE],
resolution: 'reaches-pdg',
},
{
name: 'three-link-chain',
marker: 'threeLink.inner().mid().compute(',
links: [OUTER_INNER, INNER_MID, MID_COMPUTE],
resolution: 'reaches-pdg',
},
// ── Known gaps ────────────────────────────────────────────────────────────
// Identical to the two rows above except that the field has no type
// annotation, so its type would have to be inferred from the initializer.
{
name: 'inferred-field',
marker: 'this.inferred.inner().compute(',
links: [OUTER_INNER, INNER_COMPUTE],
resolution: 'known-gap-empty-cell',
},
{
name: 'ctor-assigned-inferred',
marker: 'this.ctorUntyped.inner().compute(',
links: [OUTER_INNER, INNER_COMPUTE],
resolution: 'known-gap-empty-cell',
},
];
interface BlockCell {
readonly text: string;
readonly ids: readonly string[];
}
const repos = createTempDirPool('gn-pdg-chain-');
let blocks: readonly BlockCell[] = [];
function blocksFor(marker: string): readonly BlockCell[] {
return blocks.filter((b) => b.text.includes(marker));
}
function idsFor(marker: string): readonly string[] {
const matched = blocksFor(marker);
// Exactly one block spans each chained statement; a fixture drift that split
// or dropped it would otherwise make the id assertions vacuous.
expect(matched).toHaveLength(1);
return matched[0].ids;
}
/** The behaviour a `reaches-pdg` row has today. */
function assertChainReachesPdg(shape: ReceiverShape): void {
const ids = idsFor(shape.marker);
// Non-empty first: an unresolvable receiver drops EVERY link, so this
// separates "the chained link regressed" from "the whole cell went away".
expect(ids).not.toHaveLength(0);
expect(ids).toEqual(expect.arrayContaining([...shape.links]));
}
describe('PDG calleeIds — chained receiver calls by receiver form (known gap: #2807)', () => {
beforeAll(async () => {
const dir = repos.dir();
fs.mkdirSync(path.join(dir, path.dirname(FIXTURE_PATH)));
fs.writeFileSync(path.join(dir, FIXTURE_PATH), CHAINED_SOURCE);
const result = await runPipelineFromRepo(dir, () => {}, { pdg: true });
const collected: BlockCell[] = [];
result.graph.forEachNode((n) => {
if (n.label !== 'BasicBlock') return;
collected.push({
text: typeof n.properties.text === 'string' ? n.properties.text : '',
ids: splitCalleeIds(n.properties.calleeIds),
});
});
blocks = collected;
}, 180000);
it('every receiver shape contributes exactly one chained-call block', () => {
const counts = Object.fromEntries(
RECEIVER_SHAPES.map((s) => [s.name, blocksFor(s.marker).length]),
);
expect(counts).toEqual(Object.fromEntries(RECEIVER_SHAPES.map((s) => [s.name, 1])));
});
// The `known-gap-empty-cell` rows are deliberately absent here — an `it.fails`
// row over them would be strictly weaker than the exact pin below, since
// `it.fails` is satisfied by ANY throw, including `idsFor`'s own non-vacuity
// guard. Fixture drift that renamed a marker would keep it green while the
// premise had rotted.
for (const shape of RECEIVER_SHAPES.filter((s) => s.resolution === 'reaches-pdg')) {
it(`${shape.name}: every chain link's exact id reaches calleeIds`, () => {
assertChainReachesPdg(shape);
});
}
// Pins the CURRENT broken value, not merely that the chain fails: both known
// gaps emit an EMPTY cell — the first link (`Outer.inner`, a plainly named
// receiver) is gone too. This asserts the gap EXISTS (issue #2807), so closing
// #2807 turns it red BY DESIGN; update it together with the header table and
// the rows' `resolution` rather than loosening it.
it('KNOWN GAP (#2807): an inference-typed receiver empties the WHOLE calleeIds cell', () => {
const gaps = RECEIVER_SHAPES.filter((s) => s.resolution === 'known-gap-empty-cell');
const observed = Object.fromEntries(gaps.map((s) => [s.name, idsFor(s.marker)]));
expect(observed).toEqual({ 'inferred-field': [], 'ctor-assigned-inferred': [] });
});
});
@@ -15,13 +15,13 @@
* for its CSV dir), differing ONLY in `streamPdgEmit`, so streaming is the only
* variable.
*/
import { describe, it, expect, afterAll } from 'vitest';
import { describe, it, expect } from 'vitest';
import fs from 'fs';
import os from 'os';
import path from 'path';
import { runPipelineFromRepo } from '../../../src/core/ingestion/pipeline.js';
import { loadParseCache } from '../../../src/storage/parse-cache.js';
import type { PipelineResult } from '../../../src/types/pipeline.js';
import { createTempDirPool } from '../../helpers/temp-dir-pool.js';
const FIXTURE = path.join(__dirname, 'fixtures', 'pdg-repo');
// A `.vue` SFC importing a `.ts` module: the TS module is PDG-emitted in BOTH
@@ -36,18 +36,10 @@ const PDG_EDGE_TYPES = new Set([
'SANITIZES',
]);
const tmpDirs: string[] = [];
function freshRepo(fixture: string = FIXTURE): string {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gn-pdg-stream-'));
fs.cpSync(fixture, dir, { recursive: true });
tmpDirs.push(dir);
return dir;
}
function freshStorage(): string {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gn-pdg-store-'));
tmpDirs.push(dir);
return dir;
}
const repos = createTempDirPool('gn-pdg-stream-');
const storages = createTempDirPool('gn-pdg-store-');
const freshRepo = (fixture: string = FIXTURE): string => repos.fromFixture(fixture);
const freshStorage = (): string => storages.dir();
function pdgCounts(result: PipelineResult): { basicBlocks: number; pdgEdges: number } {
let basicBlocks = 0;
@@ -62,10 +54,6 @@ function pdgCounts(result: PipelineResult): { basicBlocks: number; pdgEdges: num
}
describe('#2202 — streaming PDG emit end-to-end', () => {
afterAll(() => {
for (const d of tmpDirs) fs.rmSync(d, { recursive: true, force: true });
});
it('streams the PDG layer out of the graph while preserving the emitted set', async () => {
// ── Baseline: --pdg on, streaming OFF (durable cache, same as streamed) ──
const baseStorage = freshStorage();
@@ -1,12 +1,12 @@
import { describe, it, expect, afterAll } from 'vitest';
import { describe, it, expect } from 'vitest';
import fs from 'fs';
import os from 'os';
import path from 'path';
import crypto from 'crypto';
import { runPipelineFromRepo } from '../../../src/core/ingestion/pipeline.js';
import type { PipelineResult } from '../../../src/types/pipeline.js';
import { decodeTaintPath } from '../../../src/core/ingestion/taint/path-codec.js';
import { fixtureTaintTotals } from '../../helpers/taint-fixture.js';
import { createTempDirPool } from '../../helpers/temp-dir-pool.js';
import { isLanguageAvailable } from '../../../src/core/tree-sitter/parser-loader.js';
import { SupportedLanguages } from '../../../src/config/supported-languages.js';
@@ -45,19 +45,10 @@ function counts(result: PipelineResult): {
return { basicBlocks, cfgEdges, reachingDefs, tainted, sanitizes, cdg };
}
const tmpDirs: string[] = [];
function freshRepo(): string {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gn-pdg-'));
fs.cpSync(FIXTURE, dir, { recursive: true });
tmpDirs.push(dir);
return dir;
}
const repos = createTempDirPool('gn-pdg-');
const freshRepo = (): string => repos.fromFixture(FIXTURE);
describe('U7 — end-to-end --pdg pipeline', () => {
afterAll(() => {
for (const d of tmpDirs) fs.rmSync(d, { recursive: true, force: true });
});
it('with --pdg on: emits BasicBlock nodes + CFG edges into the graph', async () => {
const result = await runPipelineFromRepo(freshRepo(), () => {}, { pdg: true });
const { basicBlocks, cfgEdges, reachingDefs } = counts(result);
@@ -288,11 +279,11 @@ const REMAINING_LANGS: ReadonlyArray<{
{ lang: 'Vue', fixture: 'vue-hazards.vue', hazard: 'shouldStop' }, // eventLoop: while(true)
];
const cFamilyTmpDirs: string[] = [];
// Single-file seeding, so this pool uses `dir()` rather than `fromFixture()`.
const langRepos = createTempDirPool('gn-pdg-lang-');
function freshLangRepo(fixture: string): string {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gn-pdg-lang-'));
const dir = langRepos.dir();
fs.copyFileSync(path.join(C_FAMILY_FIXTURES, fixture), path.join(dir, fixture));
cFamilyTmpDirs.push(dir);
return dir;
}
@@ -396,10 +387,6 @@ function cdgSourcedInHazardFunction(result: PipelineResult, hazardMarker: string
}
describe('U7 — C-family worker-mode --pdg pipeline', () => {
afterAll(() => {
for (const d of cFamilyTmpDirs) fs.rmSync(d, { recursive: true, force: true });
});
for (const { lang, fixture, hazard } of C_FAMILY) {
it(`${lang}: --pdg on emits BasicBlock + CFG + REACHING_DEF + CDG (> 0) via the worker`, async () => {
const result = await runPipelineFromRepo(freshLangRepo(fixture), () => {}, WORKER_PDG);
@@ -478,10 +465,6 @@ describe('U7 — C-family worker-mode --pdg pipeline', () => {
});
describe('U7 — remaining languages worker-mode --pdg pipeline (#2195 capstone)', () => {
afterAll(() => {
for (const d of cFamilyTmpDirs) fs.rmSync(d, { recursive: true, force: true });
});
for (const { lang, fixture, hazard, vendored } of REMAINING_LANGS) {
// Vendored grammars (Swift/Kotlin/Dart) may lack a prebuild on the CI
// platform — skip rather than fail when the grammar can't load (#2197 U4).
@@ -338,7 +338,7 @@ describeIfWorkerBuilt('function-local VALUES carry their own identity (#2699 A1)
// The churn this was deferred for is real and was accepted deliberately:
// it re-keys ~14,700 build-time nodes to change ~800 persisted ones,
// because `pruneLocalSymbols` deletes most locals. Hence the paired
// INCREMENTAL_SCHEMA_VERSION / parse-cache SCHEMA_BUMP bumps — without them
// schema-fingerprint changes / parse-cache SCHEMA_BUMP bumps — without them
// a warm cache or an incremental top-up replays the old un-suffixed ids.
//
// Only LOCALS move. The prefix comes from `enclosingCallablePrefix`, which
@@ -0,0 +1,202 @@
/**
* `GroupService.groupSync` reaches `syncGroup` through a DYNAMIC
* `await import('./sync.js')`. A static import would drag the six contract
* extractors — and, through them, the native tree-sitter binding — onto MCP
* server startup, which never syncs; `groupSync` is the module's only consumer.
*
* That import is the one control-flow line the change added, and EVERY
* production `group_sync` call runs it. Mocking `./sync.js` out would prove
* nothing about it: the claims worth pinning are that the specifier still
* RESOLVES and that the destructured `syncGroup` is the real function. So the
* happy path here mocks nothing and drives the real module — mutating the
* specifier (`'./sync-nope.js'`) or the destructured name turns it RED.
*
* Reaching a real `syncGroup` with no indexed repo is what the group.yaml below
* is for: `GITNEXUS_HOME` points at an empty temp home, so the registry is
* empty and every member repo lands in `missingRepos`, while a declared
* manifest link still yields synthetic-UID contracts (the same shape
* `manifest-synthetic-impact.test.ts` covers downstream). Every detector is off,
* so nothing opens a repo graph.
*
* The negative direction is covered too: `sync.js` is replaced with a module
* whose load THROWS, pinning that the failure surfaces as a rejected
* `groupSync()` the MCP dispatch layer can convert into a scoped tool error,
* and that the two guards ahead of the import still answer without ever
* resolving it.
*/
import { afterEach, describe, expect, it, vi } from 'vitest';
import fsp from 'node:fs/promises';
import path from 'node:path';
import { createTempDirPool } from '../../helpers/temp-dir-pool.js';
import { makeGroupToolPort } from '../../unit/group/fixtures.js';
import { GroupService } from '../../../src/core/group/service.js';
import { readContractRegistry } from '../../../src/core/group/storage.js';
import { causeChain } from '../../../src/lib/utils.js';
const tempDirs = createTempDirPool('gn-group-lazy-sync-');
const GROUP_NAME = 'lazy-sync';
const CONTRACT_ID = 'custom::rotateSigningKey';
const LOAD_FAILURE = 'simulated ./sync.js load failure';
/**
* A group whose two members are absent from the registry (so `syncGroup`
* reports them missing instead of opening a graph) but which declares one
* manifest link, the one input a full `syncGroup` turns into contracts without
* an indexed repo. `app/frontend` is the link's `from` with `role: consumer`,
* so `app/backend` is the provider.
*/
async function seedGroup(home: string): Promise<string> {
const groupDir = path.join(home, 'groups', GROUP_NAME);
await fsp.mkdir(groupDir, { recursive: true });
await fsp.writeFile(
path.join(groupDir, 'group.yaml'),
`version: 1
name: ${GROUP_NAME}
description: ""
repos:
app/backend: lazy-sync-backend
app/frontend: lazy-sync-frontend
links:
- from: app/frontend
to: app/backend
type: custom
contract: rotateSigningKey
role: consumer
packages: {}
detect:
http: false
grpc: false
thrift: false
topics: false
shared_libs: false
embedding_fallback: false
includes: false
workspace_deps: false
matching:
bm25_threshold: 0.7
embedding_threshold: 0.65
max_candidates_per_step: 3
`,
'utf8',
);
return groupDir;
}
/**
* Every message in an error's `cause` chain. The module runner reports a failed
* module load through its own error with the original attached as `cause`, and
* exactly where it puts it is a runner detail — flattening the chain keeps the
* assertion about the failure that happened, not about how vitest wraps it.
*/
function errorChainText(err: unknown): string {
// `causeChain` is the repo's single cause-chain traversal — its own doc asks
// callers not to re-roll the loop, because every hand-rolled copy re-decides
// the bound and they disagree. Its default depth is 5; real chains here are
// the runner's wrapper plus the original, so 2.
return [...causeChain(err)].map((link) => link.message).join(' | ');
}
/**
* A `GroupService` from a freshly re-evaluated module graph in which
* `./sync.js` cannot be loaded at all. The re-import is what makes this a
* statement about the LAZY import: a static one would have thrown here, at
* `service.js` load, rather than at the `groupSync` call below.
*/
async function serviceWithUnloadableSync(home: string): Promise<GroupService> {
vi.resetModules();
vi.doMock('../../../src/core/group/sync.js', () => {
throw new Error(LOAD_FAILURE);
});
const { GroupService: FreshGroupService } = await import('../../../src/core/group/service.js');
return new FreshGroupService(makeGroupToolPort(home));
}
describe('GroupService.groupSync — lazy ./sync.js import', () => {
afterEach(() => {
vi.doUnmock('../../../src/core/group/sync.js');
vi.unstubAllEnvs();
vi.resetModules();
});
it('resolves the real ./sync.js and returns the real syncGroup result', async () => {
const home = tempDirs.dir();
vi.stubEnv('GITNEXUS_HOME', home);
const groupDir = await seedGroup(home);
const result = await new GroupService(makeGroupToolPort(home)).groupSync({
name: GROUP_NAME,
});
// Only the REAL syncGroup produces this: two synthetic manifest contracts
// (provider + consumer) and their cross-link, with both members reported
// missing because the temp registry is empty.
expect(result).toMatchObject({
contracts: 2,
crossLinks: 1,
missingRepos: ['app/backend', 'app/frontend'],
});
// `groupDir` reached syncGroup's options too: the registry it wrote there
// carries the contracts the returned counts summarize.
await expect(readContractRegistry(groupDir)).resolves.toMatchObject({
version: 1,
missingRepos: ['app/backend', 'app/frontend'],
contracts: [
{
contractId: CONTRACT_ID,
role: 'provider',
repo: 'app/backend',
symbolUid: `manifest::app/backend::${CONTRACT_ID}`,
meta: { source: 'manifest' },
},
{
contractId: CONTRACT_ID,
role: 'consumer',
repo: 'app/frontend',
symbolUid: `manifest::app/frontend::${CONTRACT_ID}`,
meta: { source: 'manifest' },
},
],
crossLinks: [
{
contractId: CONTRACT_ID,
matchType: 'manifest',
from: { repo: 'app/frontend' },
to: { repo: 'app/backend' },
},
],
});
});
it('surfaces a ./sync.js load failure as a rejected groupSync call', async () => {
const home = tempDirs.dir();
vi.stubEnv('GITNEXUS_HOME', home);
await seedGroup(home);
const service = await serviceWithUnloadableSync(home);
// Awaited inside `groupSync`, so the caller (LocalBackend → MCP dispatch)
// gets a catchable rejection rather than a floating unhandled one. A
// `groupSync` that instead RESOLVED — swallowing the failed import into a
// fake success — reports the sentinel and fails this assertion.
const outcome = await service.groupSync({ name: GROUP_NAME }).then(
() => 'resolved: the failed ./sync.js import did not propagate',
(err: unknown) => errorChainText(err),
);
expect(outcome).toContain(LOAD_FAILURE);
});
it('answers both pre-import guards without resolving ./sync.js', async () => {
const home = tempDirs.dir();
vi.stubEnv('GITNEXUS_HOME', home);
const service = await serviceWithUnloadableSync(home);
await expect(service.groupSync({ name: ' ' })).resolves.toEqual({
error: 'name is required',
});
await expect(service.groupSync({ name: 'never-configured' })).resolves.toEqual({
error: 'Group "never-configured" not found. Run group_list to see configured groups.',
});
});
});
@@ -55,6 +55,18 @@ const SEED = [
`CREATE (rv:Const {id: 'Const:test/registry.test.ts:Registry', name: 'Registry', filePath: 'test/registry.test.ts', startLine: 3, endLine: 3, content: '', description: ''})`,
`CREATE (ru:Function {id: 'Function:src/boot.ts:boot', name: 'boot', filePath: 'src/boot.ts', startLine: 1, endLine: 5, isExported: true, content: '', description: ''})`,
`MATCH (a:Function {id:'Function:src/boot.ts:boot'}), (b:Class {id:'Class:src/registry.ts:Registry'}) CREATE (a)-[:CodeRelation {type:'CALLS', confidence:0.85, reason:'direct', step:0}]->(b)`,
// #2787 review F4 — the two shapes the Class/Interface collapse must now tell
// apart. `Panel`: ONE Class plus its Constructor (the literal #480 case) —
// exactly one candidate carries the label, so collapsing is a correct
// confident answer. `Widget`: TWO Classes in different files plus a same-named
// value binding — more than one candidate carries the label, so there is no
// right answer to pick and the caller must be told.
`CREATE (:Class {id: 'Class:src/panel.ts:Panel', name: 'Panel', filePath: 'src/panel.ts', startLine: 1, endLine: 12, isExported: true, content: '', description: ''})`,
`CREATE (:Constructor {id: 'Constructor:src/panel.ts:Panel', name: 'Panel', filePath: 'src/panel.ts', startLine: 2, endLine: 4, content: '', description: ''})`,
`CREATE (:Class {id: 'Class:src/widgets/alpha.ts:Widget', name: 'Widget', filePath: 'src/widgets/alpha.ts', startLine: 1, endLine: 20, isExported: true, content: '', description: ''})`,
`CREATE (:Class {id: 'Class:src/widgets/beta.ts:Widget', name: 'Widget', filePath: 'src/widgets/beta.ts', startLine: 1, endLine: 20, isExported: true, content: '', description: ''})`,
`CREATE (:Const {id: 'Const:test/widget.test.ts:Widget', name: 'Widget', filePath: 'test/widget.test.ts', startLine: 3, endLine: 3, content: '', description: ''})`,
];
withTestLbugDB(
@@ -151,6 +163,41 @@ withTestLbugDB(
expect(result.impactedCount).toBeGreaterThanOrEqual(1);
});
it('still collapses a lone Class against its same-named Constructor (#480, #2787 review F4)', async () => {
// Non-regression half of the F4 fix. The collapse probe went from
// `LIMIT 1` to `LIMIT 2` and now requires EXACTLY one row — one Class plus
// one Constructor still yields exactly one Class row, so the confident
// answer #480 introduced is unchanged. If this flips to `ambiguous`, the
// uniqueness guard was made too strict and every resolver-backed tool
// loses a previously confident resolution.
const result = await backend.callTool('context', { name: 'Panel' });
expect(result).toMatchObject({ status: 'found' });
expect(result.symbol).toMatchObject({ uid: 'Class:src/panel.ts:Panel', kind: 'Class' });
});
it('refuses to collapse when TWO candidates carry the Class label (#2787 review F4)', async () => {
// Collapsing returns `kind: 'ok'` — the caller never sees the scorer or
// the ambiguity report — and it is the ONLY confident path a bare name can
// take (scoreCandidate tops out at 0.60 without a file_path hint; the
// confident gate needs >= 0.95). With `LIMIT 1` the probe took whichever
// labelled row came back and answered confidently; adding `ORDER BY n.id`
// made that wrong pick REPEATABLE rather than right — `context`/`impact`
// would silently analyse src/widgets/alpha.ts and never mention beta.
// `LIMIT 2` turns the second row into a uniqueness check.
const result = await backend.callTool('context', { name: 'Widget' });
expect(result).toMatchObject({ status: 'ambiguous' });
expect((result.candidates as Array<{ uid: string }>).map((c) => c.uid).sort()).toEqual([
'Class:src/widgets/alpha.ts:Widget',
'Class:src/widgets/beta.ts:Widget',
'Const:test/widget.test.ts:Widget',
]);
// Both classes are offered, so the caller can disambiguate — the file the
// old code silently discarded is the second entry above.
expect(result.totalCandidates).toBe(3);
});
it('reports an undetermined impactedCount for an ambiguous pdg target (#2687)', async () => {
// The pdg branch has no per-candidate fan-out, so it carries no
// maxImpactedCount at all — a numeric zero here is even less correctable.
@@ -205,3 +252,73 @@ withTestLbugDB(
},
},
);
/**
* #2787 — the resolver's LIMIT 20 window must be pinned by ORDER BY.
*
* 25 Functions share one name, seeded in DESCENDING id order. Without an
* ORDER BY, LadybugDB hands back the first 20 it scans (insertion order on a
* fixture this small, an arbitrary subset on a real index) — `collide-z25`
* down to `collide-z06`. With `ORDER BY n.id` it returns the 20 lowest ids,
* `collide-z01` through `collide-z20`. The two sets are disjoint on 10
* elements, so the exact-list assertion below distinguishes them.
*
* `context` is the observable surface, not `impact`: it returns every resolver
* candidate untruncated, while `impact` slices to AMBIGUOUS_MAX_CANDIDATES and
* re-sorts by blast radius, which would hide the window difference.
*/
const COLLIDE_COUNT = 25;
const COLLIDE_IDS = Array.from(
{ length: COLLIDE_COUNT },
(_, i) => `Function:src/z${String(i + 1).padStart(2, '0')}.ts:collide`,
);
withTestLbugDB(
'resolver-window-ordering-2787',
(handle) => {
let backend: LocalBackend;
beforeAll(() => {
backend = (handle as any)._backend;
});
it('returns the 20 lowest-id candidates, not an arbitrary window (#2787)', async () => {
const result = await backend.callTool('context', { name: 'collide' });
expect(result.status).toBe('ambiguous');
// 25 is the TRUE match count (a COUNT alongside the window); 20 is the
// window. Reporting the window here claimed the cap was the total.
expect(result).toMatchObject({ totalCandidates: COLLIDE_COUNT, candidatesTruncated: true });
expect(result.message).toContain(`Found ${COLLIDE_COUNT} symbols`);
expect(result.message).toContain('showing 20');
expect((result.candidates as Array<{ uid: string }>).map((c) => c.uid)).toEqual(
COLLIDE_IDS.slice(0, 20),
);
});
},
{
// Descending: the highest id is inserted first, so an unordered scan that
// follows insertion order returns exactly the wrong half.
seed: [...COLLIDE_IDS]
.reverse()
.map(
(id, i) =>
`CREATE (:Function {id: '${id}', name: 'collide', filePath: 'src/z${String(COLLIDE_COUNT - i).padStart(2, '0')}.ts', startLine: 1, endLine: 3, isExported: true, content: '', description: ''})`,
),
poolAdapter: true,
afterSetup: async (handle) => {
vi.mocked(listRegisteredRepos).mockResolvedValue([
{
name: 'test-repo',
path: '/test/repo',
storagePath: handle.tmpHandle.dbPath,
indexedAt: new Date().toISOString(),
lastCommit: 'abc123',
stats: { files: COLLIDE_COUNT, nodes: COLLIDE_COUNT, communities: 0, processes: 0 },
},
]);
const backend = new LocalBackend();
await backend.init();
(handle as any)._backend = backend;
},
},
);
@@ -514,23 +514,38 @@ withTestLbugDB(
expect(uids).toContain('Tool:alpha');
});
it('ranks the kind-matching candidate first when kind is supplied (the --kind flag path)', async () => {
// 'alpha' is both a Function (func:alpha) and a Tool (Tool:alpha).
// kind only adds +0.20 in scoreCandidate, so 0.50 + 0.20 = 0.70 stays
// below the 0.95 confident-resolution threshold — the response is still
// ambiguous. What kind buys is ranking: the Function is promoted above
// the non-matching Tool. This exercises the scoreCandidate kind branch
// against a real DB rather than only through the mocked CLI unit test.
it('retries unfiltered and still ranks by kind when the hint matches no id prefix (the --kind flag path, #2787 review F5)', async () => {
// WHY THIS FIXTURE TAKES THE FALLBACK, NOT THE FILTER (#2787 review F5):
// `kind` is no longer a pure scoring term — it filters with
// `n.id STARTS WITH 'Function:'`, which relies on the production node-id
// convention `Label:filePath:qualifiedName`. This fixture predates that
// convention: its Function node is `func:alpha` (lowercase, in
// test/fixtures/local-backend-seed.ts), so NO row in this DB satisfies
// the `Function:` prefix and the filtered window comes back EMPTY.
// `kind` is a free-form string on the tool schema, so an empty filtered
// window must not degrade a real name to `not_found`: the resolver
// retries UNFILTERED and the hint reverts to scoreCandidate's +0.20
// ranking term. That fallback — not the filter — is what this pins.
// The filtering path is covered against production-shaped ids by the
// `resolver-kind-filter-2787` suite below.
const result = await backend.callTool('impact', {
target: 'alpha',
kind: 'Function',
direction: 'upstream',
});
expect(result).not.toHaveProperty('error');
// The unfiltered retry ran: BOTH candidates are back, including the
// Tool the filter would have excluded. Without the fallback this is
// `{ error: "Symbol 'alpha' not found" }`.
expect(result.status).toBe('ambiguous');
const candidates = result.candidates ?? [];
expect(candidates[0]?.uid).toBe('func:alpha');
expect(candidates[0]?.kind).toBe('Function');
expect(candidates.map((c: any) => c.uid).sort()).toEqual(['Tool:alpha', 'func:alpha']);
// …and the hint still ranks, exactly as it did before the filter
// existed: 0.50 + 0.20 = 0.70, below the 0.95 confident gate, so the
// response stays ambiguous with the Function promoted.
expect(candidates[0]).toMatchObject({ uid: 'func:alpha', kind: 'Function' });
const tool = candidates.find((c: any) => c.uid === 'Tool:alpha');
expect(candidates[0]?.score).toBeGreaterThan(tool?.score);
});
@@ -625,3 +640,203 @@ withTestLbugDB(
},
},
);
// ─── resolver `kind` hint FILTERS (#2787 review F5) ──────────────────────
// See `resolveSymbolCandidates` (#2787 F5) for why the id order is label-major
// and why the hint is therefore a WHERE-clause filter on BOTH the window and
// its COUNT rather than a scoring term.
//
// This suite is deliberately seeded with PRODUCTION-shaped ids, unlike the
// shared `local-backend-seed` fixture whose Function is `func:alpha` — that
// legacy shape exercises the unfiltered-retry fallback instead (see the
// "retries unfiltered…" test above).
const KIND_RUNNER_METHOD_ID = 'Method:src/pipeline.ts:Runner.run';
const KIND_RUN_FUNCTION_IDS = ['Function:src/cli/exec.ts:run', 'Function:src/pipeline.ts:run'];
withTestLbugDB(
'resolver-kind-filter-2787',
(handle) => {
describe('resolver kind hint filters rather than scores (#2787 review F5)', () => {
let backend: LocalBackend;
beforeAll(() => {
const ext = handle as typeof handle & { _backend?: LocalBackend };
if (!ext._backend) {
throw new Error('LocalBackend not initialized — afterSetup did not attach _backend');
}
backend = ext._backend;
});
it('narrows a 3-way name collision to a confident resolution', async () => {
// Three symbols are named `run`: two Functions and one Method. As a
// scoring term the Method could only reach 0.50 + 0.20 = 0.70, far
// below the 0.95 confident gate, so this came back `ambiguous` and the
// caller had to round-trip for a uid. As a filter the window contains
// exactly one row, which resolves outright.
const result = await backend.callTool('impact', {
target: 'run',
kind: 'Method',
direction: 'upstream',
});
expect(result).not.toHaveProperty('error');
expect(result.status).not.toBe('ambiguous');
expect(result.target).toMatchObject({ id: KIND_RUNNER_METHOD_ID, name: 'run' });
// The BFS still ran on the filtered pick: `boot` calls Runner.run.
expect(result.impactedCount).toBeGreaterThanOrEqual(1);
});
it('excludes non-matching kinds from the candidate set AND from the match count', async () => {
// The load-bearing difference between filtering and scoring: with
// kind:'Function' the Method must be ABSENT, not merely ranked last —
// and `totalCandidates` (the COUNT leg) must agree with the page, or a
// filtered window ships alongside the unfiltered population.
const result = await backend.callTool('context', { name: 'run', kind: 'Function' });
expect(result).toMatchObject({ status: 'ambiguous', totalCandidates: 2 });
expect((result.candidates as Array<{ uid: string }>).map((c) => c.uid).sort()).toEqual(
KIND_RUN_FUNCTION_IDS,
);
expect(result).not.toHaveProperty('totalIsLowerBound');
});
});
},
{
seed: [
`CREATE (:Function {id: 'Function:src/pipeline.ts:run', name: 'run', filePath: 'src/pipeline.ts', startLine: 1, endLine: 9, isExported: true, content: 'export function run() {}', description: 'pipeline entry'})`,
`CREATE (:Function {id: 'Function:src/cli/exec.ts:run', name: 'run', filePath: 'src/cli/exec.ts', startLine: 1, endLine: 9, isExported: true, content: 'export function run() {}', description: 'cli entry'})`,
`CREATE (:Method {id: '${KIND_RUNNER_METHOD_ID}', name: 'run', filePath: 'src/pipeline.ts', startLine: 20, endLine: 30, isExported: false, content: 'run() {}', description: 'Runner.run'})`,
`CREATE (:Function {id: 'Function:src/main.ts:boot', name: 'boot', filePath: 'src/main.ts', startLine: 1, endLine: 5, isExported: true, content: 'function boot() {}', description: 'caller of Runner.run'})`,
`MATCH (a:Function), (b:Method) WHERE a.id = 'Function:src/main.ts:boot' AND b.id = '${KIND_RUNNER_METHOD_ID}'
CREATE (a)-[:CodeRelation {type: 'CALLS', confidence: 0.9, reason: 'direct', step: 0}]->(b)`,
],
poolAdapter: true,
afterSetup: async (handle) => {
vi.mocked(listRegisteredRepos).mockResolvedValue([
{
name: 'kind-filter-repo',
path: '/kind/repo',
storagePath: handle.tmpHandle.dbPath,
indexedAt: new Date().toISOString(),
lastCommit: 'abc123',
stats: { files: 4, nodes: 4, communities: 0, processes: 0 },
},
]);
const backend = new LocalBackend();
await backend.init();
(handle as any)._backend = backend;
},
},
);
// ─── context ref window must SPREAD across categories (#2787 review F1) ──
// See the incoming-ref window in `_contextImpl` (#2787 F1) for why the 30-row
// page is keyed `ORDER BY uid, relType` and not category-major.
//
// Shape below: 40 incoming refs on one Method — 35 CALLS, 3 ACCESSES, 1 USES,
// 1 HAS_METHOD (the owning class, which is how `context` names the owner).
// relType-major → ACCESSES(3) + CALLS(27); HAS_METHOD and USES gone.
// uid-major → HAS_METHOD(1) + ACCESSES(3) + USES(1) + CALLS(25).
// The caller ids are spread through the id space on purpose (`a1` in f05, `u1`
// in f10, `a2` in f15, `a3` in f25), so the rare categories are NOT stacked at
// the front of the uid order — they survive because the key interleaves them,
// not because they were placed first.
const REF_TARGET_ID = 'Method:src/owner.ts:Owner.handle';
const REF_OWNER_ID = 'Class:src/owner.ts:Owner';
const REF_ACCESSES_IDS = [
'Function:src/f05.ts:a1',
'Function:src/f15.ts:a2',
'Function:src/f25.ts:a3',
];
const REF_USES_ID = 'Function:src/f10.ts:u1';
const refCaller = (id: string, name: string, filePath: string): string =>
`CREATE (:Function {id: '${id}', name: '${name}', filePath: '${filePath}', startLine: 1, endLine: 3, isExported: true, content: '', description: ''})`;
const refEdge = (fromLabel: string, fromId: string, relType: string): string =>
`MATCH (a:${fromLabel}), (b:Method) WHERE a.id = '${fromId}' AND b.id = '${REF_TARGET_ID}'
CREATE (a)-[:CodeRelation {type: '${relType}', confidence: 0.9, reason: 'direct', step: 0}]->(b)`;
const REF_CALLS_IDS = Array.from({ length: 35 }, (_, i) => {
const nn = String(i + 1).padStart(2, '0');
return `Function:src/f${nn}.ts:c${nn}`;
});
withTestLbugDB(
'context-ref-window-2787',
(handle) => {
describe('context incoming-ref window spreads across categories (#2787 review F1)', () => {
let backend: LocalBackend;
beforeAll(() => {
const ext = handle as typeof handle & { _backend?: LocalBackend };
if (!ext._backend) {
throw new Error('LocalBackend not initialized — afterSetup did not attach _backend');
}
backend = ext._backend;
});
it('keeps single-edge categories that a category-major ORDER BY starved', async () => {
const result = await backend.callTool('context', { uid: REF_TARGET_ID });
expect(result).not.toHaveProperty('error');
expect(result.status).toBe('found');
// The rare categories are present at all — this is the assertion the
// pre-fix key fails: HAS_METHOD and USES sort after CALLS, whose 35 rows
// overflow the window on their own.
expect(Object.keys(result.incoming).sort()).toEqual([
'accesses',
'calls',
'has_method',
'uses',
]);
expect(result.incoming.has_method.map((r: { uid: string }) => r.uid)).toEqual([
REF_OWNER_ID,
]);
expect(result.incoming.uses.map((r: { uid: string }) => r.uid)).toEqual([REF_USES_ID]);
expect(result.incoming.accesses.map((r: { uid: string }) => r.uid)).toEqual(
REF_ACCESSES_IDS,
);
// …and the window is still exactly 30 rows: the fix REDISTRIBUTES the
// page, it does not widen it. The dominant category gives up the 5 slots
// the starved ones need.
expect(result.incoming.calls).toHaveLength(25);
expect(Object.values(result.incoming).flat()).toHaveLength(30);
});
});
},
{
seed: [
`CREATE (:Method {id: '${REF_TARGET_ID}', name: 'handle', filePath: 'src/owner.ts', startLine: 10, endLine: 20, isExported: false, content: '', description: ''})`,
`CREATE (:Class {id: '${REF_OWNER_ID}', name: 'Owner', filePath: 'src/owner.ts', startLine: 1, endLine: 40, isExported: true, content: '', description: ''})`,
...REF_CALLS_IDS.map((id, i) => {
const nn = String(i + 1).padStart(2, '0');
return refCaller(id, `c${nn}`, `src/f${nn}.ts`);
}),
...REF_ACCESSES_IDS.map((id, i) => refCaller(id, `a${i + 1}`, id.split(':')[1])),
refCaller(REF_USES_ID, 'u1', 'src/f10.ts'),
...REF_CALLS_IDS.map((id) => refEdge('Function', id, 'CALLS')),
...REF_ACCESSES_IDS.map((id) => refEdge('Function', id, 'ACCESSES')),
refEdge('Function', REF_USES_ID, 'USES'),
refEdge('Class', REF_OWNER_ID, 'HAS_METHOD'),
],
poolAdapter: true,
afterSetup: async (handle) => {
vi.mocked(listRegisteredRepos).mockResolvedValue([
{
name: 'ref-window-repo',
path: '/ref/repo',
storagePath: handle.tmpHandle.dbPath,
indexedAt: new Date().toISOString(),
lastCommit: 'abc123',
stats: { files: 41, nodes: 41, communities: 0, processes: 0 },
},
]);
const backend = new LocalBackend();
await backend.init();
(handle as any)._backend = backend;
},
},
);
@@ -10,99 +10,69 @@
* level. The native binding's init can write to raw stdout in that pre-sentinel
* window and corrupt the JSON-RPC frame stream.
*
* This test locks in the fix: spawn a child Node process, import the built
* `dist/cli/mcp.js` (without invoking `mcpCommand`), and assert that
* `@ladybugdb/core` is NOT in the loaded-module set. The assertion is
* evidence-based — it checks Node's CJS module cache, which is global per
* process and tracks every native/CJS module loaded by either ESM or CJS
* importers.
* This test locks in the fix: import the built `dist/cli/mcp.js` in a child
* process (without invoking `mcpCommand`) and assert that `@ladybugdb/core` is
* NOT in the loaded-module set.
*
* The probe is `test/helpers/module-load-probe.ts`. This file used to carry its
* own copy that diffed Node's CJS module cache and nothing else. That was
* enough for the `@ladybugdb/core` headline — a native CJS module always
* surfaces in `require.cache` — but it was structurally BLIND to the ESM
* `dist/**` graph it was walking, which is where the static imports it is
* policing actually live, and it had no non-vacuity guard at all: a `cli/mcp.js`
* severed from its own imports produced an empty cache diff and passed green.
* The shared probe adds the ESM channel and REQUIRES an anchor, so "nothing
* forbidden loaded" now means something. It also spawns ONCE for the two
* assertions below, which used to pay for two separate probes of one target.
*
* `dist/mcp/stdio-context.js` is the anchor because it is the entry's only
* remaining first-party static import — the whole point of the fix — so its
* disappearance is exactly the refactor that would make both assertions vacuous.
*
* Characterization-first: this test was written before the fix landed and
* MUST fail against the pre-fix code. Run against the parent of the U1
* commit to verify the regression signal works.
*/
import { describe, it, expect } from 'vitest';
import { spawnSync } from 'node:child_process';
import path from 'node:path';
import fs from 'node:fs';
import { fileURLToPath, pathToFileURL } from 'node:url';
const __dirname = path.dirname(fileURLToPath(import.meta.url));
const REPO_ROOT = path.resolve(__dirname, '..', '..', '..');
const DIST_MCP = path.join(REPO_ROOT, 'dist', 'cli', 'mcp.js');
const DIST_MCP_URL = pathToFileURL(DIST_MCP).href;
const PROBE = `
import { createRequire } from 'node:module';
const req = createRequire(import.meta.url);
const before = new Set(Object.keys(req.cache));
await import(process.env.PROBE_TARGET);
const after = new Set(Object.keys(req.cache));
const newlyLoaded = [...after].filter((k) => !before.has(k));
process.stdout.write(JSON.stringify(newlyLoaded));
`;
import { describe, it, expect, beforeAll } from 'vitest';
import { probeModuleLoad, type ModuleLoadProbe } from '../../helpers/module-load-probe.js';
describe('MCP CLI static-import closure', () => {
it('does not load @ladybugdb/core when cli/mcp.js is imported (without invoking mcpCommand)', () => {
if (!fs.existsSync(DIST_MCP)) {
throw new Error(
`dist/cli/mcp.js missing — run \`npm run build\` first (or \`npm run test:integration\` which builds via pretest:integration).`,
);
}
let probe: ModuleLoadProbe;
const result = spawnSync(process.execPath, ['--input-type=module', '-e', PROBE], {
cwd: REPO_ROOT,
env: { ...process.env, PROBE_TARGET: DIST_MCP_URL, NODE_OPTIONS: '' },
timeout: 30_000,
encoding: 'utf8',
beforeAll(async () => {
probe = await probeModuleLoad({
entry: 'cli/mcp.js',
anchor: 'dist/mcp/stdio-context.js',
// Observed on Node 22.18 against a clean build: 4 modules. This closure is
// deliberately leaf-only, so the floor is necessarily tight — the anchor
// above is the load-bearing non-vacuity guard here.
minModules: 3,
});
}, 90_000);
if (result.status !== 0) {
throw new Error(
`probe failed (status ${result.status}):\nstderr:\n${result.stderr}\nstdout:\n${result.stdout}`,
);
}
const newlyLoaded = JSON.parse(result.stdout) as string[];
it('does not load @ladybugdb/core when cli/mcp.js is imported (without invoking mcpCommand)', () => {
// The headline assertion: @ladybugdb/core (a native CJS module) must not
// be loaded by the static-import closure of cli/mcp.js. If it is, the
// pre-sentinel stdout window the prior fix tried to close is still open.
const ladybugLoaded = newlyLoaded.filter((p) => /@ladybugdb[\\/]core/.test(p));
const ladybugLoaded = probe.matching(/@ladybugdb[\\/]core/);
expect(
ladybugLoaded,
`@ladybugdb/core was loaded at cli/mcp.js static-import time. ` +
`mcpCommand cannot install the stdout sentinel before native init runs. ` +
`Offending paths:\n${ladybugLoaded.join('\n')}\n\n` +
`Full newly-loaded set (${newlyLoaded.length} entries):\n${newlyLoaded.join('\n')}`,
`Full loaded set (${probe.modules.length} entries):\n${probe.modules.join('\n')}`,
).toEqual([]);
});
it('does not load any tree-sitter native binding (sanity check on grammar imports)', () => {
if (!fs.existsSync(DIST_MCP)) {
throw new Error(`dist/cli/mcp.js missing — run \`npm run build\` first.`);
}
const result = spawnSync(process.execPath, ['--input-type=module', '-e', PROBE], {
cwd: REPO_ROOT,
env: { ...process.env, PROBE_TARGET: DIST_MCP_URL, NODE_OPTIONS: '' },
timeout: 30_000,
encoding: 'utf8',
});
if (result.status !== 0) {
throw new Error(`probe failed: ${result.stderr}`);
}
const newlyLoaded = JSON.parse(result.stdout) as string[];
// No tree-sitter parser should load at cli/mcp.js static-import time.
// The analyze path is the only caller of warnMissingOptionalGrammars
// (which require()s each grammar); cli/mcp.ts itself does not invoke
// it, and its static-import closure is leaf-only — so importing
// dist/cli/mcp.js without invoking mcpCommand must not trigger any
// native grammar binding load.
const treeSitterNative = newlyLoaded.filter((p) => /tree-sitter-[a-z]+[\\/]build/.test(p));
const treeSitterNative = probe.matching(/tree-sitter-[a-z]+[\\/]build/);
expect(
treeSitterNative,
`tree-sitter native bindings loaded at cli/mcp.js static-import time:\n${treeSitterNative.join('\n')}`,
@@ -0,0 +1,291 @@
/**
* MCP startup must not load the analyze-only language provider registry (#2802).
*
* `mcp/local/pdg-impact.ts` once imported `core/ingestion/languages/index.ts`
* for a single extension→language lookup. That edge pulled all 16 providers,
* their extractors, and the tree-sitter native binding into every MCP server
* start: ~226 extra modules and ~130 ms, for a server that never analyzes
* anything. The finding was discovered and lost once already (during #2793)
* before #2802 re-derived it, so it gets a guard rather than a comment.
*
* The guard is a REAL MODULE-LOAD PROBE, not a source-level import walk. A
* previous regex-based version of this test (`test/unit/mcp-startup-import-
* closure.test.ts`) was defeated four separate ways: it walked from
* `local-backend.ts` instead of the actual server entry, it was structurally
* blind to eager top-level `await import(...)`, its type-only-import stripper
* lazily matched across a 16 kB window of `pdg-impact.ts` (the terminating
* `from "…"` lived inside a string literal), and its comment stripper treated
* `/*` inside a string literal as a comment opener.
*
* The probe itself now lives in `test/helpers/module-load-probe.ts`, shared with
* `test/integration/mcp/import-closure.test.ts` and
* `test/integration/optional-grammars/registry-import-closure.test.ts` — it
* spawns a child Node process, imports a built `dist/` entry, and reports every
* module the loader actually pulled in. It cannot be fooled by import syntax, a
* stale entry point, or regex drift: whatever Node evaluates, the probe sees.
* The FORBIDDEN set and its remedy stay here, because they are specific to
* #2802.
*
* Coverage note: `dist/mcp/server.js` is the entry that must be protected — it
* is what `mcpCommand` dynamically imports and what actually serves MCP.
* `dist/cli/mcp.js` is asserted too (it is the process entry, and its
* deliberately leaf-only static closure is pinned separately by
* `import-closure.test.ts`), as is `dist/mcp/local/local-backend.js` — the
* module whose import graph #2802 actually changed, and
* `dist/mcp/http-transport.js`, which is the OTHER startup entry: `gitnexus mcp
* --http` reaches it through its own `await import(...)` in `mcpCommand`, not
* through `server.js`, so nothing about `server.js` staying clean constrains it.
*
* `local-backend.js`'s closure is TODAY a strict subset of `server.js`'s (156
* of 380 modules, none of them absent from the server's), so it cannot surface
* an offender the server probe would miss. It is kept anyway, for two reasons
* that survive that measurement. The subset relation is an observation about
* the current graph and nothing enforces it: the day `server.js` stops reaching
* the local backend eagerly (remote-only default, lazy backend selection), the
* server probe's anchor — `dist/mcp/resources.js` — keeps passing while the
* module #2802 actually changed goes unobserved. Its own entry pins
* `dist/mcp/local/pdg-impact.js` as an anchor, which is coverage the server
* entry does not and cannot provide. And since the probes run concurrently, the
* marginal wall-clock cost is ~0: it finishes inside the server probe's window.
*
* Lazy `await import(...)` inside a function body remains the sanctioned escape
* hatch: it does not run at startup, so the probe does not see it. A top-level
* `await import(...)` DOES run at module evaluation, and the probe reports it —
* which is the point.
*/
import { describe, it, expect, beforeAll } from 'vitest';
import {
anchorsOf,
probeModuleLoads,
type ModuleLoadProbes,
type ModuleLoadRequest,
} from '../../helpers/module-load-probe.js';
/** Modules under this directory are the analyze-only provider registry. */
const FORBIDDEN_RE = /(^|\/)core\/ingestion\/languages\//;
/**
* The group contract extractors, and the native parser binding they reach.
*
* Same defect class as #2802, found immediately after it: `core/group/service.ts`
* statically imported `./sync.js`, which pulls all six contract extractors, five
* of which statically import `tree-sitter`. Only `group_sync` ever needs them —
* the other seven group tools do not — so a static import put the whole parser
* stack on every MCP server start. Measured cost of that one edge on a native
* filesystem: `dist/mcp/server.js` 521 ms -> 133 ms, `local-backend.js` 453 ms
* -> 66 ms.
*
* Matching the parser by its package prefix rather than a bare substring so a
* source file that merely mentions the word cannot satisfy or trip this.
*
* The separator is `[\\/]`, matching the sibling probes' `ANY_GRAMMAR_RE` and
* `OPTIONAL_GRAMMAR_RE`, NOT a bare `/`. Native bindings reach the probe through
* the `require.cache` channel as absolute paths, and `toRepoRelativePosix` only
* POSIX-normalises paths INSIDE the repo root — a hoisted `node_modules` renders
* verbatim, so on Windows this is `…\node_modules\tree-sitter\…` and a
* forward-slash-only pattern silently matches nothing. This file now runs on the
* Windows matrix, where that would have made the parser half of the assertion
* vacuous. The `core/group/extractors/` half is first-party `dist/**`, always
* in-repo and therefore already normalised.
*/
const FORBIDDEN_GROUP_RE = /(^|\/)core\/group\/extractors\/|[\\/]node_modules[\\/]tree-sitter/;
/**
* The chain the group policy above polices, and therefore ITS non-vacuity
* anchor. `core/group/service.js` is the module that statically imported
* `./sync.js` and dragged the extractors in; the fix made that edge lazy. An
* anchor on some other chain (`mcp/resources.js`, `mcp/local/pdg-impact.js`)
* plus the module-count floor both stay green when `local-backend →
* core/group/service` is severed — the obvious next lazy-load step — and the
* group assertion would then be vacuous on every row while the file still
* reported all-pass. Anchors are per-POLICY, not per-entry; see
* `test/helpers/module-load-probe.ts`.
*/
const GROUP_ANCHOR = 'dist/core/group/service.js';
/**
* Third instance of the same defect class, and the one this file could not see.
*
* `pdg-impact.ts` imported two format constants from `core/ingestion/cfg/emit.ts`.
* ESM evaluates a module to import any binding from it, so those two strings
* pulled the whole analyze-only CFG closure — `emit`, `reaching-defs`,
* `reaching-defs-graph`, `control-dependence`, `post-dominators`,
* `synthetic-escape`, `call-site-harvest` — into every MCP start. The constants
* moved to the leaf `cfg/callee-cell-format.ts`, but `emit.ts` still RE-EXPORTS
* them, so pointing the import back at `emit.js` typechecks identically and
* restores all seven modules. Neither existing policy matches
* `core/ingestion/cfg/`, so nothing was stopping that.
*
* Allowlist rather than a denylist of the seven: the failure mode is a module
* nobody has thought of yet, and a denylist only ever names the regressions
* already suffered. Everything here is a genuine LEAF — zero imports — which is
* why it can sit on the startup path at all; that is a real convention in
* `core/ingestion` (each file's header calls itself the one shared codec), and
* this is the only thing enforcing it.
*/
const CFG_ANCHOR = 'dist/core/ingestion/cfg/callee-cell-format.js';
const INGESTION_CFG_RE = /(^|\/)core\/ingestion\/cfg\//;
const CFG_LEAVES_ALLOWED: ReadonlySet<string> = new Set([
CFG_ANCHOR,
'dist/core/ingestion/cfg/reaching-def-reason-codec.js',
]);
// Observed on Node 22.18 against a clean build at the tip of this branch:
// server.js 380 distinct modules, local-backend.js 156, cli/mcp.js 4. Treat
// these as a snapshot, not a contract — they moved twice inside this branch
// alone (the group-extractor and cfg/emit closures each took ~100 and ~7 out),
// and only the FLOORS below are asserted. The floors sit well under the
// observed counts so normal dependency churn doesn't trip them, while a probe
// that silently loaded nothing still fails.
//
// `mcp/http-transport.js` is the largest startup entry — measured 516 modules,
// +136 over server.js for express, cors and the SDK's Streamable-HTTP/SSE
// transports. It is what `gitnexus mcp --http` starts and the hosted-deploy
// path, and `mcpCommand` imports it directly, not via `server.js`; because that
// edge runs one way only, a static import added inside `http-transport.ts`
// would reinstate #2802 on the HTTP path with every other row here green. The
// probes run concurrently, so its marginal wall-clock cost is ~0 — it finishes
// alongside the others rather than after them.
//
// `mcp/resources.js` (measured 57 modules, 0 offenders) and `mcp/staleness.js`
// (53, 0) are named in the #2802 write-up but deliberately get NO rows: both
// are eagerly inside the closures of `server.js` and `http-transport.js`
// (each appears in both probes' module lists), so any offender they acquired
// surfaces on those rows already. They would earn rows only if something made
// them reachable other than eagerly-from-the-server.
const ENTRIES = [
{
entry: 'mcp/server.js',
anchor: ['dist/mcp/resources.js', GROUP_ANCHOR, CFG_ANCHOR],
minModules: 100,
},
{
entry: 'mcp/http-transport.js',
anchor: ['dist/mcp/server.js', GROUP_ANCHOR, CFG_ANCHOR],
minModules: 100,
},
{
entry: 'mcp/local/local-backend.js',
anchor: ['dist/mcp/local/pdg-impact.js', GROUP_ANCHOR, CFG_ANCHOR],
minModules: 50,
},
// The one row with a single anchor, because it is subject to ONE policy. Its
// whole design is a 4-module leaf closure that reaches nothing first-party
// beyond `stdio-context → stdio-capture`, so it can never reach
// `core/group/service.js` and cannot be given the group anchor honestly. It
// is excluded from the group policy below for that reason: a row that cannot
// fail for any reason related to the policy it is listed under is exactly the
// vacuity this file's probe exists to prevent. Its leaf-only closure is
// pinned exhaustively by `import-closure.test.ts` instead.
{ entry: 'cli/mcp.js', anchor: 'dist/mcp/stdio-context.js', minModules: 3 },
] as const satisfies readonly ModuleLoadRequest[];
/**
* The rows the group policy applies to — DERIVED from the anchors, not
* hand-listed beside them, so an entry cannot join that policy without
* carrying the anchor that keeps it able to fail.
*/
const GROUP_POLICY_ENTRIES = ENTRIES.filter((request) =>
anchorsOf(request.anchor).includes(GROUP_ANCHOR),
).map((request) => request.entry);
/** Same derivation for the CFG-leaf policy. */
const CFG_POLICY_ENTRIES = ENTRIES.filter((request) =>
anchorsOf(request.anchor).includes(CFG_ANCHOR),
).map((request) => request.entry);
describe('MCP startup module-load closure (#2802)', () => {
let probes: ModuleLoadProbes;
// Every entry is probed CONCURRENTLY here, not one per test: the probes are
// independent child processes and each pays a full Node start, so running
// them in parallel cuts this file's wall clock by roughly 60% and makes each
// additional entry near-free. The helper labels every failure with its entry
// and enforces each entry's anchors and module floor, so the `it` bodies
// below are pure policy assertions.
beforeAll(async () => {
probes = await probeModuleLoads(ENTRIES);
}, 90_000);
// `%s` over the bare entries, not `$entry` over the request objects: vitest
// quotes an interpolated object property, and `importing dist/'mcp/server.js'`
// reads like a typo in CI output.
it.each(ENTRIES.map((request) => request.entry))(
'importing dist/%s loads no language provider module',
(entry) => {
const probe = probes.get(entry);
const offenders = probe.matching(FORBIDDEN_RE);
// Headline assertion: named chains, not a bare boolean, so whoever
// reintroduces the edge sees exactly which modules did it.
expect(
offenders,
`${probe.label} eagerly loads the analyze-only language provider registry. ` +
`MCP startup never analyzes anything — route the lookup through a lazy ` +
`\`await import(...)\` inside the function that needs it (see #2802). ` +
`Offending modules:\n${offenders.join('\n')}`,
).toEqual([]);
},
);
// Pins the two derivations above. The risk each guards is a policy going
// SILENT, not its exact membership: drop an anchor from every entry and the
// derived list empties, so the `it.each` registers zero cases and the whole
// policy disappears without one red test. Asserting non-emptiness catches
// exactly that; asserting the literal list would reinstate, one layer down,
// the hand-maintained list the derivation exists to remove — every entry
// added or removed would then need editing in two places.
//
// `cli/mcp.js` is pinned OUT of both policies deliberately. It is a 4-module
// leaf closure that cannot reach either policed chain, so listing it would
// give each policy a row that cannot fail — the vacuity this file exists to
// prevent. That exclusion is a real property, so it is asserted rather than
// left to the comment above.
it.each([
['group', GROUP_POLICY_ENTRIES],
['cfg-leaf', CFG_POLICY_ENTRIES],
])('the %s policy runs over a non-empty entry set that excludes cli/mcp.js', (_name, entries) => {
expect(entries.length).toBeGreaterThan(0);
expect(entries).not.toContain('cli/mcp.js');
});
// The #2802 defect class, third instance: an analyze-only closure reached
// through a constant. Allowlist, not denylist — see CFG_LEAVES_ALLOWED.
it.each(CFG_POLICY_ENTRIES)(
'importing dist/%s loads no non-leaf core/ingestion/cfg module',
(entry) => {
const probe = probes.get(entry);
const offenders = probe
.matching(INGESTION_CFG_RE)
.filter((module) => !CFG_LEAVES_ALLOWED.has(module));
expect(
offenders,
`${probe.label} eagerly loads analyze-only CFG modules. ESM evaluates a ` +
`module to import ANY binding from it, so importing a constant from ` +
`\`cfg/emit.js\` drags its whole closure onto startup — take format ` +
`constants from the leaf \`cfg/callee-cell-format.js\` instead, and add ` +
`a new module here only if it genuinely imports nothing (see #2802). ` +
`Offending modules:\n${offenders.join('\n')}`,
).toEqual([]);
},
);
it.each(GROUP_POLICY_ENTRIES)(
'importing dist/%s loads no group contract extractor or native parser',
(entry) => {
const probe = probes.get(entry);
const offenders = probe.matching(FORBIDDEN_GROUP_RE);
expect(
offenders,
`${probe.label} eagerly loads the group contract extractors and/or the ` +
`native tree-sitter binding. Only \`group_sync\` needs them, and MCP ` +
`startup never syncs — keep \`core/group/sync.js\` behind the lazy ` +
`\`await import(...)\` in \`GroupService.groupSync\`. ` +
`Offending modules:\n${offenders.join('\n')}`,
).toEqual([]);
},
);
});
@@ -17,114 +17,79 @@
* to be an always-present npm dependency.)
*
* This test locks the fix in WITHOUT needing to simulate a missing grammar:
* spawn a child Node process, import the built scope-resolution `registry.js`
* (the crash-chain root), and assert no OPTIONAL tree-sitter binding
* (swift/dart/kotlin) appears in the module cache. Pre-fix the static imports
* loaded those bindings at import time (this assertion fails); post-fix they
* are lazy (it passes). Required grammars (python/typescript/...) still load
* eagerly via their own `query.ts` — that is expected and NOT asserted against.
* import the built scope-resolution `registry.js` (the crash-chain root) in a
* child process and assert no OPTIONAL tree-sitter binding (swift/dart/kotlin/c)
* appears in the loaded-module set. Pre-fix the static imports loaded those
* bindings at import time (this assertion fails); post-fix they are lazy (it
* passes). Required grammars (python/typescript/...) still load eagerly via
* their own `query.ts` — that is expected and NOT asserted against.
*
* The probe is `test/helpers/module-load-probe.ts`. This file used to carry its
* own copy that diffed Node's CJS module cache and nothing else. That is the
* right channel for the headline — a grammar binding is native CJS and always
* surfaces there — but it was structurally BLIND to the ESM `dist/**` graph the
* registry actually is, which is why its non-vacuity guard had to be indirect
* ("at least one REQUIRED binding loaded"). The shared probe adds the ESM
* channel, so this file can now anchor DIRECTLY on
* `dist/core/ingestion/languages/swift/query.js`: the module that must be
* reached-but-lazy, whose disappearance would make the swift half of the
* headline vacuous. Both guards are kept — one pins the ESM chain to an
* optional language, the other pins the native channel to a required one.
*
* Characterization-first: this MUST fail against the pre-fix code (run against
* the parent commit to verify the regression signal works).
*/
import { describe, it, expect } from 'vitest';
import { spawnSync } from 'node:child_process';
import path from 'node:path';
import fs from 'node:fs';
import { fileURLToPath, pathToFileURL } from 'node:url';
const __dirname = path.dirname(fileURLToPath(import.meta.url));
const REPO_ROOT = path.resolve(__dirname, '..', '..', '..');
const DIST_REGISTRY = path.join(
REPO_ROOT,
'dist',
'core',
'ingestion',
'scope-resolution',
'pipeline',
'registry.js',
);
const DIST_REGISTRY_URL = pathToFileURL(DIST_REGISTRY).href;
// Import the registry, then report every newly-loaded CJS-cache key. The cache
// tracks native/.node bindings loaded by either ESM or CJS importers, which is
// exactly how a tree-sitter grammar binding surfaces.
const PROBE = `
import { createRequire } from 'node:module';
const req = createRequire(import.meta.url);
const before = new Set(Object.keys(req.cache));
await import(process.env.PROBE_TARGET);
const after = new Set(Object.keys(req.cache));
process.stdout.write(JSON.stringify([...after].filter((k) => !before.has(k))));
`;
import { describe, it, expect, beforeAll } from 'vitest';
import { probeModuleLoad, type ModuleLoadProbe } from '../../helpers/module-load-probe.js';
// `tree-sitter-c[\\/]` matches only the exact `tree-sitter-c/` package — NOT
// `tree-sitter-cpp/` or `tree-sitter-c-sharp/` (those need a non-separator after
// the `c`), so the required C++/C# eager loads are unaffected.
const OPTIONAL_GRAMMAR_RE = /tree-sitter-(swift|dart|kotlin|c)[\\/]/;
/** Any tree-sitter grammar package — required ones included. */
const ANY_GRAMMAR_RE = /tree-sitter-[a-z-]+[\\/]/;
describe('optional-grammar static-import closure (#2091/#2093, #2116)', () => {
it('importing the scope-resolution registry loads NO lazy grammar binding (swift/dart/kotlin/c)', () => {
if (!fs.existsSync(DIST_REGISTRY)) {
throw new Error(
`${DIST_REGISTRY} missing — run \`npm run build\` first (or \`npm run test:integration\`, ` +
`which builds via pretest:integration).`,
);
}
let probe: ModuleLoadProbe;
const result = spawnSync(process.execPath, ['--input-type=module', '-e', PROBE], {
cwd: REPO_ROOT,
// NODE_OPTIONS cleared so a session-pinned --max-old-space-size etc. can't
// perturb the child. The skip env is cleared so install state is probed.
env: {
...process.env,
PROBE_TARGET: DIST_REGISTRY_URL,
NODE_OPTIONS: '',
GITNEXUS_SKIP_OPTIONAL_GRAMMARS: '',
},
timeout: 60_000,
encoding: 'utf8',
beforeAll(async () => {
probe = await probeModuleLoad({
entry: 'core/ingestion/scope-resolution/pipeline/registry.js',
// The registry's closure MUST still reach the OPTIONAL languages'
// query.ts modules — that is precisely what makes "no optional binding
// loaded" meaningful rather than trivially true. If a refactor severs
// registry → swift/query.js, this fails loudly instead of letting the
// assertion below pass green on a no-longer-exercised path.
anchor: 'dist/core/ingestion/languages/swift/query.js',
// Observed on Node 22.18 against a clean build: 553 distinct modules.
minModules: 100,
// Cleared so install state — not a skip flag inherited from the caller —
// is what the child probes.
env: { GITNEXUS_SKIP_OPTIONAL_GRAMMARS: '' },
});
}, 90_000);
// Post-fix, importing the registry must not throw even though the chain
// reaches swift/dart/kotlin query.ts. (Pre-fix on a machine missing a
// grammar this would be ERR_MODULE_NOT_FOUND; here the grammar is present
// so pre-fix it would instead surface as a loaded binding below.)
if (result.status !== 0) {
// status is null when the child was killed by a signal (e.g. a native
// addon SIGSEGV) — surface the signal so that's distinguishable from a
// non-zero exit / module-not-found.
const exit =
result.status !== null ? `status ${result.status}` : `signal ${result.signal ?? 'unknown'}`;
throw new Error(
`importing the scope-resolution registry failed (${exit}):\n` +
`stderr:\n${result.stderr}\nstdout:\n${result.stdout}`,
);
}
const newlyLoaded = JSON.parse(result.stdout) as string[];
// Non-vacuity guard: the registry's static-import closure MUST still reach
// the per-language query.ts modules (which is what makes "no optional
// binding loaded" meaningful). The REQUIRED grammars (python/typescript/…)
// still import their binding eagerly in their own query.ts, so at least one
// non-optional tree-sitter binding must appear. If a future refactor severs
// the registry→query.ts edge, this fails loudly instead of letting the
// optional-binding assertion pass green on a no-longer-exercised path.
const requiredLoaded = newlyLoaded.filter(
(p) => /tree-sitter-[a-z-]+[\\/]/.test(p) && !OPTIONAL_GRAMMAR_RE.test(p),
);
it('importing the scope-resolution registry loads NO lazy grammar binding (swift/dart/kotlin/c)', () => {
// Second non-vacuity guard, on the native channel: the REQUIRED grammars
// (python/typescript/…) still import their binding eagerly in their own
// query.ts, so at least one non-optional tree-sitter binding must appear.
// Losing this would mean the probe no longer observes grammar loads at all,
// which the module-count floor alone would not catch.
const requiredLoaded = probe
.matching(ANY_GRAMMAR_RE)
.filter((p) => !OPTIONAL_GRAMMAR_RE.test(p));
expect(
requiredLoaded.length,
`Expected the registry import closure to load at least one REQUIRED tree-sitter ` +
`binding (proving the chain still reaches the per-language query.ts modules). ` +
`Newly-loaded (${newlyLoaded.length}):\n${newlyLoaded.join('\n')}`,
`Loaded (${probe.modules.length}):\n${probe.modules.join('\n')}`,
).toBeGreaterThan(0);
// Headline assertion: no lazy grammar binding (swift/dart/kotlin/c) is
// loaded at registry static-import time — they must load lazily.
const optionalLoaded = newlyLoaded.filter((p) => OPTIONAL_GRAMMAR_RE.test(p));
const optionalLoaded = probe.matching(OPTIONAL_GRAMMAR_RE);
expect(
optionalLoaded,
`Lazy tree-sitter grammar binding(s) loaded at registry static-import time. ` +
@@ -0,0 +1,353 @@
/**
* Resolver pin: a TypeScript class field whose type must be INFERRED from its
* initializer cannot act as a call receiver — the calling method emits NO CALLS
* edges at all, not even for the first, ordinary named-receiver link.
*
* ── WHERE THE SUPPORT ACTUALLY STOPS ──────────────────────────────────────────
*
* Measured against the single-file fixture below (nine receiver shapes in one
* repo). Every caller runs the same statement, `<receiver>.inner().compute(x)`;
* only the receiver FORM varies. The discriminator is whether the receiver's
* type is DECLARED, not whether it is a local or a field:
*
* receiver form CALLS edges emitted
* ------------------------------------------------------- --------------------------
* local `const o = new Outer()` Outer.inner + Inner.compute
* field `private p: Outer = new Outer()` (ANNOTATED) Outer.inner + Inner.compute
* field `private p: Outer;` + ctor `this.p = new Outer()` Outer.inner + Inner.compute
* field `private p: Outer;` + ctor param `this.p = p` Outer.inner + Inner.compute
* field `constructor(private p: Outer)` (param property) Outer.inner + Inner.compute
* result `makeOuter().inner().compute()` makeOuter + both links
* chain `o.inner().mid().compute()` (three links) all three links
* field `private p = new Outer()` (INFERRED) NONE <- known gap
* field `private p;` + ctor `this.p = new Outer()` NONE <- known gap
*
* The two gap rows do not merely lose the CHAINED link. They emit nothing: the
* caller has no outgoing CALLS edge whatsoever, so `Outer.inner` — a plainly
* named receiver call — is lost too. `KNOWN GAP` below asserts that empty value
* EXACTLY, alongside a `callerExists` probe so the assertion cannot pass
* vacuously if fixture drift or an id-scheme change moved the caller node.
*
* `only the receiver TYPE is lost` narrows where to look: the `new Outer()`
* initializer of an inferred field IS resolved (it emits its own constructor
* CALLS edge, exactly as the annotated twin does). What is missing is the step
* that turns that initializer into a type binding for the field.
*
* The annotated twins are pinned in the same file on purpose — a gap test that
* shows only the broken shape does not tell the next engineer where the
* boundary is.
*
* ── THIS PIN IS SELF-DIFFING: IT WILL GO RED ON PURPOSE ───────────────────────
*
* The gap is tracked as issue #2807 ("Inference-typed field receivers resolve to
* no CALLS edges at all"); PR #2810 is open against it at the time of writing.
* `KNOWN GAP` asserts that the gap EXISTS — no CALLS edges, exactly — so it is a
* record rather than a regression guard. When the resolver learns to infer a
* field's type from its initializer it fails with the newly resolved ids in the
* diff; that is the intended signal. The fix is to update this file (the table
* above, the rows' `resolution`, and the pin's expected value), not to relax the
* assertion into something a passing fix would also satisfy.
*
* The gap was FOUND during #2802 work but is PRE-EXISTING and independent of it:
* nothing on that branch touches receiver typing. Track the gap itself at #2807.
*
* The same fact is also observable one layer down, as an empty
* `BasicBlock.calleeIds` cell, in `test/integration/cfg/
* pdg-chained-receiver-callees.test.ts` — but that is the PDG's view of a
* RESOLVER fact, behind a full `--pdg` pipeline. Whoever closes this gap will
* be working in the resolver suite, so the fact is pinned here too, at the
* level the fix actually changes.
*/
import { describe, it, expect, beforeAll, afterAll } from 'vitest';
import path from 'node:path';
import fs from 'node:fs';
import os from 'node:os';
import {
getRelationships,
runPipelineFromRepo,
writeFixtureRepo,
type PipelineResult,
} from './helpers.js';
const FIXTURE_PATH = 'src/app.ts';
const CHAINED_SOURCE = `export class Mid {
compute(v: number): number {
return v * 3;
}
}
export class Inner {
compute(v: number): number {
return v * 2;
}
mid(): Mid {
return new Mid();
}
}
export class Outer {
inner(): Inner {
return new Inner();
}
}
export function makeOuter(): Outer {
return new Outer();
}
export function runLocalConst(x: number): number {
const localConst = new Outer();
const r = localConst.inner().compute(x);
return r;
}
export function runCallResultReceiver(x: number): number {
const r = makeOuter().inner().compute(x);
return r;
}
export function runThreeLink(x: number): number {
const threeLink = new Outer();
const r = threeLink.inner().mid().compute(x);
return r;
}
export class AnnotatedFieldCaller {
private annotated: Outer = new Outer();
runAnnotatedField(x: number): number {
const r = this.annotated.inner().compute(x);
return r;
}
}
export class CtorAssignedAnnotatedCaller {
private ctorTyped: Outer;
constructor() {
this.ctorTyped = new Outer();
}
runCtorAssignedAnnotated(x: number): number {
const r = this.ctorTyped.inner().compute(x);
return r;
}
}
export class CtorParamAnnotatedCaller {
private ctorParam: Outer;
constructor(ctorParam: Outer) {
this.ctorParam = ctorParam;
}
runCtorParamAnnotated(x: number): number {
const r = this.ctorParam.inner().compute(x);
return r;
}
}
export class ParamPropertyCaller {
constructor(private paramProp: Outer) {}
runParamProperty(x: number): number {
const r = this.paramProp.inner().compute(x);
return r;
}
}
export class InferredFieldCaller {
private inferred = new Outer();
runInferredField(x: number): number {
const r = this.inferred.inner().compute(x);
return r;
}
}
export class CtorAssignedInferredCaller {
private ctorUntyped;
constructor() {
this.ctorUntyped = new Outer();
}
runCtorAssignedInferred(x: number): number {
const r = this.ctorUntyped.inner().compute(x);
return r;
}
}
`;
// EXACT node ids — never names or substrings. `compute` alone is ambiguous
// between `Inner.compute` and `Mid.compute`, and matching on the source NAME
// would collide on `constructor` (two classes define one). `#N` is the arity
// disambiguator the resolver mints.
const OUTER_CLASS = `Class:${FIXTURE_PATH}:Outer`;
const OUTER_INNER = `Method:${FIXTURE_PATH}:Outer.inner#0`;
const INNER_COMPUTE = `Method:${FIXTURE_PATH}:Inner.compute#1`;
const INNER_MID = `Method:${FIXTURE_PATH}:Inner.mid#0`;
const MID_COMPUTE = `Method:${FIXTURE_PATH}:Mid.compute#1`;
const MAKE_OUTER = `Function:${FIXTURE_PATH}:makeOuter`;
/**
* `resolves` — every chain link becomes a CALLS edge today.
* `known-gap-no-calls-edges` — the resolver cannot type the receiver, so the
* caller emits no CALLS edge at all and even the named first link is lost.
*/
type ChainResolution = 'resolves' | 'known-gap-no-calls-edges';
interface ReceiverShape {
/** Row name; also the key of the known-gap pin below. */
readonly name: string;
/** Exact node id of the function or method holding the chained statement. */
readonly callerId: string;
/** EVERY CALLS target id this caller emits today, in any order. */
readonly targets: readonly string[];
readonly resolution: ChainResolution;
}
const RECEIVER_SHAPES: readonly ReceiverShape[] = [
{
name: 'local-const',
callerId: `Function:${FIXTURE_PATH}:runLocalConst`,
targets: [OUTER_CLASS, OUTER_INNER, INNER_COMPUTE],
resolution: 'resolves',
},
{
name: 'annotated-field-initializer',
callerId: `Method:${FIXTURE_PATH}:AnnotatedFieldCaller.runAnnotatedField#1`,
targets: [OUTER_INNER, INNER_COMPUTE],
resolution: 'resolves',
},
{
name: 'ctor-assigned-annotated',
callerId: `Method:${FIXTURE_PATH}:CtorAssignedAnnotatedCaller.runCtorAssignedAnnotated#1`,
targets: [OUTER_INNER, INNER_COMPUTE],
resolution: 'resolves',
},
{
name: 'ctor-param-annotated',
callerId: `Method:${FIXTURE_PATH}:CtorParamAnnotatedCaller.runCtorParamAnnotated#1`,
targets: [OUTER_INNER, INNER_COMPUTE],
resolution: 'resolves',
},
{
name: 'param-property',
callerId: `Method:${FIXTURE_PATH}:ParamPropertyCaller.runParamProperty#1`,
targets: [OUTER_INNER, INNER_COMPUTE],
resolution: 'resolves',
},
{
name: 'call-result-receiver',
callerId: `Function:${FIXTURE_PATH}:runCallResultReceiver`,
targets: [MAKE_OUTER, OUTER_INNER, INNER_COMPUTE],
resolution: 'resolves',
},
{
name: 'three-link-chain',
callerId: `Function:${FIXTURE_PATH}:runThreeLink`,
targets: [OUTER_CLASS, OUTER_INNER, INNER_MID, MID_COMPUTE],
resolution: 'resolves',
},
// ── Known gaps ────────────────────────────────────────────────────────────
// Identical to `annotated-field-initializer` / `ctor-assigned-annotated`
// above except that the field carries no type annotation, so its type would
// have to be inferred from the initializer.
{
name: 'inferred-field-initializer',
callerId: `Method:${FIXTURE_PATH}:InferredFieldCaller.runInferredField#1`,
targets: [],
resolution: 'known-gap-no-calls-edges',
},
{
name: 'ctor-assigned-inferred',
callerId: `Method:${FIXTURE_PATH}:CtorAssignedInferredCaller.runCtorAssignedInferred#1`,
targets: [],
resolution: 'known-gap-no-calls-edges',
},
];
describe('TypeScript chained receiver calls by field-type form (known gap: #2807)', () => {
let result: PipelineResult;
let repoDir: string | undefined;
beforeAll(async () => {
repoDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gn-ts-inferred-field-'));
writeFixtureRepo(repoDir, { [FIXTURE_PATH]: CHAINED_SOURCE });
// CALLS resolution is complete before the graph phases run and this pin
// reads nothing they produce (MRO, communities, processes), so skipping
// them narrows the run to the phase under test. Cost here is dominated by
// worker-pool startup, not by the phases, so this is about scope rather
// than speed.
result = await runPipelineFromRepo(repoDir, () => {}, { skipGraphPhases: true });
}, 120000);
afterAll(() => {
if (repoDir !== undefined) fs.rmSync(repoDir, { recursive: true, force: true });
});
/** Every CALLS target id emitted by one exact caller node, sorted. */
function callTargetsFrom(callerId: string): string[] {
return getRelationships(result, 'CALLS')
.filter((edge) => edge.rel.sourceId === callerId)
.map((edge) => edge.rel.targetId)
.sort();
}
function nodeExists(id: string): boolean {
return result.graph.getNode(id) !== undefined;
}
it('every receiver shape contributes exactly one caller node', () => {
const found = Object.fromEntries(RECEIVER_SHAPES.map((s) => [s.name, nodeExists(s.callerId)]));
expect(found).toEqual(Object.fromEntries(RECEIVER_SHAPES.map((s) => [s.name, true])));
});
// Exact set equality, not `arrayContaining`: a shape that started resolving
// something extra (or stopped resolving a link) has to show up in the diff.
for (const shape of RECEIVER_SHAPES.filter((s) => s.resolution === 'resolves')) {
it(`${shape.name}: every chain link becomes a CALLS edge`, () => {
expect(callTargetsFrom(shape.callerId)).toEqual([...shape.targets].sort());
});
}
// Pins the CURRENT broken value, not merely that the chain fails. An
// `it.fails` row here would be strictly weaker — it is satisfied by ANY
// throw, so a renamed fixture symbol would keep it green on a rotted
// premise. `callerExists` is folded into the same object so an empty
// `calls` list can never be read as "resolved fine, wrong node id".
// This asserts the gap EXISTS (issue #2807), so closing #2807 turns it red
// BY DESIGN; update it together with the header table and the rows'
// `resolution` rather than loosening it.
it('KNOWN GAP (#2807): an inference-typed field receiver emits NO CALLS edges at all', () => {
const gaps = RECEIVER_SHAPES.filter((s) => s.resolution === 'known-gap-no-calls-edges');
const observed = Object.fromEntries(
gaps.map((s) => [
s.name,
{ callerExists: nodeExists(s.callerId), calls: callTargetsFrom(s.callerId) },
]),
);
expect(observed).toEqual({
'inferred-field-initializer': { callerExists: true, calls: [] },
'ctor-assigned-inferred': { callerExists: true, calls: [] },
});
});
// Boundary evidence: the initializer is not invisible to the resolver. Both
// twins of each pair emit the `new Outer()` constructor edge; only the
// annotated one turns it into a receiver type. So the missing step is the
// initializer -> field type binding, not the initializer itself.
it('the inferred field initializer IS resolved — only the receiver TYPE is lost', () => {
const initializerCalls = {
'annotated-field-initializer': callTargetsFrom(`Class:${FIXTURE_PATH}:AnnotatedFieldCaller`),
'inferred-field-initializer': callTargetsFrom(`Class:${FIXTURE_PATH}:InferredFieldCaller`),
'ctor-assigned-annotated': callTargetsFrom(
`Method:${FIXTURE_PATH}:CtorAssignedAnnotatedCaller.constructor#0`,
),
'ctor-assigned-inferred': callTargetsFrom(
`Method:${FIXTURE_PATH}:CtorAssignedInferredCaller.constructor#0`,
),
};
expect(initializerCalls).toEqual({
'annotated-field-initializer': [OUTER_CLASS],
'inferred-field-initializer': [OUTER_CLASS],
'ctor-assigned-annotated': [OUTER_CLASS],
'ctor-assigned-inferred': [OUTER_CLASS],
});
});
});
@@ -0,0 +1,180 @@
/**
* Corpus-derived coverage for the HAND-DECLARED half of `RELATION_SCHEMA`.
*
* `test/unit/schema-pair-coverage.test.ts` derives its requirement from
* schema.ts's two rules — the scope-resolution bridge cross product, and
* `DEFINITION_ANCHOR_LABELS × ATTACHMENT_TARGET_LABELS` for the framework and
* pipeline-phase overlays — so both generated halves are covered there. What
* neither rule can reach is `STRUCTURAL_PAIR_DDL`: the containment, inheritance
* and import pairs BETWEEN TWO DEFINITION LABELS. Eleven node tables are absent
* from every rule's target side (`CodeElement`, `Impl`, `Namespace`,
* `Template`, `TypeAlias`, `Typedef`, `Union`, `Static`, `Section`, `Folder`,
* and the PDG-only `BasicBlock`), so a pair pointing at one is hand-declared or
* it does not exist. No predicate describes that surface — any container can
* hold any definition — so this asks the emitters directly: run the real
* pipeline and require every FROM/TO pair it produces to be declared.
*
* Each entry also pins the pair it exists to guard. Without that the suite is
* vacuous: `undeclared` is derived from what the pipeline emitted, so a fixture
* that stopped emitting — renamed directory, grammar that failed to load,
* swallowed parse error — yields an empty set and passes green while guarding
* nothing. `sentinels` turns each case from "nothing undeclared" into "this
* emitter still fires, and everything it emits is declared".
*
* Coverage is bounded by `NON_BRIDGE_CORPUS`: a sample, not a proof. A
* language whose fixture is absent is unguarded, so a new structural emitter
* should land with an entry here.
*
* Deliberately isolated from the resolver suites that already build three of
* these graphs (`resolvers/cobol.test.ts`, `resolvers/vue.test.ts`,
* `resolvers/php.test.ts`) — see NOTE below the corpus before re-raising that.
*/
import { describe, it, expect, vi, beforeAll, afterAll } from 'vitest';
import path from 'path';
import { NODE_TABLES } from 'gitnexus-shared';
import { FIXTURES, runPipelineFromRepo } from './resolvers/helpers.js';
import { RELATION_SCHEMA } from '../../src/core/lbug/schema.js';
import { parseRelationSchemaPairs, relPairKeyFor } from '../../src/core/lbug/rel-pair-routing.js';
import { DIST_WORKER_URL, distWorkerExists } from '../helpers/worker-parse.js';
vi.setConfig({ testTimeout: 180_000 });
const describeIfWorkerBuilt = distWorkerExists() ? describe : describe.skip;
type CorpusEntry = {
/** Fixture directory under `test/fixtures/lang-resolution`. */
readonly fixture: string;
/**
* The emitter this fixture exists to exercise, short enough that vitest does
* not truncate it out of the case title (~36 chars).
*/
readonly emitter: string;
/**
* FROM|TO pairs the fixture MUST still emit. These are the anti-vacuity
* check: they fail loudly when the fixture stops reaching the emitter,
* which is the failure mode "no undeclared pairs" cannot see.
*/
readonly sentinels: readonly string[];
};
/**
* Fixtures chosen to reach a structural emitter that no other suite drives.
*
* The sentinels are the anti-vacuity anchor: each names an emitter that must
* still fire. Two of them (`CodeElement|Property`, `Module|Namespace`) are also
* the only pairs here that no generated rule can reach; the other nine are
* rule-derived, and stay because a rule DECLARING a pair says nothing about
* whether any emitter still PRODUCES it — which is the failure this corpus
* exists to catch.
*/
const NON_BRIDGE_CORPUS = [
{
// `cobol-processor.ts`: CONTAINS/CALLS/ACCESSES over Module / Namespace /
// Record / Property / CodeElement.
fixture: 'cobol-app',
emitter: 'cobol-processor containment',
sentinels: ['CodeElement|Property', 'Module|Namespace', 'Module|Record', 'Record|Record'],
},
{
// Same processor, DECLARATIVES section: the USE-procedure Namespace
// ACCESSES a file Record.
fixture: 'cobol-declaratives',
emitter: 'cobol-processor DECLARATIVES',
sentinels: ['Namespace|Record'],
},
{
// `languages/vue/scope-resolver.ts`: the only edge whose target is a
// `File`. Function→File from `<script setup>` (App.vue), Method→File from
// the Options-API `methods:` host (OptionsHost.vue).
fixture: 'vue-basic',
emitter: 'vue BINDS_EVENT_HANDLER',
sentinels: ['Function|File', 'Method|File'],
},
{
// The inheritance pass, including trait-to-trait IMPLEMENTS.
fixture: 'php-transitive-traits',
emitter: 'inheritance-pass IMPLEMENTS',
sentinels: ['Class|Trait', 'Trait|Trait'],
},
{
// `frameworks/spring/conditionals.ts`: @ConditionalOn* on a @Bean method,
// in both Java and Kotlin.
fixture: 'spring-conditional-app',
emitter: 'spring CONDITIONAL_ON',
sentinels: ['Method|Annotation'],
},
{
// `pipeline-phases/tools.ts`: a class-based MCP tool handler.
fixture: 'mcp-tool-class',
emitter: 'tools-phase HANDLES_TOOL',
sentinels: ['Class|Tool'],
},
] as const satisfies readonly CorpusEntry[];
/*
* NOTE — why this suite runs its own pipelines instead of reusing the resolver
* suites' graphs (measured, not assumed):
*
* 1. Vitest runs with `pool: 'forks'` and default isolation, so every test
* FILE gets its own child process. A fixture-keyed result cache in
* `resolvers/helpers.ts` would be per-file module state and would share
* nothing across files — zero saving.
* 2. The graphs are not interchangeable. `resolvers/cobol.test.ts` builds
* cobol-app with `{ skipGraphPhases: true }`, which drops the phase-emitted
* edges (`Function|Process`, `Function|Community`, and the whole class of
* pair `mcp-tool-class` exists to guard: HANDLES_TOOL is a pipeline phase).
* Asserting there would silently cover LESS surface than here.
* 3. Half the corpus has no existing home anyway, and the cases run
* concurrently, so they overlap into roughly one fixture's wall time: all
* 6 measured ~6s of test time against the ~16s this file spends on
* transform+import before the first case starts. Hosting only the 3
* homeless ones measured ~5s — a ~1s saving that still cannot remove a
* file, so it cannot remove that ~16s.
*/
const DECLARED = parseRelationSchemaPairs(RELATION_SCHEMA);
const VALID_TABLES = new Set<string>(NODE_TABLES);
// Cold worker-pool startups otherwise flake against the 5s default ready budget
// on a loaded runner, failing for reasons unrelated to the schema (#1741).
// Safe under `it.concurrent`: env stubs are process-global, but `pool: 'forks'`
// gives this file its own process and every case wants the same value.
beforeAll(() => vi.stubEnv('GITNEXUS_WORKER_READY_TIMEOUT_MS', '60000'));
afterAll(() => vi.unstubAllEnvs());
/**
* The FROM/TO pairs a fixture emits, deduped.
*
* Classifies through `relPairKeyFor` — the same call the emit path routes with
* — so this guard and the router cannot disagree about which edges are skipped.
* It keys off the node IDS, which is what `assertDeclaredPair` actually sees,
* not the node's `label` field.
*/
const pairsEmittedBy = async (fixture: string): Promise<Set<string>> => {
const result = await runPipelineFromRepo(path.join(FIXTURES, fixture), () => {}, {
workerPoolSize: 1,
workerUrlForTest: DIST_WORKER_URL,
});
const emitted = new Set<string>();
for (const rel of result.graph.iterRelationships()) {
const pair = relPairKeyFor(rel.sourceId, rel.targetId, VALID_TABLES);
if (pair !== undefined) emitted.add(pair);
}
return emitted;
};
describeIfWorkerBuilt('RELATION_SCHEMA covers the non-bridge emitters', () => {
// Concurrent because the cases share nothing but cost ~5s each serially,
// almost all of it worker spawn and grammar load, which overlaps well.
it.concurrent.each(NON_BRIDGE_CORPUS)(
'$fixture emits only declared FROM/TO pairs, and still reaches $emitter',
async ({ fixture, sentinels }) => {
const emitted = await pairsEmittedBy(fixture);
// Sorted so a failure is stable and names the pair to declare.
expect({
undeclaredPairs: [...emitted].filter((pair) => !DECLARED.has(pair)).sort(),
missingSentinelPairs: sentinels.filter((pair) => !emitted.has(pair)),
}).toEqual({ undeclaredPairs: [], missingSentinelPairs: [] });
},
);
});
+58 -1
View File
@@ -149,6 +149,62 @@ describe('generateAIContextFiles', () => {
expect(content).not.toMatch(/npx gitnexus group/);
});
it('pairs mandatory MCP checks with project-local CLI fallbacks (#2671)', async () => {
const dir = await fs.mkdtemp(path.join(os.tmpdir(), 'gn-ai-ctx-transport-fallback-'));
const storage = path.join(dir, '.gitnexus');
await fs.mkdir(storage, { recursive: true });
try {
await generateAIContextFiles(
dir,
storage,
'TransportFallback',
{ nodes: 50, edges: 100, processes: 5 },
undefined,
{ defaultBranch: 'develop' },
);
for (const file of ['AGENTS.md', 'CLAUDE.md']) {
const content = await fs.readFile(path.join(dir, file), 'utf-8');
expect(content).toContain('impact({target: "symbolName", direction: "upstream"})');
expect(content).toContain(
'node .gitnexus/run.cjs impact "symbolName" --direction upstream --repo .',
);
expect(content).toContain('detect_changes({scope: "all"})');
expect(content).toContain('node .gitnexus/run.cjs detect-changes --scope all --repo .');
expect(content).toContain('detect_changes({scope: "compare", base_ref: "develop"})');
expect(content).toContain('--scope compare --base-ref "develop" --repo .');
expect(content).toContain('Never substitute grep for graph analysis');
expect(content).not.toContain('gitnexus impact --target');
}
} finally {
await fs.rm(dir, { recursive: true, force: true });
}
});
it('uses the configured runner for transport fallbacks, including PDG impact (#2671)', () => {
const content = generateGitNexusContent(
'TransportFallback',
{ nodes: 50, edges: 100, processes: 5 },
{
runnerPath: '.custom/gitnexus-runner.cjs',
hasPdg: true,
},
);
expect(content).toContain(
'node .custom/gitnexus-runner.cjs impact "symbolName" --direction upstream --repo .',
);
expect(content).toContain(
'node .custom/gitnexus-runner.cjs detect-changes --scope all --repo .',
);
expect(content).toContain(
'node .custom/gitnexus-runner.cjs impact "symbolName" --direction upstream --mode pdg --line <N> --repo .',
);
expect(content).toContain('--mode pdg');
expect(content).toContain('--line <N>');
expect(content).not.toContain('node .gitnexus/run.cjs impact');
});
it('gates the pdg_query line on hasPdg (#2086 M6 — no existing taint gate to mirror)', () => {
const stats = { nodes: 50, edges: 100, processes: 5 };
// hasPdg=true → the pdg_query line is present.
@@ -1181,7 +1237,7 @@ Indexed as **placeholder** (1 symbols, 1 relationships, 1 execution flows). Cust
// tool that does not exist.
expect(content).not.toMatch(/gitnexus_(impact|query|context|detect_changes|rename|cypher)/);
expect(content).toContain('impact({target: "symbolName", direction: "upstream"})');
expect(content).toContain('detect_changes()');
expect(content).toContain('detect_changes({scope: "all"})');
// #2175: the generated guidance must advertise the renamed param, never the
// legacy "query" key (Claude Code drops a tool arg named exactly "query").
expect(content).toContain('query({search_query: "concept"})');
@@ -1194,6 +1250,7 @@ Indexed as **placeholder** (1 symbols, 1 relationships, 1 execution flows). Cust
// raw, so it stays inside the inline code span.
const content = generateGitNexusContent('P', { nodes: 1 }, { defaultBranch: 'we"ird' });
expect(content).toContain('base_ref: "we\\"ird"');
expect(content).toContain('--base-ref "we\\"ird" --repo .');
});
it('a backtick branch cannot break the generated Markdown code span (#1996 P1)', () => {
+633 -7
View File
@@ -1,5 +1,36 @@
import fs from 'node:fs/promises';
import os from 'node:os';
import path from 'node:path';
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { JobManager } from '../../src/server/analyze-job.js';
import {
startSSEHarness,
terminalFrame,
terminalFrameCount,
type SSEHarness,
} from '../helpers/sse-harness.js';
import {
resolveEmbedRunOutcome,
withMeasuredEmbeddingCount,
type EmbeddingRunResult,
} from '../../src/server/embed-run-outcome.js';
import { mintInterruptedCheckpoint } from '../../src/core/embedding-checkpoint.js';
import {
measurePersistedEmbeddingCount,
persistedEmbeddingCountOrUndefined,
} from '../../src/core/embedding-count.js';
import { loadMeta, saveMeta, type RepoMeta } from '../../src/storage/repo-manager.js';
import { deriveEmbeddingMode } from '../../src/core/embedding-mode.js';
/**
* NOTHING in this file imports `src/server/api.ts` for behavior. That module
* pulls Express, cors, the LadybugDB native adapter and the whole MCP wiring:
* reaching three pure helpers through it cost one 30s TIMEOUT and ~20s/~22s on
* the runs that passed, against a 30s `testTimeout` (#2790 review, finding 9).
* The helpers now live in `src/server/{sse-progress,embed-run-outcome}.ts` and
* `src/core/embedding-{count,checkpoint}.ts`, none of which import a database
* or a server.
*/
describe('analyze API logic', () => {
let manager: JobManager;
@@ -36,9 +67,9 @@ describe('analyze API logic', () => {
it('SSE progress listener receives all events including terminal', () => {
const job = manager.createJob({ repoUrl: 'https://github.com/user/sse-test' });
const events: any[] = [];
const events: Array<{ phase: string; percent: number }> = [];
const unsub = manager.onProgress(job.id, (progress) => {
events.push(progress);
events.push({ phase: progress.phase, percent: progress.percent });
});
manager.updateJob(job.id, {
@@ -52,10 +83,605 @@ describe('analyze API logic', () => {
unsub();
expect(events.length).toBe(3);
expect(events[0].phase).toBe('parsing');
expect(events[1].phase).toBe('calls');
expect(events[2].phase).toBe('complete');
expect(events[2].percent).toBe(100);
expect(events).toEqual([
{ phase: 'parsing', percent: 30 },
{ phase: 'calls', percent: 50 },
{ phase: 'complete', percent: 100 },
]);
});
});
const IDENTITY = { model: 'test-model', dimensions: 384, provider: 'local' };
const CLEAN_RUN: EmbeddingRunResult = {
nodesProcessed: 412,
chunksProcessed: 900,
failedNodeIds: [],
};
/** Progress figures an in-flight checkpoint records. */
const PROGRESS = { nodesProcessed: 4, totalNodes: 12, chunksProcessed: 9 };
/**
* ── #2790: an SSE client must not be told a partial run succeeded ──────────
*
* `runEmbeddingPipeline` emits `phase: 'ready'` / 100% UNCONDITIONALLY before
* returning — including when it dropped nodes to endpoint failures — and
* /api/embed relayed that as a progress phase before it had measured anything
* or decided the outcome. The relay treated a progress PHASE STRING of
* 'complete'/'failed' as terminal, so it wrote `event: complete` with
* `error: undefined`, called `res.end()` and unsubscribed; the route's later
* `updateJob({status:'failed'})` went into a stream with no listener. The web
* client fired `onComplete` and showed "ready" while a `GET /api/embed/:jobId`
* poller saw `failed` — the two consumers of one job disagreeing about whether
* the data is complete, and a regression against the pre-#2790 behavior where
* the pipeline threw and the client received the failure.
*
* These tests drive the REAL relay over a REAL HTTP server (same harness as
* server-sse-payload.test.ts) and subscribe BEFORE the misleading event is
* emitted — subscribing after it is exactly why the previous version of this
* suite passed while the bug was live.
*/
describe('mountSSEProgress terminality (#2790)', () => {
let harness: SSEHarness;
let manager: JobManager;
let baseUrl = '';
beforeEach(async () => {
// Mirrors both production mounts in createServer().
harness = await startSSEHarness('/api/embed/:jobId/progress');
manager = harness.manager;
baseUrl = harness.baseUrl;
});
afterEach(() => harness.close());
it('a partial run reaches the client as a failure, not a success', async () => {
const job = manager.createJob({ repoPath: '/ws/embed-partial' });
manager.updateJob(job.id, {
repoName: 'embed-partial',
status: 'analyzing',
progress: { phase: 'embedding', percent: 40, message: 'Embedding nodes (40%)...' },
});
// The client is connected and listening BEFORE anything terminal-looking is
// emitted. `fetch` resolves once headers arrive, and the handler subscribes
// synchronously before that (see server-sse-payload.test.ts).
const response = await fetch(`${baseUrl}/api/embed/${job.id}/progress`);
// A progress event that CLAIMS to be terminal. Production now maps the
// pipeline's `ready` to 'finalizing' instead, but a phase string must not be
// able to end the stream no matter who sends it — that is the invariant.
manager.updateJob(job.id, {
progress: { phase: 'complete', percent: 100, message: 'Embeddings complete' },
});
// Only now does the route learn the run dropped nodes.
const outcome = resolveEmbedRunOutcome(IDENTITY, {
nodesProcessed: 10,
chunksProcessed: 24,
failedNodeIds: ['node-a', 'node-b'],
});
manager.updateJob(job.id, {
status: 'failed',
error: outcome.error,
partial: outcome.partial,
progress: { phase: 'failed', percent: 100, message: String(outcome.error) },
});
const body = await response.text();
expect(body).not.toContain('event: complete');
expect(terminalFrameCount(body)).toBe(1);
expect(terminalFrame(body, 'failed')).toMatchObject({
repoName: 'embed-partial',
repoPath: '/ws/embed-partial',
error: expect.stringContaining('finished partially') as unknown as string,
// The distinction a UI needs to offer "retry 2 nodes" instead of a bare
// red chip — carried without adding a `status` union member.
partial: { kind: 'embedding-partial', pendingNodeCount: 2, nodesProcessed: 10 },
});
});
it('a clean run produces exactly one terminal complete event', async () => {
const job = manager.createJob({ repoPath: '/ws/embed-clean' });
manager.updateJob(job.id, {
repoName: 'embed-clean',
status: 'analyzing',
progress: { phase: 'embedding', percent: 40, message: 'Embedding nodes (40%)...' },
});
const response = await fetch(`${baseUrl}/api/embed/${job.id}/progress`);
// What the route actually emits between the pipeline returning and the
// outcome being known.
manager.updateJob(job.id, {
progress: { phase: 'finalizing', percent: 100, message: 'Finalizing embeddings...' },
});
manager.updateJob(job.id, {
status: 'complete',
progress: { phase: 'complete', percent: 100, message: 'Embeddings complete' },
});
const body = await response.text();
expect(body).not.toContain('event: failed');
// Exactly one — the status update carries a `progress` too, and #2264's
// single-emit rule is what keeps that from double-writing the terminal frame.
expect(terminalFrameCount(body)).toBe(1);
expect(terminalFrame(body, 'complete')).toEqual({
repoName: 'embed-clean',
repoPath: '/ws/embed-clean',
});
// The 'finalizing' frame was relayed as ordinary progress, not swallowed.
expect(body).toContain('"phase":"finalizing"');
});
it('the analyze path still closes on its own terminal update', async () => {
// /api/analyze mounts the same relay. Its worker reports phases like
// 'parsing' and 'done' (never 'complete'), so the fix must not leave that
// stream open — it closes when the job's STATUS becomes terminal.
const job = manager.createJob({ repoPath: '/ws/reels' });
manager.updateJob(job.id, {
status: 'analyzing',
progress: { phase: 'parsing', percent: 30, message: 'Parsing' },
});
const response = await fetch(`${baseUrl}/api/embed/${job.id}/progress`);
manager.updateJob(job.id, {
progress: { phase: 'done', percent: 100, message: 'Done' },
});
manager.updateJob(job.id, { status: 'complete', repoName: 'reels' });
const body = await response.text();
expect(terminalFrameCount(body)).toBe(1);
expect(terminalFrame(body, 'complete')).toEqual({ repoName: 'reels', repoPath: '/ws/reels' });
});
it('a job that finished before the client connected replays its outcome', async () => {
const job = manager.createJob({ repoPath: '/ws/embed-late' });
const outcome = resolveEmbedRunOutcome(IDENTITY, {
nodesProcessed: 3,
chunksProcessed: 9,
failedNodeIds: ['node-a'],
});
manager.updateJob(job.id, {
status: 'failed',
repoName: 'embed-late',
error: outcome.error,
partial: outcome.partial,
});
const body = await (await fetch(`${baseUrl}/api/embed/${job.id}/progress`)).text();
expect(terminalFrameCount(body)).toBe(1);
expect(terminalFrame(body, 'failed')).toMatchObject({
error: expect.stringContaining('finished partially') as unknown as string,
partial: { kind: 'embedding-partial', pendingNodeCount: 1, nodesProcessed: 3 },
});
});
});
/**
* ── #2790: POST /api/embed must not report unqualified success ─────────
*
* The pipeline no longer throws when a sub-batch loses its endpoint — it
* deletes the affected nodes' rows and names them in `failedNodeIds`. The route
* discarded that receipt: it cleared `embeddingCheckpoint` and marked the job
* 'complete', so a partial run looked identical to a clean one and the dropped
* nodes were never retried (pre-#2790 the pipeline threw and the catch marked
* the job failed).
*/
describe('resolveEmbedRunOutcome (#2790)', () => {
it('clears the checkpoint and reports no error on a clean, measured run', () => {
const outcome = resolveEmbedRunOutcome(IDENTITY, CLEAN_RUN, { measuredEmbeddings: 412 });
expect(outcome.checkpoint).toBeUndefined();
expect(outcome.error).toBeUndefined();
expect(outcome.partial).toBeUndefined();
});
it('retains the checkpoint with the dropped ids and reports an error on a partial run', () => {
const outcome = resolveEmbedRunOutcome(IDENTITY, {
nodesProcessed: 10,
chunksProcessed: 24,
failedNodeIds: ['node-a', 'node-b'],
});
// The record of what failed survives — this is the pending set the next
// run's `forceReembedNodeIds` re-embeds.
expect(outcome.checkpoint).toMatchObject({
pendingNodeIds: ['node-a', 'node-b'],
nodesProcessed: 10,
totalNodes: 12,
chunksProcessed: 24,
model: 'test-model',
dimensions: 384,
provider: 'local',
// The run COMPLETED: these nodes provably hold zero rows, so a later
// identity mismatch may drop the set with a warning instead of wedging
// every subsequent run (repo-manager.ts).
kind: 'partial',
});
expect(outcome.error).toMatch(/2 node\(s\)/);
expect(outcome.partial).toEqual({
kind: 'embedding-partial',
pendingNodeCount: 2,
nodesProcessed: 10,
});
});
it('stamps no attempt count on a fresh partial run', () => {
const outcome = resolveEmbedRunOutcome(
IDENTITY,
{ nodesProcessed: 10, chunksProcessed: 24, failedNodeIds: ['node-a'] },
// Resumed from an in-flight marker, not a partial one.
{ resumedFrom: mintInterruptedCheckpoint(IDENTITY, PROGRESS, ['node-a']) },
);
expect(outcome.checkpoint).toMatchObject({ kind: 'partial' });
expect(outcome.checkpoint?.attempts).toBeUndefined();
});
it('advances the attempt count only when a resumed pending node fails again', () => {
const resumedFrom: RepoMeta['embeddingCheckpoint'] = {
at: new Date(0).toISOString(),
nodesProcessed: 10,
totalNodes: 12,
chunksProcessed: 24,
...IDENTITY,
kind: 'partial',
attempts: 1,
pendingNodeIds: ['node-a', 'node-b'],
};
// Same node failed again → the retry is not converging; the budget advances.
expect(
resolveEmbedRunOutcome(
IDENTITY,
{ nodesProcessed: 11, chunksProcessed: 26, failedNodeIds: ['node-a'] },
{ resumedFrom },
).checkpoint,
).toMatchObject({ kind: 'partial', attempts: 2 });
// The resumed set cleared and DIFFERENT nodes were lost → a fresh partial,
// so the budget resets. The bound exists for a node the endpoint rejects
// deterministically, not for an endpoint that is merely flaky.
expect(
resolveEmbedRunOutcome(
IDENTITY,
{ nodesProcessed: 11, chunksProcessed: 26, failedNodeIds: ['node-z'] },
{ resumedFrom },
).checkpoint?.attempts,
).toBeUndefined();
});
});
describe('the mid-run marker /api/embed writes (mintInterruptedCheckpoint, #2790)', () => {
it('stamps interrupted, so resume regenerates a possibly half-written window', () => {
const checkpoint = mintInterruptedCheckpoint(IDENTITY, PROGRESS, ['node-a', 'node-b']);
expect(checkpoint).toMatchObject({
kind: 'interrupted',
nodesProcessed: 4,
totalNodes: 12,
chunksProcessed: 9,
model: 'test-model',
dimensions: 384,
provider: 'local',
pendingNodeIds: ['node-a', 'node-b'],
});
// `attempts` bounds retries of a 'partial' set; an in-flight marker has no
// such budget because its rows may exist.
expect(checkpoint.attempts).toBeUndefined();
});
});
/**
* ── The /api/embed count omission (silent embedding loss) ──────────────
*
* The route generated embeddings and wrote `embeddingCheckpoint`, but never
* `stats.embeddings`. A repo embedded purely through the server therefore kept
* whatever count the last CLI `analyze` stamped — `0` for a repo analyzed
* without embeddings. The next CLI run reads that as `existingEmbeddingCount`,
* `deriveEmbeddingMode` sees `hasExisting: false` → `shouldLoadCache: false`,
* and `gitnexus analyze --force` wipes the database with no cache load: every
* server-generated embedding is destroyed with no warning.
*
* The route body is an inline closure inside `createServer`, so its finalize
* sequence is replayed here over the SAME helpers the route calls, with real
* meta.json I/O and the real `deriveEmbeddingMode`. The consequence is what
* these tests pin, not the field.
*/
describe('POST /api/embed records the embedding count it measured', () => {
let metaDir: string;
let seeded: RepoMeta;
beforeEach(async () => {
metaDir = await fs.mkdtemp(path.join(os.tmpdir(), 'gn-embed-count-'));
});
afterEach(async () => {
await fs.rm(metaDir, { recursive: true, force: true });
});
/** What a CLI `analyze` (plus any mid-run checkpoint) leaves on disk. */
const seedMeta = async (
embeddings: number | undefined,
embeddingCheckpoint?: RepoMeta['embeddingCheckpoint'],
): Promise<void> => {
seeded = {
repoPath: '/repo/embed-count',
lastCommit: 'abc123',
indexedAt: new Date(0).toISOString(),
stats: { nodes: 500, ...(embeddings === undefined ? {} : { embeddings }) },
embeddingCheckpoint,
};
await saveMeta(metaDir, seeded);
};
const rowsWith = (cnt: unknown) => async () => [{ cnt } as Record<string, unknown>];
/** The route's finalize sequence: measure → re-read meta → resolve → write. */
const finalizeEmbedRun = async (
runQuery: (cypher: string) => Promise<Array<Record<string, unknown>> | undefined>,
pipelineResult: EmbeddingRunResult,
): Promise<RepoMeta | null> => {
const measured = await measurePersistedEmbeddingCount(runQuery);
const finalMeta = (await loadMeta(metaDir)) ?? seeded;
const outcome = resolveEmbedRunOutcome(IDENTITY, pipelineResult, {
measuredEmbeddings: persistedEmbeddingCountOrUndefined(measured),
onDisk: finalMeta,
});
await saveMeta(
metaDir,
withMeasuredEmbeddingCount(
{ ...finalMeta, embeddingCheckpoint: outcome.checkpoint },
measured,
),
);
return loadMeta(metaDir);
};
const embeddingCountOf = (meta: RepoMeta | null): number => meta?.stats?.embeddings ?? 0;
it('writes the measured count into meta on a clean run, without disturbing the other stats', async () => {
await seedMeta(0);
const asked: string[] = [];
const written = await finalizeEmbedRun(async (cypher) => {
asked.push(cypher);
return [{ cnt: 412 }];
}, CLEAN_RUN);
expect(written).toMatchObject({ stats: { nodes: 500, embeddings: 412 } });
// A clean, MEASURED run clears the checkpoint (#2790 contract).
expect(written?.embeddingCheckpoint).toBeUndefined();
// Measured, not restated: the count comes from the live embedding table.
expect(asked).toEqual([expect.stringMatching(/MATCH \(e:\w+\) RETURN count\(e\) AS cnt/)]);
});
it('is what makes the next CLI run preserve instead of wipe', async () => {
await seedMeta(0);
// Pre-fix state: the server embedded 412 nodes but meta still says 0.
const stale = embeddingCountOf(await loadMeta(metaDir));
expect(stale).toBe(0);
expect(deriveEmbeddingMode({ force: true }, stale)).toMatchObject({
// `--force` rebuilds without loading the embedding cache → the 412
// server-generated vectors are destroyed.
shouldLoadCache: false,
preserveExistingEmbeddings: false,
});
const written = await finalizeEmbedRun(rowsWith(412), CLEAN_RUN);
const honest = embeddingCountOf(written);
expect(honest).toBe(412);
// Post-fix: `--force` loads the cache and regenerates on top of it rather
// than discarding the index. (`preserveExistingEmbeddings` is false here by
// design — `--force` upgrades to `forceRegenerateEmbeddings`; the wipe
// protection is `shouldLoadCache`.)
expect(deriveEmbeddingMode({ force: true }, honest)).toMatchObject({
shouldLoadCache: true,
forceRegenerateEmbeddings: true,
});
// A routine `analyze` preserves them outright.
expect(deriveEmbeddingMode({}, honest)).toMatchObject({
shouldLoadCache: true,
preserveExistingEmbeddings: true,
});
});
it('treats an unanswerable count query as unknown rather than 0', async () => {
// The query throws for reasons unrelated to how many rows were written.
await expect(
measurePersistedEmbeddingCount(async () => {
throw new Error('Connection closed');
}),
).resolves.toMatchObject({ kind: 'unknown', reason: 'Connection closed' });
// No row / no cell: an empty table would still answer with a 0.
await expect(measurePersistedEmbeddingCount(async () => [])).resolves.toMatchObject({
kind: 'unknown',
});
await expect(measurePersistedEmbeddingCount(async () => undefined)).resolves.toMatchObject({
kind: 'unknown',
});
// Non-numeric cell — same class of unknown.
await expect(measurePersistedEmbeddingCount(rowsWith('many'))).resolves.toMatchObject({
kind: 'unknown',
});
// A real zero is still a real answer.
await expect(measurePersistedEmbeddingCount(rowsWith(0))).resolves.toEqual({
kind: 'measured',
count: 0,
});
});
it('leaves the previous count alone when the measurement fails, never writing a fabricated 0', async () => {
await seedMeta(137);
const written = await finalizeEmbedRun(async () => {
throw new Error('Connection closed');
}, CLEAN_RUN);
expect(written).toMatchObject({ stats: { embeddings: 137 } });
// The dangerous direction is wrong-LOW: a fabricated 0 here would arm the
// wipe the test above describes.
expect(deriveEmbeddingMode({ force: true }, embeddingCountOf(written))).toMatchObject({
shouldLoadCache: true,
});
});
it('keeps the recovery marker when a clean run cannot verify its own count', async () => {
// The state that arms the silent wipe: meta records 0 embeddings (a repo
// analyzed without them, embedded through the server), the run succeeded,
// and the count query cannot answer — so no honest count can be stamped.
const midRunMarker = mintInterruptedCheckpoint(IDENTITY, PROGRESS, ['node-a']);
await seedMeta(0, midRunMarker);
const written = await finalizeEmbedRun(async () => {
throw new Error('Connection closed');
}, CLEAN_RUN);
// No fabricated value: neither a 0 nor a NaN/null lands in meta.
expect(written).toMatchObject({ stats: { nodes: 500, embeddings: 0 } });
// …and the marker this run wrote SURVIVES, so something on disk still
// records that embeddings were produced. Clearing it here would leave the
// index with zero evidence of its own embeddings.
expect(written?.embeddingCheckpoint).toMatchObject({
kind: 'interrupted',
pendingNodeIds: ['node-a'],
});
});
it('still clears the marker on an unverifiable run once meta records embeddings', async () => {
// Same unmeasurable run, but the recorded count already proves the index is
// accounted for — nothing needs preserving, so the clean-run contract wins.
await seedMeta(412, mintInterruptedCheckpoint(IDENTITY, PROGRESS, ['node-a']));
const written = await finalizeEmbedRun(async () => {
throw new Error('Connection closed');
}, CLEAN_RUN);
expect(written).toMatchObject({ stats: { embeddings: 412 } });
expect(written?.embeddingCheckpoint).toBeUndefined();
});
it('records the honest count on a partial run, alongside the pending checkpoint', async () => {
await seedMeta(0);
const written = await finalizeEmbedRun(rowsWith(300), {
nodesProcessed: 300,
chunksProcessed: 700,
failedNodeIds: ['node-a', 'node-b'],
});
// A partial index that is honest about itself survives the next run: the
// count keeps `--force` from wiping it, the checkpoint re-embeds the rest.
expect(written).toMatchObject({
stats: { embeddings: 300 },
embeddingCheckpoint: {
pendingNodeIds: ['node-a', 'node-b'],
nodesProcessed: 300,
kind: 'partial',
},
});
expect(deriveEmbeddingMode({ force: true }, embeddingCountOf(written))).toMatchObject({
shouldLoadCache: true,
});
});
});
/**
* Wiring guard for the route. Everything the helpers DECIDE is pinned
* behaviorally above; what remains is that the inline route closure inside
* `createServer` still asks them — the helper being right while the call site
* keeps writing `embeddingCheckpoint: undefined` is exactly the regression
* #2790 is about, and that closure cannot be reached without booting a server
* over a real repo + LadybugDB + embedding endpoint. Static-analysis layer of
* last resort, same precedent as api-readonly-wiring.test.ts.
*/
describe('POST /api/embed route wiring (#2790)', () => {
const readSource = () =>
fs.readFile(path.join(__dirname, '..', '..', 'src', 'server', 'api.ts'), 'utf-8');
/**
* The body of the route's `withLbugDb` callback — everything that may only
* run while the database connection is open. Sliced rather than matched with
* a character-distance regex so a comment edit cannot silently un-assert it.
*/
const insideWithLbugDb = (source: string): string => {
const start = source.indexOf('await withLbugDb(lbugPath, async () => {');
const end = source.indexOf('\n });', start);
expect(start).toBeGreaterThan(-1);
expect(end).toBeGreaterThan(start);
return source.slice(start, end);
};
it('feeds the pipeline result through resolveEmbedRunOutcome into the finalize write', async () => {
const source = await readSource();
// The result is captured, not discarded…
expect(source).toContain('const pipelineResult = await runEmbeddingPipeline(');
// …handed to the helper with the finalize context…
expect(source).toMatch(
/resolveEmbedRunOutcome\(\s*embeddingIdentity,\s*pipelineResult,\s*finalizeContext,\s*\)/,
);
// …and its checkpoint is what the finalize meta write persists (pre-fix: a
// hardcoded `embeddingCheckpoint: undefined`).
expect(source).toContain('embeddingCheckpoint: outcome.checkpoint');
expect(source).toContain('partialRunError = outcome.error;');
// A partial run does not reach `status: 'complete'`, and carries its detail.
expect(source).toMatch(
/partialRunError === undefined[\s\S]{0,400}status: 'complete'[\s\S]{0,600}status: 'failed'/,
);
expect(source).toContain('partial: partialRunDetail,');
});
it('measures after the WAL flush, inside withLbugDb, and folds the result into the write', async () => {
const source = await readSource();
const region = insideWithLbugDb(source);
// Inside the open connection — this is the route's only chance to stamp
// `stats.embeddings`, and the next CLI run's preserve-or-wipe decision
// hangs on it.
expect(region).toContain('const measuredEmbeddings = await countPersistedEmbeddings();');
expect(region).toContain('await saveMeta(entry.storagePath, embeddingMeta);');
// Ordering, without brittle character spans: flush → measure → decide →
// write. Counting before the flush would describe rows still in the WAL.
const flushed = region.lastIndexOf('await flushWAL();');
const measured = region.indexOf('const measuredEmbeddings = await countPersistedEmbeddings();');
const decided = region.indexOf('const outcome = resolveEmbedRunOutcome(');
const folded = region.indexOf('embeddingMeta = withMeasuredEmbeddingCount(', measured);
expect(flushed).toBeLessThan(measured);
expect(measured).toBeLessThan(decided);
expect(decided).toBeLessThan(folded);
expect(region.slice(folded)).toContain('measuredEmbeddings,');
});
it('measures in the post-flush checkpoint callback and nowhere else in the pipeline options', async () => {
const source = await readSource();
expect(source).toContain(
'await saveEmbeddingCheckpoint(checkpoint, [], await countPersistedEmbeddings());',
);
// The window-start callback fires before any row exists — it must pass no
// count rather than restate a stale one.
expect(source).toMatch(
/onCheckpointWindowStart: async \(\{ nodeIds, \.\.\.checkpoint \}\) => \{\s*await saveEmbeddingCheckpoint\(checkpoint, nodeIds\);\s*\},/,
);
});
it('resolves a found checkpoint through the shared resume decision', async () => {
const source = await readSource();
const region = insideWithLbugDb(source);
// The route asks the SAME decider the CLI does, instead of hard-throwing on
// any identity mismatch and ignoring `attempts` — the disagreement that let
// a CLI-written `'partial'` marker wedge every later `POST /api/embed`.
expect(region).toMatch(/decideEmbeddingResume\(priorCheckpoint, embeddingIdentity\)/);
// Every action is routed: abort fails the run, abandon warns and proceeds
// with an empty pending set, resume hands the decision's ids to the pipeline.
expect(region).toContain("if (resume?.action === 'abort') throw new Error(resume.error);");
expect(region).toMatch(/resume\?\.action === 'resume'\s*\?\s*resume\.pendingNodeIds/);
// No second copy of the gate: the route no longer authors its own message.
expect(region).not.toContain('Cannot resume embedding checkpoint:');
});
it('never maps the pipeline ready phase to a phase a client can read as terminal', async () => {
const source = await readSource();
// `ready` fires unconditionally before the route knows the outcome (#2790).
expect(source).toMatch(/p\.phase === 'ready'\s*\?\s*'finalizing'/);
expect(source).not.toMatch(/p\.phase === 'ready' \? 'complete'/);
});
});
+81 -1
View File
@@ -1,5 +1,9 @@
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { JobManager } from '../../src/server/analyze-job.js';
import {
JobManager,
isTerminalJobStatus,
type AnalyzeJobPartialOutcome,
} from '../../src/server/analyze-job.js';
describe('JobManager', () => {
let manager: JobManager;
@@ -171,4 +175,80 @@ describe('JobManager', () => {
expect(events[0].phase).toBe('complete');
});
});
/**
* #2790: terminality is a property of the job's STATUS, and a terminal
* `failed` job may still have persisted usable work. The SSE relay reads both
* — `isTerminalJobStatus` to decide the stream is over, `partial` to tell a
* client "retry these N nodes" apart from "nothing worked".
*/
describe('terminal status and partial outcomes (#2790)', () => {
it('recognizes exactly the two settled statuses', () => {
expect({
queued: isTerminalJobStatus('queued'),
cloning: isTerminalJobStatus('cloning'),
analyzing: isTerminalJobStatus('analyzing'),
loading: isTerminalJobStatus('loading'),
complete: isTerminalJobStatus('complete'),
failed: isTerminalJobStatus('failed'),
}).toEqual({
queued: false,
cloning: false,
analyzing: false,
loading: false,
complete: true,
failed: true,
});
});
const partial: AnalyzeJobPartialOutcome = {
kind: 'embedding-partial',
pendingNodeCount: 2,
nodesProcessed: 10,
};
it('records a partial outcome on a failed job without changing its status', () => {
const job = manager.createJob({ repoPath: '/tmp/embed-partial' });
manager.updateJob(job.id, { status: 'analyzing' });
manager.updateJob(job.id, {
status: 'failed',
error: 'Embedding generation finished partially: 2 node(s) lost their embeddings.',
partial,
});
expect(manager.getJob(job.id)).toMatchObject({
status: 'failed',
error: expect.stringContaining('finished partially'),
partial: { kind: 'embedding-partial', pendingNodeCount: 2, nodesProcessed: 10 },
});
});
it('leaves a clean job with no partial marker at all', () => {
const job = manager.createJob({ repoPath: '/tmp/embed-clean' });
manager.updateJob(job.id, { status: 'analyzing' });
manager.updateJob(job.id, { status: 'complete' });
// Absent, not `null`/`false` — JSON.stringify omits it, so the wire shape
// of every non-partial job is unchanged.
expect(manager.getJob(job.id)).toMatchObject({ status: 'complete' });
expect(manager.getJob(job.id)?.partial).toBeUndefined();
});
it('still emits exactly one terminal event when a partial outcome is attached', () => {
const job = manager.createJob({ repoPath: '/tmp/embed-partial' });
const events: Array<{ phase: string; message: string }> = [];
manager.updateJob(job.id, { status: 'analyzing' });
manager.onProgress(job.id, (data) => events.push(data));
manager.updateJob(job.id, {
status: 'failed',
error: 'Embedding generation finished partially: 2 node(s) lost their embeddings.',
partial,
progress: { phase: 'failed', percent: 100, message: 'partial' },
});
expect(events).toEqual([
{ phase: 'failed', percent: 100, message: expect.stringContaining('finished partially') },
]);
});
});
});
@@ -0,0 +1,169 @@
/**
* Tests for the undeclared FROM→TO label-pair failure path in the
* `analyzeCommand` CLI (#2789).
*
* `assertDeclaredPair` aborts the run when an extracted edge's endpoint-label
* pair is missing from GitNexus's own relation DDL — deliberately, because the
* alternative is a late `COPY` failure that silently drops edges. Before this
* branch existed the user got `Analysis failed` plus a stack trace through
* GitNexus internals: no file, no relationship, and an implicit "try again"
* that can never work. The CLI must instead name the pair, the relationship
* type and the offending file, and point at an issue report.
*
* The CLI does NOT compose that text: it prints `err.message` indented (the
* `LbugWipeError` idiom in the same catch block), because the message is
* self-contained — `gitnexus serve` forwards only `err.message` over worker
* IPC, so anything rendered here instead would be invisible to serve users.
* The needles below are therefore the SAME strings
* `test/unit/rel-pair-routing.test.ts` pins on the message itself.
*
* Mirrors analyze-http-endpoint-error.test.ts:
* - vi.mock the heavy dependencies so no real DB / git is touched
* - drive `analyzeCommand` with a mocked `runFullAnalysis` that rejects
* - assert on process.exitCode and the captured logger records
*/
import { beforeEach, describe, expect, it, vi } from 'vitest';
const runFullAnalysisMock = vi.fn();
vi.mock('../../src/core/run-analyze.js', () => ({
runFullAnalysis: runFullAnalysisMock,
}));
vi.mock('../../src/core/lbug/lbug-adapter.js', () => ({
closeLbug: vi.fn(async () => undefined),
closeLbugBeforeExit: vi.fn(async () => undefined),
isLbugReady: vi.fn(() => false),
LbugWipeError: class LbugWipeError extends Error {},
}));
vi.mock('../../src/storage/repo-manager.js', () => ({
getStoragePaths: vi.fn(() => ({ storagePath: '.gitnexus', lbugPath: '.gitnexus/lbug' })),
getGlobalRegistryPath: vi.fn(() => 'registry.json'),
RegistryNameCollisionError: class RegistryNameCollisionError extends Error {},
AnalysisNotFinalizedError: class AnalysisNotFinalizedError extends Error {},
assertAnalysisFinalized: vi.fn(async () => undefined),
}));
vi.mock('../../src/storage/git.js', () => ({
getGitRoot: vi.fn(() => '/repo'),
hasGitDir: vi.fn(() => true),
}));
vi.mock('../../src/core/ingestion/utils/max-file-size.js', () => ({
getMaxFileSizeBannerMessage: vi.fn(() => null),
}));
// analyze.ts imports isHfDownloadFailure from hf-env.js — mock it to break the
// transitive gitnexus-shared chain (same reason as the sibling suite).
vi.mock('../../src/core/embeddings/hf-env.js', () => ({
isHfDownloadFailure: vi.fn(() => false),
}));
const PAIR_ERROR_ARGS = [
'Method|Annotation',
'ANNOTATED_BY',
'Method:src/main/java/app/BeanConfig.java:BeanConfig.dataSource#42',
'Annotation:src/main/java/app/BeanConfig.java:ConditionalOnMissingBean',
] as const;
describe('analyzeCommand undeclared relation-pair handling (#2789)', () => {
beforeEach(() => {
vi.resetModules();
runFullAnalysisMock.mockReset();
process.exitCode = undefined;
process.env.NODE_OPTIONS = `${process.env.NODE_OPTIONS ?? ''} --max-old-space-size=8192`.trim();
});
it('renders an actionable schema-gap message naming the pair, relationship, ids and file', async () => {
const { UndeclaredRelationPairError } = await import('../../src/core/lbug/rel-pair-routing.js');
runFullAnalysisMock.mockRejectedValue(new UndeclaredRelationPairError(...PAIR_ERROR_ARGS));
const { _captureLogger } = await import('../../src/core/logger.js');
const cap = _captureLogger();
const { analyzeCommand } = await import('../../src/cli/analyze.js');
await analyzeCommand(undefined, {});
expect(process.exitCode).toBe(1);
const record = cap.records().find((r) => r.recoveryHint === 'undeclared-relation-pair');
cap.restore();
expect(record).toMatchObject({
recoveryHint: 'undeclared-relation-pair',
labelPair: 'Method|Annotation',
relationType: 'ANNOTATED_BY',
sourceFile: 'src/main/java/app/BeanConfig.java',
});
// The CLI renders `err.message` indented (the `LbugWipeError` idiom) rather
// than re-formatting the structured fields, so these needles are the ONE
// wording — `test/unit/rel-pair-routing.test.ts` pins the same strings on
// the message itself and a reword updates one place, not two.
// Filter-to-empty, not an array of booleans: the failure output NAMES the
// missing string instead of making you count positions.
const text = typeof record?.msg === 'string' ? record.msg : '';
const required = [
'Method → Annotation',
'ANNOTATED_BY',
'src/main/java/app/BeanConfig.java',
'Method:src/main/java/app/BeanConfig.java:BeanConfig.dataSource#42',
'Annotation:src/main/java/app/BeanConfig.java:ConditionalOnMissingBean',
// Names it as a GitNexus gap, tells the user a re-run is pointless, and
// points at both actionable next steps.
"gap in GitNexus's own relation schema",
're-running the analysis will fail',
'https://github.com/abhigyanpatwari/GitNexus/issues/new',
'.gitnexusignore',
];
expect(required.filter((needle) => !text.includes(needle))).toEqual([]);
});
it('still fires when the ingestion phase runner has rewrapped it as a cause', async () => {
// The guard throws inside an emit phase, and the phase runner rewraps every
// phase failure as `new Error("Phase 'X' failed: …", { cause })` — a bare
// instanceof check at the CLI boundary would miss it entirely.
const { UndeclaredRelationPairError } = await import('../../src/core/lbug/rel-pair-routing.js');
const original = new UndeclaredRelationPairError(...PAIR_ERROR_ARGS);
// The wrapper deliberately does NOT interpolate the cause's message: the
// branch must render `undeclaredPair.message`, i.e. the message of the link
// it FOUND in the chain, not the outer wrapper's message.
runFullAnalysisMock.mockRejectedValue(
new Error(`Phase 'graph-emit' failed`, { cause: original }),
);
const { _captureLogger } = await import('../../src/core/logger.js');
const cap = _captureLogger();
const { analyzeCommand } = await import('../../src/cli/analyze.js');
await analyzeCommand(undefined, {});
expect(process.exitCode).toBe(1);
const records = cap.records();
cap.restore();
const record = records.find((r) => r.recoveryHint === 'undeclared-relation-pair');
expect(record).toMatchObject({ labelPair: 'Method|Annotation' });
const text = typeof record?.msg === 'string' ? record.msg : '';
expect(
['Method → Annotation', '.gitnexusignore'].filter((needle) => !text.includes(needle)),
).toEqual([]);
// The generic large-repo / module-not-found guidance must not also appear.
expect(records.some((r) => r.recoveryHint === 'large-repo')).toBe(false);
expect(records.some((r) => r.recoveryHint === 'module-not-found')).toBe(false);
});
it('does not claim an unrelated failure', async () => {
runFullAnalysisMock.mockRejectedValue(new Error('LadybugDB write failed'));
const { _captureLogger } = await import('../../src/core/logger.js');
const cap = _captureLogger();
const { analyzeCommand } = await import('../../src/cli/analyze.js');
await analyzeCommand(undefined, {});
expect(process.exitCode).toBe(1);
const records = cap.records();
cap.restore();
expect(records.some((r) => r.recoveryHint === 'undeclared-relation-pair')).toBe(false);
});
});
@@ -0,0 +1,311 @@
/**
* Locally linked dev dependencies in the analyzer identity receipt (#2798).
*
* `dependencyNames` admits a devDependency whose declared SPECIFIER is
* checkout-local (`file:`/`link:`/`workspace:`/`portal:` and bare local paths).
* `npm link <pkg>` leaves the specifier a registry range and only replaces the
* `node_modules` entry with a symlink into a checkout, so the specifier check is
* blind to it while the linked code is just as load-bearing for analyzer
* semantics as a declared `file:` sibling.
*
* The second, RESOLVED-LOCATION half closes that: the resolver already returns a
* realpath, so a linked package reports a root carrying no `node_modules`
* segment. These tests pin the three properties that make it affordable and
* safe — root-only scoping, the pnpm-store exclusion, and the admission cap —
* plus the declared half it does not replace.
*/
import { mkdir, rm, symlink, writeFile } from 'node:fs/promises';
import path from 'node:path';
import { pathToFileURL } from 'node:url';
import { describe, expect, it } from 'vitest';
import {
_clearAnalyzerIdentityProcessCacheForTests,
_hasNodeModulesSegmentForTests,
resolveAnalyzerRunnerIdentity,
} from '../../src/core/analyzer-identity.js';
import type { AnalyzerRunnerIdentity } from '../../src/storage/repo-manager.js';
import { createTempDir } from '../helpers/test-db.js';
type Fixture = { root: string; modulePath: string };
type FixtureOptions = {
/** Root `devDependencies`, verbatim. */
devDependencies?: Record<string, string>;
/** `devDependencies` for the nested runtime dependency (root-only scoping). */
nestedDevDependencies?: Record<string, string>;
};
/**
* A package root with one ordinary resolvable runtime dependency, so every
* fixture starts from `packageCount: 2` (root + `runtime-package`).
*/
async function createFixture(root: string, options: FixtureOptions = {}): Promise<Fixture> {
const modulePath = path.join(root, 'src', 'core', 'analyzer.ts');
const runtimeRoot = path.join(root, 'node_modules', 'runtime-package');
await mkdir(path.dirname(modulePath), { recursive: true });
await mkdir(runtimeRoot, { recursive: true });
await writeFile(
path.join(root, 'package.json'),
JSON.stringify({
name: 'fixture-analyzer',
version: '9.8.7',
dependencies: { 'runtime-package': '1.0.0' },
...(options.devDependencies ? { devDependencies: options.devDependencies } : {}),
}),
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
await writeFile(
path.join(runtimeRoot, 'package.json'),
JSON.stringify({
name: 'runtime-package',
version: '1.0.0',
...(options.nestedDevDependencies ? { devDependencies: options.nestedDevDependencies } : {}),
}),
);
await writeFile(path.join(runtimeRoot, 'runtime.js'), 'export const runtime = 1;\n');
return { root, modulePath };
}
/** A checkout-local package: a real directory OUTSIDE any `node_modules` tree. */
async function createCheckout(root: string, name: string, payload: string): Promise<string> {
const checkout = path.join(root, 'checkouts', name);
await mkdir(checkout, { recursive: true });
await writeFile(path.join(checkout, 'package.json'), JSON.stringify({ name, version: '1.0.0' }));
await writeFile(path.join(checkout, 'tool.js'), payload);
return checkout;
}
/**
* Resolve cold. Each call gets its own cache directory AND drops the in-process
* LRU, whose key does not include the cache directory — without both, a second
* resolution in the same test would echo the first receipt instead of
* recomputing it, which is precisely what these assertions must not do.
*/
function resolveCold(
fixture: Fixture,
run: number,
onGuardCount?: (guardCount: number) => void,
): AnalyzerRunnerIdentity {
_clearAnalyzerIdentityProcessCacheForTests();
return resolveAnalyzerRunnerIdentity(pathToFileURL(fixture.modulePath).href, {
cacheDirectory: path.join(fixture.root, `identity-cache-${run}`),
onCacheValidationPass: ({ guardCount }) => onGuardCount?.(guardCount),
});
}
describe('analyzer identity resolved-location dev dependencies (#2798)', () => {
// Every case here needs a symbolic link to exist; Windows runners without the
// developer-mode privilege cannot create one, so the whole block is skipped
// rather than branching inside test bodies.
describe.skipIf(process.platform === 'win32')('npm link shape', () => {
it('admits a registry-specifier dev dependency symlinked to a checkout', async () => {
const temp = await createTempDir();
try {
const fixture = await createFixture(temp.dbPath, {
// A registry RANGE: `isLocallyLinkedSpecifier` rejects this, so only
// the resolved-location half can admit the package.
devDependencies: { 'linked-tool': '^1.0.0' },
});
const checkout = await createCheckout(temp.dbPath, 'linked-tool', 'export const v = 1;\n');
await symlink(checkout, path.join(temp.dbPath, 'node_modules', 'linked-tool'), 'dir');
const first = resolveCold(fixture, 1);
// root + runtime-package + the linked checkout.
expect(first.dependencyRuntime.packageCount).toBe(3);
// The regression this closes: a SEMANTIC-ONLY edit inside the linked
// checkout moved neither digest before the resolved-location half.
await writeFile(path.join(checkout, 'tool.js'), 'export const v = 2;\n');
const second = resolveCold(fixture, 2);
expect(second.dependencyRuntime.digest).not.toBe(first.dependencyRuntime.digest);
} finally {
await temp.cleanup();
}
});
it('excludes a pnpm virtual-store link that stays inside node_modules', async () => {
const temp = await createTempDir();
try {
const fixture = await createFixture(temp.dbPath, {
devDependencies: { 'pnpm-tool': '^1.0.0' },
});
const store = path.join(
temp.dbPath,
'node_modules',
'.pnpm',
'pnpm-tool@1.0.0',
'node_modules',
'pnpm-tool',
);
await mkdir(store, { recursive: true });
await writeFile(
path.join(store, 'package.json'),
JSON.stringify({ name: 'pnpm-tool', version: '1.0.0' }),
);
await writeFile(path.join(store, 'tool.js'), 'export const v = 1;\n');
// pnpm's own shape: `node_modules/<pkg>` IS a symlink, but it points
// back inside `node_modules`, so the realpath keeps the segment.
await symlink(
path.join('.pnpm', 'pnpm-tool@1.0.0', 'node_modules', 'pnpm-tool'),
path.join(temp.dbPath, 'node_modules', 'pnpm-tool'),
'dir',
);
const first = resolveCold(fixture, 1);
expect(first.dependencyRuntime.packageCount).toBe(2);
await writeFile(path.join(store, 'tool.js'), 'export const v = 2;\n');
const second = resolveCold(fixture, 2);
expect(second.dependencyRuntime.digest).toBe(first.dependencyRuntime.digest);
} finally {
await temp.cleanup();
}
});
it('scopes the resolved-location probe to the root package', async () => {
const bare = await createTempDir();
const nested = await createTempDir();
try {
// Identical trees except that `runtime-package` — a NON-root package —
// declares dev-only names that resolve to checkout-local siblings. If
// the root-only scope is ever dropped, probing them costs extra path
// guards (re-probed on every warm validation) and folds three more
// packages into the receipt. Comparing the two runs pins the scope
// without hard-coding a guard total that unrelated work would churn.
const bareFixture = await createFixture(bare.dbPath);
const nestedFixture = await createFixture(nested.dbPath, {
nestedDevDependencies: {
'nested-a': '^1.0.0',
'nested-b': '^1.0.0',
'nested-c': '^1.0.0',
},
});
for (const name of ['nested-a', 'nested-b', 'nested-c']) {
const checkout = await createCheckout(nested.dbPath, name, 'export const v = 1;\n');
await symlink(checkout, path.join(nested.dbPath, 'node_modules', name), 'dir');
}
let bareGuards = 0;
let nestedGuards = 0;
const bareIdentity = resolveCold(bareFixture, 1, (count) => {
bareGuards = count;
});
const nestedIdentity = resolveCold(nestedFixture, 1, (count) => {
nestedGuards = count;
});
expect({
guards: nestedGuards,
packages: nestedIdentity.dependencyRuntime.packageCount,
}).toEqual({ guards: bareGuards, packages: bareIdentity.dependencyRuntime.packageCount });
// Pin the shared value too, so an accidental collapse to zero guards on
// both sides cannot make the comparison vacuous.
expect(bareGuards).toBeGreaterThan(0);
} finally {
await bare.cleanup();
await nested.cleanup();
}
});
it('disables the resolved-location channel past the admission cap', async () => {
const under = await createTempDir();
const over = await createTempDir();
try {
const names = ['tool-a', 'tool-b', 'tool-c', 'tool-d', 'tool-e'];
const link = async (root: string, count: number): Promise<void> => {
for (const name of names.slice(0, count)) {
const checkout = await createCheckout(root, name, 'export const v = 1;\n');
await symlink(checkout, path.join(root, 'node_modules', name), 'dir');
}
};
const devDependencies = (count: number): Record<string, string> =>
Object.fromEntries(names.slice(0, count).map((name) => [name, '^1.0.0']));
const underFixture = await createFixture(under.dbPath, {
devDependencies: devDependencies(4),
});
await link(under.dbPath, 4);
const overFixture = await createFixture(over.dbPath, {
devDependencies: devDependencies(5),
});
await link(over.dbPath, 5);
// At the cap every link is admitted; one past it the channel is dropped
// WHOLESALE rather than admitting an arbitrary prefix, because a
// mis-firing proxy folds the entire dev tree in and
// runtimePackages/runtimeEntries/runtimeBytes THROW rather than degrade.
expect({
under: resolveCold(underFixture, 1).dependencyRuntime.packageCount,
over: resolveCold(overFixture, 1).dependencyRuntime.packageCount,
}).toEqual({ under: 6, over: 2 });
} finally {
await under.cleanup();
await over.cleanup();
}
});
});
it('keeps enumerating a declared file: dev link whose checkout is absent', async () => {
const temp = await createTempDir();
try {
const fixture = await createFixture(temp.dbPath, {
devDependencies: { 'declared-link': 'file:./declared-link' },
});
// Nothing resolves: the declared half is the only thing that enumerates
// the name at all, and it contributes a `<missing>` edge.
const absent = resolveCold(fixture, 1);
expect(absent.dependencyRuntime.packageCount).toBe(2);
// Materialize it as a real DIRECTORY under node_modules — a copied
// `file:` install. Its realpath still carries a `node_modules` segment,
// so the resolved-location half provably cannot admit it and this
// transition isolates the declared half.
const materialized = path.join(temp.dbPath, 'node_modules', 'declared-link');
await mkdir(materialized, { recursive: true });
await writeFile(
path.join(materialized, 'package.json'),
JSON.stringify({ name: 'declared-link', version: '1.0.0' }),
);
await writeFile(path.join(materialized, 'tool.js'), 'export const v = 1;\n');
const present = resolveCold(fixture, 2);
expect(present.dependencyRuntime.packageCount).toBe(3);
expect(present.dependencyRuntime.digest).not.toBe(absent.dependencyRuntime.digest);
// Removing it returns to the `<missing>` receipt rather than to a
// silently-dropped name.
await rm(materialized, { recursive: true });
const removed = resolveCold(fixture, 3);
expect(removed.dependencyRuntime.digest).toBe(absent.dependencyRuntime.digest);
} finally {
await temp.cleanup();
}
});
it('treats node_modules as a whole path segment, per platform separator', () => {
expect({
checkout: _hasNodeModulesSegmentForTests('/home/u/checkouts/tool', path.posix),
installed: _hasNodeModulesSegmentForTests('/home/u/app/node_modules/tool', path.posix),
substring: _hasNodeModulesSegmentForTests('/home/u/node_modules_old/tool', path.posix),
nested: _hasNodeModulesSegmentForTests(
'/a/node_modules/.pnpm/x@1/node_modules/x',
path.posix,
),
// `\` is a legal POSIX filename character, so it is NOT a boundary there
// — but it is the separator win32 realpaths come back with.
posixBackslash: _hasNodeModulesSegmentForTests('/a/node_modules\\x/tool', path.posix),
win32Backslash: _hasNodeModulesSegmentForTests('C:\\app\\node_modules\\tool', path.win32),
win32Substring: _hasNodeModulesSegmentForTests('C:\\app\\node_modulesx\\tool', path.win32),
}).toEqual({
checkout: false,
installed: true,
substring: false,
nested: true,
posixBackslash: false,
win32Backslash: true,
win32Substring: false,
});
});
});
@@ -0,0 +1,292 @@
/**
* Symbolic-link handling in the analyzer runtime-payload scan (#2798).
*
* `collectArtifacts` used to fuse two unrelated facts into one condition:
* "this name is never runtime payload" and "a symlink must not reach the
* file-payload branch". Only the four pruned names got the symlink half, so any
* OTHER symlinked directory inside a scanned package root — `dist -> build`, a
* vendored-grammar link, anything in a workspace-linked sibling checkout — fell
* through to `snapshotReadableFile`, which stats the target, sees a directory,
* and throws `Analyzer identity input is not a file`, aborting the whole
* analyze. Workspace-linked packages became scannable on this branch, so the
* crash is newly reachable (this worktree's own `gitnexus-shared/node_modules`
* is a symlink).
*
* These tests pin the split: prune by NAME alone, and route every symlink that
* does not resolve to a regular file into a link-text artifact instead of the
* payload branch.
*/
import { mkdir, symlink, unlink, writeFile } from 'node:fs/promises';
import path from 'node:path';
import { pathToFileURL } from 'node:url';
import { describe, expect, it } from 'vitest';
import {
_clearAnalyzerIdentityProcessCacheForTests,
resolveAnalyzerRunnerIdentity,
} from '../../src/core/analyzer-identity.js';
import { createTempDir } from '../helpers/test-db.js';
type Fixture = {
root: string;
modulePath: string;
cacheDirectory: string;
packageRoot: string;
};
/** A package root with one resolvable dependency whose payload tree we mutate. */
async function createFixture(root: string): Promise<Fixture> {
const modulePath = path.join(root, 'src', 'core', 'analyzer.ts');
const packageRoot = path.join(root, 'node_modules', 'runtime-package');
await mkdir(path.dirname(modulePath), { recursive: true });
await mkdir(path.join(packageRoot, 'build'), { recursive: true });
await writeFile(
path.join(root, 'package.json'),
JSON.stringify({
name: 'fixture-analyzer',
version: '9.8.7',
dependencies: { 'runtime-package': '1.0.0' },
}),
);
await writeFile(modulePath, 'export const analyzer = 1;\n');
await writeFile(
path.join(packageRoot, 'package.json'),
JSON.stringify({ name: 'runtime-package', version: '1.0.0' }),
);
await writeFile(path.join(packageRoot, 'runtime.js'), 'export const runtime = 1;\n');
await writeFile(path.join(packageRoot, 'build', 'native.node'), 'native-v1');
return { root, modulePath, cacheDirectory: path.join(root, 'identity-cache'), packageRoot };
}
/**
* Create a symbolic link, reporting whether the platform allowed it. Windows
* runners without the developer-mode privilege cannot create links at all;
* mirrors the guard used by the sibling analyzer-identity suite.
*/
async function trySymlink(
target: string,
linkPath: string,
type: 'dir' | 'file',
): Promise<boolean> {
try {
await symlink(target, linkPath, type);
return true;
} catch (error) {
if (['EPERM', 'EACCES'].includes((error as NodeJS.ErrnoException).code ?? '')) return false;
throw error;
}
}
describe('analyzer identity runtime-payload symbolic links (#2798)', () => {
it('records a symlinked directory instead of aborting the scan', async () => {
const temp = await createTempDir();
try {
const fixture = await createFixture(temp.dbPath);
// NOT one of the four pruned names: this is the case that used to throw
// `Analyzer identity input is not a file` and abort the entire analyze.
const linked = await trySymlink(
path.join(fixture.packageRoot, 'build'),
path.join(fixture.packageRoot, 'dist'),
'dir',
);
if (!linked) return;
const identity = resolveAnalyzerRunnerIdentity(pathToFileURL(fixture.modulePath).href, {
cacheDirectory: fixture.cacheDirectory,
});
// runtime.js + build/native.node + the recorded `dist` link.
expect(identity.dependencyRuntime).toMatchObject({
packageCount: 2,
artifactCount: 3,
digest: expect.stringMatching(/^sha256:[a-f0-9]{64}$/),
});
} finally {
await temp.cleanup();
}
});
it('moves the receipt when a recorded directory link is retargeted', async () => {
const temp = await createTempDir();
try {
const fixture = await createFixture(temp.dbPath);
await mkdir(path.join(fixture.packageRoot, 'build-next'));
await writeFile(path.join(fixture.packageRoot, 'build-next', 'native.node'), 'native-v1');
const linkPath = path.join(fixture.packageRoot, 'dist');
if (!(await trySymlink(path.join(fixture.packageRoot, 'build'), linkPath, 'dir'))) return;
const first = resolveAnalyzerRunnerIdentity(pathToFileURL(fixture.modulePath).href, {
cacheDirectory: fixture.cacheDirectory,
});
await unlink(linkPath);
await symlink(path.join(fixture.packageRoot, 'build-next'), linkPath, 'dir');
_clearAnalyzerIdentityProcessCacheForTests();
const retargeted = resolveAnalyzerRunnerIdentity(pathToFileURL(fixture.modulePath).href, {
cacheDirectory: fixture.cacheDirectory,
});
// Both targets hold byte-identical payloads, so only the link TEXT
// distinguishes them. Recording it is what keeps the retarget visible.
expect(retargeted.dependencyRuntime.digest).not.toBe(first.dependencyRuntime.digest);
} finally {
await temp.cleanup();
}
});
it('reuses the warm cache for a recorded directory link', async () => {
const temp = await createTempDir();
try {
const fixture = await createFixture(temp.dbPath);
const linked = await trySymlink(
path.join(fixture.packageRoot, 'build'),
path.join(fixture.packageRoot, 'dist'),
'dir',
);
if (!linked) return;
const cold = resolveAnalyzerRunnerIdentity(pathToFileURL(fixture.modulePath).href, {
cacheDirectory: fixture.cacheDirectory,
});
// Drop the in-process reuse so the second call must load, validate, and
// accept the persisted cache — including the link artifact's guard.
_clearAnalyzerIdentityProcessCacheForTests();
let work = 0;
let hashes = 0;
const warm = resolveAnalyzerRunnerIdentity(pathToFileURL(fixture.modulePath).href, {
cacheDirectory: fixture.cacheDirectory,
onCacheMissWork: () => {
work += 1;
},
onHashedInput: () => {
hashes += 1;
},
});
expect(warm).toEqual(cold);
expect({ work, hashes }).toEqual({ work: 0, hashes: 0 });
} finally {
await temp.cleanup();
}
});
it('records a dangling link rather than failing the whole analyze', async () => {
const temp = await createTempDir();
try {
const fixture = await createFixture(temp.dbPath);
const linkPath = path.join(fixture.packageRoot, 'dangling.js');
if (!(await trySymlink(path.join(fixture.packageRoot, 'absent.js'), linkPath, 'file')))
return;
const identity = resolveAnalyzerRunnerIdentity(pathToFileURL(fixture.modulePath).href, {
cacheDirectory: fixture.cacheDirectory,
});
expect(identity.dependencyRuntime.artifactCount).toBe(3);
// Creating the target promotes the link to a content-hashed payload.
await writeFile(path.join(fixture.packageRoot, 'absent.js'), 'export const late = 1;\n');
_clearAnalyzerIdentityProcessCacheForTests();
const resolved = resolveAnalyzerRunnerIdentity(pathToFileURL(fixture.modulePath).href, {
cacheDirectory: fixture.cacheDirectory,
});
expect(resolved.dependencyRuntime.digest).not.toBe(identity.dependencyRuntime.digest);
} finally {
await temp.cleanup();
}
});
it('does not follow a self-referential link into the depth limit', async () => {
const temp = await createTempDir();
try {
const fixture = await createFixture(temp.dbPath);
// Following links would recurse here until `runtimeDepth` THREW — trading
// one hard abort for another. Recording the link text is cycle-free.
if (!(await trySymlink(fixture.packageRoot, path.join(fixture.packageRoot, 'self'), 'dir'))) {
return;
}
const identity = resolveAnalyzerRunnerIdentity(pathToFileURL(fixture.modulePath).href, {
cacheDirectory: fixture.cacheDirectory,
traversalLimits: { runtimeDepth: 4 },
});
expect(identity.dependencyRuntime.artifactCount).toBe(3);
} finally {
await temp.cleanup();
}
});
it('still hashes the target content behind a link to a regular file', async () => {
const temp = await createTempDir();
try {
const fixture = await createFixture(temp.dbPath);
const target = path.join(fixture.packageRoot, 'build', 'native.node');
if (!(await trySymlink(target, path.join(fixture.packageRoot, 'linked.node'), 'file')))
return;
const first = resolveAnalyzerRunnerIdentity(pathToFileURL(fixture.modulePath).href, {
cacheDirectory: fixture.cacheDirectory,
});
expect(first.dependencyRuntime.artifactCount).toBe(3);
// Only the TARGET's bytes change; the link text and its lstat are
// untouched. A link-text-only recording would go blind here, so this is
// the guard that the file-payload branch still owns resolvable links.
await writeFile(target, 'native-v2-changed');
_clearAnalyzerIdentityProcessCacheForTests();
const changed = resolveAnalyzerRunnerIdentity(pathToFileURL(fixture.modulePath).href, {
cacheDirectory: fixture.cacheDirectory,
});
expect(changed.dependencyRuntime.digest).not.toBe(first.dependencyRuntime.digest);
} finally {
await temp.cleanup();
}
});
it('prunes the VCS/nested-install names by name alone, whatever their type', async () => {
const temp = await createTempDir();
try {
const fixture = await createFixture(temp.dbPath);
// A `.git` FILE is what a submodule or linked worktree checkout carries;
// it is a gitdir pointer, never analyzer payload, and it churns whenever
// the checkout moves. Pruning on the name alone keeps it out.
const gitPointer = path.join(fixture.packageRoot, '.git');
await writeFile(gitPointer, 'gitdir: /elsewhere/.git/worktrees/one\n');
const first = resolveAnalyzerRunnerIdentity(pathToFileURL(fixture.modulePath).href, {
cacheDirectory: fixture.cacheDirectory,
});
expect(first.dependencyRuntime.artifactCount).toBe(2);
await writeFile(gitPointer, 'gitdir: /moved/.git/worktrees/two\n');
_clearAnalyzerIdentityProcessCacheForTests();
const moved = resolveAnalyzerRunnerIdentity(pathToFileURL(fixture.modulePath).href, {
cacheDirectory: fixture.cacheDirectory,
});
expect(moved.dependencyRuntime.digest).toBe(first.dependencyRuntime.digest);
} finally {
await temp.cleanup();
}
});
it('keeps pruning a linked node_modules tree it never owned', async () => {
const temp = await createTempDir();
try {
const fixture = await createFixture(temp.dbPath);
const shared = path.join(temp.dbPath, 'shared-store');
await mkdir(path.join(shared, 'nested'), { recursive: true });
await writeFile(path.join(shared, 'nested', 'payload.js'), 'export const nested = 1;\n');
// The shape this worktree ships: a workspace checkout whose
// `node_modules` is a symbolic link into a shared store.
if (!(await trySymlink(shared, path.join(fixture.packageRoot, 'node_modules'), 'dir')))
return;
const identity = resolveAnalyzerRunnerIdentity(pathToFileURL(fixture.modulePath).href, {
cacheDirectory: fixture.cacheDirectory,
});
expect(identity.dependencyRuntime.artifactCount).toBe(2);
} finally {
await temp.cleanup();
}
});
});
@@ -76,6 +76,16 @@ describe('analyzer runner identity', () => {
expect(second.invokedArtifact.digest).toBe(first.invokedArtifact.digest);
expect(second.build.digest).not.toBe(first.build.digest);
expect(second.dependencyRuntime.digest).toBe(first.dependencyRuntime.digest);
// THE #2798 INVARIANT: the build digest moved while nothing else did.
// Node-id formats, wire formats, resolution tiers and emit ordering live in
// analyzer code, not in DDL, so `SCHEMA_FINGERPRINT` (lbug/schema.ts) is
// structurally incapable of firing on a change shaped like this one — it is
// a digest of the node+relation DDL and of nothing else. #2798 deleted the
// hand-incremented INCREMENTAL_SCHEMA_VERSION ladder, and roughly 30 of its
// ~35 bumps were exactly this shape: semantic, no DDL. This receipt is their
// only remaining cover, so a moved build digest MUST refuse index reuse.
// (call-summary-schema-version.test.ts holds the DDL-blind half of the split.)
expect(analyzerRunnerIdentitiesEqual(second, first)).toBe(false);
} finally {
await fixture.cleanup();
}
@@ -1282,6 +1292,12 @@ describe('analyzer runner identity', () => {
identity,
),
).toBe(false);
// Fail-closed on a receipt that cannot be read at all — the same posture as
// an absent schemaFingerprint. An index predating the field stamps nothing
// (undefined) and a cleared/legacy field reads back as null; neither is ever
// grandfathered into an incremental top-up (#2798).
expect(analyzerRunnerIdentitiesEqual(undefined, identity)).toBe(false);
expect(analyzerRunnerIdentitiesEqual(null, identity)).toBe(false);
await writeFile(
path.join(sourceRoot, 'new-semantic-input.ts'),
@@ -9,7 +9,7 @@
* 3. bulk COPY column list (getCopyQuery('BasicBlock'))
* 4. single-node CREATE (insertNodeToLbug)
* 5. incremental MERGE (batchInsertNodesToLbug)
* plus the INCREMENTAL_SCHEMA_VERSION 2 → 3 bump (KTD5).
* plus its presence in the fingerprinted DDL set (KTD5).
*
* `calleeIds` is added LAST in the CSV/COPY/CREATE/MERGE tuple, so the column
* order MUST stay identical across header, COPY list, and row array — the
@@ -24,8 +24,7 @@ import { afterEach, describe, expect, it, vi } from 'vitest';
import type { GraphNode, NodeProperties } from 'gitnexus-shared';
import { BASICBLOCK_CSV_HEADER, buildBasicBlockRow } from '../../src/core/lbug/csv-generator.js';
import { getCopyQuery } from '../../src/core/lbug/lbug-adapter.js';
import { BASICBLOCK_SCHEMA } from '../../src/core/lbug/schema.js';
import { INCREMENTAL_SCHEMA_VERSION } from '../../src/storage/repo-manager.js';
import { BASICBLOCK_SCHEMA, NODE_SCHEMA_QUERIES } from '../../src/core/lbug/schema.js';
// ── helpers ─────────────────────────────────────────────────────────────────
@@ -117,19 +116,21 @@ describe('BasicBlock calleeIds — header/COPY/row column parity', () => {
});
});
// ── 3. schema DDL + incremental version bump (pure) ───────────────────────────
// ── 3. schema DDL + fingerprint coverage (pure) ───────────────────────────────
describe('BasicBlock calleeIds — schema DDL + version bump', () => {
describe('BasicBlock calleeIds — schema DDL + fingerprint coverage', () => {
it('BASICBLOCK_SCHEMA declares the calleeIds STRING column', () => {
expect(BASICBLOCK_SCHEMA).toContain('callees STRING');
expect(BASICBLOCK_SCHEMA).toContain('calleeIds STRING');
});
it('INCREMENTAL_SCHEMA_VERSION is at least 3 (calleeIds column bump, KTD5)', () => {
// The exact value advances as later milestones add re-index-forcing changes
// (v4 = CALL_SUMMARY, PDG FU-C). This guard pins the floor the calleeIds
// column established; the v3→4 reuse-gate guard lives in its own test.
expect(INCREMENTAL_SCHEMA_VERSION).toBeGreaterThanOrEqual(3);
it('the calleeIds column is inside the schema fingerprint, so a pre-column index cannot be reused', () => {
// This replaces the old `INCREMENTAL_SCHEMA_VERSION >= 3` floor (#2798).
// The hand-incremented integer is gone: reuse is now gated on a digest of
// the DDL itself, so the invariant to pin is that BASICBLOCK_SCHEMA is one
// of the strings that digest covers. An index built before the column
// existed therefore carries a different fingerprint and is rebuilt.
expect(NODE_SCHEMA_QUERIES).toContain(BASICBLOCK_SCHEMA);
});
});
@@ -1,18 +1,37 @@
/**
* PDG FU-C (U-C1 / U-C5) — CALL_SUMMARY relation-type posture + the v3→4
* incremental reuse gate.
* PDG FU-C (U-C1) — CALL_SUMMARY relation-type posture — plus the index-reuse
* gates that decide whether an existing index may be topped up incrementally
* (U-C5, #2798).
*
* CALL_SUMMARY is an INTERNAL PDG-engine edge: like the taint substrate edges
* (TAINTED / TAINT_PATH / CDG / REACHING_DEF / CFG) it must stay OUT of
* `VALID_RELATION_TYPES` so it never enters impact-style symbol-space traversal,
* and the impact relType allowlists (local-backend.ts ~:4373 / ~:5674) that gate
* on `VALID_RELATION_TYPES` therefore never surface it. The v4 bump forces a
* full re-analyze on a pre-v4 index (which has no CALL_SUMMARY edges, so an
* incremental top-up would silently under-report return-value ascent).
* on `VALID_RELATION_TYPES` therefore never surface it.
*
* The reuse gates below are split by what each one can SEE, and that split is
* the point of this file:
*
* • `SCHEMA_FINGERPRINT` (lbug/schema.ts) is a digest of the node + relation
* DDL. It fires exactly when a table shape changes — and is structurally
* blind to everything else.
* • the analyzer runner-identity receipt (analyzer-identity.ts) hashes the
* analyzer BUILD, so it — and only it — covers SEMANTIC changes that touch
* no DDL: node-id formats, wire formats, resolution tiers, emit ordering.
*
* That second gate became load-bearing in #2798. The hand-incremented
* `INCREMENTAL_SCHEMA_VERSION` it replaced was bumped ~35 times, and roughly 30
* of those bumps changed NO DDL — they were semantic. A DDL digest cannot fire
* on any of them. The runner-identity receipt is their only remaining cover, so
* this file names that split instead of leaving it implicit: it owns the
* DDL-blind half (the fingerprint below) plus a source anchor proving
* run-analyze.ts still consults the receipt. The receipt predicate's own
* behaviour is asserted against the real function in analyzer-identity.test.ts.
*/
import { describe, it, expect } from 'vitest';
import { readFileSync } from 'node:fs';
import { createHash } from 'node:crypto';
import { fileURLToPath } from 'node:url';
import path from 'node:path';
import {
@@ -20,11 +39,18 @@ import {
EPISTEMIC_HERITAGE_RELATION_TYPES,
EPISTEMIC_CONSUMER_RELATION_TYPES,
} from '../../src/mcp/local/local-backend.js';
import { INCREMENTAL_SCHEMA_VERSION } from '../../src/storage/repo-manager.js';
import {
schemaFingerprintMismatch,
NODE_SCHEMA_QUERIES,
REL_SCHEMA_QUERIES,
SCHEMA_FINGERPRINT,
} from '../../src/core/lbug/schema.js';
const here = path.dirname(fileURLToPath(import.meta.url));
const repoRoot = path.resolve(here, '..', '..');
const runAnalyzeSource = readFileSync(path.join(repoRoot, 'src', 'core', 'run-analyze.ts'), 'utf8');
describe('CALL_SUMMARY relation-type exclusion (U-C1)', () => {
it('is NOT in VALID_RELATION_TYPES (never enters impact symbol-space traversal)', () => {
expect(VALID_RELATION_TYPES.has('CALL_SUMMARY')).toBe(false);
@@ -72,158 +98,50 @@ describe('CALL_SUMMARY relation-type exclusion (U-C1)', () => {
});
});
describe('CALL_SUMMARY incremental reuse gate (U-C5)', () => {
it('INCREMENTAL_SCHEMA_VERSION is bumped to 34 (Spring AOP relation pairs #2416, then receiver-chain wire format v2)', () => {
// Moves with every bump BY DESIGN — that is the point of pinning it. A
// change that alters emitted ids or edges without bumping would otherwise
// ship silently, and an existing index would keep serving the old graph
// through the reuse gate below.
expect(INCREMENTAL_SCHEMA_VERSION).toBe(34);
describe('incremental reuse gate — schema fingerprint (U-C5, #2798)', () => {
// Calls the real predicate the production gates call. Before #2798 this file
// pinned `expect(INCREMENTAL_SCHEMA_VERSION).toBe(35)`, a literal that failed
// CI on every bump by design; a digest has no literal to pin, so what is
// pinned instead is the decision the digest drives.
it.each([
{ stamped: SCHEMA_FINGERPRINT, mismatch: false, why: "this build's own DDL" },
{ stamped: 'a0b1c2d3e4f5', mismatch: true, why: 'a well-formed digest from another build' },
{ stamped: undefined, mismatch: true, why: 'an index predating the field' },
{ stamped: '', mismatch: true, why: 'an empty stamp' },
])('treats $why as mismatch=$mismatch', ({ stamped, mismatch }) => {
expect(schemaFingerprintMismatch(stamped)).toBe(mismatch);
});
it('a pre-current stamp fails the `=== INCREMENTAL_SCHEMA_VERSION` reuse gate → forces full re-analyze', () => {
// The reuse gate at run-analyze.ts:920 is exactly this strict equality on
// the persisted `existingMeta.schemaVersion` (a plain number, possibly
// absent on a legacy stamp). Replicate it as a typed predicate.
const passesReuseGate = (stampedSchemaVersion: number | undefined): boolean =>
stampedSchemaVersion === INCREMENTAL_SCHEMA_VERSION;
// A pre-v4 (v3) index has no CALL_SUMMARY edges → must NOT reuse.
expect(passesReuseGate(3)).toBe(false);
// A pre-v5 (v4) index predates the multi-verb Route identity change → its
// persisted Route nodes use the old url-only ids, so an incremental top-up
// would strand them → must NOT reuse.
expect(passesReuseGate(4)).toBe(false);
// A legacy stamp with no schemaVersion at all is likewise rejected.
expect(passesReuseGate(undefined)).toBe(false);
// A pre-v6 (v5) index predates the uniform 0-based line-storage flip → its
// COBOL/JCL/markdown/scope rows are still 1-based, so an incremental top-up
// would mix bases → must NOT reuse.
expect(passesReuseGate(5)).toBe(false);
// A pre-v7 (v6) index predates the callable-value-flow edges (#2437/#2522)
// — new edges between unchanged files would never enter the incremental
// write set → must NOT reuse.
expect(passesReuseGate(6)).toBe(false);
// A pre-v8 (v7) index predates the Java anonymous-class instance model
// (#2550) — `Worker.run`-keyed Method nodes would be stranded alongside
// the re-keyed `Worker$N.run` ones on unchanged files → must NOT reuse.
expect(passesReuseGate(7)).toBe(false);
// A pre-v9 (v8) index predates enum constant bodies + JLS 13.1
// immediate-host naming (#2555) — `E.hook`-keyed Method nodes and
// topmost-anchored `EnumWrap$1`-style ids would be stranded alongside
// the re-keyed ones on unchanged files → must NOT reuse.
expect(passesReuseGate(8)).toBe(false);
// A pre-v10 (v9) index predates the Java record container-node fix
// (#2564) — a record's methods would keep being ownerless Method nodes
// with no HAS_METHOD edge on unchanged files → must NOT reuse.
expect(passesReuseGate(9)).toBe(false);
// A pre-v11 (v10) index predates the Rust dyn-trait-object dispatch fix
// (#2604) — abstract trait methods would keep being uncaptured (no
// ownerId/CALLS resolution) on unchanged Rust trait files → must NOT reuse.
expect(passesReuseGate(10)).toBe(false);
// A pre-v12 (v11) index predates the #2514 Rust range-binding fix — the
// ambiguity latch removes spurious cross-file CALLS edges and the
// import-disambiguated resolution adds new ones on unchanged Rust files,
// neither of which reach an incremental write set → must NOT reuse.
expect(passesReuseGate(11)).toBe(false);
// A pre-v13 (v12) index predates javac-compatible Java local-type
// identities and lexical visibility scopes (#2562), so unchanged
// simple-name-keyed type/member ids must not survive.
expect(passesReuseGate(12)).toBe(false);
// A pre-v14 (v13) index predates the C#/Kotlin instance-ownership gate,
// so unchanged files may retain spurious same-file CALLS edges.
expect(passesReuseGate(13)).toBe(false);
// A pre-v15 (v14) index predates the #2687 const-arrow twin removal — an
// edgeless `Const:<file>:X` twin survives beside its `Function` node on
// every unchanged TS/JS file, and the incremental write set never touches
// those files → must NOT reuse.
expect(passesReuseGate(14)).toBe(false);
// A pre-v16 (v15) index predates #2693: calls through a closure-valued
// binding do not resolve in Kotlin/Swift/Dart, and the incremental write
// set never revisits unchanged files, so those symbols would keep reporting
// a zero blast radius → must NOT reuse.
expect(passesReuseGate(15)).toBe(false);
// A pre-v17 (v16) index predates #2701: `this` inside an ordinary JS/TS
// `function` still resolves to the enclosing class, so every unchanged
// TS/JS file keeps its fabricated `this` edges → must NOT reuse.
expect(passesReuseGate(16)).toBe(false);
// A pre-v18 (v17) index predates #2699: a function-local callable still
// shares a node id with a same-named file-level one, and the incremental
// write set would mix old and new ids → must NOT reuse.
expect(passesReuseGate(17)).toBe(false);
// A pre-v19 (v18) index holds the WRONG Java anonymous-class ids — v18 bounded the
// enclosing-callable walk on class DECLARATIONS only, so `Worker$1.run` was re-keyed
// as `Worker.makeHandler.run@7:12`. Reusing it would keep those on unchanged files.
expect(passesReuseGate(18)).toBe(false);
// A pre-v20 (v19) index holds the false CALLS/ACCESSES a NAMED explicit receiver
// used to mint through the lexical chain (`options.baseUrl` → a function-local
// `const baseUrl`) — 709 of them on a 762-file corpus. Reusing it would keep
// every one on unchanged files.
expect(passesReuseGate(19)).toBe(false);
// A pre-v21 (v20) index predates closure bindings becoming call SOURCES in
// PHP/Rust/Kotlin/Ruby/Dart, the Rust graph node for `let f = || …`, the Dart
// closure scope + enclosing-callable identity, and position-qualified
// function-local VALUES. All of those change emitted ids and edges on files
// that did not themselves change, so reusing a v20 index keeps serving the
// old attribution — including the Dart case where two same-named closures
// collapsed onto one node and asserted a CALLS edge present nowhere in the
// source.
expect(passesReuseGate(20)).toBe(false);
// A pre-v22 (v21) index predates CommonJS export indexing (#2723): every
// unchanged CJS file would keep its pre-fix graph → must NOT reuse.
expect(passesReuseGate(21)).toBe(false);
// A pre-v23 (v22) index predates Rust module-qualified call resolution
// (#2730): every unchanged Rust file would keep the same-name self-loop and
// keep reporting the real callee as unreached → must NOT reuse.
expect(passesReuseGate(22)).toBe(false);
// A pre-v24 (v23) index predates the #2708 inline-constructor receivers.
expect(passesReuseGate(23)).toBe(false);
// A pre-v25 (v24) index predates `unresolvedReceiverMembers` (#2744). An
// absent summary is indistinguishable from "nothing was dropped", so a
// top-up would keep reporting `epistemic: 'exact'` for exactly the symbols
// whose callers were dropped → must NOT reuse.
expect(passesReuseGate(24)).toBe(false);
// A pre-v26 (v25) index typed receivers from source TEXT, so `svc?.m().n()`,
// `svc!.m().n()` and `svc.m<T>().n()` emitted no CALLS edge — and two of the
// three recorded no drop either, so the count still claimed `exact`. A
// changed-files top-up keeps both the missing edge and the false confidence
// for every unchanged file → must NOT reuse.
expect(passesReuseGate(25)).toBe(false);
// A pre-v27 (v26) index was stamped by an intermediate build of this same
// series: structural typing was TypeScript-only at that point, and the fold
// still typed a bare identifier that merely shadowed a class name as that
// class — so such an index carries both pre-rollout edges for 13 languages
// and fabricated ones. The gate is a strict `===`, so it must NOT reuse.
expect(passesReuseGate(26)).toBe(false);
// A pre-v31 index predates the receiver-chain wire format v2, so its
// persisted chains carry the v1 prefix a v2 decoder refuses by design.
// Original note: a pre-v28 (v27) index was stamped mid-series: TypeScript-only structural
// typing, and the fold still typed a local that merely shadowed a class name
// as that class — so it carries pre-rollout edges AND fabricated ones.
expect(passesReuseGate(27)).toBe(false);
// A pre-v29 (v28) index lacks Class→CodeElement relation schema support,
// so Spring @Bean injection edges (#2413) would be dropped during
// persistence → must NOT reuse.
expect(passesReuseGate(28)).toBe(false);
// A pre-v30 (v29) index keeps wrapper-line startLines for multi-line closure
// bindings (#2735), so the graph-to-scope join still drops the CALLS edge.
expect(passesReuseGate(29)).toBe(false);
// A pre-v31 (v30) index treats `from pkg import models` as a named package
// import, so unchanged files retain the old missing qualified CALLS edges.
expect(passesReuseGate(30)).toBe(false);
// A pre-v32 (v31) index predates the Rust impl/trait, JS/TS object-literal
// and Swift member-containment relation pairs (#2769), so an incremental
// top-up emitting one of those edges would fail the bulk COPY (or silently
// drop it on the streamed path) → must NOT reuse.
expect(passesReuseGate(31)).toBe(false);
// A pre-v33 (v32) index predates the Spring AOP Interface→CodeElement
// relation pair (#2416), so it cannot persist all evidence edges.
expect(passesReuseGate(32)).toBe(false);
// A pre-v34 (v33) index carries `receiverChain` strings in wire format v1,
// which the v2 decoder refuses by design (#2766) — an incremental top-up
// would silently fall back to the text cascade for every chain-carrying
// site → must NOT reuse.
expect(passesReuseGate(33)).toBe(false);
// The current stamp passes the gate (incremental top-up eligible).
expect(passesReuseGate(34)).toBe(true);
it('is a digest of the node+relation DDL and of nothing else', () => {
// Pins the INPUT SET, not the algorithm: the fingerprint is a pure function
// of the DDL, which is why it cannot fire on a semantic change (see the
// runner-identity describe below) and why EMBEDDING_SCHEMA — whose FLOAT[N]
// width comes from GITNEXUS_EMBEDDING_DIMS at module load — must stay out,
// or the same build under different env would disagree with itself.
// schema-fingerprint.test.ts owns the digest's other properties.
expect(SCHEMA_FINGERPRINT).toBe(
createHash('sha256')
.update([...NODE_SCHEMA_QUERIES, ...REL_SCHEMA_QUERIES].join('\n'))
.digest('hex')
.slice(0, 12),
);
});
});
describe('semantic (non-DDL) analyzer changes ride the runner-identity receipt (#2798)', () => {
it('run-analyze.ts still forces a full rebuild when the stamped runner identity differs', () => {
// The invariant the INCREMENTAL_SCHEMA_VERSION ladder used to backstop. It
// is implicit nowhere else: no other gate observes analyzer code that emits
// no DDL. Deleting this block silently re-opens same-commit top-ups across
// an analyzer that changed how the graph is shaped.
//
// Source-anchored on purpose: the wiring has no extracted predicate to call,
// so the only way to assert the gate still exists is to read run-analyze.ts.
// The predicate's OWN behaviour — a moved build digest with unmoved DDL, an
// absent/null/legacy/malformed receipt, an alternate diagnostic entrypoint —
// is asserted against the real function in analyzer-identity.test.ts.
expect(runAnalyzeSource).toMatch(
/!analyzerRunnerIdentitiesEqual\(\s*existingMeta\.runnerIdentity,\s*runnerIdentity,?\s*\)[\s\S]{0,900}?options = \{ \.\.\.options, force: true \};/,
);
});
});
+555 -1
View File
@@ -7,7 +7,7 @@
* These are pure unit tests that mock the LadybugDB layer to test
* the dispatch and error handling logic in isolation.
*/
import { describe, it, expect, vi, beforeEach, afterAll } from 'vitest';
import { describe, it, expect, vi, beforeEach, afterEach, afterAll } from 'vitest';
import type { StalenessInfo } from '../../src/core/git-staleness.js';
import { mkdirSync, mkdtempSync, rmSync, writeFileSync } from 'fs';
import fsPromises from 'fs/promises';
@@ -938,6 +938,560 @@ describe('LocalBackend.callTool', () => {
expect(result.candidates[0].score).toBeGreaterThanOrEqual(result.candidates[1].score);
});
// #2787 — LadybugDB returns an ARBITRARY subset when a LIMIT has no ORDER BY,
// and a different subset from one process to the next. With 92 nodes named
// `constructor` in this repo's own index, the resolver's LIMIT 20 window
// moved every run, so `impact`/`context` resolved a different symbol each
// time and the HIGH/CRITICAL warning the agent workflow relies on fired at
// random. These assert the emitted SQL, which is the only shape that fails
// deterministically — a run-N-times-and-compare test would pass by luck at
// the ~5-8% flip rate actually measured.
/** N same-named Function rows, ids ascending so the window order is obvious. */
const collideRows = (count: number) =>
Array.from({ length: count }, (_, i) => ({
id: `Function:src/f${String(i).padStart(2, '0')}.ts:collide`,
name: 'collide',
type: 'Function',
filePath: `src/f${String(i).padStart(2, '0')}.ts`,
startLine: 1,
endLine: 3,
}));
// The window must come back FULL (20 rows). The resolver only issues the COUNT
// when the window saturated its own LIMIT — a short page already proves the
// exact total, so counting again would be a second unlabeled full scan for a
// number we hold. A 0-row fixture would therefore assert the wrong shape.
const resolverQueriesFor = async (params: Record<string, unknown>): Promise<string[]> => {
(executeParameterized as any).mockClear();
(executeParameterized as any).mockResolvedValue(collideRows(20));
await backend.callTool('context', params);
return (executeParameterized as any).mock.calls
.map((c: unknown[]) => String(c[1]))
.filter((q: string) => q.includes('$symName'));
};
// Two queries share `$symName`: the ordered 20-row window, and the COUNT that
// reports the TRUE match total (the window length would report the cap — 20
// when 92 match). Both halves are pinned; the COUNT must carry no LIMIT of
// its own or it would just re-report the cap.
const expectWindowAndCount = (queries: string[]): void => {
expect(queries).toHaveLength(2);
expect(queries.filter((q) => /ORDER BY n\.id LIMIT 20/.test(q))).toHaveLength(1);
expect(
queries.filter((q) => /RETURN COUNT\(\*\) AS total/.test(q) && !/LIMIT/.test(q)),
).toHaveLength(1);
};
// One case per WHERE-clause shape the resolver builds — all three must carry
// the ordered window and its COUNT.
it.each([
['bare name', { name: 'main' }],
['file_path hint', { name: 'main', file_path: 'src/a.ts' }],
['qualified name', { name: 'src/a.ts:main' }],
] as Array<[string, Record<string, unknown>]>)(
'resolver window is pinned by ORDER BY n.id, with a COUNT for the true total — %s (#2787)',
async (_label, params) => {
expectWindowAndCount(await resolverQueriesFor(params));
},
);
it('ambiguous context reports the COUNT as the match total, not the capped window (#2787)', async () => {
const windowRows = Array.from({ length: 20 }, (_, i) => ({
id: `Function:src/f${String(i).padStart(2, '0')}.ts:collide`,
name: 'collide',
type: 'Function',
filePath: `src/f${String(i).padStart(2, '0')}.ts`,
startLine: 1,
endLine: 3,
}));
(executeParameterized as any).mockClear();
(executeParameterized as any).mockImplementation(async (_repo: string, query: string) =>
/RETURN COUNT\(\*\) AS total/.test(query) ? [{ total: 92 }] : windowRows,
);
const result = await backend.callTool('context', { name: 'collide' });
(executeParameterized as any).mockReset();
(executeParameterized as any).mockResolvedValue([]);
expect(result).toMatchObject({
status: 'ambiguous',
totalCandidates: 92,
candidatesTruncated: true,
});
expect(result.candidates).toHaveLength(20);
expect(result.message).toContain('Found 92 symbols');
expect(result.message).toContain('showing 20');
});
it('no multi-row LIMIT on the context path is left unordered (#2787)', async () => {
(executeParameterized as any).mockClear();
(executeParameterized as any).mockResolvedValue([
{
id: 'Class:src/a.ts:Widget',
name: 'Widget',
type: 'Class',
filePath: 'src/a.ts',
startLine: 1,
endLine: 9,
},
]);
await backend.callTool('context', { name: 'Widget' });
const captured: string[] = (executeParameterized as any).mock.calls.map((c: unknown[]) =>
String(c[1]),
);
// Guard against a vacuous pass: a resolver that bailed early would capture
// one query and satisfy the invariant below trivially.
expect(captured.length).toBeGreaterThan(1);
// `LIMIT 1` anchored on a unique id is a singleton lookup, not a window —
// which rows come back cannot vary. Every other cap must be ordered.
const unordered = captured
.filter((q) => /\bLIMIT\s+\d+/.test(q))
.filter((q) => !/LIMIT\s+1\b/.test(q))
.filter((q) => !/ORDER BY/.test(q));
expect(unordered).toEqual([]);
});
// ── #2787 review fixes ────────────────────────────────────────────────
// Each of these pins a behaviour the ORDER BY work itself introduced or left
// exposed. `executeParameterized` is routed on QUERY TEXT (mock-internal
// `if`s, the established pattern in this file) so a single leg can be made to
// fail or return a shaped page without touching the others.
describe('#2787 review fixes', () => {
/** Restore the file-wide default (`executeParameterized` → `[]`). */
const restoreQueryMock = (): void => {
(executeParameterized as any).mockReset();
(executeParameterized as any).mockResolvedValue([]);
};
// Every test below installs its own query routing. `afterEach` puts the
// file-wide default back — on the throwing path as well as the clean one,
// exactly like the per-test `finally` blocks it replaces — so a failure
// here still cannot leak a mock into the rest of the suite.
afterEach(restoreQueryMock);
/** The two `$symName` legs the resolver emits, with their bound params. */
const resolverCalls = (): Array<{ query: string; params: Record<string, unknown> }> =>
(executeParameterized as any).mock.calls
.filter((c: unknown[]) => String(c[1]).includes('$symName'))
.map((c: unknown[]) => ({
query: String(c[1]),
params: (c[2] ?? {}) as Record<string, unknown>,
}));
it('keys every multi-relType ref window uid-major, not category-major (#2787 review F1)', async () => {
// Supplement to the real-DB spread test in
// test/integration/local-backend-calltool.test.ts: that one proves the
// BEHAVIOUR on the primary incoming window; this one proves the same key
// reaches the four windows a single fixture cannot exercise at once (the
// Class-only Constructor / File / typed-Property expansions, plus outgoing).
// See the incoming-ref window in `_contextImpl` (#2787 F1) for why a
// category-major key starves whole buckets silently.
(executeParameterized as any).mockImplementation(async (_repo: string, query: string) =>
query.includes('$uid')
? [
{
id: 'Class:src/a.ts:Widget',
name: 'Widget',
type: 'Class',
filePath: 'src/a.ts',
startLine: 1,
endLine: 9,
},
]
: [],
);
// A Class target opens the #480 expansion windows as well as the two
// primary ones.
await backend.callTool('context', { uid: 'Class:src/a.ts:Widget' });
const windows = (executeParameterized as any).mock.calls
.map((c: unknown[]) => String(c[1]))
// The single-relType ADVISED_BY windows are keyed `ORDER BY uid` alone
// (one category, nothing to starve) and are excluded by `r.type IN [`.
.filter((q: string) => /RETURN r\.type AS relType/.test(q) && /r\.type IN \[/.test(q));
expect(windows).toHaveLength(5);
expect(windows.filter((q: string) => /ORDER BY uid, relType/.test(q))).toHaveLength(5);
});
it('marks the match total as a LOWER BOUND when only the COUNT leg fails (#2787 review F3)', async () => {
// The COUNT rides alongside the window so the response can report the TRUE
// match count instead of the cap. When that leg fails the code falls back to
// the window length — which, un-marked, is byte-identical to a genuine
// N-match result and silently reinstates the pre-PR undercount. The failure
// must therefore be BOTH marked on the payload and logged.
(executeParameterized as any).mockImplementation(async (_repo: string, query: string) => {
// Both legs carry `$symName`, so the COUNT must be matched first.
if (/RETURN COUNT\(\*\) AS total/.test(query)) throw new Error('count leg exploded');
if (query.includes('$symName')) return collideRows(20);
return [];
});
const cap = _captureLogger();
try {
const result = await backend.callTool('context', { name: 'collide' });
expect(result).toMatchObject({
status: 'ambiguous',
totalCandidates: 20,
totalIsLowerBound: true,
});
expect(result.candidates).toHaveLength(20);
// The prose is what an agent actually reads, so the hedge has to be there
// too — "Found 20 symbols" asserts an exactness the resolver no longer has.
expect(result.message).toContain("Found at least 20 symbols matching 'collide'");
// `candidatesTruncated` is driven by `total > candidates.length`, and with
// the COUNT dead the total floors at the window length — so the flag is
// absent here and CANNOT stand in for the lower-bound marker. That is the
// whole point: without `totalIsLowerBound` this response is byte-identical
// to a genuine, exactly-20-match result.
expect(result).not.toHaveProperty('candidatesTruncated');
// …and the swallowed failure is observable in telemetry, not silent.
expect(
cap.records().filter((r) => String(r.context) === 'resolve:candidate-count'),
).toHaveLength(1);
} finally {
cap.restore();
}
});
it('a successful COUNT leg reports an EXACT total with no lower-bound marker (#2787 review F3)', async () => {
// Negative control for the test above: the marker must be absent on the
// healthy path, or it degrades into noise that consumers learn to ignore.
(executeParameterized as any).mockImplementation(async (_repo: string, query: string) => {
if (/RETURN COUNT\(\*\) AS total/.test(query)) return [{ total: 92 }];
if (query.includes('$symName')) return collideRows(20);
return [];
});
const result = await backend.callTool('context', { name: 'collide' });
expect(result).toMatchObject({ status: 'ambiguous', totalCandidates: 92 });
expect(result).not.toHaveProperty('totalIsLowerBound');
expect(result.message).toContain("Found 92 symbols matching 'collide'");
expect(result.message).not.toContain('at least');
});
it('a kind hint filters in the WHERE clause on BOTH legs, it does not merely score (#2787 review F5)', async () => {
// See `resolveSymbolCandidates` (#2787 F5) for why the id order is
// label-major and why the hint therefore has to filter, not merely score.
(executeParameterized as any).mockImplementation(async (_repo: string, query: string) =>
query.includes('$symName') ? collideRows(20) : [],
);
await backend.callTool('context', { name: 'collide', kind: 'Function' });
const calls = resolverCalls();
// The filtered window returned rows, so the unfiltered fallback stays out:
// exactly the window + its COUNT.
expect(calls).toHaveLength(2);
expect(calls.filter((c) => /AND n\.id STARTS WITH \$kindPrefix/.test(c.query))).toHaveLength(
2,
);
// The COUNT must carry the SAME filter, or `totalCandidates` reports the
// unfiltered population next to a filtered page.
expect(calls.map((c) => c.params.kindPrefix)).toEqual(['Function:', 'Function:']);
});
it('a kind hint on a qualified name keeps the id/name OR-clause parenthesised (#2787 review F5)', async () => {
// `AND` binds tighter than `OR`: an unparenthesised
// `n.id = $symName OR n.name = $symName AND n.id STARTS WITH $kindPrefix`
// applies the kind filter to the name branch ONLY, so a qualified-id lookup
// silently ignores the hint.
(executeParameterized as any).mockImplementation(async (_repo: string, query: string) =>
query.includes('$symName') ? collideRows(20) : [],
);
await backend.callTool('context', { name: 'src/a.ts:collide', kind: 'Function' });
const parenthesised =
/WHERE \(n\.id = \$symName OR n\.name = \$symName\) AND n\.id STARTS WITH \$kindPrefix/;
const calls = resolverCalls();
expect(calls).toHaveLength(2);
expect(calls.filter((c) => parenthesised.test(c.query))).toHaveLength(2);
});
it('retries UNFILTERED when the kind hint matches no label prefix (#2787 review F5)', async () => {
// `kind` is a free-form string on the tool schema. A miscased or
// repo-absent kind must not turn a real name into `not_found` — the
// resolver falls back to the unfiltered window and treats the hint as a
// ranking term again. Modelled by failing the `$kindPrefix` leg to zero rows.
(executeParameterized as any).mockImplementation(async (_repo: string, query: string) => {
if (query.includes('$kindPrefix')) return [];
if (query.includes('$symName')) return collideRows(20);
return [];
});
const result = await backend.callTool('context', { name: 'collide', kind: 'function' });
// Filtered window (empty), THEN unfiltered window + its COUNT. The filtered
// leg issues NO count: a window shorter than its own LIMIT already proves
// the total, so counting again would be a second unlabeled full scan for a
// number we hold — here, zero.
expect(resolverCalls().map((c) => c.query.includes('$kindPrefix'))).toEqual([
true,
false,
false,
]);
// Not `{ error: "Symbol 'collide' not found" }`.
expect(result).toMatchObject({ status: 'ambiguous' });
expect(result.candidates).toHaveLength(20);
});
// #2787 review F6 — two distinct entry points can collide on (total_hits,
// filePath, name); equal `total_hits` is the norm. The old three-key
// comparator therefore tied, and a tie in `Array.prototype.sort` (stable in
// V8) falls through to `Map` insertion order — i.e. raw DB row order, the
// exact nondeterminism this issue is about. The entry-point id (the map key)
// is the unique key that closes the order.
const COLLIDING_ENTRY_POINT_ROWS = [
{
pId: 'proc:zeta-flow',
name: 'Zeta Flow',
processType: 'intra_community',
entryPointId: 'ep:zeta',
hits: 3,
minStep: 5,
stepCount: 4,
epName: 'step',
epType: 'Function',
epFilePath: 'src/hooks/useSigma.ts',
},
{
pId: 'proc:alpha-flow',
name: 'Alpha Flow',
processType: 'intra_community',
entryPointId: 'ep:alpha',
hits: 3,
minStep: 2,
stepCount: 4,
epName: 'step',
epType: 'Method',
epFilePath: 'src/hooks/useSigma.ts',
},
];
const impactWithCollidingEntryPoints = async (): Promise<any> => {
(executeParameterized as any).mockImplementation(async (_repo: string, query: string) => {
if (query.includes('$symName')) {
return [
{
id: 'func:main',
name: 'main',
type: 'Function',
filePath: 'src/index.ts',
startLine: 1,
endLine: 5,
},
];
}
if (query.includes('$frontierIds')) {
return [
{
sourceId: 'func:main',
id: 'func:caller',
name: 'caller',
type: 'Function',
filePath: 'src/caller.ts',
relType: 'CALLS',
confidence: 0.9,
},
];
}
// The chunked process/entry-point aggregation.
if (query.includes('p.entryPointId')) return COLLIDING_ENTRY_POINT_ROWS;
return [];
});
return backend.callTool('impact', { target: 'main', direction: 'upstream' });
};
it('orders affected_processes by entry-point id when name/filePath/hits all tie (#2787 review F6)', async () => {
const result = await impactWithCollidingEntryPoints();
// Both entry points are named `step`, live in the same file and carry the
// same hit count, so ONLY the id tiebreak can decide: `ep:alpha` before
// `ep:zeta`, regardless of the row order the DB handed back (here, zeta
// first). The id itself stays out of the payload — `earliest_broken_step`
// and `type` are what make the order observable.
expect(result.affected_processes).toEqual([
{
name: 'step',
type: 'Method',
filePath: 'src/hooks/useSigma.ts',
affected_process_count: 1,
total_hits: 3,
earliest_broken_step: 2,
},
{
name: 'step',
type: 'Function',
filePath: 'src/hooks/useSigma.ts',
affected_process_count: 1,
total_hits: 3,
earliest_broken_step: 5,
},
]);
});
it('pins the process-chunk row order with ORDER BY pId (#2787 review F6)', async () => {
await impactWithCollidingEntryPoints();
// The comparator tiebreak above only fixes the FINAL order. Aggregation
// (`affected_process_count`, `total_hits`, `Math.min` on the step) reads
// rows as they arrive and the map is keyed on first sight, so the chunk
// query needs its own total order too.
const chunkQueries = (executeParameterized as any).mock.calls
.map((c: unknown[]) => String(c[1]))
.filter((q: string) => q.includes('p.entryPointId'));
expect(chunkQueries).toHaveLength(1);
expect(chunkQueries[0]).toMatch(/ORDER BY pId/);
});
});
// ── #2787 — the impact BFS frontier query was the only `ORDER BY` in the
// backend with no `LIMIT` to escape into (every other one pairs its key with a
// small limit, so the engine answers from a bounded top-k heap). It now
// returns rows unordered and the traversal re-establishes, in JS, exactly the
// two properties that key provided. Nothing downstream re-establishes them for
// it: `byDepth` slices `impacted` without re-sorting, and the process/module
// enrichment reads a positional prefix of the same array.
describe('#2787 impact BFS frontier ordering', () => {
const TARGET_ROW = {
id: 'Function:src/index.ts:main',
name: 'main',
type: 'Function',
filePath: 'src/index.ts',
startLine: 1,
endLine: 5,
};
const ROOT = TARGET_ROW.id;
interface FrontierRow {
sourceId: string;
id: string;
name: string;
type: string;
filePath: string;
relType: string;
confidence: number;
}
/** One frontier-query row. `id` is `Label:filePath:name`, as in a real index. */
const edge = (
sourceId: string,
id: string,
relType: string,
confidence: number,
): FrontierRow => {
const [, filePath, name] = id.split(':');
return { sourceId, id, name, type: 'Function', filePath, relType, confidence };
};
/**
* Drive `_runImpactBFS` with a scripted frontier: `levels[d - 1]` is what the
* depth-`d` query returns, in exactly that row order. Every other query
* (process/module enrichment, epistemic probe) returns nothing, so `byDepth`
* is precisely what the traversal produced.
*/
const impactOverFrontier = async (levels: FrontierRow[][]): Promise<any> => {
let level = 0;
(executeParameterized as any).mockImplementation(async (_repo: string, query: string) => {
if (query.includes('$symName')) return [TARGET_ROW];
if (query.includes('$frontierIds')) return levels[level++] ?? [];
return [];
});
return backend.callTool('impact', { target: 'main', direction: 'upstream' });
};
const stamped = (items: any[]): unknown[] =>
items.map((it) => ({ id: it.id, relationType: it.relationType, confidence: it.confidence }));
it('leaves the frontier query unordered — the sort has no top-k escape', async () => {
await impactOverFrontier([]);
const frontierQueries = (executeParameterized as any).mock.calls
.map((c: unknown[]) => String(c[1]))
.filter((q: string) => q.includes('$frontierIds'));
expect(frontierQueries.length).toBeGreaterThan(0);
expect(frontierQueries.filter((q: string) => /ORDER BY/.test(q))).toEqual([]);
// …and equally: no LIMIT was added in its place. The traversal must see
// every neighbour edge; only the ORDERING moved.
expect(frontierQueries.filter((q: string) => /LIMIT/.test(q))).toEqual([]);
});
// (a) Diamond: `delta` and `epsilon` are each reached at depth 2 from BOTH
// depth-1 frontier nodes. The DB key made the stamped relationType/confidence
// the argmax under `relType ASC, confidence DESC, sourceId ASC`; the JS
// argmax must pick the same edge, and must pick it from either permutation.
const DIAMOND_LEVEL_1 = [
edge(ROOT, 'Function:src/alpha.ts:alpha', 'CALLS', 0.9),
edge(ROOT, 'Function:src/beta.ts:beta', 'CALLS', 0.9),
];
// `delta` discriminates relType (CALLS < IMPORTS) AGAINST confidence — the
// weaker-confidence CALLS edge wins, so a "highest confidence" shortcut fails
// here. `epsilon` ties on relType and is decided by `confidence DESC`, across
// the 0.8 `fuzzy` boundary the tool description publishes.
const DIAMOND_LEVEL_2 = [
edge('Function:src/alpha.ts:alpha', 'Function:src/delta.ts:delta', 'IMPORTS', 0.6),
edge('Function:src/beta.ts:beta', 'Function:src/delta.ts:delta', 'CALLS', 0.55),
edge('Function:src/alpha.ts:alpha', 'Function:src/epsilon.ts:epsilon', 'CALLS', 0.7),
edge('Function:src/beta.ts:beta', 'Function:src/epsilon.ts:epsilon', 'CALLS', 0.85),
];
const DIAMOND_ARGMAX = [
{ id: 'Function:src/delta.ts:delta', relationType: 'CALLS', confidence: 0.55 },
{ id: 'Function:src/epsilon.ts:epsilon', relationType: 'CALLS', confidence: 0.85 },
];
it('stamps the argmax edge on a diamond-reached node', async () => {
const result = await impactOverFrontier([DIAMOND_LEVEL_1, DIAMOND_LEVEL_2]);
expect(stamped(result.byDepth[2])).toEqual(DIAMOND_ARGMAX);
});
it('stamps the same argmax edge when the rows arrive reversed', async () => {
// The second of two fixed permutations — not a randomised or repeat-N
// probe. Without the JS argmax the surviving edge is whichever row landed
// first, so this permutation would stamp IMPORTS/0.6 on `delta` and 0.7 on
// `epsilon` instead.
const result = await impactOverFrontier([
[...DIAMOND_LEVEL_1].reverse(),
[...DIAMOND_LEVEL_2].reverse(),
]);
expect(stamped(result.byDepth[2])).toEqual(DIAMOND_ARGMAX);
});
// (b) `impacted` — and therefore `byDepth`, which slices it — is ordered by
// node id ascending in UTF-16 code units, whatever order the rows arrive in.
// The ids below also separate code-unit order from `localeCompare`: 'Z'
// (0x5A) precedes 'a' (0x61) in code units, while a locale collator sorts
// `apple` and `mango` ahead of `Zebra`.
const SCRAMBLED_LEVEL_1 = [
edge(ROOT, 'Function:src/mango.ts:mango', 'CALLS', 0.9),
edge(ROOT, 'Function:src/Zebra.ts:Zebra', 'CALLS', 0.9),
edge(ROOT, 'Function:src/apple.ts:apple', 'CALLS', 0.9),
];
const SCRAMBLED_LEVEL_2 = [
edge('Function:src/mango.ts:mango', 'Function:src/quince.ts:quince', 'CALLS', 0.9),
edge('Function:src/apple.ts:apple', 'Function:src/Fig.ts:Fig', 'CALLS', 0.9),
];
const ID_ASCENDING_1 = [
'Function:src/Zebra.ts:Zebra',
'Function:src/apple.ts:apple',
'Function:src/mango.ts:mango',
];
const ID_ASCENDING_2 = ['Function:src/Fig.ts:Fig', 'Function:src/quince.ts:quince'];
it('orders byDepth by node id ascending, whatever order the rows arrive in', async () => {
const result = await impactOverFrontier([SCRAMBLED_LEVEL_1, SCRAMBLED_LEVEL_2]);
expect(result.byDepth[1].map((it: any) => it.id)).toEqual(ID_ASCENDING_1);
expect(result.byDepth[2].map((it: any) => it.id)).toEqual(ID_ASCENDING_2);
expect(result.byDepthCounts).toEqual({ 1: 3, 2: 2 });
});
it('orders byDepth identically when the rows arrive reversed', async () => {
const result = await impactOverFrontier([
[...SCRAMBLED_LEVEL_1].reverse(),
[...SCRAMBLED_LEVEL_2].reverse(),
]);
expect(result.byDepth[1].map((it: any) => it.id)).toEqual(ID_ASCENDING_1);
expect(result.byDepth[2].map((it: any) => it.id)).toEqual(ID_ASCENDING_2);
});
});
it('context tool ranks file_path match higher than non-match (#470)', async () => {
(executeParameterized as any).mockResolvedValue([
{
@@ -0,0 +1,133 @@
/**
* Pins the weight-aware split behind the cross-platform matrix (#2449).
*
* The regression this guards is specific and was expensive: three CHEAP files
* were registered in `SPAWN_CLI`, vitest re-partitioned the list by file COUNT,
* and the reshuffle clustered `cli-e2e` (361 s on Windows) with `cli-limit-e2e`
* (75 s) and `analyze-heap-oom-e2e` (23 s) on one shard, which then blew the
* 20-minute watchdog. The added files cost nothing; the COUNT-split did it.
*
* So the load-bearing case here is not "the split is even" — it is
* "adding a cheap file does not move a heavy one". A partition that merely
* balanced totals could still reshuffle everything on every insertion and would
* reproduce the outage exactly.
*/
import { describe, it, expect } from 'vitest';
import {
shardFiles,
shardWeight,
weightOf,
WINDOWS_WEIGHTS_SEC,
} from '../../scripts/cross-platform-shard.js';
import { ALL_CROSS_PLATFORM } from '../../scripts/cross-platform-tests.js';
const SHARD_TOTAL = 3;
/** Every shard of a split, as file lists. */
const allShards = (files: readonly string[], total: number): readonly (readonly string[])[] =>
Array.from({ length: total }, (_unused, i) => shardFiles(files, i + 1, total));
describe('cross-platform shard partition', () => {
it('covers every file exactly once, with no overlap between shards', () => {
const shards = allShards(ALL_CROSS_PLATFORM, SHARD_TOTAL);
const seen = shards.flatMap((s) => [...s]);
expect(seen.slice().sort()).toEqual([...ALL_CROSS_PLATFORM].sort());
expect(new Set(seen).size).toBe(ALL_CROSS_PLATFORM.length);
});
it('keeps each shard within a shard of the ideal weight', () => {
const shards = allShards(ALL_CROSS_PLATFORM, SHARD_TOTAL);
const weights = shards.map(shardWeight);
const ideal = shardWeight(ALL_CROSS_PLATFORM) / SHARD_TOTAL;
// LPT's guarantee is 4/3 of optimal, and optimal is at least the ideal
// average. A hard 1.34x ceiling on the busiest shard is what keeps the
// matrix inside its watchdog no matter how the list is edited.
expect(Math.max(...weights)).toBeLessThanOrEqual(ideal * 1.34);
});
it('never puts the two heaviest suites on the same shard', () => {
// The exact shape of the outage: cli-e2e and worker-pool are 361 s and
// 222 s, so together they are most of a shard's budget before anything else
// is scheduled.
const shards = allShards(ALL_CROSS_PLATFORM, SHARD_TOTAL);
const withBoth = shards.filter(
(s) =>
s.includes('test/integration/cli-e2e.test.ts') &&
s.includes('test/integration/worker-pool.test.ts'),
);
expect(withBoth).toEqual([]);
});
it('does not move a heavy file when a cheap file is added — the #2449 regression', () => {
const heavy = Object.keys(WINDOWS_WEIGHTS_SEC);
const placementOf = (files: readonly string[]): ReadonlyMap<string, number> => {
const shards = allShards(files, SHARD_TOTAL);
return new Map(
heavy
.map((f) => [f, shards.findIndex((s) => s.includes(f))] as const)
.filter(([, i]) => i >= 0),
);
};
const before = placementOf(ALL_CROSS_PLATFORM);
// The inserted names sort EARLY, and there is a case that is NOT a multiple
// of the shard count. Both details are load-bearing, and getting them wrong
// made earlier versions of this test vacuous:
// - names that sort last cannot disturb anything under any scheme;
// - adding exactly `total` files leaves an equal-weight round-robin in the
// same rotation, so a count-split would pass too.
// Under the real weighted split, heavy files are scheduled before every
// light one, so no number of cheap insertions can move them.
const afterOne = placementOf([...ALL_CROSS_PLATFORM, 'test/aaa-new-cheap-a.test.ts']);
const afterTwo = placementOf([
...ALL_CROSS_PLATFORM,
'test/aaa-new-cheap-a.test.ts',
'test/aaa-new-cheap-b.test.ts',
]);
expect(Object.fromEntries(afterOne)).toMatchObject(Object.fromEntries(before));
expect(Object.fromEntries(afterTwo)).toMatchObject(Object.fromEntries(before));
});
it('is deterministic, so every runner computes the same split independently', () => {
// Each matrix job resolves its own slice on its own machine with no shared
// state, so an unstable sort would silently drop or duplicate files.
const once = allShards(ALL_CROSS_PLATFORM, SHARD_TOTAL).map((s) => [...s]);
const twice = allShards([...ALL_CROSS_PLATFORM].reverse(), SHARD_TOTAL).map((s) =>
[...s].sort(),
);
expect(twice).toEqual(once.map((s) => [...s].sort()));
});
it('returns every file for a single-shard run, and rejects an out-of-range shard', () => {
expect(shardFiles(ALL_CROSS_PLATFORM, 1, 1)).toEqual([...ALL_CROSS_PLATFORM]);
expect(() => shardFiles(ALL_CROSS_PLATFORM, 0, 3)).toThrow(/shard index/);
expect(() => shardFiles(ALL_CROSS_PLATFORM, 4, 3)).toThrow(/shard index/);
expect(() => shardFiles(ALL_CROSS_PLATFORM, 1, 0)).toThrow(/shard total/);
});
it('charges every file the per-file floor, so light files are never free', () => {
// Without this, the balancer isolates the monsters and then piles all the
// light files onto the remaining shards — a count imbalance that costs just
// as much wall clock as the runtime one it just fixed.
expect(weightOf('test/unit/zzz-does-not-exist.test.ts')).toBeGreaterThan(0);
expect(weightOf('test/integration/cli-e2e.test.ts')).toBeGreaterThan(
WINDOWS_WEIGHTS_SEC['test/integration/cli-e2e.test.ts'] ?? 0,
);
});
it('weights only files that are actually registered', () => {
// A weight entry for a file no longer in the list is dead config that the
// balancer silently ignores; catching it here keeps the table honest.
const registered = new Set(ALL_CROSS_PLATFORM);
const stale = Object.keys(WINDOWS_WEIGHTS_SEC).filter((f) => !registered.has(f));
expect(stale).toEqual([]);
});
});
@@ -0,0 +1,152 @@
import { describe, expect, it } from 'vitest';
import {
EMBEDDING_RESUME_MAX_ATTEMPTS,
checkpointKind,
decideEmbeddingResume,
nextAttemptCount,
} from '../../src/core/embedding-checkpoint.js';
import type { EmbeddingCheckpoint } from '../../src/core/embedding-checkpoint.js';
/**
* ── #2790 review: ONE owner of `RepoMeta.embeddingCheckpoint` ──────────
*
* The record was minted at five sites across two layers and READ by two gates
* that disagreed: only the CLI implemented `kind`, so a `'partial'` marker
* written by `gitnexus analyze` and resumed through `POST /api/embed` hit the
* permanent wedge `kind` exists to remove. These pin the shared decision both
* callers now route through.
*/
const identity = { model: 'text-embedding-3-small', dimensions: 1536, provider: 'http:abc' };
const foreign = { model: 'all-MiniLM-L6-v2', dimensions: 384, provider: 'local' };
const checkpoint = (over: Partial<EmbeddingCheckpoint> = {}): EmbeddingCheckpoint => ({
at: '2026-08-02T00:00:00.000Z',
nodesProcessed: 10,
totalNodes: 12,
chunksProcessed: 40,
model: identity.model,
dimensions: identity.dimensions,
provider: identity.provider,
pendingNodeIds: ['node-a', 'node-b'],
...over,
});
describe('checkpointKind', () => {
it('defaults an absent kind to interrupted, so pre-#2790 markers keep the strict path', () => {
expect(checkpointKind(checkpoint({ kind: undefined }))).toBe('interrupted');
});
});
describe('decideEmbeddingResume — identity gate', () => {
it('resumes a matching identity', () => {
expect(decideEmbeddingResume(checkpoint({ kind: 'interrupted' }), identity)).toMatchObject({
action: 'resume',
pendingNodeIds: new Set(['node-a', 'node-b']),
});
});
it('fails closed on a foreign identity for an interrupted marker', () => {
expect(decideEmbeddingResume(checkpoint({ kind: 'interrupted' }), foreign)).toMatchObject({
action: 'abort',
error: expect.stringMatching(/provider configuration differs/i),
});
});
/**
* REGRESSION (caught by an apply agent, and by `run-analyze.test.ts`): the
* first cut of `decideEmbeddingResume` short-circuited on `pending === 0`
* BEFORE the identity gate, assuming an empty set meant the
* 'unverified-count' marker. It does not — `onCheckpoint` mints an
* 'interrupted' marker with `pendingNodeIds: []` after every post-window
* save — so an interrupted marker under a foreign provider was silently
* cleared instead of failing closed.
*/
it('fails closed on a foreign identity even when the interrupted marker names no pending nodes', () => {
expect(
decideEmbeddingResume(checkpoint({ kind: 'interrupted', pendingNodeIds: [] }), foreign),
).toMatchObject({ action: 'abort' });
});
it('fails closed for a legacy marker that omits both kind and pendingNodeIds', () => {
expect(
decideEmbeddingResume(checkpoint({ kind: undefined, pendingNodeIds: undefined }), foreign),
).toMatchObject({ action: 'abort' });
});
it('still resumes an empty interrupted marker once the identity matches', () => {
expect(
decideEmbeddingResume(checkpoint({ kind: 'interrupted', pendingNodeIds: [] }), identity),
).toMatchObject({ action: 'resume', pendingNodeIds: new Set() });
});
it('abandons a partial marker under a foreign identity instead of wedging the repo', () => {
expect(decideEmbeddingResume(checkpoint({ kind: 'partial' }), foreign)).toMatchObject({
action: 'abandon',
log: expect.stringMatching(/hold no embedding rows/i),
});
});
});
describe('decideEmbeddingResume — unverified-count', () => {
it('abandons without consulting the identity, having nothing to corrupt', () => {
expect(
decideEmbeddingResume(
checkpoint({ kind: 'unverified-count', pendingNodeIds: [] }),
undefined,
),
).toMatchObject({ action: 'abandon', log: expect.stringMatching(/re-deriving/i) });
});
});
describe('decideEmbeddingResume — flags and the retry bound', () => {
it.each([
['dropEmbeddings', { dropEmbeddings: true }, /--drop-embeddings/],
['force', { force: true }, /--force/],
])('discards on %s', (_name, options, pattern) => {
expect(decideEmbeddingResume(checkpoint(), identity, options)).toMatchObject({
action: 'discard',
log: expect.stringMatching(pattern),
});
});
it('abandons a partial set that has spent its attempt budget', () => {
expect(
decideEmbeddingResume(
checkpoint({ kind: 'partial', attempts: EMBEDDING_RESUME_MAX_ATTEMPTS }),
identity,
),
).toMatchObject({ action: 'abandon', log: expect.stringMatching(/consecutive/i) });
});
it('still resumes one attempt below the budget', () => {
expect(
decideEmbeddingResume(
checkpoint({ kind: 'partial', attempts: EMBEDDING_RESUME_MAX_ATTEMPTS - 1 }),
identity,
),
).toMatchObject({ action: 'resume' });
});
});
describe('nextAttemptCount', () => {
it('advances only when a node it was handed fails again', () => {
expect(nextAttemptCount(checkpoint({ kind: 'partial', attempts: 1 }), ['node-a'])).toBe(2);
});
it('resets when the resume cleared its set and lost different nodes', () => {
expect(nextAttemptCount(checkpoint({ kind: 'partial', attempts: 2 }), ['node-z'])).toBe(
undefined,
);
});
it('does not advance off an interrupted marker — the bound is a partial-only rule', () => {
expect(nextAttemptCount(checkpoint({ kind: 'interrupted' }), ['node-a'])).toBe(undefined);
});
it('starts a fresh chain at 1 when the prior marker carried no count', () => {
expect(nextAttemptCount(checkpoint({ kind: 'partial', attempts: undefined }), ['node-a'])).toBe(
1,
);
});
});
@@ -0,0 +1,65 @@
import { describe, expect, it } from 'vitest';
import type { PersistedEmbeddingCount } from '../../src/core/embedding-count.js';
/**
* ── #2790 review, finding 2: ONE counter, three call sites ─────────────
*
* `run-analyze.ts` (Phase 5 + the mid-run checkpoint) and `server/api.ts` both
* publish `RepoMeta.stats.embeddings`. Their hand-copied bodies had already
* drifted inside a single change — one ended `?? 0`, the other `?? Number.NaN`
* — under a comment asserting they measured it "the same way". These pin the
* shared implementation both now call, on the three inputs that told the two
* copies apart.
*/
describe('measurePersistedEmbeddingCount (#2790 review)', () => {
const answer = async (
rows: Array<Record<string, unknown>> | undefined,
): Promise<PersistedEmbeddingCount> => {
const { measurePersistedEmbeddingCount } = await import('../../src/core/embedding-count.js');
return await measurePersistedEmbeddingCount(async () => rows);
};
it('reports a real count when the query answers with one', async () => {
expect(await answer([{ cnt: 41 }])).toEqual({ kind: 'measured', count: 41 });
});
it('reports an empty table as a measured zero', async () => {
// The distinction the whole tri-state rests on: an EMPTY table still
// answers with one row holding 0. That is a measurement, not an absence.
expect(await answer([{ cnt: 0 }])).toEqual({ kind: 'measured', count: 0 });
});
it('reports a no-row answer as unknown, not as zero', async () => {
// `?? 0` made this a measured zero — and `Number.isFinite(0)` is true, so
// the unknown branch became unreachable for exactly this case.
expect(await answer([])).toMatchObject({ kind: 'unknown' });
});
it('reports a missing cell as unknown', async () => {
expect(await answer([{}])).toMatchObject({ kind: 'unknown' });
});
it('reports a non-numeric cell as unknown', async () => {
expect(await answer([{ cnt: 'x' }])).toMatchObject({ kind: 'unknown' });
});
it('reports an undefined result set as unknown', async () => {
expect(await answer(undefined)).toMatchObject({ kind: 'unknown' });
});
it('reports a thrown query as unknown, carrying the reason', async () => {
const { measurePersistedEmbeddingCount } = await import('../../src/core/embedding-count.js');
expect(
await measurePersistedEmbeddingCount(async () => {
throw new Error('Binder exception: Table CodeEmbedding does not exist.');
}),
).toEqual({ kind: 'unknown', reason: 'Binder exception: Table CodeEmbedding does not exist.' });
});
it('collapses to the optional number the callers persist', async () => {
const { persistedEmbeddingCountOrUndefined } =
await import('../../src/core/embedding-count.js');
expect(persistedEmbeddingCountOrUndefined({ kind: 'measured', count: 0 })).toBe(0);
expect(persistedEmbeddingCountOrUndefined({ kind: 'unknown', reason: 'nope' })).toBeUndefined();
});
});

Some files were not shown because too many files have changed in this diff Show More