Compare commits

...
Author SHA1 Message Date
gitnexus-release-bot[bot] d6173988e1 release: v1.6.9-rc.34 2026-07-01 06:56:36 +00:00
Gergő Magyar 905b7dfa21 feat(embeddings): compact, description-forward embedding text (#2333) (#2334) 2026-07-01 07:37:16 +01:00
Parafee41 f5a2e6a248 fix(search): make vector distance threshold configurable (#2330) 2026-07-01 05:37:46 +01:00
Gergő MagyarandClaude Opus 4.8 e148bc089a fix(group): replace LadybugDB-incompatible multi-label Cypher (#2325) (#2327)
* fix(group): use labels(n) IN allowlist instead of LadybugDB-incompatible multi-label Cypher (#2325)

manifest-extractor and http-route-extractor built Cypher with the openCypher
label disjunction `MATCH (n:A|B|C)`, which LadybugDB's parser rejects. The
error was swallowed by try/catch, so manifest contracts silently fell back to
synthetic UIDs with empty filePath and http-route cross-file handler
resolution silently returned null.

Replace all 7 queries with `MATCH (n) WHERE labels(n) IN [...]`. LadybugDB
returns labels(n) as a single string, so this is an exact allowlist — a 1:1
behavior-preserving syntax translation (validated against LadybugDB 0.17.1).
Export the two http-route query constants so integration tests can run the
exact production strings against a real DB, and add per-branch real-DB
regression coverage (the bug shipped because no test exercised these queries).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(group): import CypherExecutor from contract-extractor in #2325 test

The new manifest regression test imported `CypherExecutor` from
`group/types.js`, which does not export it — the type is defined only in
`group/contract-extractor.js` (as all production extractors import it).
This was a real TS2305 under `tsc -p tsconfig.test.json`, masked from CI
because the default tsconfig excludes `test/` and `import type` is erased
at runtime. Split the import so the type resolves from its real module.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(group): run #2325 native-LadybugDB tests in the lbug-db project

Per TESTING.md, every test that opens a real `@ladybugdb/core` handle must
be registered in the sequential `lbug-db` Vitest project (and excluded from
`default`) to avoid native-mmap file-lock conflicts across parallel forks on
Windows. The two new group integration tests use `withTestLbugDB`/pool-adapter
but were in neither list, so they ran under the parallel `default` project.
Add both to `lbug-db.include` and `default.exclude`, matching every sibling.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(group): export custom-contract resolve query for #2325 test

The #2325 integration test hand-copied the 21-label `custom`-branch
resolve query into a local `LABELS_CUSTOM_QUERY` constant, so editing the
production allowlist would silently desync the canary. Promote the query to
an exported `CUSTOM_CONTRACT_RESOLVE_QUERY` (mirroring http-route-extractor's
exported query strings) and import it in the test, so the canary always runs
the exact production query. Behavior unchanged — same query string.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(group): de-brittle the #2325 custom-query label assertion

The unit test asserted a fixed 7-label ordered substring of the 21-label
custom-branch allowlist, coupling it to label order and no-space formatting —
a harmless reorder would have broken it. Replace with order/spacing-tolerant
membership checks for a spread of individual labels, keeping the unconditional
`not.toContain('Function|Method')` guard as the real regression check.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(group): correct #2325 http-route docstring + add real-trigger canary

The http-route test claimed `MATCH (n:Function|Method|CodeElement)` "which
LadybugDB rejects" — but that 3-label disjunction actually PARSES. Verified
against the real parser, the genuine #2325 trigger is a *reserved-keyword*
label in the disjunction: `Macro` and `Union` both are, and only the manifest
custom branch (21-label list) and the lib branch (missing `Package` table)
actually threw. The http-route conversion to `labels(n) IN [...]` was a
consistency change, not a parser fix.

Correct the misleading docstring and add a rejection canary pinned to the real
cause (`MATCH (n:Function|Macro|Union)` rejects), so a future query that
reintroduces a reserved-keyword disjunction is caught.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(group): cover the thrift package-strip path against a real LadybugDB

The thrift-only branch of resolveSymbol strips a `package.` prefix from the
service name (`com.example.AuthService` -> `AuthService`) before the
Class/Interface lookup — previously exercised only with a mocked executor.
Add a service-contract integration case (no method, so it takes the
package-strip path, not the grpc-identical method path) that resolves the real
`cls:AuthService`. Without the strip the lookup matches nothing and falls back
to a synthetic uid, so this is a non-vacuous guard for the strip.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(group): drop vestigial 'Package' label from lib contract lookup

The `lib` branch allowlisted `labels(n) IN ['Package','Module']`, but there is
no `Package` node table (see NODE_TABLES) — the entry only ever matched
nothing. Restrict to `['Module']`, the label libraries actually resolve to.
Behavior-neutral: the lib integration case still resolves its Module symbol.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(group): update PIPELINE label-scoped queries to labels(n) IN form

The resolveSymbol label-scoping bullets still showed the banned
`MATCH (n:A|B)` disjunction; a contributor copying them would reintroduce
#2325. Rewrite them in the actual `labels(n) IN [...]` form, note the real
trigger (LadybugDB rejects a disjunction naming a reserved keyword such as
`Macro`/`Union`), and reflect the lib allowlist as `['Module']` after dropping
the vestigial `Package` label.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(group): correct #2325 root-cause comments in the extractors

The production comments claimed LadybugDB rejects the `MATCH (n:A|B)`
disjunction "outright". Verified against the real parser, it rejects only when
a label is a reserved keyword (`Macro`, `Union`) or names a missing node
table. So only the manifest `custom` branch (reserved keywords in its 21-label
list) and the `lib` branch (missing `Package` table) actually threw; the
http-route/grpc/thrift/topic disjunctions parse fine and were converted to
`labels(n) IN [...]` for consistency and future-proofing, not because they were
broken. Rewrite the comments to say so accurately. No behavior change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(group): make #2325 test prose name the real reserved-keyword trigger

The manifest test docstring/title and the unit-test comment said LadybugDB
rejects the `MATCH (n:A|B)` disjunction generally. It rejects only when a label
is a reserved keyword (`Macro`/`Union`) or a missing table. Reword the docstring
(custom + lib branches threw; others parsed), retitle the rejection canary to
"its list names reserved keywords Macro/Union", and correct the unit-test
comment. The rejection canary still passes — the custom 21-label list does
contain Macro/Union. No behavior change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-30 15:08:43 +01:00
dependabot[bot]andAbhigyan Patwari 9c5a174303 chore(deps)(deps): bump commander from 14.0.3 to 15.0.0 in /gitnexus (#2322)
Bumps [commander](https://github.com/tj/commander.js) from 14.0.3 to 15.0.0.
- [Release notes](https://github.com/tj/commander.js/releases)
- [Changelog](https://github.com/tj/commander.js/blob/master/CHANGELOG.md)
- [Commits](https://github.com/tj/commander.js/compare/v14.0.3...v15.0.0)

---
updated-dependencies:
- dependency-name: commander
  dependency-version: 15.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-06-30 08:26:48 +01:00
dependabot[bot]andAbhigyan Patwari 15583fc9e9 chore(deps)(deps): bump onnxruntime-node in /gitnexus (#2321)
Bumps [onnxruntime-node](https://github.com/Microsoft/onnxruntime) from 1.26.0 to 1.27.0.
- [Release notes](https://github.com/Microsoft/onnxruntime/releases)
- [Changelog](https://github.com/microsoft/onnxruntime/blob/main/docs/ReleaseManagement.md)
- [Commits](https://github.com/Microsoft/onnxruntime/compare/v1.26.0...v1.27.0)

---
updated-dependencies:
- dependency-name: onnxruntime-node
  dependency-version: 1.27.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-06-30 08:13:12 +01:00
028bd11053 fix(group): cache read-only bridge handle to fix Windows @group reopen (#2274) (#2313)
* fix(group): cache read-only bridge handle to fix Windows @group reopen (#2274)

A long-lived MCP server opened bridge.lbug read-only, queried, and closed
it on every @group trace/impact call. On Windows the in-process reopen of the
same file fails (the OS handle is not fully released before the next open races
in), so repeated @group calls broke. #2269 fixed Linux/macOS by skipping
CHECKPOINT on read-only handles; Windows stayed broken.

Instead of fighting LadybugDB's Windows close/reopen timing: cache one
read-only handle per groupDir and reuse it across calls (open-once-per-process
already works on Windows). getCachedBridgeReadOnly:
  - reuses a single handle keyed by resolved groupDir,
  - invalidates on mtime change (external writer / re-sync),
  - invalidates explicitly before same-process writes (writeBridge),
  - guards concurrent first-open with an in-flight promise (no handle leak),
  - closes all handles on process exit.

closeBridgeDb now no-ops for the cached handle (cache owns its lifetime);
uncached/writable handles are unaffected. ensureBridgeReady uses the cache.

The in-process write->read reopen of the same bridge.lbug file remains a known
LadybugDB Windows limitation, so the existing reopen tests stay win32-skipped.
A new cache-aware itCacheReopen gate applies to the 3 new tests whose setup
requires write-then-read in the same process (same class as itLbugReopen). The
cache itself exercises read->read reuse and is unaffected.

* fix(group): harden bridge RO-handle cache for concurrency, lifetime & Windows (#2313 review)

Addresses the tri-review + Copilot findings on the read-only bridge-handle cache:

- P1 (F2): serialize queryBridge per cached handle via a per-handle FIFO lock
  (the conn-lock.ts chain mechanic, keyed per cache entry, not the global lock).
  Two concurrent @group callers sharing one lbug.Connection can no longer
  dispatch two queries at once (the heap-corruption hazard). Uncached/writable
  handles skip the lock at zero cost.
- P1 (F3): refcount lease — getCachedBridgeReadOnly acquires, closeBridgeDb
  releases (no caller change). The native close is deferred until in-flight
  readers drain (refs===0) and runs exactly once (closeStarted guard).
  invalidateBridgeCache and the mtime-evict path share one evict/close path.
- Windows: bounded drain in evictBridgeEntry — a concurrent group_sync waits
  (<= WINDOWS_DRAIN_TIMEOUT_MS) for readers to release before the atomic rename
  on win32 so it stays clean; POSIX remains fully non-blocking; single-threaded
  sync still closes-before-rename on all platforms.
- P0 (F1/F6): gate the mtime cache test with itCacheReopen (win32-skipped) and
  drop the manual invalidate so writeBridge self-invalidation is under test;
  add an external-writer (fsp.utimes) reopen case.
- Windows coverage (F9): new cross-process integration test seeds bridge.lbug
  in a separate tsx process, so read->read handle reuse is proven on win32 CI
  (not skipped). Plus concurrent cold-open dedupe coverage.
- P2/P3: scope the Windows NOTE to read->read (F4); JSDoc the closeBridgeDb
  release/close contract (F5); drop the if-branch in the B2 probe (F7); revert
  incidental Prettier churn in cross-impact.ts (F14); fix the stale describe
  header (F15); document the beforeExit/signal and ENOENT-mtime behavior
  (F11/F13).

tsc clean; group unit + integration suites green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(group): run the B2 rename-clash probe on win32 via cross-process seed (#2313 review)

Moves the B2 "external rename while a cached RO handle is held" probe out of the
unit suite (where it was win32-skipped, because its in-process writeBridge->RO-open
is the unfixed Windows reopen) into the cross-process integration test, where a
separate-process seed makes the RO open clean. The probe now RUNS ON WIN32 CI and
empirically answers whether an open RO handle blocks an external atomic rename over
bridge.lbug — the assumption under writeBridge's invalidate-before-rename and the
win32 drain.

Hardened (per adversarial review) so a win32 RED is the real steady-state share-mode
signal, not an artifact:
- use production retryRename (not bare fsp.rename) so transient EBUSY/EPERM from the
  Windows AV/indexer scanning the fresh temp file is absorbed; a RED then means the
  rename is blocked even after retries (FILE_SHARE_DELETE absent -> invalidate-before-
  rename is load-bearing).
- stage the byte-identical replacement BEFORE opening the RO handle, so no second OS
  handle touches bridge.lbug while LadybugDB holds it (avoids a FILE_SHARE_READ red for
  the wrong question).
- drop the post-rename query (handle survival is covered by the reuse test); the probe's
  sole verdict is whether the rename is blocked.

Removes the old win32-skipped unit B2 (a strict subset of the new probe).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-30 07:49:47 +01:00
dependabot[bot] c1a1b2a553 chore(deps)(deps): bump onnxruntime-common in /gitnexus (#2320)
Bumps [onnxruntime-common](https://github.com/Microsoft/onnxruntime) from 1.26.0 to 1.27.0.
- [Release notes](https://github.com/Microsoft/onnxruntime/releases)
- [Changelog](https://github.com/microsoft/onnxruntime/blob/main/docs/ReleaseManagement.md)
- [Commits](https://github.com/Microsoft/onnxruntime/compare/v1.26.0...v1.27.0)

---
updated-dependencies:
- dependency-name: onnxruntime-common
  dependency-version: 1.27.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-30 07:46:38 +01:00
azizur100389 8ad4469e96 fix(test): stabilize local Windows gate baselines (#2314) 2026-06-29 22:27:50 +01:00
Parafee41 a7df8f861a fix(search): make FTS stemmer configurable (#2307) 2026-06-29 05:13:31 +01:00
Parafee41 7ca7166b8e fix(fastapi): apply APIRouter constructor prefixes (#2312) 2026-06-28 13:37:31 +01:00
Gergő MagyarandClaude Opus 4.8 57e4afa4c8 fix(mcp): stabilize api_impact response shape for same-URL multi-verb routes (#2308) (#2309)
* fix(mcp): stabilize api_impact response shape for same-URL multi-verb routes

After #2302 made Route identity method-aware, a same URL exposes one Route
node per HTTP verb, so a bare-URL api_impact lookup could silently flip from a
direct route object to the wrapped { routes, total } envelope. Surface each
route's `method` (via the shared fetch) so multi-verb results are
distinguishable, and add an optional `method` selector that narrows a
multi-verb URL/file to one verb and forces the singular shape. A verb that
matches no route returns a clear error. Document the match-count contract in
the tool schema.

Refs #2308

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(mcp): cover same-URL multi-verb api_impact contract

Regression coverage for #2308: bare-URL and bare-file lookups of a same-URL
GET+POST pair return the wrapped form with distinct per-route methods; the
method selector collapses to the singular shape (case-insensitively); an
unmatched verb returns a verb-not-found error; and verbless routes surface a
null method.

Refs #2308

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(review): apply autofix feedback

- tools.ts: correct api_impact contract docs — `method` narrows to one verb
  but the singular shape only holds when exactly one route remains after
  filtering (substring route/file matches can still wrap); cover file lookups;
  enumerate verbs.
- local-backend.ts: surface `method` in route_map and shape_check output (the
  shared fetch already returns it; agents discover verbs there before
  api_impact).
- local-backend.ts: compute routeCountByHandler from the unfiltered match so a
  method-scoped api_impact still flags a multi-verb handler's partial middleware.
- tests: add file+method and verbless-exclusion cases; assert unconditionally
  via toMatchObject; lowercase the verb-not-found input to exercise error
  uppercasing.

Refs #2308

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(mcp): treat wildcard '*' routes as matching any api_impact method selector (#2308)

Method-agnostic routes (Django function views) persist with Route method
'*', not null. The api_impact method selector used exact verb equality, so
'*' routes were excluded and api_impact({route, method:'POST'}) falsely
reported 'No routes found' for a route that handles every verb. Treat '*'
as matching any requested verb, and correct the comment + tool-description
strings that wrongly grouped Django wildcards with null/verbless routes.

* fix(mcp): harden api_impact method input against non-string and empty values (#2308)

The MCP envelope is not schema-validated, so a non-string `method` reached
`.toUpperCase()` and threw a TypeError. Widen the param to `unknown` and guard
it with a typeof check that returns a structured error (mirroring the
resolveAliasString pattern from #2175), and collapse empty/whitespace verbs to
no selector.

* fix(mcp): distinguish url-not-found from verb-not-found in api_impact error (#2308)

The verb-not-found error appended 'with method "X"' even when the URL/file
itself did not exist, implying the URL exists with other verbs. Gate the verb
clause on matched.length > 0 so a non-existent URL/file gets the plain message.

* fix(mcp): clarify api_impact middlewareNote wording for verbless siblings (#2308)

The partial-middleware note claimed 'other methods in this handler' even when
the co-located sibling is a verbless (null) route rather than another HTTP
verb. Refer to 'other route exports' instead, which covers both cases.

* docs(mcp): document and test the method field on route_map and shape_check (#2308)

The shared fetchRoutesWithConsumers change surfaced a method key on route_map
and shape_check responses too, but their tool descriptions never mentioned it
and no test covered it. Document the field on both descriptions and add unit
tests asserting it (shape_check rows carry responseKeys + a consumer so they
survive shape_check's keys-and-consumers filter).

* test(mcp): cover middlewareDetection 'partial' survival under a method filter (#2308)

The diff's core behavioral line counts verbs-per-handler from the unfiltered
match set so a method-scoped query still flags a multi-verb handler's partial
middleware, but no test exercised it (every verbRow hardcoded middleware:null).
Add a middleware param to verbRow and a test that fails if the count is taken
from the post-filter set instead. Verified via mutation: matched->routes fails it.

* test(mcp): add live-LadybugDB integration coverage for route method round-trip (#2308)

The new n.method query column was only unit-mocked. Add a self-contained
integration suite that seeds GET+POST /api/orders and a method-agnostic '*'
Django route, then asserts api_impact surfaces method, narrows by verb, and
matches the '*' route end-to-end (the U1 fix), plus route_map surfacing.
Own seed + no FTS so it neither perturbs api-impact-e2e nor silently skips.

* refactor(mcp): type the api_impact response shape instead of Promise<any> (#2308)

Replace apiImpact's Promise<any> with an explicit ApiImpactResult union
(single route | wrapped { routes, total } | { error }) and a typed
ApiImpactRoute. The results.map is annotated so the response builder is
checked against the declared shape. Behavior unchanged; sibling MCP methods
keep their Promise<any> convention.

* fix(mcp): express the route-or-file requirement in the api_impact schema (#2308)

The inputSchema left route/file as bare optionals, so the 'at least one of
route/file' rule the handler enforces was invisible to clients. Add an optional
anyOf to ToolDefinition (forwarded verbatim by the ListTools handler) and an
anyOf:[{required:[route]},{required:[file]}] on api_impact. Matches runtime
(both allowed, route wins); 'at least one' not 'exactly one'.

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-26 19:50:10 +01:00
Parafee41 c45d38f27a fix(ingestion): index generator function declarations (#2305) 2026-06-26 13:37:02 +01:00
8bef64baa1 feat(ingestion/routes): give Route nodes a (method, url) identity (#2289) (#2302)
* feat(ingestion/routes): give Route nodes a (method, url) identity (#2289)

Route node identity was URL-only, so a same-URL multi-verb pair
(GET /x + POST /x) collapsed into a single node and the second verb's
handler and execution flow were silently lost. Route identity is now
(method, url) via routeNodeKey(method, url): a known, specific verb keys
as "METHOD url", while a method-less route (filesystem routes — Next.js /
Expo / PHP — and Laravel resource/apiResource) or a wildcard "*" route
(e.g. Django function views) falls back to URL-only. The fallback is
byte-identical to the previous URL-only ids, so only genuine
declaration-style multi-verb routes split into separate nodes.

The identity key is shared across the three phases that must agree on the
Route node id:
- routes phase: registry key + node id + handler-symbol lookup; the Route
  node still carries the bare URL as its display name.
- call-processor: resolveRouteHandlerSymbols re-keyed by identity so each
  verb resolves its own handler; a verb-less fetch() consumer matches by
  URL and connects to every Route node at that URL (one per verb).
- processes phase: ENTRY_POINT_OF targets the identity-keyed node id.

Bumps INCREMENTAL_SCHEMA_VERSION 4 -> 5: persisted pre-v5 Route nodes use
the old url-only ids, so an incremental top-up would strand them alongside
new composite-keyed nodes — force a full re-analyze instead.

Part of #2280.

* fix(ingestion/routes): address PR #2302 review (P1/P2/P3)

P1 — Schema v5 fast-path bypass (run-analyze.ts):
  Adds a schemaVersion-mismatch guard above the alreadyUpToDate early-return,
  mirroring the pdgModeMismatch slot. Without it, a same-commit re-analyze on
  a pre-v5 stamp returned alreadyUpToDate without ever reaching the
  isIncremental gate, defeating the v5 schema bump's migration intent.
  Regression test covers: analyze (stamps v5) → meta downgrade to v4 → same
  commit re-analyze must NOT early-return and meta restamps to v5.

P2 — ENTRY_POINT_OF handler-aware linking (processes.ts):
  Pre-fix routesByFile fanned every same-file Route to every same-file
  process, cross-wiring same-file GET/POST handlers. Now reads handlerSymbolId
  off the Route graph node (the source of truth routes.ts stamps) into
  routesByHandlerId, with a routesWithoutHandlerByFile fallback — mirrors the
  Tool linking precedent 10 lines below. Two regression tests: weak form
  (only one handler has a process; sibling verb does not get spuriously
  attached) and strong form (both handlers form distinct processes; each
  Route links to exactly its own entryPoint, 2 edges not pre-fix 4).

P2 — Roundtrip composite-id (route-{method,handler-symbol}-roundtrip):
  Both tests now seed the Route node with
  generateId('Route', routeNodeKey('POST', '/api/orders')) and run the
  Cypher MATCH against the composite id, exercising the literal-space-in-id
  through CSV→COPY→HANDLES_ROUTE_QUERY. A space-in-id escape regression
  would surface here instead of being silently swallowed by the extractor's
  catch.

P3 — doc-drift + test if:
  - route-path.ts:4 — header updated to "(method, url) via routeNodeKey"
  - java.ts:684 — drop "Route nodes are URL-keyed"; #2289 closes that gap
  - manifest-extractor.ts:196 — explicit that Route node *id* is composite
    while route.name remains the bare URL
  - multi-verb-route-identity.test.ts:88 — forEachRelationship+if rewritten
    as a .filter().map() chain (no test-level conditional). New
    route-process-linking tests are also if-free.

Validation: tsc clean, prettier clean, 9 touched suites / 43 tests pass.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(ingestion/routes): drop routes.ts re-export, fix CI fast-path tests

Two follow-ups on PR #2302's CHANGES_REQUESTED review:

1. Drop `routes.ts` re-export of `normalizeExtractedRoutePath` /
   `normalizeRouteMethod` / `routeNodeKey` (per @magyargergo's inline
   comment at routes.ts:153 — the symbols already live in
   `route-extractors/route-path.ts` and consumers should import them
   from the source, not via a routes-phase indirection that was kept
   only as a compat shim during the #2289 refactor). Updated the two
   remaining callers (blade-template-routes / spring-route-extractor-
   parity tests) to import directly from `route-extractors/route-path.js`.
   `call-processor.ts` and `processes.ts` already import from the source.

2. Fix two `run-analyze.test.ts` fast-path tests that started failing
   on CI after the schema-version mismatch guard landed (
   "creates .gitnexus/.gitignore on the already-up-to-date fast path"
   and "reports isPrimaryBranch false for an up-to-date non-primary
   branch"). The test fixtures hand-built a RepoMeta with NO
   schemaVersion field; with the guard now checking
   `existingMeta.schemaVersion !== INCREMENTAL_SCHEMA_VERSION`, that
   pre-versioning shape was treated as a mismatch and forced a rebuild,
   short-circuiting the fast path the tests exercise. Stamp the current
   schemaVersion on those fixtures so they reflect the post-#2289 meta
   shape production actually writes (`runFullAnalysis` always stamps
   the field on git repos — see meta save site).

Validation: tsc clean, prettier clean, 11 touched suites / 80 tests pass.

---------

Co-authored-by: henry <zhangwei2017@unipus.cn>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-26 07:59:46 +01:00
Gergő Magyar 576e81442e fix(search): index description field for FTS so doc comments are keyword-searchable (#2300)
* fix(search): index description column for FTS so doc comments are keyword-searchable

Closes #2299. descriptionExtractor (#2286) populates the `description`
column for every symbol table, but FTS only indexed name+content on 5
tables, so doc-comment keywords (Javadoc/KDoc/godoc/Rust ///) were
invisible to BM25 keyword search.

- Add `description` to the Function/Class/Method/Interface FTS indexes
  (File has no description column, left as name+content).
- Add FTS indexes for the remaining EMBEDDABLE_LABELS symbol tables
  (Struct, Enum, Trait, Impl, Macro, Namespace, Constructor, TypeAlias,
  Typedef, Const, Property, Record, Union, Static, Variable).
- createSearchFTSIndexes now drops-then-creates each index so the schema
  change reaches existing DBs on incremental re-analyze and --repair-fts
  (createFTSIndex is idempotent-by-name and would otherwise skip stale
  indexes).

Tests: fts-schema column-subset + coverage guards; drop-before-create
order; e2e doc-comment keyword search (Java class + Rust struct found by
description-only terms). bm25-search assertions derive from FTS_INDEXES.

* fix(review): apply autofix feedback

- Guard the --repair-fts path on FTS-extension availability before
  createSearchFTSIndexes drops-then-creates indexes (P1 regression:
  without the gate, an unavailable extension could drop existing indexes
  then fail to recreate them, leaving the DB index-less). Mirrors the
  analyze path's ftsAvailable gate and fails loudly first.
- Add a re-analyze upgrade integration test: seed an old name+content-only
  DB (no Struct index), run the real createSearchFTSIndexes(), and assert
  description keyword search + the previously un-indexed Struct now resolve.
  Proves drop-then-create upgrades a live stale index end-to-end.

* fix(ci): add loadFTSExtension to --repair-fts test mocks

The R3 review fix added a loadFTSExtension availability gate to the
--repair-fts path, but run-analyze-fts-repair.test.ts mocked the lbug
adapter without that export, so both repair tests threw `No
"loadFTSExtension" export`. Add loadFTSExtension to the two mocks
(returning true to preserve their original intent) and add a dedicated
test proving the guard fails loudly — and does NOT drop any index —
when the extension is unavailable.

* test(fts): run fts-description-search in the sequential lbug-db project

It was the only FTS-index-creating integration test left in the parallel
`default` vitest project; every other ftsIndexes-using test (search-core,
search-pool, augmentation, …) runs in the `lbug-db` project, which forces
fileParallelism: false to avoid LadybugDB native mmap file-lock conflicts
in parallel forks (Windows). Add it to the lbug-db include list and the
default exclude list to match the convention and remove the flake risk.

* test(ci): fail loudly when FTS extension is unavailable, never silently skip

FTS-dependent lbug integration suites (search-core, search-pool,
augmentation, fts-description-search, …) self-skip via ctx.skip() when the
LadybugDB FTS extension can't load, emitting only a console.warn while the
job stays green. That means a broken/missing FTS extension in CI would make
these integration tests silently vanish with no signal — false confidence.

withTestLbugDB now honors GITNEXUS_REQUIRE_FTS=1: when set and the extension
is unavailable, setup() throws instead of skipping, so the suite fails
loudly. The CI test jobs (ubuntu coverage + windows/macOS cross-platform)
set the flag; local/offline runs leave it unset and keep skipping
gracefully. (Verified the extension currently loads on all three runners,
so this is a guard against regression, not a behavior change today.)

* test(ci): run fts-description-search on macOS/Windows cross-platform jobs

The new FTS description-search suite was registered in the sequential
lbug-db vitest project (ubuntu/coverage) but absent from LBUG_NATIVE, so
the macOS/Windows platform-sensitive jobs (which run only the explicit
ALL_CROSS_PLATFORM allowlist via run-cross-platform.ts) never executed it.
The GITNEXUS_REQUIRE_FTS=1 hardening on those jobs guarded the old FTS
fixtures but not the new 20-index/description path. Add the suite to
LBUG_NATIVE so the new path is validated cross-platform too.

Refs #2299.

* fix(search): verify FTS indexes cover description, not just queryability

verifySearchFTSIndexes probed each index with QUERY_FTS_INDEX and treated
'queryable' as 'present'. A stale name+content-only index left on a
pre-#2299 DB stays queryable yet silently misses the description column, so
verification would pass green while doc-comment search stayed broken.

Switch to a single CALL SHOW_INDEXES() that exposes property_names per
index, and report an index as missing when it is absent OR does not cover
its configured columns. Return contract (string[] of table.indexName) is
unchanged, so both run-analyze.ts call sites are untouched. The per-index
string interpolation is gone, so the now-dead safeIdentifier helper is
removed.

The real caller of the live function in tests is bm25-search.test.ts (the
repair test mocks verifySearchFTSIndexes wholesale); its two probe-shaped
cases are rewritten to feed SHOW_INDEXES rows and now assert column
coverage, plus an absent-index case.

Refs #2299.

* test(search): assert description search via the public query surface

The #2299 integration suite only exercised the searchFTSFromLbug helper.
Add a third block that drives the public LocalBackend.callTool('query')
path — which resolves the repo via the registry and routes BM25 through the
pool adapter (a different connection context than the core-adapter helper) —
and asserts a description-only keyword returns the seeded class. Reuses the
existing description-only SEED and production FTS_INDEXES; partial-mocks
repo-manager so listRegisteredRepos points at the test DB while
cleanupOldKuzuFiles and the rest stay real.

Refs #2299.

* test(search): make lbug-core-adapter FTS gate honor GITNEXUS_REQUIRE_FTS

lbug-core-adapter.test.ts has its own per-test FTS gate (skipUnlessFtsAvailable)
that called ctx.skip() whenever the extension could not load — bypassing the
GITNEXUS_REQUIRE_FTS=1 hardening that withTestLbugDB already honors. Since this
file is in LBUG_NATIVE it runs on the ubuntu/macOS/windows jobs that all set
GITNEXUS_REQUIRE_FTS=1, so an FTS regression on a runner would have let these
FTS-primitive tests silently vanish from a green run — the exact gap #2299's
test-infra hardening set out to close.

Make the helper mirror withTestLbugDB: when GITNEXUS_REQUIRE_FTS=1 and the
extension is unavailable, throw (hard fail) instead of skipping. Offline/local
runs (no env var) still skip gracefully.

Refs #2299.
2026-06-25 14:21:44 +01:00
henry201605 d7ff76e6e9 fix(ingestion/routes): resolve Spring interface-inherited routes (#2288) (#2290) 2026-06-25 09:22:20 +01:00
dependabot[bot] 269737982e chore(deps)(deps): bump @langchain/openai in /gitnexus-web (#2291)
Bumps [@langchain/openai](https://github.com/langchain-ai/langchainjs) from 1.4.5 to 1.5.0.
- [Release notes](https://github.com/langchain-ai/langchainjs/releases)
- [Commits](https://github.com/langchain-ai/langchainjs/commits/@langchain/openai@1.5.0)

---
updated-dependencies:
- dependency-name: "@langchain/openai"
  dependency-version: 1.4.7
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-25 06:53:56 +01:00
5f667c32a3 chore(deps): bump actions/checkout from 6.0.3 to 7.0.0 (#2292)
* chore(deps): bump actions/checkout from 6.0.3 to 7.0.0

Bumps [actions/checkout](https://github.com/actions/checkout) from 6.0.3 to 7.0.0.
- [Release notes](https://github.com/actions/checkout/releases)
- [Changelog](https://github.com/actions/checkout/blob/main/CHANGELOG.md)
- [Commits](https://github.com/actions/checkout/compare/df4cb1c069e1874edd31b4311f1884172cec0e10...9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0)

---
updated-dependencies:
- dependency-name: actions/checkout
  dependency-version: 7.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>

* ci: set persist-credentials: false on read-only checkouts (zizmor artipacked)

Adds `persist-credentials: false` to the 9 checkout steps flagged by
zizmor's credential-persistence (artipacked) rule on PR #2292. All are
read-only CI/test/quality jobs that never use the git token afterward, so
not persisting it removes the leak surface. Checkouts that push (publish,
pr-autofix, commit-fork-prebuilds, etc.) keep credentials and are untouched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-25 06:53:47 +01:00
dependabot[bot] 9b8d31a1f2 chore(deps)(deps): bump langchain from 1.4.4 to 1.4.6 in /gitnexus-web (#2294)
Bumps [langchain](https://github.com/langchain-ai/langchainjs) from 1.4.4 to 1.4.6.
- [Release notes](https://github.com/langchain-ai/langchainjs/releases)
- [Commits](https://github.com/langchain-ai/langchainjs/compare/@langchain/openai@1.4.4...langchain@1.4.6)

---
updated-dependencies:
- dependency-name: langchain
  dependency-version: 1.4.6
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-25 06:28:09 +01:00
dependabot[bot]andGergő Magyar e6f2296d00 chore(deps)(deps-dev): bump @vitest/coverage-v8 in /gitnexus-web (#2297)
Bumps [@vitest/coverage-v8](https://github.com/vitest-dev/vitest/tree/HEAD/packages/coverage-v8) from 4.1.8 to 4.1.9.
- [Release notes](https://github.com/vitest-dev/vitest/releases)
- [Changelog](https://github.com/vitest-dev/vitest/blob/main/docs/releases.md)
- [Commits](https://github.com/vitest-dev/vitest/commits/v4.1.9/packages/coverage-v8)

---
updated-dependencies:
- dependency-name: "@vitest/coverage-v8"
  dependency-version: 4.1.9
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-06-25 06:27:43 +01:00
dependabot[bot] a05a1659bd chore(deps): bump release-drafter/release-drafter from 7.3.1 to 7.4.0 (#2295)
Bumps [release-drafter/release-drafter](https://github.com/release-drafter/release-drafter) from 7.3.1 to 7.4.0.
- [Release notes](https://github.com/release-drafter/release-drafter/releases)
- [Commits](https://github.com/release-drafter/release-drafter/compare/693d20e7c1ce1a81d3a41962f85914253b518449...ed4bc48ec97379be2258e7b7ac2624a3e26ab809)

---
updated-dependencies:
- dependency-name: release-drafter/release-drafter
  dependency-version: 7.4.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-25 06:25:41 +01:00
dependabot[bot] ba071b5bb3 chore(deps)(deps): bump lru-cache from 11.3.6 to 11.5.1 in /gitnexus-web (#2298) 2026-06-25 01:15:40 +01:00
dependabot[bot] 5165686798 chore(deps)(deps): bump @langchain/core in /gitnexus-web (#2293) 2026-06-24 22:15:21 +01:00
9aa65ae3f8 feat: ✨ resolve Nuxt/Nitro auto-imports in TypeScript scope resolver (#2026)
* feat: ✨ resolve Nuxt/Nitro auto-imports in TypeScript scope resolver

* fix: 🐛 skip self-referential edges in Nuxt auto-import emission

* fix: 🐛 address Sourcery review -- gate Nitro scan on imports.d.ts and pre-index explicit imports

* fix: scope Nuxt auto-import resolution

* fix: address Nuxt auto-import review follow-ups

* fix(ingestion): capture only LHS binding names in Nitro server-util exports

The Nuxt server-util export scanner ran a declarator regex over the whole
`export const …` right-hand side, so it registered RHS tokens as auto-import
names: arrow-function parameters (`export const f = (event) => …` → `event`),
object-literal keys (`export const c = { onError } ` → `onError`), and bare
operands. It also dropped generic-typed declarators
(`export const x: Map<a, b> = …`) because the type-annotation skip broke at the
comma inside the generic. Both produced wrong/missing auto-import CALLS edges.

Capture only the leading binding name of each top-level declarator via a
depth-aware comma splitter (tracks (), [], {}, <>), skipping destructuring
patterns. Nitro auto-imports only surface top-level binding names, so the RHS
is never parsed. Adds unit coverage for the param/object-key/operand/generic
and multi-declarator forms.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4W25WLfYD1JNy8icxeLPU

* fix(ingestion): stop Nitro server callers resolving client composables

`getNuxtAutoImportEntry` fell back to the client composable map when a
`server/api|routes|middleware` caller's name had no `server/utils` entry. But
Nitro only auto-imports `server/utils/**` into the server context — app
`composables/` are Vue-app-only — so that fallback minted CALLS/IMPORTS edges
Nitro never creates (e.g. a server route "calling" a composable it cannot see
without an explicit import).

Server callers now resolve the server map only. Restructure the barrel-directory
integration test to use a client caller (which legitimately auto-imports the
composable) so `index.*` resolution stays covered, and add a negative assertion
that `server/api/route.ts` emits no edge to `composables/*` while its real
`server/utils` call still resolves. Unit test locks that a server caller does
not fall back to a client-only name.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4W25WLfYD1JNy8icxeLPU

* fix(ingestion): let unresolved explicit imports shadow Nuxt auto-imports

The explicit-import suppression index only recorded import local names whose
edge resolved to a file (`edge.targetFile !== null`). An explicit import from an
unresolved external package — `import { useAuto } from '@vueuse/core'; useAuto()`
— therefore escaped suppression, and the post-resolution hook emitted a spurious
Nuxt auto-import CALLS edge for a name the file already imports explicitly.

Record the local name regardless of whether the import resolved: an explicit
import is authoritative shadowing intent. Adds an integration fixture importing
from an external package and a (non-vacuous) assertion that it emits no nuxt
edge.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4W25WLfYD1JNy8icxeLPU

* fix(ingestion): let type-annotated params shadow Nuxt auto-imports

hasLocalBindingInScopeChain only consulted scope.bindings, but type-annotated
function parameters live in scope.typeBindings (the TS scope query records them
as `@type-binding.parameter`, not `@declaration`). A parameter named like a
composable therefore failed to suppress the auto-import, leaking a spurious
CALLS edge.

Also check scope.typeBindings for the name (same-file scopes only). typeBindings
holds value-space binders' type facts (parameter annotations, `self`, variable
annotations) and never a pure type that belongs to callable space, so this
cannot over-suppress a real auto-import.

Documents the residual: function-typed params (`p: () => void`), untyped params,
destructured locals, and catch-clause vars are captured by neither map and still
leak — closing that needs shared scope-query changes beyond this feature, left
as a follow-up. Also adds a no-vacuous-pass guard to the shadowing/noise test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4W25WLfYD1JNy8icxeLPU

* fix(ingestion): treat server/plugins and server/tasks as Nitro runtime

isNitroServerRuntimeFile only matched server/api, server/routes, and
server/middleware. Nitro also auto-imports server/utils into server/plugins
and (since Nitro 2.6) server/tasks, so callers there were misrouted to the
client composable map. Extend the prefix set (now a named constant) to cover
them.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4W25WLfYD1JNy8icxeLPU

* fix(ingestion): merge duplicate JSDoc on collectImportsDts

Two consecutive JSDoc blocks preceded collectImportsDts; tooling (IDEs,
TypeDoc) attaches only the last one, silently dropping the descriptive block.
Fold the "returns true when read" line into the descriptive block as a
`@returns` tag.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4W25WLfYD1JNy8icxeLPU

* fix(ingestion): contain .nuxt/imports.d.ts source resolution to the repo

A crafted `.nuxt/imports.d.ts` source such as `from '../../../../etc/passwd'`
passes the project-local relative-path check but resolves outside the analyzed
repo, causing fs.stat probes against arbitrary host paths. Skip any source that
resolves outside repoRoot before touching the filesystem.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4W25WLfYD1JNy8icxeLPU

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 12:32:51 +01:00
Gergő MagyarandClaude Opus 4.8 8886d55008 feat(ingestion): make doc comments searchable across all languages (#2286)
* feat(ingestion): add shared leading-doc-comment description extractor (#2270)

Add `extractLeadingDocComment` plus a language-neutral
`createLeadingDocDescriptionExtractor` factory and a shared
`DOC_BEARING_LABELS` set to `utils/ast-helpers.ts`. The helper pulls the
normalized text of a leading doc comment off a definition node's preceding
named sibling, covering both block doc comments (Javadoc/KDoc/JSDoc/PHPDoc/
Doxygen, opened by double-star or bang) and runs of line doc comments
(triple-slash, bang-slash, or caller-supplied prefixes such as Go's
double-slash or Ruby's hash). Grammar-agnostic by prefix match; widens
`getDefinitionNodeFromCaptures` to accept the optional-valued capture map.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eh2UmA6f2p25F3ow75Hjzx

* feat(languages): surface leading doc comments as description for all languages (#2270)

Register the leading-doc `descriptionExtractor` on every documentable
provider so Javadoc/KDoc/JSDoc/Doxygen/godoc/RDoc/`///` doc text lands in the
`description` column and reaches the embedding metadata header — making
methods/types semantically searchable by doc-only terms, matching the
behavior Python (docstring) and PHP (Eloquent) already had.

- Java, Kotlin, TypeScript, JavaScript, C, C++, C#, Dart, Rust, Swift: default
  config (block + triple-slash/bang-slash doc comments).
- Go: godoc double-slash leading comments.
- Ruby: leading hash (RDoc/YARD) comments.
- PHP: existing Eloquent metadata takes precedence, else PHPDoc docblock.

Field/property/variable/const docs are intentionally out of scope.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eh2UmA6f2p25F3ow75Hjzx

* fix(review): apply autofix feedback (#2270)

Code-review autofix pass on the leading-doc-comment extractor:
- Enforce start-row adjacency in the line-comment run so a doc run stops at a
  blank line (godoc/RDoc/rustdoc semantics). Prevents a Go license/earlier `//`
  block or a Ruby shebang + `# frozen_string_literal:` magic comment, separated
  by a blank line, from being absorbed into the first declaration's
  description. Adjacency uses startPosition.row (reliable across grammars).
- Fix the degenerate empty comment `/**/` producing a spurious `/` description.
- PHP: compose createLeadingDocDescriptionExtractor() as the docblock fallback
  instead of duplicating its body, and widen the param to CaptureMap to match
  the LanguageProvider hook contract.
- Drop the factory's unused `labels` option (no consumer overrides it).
- Add tests: degenerate `/**/`, multi-line `///` run, `//!` inner doc, `/*!`
  Doxygen block, Go/Ruby blank-line non-attachment + two-block adjacency, and
  PHP Eloquent-metadata-wins-over-docblock ordering.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eh2UmA6f2p25F3ow75Hjzx

* fix(ingestion): resolve exported TS/JS JSDoc via export_statement wrapper

Exported TS/JS declarations dropped their JSDoc: the TS query captures the
inner function_declaration/class_declaration, whose previousNamedSibling is
null because the JSDoc precedes the wrapping export_statement (PR #2286 review,
reproduced). Add a wrapperNodeTypes option to extractLeadingDocComment (folded
into a LeadingDocCommentOptions object threaded through the factory); when the
captured node yields no doc and its parent type is a configured wrapper, retry
from the parent. TS/JS providers pass ['export_statement']. Language config
stays at the call site (RFC #909). Mirrors the existing walk-up in
languages/javascript/captures.ts for JSDoc params.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ingestion): bound DOC_BEARING_LABELS to embeddable labels

Module/Delegate/Annotation were doc-bearing but absent from EMBEDDABLE_LABELS,
so their descriptions were extracted and written to the DB yet never embedded
or searchable (PR #2286 review) — wasted work, and the factory JSDoc overstated
"becomes semantically searchable". Remove those three labels so DOC_BEARING_LABELS
is a subset of EMBEDDABLE_LABELS, narrow the JSDoc, and add a subset-invariant
unit test to guard against drift. Making those labels (and C++ `Template`)
searchable needs an embedding-pipeline/schema change and is left as a follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ingestion): skip file-top license/header blocks as descriptions

A file-top /** … */ license/copyright/overview block has no package/import
sibling to shield it from the first declaration, so it was absorbed as that
symbol's description and polluted the embedding text (PR #2286 review). The
block-comment branch already cannot use a strict row-adjacency check (grammars
fold the trailing newline into the comment node), so match header markers
instead — SPDX-License-Identifier, @license/@file/@fileoverview, "Licensed
under", and copyright-with-(c)/year. Markers are specific enough not to fire on
an ordinary doc that merely mentions the word "copyright" (over-fire guard test).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ingestion): ignore Go/Ruby directive & magic comments in doc runs

Go build/tool directives (//go:build, //go:generate, // +build, //nolint, //line)
and Ruby magic comments / shebang (# frozen_string_literal:, # encoding:, # -*-,
#!, …) sitting directly above a symbol were folded into its description and
polluted the embedding text (PR #2286 review). Add a lineDirectivePrefixes option;
a matching line is skipped in the doc run (skip-and-continue, so a real doc above
an interleaved directive is still collected — godoc/RDoc semantics). Go and Ruby
providers supply their own directive prefixes (RFC #909 — config at the call site).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ingestion): guard descriptionExtractor call against throws

A throw inside any provider's descriptionExtractor escaped processFileGroup to
the language-group catch, which treats any throw as "parser unavailable" and
silently drops every remaining file in the group (PR #2286 review). Wrap the
call in try/catch + reportWarning, mirroring the adjacent extractTemplateConstraints
guard. Defensive parity — no behavior change on the success path.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ingestion): treat Rust //! and /*! as inner docs

Rust //! and /*! are INNER doc comments (they document the enclosing item/module),
not the following item, but the shared helper attached them to the next definition
(PR #2286 review; a test even enshrined the wrong behavior). Add a blockDocPrefixes
option (default ['/**','/*!']); the Rust provider opts out of both inner-doc markers
(lineCommentPrefixes ['///'], blockDocPrefixes ['/**']). Doxygen //! and /*! keep
working for C/C++ via the defaults. Flip the Rust //! test to a negative assertion
and add a Rust /*! negative case.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ingestion): strip bidi/zero-width controls from doc descriptions

Doc-comment text is attacker-influenceable (any indexed repo) and is returned
verbatim to MCP clients, so a description could smuggle Trojan-Source-style bidi
overrides or zero-width characters (PR #2286 review). Strip U+202A–202E,
U+2066–2069, U+200B–200D and U+FEFF in the doc-comment normalization path
(block + line). Scoped to the description path only — global sanitizeUTF8 is
deliberately left alone (pre-existing, affects all fields). Implemented with a
code-point predicate so no literal invisible bytes live in the source.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(ingestion): doc-comment helper maintainability cleanups

PR #2286 review nits (no behavior change): drop the unused `export` on
DEFAULT_LINE_DOC_PREFIXES (no importer outside ast-helpers.ts); widen
getLabelFromCaptures' captureMap param to `Record<string, SyntaxNode | undefined>`
to match getDefinitionNodeFromCaptures (all accesses are truthiness-guarded); and
merge the split ast-helpers import statements in dart/ruby/rust into one each.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(architecture): document descriptionExtractor LanguageProvider hook

descriptionExtractor is now a near-universal LanguageProvider field (issue #2270)
but was missing from the architecture "Key fields" table (PR #2286 review). Add a
row describing it and the shared createLeadingDocDescriptionExtractor factory.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(ingestion): end-to-end description searchability for exported symbols

The unit tests stop at the descriptionExtractor hook; nothing proved a doc
comment survives the full parse pipeline into node.properties.description (the
field the embedding metadata header reads) — the exact gap that hid the exported
TS/JS regression (PR #2286 review). Add an integration test running the real
worker pipeline over an exported, JSDoc'd TS function and asserting its node
description carries the doc text. Verified locally against a built worker
(20s); runs in CI via pretest:integration build.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(ingestion): prettier-wrap a long line in the doc-comment test

Formatting-only follow-up to the U3/U7 test additions so `quality / format` is
green. No behavior change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 08:36:55 +01:00
0936553d63 fix(ingestion/routes): recognise Spring method-level array-form route mappings (#2281)
* feat(routes): extract Spring method-level array-form routes in ingestion + extractor parity test (#2138 follow-up)

ingestion's `extractSpringRoutes` (route-extractors/spring.ts) matched only a
single string literal on `@(Get|...)Mapping`, so the array form
`@GetMapping({"/a","/b"})` produced no graph Route node — while the group-layer
`java.ts` scan did match it. That divergence was the root of the #2265 array-form
parse-skip gap.

- spring.ts: add the array-form alternation
  `[(string_literal) @value (element_value_array_initializer (string_literal) @value)]`
  to the two method-declaration query branches (positional + `path=`/`value=`),
  mirroring the group query. A multi-element array yields one match per element,
  so the Phase 2 loop emits one route per path with no other change. Class-level
  `@RequestMapping` array prefixes remain single-literal (rare; left to a
  follow-up).
- test: spring-route-parity runs one shared Java fixture through BOTH extractors
  (ingestion `extractSpringRoutes` + group `JAVA_HTTP_PLUGIN.scan`) and asserts
  identical provider {method,path} sets — the parity guard the maintainer asked
  for in #2078, so the two Spring extractors can't silently drift again
  (verified: reverting the array branch turns the parity test red).

* fix(ingestion/routes): suppress wrong unprefixed route under class-array @RequestMapping; cover named-array + class-array parity

Addresses PR review on #2281:
- P2 class-array wrong-route: class branches now match the array form only to detect it; a method-level array route under a class-level array-form @RequestMapping is suppressed rather than emitted with a dropped prefix, so ingestion stays a strict subset of the group scan. Scalar method paths under an array class prefix are unchanged (pre-existing). Full class-array cross-product support tracked in a follow-up.
- P2 named-array coverage: added value={...}/path={...} parity cases, a consumes/produces array false-positive case, and a dedicated empty-provider-set assertion.
- P3 stale comments: updated the routeCoverage comment in java.ts and the route-parse-skip test note; narrowed the parity test drift claim.

routeCoverage stays 'partial'.

---------

Co-authored-by: henry <zhangwei2017@unipus.cn>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-06-24 07:10:37 +01:00
dependabot[bot] ca396e38bc chore(deps)(deps): bump uuid from 14.0.0 to 14.0.1 in /gitnexus (#2285) 2026-06-24 06:44:25 +01:00
Gergő MagyarandClaude Opus 4.8 47477e5554 fix(mcp): tolerate adapter-materialized line:0 in impact callgraph mode (#2279) (#2283)
* fix(mcp): tolerate adapter-materialized line:0 in impact callgraph mode (#2279)

Some MCP client/agent adapters serialize an omitted optional numeric
field as `0` rather than dropping it, so callgraph `impact` calls arrive
carrying a spurious `line: 0`. `line` is a PDG-only statement anchor and
is meaningless on the callgraph path, so the backend rejected the call
("'line' is only supported with mode:'pdg'") and strict clients rejected
it client-side against the advertised `minimum: 1`.

Treat a literal `line: 0` as omitted in `_impactImpl` when mode !== 'pdg'
and let the normal symbol→symbol BFS run. The coercion is deliberately
narrow: only the literal 0, only on the callgraph path. A genuine
positive `line` on callgraph still errors (real mode mistake), negative/
fractional values still error, and pdg mode is untouched — `line: 0`
there is still rejected (there is no 1-based source line 0 to anchor on).

Regression tests pin the full matrix: callgraph + line:0 runs the BFS and
is byte-identical to omitting line; pdg + line:0 still errors; positive
line on callgraph still errors.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(mcp): log swallowed best-effort query degradations at warn, not error

`logQueryError` is the shared handler for query failures that every caller
catches and degrades past with a safe fallback (the operation still returns
a result). It logged all of them at `logger.error` (level 50) — the same
severity as fatal failures — so a gracefully-handled degradation raised a
false alarm and drowned genuine errors. This surfaced as an ERROR-level log
firing during a passing unit test that intentionally injects a slice-callees
query failure to verify the degrade path.

Make the severity match reality:
  - benign missing optional table/label/column (a repo analyzed without
    processes/communities, or a pre-v3 PDG index lacking the `calleeIds`
    column — a query that fails on every pdg-downstream impact for such an
    index) → debug, the normal-configuration case.
  - any other swallowed failure → warn (handled degradation, still observable).
  - error is reserved for failures that actually abort an operation, which
    log directly rather than through this helper.

Also fix the sibling bm25/FTS fallback, which logged its swallowed
"FTS indexes may not exist" degradation at error while its own import-failure
fallback already used warn.

The slice-callees degradation test now captures the log and asserts it lands
at warn (40), not error (50), pinning the severity against regression.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(mcp): relax impact `line` schema minimum to 0 for adapter compatibility (#2279)

Strict MCP clients/agents validate against the advertised input schema and
reject a request before sending it. With `line` declaring `minimum: 1`, a
client that materializes the omitted optional `line` as `0` rejects a
perfectly valid callgraph impact call client-side — so the backend tolerance
added in the previous commit never gets a chance to run.

Lower the advertised `line.minimum` to 0 and document that 0 (or omission)
means "no statement anchor" while mode:'pdg' still requires a positive line.
The advertised schema is advisory (the backend self-validates and is the real
gate), so this cannot loosen any enforced contract — it only stops strict
clients from pre-rejecting `line: 0`. Negative lines are still rejected at the
client boundary.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(review): apply autofix feedback

Code-review autofix pass on the #2279 branch:
- Replace a newly-introduced `mode as any` cast in the #2279 it.each with the
  narrow `mode as 'callgraph' | undefined` (strict-typing-no-any).
- Add a degradation test for the new logQueryError benign-missing-table → debug
  branch (asserts no warn/error record surfaces, i.e. it routed to debug).
- Pin the bm25/FTS error→warn severity change with a _captureLogger assertion
  in the existing #1489 test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(mcp): make swallowed-failure callers surface degradation; narrow benign-error match (#2283)

Tri-review (#2283) found the `error → {debug|warn}` rework reduced telemetry
for `logQueryError` callers that do NOT degrade safely, while the docstring
over-claimed "every caller degrades to a safe fallback". Address the substance
rather than only the log level:

- rename apply-edit: track failed writes and return status:'partial' with
  `failed_files` instead of reporting `status:'success'` when a write was
  swallowed. A partial rename is no longer indistinguishable from a clean one.
- detect_changes: a swallowed symbol/process query failure now sets
  `partial:true` (rendered by the existing eval-server partial path) so the
  pre-commit safety gate can't return a false-clean `risk_level:'low'` no-op.
- isBenignMissingTableError: scope the `not (defined|found)` arm to a schema
  object (table/label/rel/column/property), mirroring lbug-adapter's
  isMissingColumnError. An unscoped "not found" matched operation failures like
  `rg: not found` / `Symbol not found` and silently demoted them to debug.
- logQueryError docstring: state the contract honestly — level reflects
  telemetry severity, and mutating/safety-critical callers MUST also surface a
  result-level degradation signal; `warn` alone is not a substitute.
- pdg dispatch: pass the normalized `effectiveLine` (not raw params.line) so
  the validation gate and engine share one source of truth (identity today).

Tests:
- _captureLogger(level?) lets tests capture below info; the benign-missing-table
  test now asserts the record IS emitted at debug (20), not merely absent —
  no longer a vacuous pass if the call were deleted.
- new: a non-schema "not found" failure logs at warn (regex-narrowing guard);
  rename write-failure degrades to status:'partial'+failed_files; line:-1 on
  the callgraph path still errors (line:0 coercion is narrow); typed the
  it.each tuple to drop a `mode as` cast.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(mcp): fix impact `line` description contradiction for whole-symbol pdg (#2283)

The new `line` schema description said "mode:'pdg' requires a positive line",
which contradicted the top-level impact description ("Without 'line', pdg
returns whole-symbol inter-procedural reach plus local whole-symbol PDG
diagnostics"). A pdg call without a line is a valid (degraded whole-symbol)
call, not an error — the old wording could push an agent to avoid valid no-line
pdg calls or synthesize line:0 (which then hard-errors).

Reword to: omit line for whole-symbol pdg; a positive line anchors a statement
slice; literal 0 is tolerated only as an omitted-line compatibility sentinel on
the callgraph path and is rejected for mode:'pdg'. Update the schema test to
pin the new, non-contradictory wording and assert "requires a positive line" is
gone.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 20:35:49 +01:00
Gergő Magyar 698f5efc82 feat(group): resolve inline HTTP provider handlers via call-site line (#2276) (#2282)
* feat(group): resolve Go inline provider handlers via line containment (#2276)

Widen the Go HandleFunc + framework-route handler capture to match
func literals and emit name:null + call-site line for them, so an
inline handler resolves to its containing/closure symbol instead of
file-level. Named identifier handlers keep resolving by name.

* feat(group): resolve Laravel closure provider handlers via line containment (#2276)

Capture the Laravel route handler argument; a closure (anonymous
function or arrow fn) now emits name:null + the registration line so it
resolves to its containing symbol (service-provider boot, controller
method) by containment. Named-controller routes keep the 'route' label.
File-scope closures stay file-level (PHP closures not yet indexed).

* feat(group): wire call-site line on FastAPI provider emits (#2276)

Set line on the FastAPI @app/@router provider detections (already
name:null) so the source-scan fallback resolves the decorated handler
by line-span containment. Best-effort: FastAPI routes are graph-backed
and the function span starts at def, so this lands the single-decorator
case. Flask add_url_rule already carried line.

* feat(group): wire call-site line on Kotlin/Java Spring provider emits (#2276)

Add line to the Kotlin and Java Spring @*Mapping provider detections
for parity with the consumer emits and a future inline DSL. Inert for
current resolution: a named Spring controller method resolves by name
and never falls through to line-span containment.

* fix(review): apply autofix feedback

Pin two documented limitations with tests: a file-scope Laravel closure
and a multi-decorator FastAPI handler both degrade to file-level rather
than mis-attributing (#2276 ce-code-review autofix).

* test(group): lock named gin framework-route resolves by name not registrar (#2276)

Reviewer verified named Go handlers still resolve by name across the
widened queries; the HandleFunc path was already pinned, this adds the
framework-route (gin/echo) path with a DB + enclosing registrar whose
span covers the registration line, proving the emitted line never
diverts a named provider to its registrar via containment.

* test(group): end-to-end inline Go provider resolution against real LadybugDB (#2276)

Closes the validation gap that all prior coverage mocked CONTAINING_QUERY:
runs the real pipeline over a Go file with an inline http.HandleFunc
func-literal handler, persists into a real LadybugDB, and runs the
production HttpRouteExtractor against the real executor — proving the
emitted call-site line lands inside main()'s real 0-based span and yields
source_scan_resolved, not the file-level fallback.

* fix(test): use fs.mkdtemp to satisfy CodeQL insecure-temporary-file gate (#2276)

The new integration test created its temp base via a predictable
os.tmpdir()+name join, which CodeQL flags as js/insecure-temporary-file
(1 high). Switch to fs.mkdtemp for an atomic, randomly-named base dir.

* fix(group): anchor Go provider @handler to the trailing argument (#2276)

The widened framework-route and HandleFunc handler captures
(`[(identifier) (func_literal)] @handler`) were unanchored, so a variadic
middleware route `r.GET("/x", mw, func(){})` produced two provider
detections — one for the middleware identifier and one for the closure.
The contractId-only merge then kept the middleware detection and
mis-attributed the route to it (and the pre-existing `mw, namedHandler`
shape had the same defect), silently neutralizing the inline-handler
containment resolution from #2276.

Add a trailing tree-sitter anchor (`@handler .`) so the handler binds the
LAST argument of the call, leaving middleware args before it unconstrained.
Verified against tree-sitter-go: the multi-arg shapes now yield exactly one
detection (the real handler) while every 2-arg case is unchanged. Adds two
regression tests pinning that a middleware + inline closure resolves to its
containing function and a middleware + named handler resolves by name.

* test(group): cover FastAPI @router inline-handler containment (#2276)

The @router/APIRouter provider emit gained a call-site `line` in #2276 but
only the @app path was tested; the existing @router tests call
`extract(null, …)` so the resolver/containment path never ran for @router.
Add two tests mirroring the @app cases: a single-decorator @router handler
resolves to its function via source_scan_resolved (which fails if `line` is
dropped), and a multi-decorator one degrades to file-level.

* fix(group): treat synthetic 'route' label as anonymous in cross-trace (#2276)

After #2276 an unresolved file-scope Laravel closure emits name:null, so its
persisted symbolName falls back to 'handler' — which providerLabel already
anonymizes to '<contractId handler>'. But an unresolved named-controller
route still carries the synthetic 'route' placeholder, which the sentinel
did NOT cover, so group_trace/group_cross_impact rendered it as the literal
'route' while equivalent closures showed '<... handler>'.

'route' is only ever the synthetic Laravel placeholder (php.ts), never a
resolved handler name, so add it to the unresolved-generic sentinel set
alongside 'handler'/'fetch'. The resolved branch is untouched, so a real
symbol genuinely named 'route' still displays its name. Adds a cross-trace
test pinning the anonymized label.

* fix(group): gate Spring provider line on a present method name (#2276)

The Java/Kotlin Spring @*Mapping provider emits set `line` unconditionally
while the method name is typed string|null. The 'a named provider never
reaches containment' guarantee held only because the grammar always captures
a method name — the type did not enforce it. A (grammar-impossible) null name
would emit name:null + line and resolve by containment to the enclosing class
body instead of staying file-level.

Emit `line` only when the method name is truthy, so a nameless provider
degrades to file-level (the safe no-mis-attribution outcome). Behavior is
unchanged for every real Spring route (name is always present), but the
inertness is now enforced rather than incidental.
2026-06-23 17:51:11 +01:00
Gergő Magyar 49ffd8e316 feat(group): resolve cross-file named HTTP handlers (#2275) (#2277)
* feat(group): resolve cross-file named HTTP handlers via unique repo-wide lookup

U1 of #2275. When a provider's named handler is defined in a file other than its
route registration (e.g. router.get('/x', listUsers) with listUsers imported),
the registration file's symbols don't contain it, so resolution fell back to the
file-level boundary. Add a repo-wide name query (RESOLVE_BY_NAME_QUERY, the
label-union pattern from manifest-extractor) consulted only after the file-scoped
lookup misses, and honored ONLY when exactly one Function/Method/CodeElement
carries that name (zero/many → keep the file fallback, no wrong-symbol
attribution). Provider-only, cached by name. 4 unit tests; 743 group tests pass.

* test(bench): cross-file named handler scenario (end-to-end proof of #2275)

U2 of #2275. Adds a fifth bench scenario: a backend route whose handler
(listUsers) is imported from another file than its registration, with a frontend
consumer. Asserts the provider resolves to the handler via the repo-wide unique
name lookup (sym=listUsers, uid set) and that the cross-repo trace is symbol-
precise (no file-level fallback). verify.mjs now 12/12 on the real pipeline.

* fix(review): apply autofix feedback

ce-code-review (autofix) — no correctness/security findings; applied test-coverage
+ robustness fixes: repo-wide query throw -> empty (no exception); by-name lookup
cache fires once across same-named handlers; consumers never consult the repo-wide
lookup; same-file-wins now asserts the global path is bypassed; bench provider find
scoped by contractId; clarified the uniqueness-guard comment. 167 extractor tests.

* fix(group): tri-review fixes for cross-file handler resolution

Two-engine PR tri-review (Claude swarm+ce, Codex gpt-5.5 swarm+ce+adversarial)
on #2277. Correctness/security clean (injection refuted, bind-param). Fixes:

- Named-provider wrapper-attach (Codex swarm P1 + Claude ce-adversarial,
  cross-engine): a named handler that fails both name lookups no longer falls
  through to line-span containment, which attached the route to the enclosing
  registrar (e.g. a setupRoutes() wrapper) instead of leaving it empty.
  Containment now applies only to consumers and inline-arrow providers.
- CodeElement/ORM empty-file nodes (Claude ce-adversarial reproduced +
  ce-maintainability): RESOLVE_BY_NAME_QUERY gains 'AND n.filePath <> ""' so a
  handler name colliding with a synthetic ORM model node (orm.ts emits
  filePath:'') neither resolves to an edge-less node nor inflates the uniqueness
  count and masks the real handler; + a defensive empty-filePath guard in
  resolveSymbolByNameUnique. Added LIMIT 2 (Codex swarm P3 + ce-maintainability)
  to bound homonym materialization (count guard stays exact).
- Documented the aliased-import limitation (Codex adversarial): the route-site
  identifier is the local alias, fix deferred to #2275 import narrowing.
- README expected verdict 9/9 -> 12/12 (Codex swarm+ce P3).

Tests: +3 (wrapper-no-attach, empty-filePath reject, empty-registration-file
resolves) covering the cross-engine gaps. 170 extractor / 748 group+integration
pass; bench 12/12 end-to-end.

* feat(group): import-pinned handler resolution (fixes deferred alias case)

Resolves the tri-review's deferred item: cross-file named handlers are now pinned
to their import's target module instead of resolved by name alone, so aliases and
names that collide with a local symbol resolve correctly.

- node.ts builds a local-binding -> {declared name, module} map from the file's
  named imports; the express handler emits the DECLARED name + a handlerImport
  {name, module} (HttpDetection gains the optional field).
- resolveDetectionSymbol gains an imported-handler rung: resolveImportedSymbol
  pins to the import's target file via RESOLVE_IN_MODULE_QUERY
  (n.name= AND filePath STARTS WITH the resolved module path), unique-match
  only. An imported handler never uses file-scoped lookup (it is defined
  elsewhere); on a module miss it falls back to a unique repo-wide name match on
  the DECLARED name, then null. Relative imports only; bare/non-relative imports
  keep the repo-wide fallback. Cached by (module-prefix, name).
- Closes the Codex-adversarial alias finding: import { listUsers as handleUsers }
  + an unrelated handleUsers no longer mis-resolves — the route resolves to the
  imported listUsers in its module, and the alias is never looked up.
- Shared toResolvedSymbol helper (dedups the row->symbol + empty-filePath guard).

Tests: alias-resolves-to-declared-name + module-pin-resolves-ambiguous-name unit
tests; same-file-wins reworked to a genuinely LOCAL handler. Bench scenario 6
(aliased import with a decoy) proves it end-to-end. 172 extractor / 751
group+integration pass; bench 14/14.

* feat(group): import-pinned resolution for Python aliased handlers

Extends the JS/TS import-pinning to Python. The Python analog of express
router.get(path, handler) is Flask's imperative add_url_rule(view_func=...),
whose view is often an imported (aliased) symbol.

- New Flask add_url_rule provider pattern (path + view_func handler + methods;
  default GET, methods=[...] honored). High Flask-specificity keeps false
  positives low — unlike bare path()/Route(), which the plugin deliberately
  leaves to graph Route nodes.
- buildPythonImportMap resolves 'from .mod import name as alias' (and plain
  'from mod import name') to the declared name + raw module spec.
- resolveModuleBase generalized to two relative-import dialects: path-style
  (JS './h/users') and dotted (Python '.handlers.users', '..pkg.users' — leading
  dots are package levels). Bare/absolute imports keep the repo-wide fallback.
- Django stays graph-resolved (handlerSymbolId); FastAPI/Flask decorators stay
  same-file (decorated function). This only adds the imperative imported-view
  case Python lacked.

Tests: Flask aliased add_url_rule unit test (relative dotted module pinned, alias
never queried) + bench scenario 7 (end-to-end, 16/16). 173 extractor / 752
group+integration pass.
2026-06-23 12:12:49 +01:00
d27fd11c4b fix(lang-kotlin): support fun interface extraction via tree-sitter-kotlin re-vendor (#2271)
* fix(lang-kotlin): support `fun interface` extraction via tree-sitter-kotlin re-vendor

Vendored tree-sitter-kotlin@0.3.8 (fwcd) parsed `fun interface Foo` as an
ERROR node and dropped the declaration plus its abstract method, so functional
(SAM) interfaces were never extracted. The fix landed upstream in
fwcd/tree-sitter-kotlin#169 (closes #87), merged to main 2025-04-25, but is not
in any npm release (latest tag 0.3.8; main is the unreleased 0.4.0).

Re-vendor the grammar from the unreleased fwcd main commit c8ac3d26:
- refresh src/{parser.c,scanner.c,node-types.json,tree_sitter/*.h} and
  bindings/node/index.js; bump the vendor version 0.3.8 -> 0.4.0; record the
  pinned SHA + rationale in _vendoredBy and the vendor README.
- switch the prebuild workflow's kotlin registry kind 'npm' -> 'vendored' (the
  fix is unreleased on npm, so prebuilds must build from the vendored C source,
  like swift/dart/proto).
- add a hold to .github/vendored-grammars.json so the weekly auto-update
  monitor does not strict-inequality-revert the pin to the broken npm 0.3.8
  (isNewer compares 0.3.8 != 0.4.0).
- add 3 regression tests + a fixture asserting fun interfaces extract as
  Interface nodes with their abstract methods, and that plain-interface
  heritage still resolves.

Existing KOTLIN_QUERIES need no change: the new grammar models `fun interface`
as a class_declaration with an "interface" keyword child (plus an extra "fun"
modifier child), which the existing interface rule already matches. Full Kotlin
suite green against the new grammar (300 unit/cfg/resolver + 233 integration).

NOTE: prebuilds/ are intentionally not in this commit. The version bump
auto-triggers .github/workflows/build-tree-sitter-prebuilds.yml, which
regenerates all 6 platform binaries from the vendored source in a separate PR.
Until that lands, CI loads the committed 0.3.8 prebuild, so the new kotlin
tests are red and the grammar change is inert at runtime. Merge the prebuild PR
first or together.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(ci): count kotlin's vendored hold as a 0.25-readiness blocker

The kotlin `hold` added in the previous commit makes the tree-sitter
upgrade-readiness report count it as a blocker — the report treats every
held vendored grammar as frozen below a runtime upgrade (same as the
intentionally-pinned tree-sitter-cpp and the ABI-held tree-sitter-c),
"in-range ABI or not". So the report's blocker count goes 2 -> 3.

Update the hardcoded count in
test_issue_update_summary_regex_matches_current_report (and the
_render_report docstring) accordingly — exactly as that test instructs:
"if a grammar is added/removed or a pin/hold changes, update the expected
counts". kotlin's ABI (14) is in range; the hold is what flags it, with the
reason recorded in .github/vendored-grammars.json.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(ci): refresh kotlin baselines for the grammar bump

Two committed baselines pinned the pre-bump kotlin state and broke when the
grammar was re-vendored (0.3.8 -> 0.4.0):

- cli-commands.test.ts pinned the vendored kotlin package version at 0.3.8 ->
  update to 0.4.0.
- bench/scope-capture/baselines.json: the new kotlin-fun-interface fixture joins
  the lang-resolution/kotlin-* corpus AND the new grammar parses `fun interface`
  as a class_declaration (not an ERROR node), so the capture fingerprint drifts.
  Rebaselined to the NEW grammar's fingerprint (verified by building the vendored
  parser.c against tree-sitter@0.21.1 and running measure.mjs --check); scaling
  ~0.83 (linear).

Like the fun-interface integration tests, the scope-capture --check passes only
once the regenerated prebuilds land; until then CI loads the committed 0.3.8
binary, so it stays red.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(prebuilds): rebuild + commit grammar prebuilds into the PR on vendored-source change

build-tree-sitter-prebuilds.yml previously rebuilt a grammar's native prebuilds
only when its package.json VERSION bumped, and delivered them via a separate bot
PR. Now any change to the vendored grammar source re-cuts the prebuilds and they
ride into the same PR.

- Trigger on any build-affecting change under gitnexus/vendor/tree-sitter-*/**
  (parser.c, grammar.js, binding.gyp, scanner, bindings), not just version bumps.
  The prebuilds/ subtree is negated in the paths filter AND excluded from the
  guard's source diff, so the bot's own prebuild commit can never retrigger the
  workflow (no build -> commit -> build loop).
- The guard builds a grammar when its recorded version changed OR its vendored
  source changed vs the PR base.
- Same-repo PRs get the rebuilt prebuilds committed straight onto their own head
  branch (included in the SAME PR) via a non-force push that only adds a commit
  on top of head. Manual dispatch still opens a fresh chore/ PR; fork PRs stay
  artifacts-only (a bot cannot push into a fork branch).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(prebuilds): deliver rebuilt prebuilds to fork PRs via a trusted workflow_run stage

A fork PR's producer run has a read-only token and no secrets, so it can build and
validate the prebuilds but can't commit them. Add the safe two-stage handoff that
mirrors the pr-autofix producer/publish split.

- build-tree-sitter-prebuilds.yml (untrusted producer): on a fork PR, upload a
  pr-meta artifact (schema, pr_number, head_sha, head_ref, head_repo, base_repo)
  alongside the prebuild artifacts. Values flow through env + jq, never
  interpolated into a shell.
- commit-fork-prebuilds.yml (trusted, workflow_run): downloads ONLY the artifacts
  (never executes fork code — it checks out the pinned HEAD SHA solely to add
  files), allowlist-validates every metadata field, cross-checks identity against
  the workflow_run authority (head_sha / head_repo / pr_number, via
  commits/{sha}/pulls for forks), then pushes the prebuilds onto the fork head
  branch with --force-with-lease + http.extraheader auth. No PAT: this works when
  the contributor left "Allow edits by maintainers" on; on push failure it posts a
  sticky comment telling them to enable it or commit the downloaded artifacts.

zizmor: allowlist commit-fork-prebuilds.yml's workflow_run dangerous-trigger with
the documented mitigation, matching the existing ci-report / pr-autofix entries.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(vendor): rebuild tree-sitter-kotlin prebuilds for the re-vendored fun-interface grammar

The fun-interface re-vendor changed vendor/tree-sitter-kotlin source but left main's
old (0.3.8) prebuilds in place, so all 6 platform binaries were stale relative to the
new parser. Replace them with the freshly cross-built + ABI-validated binaries from
build-tree-sitter-prebuilds run 28010841458 — each .node was require()-loaded and
parsed a snippet on its target platform-arch before upload.

This is the manual equivalent of the commit-fork-prebuilds.yml delivery, which can't
run for this fork PR until it lands on main.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(lang-kotlin): read extension-function receiverType from the re-vendored grammar's `receiver` field

The fun-interface re-vendor changed the kotlin AST: an extension function's
receiver is now a `receiver_type` exposed via a named `receiver` field, where the
old grammar emitted a bare user_type before the name. extractReceiverType only
matched the old shape, so receiverType came back null
(method-extraction.test.ts > Kotlin MethodExtractor > extracts receiverType).
Prefer the `receiver` field (unwrapping it), and keep the old child-scan — now
also recognizing `receiver_type` — as a fallback.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-06-23 10:01:28 +01:00
Gergő Magyar 1a03c8527a feat(group): cross-repo call trace using PDG (#2269)
* refactor(group): extract shared resolveBridgeNeighbors from cross-impact

Lift the uid-filtered consumer<->provider ContractLink join (direction +
queryBridge + row normalization + confidence sort) out of runGroupImpact's
inline Phase-2 block into an exported resolveBridgeNeighbors helper. Behavior
is unchanged for impact; the helper becomes the single shared bridge join so
the upcoming cross-repo trace path never forks its own copy of the neighbor
Cypher. Empty uid sets short-circuit without a DB round-trip.

Adds direct coverage (real bridge via writeBridge/openBridgeDbReadOnly) for
both directions plus the empty-set and unknown-uid edges.

* feat(group): cross-repo trace stitching (groupTrace + runGroupTrace)

Add GroupService.groupTrace and the pure runGroupTrace engine that stitches
per-repo CALLS/HAS_METHOD trace segments across one ContractLink boundary in
the group bridge:

  from --(local trace)--> consumer --(ContractLink)--> provider --(local trace)--> to

- Resolves from/to across all members (symbol node id == bridge symbolUid);
  same-repo endpoints delegate to a single local trace with no crossing.
- Single boundary crossing (MAX_SUPPORTED_CROSS_DEPTH); deeper crossDepth is
  clamped with a note, mirroring cross-impact.
- Discriminated GroupTraceResult union (ok|not_found|ambiguous|error) with
  per-hop repo tags, a typed crossings[] entry, and centralized degraded-state
  note constants (TRACE_NOTES). No .
- Trace-specific pair query (keeps BOTH crossing endpoints) lives in this
  module; the uid-filtered neighbor join (resolveBridgeNeighbors) is reused
  where it fits. ensureBridgeReady exported for reuse.
- New GroupToolPort methods (trace/resolveSymbol/pdgFlows) are optional so
  existing port mocks keep type-checking; runGroupTrace guards on presence.

PDG enrichment is wired as an opt-in hook (enrichSegment) — the port method is
stubbed until U4. Covered by unit tests over a real bridge + mocked port.

* feat(group): route trace tool to groupTrace on @group syntax

Wire the cross-repo trace through the existing @group dispatch:
- callTool routes trace with an @-prefixed repo to callToolAtGroupRepo, which
  forwards from/to/uid/file/maxDepth/includeTests plus the experimental
  pdg/crossDepth flags to GroupService.groupTrace. Member path in @group/path
  is advisory for trace (resolution is whole-group).
- Port gains trace/resolveSymbol/pdgFlows adapters. resolveSymbolForGroup wraps
  the shared resolveSymbolCandidates so groupTrace can locate the member repo
  and recover each endpoint node id (== bridge symbolUid). pdgFlowsForGroup is
  a degraded stub here (call-level only); U4 implements the REACHING_DEF walk.
- trace tool schema documents the @group entry point, pdg, and crossDepth.

Single-repo trace is untouched. Covered by dispatch-routing tests (@group ->
groupTrace, non-group stays local) and tool-schema assertions.

* feat(group): opt-in PDG data-flow enrichment for cross-repo trace

Implement _pdgFlowsForGroupImpl: the real REACHING_DEF anchor walk that backs
the port pdgFlows adapter (replacing the U3 call-level stub). When pdg:true and
the segment repo has a flows PDG layer, the boundary-adjacent segments carry
their intra-procedural def->use hops:

- Anchors by the boundary symbol UID (precise; avoids the by-name ambiguity the
  resolveBlockAnchor path can hit), then reuses the same span-anchored,
  bind-param-only flows query as pdg_query (BasicBlock id-prefix + [start+1,
  end+1] line window; no rel-property index, so the anchor IS the bound).
- Stays intra-procedural: data flow never crosses the repo boundary.
- pdgStampForMode probe: false -> available:false (degrade with note); the
  trace stays ok. Any query failure is swallowed (enrichment is auxiliary).

Covered by runGroupTrace enrichment tests: dataFlow attached on opt-in,
degraded note when no layer, and no pdgFlows call when pdg is omitted.

* test(group): evaluation-first cross-repo trace e2e (two real indexes)

End-to-end gate for the cross-repo trace: stands up two real LadybugDB indexes
(consumer 'frontend' + provider 'backend'), a real ContractLink bridge, and a
real LocalBackend with both repos registered, then drives the public
callTool('trace', { repo: '@grp', pdg: true }) and asserts:
  - the stitched checkout -> callUsers -(CONTRACT_LINK)-> handleUsers -> getUsers
    path, each hop tagged with its member repo
  - real REACHING_DEF data-flow enrichment of the consumer segment (userId)
  - a degraded 'No PDG layer in app/backend' note (provider has no PDG layer)
  - single-repo trace against one member is unchanged (no crossings)

Hand-persists the minimal real graph (deterministic; a full two-repo analyze is
heavier than this gate needs) and exercises real Cypher across
resolveSymbolCandidates, _traceImpl, the bridge pair query, and
_pdgFlowsForGroupImpl. Windows-skipped (describeReopen) and registered in the
cross-platform native-lbug set.

Scoped to a single @group call: opening bridge.lbug read-only a SECOND time in
one process currently fails (shared bridge open/close lifecycle, also affects
impact @group) — the pdg-omitted/clamp variants are unit-covered.

* docs(group): document cross-repo trace + PDG enrichment

ARCHITECTURE.md: trace is now group-aware; describe the @group cross-repo
stitch over a single ContractLink boundary (CONTRACT_LINK hop, crossings[],
crossDepth clamp), the opt-in experimental PDG REACHING_DEF enrichment of
boundary-adjacent segments, the symbolUid-grain join between the two stores,
and the deferred full cross-program (SDG-like) data flow. PIPELINE.md: add the
cross-trace consumer of the bridge with its pair-query rationale.

Does not touch gitnexus/CHANGELOG.md (release-owned).

* fix(review): apply autofix feedback

Apply safe_auto findings from ce-code-review (run 20260622-094243):
- local-backend.ts: drop (r: any) in _pdgFlowsForGroupImpl row map; coerce
  hop line via Number() so a nullish LadybugDB cell can't surface NaN.
- tools.ts: advertise the forwarded  param in the trace schema and add
  crossDepth maximum:10 (schema now matches what groupTrace reads).
- cross-trace.ts: parallelize per-member resolveSymbol/resolveRepo with
  order-preserving Promise.all (matches groupContext/groupQuery); add a note
  when pdg:true is passed to a same-repo trace (PDG only enriches at a
  cross-repo boundary).
- tests: remove  / tighten  (no-any rule).

Residual gated_auto/manual findings (unbounded crossing query + loop,
whole-file PDG widening on absent span, error-vs-no_path masking, top-level
try/catch parity, helper dedupe, branch-coverage gaps) are recorded in the run
artifact for the PR body.

* fix(group): skip CHECKPOINT on read-only bridge close so it can reopen

Root cause of the in-process bridge.lbug reopen failure (which broke repeated
@group impact/trace calls in a long-lived MCP server): closeBridgeDb issued
CHECKPOINT on EVERY handle, including read-only ones. A CHECKPOINT on a
read-only connection has nothing to flush but leaves a WAL/shadow lock artifact
that makes the next read-only open of the same path fail (openBridgeDbReadOnly
returns null -> 'Could not open bridge.lbug read-only'). Reproduced: open ->
query -> closeBridgeDb -> open again returned null only when the close ran
CHECKPOINT; a non-checkpoint close reopened fine, and the raw native
open/close cycle was never the problem.

Fix: tag read-only handles (BridgeHandle._readOnly, set by openBridgeDbReadOnly)
and skip CHECKPOINT for them in closeBridgeDb. Writable handles are unchanged
(they still flush before close). This is the shared bridge-db close path, so
impact @group benefits identically.

- Regression test in bridge-db.test.ts: open/query/close/open/query/open in one
  process now succeeds.
- Re-enabled the second @group call in cross-trace-e2e.test.ts (was scoped to a
  single call for this very limitation).

* fix(group): bring bridge-db close to parity with the core adapter safeClose

The bridge open/close cycle was less robust than the main graph DB's: closeBridgeDb
closed the connection/database but skipped the post-close steps the core adapter's
safeClose performs, so a rapid in-process reopen could race the OS handle release
(Windows) or an orphaned WAL sidecar. That gap is why the close-then-reopen tests
had to skip Windows.

closeBridgeDb now mirrors safeClose after closing the handle:
- waitForWindowsHandleRelease(dbPath): probe the file (+ .wal) until the residual
  Windows lock clears, so the next open does not race (warns if the budget is
  exhausted, matching the core adapter).
- finalizeLbugSidecarsAfterClose(dbPath): quarantine an orphaned WAL (shadow
  missing) so the next open replays a consistent file.

Both helpers are the same ones safeClose uses (Windows-proven via the core adapter
CI), and the bridge read open already retries transient locks. Combined with the
read-only CHECKPOINT skip, the bridge reopen is now robust on every platform, so
the close-then-reopen tests run on all platforms (Windows CI exercises them via the
cross-platform subset). No write-path behavior change; Linux/macOS unaffected.

* fix(group): bound cross-repo crossing fan-out (LIMIT + segment memoization)

Address the top review residual: the bridge crossing query was unbounded and the
crossing-selection loop could run an O(2*N) sequential trace-BFS over every
ContractLink between a repo pair.

- CY_CROSSINGS_BETWEEN now ORDERs BY confidence DESC and LIMITs to
  MAX_CROSSINGS_TO_TRY + 1; listCrossingsBetween slices to the cap and reports
  truncation. Exceeding the cap surfaces a note (no silent truncation), keeping
  the highest-confidence crossings. Aligns with the repo's anchored+LIMIT-bounded
  query discipline (LadybugDB has no rel-property index).
- The home-repo segment (from -> consumer) depends only on the consumer uid and
  the target-repo segment (provider -> to) only on the provider uid, so each is
  memoized by that uid. Many crossings sharing a consumer/provider (one client
  call linked to several providers) now cost one trace per distinct endpoint
  instead of one per crossing. A consumer whose segment already failed is skipped
  for every later crossing that shares it.

Test: two links sharing a consumer (first provider unreachable, second reachable)
assert the from->consumer segment is traced exactly once and the second crossing
wins.

* fix(group): restore Windows skip for bridge reopen tests; drop ineffective close-side probe

The previous commit flipped the bridge close-then-reopen tests to run on Windows,
betting that a close-side waitForWindowsHandleRelease + finalizeLbugSidecarsAfterClose
probe (mirroring the core adapter safeClose) would make the in-process reopen work
there. Windows CI proved otherwise: 4 writeBridge->openBridgeDbReadOnly tests fail
('expected null not to be null' — the read open returns null). The writable-close ->
read-open handoff plus writeBridge's atomic sidecar rename does not release the OS
file handle before the read open races, and the existing open-side LBUG_OPEN_RETRY
only retries lock-pattern errors, not the post-rename sidecar database-id mismatch.
macOS passes; the core adapter's own reopen also passes — this is bridge+Windows
specific.

- Revert itLbugReopen to the Windows skip (the pre-existing, correct state).
- Remove the close-side probe + finalize from closeBridgeDb: it did NOT close the
  Windows gap, and reviewers flagged it for hot-path latency (finalize ran on every
  close, all platforms) and safeClose duplication.
- KEEP the load-bearing fix — skipping CHECKPOINT on read-only handles — which fixed
  the reproduced Linux/macOS in-process reopen artifact (the real bug).

Net: Linux/macOS repeated @group impact/trace works in-process; Windows in-process
bridge reopen remains a documented limitation (unchanged from before this PR).

* fix(group): surface degraded members + cap truncation; honest crossDepth schema

Address the cross-engine-corroborated tri-review findings (Codex + Claude):
- resolveAcrossMembers / runGroupTrace now track member repos that could NOT be
  queried (resolveRepo or resolveSymbol threw) and, when the result is not_found,
  attach a degraded-member note. A transient/corrupt member DB is no longer
  silently reported as a clean 'symbol absent' not_found. (Codex B1+B3 + ce-reliability.)
- The cross-repo not_found now carries a programmatic truncated:true flag (and a
  clearer suggestion) when the MAX_CROSSINGS_TO_TRY cap was hit, so a consumer can
  distinguish 'no path' from 'cap may have hidden a connecting ContractLink'.
  (Codex B3 + ce-adversarial + ce-api-contract.)
- trace tool schema: crossDepth maximum 10 -> 1 to match the implementation's
  single-hop clamp (the schema previously advertised an unsupported 2-10 range).
  (ce-api-contract, conf 100.)

Test: a member whose resolveSymbol throws yields not_found WITH a degraded note
naming the unreachable repo (if-free responder map).

* docs(group): clarify trace @group/memberPath is advisory (resolves all members)

Tri-review (Codex ce, conf 100) caught a doc/impl inconsistency: ARCHITECTURE.md
lumped trace with query/context/impact as honoring @group/memberPath member
scoping, but cross-repo trace resolves from/to across ALL members (the member
path is advisory). Clarify the behavior and point to from_uid/to_uid for
disambiguating same-named symbols across members.

* feat(group): file-level boundary fallback so cross-repo trace works on HTTP contracts

Benchmark (bench/cross-repo-trace/) running the REAL pipeline (runFullAnalysis
--pdg -> real syncGroup -> trace @group) found that cross-repo trace returned
not_found for real HTTP links even though sync built the correct ContractLinks:
HTTP (and other source-scan) contracts hardcode symbolUid:'' (http-route-extractor),
and both cross-trace AND cross-impact join crossings by Contract.symbolUid, which
never matches an empty uid. (Pre-existing — impact @group has the same gap.)

Fix: when a crossing's symbolUid is empty, fall back to the contract's FILE — if
the user's from/to resolves into the contract file, that endpoint anchors the
boundary. CY_CROSSINGS_BETWEEN now returns consumer/provider filePath; a crossing
is kept if it can be anchored by uid OR file on each side; a fileBoundaryFallback
note flags that the boundary is file-level, not symbol-precise. This makes the
common 'trace from=<calling fn> to=<handler fn>' case work end-to-end (verified:
fetchUsers -> listUsers stitches with a CONTRACT_LINK hop + PDG enrichment, 2/2).

Limits (documented in the bench README + the note): anonymous handlers have no
named target; when several contracts share files the file fallback may attach the
wrong contractId to a correct path. The proper upstream fix is to populate
symbolUid in the HTTP extraction (benefits impact too) — the bench is its gate.

Adds a unit test pinning the empty-symbolUid file-fallback stitch.

* fix(group): resolve HTTP contract symbolUid by containment (fixes cross-repo trace + impact)

Addresses the root cause behind the cross-repo trace file-fallback: HTTP
contracts hardcoded symbolUid:'' (http-route-extractor), so both cross-trace and
cross-impact — which join crossings on Contract.symbolUid — could not traverse
HTTP links. (Also found: the pre-existing graph-assisted resolution queried the
wrong edge, CONTAINS instead of DEFINES, so it never resolved a uid either.)

Now the extractor resolves each detection to a real symbol:
- HttpDetection carries the call-site line (node.ts sets it on every express/
  fetch/axios/jquery/nest detection; express also captures the handler arg).
- resolveDetectionSymbol resolves the named handler first, else the innermost
  Function/Method whose line span encloses the call (consumer = the function
  containing the fetch; provider = the named/inline handler), over the correct
  File-[DEFINES]->symbol edge. Base-tolerant (0- vs 1-based startLine).
- Wired into both source-scan and graph-assisted provider/consumer paths.

Verified end-to-end (bench/cross-repo-trace): all 4 contracts now carry real
uids, trace is symbol-precise (GET pair -> http::GET, POST -> http::POST, no
file-fallback note), and impact @group fans out (cross_repo_hits 0 -> 1). The
cross-trace file-level fallback remains as the secondary path for truly
anonymous handlers. Adds 2 containment unit tests; 738 group/integration pass.

Languages other than JS/TS still resolve providers by handler name; their
consumers fall through to the file fallback until their plugins set the line.

* fix(group): extend HTTP symbolUid containment to all languages + nested methods

Completes the symbolUid resolution across every bundled HTTP plugin: Python, Go,
PHP, Kotlin and Java now set the call-site line on their consumer (and Feign/
named) detections, so their HTTP contracts resolve to the containing function
the same way Node/TS already did.

Also generalizes the containment query: it now matches Function/Method/CodeElement
by filePath (UNION ALL) instead of File-[DEFINES]->symbol. The DEFINES edge only
reaches a file's TOP-LEVEL symbols, so methods nested in classes (Java/Kotlin —
File defines the class, the class defines the method) were invisible; matching by
filePath reaches them. Verified against a real index (LadybugDB supports the
UNION); JS/TS still fully symbol-precise (bench 2/2), 709 group tests pass.

Residual is now only the inherent case — a fully anonymous handler with no named
callee — which keeps the cross-trace file-level fallback.

* feat(group): destination trace — follow a consumer to an anonymous handler

Handles the one inherent residual: an anonymous route handler
(`router.get('/x', (req,res) => …)`) has no symbol node at all (the file holds
only a Const + PDG BasicBlocks), so it can never be named as a trace `to`.

Adds a DESTINATION TRACE: omit to/to_uid/to_file on an @group trace and
`trace from=<consumer>` follows the consumer's outgoing HTTP call across the
bridge and reports where it lands — by route + file:line, with a notes[] entry
flagging the handler as anonymous. Implemented as a new branch in runGroupTrace
(p.destination) backed by CY_CROSSINGS_FROM (all ContractLinks leaving the
consumer repo) + stitchToDestination; the provider endpoint is labelled
'<METHOD /path handler>' when its symbolName is a generic token/file basename.

The MCP routing already omitted an absent `to`, so only the schema docs changed.
parseTraceParams now treats a missing `to` as a destination trace instead of an
error. Verified end-to-end: anonymous fixture reports
'app/frontend:fetchUsers -> app/backend:<http::GET::/api/users handler>'; named
fixture lands at the real function. Adds 2 unit tests; 915 group tests pass.

* fix(group): tri-review fixes for cross-repo trace + symbolUid resolution

Two-engine tri-review (Claude swarm+ce + Codex GPT-5.5 swarm+ce+adversarial)
surfaced these; cross-engine-corroborated unless noted.

Correctness (P1, all four lanes): destination trace reported the WRONG endpoint
— an empty-uid consumer made trace(from->from) trivially succeed, so the highest-
confidence same-file crossing won regardless of which call `from` makes.
stitchToDestination now collects ALL connecting crossings, prefers symbol-precise
hits, and returns `ambiguous` (with candidates) when it cannot disambiguate.

Correctness (P1, Codex): resolveDetectionSymbol early-returned null when
d.line==null, blocking NAME resolution for named providers that set no line
(Spring/Go/etc.). Name resolution now runs first; only containment needs a line.

Correctness (P2): resolveContainingSymbol OR-ed `line` and `line-1`, which could
mis-pick a one-line sibling. It now probes the base-correct `line-1` first and
falls back to `line` only if nothing matches.

Correctness (Codex): anonymous Express handlers emitted name:'handler' and could
attach to an unrelated fn literally named `handler`. node.ts now emits name:null
for non-identifier handlers (containment-only).

Robustness: drop the first-symbol-in-file pickSymbolUid guess from the graph
consumer/provider paths (a wrong uid would win the contractId merge); remove the
dead CONTAINS_QUERY fallback (CONTAINS is File->Folder, never a symbol) + the now
-unused pickSymbolUid/handlerName; seed destination notes with degraded-member
notes so a successful trace still surfaces them; providerLabel takes providerUid
so a resolved fn named `handler` is not mislabeled anonymous, and only true file
basenames (known extensions) — not any dotted name — count as anonymous.

API contract: a single-repo trace with no `to` now returns an actionable error
(destination trace is @group-only) instead of "symbol 'undefined' not found".

Maintainability/tests: narrow asLocalTrace per-field (drop as-unknown-as); fix the
PR's lone as-any (vi.mocked); if-free e2e teardown; qualify the bench README.

Adds ambiguous-destination, anonymous-handler-no-false-name, and single-repo-no-to
tests; redirects graph mocks CONTAINS->UNION ALL. 918 group/integration pass.

* fix(group): carry degraded-member notes through SUCCESSFUL group traces

A reviewer (koriyoshi2041, PR #2269) correctly flagged that degraded-member
resolution was surfaced only on not_found, not on a successful ok result. Group
trace resolves names across ALL members, so an ok is 'unique among the members
we could query' — if a member that threw during resolveSymbol also holds from/to,
the real answer could be ambiguous. The destination path already seeded the note
(prior commit); this extends it to the same-repo and cross-repo success paths by
seeding the dispatch notes with degradedNotes([...fromRes.degraded, ...toRes.degraded]).

Adds a regression test: reg-be throws while a same-repo trace succeeds in reg-fe;
the ok result now carries the 'could not be queried' degraded note (app/backend).

* test(bench): cover all implemented cross-repo trace cases in one runner

Replace the single named-handler script with a self-contained verify.mjs that
generates each fixture inline and exercises every implemented end-to-end case
against the real analyze -> sync -> trace/impact pipeline, asserting PASS/FAIL
(exit non-zero on failure). 10 checks across 4 scenarios:
- named handlers: 4/4 symbolUid resolved; symbol-precise GET vs POST crossing
  selection; destination trace lands at the named handler.
- anonymous handler: empty symbolUid; destination trace reports it by route with
  the anonymous note.
- impact @group fan-out (cross_repo_hits >= 1).
- multi-language (Python Flask + requests): link built, cross-repo trace stitches,
  and the file-level boundary fallback is exercised when the provider has no uid.

Ambiguous-destination and degraded-member paths need synthetic inputs the real
analyzer cannot produce, so they stay in the unit suite (documented in the README
+ script header). Removes verify-named.mjs + fixtures-named/ (folded inline).

* test(group): pin destination degraded-success + precise-tier ambiguity

Adds the two regression guards koriyoshi2041 requested on PR #2269 after the
degraded-on-success fix:
- destination trace success with a degraded member: reg-fe resolves from and
  follows the link to an anonymous handler while reg-be throws; the ok result
  carries the anonymous endpoint AND the 'could not be queried' degraded note, so
  the no-to path stays aligned with explicit to traces.
- multiple PRECISE destination hits: one from reaches two consumers with resolved
  uids linked to different routes; the result is ambiguous (role: to) with both
  route candidates. Distinct from the existing file-level ambiguous test, this
  pins the stronger precise tier against a future change silently picking the
  highest-confidence destination.

Both already pass against current behavior; 716 group tests pass.
2026-06-23 07:54:13 +01:00
b16ec344f7 perf(group/http): skip source parse for graph-covered route files (#2138 Part 2) (#2265)
* feat(routes): resolve + persist handler symbol on Route nodes (#2138 part 2, WIP)

Part 2 groundwork for #2138: give the graph-assisted HTTP provider path the
handler symbol directly, so it no longer re-parses source to recover the
handler name. (The remaining parse-skip in extract() + a call-count benchmark
land in a follow-up commit.)

- ExtractedDecoratorRoute gains `handlerName`; the Spring extractor captures
  the decorated method's name (the method_declaration node is in hand).
- New `resolveRouteHandlerSymbols` (call-processor) resolves each route's
  handler to a real symbol UID, keyed by normalized route URL — Laravel
  framework routes (controller + method) and decorator routes (Spring/FastAPI)
  both reduce to `(filePath, name) -> nodeId`. Threaded through the parse phase
  onto `ParseOutput.routeHandlerSymbols`.
- routes phase stamps `Route.handlerSymbolId`; persisted end-to-end (schema +
  Route CSV row + getCopyQuery COPY columns), mirroring Part 1's `method`.
- HttpRouteExtractor: `HANDLES_ROUTE_QUERY` returns `handlerSymbolId`;
  `extractProvidersGraph` uses it as the authoritative symbol and SKIPS
  `getDetections()` for resolved rows (CONTAINS is a cheap graph lookup for the
  display name only — no tree-sitter parse). Fully backward compatible: an
  unresolved/old-index route with no `handlerSymbolId` keeps the source-scan
  fallback.
- Extracted `normalizeExtractedRoutePath` to `route-extractors/route-path.ts`
  (shared by routes phase + resolver without an import cycle).
- SCHEMA_BUMP 6->7 (ParseWorkerResult gained `handlerName`); regenerated the
  emit-persistence byte-identity baseline (route.csv header gained two columns).
- Tests: Spring pipeline asserts the Route node carries a handlerSymbolId
  resolving to the handler method; extractor fast-path test proves the handler
  resolves with zero source detections.

Refs #2138

* perf(group/http): skip source parse for graph-covered route files (#2138 Part 2)

Builds on the persisted Route.handlerSymbolId (U0–U3a). When a file's
HANDLES_ROUTE rows all resolve a handler symbol AND its language plugin
declares routeCoverage: 'complete' (Java/Python/PHP), the graph is
authoritative for that file's providers, so the source scan + tree-sitter
parse can be skipped — the scan would only re-discover routes the graph
already has. This is the measurable parse reduction #2167 could not show.

Consumer safety: routeCoverage: 'complete' asserts *provider* Route-node
completeness only. The scan() of those same languages also emits consumer
detections (RestTemplate/WebClient/OkHttp/Feign, Guzzle/Http::,
requests/httpx), and ingestion's FETCHES edges are JS/TS-only — so the
graph cannot back up server-side consumers. A provider-covered controller
that also calls out would otherwise lose its consumer contract. Guarded by
a cheap, parse-free text gate.

- types: HttpLanguagePlugin gains
    - routeCoverage?: 'complete' | 'partial' (default 'partial')
    - hasConsumerSignals?(content): false only when the raw source provably
      has no outbound-HTTP call this plugin detects (conservative).
- java/python/php: mark routeCoverage 'complete' + implement
  hasConsumerSignals with a token regex over their consumer idioms.
- http-route-extractor: run the graph provider pass first to build a
  coveredFiles set; then keep a file covered only when
  hasConsumerSignals(content) === false (read via readSafe, no parse).
  scanFiles = files not covered → drives collectProjectDetections + both
  source scans. Fail-open per file: any unresolved row, a 'partial'
  language, a positive consumer signal, a missing hook, or an unreadable
  file leaves the file in the scan set. The orchestrator names no
  languages — token knowledge stays in the plugins.

Net: pure-provider controllers skip the parse (the win); controllers that
also call out are still parsed (no consumer loss); partial-coverage
languages and graph-less runs are unchanged.

- test: route-parse-skip integration test spies the real parseSourceSafe to
  COUNT parses over a temp repo of Spring controllers with a mock DB —
  baseline (every file parsed), fully-covered (0 parses), mixed (unresolved
  file falls back, resolved stays skipped), and provider+consumer (a covered
  controller that also calls restTemplate is parsed; its consumer contract
  survives).

* fix(group/http): cover Spring HTTP Interface @*Exchange in Java consumer-signal gate

#2254 (merged) added Spring 6 HTTP Interface `@(Get|...)Exchange` /
`@HttpExchange` as a new Java *consumer* idiom. The #2138 parse-skip
consumer-safety gate must recognize it, or a provider-covered file carrying
an `@GetExchange` could be parse-skipped and lose that consumer contract.
Add `Exchange` to JAVA_HTTP_PLUGIN.hasConsumerSignals (conservative; also
matches `restTemplate.exchange(`).

* style(group/http): prettier formatting for #2138 Part 2 files

* style(ingestion): prettier formatting for call-processor.ts (#2138 Part 2)

* fix(group/http): P1 (Java over-claim) + P2 (handler mis-attribution) on top of #2268 (#2138 Part 2)

Re-applied on the maintainer's #2268 (expanded Java/Kotlin consumer
extraction) base.

P1 — `routeCoverage: 'complete'` over-claimed for Java: the graph provider
set is a strict subset of the group scan (array-form `@GetMapping({...})`,
interface-inherited routes, same-URL multi-verb have no graph Route node),
so parse-skip could drop those group-only providers.
- java/python → default 'partial' (always source-scanned). Java flips to
  'complete' only once ingestion provider extraction matches the group scan
  (a separate follow-up). Python was a no-op anyway (no handlerName resolved);
  'complete' was a latent trap. PHP stays 'complete' (Laravel ingestion ⊇ the
  group scan, the one language the skip engages for).
- python hasConsumerSignals widened to a true superset of scan() (uri=/url=
  wrapper, aiohttp, urllib). Java's gate already covers #2268's consumer set
  (same receivers; the @*Exchange token is present).

P2 — resolveRouteHandlerSymbols: reserve the URL slot on first encounter even
when unresolved (mirrors addRoute first-writer-wins, so a later same-URL route
can't stamp the node-winner's slot); refuse to guess on an ambiguous same-name
lookup (exactly one match → use it; zero/many → fail-open, never a wrong
handler). The cross-source case (filesystem route winning a URL a framework
route also normalizes to) is unchanged — the resolver never receives
filesystem routes — and stays fail-open.

Tests:
- route-parse-skip rewritten: the parse-skip win is proven on PHP (fully
  covered → 0 parses; mixed fallback; consumer-covered file still parsed), plus
  three Java P1 regression guards (array-form / interface-inherited / multi-verb)
  asserting the group-only routes survive — verified they go red if Java is
  flipped back to 'complete'.
- resolve-route-handler-symbols: direct unit tests (the fn had none) — unique
  resolve, ambiguous/unknown fail-open, same-URL reservation, first-writer-wins.
- http-consumer-signals: each plugin's hasConsumerSignals is a superset of its
  scan() consumer idioms; pure providers return false.
- route-handler-symbol-roundtrip: real-LadybugDB CSV→COPY→query for
  Route.handlerSymbolId.

---------

Co-authored-by: henry <zhangwei2017@unipus.cn>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-06-23 07:22:43 +01:00
204 changed files with 903874 additions and 663730 deletions
@@ -61,9 +61,10 @@ def _physical_vendor_grammars() -> set[str]:
def _render_report() -> tuple[str, int]:
"""Run main() with network mocked to mirror PRODUCTION; return (md, exit_code).
- npm grammars resolve to a permissive "Ready" peer dep, so the ONLY blocker
left is the held vendored tree-sitter-c — letting us assert the hold is
load-bearing (exit code stays non-zero because of it).
- npm grammars resolve to a permissive "Ready" peer dep, so the only blockers
left are the held vendored grammars (tree-sitter-c, tree-sitter-kotlin) plus
the intentionally-pinned tree-sitter-cpp — letting us assert holds are
load-bearing (exit code stays non-zero because of them).
- npm_view_json records its calls so we can prove vendored grammars are never
npm-queried.
- fetch_text mirrors the real workflow: upstream parser.c resolves to a real
@@ -363,13 +364,15 @@ class ReportRendering(TestCase):
# Counts are derived from _render_report()'s mock corpus (all npm peer
# deps mocked permissive): of the 10 npm-installed grammars, 9 render
# Ready and 1 — tree-sitter-cpp — is the intentional pin (#1242), so it is
# not counted ready. The 2 blockers are that same pinned tree-sitter-cpp
# plus the vendored, ABI-held tree-sitter-c (the only out-of-range
# vendored grammar). If a grammar is added/removed or a pin/hold changes,
# not counted ready. The 3 blockers are that same pinned tree-sitter-cpp
# plus two held vendored grammars: ABI-held tree-sitter-c (#1242/#858) and
# tree-sitter-kotlin (pinned to an unreleased fwcd main commit for `fun
# interface` support — ABI 14 is in range, but a hold counts as a blocker
# until it is lifted). If a grammar is added/removed or a pin/hold changes,
# update _render_report()'s mock AND these expected counts together; a
# mismatch here means the report prose drifted, not the regex.
self.assertEqual(ready.groups(), ("9", "10"))
self.assertEqual(blockers.group(1), "2")
self.assertEqual(blockers.group(1), "3")
def _matrix_row(self, name: str) -> str:
for line in self.report.splitlines():
+2 -1
View File
@@ -12,7 +12,8 @@
},
"kotlin": {
"name": "tree-sitter-kotlin",
"upstream": { "npm": "tree-sitter-kotlin" }
"upstream": { "npm": "tree-sitter-kotlin" },
"hold": "pinned to unreleased fwcd main commit c8ac3d26 for `fun interface` support (fwcd/tree-sitter-kotlin#169, closes #87) — npm latest (0.3.8) lacks the fix, so the monitor must NOT auto-revert (isNewer is strict-inequality: 0.3.8 != 0.4.0). Drop this hold and bump when upstream cuts a release that includes the fix"
},
"dart": {
"name": "tree-sitter-dart",
+145 -42
View File
@@ -14,8 +14,9 @@ name: Build tree-sitter prebuilds
# REQUIRED grammar)
# - tree-sitter-dart (vendored source; built from gitnexus/vendor/)
# - tree-sitter-proto (vendored source; built from gitnexus/vendor/)
# - tree-sitter-kotlin (vendored source; built from the published npm package —
# upstream ships source only)
# - tree-sitter-kotlin (vendored source; built from gitnexus/vendor/ — pinned to
# an unreleased main commit for `fun interface` support
# (#169) that no npm release carries yet)
# - tree-sitter-swift (vendored source; built from gitnexus/vendor/ — its
# prebuilds were originally upstream-shipped, now
# GitNexus-cross-built like the rest for uniformity)
@@ -28,12 +29,20 @@ name: Build tree-sitter prebuilds
# incl. macOS + arm64). It is DELIBERATELY NOT wired into normal PR/push CI. It
# runs only:
# 1. on manual dispatch (workflow_dispatch); or
# 2. when a covered grammar's recorded version actually CHANGES — the `guard`
# job is the real gate (it diffs the recorded version vs the PR base); the
# `paths:` filter below only makes ordinary code PRs cost ZERO matrix time.
# Net effect: an ordinary code PR triggers nothing; bumping one grammar costs
# exactly one matrix run for that grammar, which opens a PR committing its rebuilt
# binaries.
# 2. when a covered grammar's VENDORED SOURCE changes in a PR — a version bump
# OR an edit to the grammar's build-affecting source (parser.c / grammar.js /
# binding.gyp / scanner / bindings). The `guard` job is the real gate (it
# diffs BOTH the recorded version AND the source files vs the PR base); the
# `paths:` filter below keeps ordinary code PRs at ZERO matrix time and
# excludes the prebuilds the job commits back, so it never retriggers itself.
# Net effect: an ordinary code PR triggers nothing; touching one grammar's source
# costs exactly one matrix run for that grammar. Delivery of the rebuilt binaries:
# - same-repo PR -> committed straight onto the PR's own branch (in the SAME PR);
# - manual dispatch (open_pr=true) -> a fresh chore/ PR;
# - fork PR -> the trusted commit-fork-prebuilds.yml (workflow_run) pushes them
# onto the fork branch when "Allow edits by maintainers" is on, else
# comments download-and-commit instructions. That consumer must be
# on the DEFAULT branch to run, so it activates once merged to main.
#
# Concurrency convention: see CONTRIBUTING.md -> "GitHub Actions — Concurrency Convention".
#
@@ -68,13 +77,16 @@ on:
pull_request:
branches: [main]
paths:
# Vendored grammars: their version lives in the vendor snapshot package.json.
- 'gitnexus/vendor/tree-sitter-c/package.json'
- 'gitnexus/vendor/tree-sitter-dart/package.json'
- 'gitnexus/vendor/tree-sitter-proto/package.json'
- 'gitnexus/vendor/tree-sitter-kotlin/package.json'
- 'gitnexus/vendor/tree-sitter-swift/package.json'
# Transition window: kotlin's pin still lives here until it is vendored.
# Any build-affecting change under a vendored grammar triggers a rebuild —
# not just a version bump — so editing the vendored source (parser.c,
# grammar.js, binding.gyp, scanner, bindings) re-cuts the prebuilds too.
# The prebuilds we commit back are EXCLUDED (negated last) so the bot's own
# in-PR commit can never retrigger this workflow (no build->commit->build loop).
- 'gitnexus/vendor/tree-sitter-*/**'
- '!gitnexus/vendor/tree-sitter-*/prebuilds/**'
# Self-test: re-run the guard if a future grammar pin is reintroduced in
# the main package.json (optionalDependencies fallback). No-op otherwise —
# all five grammars are now fully vendored (kotlin included).
- 'gitnexus/package.json'
# Self-test: re-run the guard (normally a no-op) when the recipe changes.
- '.github/workflows/build-tree-sitter-prebuilds.yml'
@@ -102,7 +114,7 @@ jobs:
matrix: ${{ steps.decide.outputs.matrix }}
release_app: ${{ steps.relapp.outputs.configured }}
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
fetch-depth: 0 # need base history to diff recorded versions
persist-credentials: false
@@ -135,7 +147,12 @@ jobs:
c: { name: 'tree-sitter-c', kind: 'npm' },
dart: { name: 'tree-sitter-dart', kind: 'vendored' },
proto: { name: 'tree-sitter-proto', kind: 'vendored' },
kotlin: { name: 'tree-sitter-kotlin', kind: 'npm' },
// kotlin is vendored WITH its source (parser.c/scanner.c/binding.gyp),
// so it builds from gitnexus/vendor/ like dart/proto/swift. It was
// 'npm' while tracking released versions, but is now pinned to an
// unreleased main commit for `fun interface` support (#169) that no
// npm release carries yet — so it must build from the vendored source.
kotlin: { name: 'tree-sitter-kotlin', kind: 'vendored' },
// swift is vendored WITH its source (parser.c/scanner.c/binding.gyp),
// so it builds from gitnexus/vendor/ like dart/proto. Its prebuilds
// were originally upstream-shipped; rebuilding them here unifies it.
@@ -184,8 +201,13 @@ jobs:
// Resolve the base-ref recorded versions (pull_request only) so we can
// diff. On dispatch, base is irrelevant (manual intent / force wins).
const baseRoot = `${process.env.RUNNER_TEMP}/base`;
const baseSha = process.env.BASE_SHA;
// Defense in depth: baseSha is interpolated into git commands below, so
// reject anything that is not a plain commit-ish before we touch a shell.
if (event === 'pull_request' && baseSha && !/^[0-9a-fA-F]{7,40}$/.test(baseSha)) {
throw new Error(`unexpected base sha '${baseSha}'`);
}
if (event === 'pull_request') {
const baseSha = process.env.BASE_SHA;
for (const s of selected) {
const name = REGISTRY[s].name;
for (const rel of [`gitnexus/vendor/${name}/package.json`, `gitnexus/package.json`]) {
@@ -219,9 +241,24 @@ jobs:
if (event === 'workflow_dispatch') {
build = true; // manual intent (force toggles only the unchanged-guard, which is bypassed here)
} else {
// pull_request: build when the recorded version changed OR any
// build-affecting source file under the vendored grammar changed vs
// the PR base. The prebuilds/ subtree is excluded from the diff so
// the bot's own in-PR commit (which adds ONLY prebuilds) never reads
// as a source change — this is the other half of the no-loop guard.
const base = recordedVersion(baseRoot, name);
build = !!head && head !== base;
console.log(`${short}: head='${head || '<absent>'}' base='${base || '<absent>'}' -> ${build ? 'BUILD' : 'skip'}`);
const versionChanged = !!head && head !== base;
let sourceChanged = false;
try {
const diff = execSync(
`git diff --name-only ${baseSha} -- gitnexus/vendor/${name} ` +
`':(exclude)gitnexus/vendor/${name}/prebuilds/**'`,
{ stdio: ['ignore', 'pipe', 'ignore'] },
).toString().trim();
sourceChanged = diff.length > 0;
} catch { /* base unavailable -> fall back to the version gate */ }
build = versionChanged || sourceChanged;
console.log(`${short}: version ${versionChanged ? 'changed' : 'same'}, source ${sourceChanged ? 'changed' : 'same'} -> ${build ? 'BUILD' : 'skip'}`);
}
if (force) build = true;
if (!build) continue;
@@ -255,6 +292,47 @@ jobs:
echo "::notice::Release GitHub App secrets (RELEASE_APP_ID / RELEASE_APP_PRIVATE_KEY) are not configured — prebuilds will build and upload as artifacts, but the auto-PR is skipped. Provision the App, or run with open_pr=false to suppress this notice."
fi
# ── Fork PRs: emit the PR identity so the trusted `commit-fork-prebuilds`
# workflow_run job can push the rebuilt prebuilds back onto the fork's
# branch. That job has no PR context of its own (workflow_run.pull_requests
# is empty for forks), so it reads this. Same-repo PRs don't need it — the
# aggregate job below commits straight onto their branch. This artifact is
# untrusted producer output: every field is allowlist-validated again on
# the consumer side AND cross-checked against the workflow_run authority.
- name: Record fork PR identity
id: forkmeta
if: github.event_name == 'pull_request' && github.event.pull_request.head.repo.fork == true && steps.decide.outputs.any == 'true'
env:
PR_NUMBER: ${{ github.event.pull_request.number }}
HEAD_SHA: ${{ github.event.pull_request.head.sha }}
HEAD_REF: ${{ github.event.pull_request.head.ref }}
HEAD_REPO: ${{ github.event.pull_request.head.repo.full_name }}
BASE_REPO: ${{ github.repository }}
run: |
set -euo pipefail
mkdir -p "$RUNNER_TEMP/pr-meta"
# Values flow through env + jq so an exotic head_ref is quoted, never
# interpolated into a shell command.
jq -n \
--arg schema "gitnexus.ts-prebuild/v1" \
--argjson pr_number "$PR_NUMBER" \
--arg head_sha "$HEAD_SHA" \
--arg head_ref "$HEAD_REF" \
--arg head_repo "$HEAD_REPO" \
--arg base_repo "$BASE_REPO" \
'{schema:$schema, pr_number:$pr_number, head_sha:$head_sha, head_ref:$head_ref, head_repo:$head_repo, base_repo:$base_repo}' \
> "$RUNNER_TEMP/pr-meta/metadata.json"
cat "$RUNNER_TEMP/pr-meta/metadata.json"
- name: Upload fork PR meta
if: github.event_name == 'pull_request' && github.event.pull_request.head.repo.fork == true && steps.decide.outputs.any == 'true'
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: pr-meta
path: ${{ runner.temp }}/pr-meta/metadata.json
if-no-files-found: error
retention-days: 7
# ── Build one native prebuild per (grammar, platform-arch). No cross-compile. ─
build:
name: ${{ matrix.grammar }} ${{ matrix.platform_arch }}
@@ -270,7 +348,7 @@ jobs:
# and compiling them under emulation on the arm runners is slow.
timeout-minutes: 45
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false # this job uploads artifacts (artipacked)
@@ -402,16 +480,18 @@ jobs:
if-no-files-found: error
retention-days: 7
# ── Aggregate every grammar's six prebuilds, assert completeness, open a PR. ─
# ── Aggregate every grammar's six prebuilds, assert completeness, deliver them. ─
aggregate:
name: Vendor prebuilds + open PR
name: Vendor prebuilds + deliver
needs: [guard, build]
# Open the prebuild PR on a non-fork pull_request that bumped a grammar
# version (the documented version-change -> prebuild-PR flow), or on a manual
# dispatch with open_pr=true. Event-gating is explicit so we never rely on
# GHA coercing a null `inputs.open_pr` on pull_request events (Codex F4):
# `inputs.open_pr` is null off-dispatch, and `null != false` is direction-
# ambiguous, so `open_pr` is only consulted on workflow_dispatch.
# Runs on a non-fork pull_request whose vendored grammar source changed — the
# rebuilt prebuilds are committed straight onto that PR's own branch (same PR)
# — or on a manual dispatch with open_pr=true, which opens a fresh chore/ PR.
# Fork PRs are excluded: a bot cannot push into a fork branch, so they get
# artifacts only. Event-gating is explicit so we never rely on GHA coercing a
# null `inputs.open_pr` on pull_request events (Codex F4): `inputs.open_pr` is
# null off-dispatch, and `null != false` is direction-ambiguous, so `open_pr`
# is only consulted on workflow_dispatch.
if: >-
needs.guard.outputs.any == 'true' &&
needs.guard.outputs.release_app == 'true' &&
@@ -431,9 +511,13 @@ jobs:
app-id: ${{ secrets.RELEASE_APP_ID }}
private-key: ${{ secrets.RELEASE_APP_PRIVATE_KEY }}
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
token: ${{ steps.app-token.outputs.token }}
# On a (non-fork) PR, check out the PR's HEAD branch — not the merge ref —
# so the rebuilt-prebuilds commit lands on the PR's own branch (same PR).
# Empty on manual dispatch -> the workflow's default ref.
ref: ${{ github.event_name == 'pull_request' && github.event.pull_request.head.ref || '' }}
persist-credentials: false
- name: Download all prebuild artifacts
@@ -481,7 +565,7 @@ jobs:
with:
subject-path: 'gitnexus/vendor/tree-sitter-*/prebuilds/**/*.node'
- name: Create or update PR
- name: Deliver rebuilt prebuilds
uses: actions/github-script@3a2844b7e9c422d3c10d287c895573f7108da1b3 # v9.0.0
env:
GRAMMARS: ${{ steps.place.outputs.grammars }}
@@ -493,8 +577,8 @@ jobs:
const { execSync } = require('node:child_process');
const run = (c) => execSync(c, { stdio: ['ignore', 'pipe', 'inherit'] }).toString().trim();
const grammars = process.env.GRAMMARS;
const slug = grammars.replace(/[^a-z0-9]+/gi, '-');
const branch = `chore/vendor-ts-prebuilds-${slug}-${context.runId}`;
const { owner, repo } = context.repo;
const remote = `https://x-access-token:${process.env.GH_TOKEN}@github.com/${owner}/${repo}.git`;
run('git add gitnexus/vendor/tree-sitter-*/prebuilds');
if (!run('git status --porcelain -- gitnexus/vendor/tree-sitter-*/prebuilds')) {
@@ -503,16 +587,35 @@ jobs:
}
run('git config user.name "gitnexus-release-bot[bot]"');
run('git config user.email "gitnexus-release-bot[bot]@users.noreply.github.com"');
run(`git commit -m "chore(vendor): rebuild native prebuilds (${grammars})" -m "Built by ${process.env.RUN_URL}"`);
// ── Same-repo PR: ride the rebuilt prebuilds into the SAME PR by
// pushing one commit onto its head branch. The aggregate checkout
// used `ref: head.ref`, so HEAD is the PR branch tip (NOT the merge
// ref) and this is a clean fast-forward of exactly our new commit.
// Plain push (NOT --force): we only ever ADD on top of head, so we
// must never clobber the contributor's commits. If the branch
// advanced mid-build the push is rejected — and the PR's
// cancel-in-progress concurrency will already have started a fresher
// run against the new head — so a rejection is a no-op we just note.
if (context.eventName === 'pull_request') {
const headRef = context.payload.pull_request.head.ref;
try {
run(`git push "${remote}" "HEAD:${headRef}"`);
core.notice(`Pushed rebuilt prebuilds onto PR branch '${headRef}' (included in this PR).`);
} catch (e) {
core.warning(`Could not fast-forward '${headRef}' (it likely advanced mid-build); a fresher run will rebuild. ${e.message}`);
}
return;
}
// ── Manual dispatch: there is no PR to attach to, so open a fresh one
// off an ephemeral, run-unique branch. Plain --force is safe here:
// the branch is keyed by context.runId and written ONLY by this job,
// so there is no concurrent writer to protect against.
const slug = grammars.replace(/[^a-z0-9]+/gi, '-');
const branch = `chore/vendor-ts-prebuilds-${slug}-${context.runId}`;
run(`git checkout -b "${branch}"`);
run(`git commit -m "chore(vendor): rebuild native prebuilds (${grammars})\n\nBuilt by ${process.env.RUN_URL}"`);
const { owner, repo } = context.repo;
const remote = `https://x-access-token:${process.env.GH_TOKEN}@github.com/${owner}/${repo}.git`;
// Plain --force, not --force-with-lease: the branch is ephemeral and
// unique per run (keyed by context.runId), written ONLY by this job, so
// there is no concurrent writer to protect against. --force-with-lease
// would compare against a remote-tracking ref this fresh checkout never
// fetched, so re-running the SAME run (branch already pushed by attempt
// 1) fails with "stale info" instead of overwriting.
run(`git push --force "${remote}" "HEAD:${branch}"`);
const body = [
`Rebuilt the vendored native prebuilds for: **${grammars}**.`,
+2 -2
View File
@@ -36,7 +36,7 @@ jobs:
# persist-credentials: false — this job only reads (tests and syntax
# checks) and never pushes. The setting keeps GITHUB_TOKEN out of
# .git/config, which zizmor flags as the "artipacked" issue.
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
@@ -57,7 +57,7 @@ jobs:
# persist-credentials: false — this is a read-only build smoke that
# never pushes. The setting keeps GITHUB_TOKEN out of .git/config,
# which zizmor flags as the "artipacked" issue.
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
+6 -2
View File
@@ -14,7 +14,9 @@ jobs:
outputs:
web_changed: ${{ steps.filter.outputs.web }}
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
- uses: dorny/paths-filter@fbd0ab8f3e69293af611ebaee6363fc25e6d187d # v3
id: filter
with:
@@ -29,7 +31,9 @@ jobs:
runs-on: ubuntu-latest
timeout-minutes: 20
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
- name: Configure e2e GitNexus home
run: echo "GITNEXUS_HOME=${RUNNER_TEMP}/gitnexus-home" >> "$GITHUB_ENV"
+15 -5
View File
@@ -11,7 +11,9 @@ jobs:
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
with:
node-version: 22
@@ -24,7 +26,9 @@ jobs:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
with:
node-version: 22
@@ -37,7 +41,9 @@ jobs:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
- uses: ./.github/actions/setup-gitnexus
- run: npx tsc --noEmit
working-directory: gitnexus
@@ -46,7 +52,9 @@ jobs:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
- uses: ./.github/actions/setup-gitnexus-web
- run: npx tsc -b --noEmit
working-directory: gitnexus-web
@@ -67,7 +75,9 @@ jobs:
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
- name: Validate workflow concurrency convention
shell: bash
run: |
+1 -1
View File
@@ -125,7 +125,7 @@ jobs:
- name: Checkout (for vitest config)
if: steps.meta.outputs.skip != 'true'
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
sparse-checkout: gitnexus/vitest.config.ts
sparse-checkout-cone-mode: false
+15 -5
View File
@@ -11,12 +11,16 @@ jobs:
name: ubuntu / coverage
runs-on: ubuntu-latest
timeout-minutes: 25
# Fail loudly (don't silently skip) if the FTS extension is unavailable, so
# FTS-dependent lbug integration suites are guaranteed to run in CI.
env:
GITNEXUS_REQUIRE_FTS: '1'
steps:
# persist-credentials: false — this job runs tests and uploads a
# test-reports artifact (if: always()). The default-persisted token in
# .git/config must not be capturable through that upload (zizmor
# credential-persistence / artipacked audit). The job never pushes.
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
- uses: ./.github/actions/setup-gitnexus
@@ -77,10 +81,14 @@ jobs:
os: [windows-latest, macos-latest]
runs-on: ${{ matrix.os }}
timeout-minutes: 20
# Same guarantee on the platform-sensitive runners: FTS-dependent suites in
# the cross-platform subset must run, not silently skip.
env:
GITNEXUS_REQUIRE_FTS: '1'
steps:
# persist-credentials: false — runs tests only, never pushes (zizmor
# credential-persistence / artipacked audit).
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
- uses: ./.github/actions/setup-gitnexus
@@ -106,7 +114,9 @@ jobs:
runs-on: ${{ matrix.os }}
timeout-minutes: 20
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
- uses: ./.github/actions/setup-gitnexus
with:
build: 'true'
@@ -138,7 +148,7 @@ jobs:
# from a tarball and never pushes back; the token in .git/config would
# be at risk of leaking through any future artifact-upload step
# (zizmor artipacked audit). Disable upfront.
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
- uses: ./.github/actions/setup-gitnexus
@@ -246,7 +256,7 @@ jobs:
# and never pushes; the default-persisted token in .git/config would be at
# risk of leaking through an artifact upload (zizmor credential-persistence
# / artipacked audit). Mirrors the packaged-install-smoke job below.
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
- uses: ./.github/actions/setup-gitnexus
+1 -1
View File
@@ -129,7 +129,7 @@ jobs:
core.setOutput('code_review', isCodeReview ? 'true' : 'false');
- name: Checkout repository
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
repository: ${{ steps.pr.outputs.is_pr == 'true' && steps.pr.outputs.repo || github.repository }}
ref: ${{ steps.pr.outputs.is_pr == 'true' && steps.pr.outputs.sha || '' }}
+1 -1
View File
@@ -42,7 +42,7 @@ jobs:
steps:
- name: Checkout
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
# Don't leave GITHUB_TOKEN in .git/config for downstream steps to read.
persist-credentials: false
+351
View File
@@ -0,0 +1,351 @@
name: Commit fork prebuilds
# TRUSTED HALF of the vendored-grammar prebuild pipeline — FORK PRs only.
#
# `build-tree-sitter-prebuilds.yml` runs in the UNTRUSTED `pull_request`
# context. On a fork PR it has a read-only token and no secrets, so it can
# build + validate the native prebuilds and upload them as artifacts, but it
# cannot commit them back. This workflow is the trusted consumer: triggered by
# `workflow_run`, it runs from the DEFAULT BRANCH's copy of this file (the trust
# anchor) with a writable token, downloads ONLY the artifacts (data — the
# already-built-and-validated `.node` files + a small metadata.json), verifies
# the metadata against the GitHub-controlled workflow_run authority, then pushes
# the prebuilds onto the fork PR's head branch.
#
# It NEVER checks out or executes fork-controlled code: the producer already
# `require()`-loaded + parsed each `.node` on its target platform in the
# untrusted half (the correct place to run untrusted code). Here we only move
# bytes and run git. The prebuilds touch ONLY gitnexus/vendor/<g>/prebuilds/**,
# never .github/ — so the GITHUB_TOKEN's lack of `workflows` scope is irrelevant.
#
# Pushing to a fork branch with the GITHUB_TOKEN works only when the contributor
# left "Allow edits by maintainers" enabled (the PR default) — the same
# constraint as pr-autofix-apply.yml. When it's off we fall back to a comment.
#
# Same-repo PRs do NOT come here: they have secrets in the producer run, so the
# `aggregate` job in build-tree-sitter-prebuilds.yml commits straight onto their
# branch. This workflow's `if:` filters to forks.
on:
workflow_run:
workflows: ['Build tree-sitter prebuilds']
types: [completed]
concurrency:
# Per-PR identity, NOT workflow_run.id (which is per-run unique and would
# defeat serialization). Fork PRs have an empty pull_requests[] in the
# workflow_run payload, so fall back to head-repo + head-branch.
group: ${{ github.workflow }}-${{ github.event.workflow_run.pull_requests[0].number || format('{0}/{1}', github.event.workflow_run.head_repository.full_name, github.event.workflow_run.head_branch) }}
cancel-in-progress: false
permissions: {}
jobs:
deliver:
name: deliver-fork-prebuilds
# Only a SUCCESSFUL fork pull_request producer run. Same-repo PRs
# (head_repository == base) are handled by the producer's aggregate job.
if: >-
github.event.workflow_run.event == 'pull_request'
&& github.event.workflow_run.conclusion == 'success'
&& github.event.workflow_run.head_repository.full_name != github.repository
runs-on: ubuntu-latest
timeout-minutes: 15
permissions:
contents: write # push the prebuilds commit to the fork PR head branch
pull-requests: write # comment the delivery outcome
actions: read # download artifacts produced by the producer run
steps:
# Pinned to v8.0.1 (same SHA used across this repo's workflows).
- name: Download prebuild artifacts
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
continue-on-error: true
with:
run-id: ${{ github.event.workflow_run.id }}
github-token: ${{ secrets.GITHUB_TOKEN }}
pattern: ts-prebuild-*
path: prebuilds-in
- name: Download PR meta
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
continue-on-error: true
with:
name: pr-meta
run-id: ${{ github.event.workflow_run.id }}
github-token: ${{ secrets.GITHUB_TOKEN }}
path: meta-in
- name: Read and validate metadata
id: meta
shell: bash
run: |
set -euo pipefail
# No meta => this producer run had no fork-PR prebuilds to deliver
# (nothing changed, or it wasn't a fork). Exit cleanly.
if [ ! -f meta-in/metadata.json ]; then
echo "No pr-meta artifact — nothing to deliver."
echo "deliver=false" >> "$GITHUB_OUTPUT"
exit 0
fi
# No prebuild artifacts => same (defensive; producer uploads both together).
if ! ls prebuilds-in/ts-prebuild-* >/dev/null 2>&1; then
echo "No ts-prebuild-* artifacts — nothing to deliver."
echo "deliver=false" >> "$GITHUB_OUTPUT"
exit 0
fi
jq . meta-in/metadata.json
# The artifact comes from the untrusted producer running fork code.
# Allowlist EVERY field before it flows into $GITHUB_OUTPUT — a newline
# in head_ref would otherwise inject a second output line and redirect
# this job's write-scoped push/comment onto a victim PR.
assert_field() {
local key="$1" pattern="$2" value
value=$(jq -r ".${key} // empty" meta-in/metadata.json)
if [ -z "$value" ] || ! [[ "$value" =~ $pattern ]]; then
echo "::error::metadata.${key} failed allowlist (got: $(printf '%q' "$value"))"
exit 1
fi
printf '%s' "$value"
}
SCHEMA=$(assert_field schema '^gitnexus\.ts-prebuild/v[0-9]+$')
PR_NUMBER=$(assert_field pr_number '^[0-9]+$')
HEAD_SHA=$(assert_field head_sha '^[0-9a-f]{40}$')
HEAD_REF=$(assert_field head_ref '^[A-Za-z0-9._/-]+$')
HEAD_REPO=$(assert_field head_repo '^[A-Za-z0-9._-]+/[A-Za-z0-9._-]+$')
BASE_REPO=$(assert_field base_repo '^[A-Za-z0-9._-]+/[A-Za-z0-9._-]+$')
# Defence-in-depth: refuse to act if the artifact claims another repo.
if [ "$BASE_REPO" != "${GITHUB_REPOSITORY}" ]; then
echo "::error::Artifact base_repo does not match \$GITHUB_REPOSITORY — refusing to deliver."
exit 1
fi
{
echo "deliver=true"
echo "schema=${SCHEMA}"
echo "pr_number=${PR_NUMBER}"
echo "head_sha=${HEAD_SHA}"
echo "head_ref=${HEAD_REF}"
echo "head_repo=${HEAD_REPO}"
} >> "$GITHUB_OUTPUT"
# Cross-verify the artifact's claimed identity against the GitHub-controlled
# workflow_run event. The allowlist above only proves the fields are
# well-formed — not that they refer to the PR/SHA that actually triggered
# us. A fork-controlled build could mutate metadata.json to reference
# another PR/SHA and redirect our write-scoped push. Authority sources are
# all server-controlled: workflow_run.head_sha, head_repository.full_name,
# and pull_requests[].number (empty on forks -> commits/{sha}/pulls).
- name: Verify metadata against workflow_run authority
if: steps.meta.outputs.deliver == 'true'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
GH_REPO: ${{ github.repository }}
META_PR_NUMBER: ${{ steps.meta.outputs.pr_number }}
META_HEAD_SHA: ${{ steps.meta.outputs.head_sha }}
META_HEAD_REPO: ${{ steps.meta.outputs.head_repo }}
WF_HEAD_SHA: ${{ github.event.workflow_run.head_sha }}
WF_HEAD_REPO: ${{ github.event.workflow_run.head_repository.full_name }}
WF_PR_NUMBERS: ${{ toJSON(github.event.workflow_run.pull_requests.*.number) }}
shell: bash
run: |
set -euo pipefail
# 1) head_sha must match exactly — the commit GitHub ran the producer against.
if [ "${META_HEAD_SHA}" != "${WF_HEAD_SHA}" ]; then
echo "::error::Artifact head_sha (${META_HEAD_SHA}) != workflow_run.head_sha (${WF_HEAD_SHA}) — refusing."
exit 1
fi
# 2) head_repo must match exactly.
if [ "${META_HEAD_REPO}" != "${WF_HEAD_REPO}" ]; then
echo "::error::Artifact head_repo (${META_HEAD_REPO}) != workflow_run.head_repository (${WF_HEAD_REPO}) — refusing."
exit 1
fi
# 3) pr_number must reference an open PR with this head SHA. Forks have
# an empty pull_requests[] by design — fall back to commits/{sha}/pulls.
allowed_numbers=$(jq -c '.' <<< "${WF_PR_NUMBERS}")
if [ "${allowed_numbers}" = "[]" ]; then
echo "workflow_run.pull_requests empty (fork) — using commits/{sha}/pulls."
allowed_numbers=$(gh api "repos/${GH_REPO}/commits/${WF_HEAD_SHA}/pulls" \
--jq '[.[] | select(.state == "open") | .number]' 2>/dev/null || echo "[]")
if [ "${allowed_numbers}" = "[]" ]; then
echo "::error::No open PR for head ${WF_HEAD_SHA} — refusing."
exit 1
fi
fi
if ! jq -e --argjson n "${META_PR_NUMBER}" 'index($n) != null' <<< "${allowed_numbers}" >/dev/null; then
echo "::error::Artifact pr_number (${META_PR_NUMBER}) not in authoritative list (${allowed_numbers}) — refusing."
exit 1
fi
echo "Verified identity: PR=${META_PR_NUMBER} head_sha=${META_HEAD_SHA} head_repo=${META_HEAD_REPO}."
# Pinned to v6.0.3 (same SHA used by build-tree-sitter-prebuilds.yml).
# persist-credentials: false — push auth is provided inline at push time,
# never written to .git/config on disk.
- name: Checkout fork PR head
if: steps.meta.outputs.deliver == 'true'
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
repository: ${{ steps.meta.outputs.head_repo }}
ref: ${{ steps.meta.outputs.head_sha }}
token: ${{ secrets.GITHUB_TOKEN }}
persist-credentials: false
fetch-depth: 0
path: pr-checkout
- name: Place prebuilds into the fork checkout
if: steps.meta.outputs.deliver == 'true'
env:
DL: prebuilds-in
CHECKOUT: pr-checkout
shell: bash
run: |
set -euo pipefail
node --input-type=module - <<'NODE'
import fs from 'node:fs';
import { execSync } from 'node:child_process';
const dl = process.env.DL;
const checkout = process.env.CHECKOUT;
const PLATFORMS = ['linux-x64', 'linux-arm64', 'darwin-arm64', 'darwin-x64', 'win32-x64', 'win32-arm64'];
// Reconstruct {grammar -> archs} from the downloaded artifact dir names
// (ts-prebuild-<grammar>-<platform-arch>; grammar shortnames are dash-free).
const byGrammar = {};
for (const d of (fs.existsSync(dl) ? fs.readdirSync(dl) : [])) {
const m = d.match(/^ts-prebuild-([a-z0-9]+)-(.+)$/);
if (m) (byGrammar[m[1]] ||= []).push(m[2]);
}
const grammars = Object.keys(byGrammar);
if (grammars.length === 0) throw new Error('no ts-prebuild-* artifacts present');
const changed = [];
for (const grammar of grammars) {
const name = `tree-sitter-${grammar}`;
const dest = `${checkout}/gitnexus/vendor/${name}/prebuilds`;
// A grammar with 5/6 prebuilds silently breaks node-gyp-build on the
// 6th platform — refuse a partial result.
for (const pa of PLATFORMS) {
const art = `${dl}/ts-prebuild-${grammar}-${pa}/${name}.node`;
if (!fs.existsSync(art)) throw new Error(`missing ${grammar} prebuild for ${pa}`);
fs.mkdirSync(`${dest}/${pa}`, { recursive: true });
fs.copyFileSync(art, `${dest}/${pa}/${name}.node`);
}
execSync(`cd ${dest} && find . -name "*.node" | sort | xargs sha256sum > SHA256SUMS`);
changed.push(name);
}
console.log('Placed prebuilds for:', changed.join(', '));
NODE
- name: Commit and push to the fork branch
id: push
if: steps.meta.outputs.deliver == 'true'
working-directory: pr-checkout
env:
HEAD_REF: ${{ steps.meta.outputs.head_ref }}
HEAD_REPO: ${{ steps.meta.outputs.head_repo }}
HEAD_SHA: ${{ steps.meta.outputs.head_sha }}
# Push auth only — supplied via env, never interpolated into the command.
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
shell: bash
run: |
set -euo pipefail
git add gitnexus/vendor/tree-sitter-*/prebuilds
if git diff --cached --quiet; then
echo "Prebuilds byte-identical to the fork branch — nothing to commit."
echo "result=nothing-to-commit" >> "$GITHUB_OUTPUT"
exit 0
fi
# Loop guard: if HEAD is already our prebuild bot commit, don't stack
# another. (The producer's paths filter already excludes prebuilds/**,
# so a prebuild-only push cannot retrigger it — this is defence in depth.)
head_author=$(git log -1 --format='%ae' HEAD)
head_subject=$(git log -1 --format='%s' HEAD)
if [ "${head_author}" = "41898282+github-actions[bot]@users.noreply.github.com" ] \
&& [[ "${head_subject}" =~ ^chore\(vendor\) ]]; then
echo "::warning::HEAD is already a prebuild bot commit — refusing to re-apply."
echo "result=loop-prevented" >> "$GITHUB_OUTPUT"
exit 0
fi
grammars=$(git diff --cached --name-only \
| sed -n 's#gitnexus/vendor/\(tree-sitter-[a-z0-9]*\)/.*#\1#p' | sort -u | paste -sd, -)
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
git config user.name "github-actions[bot]"
git commit -q -m "chore(vendor): rebuild native prebuilds (${grammars})" \
-m "Built + validated by ${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}/actions/runs/${GITHUB_RUN_ID}"
# Push to the fork head with a lease against the resolved SHA, so a
# contributor force-push during the build surfaces as lease-failed (not
# push-failed, which would mislead them into the maintainer-edit fix).
# Auth via per-invocation http.extraheader (never persisted, never in
# the process args / git remote -v). Base64-encoded form is masked too.
push_url="${GITHUB_SERVER_URL}/${HEAD_REPO}.git"
auth_header="Authorization: Basic $(printf 'x-access-token:%s' "${GITHUB_TOKEN}" | base64 -w0)"
echo "::add-mask::${auth_header}"
push_stderr=$(mktemp)
if git -c http.extraheader="${auth_header}" \
push --force-with-lease="refs/heads/${HEAD_REF}:${HEAD_SHA}" \
"${push_url}" "HEAD:${HEAD_REF}" 2>"$push_stderr"; then
echo "result=applied" >> "$GITHUB_OUTPUT"
else
cat "$push_stderr" >&2
if grep -qE "stale info|force-with-lease|rejected.*non-fast-forward|remote rejected|! \[rejected\]" "$push_stderr"; then
echo "::error::Push lease failed — fork branch moved during build."
echo "result=lease-failed" >> "$GITHUB_OUTPUT"
else
echo "::error::Push failed — likely a fork without 'Allow edits by maintainers'."
echo "result=push-failed" >> "$GITHUB_OUTPUT"
fi
exit 0
fi
- name: Comment delivery outcome
if: always() && steps.meta.outputs.deliver == 'true' && steps.push.outcome != 'skipped'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
GH_REPO: ${{ github.repository }}
PR: ${{ steps.meta.outputs.pr_number }}
RESULT: ${{ steps.push.outputs.result }}
RUN_URL: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}
shell: bash
run: |
set -euo pipefail
marker="<!-- gitnexus:ts-prebuild-fork -->"
case "${RESULT}" in
applied)
body="${marker}
✅ **Rebuilt native prebuilds pushed to this PR branch.** A grammar source change re-cut the vendored \`tree-sitter\` prebuilds for all 6 platforms and they're now committed on your branch. ([builder run](${RUN_URL}))" ;;
nothing-to-commit)
body="${marker}
✅ Native prebuilds are already up to date on this branch — nothing to push." ;;
loop-prevented)
body="${marker}
🔁 Skipping prebuild push: the branch HEAD is already an automated prebuild commit." ;;
lease-failed)
body="${marker}
⏳ The PR head moved while the prebuilds were building, so they weren't pushed. Push another commit (or wait for the next build) and they'll be re-cut. ([builder run](${RUN_URL}))" ;;
push-failed)
body="${marker}
⚠️ Rebuilt native prebuilds are ready but **couldn't be pushed to your fork branch**. Tick **Allow edits by maintainers** in the PR sidebar so CI can commit them — or download them from the [builder run](${RUN_URL}) artifacts (\`ts-prebuild-*\`) and commit them under \`gitnexus/vendor/<grammar>/prebuilds/\` yourself." ;;
*)
body="${marker}
❓ Prebuild delivery finished in an unexpected state (\`${RESULT:-unknown}\`). See the [builder run](${RUN_URL})." ;;
esac
# Strip the YAML block indent so the rendered comment starts at column 0.
body="$(printf '%s\n' "$body" | sed 's/^ //')"
# Upsert a single sticky comment keyed by the marker; only ever edit our
# own bot comment (PATCH on someone else's 403s and would abort).
existing=$(gh api "repos/${GH_REPO}/issues/${PR}/comments" --paginate \
--jq ".[] | select(.user.login == \"github-actions[bot]\" and (.body | contains(\"${marker}\"))) | .id" \
| head -n1 || true)
if [ -n "${existing}" ]; then
gh api -X PATCH "repos/${GH_REPO}/issues/comments/${existing}" -f body="${body}" >/dev/null
echo "Updated comment ${existing}."
else
gh api -X POST "repos/${GH_REPO}/issues/${PR}/comments" -f body="${body}" >/dev/null
echo "Created delivery comment."
fi
+1 -1
View File
@@ -28,7 +28,7 @@ jobs:
steps:
- name: Checkout
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
+1 -1
View File
@@ -101,7 +101,7 @@ jobs:
# When triggered by workflow_call the caller passes the RC tag as an input;
# we check out that tag so the Dockerfile and package.json match the built image.
# For tag-push events github.ref is already the tag ref — no override needed.
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
ref: ${{ inputs.tag || github.ref }}
+1 -1
View File
@@ -29,7 +29,7 @@ jobs:
steps:
- name: Checkout
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
# Full history needed for the on-push full-history scan; on PRs the
# action diffs against the base ref so the cost is bounded by the PR.
+1 -1
View File
@@ -44,7 +44,7 @@ jobs:
permissions:
contents: read
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
@@ -37,7 +37,7 @@ jobs:
# artifact and never pushes; the default-persisted token in .git/config
# must not be capturable through that upload (zizmor credential-persistence
# / artipacked audit). Mirrors ci-tests.yml.
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
+1 -1
View File
@@ -336,7 +336,7 @@ jobs:
# Push auth is provided inline at push time via the URL.
- name: Checkout PR head
if: steps.locate.outputs.found == 'true'
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v5.0.4
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v5.0.4
with:
repository: ${{ steps.locate.outputs.head_repo }}
ref: ${{ steps.locate.outputs.head_sha }}
+1 -1
View File
@@ -51,7 +51,7 @@ jobs:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
# PR head commit (not the synthetic merge ref) — we need the
# exact tree the contributor pushed so suggestions line up.
+1 -1
View File
@@ -108,7 +108,7 @@ jobs:
# Pinned to v7.2.0. Verify SHA via:
# gh api repos/release-drafter/release-drafter/git/refs/tags/v7.2.0
# v7 removed `disable-releaser`; use `dry-run: true` to only autolabel.
- uses: release-drafter/release-drafter@693d20e7c1ce1a81d3a41962f85914253b518449 # v7.3.1
- uses: release-drafter/release-drafter@ed4bc48ec97379be2258e7b7ac2624a3e26ab809 # v7.4.0
with:
config-name: release-drafter.yml
dry-run: true
+3 -3
View File
@@ -162,7 +162,7 @@ jobs:
should_run: ${{ steps.decide.outputs.should_run }}
head_sha: ${{ steps.decide.outputs.head_sha }}
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
fetch-depth: 0
fetch-tags: true
@@ -332,7 +332,7 @@ jobs:
# on the RC path.
- name: Checkout (RC)
if: needs.route.outputs.mode == 'rc'
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
fetch-depth: 0
fetch-tags: true
@@ -349,7 +349,7 @@ jobs:
- name: Checkout (stable)
if: needs.route.outputs.mode == 'stable'
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
# No `token:` — actions/checkout uses GITHUB_TOKEN by default. Stable
# path performs no git pushes; the default scope is sufficient.
with:
+1 -1
View File
@@ -33,7 +33,7 @@ jobs:
steps:
- name: Checkout
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
@@ -52,7 +52,9 @@ jobs:
report: ${{ steps.readiness.outputs.report }}
exit_code: ${{ steps.readiness.outputs.exit_code }}
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
- uses: ./.github/actions/setup-gitnexus
with:
+1 -1
View File
@@ -59,7 +59,7 @@ jobs:
timeout-minutes: 30
steps:
- name: Checkout repository
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
sparse-checkout: .github/scripts/triage
sparse-checkout-cone-mode: false
+1 -1
View File
@@ -45,7 +45,7 @@ jobs:
steps:
- name: Checkout
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
+2 -2
View File
@@ -31,7 +31,7 @@ jobs:
contents: read
steps:
- name: Checkout
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
@@ -53,7 +53,7 @@ jobs:
steps:
- name: Checkout
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false
+12
View File
@@ -23,6 +23,18 @@ rules:
# comment in the file documents the split.
- pr-autofix-publish.yml
# workflow_run is the trusted half of the vendored-grammar prebuild
# pipeline (commit-fork-prebuilds.yml). The untrusted producer
# (build-tree-sitter-prebuilds.yml on a fork pull_request) builds +
# validates the .node prebuilds and uploads them as artifacts. This
# consumer downloads ONLY those artifacts + metadata.json,
# allowlist-validates every metadata field, cross-checks identity against
# the workflow_run authority (head_sha / head_repo / pr_number), and
# checks out the fork head pinned to that HEAD SHA solely to ADD prebuild
# files (never executes fork code) before pushing. Header comment in the
# file documents the split.
- commit-fork-prebuilds.yml
# pull_request_target needed by claude-code-action to access secrets
# and post review comments on fork PRs. Mitigated by: PR checkouts pin
# the fork's HEAD SHA (not the branch ref) to prevent TOCTOU races,
+6 -2
View File
@@ -38,7 +38,7 @@ Monorepo: **CLI/MCP** (`gitnexus/`) + **browser UI** (`gitnexus-web/`).
| `detect_changes` | Map git diffs to affected symbols and processes |
| `rename` | Graph-assisted multi-file rename with `dry_run` preview |
| `api_impact` | Pre-change impact report for an API route handler |
| `trace` | Shortest directed path between two symbols (call + class-member edges) |
| `trace` | Shortest directed path between two symbols (call + class-member edges); group-aware (`repo: "@<group>"`) for cross-repo traces |
| `route_map` | API route → handler → consumer mappings |
| `tool_map` | MCP/RPC tool definitions and handlers |
| `shape_check` | Response shape vs consumer property access mismatches |
@@ -47,7 +47,9 @@ Monorepo: **CLI/MCP** (`gitnexus/`) + **browser UI** (`gitnexus-web/`).
| `group_list` | List repo groups or details for one group |
| `group_sync` | Rebuild group Contract Registry (`contracts.json`) and bridge graph |
`query`, `context`, and `impact` are group-aware: pass `repo: "@<groupName>"` (or `"@<groupName>/<memberPath>"` to scope to one member) plus optional `service: "<monorepo/path>"`. Group-mode `query` merges per-repo results via Reciprocal Rank Fusion; group-mode `impact` runs the local walk in the chosen member and fans out across boundaries via the Contract Bridge (`gitnexus/src/core/group/cross-impact.ts`). The previously-planned `group_query`, `group_context`, `group_impact`, `group_contracts`, `group_status` MCP tools are intentionally not introduced — group-level state is exposed via resources instead:
`query`, `context`, and `impact` are group-aware: pass `repo: "@<groupName>"` (or `"@<groupName>/<memberPath>"` to scope to one member) plus optional `service: "<monorepo/path>"`. Group-mode `query` merges per-repo results via Reciprocal Rank Fusion; group-mode `impact` runs the local walk in the chosen member and fans out across boundaries via the Contract Bridge (`gitnexus/src/core/group/cross-impact.ts`). `trace` is also group-aware via `repo: "@<groupName>"` — but, unlike the others, it resolves `from`/`to` across **all** members (a `@<groupName>/<memberPath>` suffix is advisory for trace, not a scope); pass `from_uid`/`to_uid` to disambiguate a symbol name that occurs in more than one member.
Group-mode `trace` (`gitnexus/src/core/group/cross-trace.ts`) stitches a path that crosses repositories: it resolves `from`/`to` across all members, and when they live in different repos it joins the home-repo segment to the target-repo segment over a single `ContractLink` boundary (an HTTP consumer→provider link, joined on `Contract.symbolUid`), reported as a `CONTRACT_LINK` hop in `crossings[]`. The crossing is clamped to one boundary (`MAX_SUPPORTED_CROSS_DEPTH`, shared with cross-impact); deeper `crossDepth` is reported via `notes[]`. With `pdg: true` (experimental, opt-in), each boundary-adjacent segment is enriched with its intra-procedural REACHING_DEF data-flow when that repo was indexed with `--pdg` (reusing the same anchored `flows` query as `pdg_query`); data flow never crosses the repo boundary, and a missing PDG layer degrades to call-level hops with a note. Two stores meet only at the `symbolUid` grain — the per-repo PDG/call graph and the group bridge — so this is the documented join; full cross-program (SDG-like) data flow across the boundary remains deferred (see `docs/plans/2026-06-18-002-feat-unified-pdg-impact-evaluation-plan.md`). The previously-planned `group_query`, `group_context`, `group_impact`, `group_contracts`, `group_status` MCP tools are intentionally not introduced — group-level state is exposed via resources instead:
| Resource URI | Purpose |
|--------------|---------|
@@ -216,6 +218,7 @@ On a `--pdg` run the parse worker builds a per-function control-flow graph from
- **M3/M4 — TAINTED / SANITIZES / TAINT_PATH** (#2083–#2084): intra- and inter-procedural taint (source→sink) — the `explain` tool's data.
- **M5 — CDG** (#2085): Ferrante control dependence over a Cooper–Harvey–Kennedy post-dominator tree (the EXIT-rooted reverse CFG); branch sense (`'T'`/`'F'`) rides `reason`. A CFG whose EXIT is unreachable from some block is skipped for CDG (post-dominance would be unsound) while its CFG/REACHING_DEF layers are kept.
- **M6 — read surface** (#2086): the `pdg_query` MCP tool answers "what gates X?" (CDG, `mode: controls`) and "where does Y flow?" (REACHING_DEF, `mode: flows`); `explain` is the taint consumer. Both are always anchored + `LIMIT`-bounded (LadybugDB has no rel-property index) and share one `resolveBlockAnchor` helper. These PDG edge types are deliberately kept out of the default `VALID_RELATION_TYPES` / web schema.
- **Cross-repo trace enrichment**: group-mode `trace` (`pdg: true`) reuses the same anchored REACHING_DEF `flows` query to annotate a boundary-adjacent segment with how a value reaches the cross-repo call — strictly intra-procedural (data flow never crosses the repo boundary). See the group-aware tools note above.
See `core/ingestion/cfg/` (emit + the pure CFG / post-dominator / control-dependence / reaching-defs / taint passes) and `mcp/local/local-backend.ts` (`_pdgQueryImpl`, `_explainImpl`, the shared `resolveBlockAnchor`).
@@ -300,6 +303,7 @@ Each language implements `LanguageProvider` (`language-provider.ts`). Key fields
| `exportChecker` | Public/exported symbol detection |
| `typeConfig` | Type annotation extraction rules |
| `mroStrategy` | `first-wins` / `c3` / `none` |
| `descriptionExtractor` | Optional hook returning a symbol's doc-comment text as its `description`; feeds the embedding metadata header so doc-only terms are semantically searchable (issue #2270). Most languages register `createLeadingDocDescriptionExtractor` (shared, language-neutral; per-language comment/wrapper config passed at the call site) |
16 providers in `languages/index.ts` via `satisfies Record<SupportedLanguages, LanguageProvider>` — missing a language is a compile error.
+1
View File
@@ -327,6 +327,7 @@ Most `analyze` knobs are also CLI flags (`--workers`, `--worker-timeout`, `--max
| `PROF_LBUG_LOAD` | unset | When `1`, emits one `[lbug-load prof]` summary line per `loadGraphToLbug` call breaking the graph-DB persistence wall into stages (`csv-emit` / `copy-nodes` / `copy-rels` / `fallback` / `total`) plus node & edge counts. Zero-cost when unset. | Attributing large-repo analyze wall time across CSV generation vs. LadybugDB `COPY` (issue #2203) — the analyze "emit" timing is the scope-resolution bucket, not this DB-write path. |
| `GITNEXUS_MAX_FILE_SIZE` | `512` (KB) | Walker skip threshold in KB. Hard cap is `32768` (tree-sitter buffer ceiling). Equivalent to `--max-file-size <kb>`. | Indexing repos with intentionally-large source files (generated parsers, vendored bundles) that should still be parsed. |
| `GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS` | `30000` | Worker idle timeout in milliseconds before retry/fallback. Equivalent to `--worker-timeout <seconds>` × 1000. | Slow-parsing files (large minified JS, deeply-nested TS types) that legitimately need more than 30s. |
| `GITNEXUS_FTS_STEMMER` | `porter` | Stemmer used when rebuilding BM25/FTS indexes. Use `none` for CJK-heavy repositories, or a language stemmer such as `german`, `french`, or `spanish` for matching repository comments. Re-run `gitnexus analyze --repair-fts` after changing it. | Keyword search quality is poor for non-English comments or identifiers under English stemming. |
| `GITNEXUS_WAL_CHECKPOINT_THRESHOLD` | `67108864` (64 MiB) | LadybugDB WAL auto-checkpoint threshold in bytes. Equivalent to `--wal-checkpoint-threshold <bytes>`. `-1` keeps LadybugDB's stock threshold (~16 MiB). Larger thresholds reduce checkpoint frequency but increase the WAL size at rotation time — choose a smaller value on disk-constrained environments. | You need a larger or smaller WAL auto-checkpoint threshold for your analyze workload. |
| `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` | `8388608` (8 MB) | Per-job byte budget the pool will send to a worker in one `postMessage`. | Very large individual files; mostly diagnostic — bumping past 8 MB risks structured-clone memory pressure. |
| `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT` | `3` | Max replacement spawns per worker slot before the slot is dropped from the active rotation. Bounds respawn loops on a chronically-crashing slot. | Hosts where a flaky worker should retry more (raise) or fail-fast (lower) before the slot is dropped. |
+87 -78
View File
@@ -9,11 +9,11 @@
"version": "0.0.0",
"dependencies": {
"@langchain/anthropic": "^1.3.29",
"@langchain/core": "^1.1.44",
"@langchain/core": "^1.1.49",
"@langchain/google-genai": "^2.1.30",
"@langchain/langgraph": "^1.4.1",
"@langchain/ollama": "^1.2.7",
"@langchain/openai": "^1.4.5",
"@langchain/openai": "^1.5.0",
"@sigma/edge-curve": "^3.1.0",
"@tailwindcss/vite": "^4.3.0",
"axios": "^1.16.1",
@@ -28,8 +28,8 @@
"graphology-utils": "^2.3.0",
"i18next": "^26.3.0",
"i18next-browser-languagedetector": "^8.2.1",
"langchain": "^1.4.4",
"lru-cache": "^11.2.4",
"langchain": "^1.4.6",
"lru-cache": "^11.5.1",
"lucide-react": "^1.17.0",
"mermaid": "^11.15.0",
"mnemonist": "^0.40.4",
@@ -59,7 +59,7 @@
"@types/react-syntax-highlighter": "^15.5.13",
"@vercel/node": "^5.8.12",
"@vitejs/plugin-react": "^5.1.4",
"@vitest/coverage-v8": "^4.1.8",
"@vitest/coverage-v8": "^4.1.9",
"jsdom": "^29.1.1",
"tree-sitter-wasms": "^0.1.13",
"typescript": "^5.4.5",
@@ -1357,9 +1357,9 @@
}
},
"node_modules/@langchain/core": {
"version": "1.1.48",
"resolved": "https://registry.npmjs.org/@langchain/core/-/core-1.1.48.tgz",
"integrity": "sha512-fQU6Guyb1pwc2fEplmA8FPbKfOMAofjnyJzExevro0FxEiuGHE18Ov/ZHmT9trWCDTZRI9eW1VIc6aChxV8pAQ==",
"version": "1.2.1",
"resolved": "https://registry.npmjs.org/@langchain/core/-/core-1.2.1.tgz",
"integrity": "sha512-NNG/cC5FGuHDOAP56h0ddp8Rfk8p+othWzEK5RV9JIG6RvnF5vGa5r0AEGtKfQieed7s1kC42GuIzVOBvMBL/g==",
"license": "MIT",
"dependencies": {
"@cfworker/json-schema": "^4.0.2",
@@ -1510,20 +1510,20 @@
}
},
"node_modules/@langchain/openai": {
"version": "1.4.5",
"resolved": "https://registry.npmjs.org/@langchain/openai/-/openai-1.4.5.tgz",
"integrity": "sha512-bQ2WMIZfSh02trJLYSAtiIcD3j6EBCiAm9nw0dZWQsVaUxmWc3JJqs8uUte6AkMazmLHzcUIw+14UkXO5fRJvQ==",
"version": "1.5.0",
"resolved": "https://registry.npmjs.org/@langchain/openai/-/openai-1.5.0.tgz",
"integrity": "sha512-ooC02qF3wnQ5m0WyibVPO5vCkgyZwjWPgNrpGFSTv3ZLnKfW1yC4k2Fp4qOf6qoVmwTeYSW4C+wNiiZ3PXshMA==",
"license": "MIT",
"dependencies": {
"js-tiktoken": "^1.0.12",
"openai": "^6.34.0",
"openai": "^6.41.0",
"zod": "^3.25.76 || ^4"
},
"engines": {
"node": ">=20"
},
"peerDependencies": {
"@langchain/core": "^1.1.42"
"@langchain/core": "^1.2.0"
}
},
"node_modules/@langchain/protocol": {
@@ -3044,14 +3044,14 @@
}
},
"node_modules/@vitest/coverage-v8": {
"version": "4.1.8",
"resolved": "https://registry.npmjs.org/@vitest/coverage-v8/-/coverage-v8-4.1.8.tgz",
"integrity": "sha512-lt3kovsyHwYe00wq4D1ti0Z974fWj4NLp6siqiyEufUpyFwK9Yhi7rBhac9JL5aA0zoMrJqc4vYPZRUnI7l7nw==",
"version": "4.1.9",
"resolved": "https://registry.npmjs.org/@vitest/coverage-v8/-/coverage-v8-4.1.9.tgz",
"integrity": "sha512-G9/lgqibheLVBDRuya45EbsEXTYcWoSG+TLg7i2axuzx0Eq62eXn+aWXyaVdV5vKvFSWd6ywcX8hA7la9Pvu8g==",
"dev": true,
"license": "MIT",
"dependencies": {
"@bcoe/v8-coverage": "^1.0.2",
"@vitest/utils": "4.1.8",
"@vitest/utils": "4.1.9",
"ast-v8-to-istanbul": "^1.0.0",
"istanbul-lib-coverage": "^3.2.2",
"istanbul-lib-report": "^3.0.1",
@@ -3065,8 +3065,8 @@
"url": "https://opencollective.com/vitest"
},
"peerDependencies": {
"@vitest/browser": "4.1.8",
"vitest": "4.1.8"
"@vitest/browser": "4.1.9",
"vitest": "4.1.9"
},
"peerDependenciesMeta": {
"@vitest/browser": {
@@ -3075,16 +3075,16 @@
}
},
"node_modules/@vitest/expect": {
"version": "4.1.8",
"resolved": "https://registry.npmjs.org/@vitest/expect/-/expect-4.1.8.tgz",
"integrity": "sha512-h3nDO677RDLEGlBxyQ5CW8RlMThSKSRLUePLOx09gNIWRL40edgA1GCZSZgf1W55MFAG6/Sw14KeaAnqv0NKdQ==",
"version": "4.1.9",
"resolved": "https://registry.npmjs.org/@vitest/expect/-/expect-4.1.9.tgz",
"integrity": "sha512-vl/rYsUKcBr3SnQn166+XR5ZQcgMx3DQhFWdfli/cWpLnLUmbxZvyrJZotLFUryib+LtArYMSTJ5RbQ57ZqrlA==",
"dev": true,
"license": "MIT",
"dependencies": {
"@standard-schema/spec": "^1.1.0",
"@types/chai": "^5.2.2",
"@vitest/spy": "4.1.8",
"@vitest/utils": "4.1.8",
"@vitest/spy": "4.1.9",
"@vitest/utils": "4.1.9",
"chai": "^6.2.2",
"tinyrainbow": "^3.1.0"
},
@@ -3093,13 +3093,13 @@
}
},
"node_modules/@vitest/mocker": {
"version": "4.1.8",
"resolved": "https://registry.npmjs.org/@vitest/mocker/-/mocker-4.1.8.tgz",
"integrity": "sha512-LEiN/xe4OSIbKe9HQIp5OC24agGD9J5CnmMgsLohVVoOPWL9a2sBoR6VBx43jQZb7Kr1l4RCuyCJzcAa0+dojw==",
"version": "4.1.9",
"resolved": "https://registry.npmjs.org/@vitest/mocker/-/mocker-4.1.9.tgz",
"integrity": "sha512-EVkXzBjrPGM+cK8/ANWgBrkUCfJfb38/EfTSO8h7pWvKkyPkpWxvR7BkD2MyItMF62C97zAEoqdpUixwR/e+Rw==",
"dev": true,
"license": "MIT",
"dependencies": {
"@vitest/spy": "4.1.8",
"@vitest/spy": "4.1.9",
"estree-walker": "^3.0.3",
"magic-string": "^0.30.21"
},
@@ -3130,9 +3130,9 @@
}
},
"node_modules/@vitest/pretty-format": {
"version": "4.1.8",
"resolved": "https://registry.npmjs.org/@vitest/pretty-format/-/pretty-format-4.1.8.tgz",
"integrity": "sha512-9GasEBxpZ1VYIpqHf/0+YGg121uSNwCKOJqIrTwWP/TB7DmFCiaBpNl3aPZzoLWfWkuqhbH8vJIVobZkvdo2cA==",
"version": "4.1.9",
"resolved": "https://registry.npmjs.org/@vitest/pretty-format/-/pretty-format-4.1.9.tgz",
"integrity": "sha512-s0iufns3iIFitdgm+YR7g1whCAaGtXz459VS9/PqyKDEEFgYIhsHOQmXgIgDuYCt7DeQmiZT0Qe2OA2p4ZPu5A==",
"dev": true,
"license": "MIT",
"dependencies": {
@@ -3143,13 +3143,13 @@
}
},
"node_modules/@vitest/runner": {
"version": "4.1.8",
"resolved": "https://registry.npmjs.org/@vitest/runner/-/runner-4.1.8.tgz",
"integrity": "sha512-EmVxeBAfMJvycdjd6Hm+RbFBbA9fKvo0Kx37hNpBYoYeavH3RNsBXWDooR1mgD52dCrxIIuP7UotpfiwOikvcg==",
"version": "4.1.9",
"resolved": "https://registry.npmjs.org/@vitest/runner/-/runner-4.1.9.tgz",
"integrity": "sha512-KXLMDtc7oe70+3mJfGrPUWPesswH+3sTxAMAMl8DG7I8IUQT4XW718dY5ID3vPUcmlu27CcKfY4P3h3I29SLJg==",
"dev": true,
"license": "MIT",
"dependencies": {
"@vitest/utils": "4.1.8",
"@vitest/utils": "4.1.9",
"pathe": "^2.0.3"
},
"funding": {
@@ -3157,14 +3157,14 @@
}
},
"node_modules/@vitest/snapshot": {
"version": "4.1.8",
"resolved": "https://registry.npmjs.org/@vitest/snapshot/-/snapshot-4.1.8.tgz",
"integrity": "sha512-acfZboRmAIf05DEKcBQy33VXojFJjtUdLyo7oOmV9kebb2xdU01UknNiPuPZoJZQyO7DF0gZdTGTpeAzET9QPQ==",
"version": "4.1.9",
"resolved": "https://registry.npmjs.org/@vitest/snapshot/-/snapshot-4.1.9.tgz",
"integrity": "sha512-Jc7RKGNBo8Z28WYIm0Niej4xdSPByRf6mU58VpHQkd6Zh05rlnA+twjbK5HyeIGHxrzsc3mJgS43uM0CZKzaIA==",
"dev": true,
"license": "MIT",
"dependencies": {
"@vitest/pretty-format": "4.1.8",
"@vitest/utils": "4.1.8",
"@vitest/pretty-format": "4.1.9",
"@vitest/utils": "4.1.9",
"magic-string": "^0.30.21",
"pathe": "^2.0.3"
},
@@ -3173,9 +3173,9 @@
}
},
"node_modules/@vitest/spy": {
"version": "4.1.8",
"resolved": "https://registry.npmjs.org/@vitest/spy/-/spy-4.1.8.tgz",
"integrity": "sha512-6EevtBp6OZOPF7bmz36HrGMeP3txgVSrgebWxHOafDXGkhIzfXK14f8KF6MuFfgXXUeHxmpD3BQxkV00/3s5mA==",
"version": "4.1.9",
"resolved": "https://registry.npmjs.org/@vitest/spy/-/spy-4.1.9.tgz",
"integrity": "sha512-fHpsS6mIi+PiEW+vcRVOMkX1oSaPKne3VOclSFICPcGOmfKgXPU5iAah+wcNcj2xPrCCmfq99IDGf+EojhhvhA==",
"dev": true,
"license": "MIT",
"funding": {
@@ -3183,13 +3183,13 @@
}
},
"node_modules/@vitest/utils": {
"version": "4.1.8",
"resolved": "https://registry.npmjs.org/@vitest/utils/-/utils-4.1.8.tgz",
"integrity": "sha512-uOJamYALNhfJ6iolExyQM40yIQwDqYnkKtQ5VCiSe17E33H0aQ/u+1GlRuz4LZBk6Mm3sg90G9hEbmEt37C1Zg==",
"version": "4.1.9",
"resolved": "https://registry.npmjs.org/@vitest/utils/-/utils-4.1.9.tgz",
"integrity": "sha512-A51o8ymO5PpqlWNnBP9ZHPXDIpuMtTLlGSjN7la4US+LJzoUMyhwjA5QXlm39JexgwHKW4Xjs8Z2d3dLCXOeuA==",
"dev": true,
"license": "MIT",
"dependencies": {
"@vitest/pretty-format": "4.1.8",
"@vitest/pretty-format": "4.1.9",
"convert-source-map": "^2.0.0",
"tinyrainbow": "^3.1.0"
},
@@ -5707,13 +5707,13 @@
"integrity": "sha512-Ls993zuzfayK269Svk9hzpeGUKob/sIgZzyHYdjQoAdQetRKpOLj+k/QQQ/6Qi0Yz65mlROrfd+Ev+1+7dz9Kw=="
},
"node_modules/langchain": {
"version": "1.4.4",
"resolved": "https://registry.npmjs.org/langchain/-/langchain-1.4.4.tgz",
"integrity": "sha512-tepOCwUDaIZOYJ9Eo0O6o5dXEN/0KJheiFDnHHFL8Tx8rfkDLL4cOTSTln4Vpn9LpWzXYkjQ8lkHnnNDQWZPeg==",
"version": "1.4.6",
"resolved": "https://registry.npmjs.org/langchain/-/langchain-1.4.6.tgz",
"integrity": "sha512-pwuFmGOyiMezptLVLrpb5jILirvYPGHI5uJCFHL5K5WPxMy2XuPLI5QNMKtoHkdiL6a2dLebqugKw87cneaESw==",
"license": "MIT",
"dependencies": {
"@langchain/langgraph": "^1.3.2",
"@langchain/langgraph-checkpoint": "^1.0.1",
"@langchain/langgraph": "^1.3.4",
"@langchain/langgraph-checkpoint": "^1.0.4",
"langsmith": ">=0.5.0 <1.0.0",
"zod": "^3.25.76 || ^4"
},
@@ -5721,7 +5721,7 @@
"node": ">=20"
},
"peerDependencies": {
"@langchain/core": "^1.1.48"
"@langchain/core": "^1.2.0"
}
},
"node_modules/langsmith": {
@@ -6050,9 +6050,9 @@
}
},
"node_modules/lru-cache": {
"version": "11.3.6",
"resolved": "https://registry.npmjs.org/lru-cache/-/lru-cache-11.3.6.tgz",
"integrity": "sha512-Gf/KoL3C/MlI7Bt0PGI9I+TeTC/I6r/csU58N4BSNc4lppLBeKsOdFYkK+dX0ABDUMJNfCHTyPpzwwO21Awd3A==",
"version": "11.5.1",
"resolved": "https://registry.npmjs.org/lru-cache/-/lru-cache-11.5.1.tgz",
"integrity": "sha512-RPimw/7aMdv2oqRrxKwvZXcPfwBrn/JZ2xYcY9Hus/6LaS3VOAKVWKWgNLCFSiOm1ESXinjsDlidVU7JlnCN2A==",
"license": "BlueOak-1.0.0",
"engines": {
"node": "20 || >=22"
@@ -7307,18 +7307,27 @@
}
},
"node_modules/openai": {
"version": "6.34.0",
"resolved": "https://registry.npmjs.org/openai/-/openai-6.34.0.tgz",
"integrity": "sha512-yEr2jdGf4tVFYG6ohmr3pF6VJuveP0EA/sS8TBx+4Eq5NT10alu5zg2dmxMXMgqpihRDQlFGpRt2XwsGj+Fyxw==",
"version": "6.45.0",
"resolved": "https://registry.npmjs.org/openai/-/openai-6.45.0.tgz",
"integrity": "sha512-5DQVNErssk0afNpTTHUm/qZPU4iKR9OYdNid8Ib4puq4gHNNvGWZht2zY4h9a8JMF949Ik6m8gQutllVPbjdnw==",
"license": "Apache-2.0",
"bin": {
"openai": "bin/cli"
},
"peerDependencies": {
"@aws-sdk/credential-provider-node": ">=3.972.0 <4",
"@smithy/hash-node": ">=4.3.0 <5",
"@smithy/signature-v4": ">=5.4.0 <6",
"ws": "^8.18.0",
"zod": "^3.25 || ^4.0"
},
"peerDependenciesMeta": {
"@aws-sdk/credential-provider-node": {
"optional": true
},
"@smithy/hash-node": {
"optional": true
},
"@smithy/signature-v4": {
"optional": true
},
"ws": {
"optional": true
},
@@ -8789,19 +8798,19 @@
}
},
"node_modules/vitest": {
"version": "4.1.8",
"resolved": "https://registry.npmjs.org/vitest/-/vitest-4.1.8.tgz",
"integrity": "sha512-flY6ScbCIt9HThs+C5HS7jvGOB560DJtk/Z15IQROTA6zEy49Nh8T/dofWTQL+n3vswqn87sbJNiuqw1SDp5Ig==",
"version": "4.1.9",
"resolved": "https://registry.npmjs.org/vitest/-/vitest-4.1.9.tgz",
"integrity": "sha512-nE3/LEyc0z87uHYLZebqCUOaJr2hdtuPp7BQ4BosVFnfltxgAvMG08NyrSGlPpOUWvR27c5flSmYFTNr78L9GQ==",
"dev": true,
"license": "MIT",
"dependencies": {
"@vitest/expect": "4.1.8",
"@vitest/mocker": "4.1.8",
"@vitest/pretty-format": "4.1.8",
"@vitest/runner": "4.1.8",
"@vitest/snapshot": "4.1.8",
"@vitest/spy": "4.1.8",
"@vitest/utils": "4.1.8",
"@vitest/expect": "4.1.9",
"@vitest/mocker": "4.1.9",
"@vitest/pretty-format": "4.1.9",
"@vitest/runner": "4.1.9",
"@vitest/snapshot": "4.1.9",
"@vitest/spy": "4.1.9",
"@vitest/utils": "4.1.9",
"es-module-lexer": "^2.0.0",
"expect-type": "^1.3.0",
"magic-string": "^0.30.21",
@@ -8829,12 +8838,12 @@
"@edge-runtime/vm": "*",
"@opentelemetry/api": "^1.9.0",
"@types/node": "^20.0.0 || ^22.0.0 || >=24.0.0",
"@vitest/browser-playwright": "4.1.8",
"@vitest/browser-preview": "4.1.8",
"@vitest/browser-webdriverio": "4.1.8",
"@vitest/coverage-istanbul": "4.1.8",
"@vitest/coverage-v8": "4.1.8",
"@vitest/ui": "4.1.8",
"@vitest/browser-playwright": "4.1.9",
"@vitest/browser-preview": "4.1.9",
"@vitest/browser-webdriverio": "4.1.9",
"@vitest/coverage-istanbul": "4.1.9",
"@vitest/coverage-v8": "4.1.9",
"@vitest/ui": "4.1.9",
"happy-dom": "*",
"jsdom": "*",
"vite": "^6.0.0 || ^7.0.0 || ^8.0.0"
+5 -5
View File
@@ -19,11 +19,11 @@
},
"dependencies": {
"@langchain/anthropic": "^1.3.29",
"@langchain/core": "^1.1.44",
"@langchain/core": "^1.1.49",
"@langchain/google-genai": "^2.1.30",
"@langchain/langgraph": "^1.4.1",
"@langchain/ollama": "^1.2.7",
"@langchain/openai": "^1.4.5",
"@langchain/openai": "^1.5.0",
"@sigma/edge-curve": "^3.1.0",
"@tailwindcss/vite": "^4.3.0",
"axios": "^1.16.1",
@@ -38,8 +38,8 @@
"graphology-utils": "^2.3.0",
"i18next": "^26.3.0",
"i18next-browser-languagedetector": "^8.2.1",
"langchain": "^1.4.4",
"lru-cache": "^11.2.4",
"langchain": "^1.4.6",
"lru-cache": "^11.5.1",
"lucide-react": "^1.17.0",
"mermaid": "^11.15.0",
"mnemonist": "^0.40.4",
@@ -69,7 +69,7 @@
"@types/react-syntax-highlighter": "^15.5.13",
"@vercel/node": "^5.8.12",
"@vitejs/plugin-react": "^5.1.4",
"@vitest/coverage-v8": "^4.1.8",
"@vitest/coverage-v8": "^4.1.9",
"jsdom": "^29.1.1",
"tree-sitter-wasms": "^0.1.13",
"typescript": "^5.4.5",
+4
View File
@@ -363,6 +363,7 @@ Configure the behavior with two environment variables:
| -------------------------------------------- | ---------------------------- | ------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `GITNEXUS_LBUG_EXTENSION_INSTALL` | `auto`, `load-only`, `never` | `auto` | `auto` runs one bounded INSTALL if LOAD fails. `load-only` only uses already-installed extensions (recommended for offline / firewalled environments). `never` skips optional extensions entirely. |
| `GITNEXUS_LBUG_EXTENSION_INSTALL_TIMEOUT_MS` | positive integer | `15000` | Wall-clock budget for the out-of-process `INSTALL` child before it is killed. |
| `GITNEXUS_FTS_STEMMER` | supported LadybugDB stemmer | `porter` | Stemmer used when rebuilding BM25/FTS indexes. Use `none` for CJK-heavy repositories, or a language stemmer such as `german`, `french`, or `spanish` when that better matches repository comments and identifiers. Re-run `gitnexus analyze --repair-fts` after changing it. |
| `GITNEXUS_WAL_CHECKPOINT_THRESHOLD` | integer `>= -1` | `67108864` (64 MiB) | LadybugDB WAL auto-checkpoint threshold during analyze (bytes). Auto-checkpoint remains enabled; `-1` keeps Ladybug's stock ~16 MiB. Larger thresholds reduce checkpoint frequency but increase the WAL size at rotation time — choose a smaller value on disk-constrained environments. |
```bash
@@ -371,6 +372,9 @@ GITNEXUS_LBUG_EXTENSION_INSTALL=load-only npx gitnexus analyze
# Slow network: give extension downloads more time
GITNEXUS_LBUG_EXTENSION_INSTALL_TIMEOUT_MS=30000 npx gitnexus analyze
# CJK-heavy codebase: rebuild keyword indexes without English stemming
GITNEXUS_FTS_STEMMER=none npx gitnexus analyze --repair-fts
```
### Analysis runs out of memory
+102
View File
@@ -0,0 +1,102 @@
# Cross-repo trace — end-to-end verification
Verifies the cross-repo `trace` MCP tool against the **real pipeline** (not
hand-persisted graphs): `runFullAnalysis(--pdg)` on two repos → real `syncGroup`
HTTP contract extraction + bridge build → `callTool('trace', { repo: '@group' })`.
Run from `gitnexus/` (needs a current build for the parse worker):
```bash
node scripts/build.js
node bench/cross-repo-trace/verify.mjs
```
`verify.mjs` is self-contained — it generates each fixture inline, runs the real
analyze → sync → trace/impact pipeline, and prints PASS/FAIL per assertion
(exit non-zero on any failure). Expected verdict: **16/16 checks passed**.
## Cases covered (one scenario each)
1. **Named handlers, same file** — a frontend with named `fetch` wrappers
(`fetchUsers`, `createUserReq`) and a backend with named express handlers
(`listUsers`, `createUser`) on `/api/users` GET/POST. Asserts: all four
contracts resolve a `symbolUid`; `trace` is **symbol-precise** (the GET pair
selects `http::GET`, the POST pair `http::POST`, no file-fallback note); the
destination trace lands at `listUsers`.
2. **Anonymous handler** — `router.get('/api/ping', (req,res) => …)`. Asserts the
provider contract has an empty `symbolUid`, and the **destination trace**
(omit `to`) reaches it, reported as `<http::GET::/api/ping handler>` with an
anonymous note.
3. **Cross-repo `impact` fan-out** — `impact @group` on `fetchUsers` crosses the
boundary (`cross_repo_hits >= 1`); the same `symbolUid` join was 0 before.
4. **Multi-language (Python)** — a Flask provider + `requests` consumer; asserts
the Python line wiring resolves the consumer and the cross-repo `trace`
stitches `fetch_items -> list_items`.
5. **Cross-file named handler** (#2275) — a route whose handler (`listUsers`) is
imported from another file than its registration. Asserts the provider
resolves to the handler via the import-pinned module lookup, and the trace is
symbol-precise (no file-level fallback).
6. **Aliased cross-file import** (#2275) — `import { listUsers as handleUsers }`
with an unrelated decoy `handleUsers` elsewhere. Asserts the route resolves
through the import to the declared `listUsers` (not the alias or the decoy),
proving import-pinned resolution.
7. **Python aliased import** (#2275) — a Flask `add_url_rule('/api/users',
view_func=handle_users)` whose view is `from .handlers.users import list_users
as handle_users`. Asserts the handler resolves through Python's dotted
relative module to `list_users`, symbol-precise.
The **ambiguous-destination** (a file making several HTTP calls whose consumer
contracts have no resolved uid) and **degraded-member** (a member DB that throws
mid-resolution) paths need synthetic inputs the real analyzer cannot produce, so
they live in the unit suite (`test/unit/group/cross-trace.test.ts`).
## What it proves
- `analyze` + `syncGroup` build the correct `ContractLink`s (exact HTTP match).
- HTTP contracts carry a **real `symbolUid` whenever the endpoint resolves** —
the extractor binds each detection to the function it lives in (the function
CONTAINING the `fetch`; the named handler, or the inline handler by line-span
containment, for a route). A handler/consumer that resolves to no named symbol
(a fully anonymous handler, or a language plugin that does not yet set the
call-site line) keeps an empty uid and degrades to the file/destination
fallback. When resolved, contracts report
`extractionStrategy: 'source_scan_resolved'` / `'graph_assisted'` with a uid.
- `trace @group from=<calling fn> to=<handler fn>` **stitches the cross-repo
path** (`fetchUsers → listUsers`), reporting the `CONTRACT_LINK` hop and
(with `pdg:true`) the data-flow enrichment, **symbol-precise** (GET pair →
`http::GET` contract, POST → `http::POST`), with no file-fallback note.
- The same `symbolUid` fix makes `impact @group` fan out across the boundary
(it was 0 cross-repo hits before — both tools join crossings on `symbolUid`).
## Resolution precedence & residual limits
The extractor resolves `symbolUid` in this order, falling through on a miss:
1. **Named handler** — `router.get('/x', listUsers)` resolves `listUsers` by name.
2. **Containment** — the innermost `Function`/`Method` whose line span encloses
the call/registration line (consumers; inline-arrow providers).
3. **File-level boundary fallback** (in `cross-trace`) — only when 1–2 leave the
uid empty: if the user's `from`/`to` resolves into the contract's file, that
endpoint anchors the boundary. A `notes[]` entry flags it as file-level, not
symbol-precise.
The call-site line is set by all bundled language plugins (Node/TS, Python, Go,
PHP, Kotlin, Java), and containment matches symbols by `filePath` across
`Function`/`Method`/`CodeElement`, so it also resolves methods nested in classes
(Java/Kotlin), not just top-level functions.
### Anonymous handlers — the destination trace
A **fully anonymous handler** (`router.get('/x', (req,res) => res.json(...))`)
has no symbol node at all, so it cannot be named as a `to` target. This is
handled by the **destination trace**: omit `to`/`to_uid`/`to_file` on an
`@group` trace and `trace from=<consumer>` follows the consumer's outgoing HTTP
call across the bridge and reports where it lands — by route + file:line, with a
`notes[]` entry flagging the handler as anonymous:
```
app/frontend:fetchUsers → app/backend:<http::GET::/api/users handler> [CONTRACT_LINK]
```
To go deeper into an anonymous handler, trace to a named function it calls (the
provider segment then resolves normally).
+452
View File
@@ -0,0 +1,452 @@
/**
* Cross-repo trace — comprehensive end-to-end verification.
*
* Drives the REAL pipeline (runFullAnalysis --pdg -> real syncGroup -> trace /
* impact via a LocalBackend) over inline fixtures, one scenario per implemented
* case, and reports PASS/FAIL per assertion. Run from gitnexus/ (needs a current
* build for the parse worker):
*
* node scripts/build.js
* node bench/cross-repo-trace/verify.mjs
*
* Cases covered: symbolUid containment resolution (named, same-file + nested),
* symbol-precise crossing selection, the destination trace (named + anonymous
* endpoint), cross-repo impact fan-out, and multi-language (Python) resolution.
* (Ambiguous-destination and degraded-member paths need synthetic inputs the
* real analyzer can't produce; those are covered in the unit suite.)
*/
import fs from 'node:fs';
import path from 'node:path';
import os from 'node:os';
const REPO = path.resolve('.');
const { runFullAnalysis } = await import(path.join(REPO, 'dist/core/run-analyze.js'));
const { getGroupDir } = await import(path.join(REPO, 'dist/core/group/storage.js'));
const { loadGroupConfig } = await import(path.join(REPO, 'dist/core/group/config-parser.js'));
const { syncGroup } = await import(path.join(REPO, 'dist/core/group/sync.js'));
const { LocalBackend } = await import(path.join(REPO, 'dist/mcp/local/local-backend.js'));
const cb = { onProgress: () => {}, onLog: () => {} };
const ANALYZE = { pdg: true, skipSkills: true, embeddings: false, force: true };
const line = (s = '') => console.log(s);
const results = [];
const check = (pass, label, detail = '') => {
results.push({ pass, label });
line(` [${pass ? 'PASS' : 'FAIL'}] ${label}${detail ? ` — ${detail}` : ''}`);
};
function writeFiles(dir, files) {
for (const [rel, content] of Object.entries(files)) {
const p = path.join(dir, rel);
fs.mkdirSync(path.dirname(p), { recursive: true });
fs.writeFileSync(p, content);
}
}
function groupYaml(name, repos) {
const lines = Object.entries(repos)
.map(([k, v]) => ` ${k}: ${v}`)
.join('\n');
return `version: 1
name: ${name}
description: ""
repos:
${lines}
links: []
packages: {}
detect:
http: true
matching:
bm25_threshold: 0.7
embedding_threshold: 0.65
max_candidates_per_step: 3
`;
}
/** Analyze each repo, sync the group, return a ready LocalBackend + sync result. */
async function setup(tag, repos, groupName, groupRepos) {
const home = fs.mkdtempSync(path.join(os.tmpdir(), `gn-bench-${tag}-`));
process.env.GITNEXUS_HOME = home;
for (const [reg, files] of Object.entries(repos)) {
const dir = path.join(home, reg);
writeFiles(dir, files);
await runFullAnalysis(dir, ANALYZE, cb);
}
const gd = getGroupDir(home, groupName);
fs.mkdirSync(gd, { recursive: true });
fs.writeFileSync(path.join(gd, 'group.yaml'), groupYaml(groupName, groupRepos));
const sync = await syncGroup(await loadGroupConfig(gd), { groupDir: gd });
const backend = new LocalBackend();
await backend.init();
return { home, sync, backend };
}
const hasNote = (r, frag) => (r.notes ?? []).some((n) => n.includes(frag));
const crossingId = (r) => r.crossings?.[0]?.contractId;
// ── Scenario 1+3: named handlers (precise trace, destination, impact fan-out) ──
line('## Scenario: named handlers (same-file) — symbolUid precise');
{
const { sync, backend, home } = await setup(
'named',
{
'named-backend': {
'src/routes.ts': `import { Router } from 'express';
const router = Router();
export function listUsers(req: { body: unknown }, res: { json: (v: unknown) => void }) { res.json([]); }
export function createUser(req: { body: unknown }, res: { json: (v: unknown) => void }) { res.json({}); }
router.get('/api/users', listUsers);
router.post('/api/users', createUser);
export default router;
`,
'package.json': '{ "name": "named-backend", "version": "1.0.0" }',
},
'named-frontend': {
'src/api.ts': `export async function fetchUsers() {
const r = await fetch('/api/users');
return r.json();
}
export async function createUserReq(data: { name: string }) {
const r = await fetch('/api/users', { method: 'POST', body: JSON.stringify(data) });
return r.json();
}
`,
'package.json': '{ "name": "named-frontend", "version": "1.0.0" }',
},
},
'named-group',
{ 'app/backend': 'named-backend', 'app/frontend': 'named-frontend' },
);
const resolved = sync.contracts.filter((c) => c.symbolUid).length;
check(resolved >= 4, `all 4 contracts resolve a symbolUid (got ${resolved}/4)`);
const get = await backend.callTool('trace', {
repo: '@named-group',
from: 'fetchUsers',
to: 'listUsers',
pdg: true,
});
check(
get.status === 'ok' && crossingId(get) === 'http::GET::/api/users' && !hasNote(get, 'file'),
'GET trace is symbol-precise (fetchUsers -> listUsers over http::GET::/api/users, no file fallback)',
`status=${get.status} crossing=${crossingId(get)}`,
);
const post = await backend.callTool('trace', {
repo: '@named-group',
from: 'createUserReq',
to: 'createUser',
pdg: true,
});
check(
post.status === 'ok' && crossingId(post) === 'http::POST::/api/users',
'POST trace selects the POST crossing (no GET/POST confusion)',
`crossing=${crossingId(post)}`,
);
const dest = await backend.callTool('trace', { repo: '@named-group', from: 'fetchUsers' });
check(
dest.status === 'ok' &&
dest.to?.name === 'listUsers' &&
crossingId(dest) === 'http::GET::/api/users',
'destination trace (no `to`) lands at the named handler listUsers',
`to=${dest.to?.name}`,
);
const imp = await backend.callTool('impact', {
repo: '@named-group/app/frontend',
target: 'fetchUsers',
direction: 'downstream',
});
const hits = imp.summary?.cross_repo_hits ?? (Array.isArray(imp.cross) ? imp.cross.length : 0);
check(hits >= 1, `impact @group fans out across the boundary (cross_repo_hits=${hits})`);
fs.rmSync(home, { recursive: true, force: true });
}
// ── Scenario 2: anonymous handler — destination reports endpoint by route ──
line('\n## Scenario: anonymous handler — destination trace');
{
const { sync, backend, home } = await setup(
'anon',
{
'anon-backend': {
'src/routes.ts': `import { Router } from 'express';
const router = Router();
router.get('/api/ping', (req: unknown, res: { json: (v: unknown) => void }) => { res.json({ ok: true }); });
export default router;
`,
'package.json': '{ "name": "anon-backend", "version": "1.0.0" }',
},
'anon-frontend': {
'src/ping.ts': `export async function ping() {
const r = await fetch('/api/ping');
return r.json();
}
`,
'package.json': '{ "name": "anon-frontend", "version": "1.0.0" }',
},
},
'anon-group',
{ 'app/backend': 'anon-backend', 'app/frontend': 'anon-frontend' },
);
const provider = sync.contracts.find((c) => c.role === 'provider');
check(
provider !== undefined && !provider.symbolUid,
'anonymous provider has an empty symbolUid (no named symbol to resolve)',
`uid=${provider?.symbolUid || 'empty'}`,
);
const dest = await backend.callTool('trace', { repo: '@anon-group', from: 'ping' });
check(
dest.status === 'ok' &&
dest.to?.name === '<http::GET::/api/ping handler>' &&
hasNote(dest, 'anonymous'),
'destination trace reaches the anonymous handler, reported by route + anonymous note',
`to=${dest.to?.name}`,
);
fs.rmSync(home, { recursive: true, force: true });
}
// ── Scenario 4: multi-language (Python) — symbolUid resolution beyond TS ──
line('\n## Scenario: multi-language (Python) — line wiring + resolution');
{
const { sync, backend, home } = await setup(
'py',
{
'py-backend': {
'app.py': `from flask import Flask
app = Flask(__name__)
@app.route('/api/items')
def list_items():
return []
`,
},
'py-frontend': {
'client.py': `import requests
def fetch_items():
return requests.get('/api/items').json()
`,
},
},
'py-group',
{ 'app/backend': 'py-backend', 'app/frontend': 'py-frontend' },
);
line(
` (py contracts: ${sync.contracts
.map((c) => `${c.role}:${c.symbolName}:${c.symbolUid ? 'uid' : 'empty'}`)
.join(' ')} | crossLinks=${sync.crossLinks.length})`,
);
check(
sync.crossLinks.length >= 1,
`Python HTTP link built (crossLinks=${sync.crossLinks.length})`,
);
const tr = await backend.callTool('trace', {
repo: '@py-group',
from: 'fetch_items',
to: 'list_items',
});
check(
tr.status === 'ok' && crossingId(tr) === 'http::GET::/api/items',
'Python cross-repo trace stitches fetch_items -> list_items',
`status=${tr.status} crossing=${crossingId(tr) ?? tr.role}`,
);
// The Flask provider resolves no symbol here, so the provider boundary is
// anchored by the contract FILE (to=list_items lives in the provider file).
// This exercises the file-level fallback path end-to-end.
check(
hasNote(tr, 'FILE'),
'provider boundary uses the file-level fallback when the provider has no uid',
`notes=${(tr.notes ?? []).length}`,
);
fs.rmSync(home, { recursive: true, force: true });
}
// ── Scenario: Python Flask add_url_rule with an ALIASED relative import —
// import-pinned resolution across Python's dotted module syntax. ───────────
line('\n## Scenario: Python aliased import (Flask add_url_rule) — import-pinned');
{
const { sync, backend, home } = await setup(
'pyalias',
{
'pyalias-backend': {
'app/handlers/users.py': `def list_users():
return []
`,
'app/routes.py': `from flask import Flask
from .handlers.users import list_users as handle_users
app = Flask(__name__)
app.add_url_rule('/api/users', view_func=handle_users)
`,
},
'pyalias-frontend': {
'client.py': `import requests
def fetch_users():
return requests.get('/api/users').json()
`,
},
},
'pyalias-group',
{ 'app/backend': 'pyalias-backend', 'app/frontend': 'pyalias-frontend' },
);
const provider = sync.contracts.find(
(c) => c.role === 'provider' && c.contractId === 'http::GET::/api/users',
);
check(
provider?.symbolName === 'list_users',
'Python Flask aliased view resolves through the relative import to list_users',
`sym=${provider?.symbolName} uid=${provider?.symbolUid ? 'set' : 'empty'}`,
);
const tr = await backend.callTool('trace', {
repo: '@pyalias-group',
from: 'fetch_users',
to: 'list_users',
});
check(
tr.status === 'ok' && crossingId(tr) === 'http::GET::/api/users' && !hasNote(tr, 'FILE'),
'Python aliased-import trace is symbol-precise (no file-level fallback)',
`status=${tr.status} crossing=${crossingId(tr)}`,
);
fs.rmSync(home, { recursive: true, force: true });
}
// ── Scenario: cross-file named handler (#2275) — repo-wide unique resolution ──
line('\n## Scenario: cross-file named handler — repo-wide unique resolution');
{
const { sync, backend, home } = await setup(
'xfile',
{
'xfile-backend': {
'src/handlers/users.ts': `export function listUsers(req: { body: unknown }, res: { json: (v: unknown) => void }) {
res.json([]);
}
`,
'src/routes.ts': `import { Router } from 'express';
import { listUsers } from './handlers/users';
const router = Router();
router.get('/api/users', listUsers);
export default router;
`,
'package.json': '{ "name": "xfile-backend", "version": "1.0.0" }',
},
'xfile-frontend': {
'src/api.ts': `export async function fetchUsers() {
const r = await fetch('/api/users');
return r.json();
}
`,
'package.json': '{ "name": "xfile-frontend", "version": "1.0.0" }',
},
},
'xfile-group',
{ 'app/backend': 'xfile-backend', 'app/frontend': 'xfile-frontend' },
);
const provider = sync.contracts.find(
(c) => c.role === 'provider' && c.contractId === 'http::GET::/api/users',
);
check(
Boolean(provider?.symbolUid) && provider?.symbolName === 'listUsers',
'cross-file provider resolves to the handler defined in another file (repo-wide unique)',
`sym=${provider?.symbolName} uid=${provider?.symbolUid ? 'set' : 'empty'}`,
);
const tr = await backend.callTool('trace', {
repo: '@xfile-group',
from: 'fetchUsers',
to: 'listUsers',
pdg: true,
});
check(
tr.status === 'ok' && crossingId(tr) === 'http::GET::/api/users' && !hasNote(tr, 'FILE'),
'cross-file trace is symbol-precise (no file-level fallback)',
`status=${tr.status} crossing=${crossingId(tr)}`,
);
fs.rmSync(home, { recursive: true, force: true });
}
// ── Scenario: ALIASED cross-file import — resolved through the import to the
// declared symbol, not the local alias (and not a same-named decoy). ───────
line('\n## Scenario: aliased cross-file import — import-pinned resolution');
{
const { sync, backend, home } = await setup(
'alias',
{
'alias-backend': {
'src/handlers/users.ts': `export function listUsers(req: { body: unknown }, res: { json: (v: unknown) => void }) {
res.json([]);
}
`,
// Decoy: a DIFFERENT, unrelated symbol named handleUsers. Name-only
// resolution of the local alias would wrongly pick this one.
'src/util.ts': `export function handleUsers() {
return 1;
}
`,
'src/routes.ts': `import { Router } from 'express';
import { listUsers as handleUsers } from './handlers/users';
const router = Router();
router.get('/api/users', handleUsers);
export default router;
`,
'package.json': '{ "name": "alias-backend", "version": "1.0.0" }',
},
'alias-frontend': {
'src/api.ts': `export async function fetchUsers() {
const r = await fetch('/api/users');
return r.json();
}
`,
'package.json': '{ "name": "alias-frontend", "version": "1.0.0" }',
},
},
'alias-group',
{ 'app/backend': 'alias-backend', 'app/frontend': 'alias-frontend' },
);
const provider = sync.contracts.find(
(c) => c.role === 'provider' && c.contractId === 'http::GET::/api/users',
);
check(
provider?.symbolName === 'listUsers',
'aliased handler resolves through the import to the declared symbol (not the alias/decoy)',
`sym=${provider?.symbolName} uid=${provider?.symbolUid ? 'set' : 'empty'}`,
);
const tr = await backend.callTool('trace', {
repo: '@alias-group',
from: 'fetchUsers',
to: 'listUsers',
pdg: true,
});
check(
tr.status === 'ok' && crossingId(tr) === 'http::GET::/api/users' && !hasNote(tr, 'FILE'),
'aliased-import trace is symbol-precise (no file-level fallback)',
`status=${tr.status} crossing=${crossingId(tr)}`,
);
fs.rmSync(home, { recursive: true, force: true });
}
// ── Summary ────────────────────────────────────────────────────────────────
const passed = results.filter((r) => r.pass).length;
line(`\n## Verdict: ${passed}/${results.length} checks passed`);
if (passed !== results.length) {
line(' FAILED:');
for (const r of results.filter((x) => !x.pass)) line(` - ${r.label}`);
}
process.exit(passed === results.length ? 0 : 1);
@@ -1,5 +1,5 @@
{
"fingerprint": "4cc418ea87b6d20a68b5c1139f35d81820b715c63de0ec812e73e2135f5b00b1",
"fingerprint": "b169463b7d02185d757b6d8601db6215ac6e7b2a20e52fb0f1276cc153836bd4",
"scaling_budget": 1.8,
"max_ms_large": 1000,
"_note": "fingerprint = sha256 over per-file digests (filename + sha256(file bytes)), entry list sorted — binds each emitted line to its file so a row routed to the WRONG pair file changes the hash, AND catches within-file row reordering (file bytes hashed as-written). Byte-identity gate for #2203 U2/U3. NOTE: a future change that legitimately reorders emit (without changing the node/edge SET) will trip --check; regenerate then. scaling_budget bounds (t_large/t_small)/(LARGE/SMALL): observed ~0.95-1.05 (linear); 1.8 tolerates disk-I/O timing noise on CI while still catching an O(n^2) re-regression (~4x). max_ms_large=1000ms is a coarse absolute backstop (observed ~200ms) that catches a gross uniform slowdown the ratio gate misses; generous so CI host noise won't flake it. Regenerate via `node --import tsx bench/emit-persistence/measure.mjs`."
+3 -2
View File
@@ -80,9 +80,10 @@
"_rebaselined": "#1956 synth-widening: + javascript-qualified-base fixture; synthesizeJsInheritanceReferences now handles a member_expression base (class S extends ns.Base -> Base), matching the #1940 legacy leg + the TS terminalTsTypeNameNode property_identifier case, at parity. Linear (~1.05). | #942: scope-resolution-only cleanup reworded fixture comments; capture byte-positions shift, capture LOGIC unchanged."
},
"kotlin": {
"fingerprint": "90aa832978d9744e50058e77a04748390a7e34e36b309f6c1d178eb07280b7ea",
"fingerprint": "4900431791f2b9280009deb2b82659c26ead8aa6fb8731190a7c505dec5a9041",
"scaling_budget": 1.5,
"_added": "#1951: bench coverage added (was ungated); scale source heritage-bearing (: Base()); js/kotlin O(n^2) findNodeAtRange-per-match fixed to threaded captured node, now linear.",
"_rebaselined": "#1919 review CF3 fix: extended kotlin-local-property-owner (init/accessor destructuring) + new dart-accessor-owner fixture (getter/setter ownership). Fingerprint-only corpus drift; scaling ~1.0."
"_rebaselined": "#1919 review CF3 fix: extended kotlin-local-property-owner (init/accessor destructuring) + new dart-accessor-owner fixture (getter/setter ownership). Fingerprint-only corpus drift; scaling ~1.0.",
"_rebaselined_2271": "PR #2271: re-vendored tree-sitter-kotlin 0.3.8 -> unreleased fwcd main c8ac3d26 for `fun interface` support + new kotlin-fun-interface fixture in the corpus. Drift is both corpus-additive (the fixture) and grammar-driven (the new grammar parses `fun interface` as a class_declaration, not an ERROR node). Baselined to the NEW grammar's fingerprint, so this --check passes only once the regenerated prebuilds land — until then CI loads the committed 0.3.8 binary and the bench is red, same as the kotlin fun-interface integration tests. scaling ~0.83 (linear)."
}
}
+23 -17
View File
@@ -1,12 +1,12 @@
{
"name": "gitnexus",
"version": "1.6.8",
"version": "1.6.9-rc.34",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "gitnexus",
"version": "1.6.8",
"version": "1.6.9-rc.34",
"hasInstallScript": true,
"license": "PolyForm-Noncommercial-1.0.0",
"dependencies": {
@@ -16,7 +16,7 @@
"@scarf/scarf": "^1.4.0",
"busboy": "^1.6.0",
"cli-progress": "^3.12.0",
"commander": "^14.0.3",
"commander": "^15.0.0",
"cors": "^2.8.5",
"express": "^5.2.1",
"express-rate-limit": "^8.4.1",
@@ -2527,12 +2527,12 @@
}
},
"node_modules/commander": {
"version": "14.0.3",
"resolved": "https://registry.npmjs.org/commander/-/commander-14.0.3.tgz",
"integrity": "sha512-H+y0Jo/T1RZ9qPP4Eh1pkcQcLRglraJaSLoyOtHxu6AapkjWVCy2Sit1QQ4x3Dng8qDlSsZEet7g5Pq06MvTgw==",
"version": "15.0.0",
"resolved": "https://registry.npmjs.org/commander/-/commander-15.0.0.tgz",
"integrity": "sha512-z67u4ZhzCL/Tydu1lJARtEZYWbWaN7oYLHbsuzocr6y4N6WZAagG3RQ4FW61V1/0+jImpj293XfrcYnd1qxtPg==",
"license": "MIT",
"engines": {
"node": ">=20"
"node": ">=22.12.0"
}
},
"node_modules/content-disposition": {
@@ -4200,15 +4200,15 @@
}
},
"node_modules/onnxruntime-common": {
"version": "1.26.0",
"resolved": "https://registry.npmjs.org/onnxruntime-common/-/onnxruntime-common-1.26.0.tgz",
"integrity": "sha512-qVyMR4lcWgbkc4getFV+GQijsTnbg/siteoqcDwa3sI/LxbrMSNw4ePyvCq/ymdQaRomCA7YuWmhzsswxvymdw==",
"version": "1.27.0",
"resolved": "https://registry.npmjs.org/onnxruntime-common/-/onnxruntime-common-1.27.0.tgz",
"integrity": "sha512-3KxL5wIVqa8Ex08jxSzncm9CMgw8CjOFyOQ7SxvG9o0cVLlhTNKXyIQuTbtX4tGPJEf73OER2xrjt4HJSBL4ow==",
"license": "MIT"
},
"node_modules/onnxruntime-node": {
"version": "1.26.0",
"resolved": "https://registry.npmjs.org/onnxruntime-node/-/onnxruntime-node-1.26.0.tgz",
"integrity": "sha512-OHl6PiOEOqxaLHL0N9eFrbzS7IGmu3BtJNH3RTEnRAheCIkfc3gjcjl4sGcjp9C22ZC9YTquDOxSdT/stBQ6BQ==",
"version": "1.27.0",
"resolved": "https://registry.npmjs.org/onnxruntime-node/-/onnxruntime-node-1.27.0.tgz",
"integrity": "sha512-QEzGwrvNBgv4uPVdnbHsOGG4G6T96mdlcFI8aAKPjMU8wOPpVocPXb6k3QGkaZagVTv2G9Bnnbo6Z3JdXr1fQw==",
"hasInstallScript": true,
"license": "MIT",
"os": [
@@ -4219,9 +4219,15 @@
"dependencies": {
"adm-zip": "^0.5.16",
"global-agent": "^4.1.3",
"onnxruntime-common": "1.26.0"
"onnxruntime-common": "1.27.0"
}
},
"node_modules/onnxruntime-node/node_modules/onnxruntime-common": {
"version": "1.26.0",
"resolved": "https://registry.npmjs.org/onnxruntime-common/-/onnxruntime-common-1.26.0.tgz",
"integrity": "sha512-qVyMR4lcWgbkc4getFV+GQijsTnbg/siteoqcDwa3sI/LxbrMSNw4ePyvCq/ymdQaRomCA7YuWmhzsswxvymdw==",
"license": "MIT"
},
"node_modules/onnxruntime-web": {
"version": "1.26.0-dev.20260416-b7804b056c",
"resolved": "https://registry.npmjs.org/onnxruntime-web/-/onnxruntime-web-1.26.0-dev.20260416-b7804b056c.tgz",
@@ -5424,9 +5430,9 @@
"license": "MIT"
},
"node_modules/uuid": {
"version": "14.0.0",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-14.0.0.tgz",
"integrity": "sha512-Qo+uWgilfSmAhXCMav1uYFynlQO7fMFiMVZsQqZRMIXp0O7rR7qjkj+cPvBHLgBqi960QCoo/PH2/6ZtVqKvrg==",
"version": "14.0.1",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-14.0.1.tgz",
"integrity": "sha512-6ZxzVpzDXDa3bJWaHilVayA+BH/1zmxCJoVgvmqJnid/gPoKHxUrS/aC/T6LGQtNHT+XHG9fXPJB4d+IrU30Ew==",
"funding": [
"https://github.com/sponsors/broofa",
"https://github.com/sponsors/ctavan"
+2 -2
View File
@@ -1,6 +1,6 @@
{
"name": "gitnexus",
"version": "1.6.8",
"version": "1.6.9-rc.34",
"description": "Graph-powered code intelligence for AI agents. Index any codebase, query via MCP or CLI.",
"author": "Abhigyan Patwari",
"license": "PolyForm-Noncommercial-1.0.0",
@@ -61,7 +61,7 @@
"@scarf/scarf": "^1.4.0",
"busboy": "^1.6.0",
"cli-progress": "^3.12.0",
"commander": "^14.0.3",
"commander": "^15.0.0",
"cors": "^2.8.5",
"express": "^5.2.1",
"express-rate-limit": "^8.4.1",
+6
View File
@@ -36,6 +36,7 @@ const PLATFORM_LOGIC = [
'test/unit/lbug-pool-fts-load.test.ts',
'test/unit/repo-manager.test.ts',
'test/unit/repo-manager-finalize-invariant.test.ts',
'test/unit/git-utils.test.ts',
'test/unit/hooks.test.ts',
'test/unit/hook-db-lock-probe.test.ts',
'test/unit/cursor-hook.test.ts',
@@ -62,10 +63,15 @@ const LBUG_NATIVE = [
'test/integration/lbug-orphan-sidecar-recovery.test.ts',
'test/integration/lbug-readonly-init.test.ts',
'test/integration/lbug-non-ascii-path.test.ts',
// Cross-repo trace e2e: builds two real lbug indexes + a real bridge and
// opens them through the pool adapter (native addon + bridge file locking).
// Windows is skipped in-file (describeReopen) due to the bridge reopen lock.
'test/integration/group/cross-trace-e2e.test.ts',
'test/integration/local-backend.test.ts',
'test/integration/local-backend-calltool.test.ts',
'test/integration/search-core.test.ts',
'test/integration/search-pool.test.ts',
'test/integration/fts-description-search.test.ts',
'test/integration/staleness-and-stability.test.ts',
'test/integration/analyze-wal-checkpoint-failure.test.ts',
];
+1 -1
View File
@@ -285,5 +285,5 @@ export const en = {
'help.option.group.contracts.repo': 'Filter by repo',
'help.option.group.contracts.unmatched': 'Show only unmatched contracts',
'help.analyze.environment':
'\nEnvironment variables:\n GITNEXUS_NO_GITIGNORE=1 Skip .gitignore parsing (still reads .gitnexusignore)\n GITNEXUS_MAX_FILE_SIZE=N Override large-file skip threshold (KB). Default 512, max 32768.\n GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS=N Worker idle timeout in milliseconds. Default 30000.\n GITNEXUS_WAL_CHECKPOINT_THRESHOLD=N LadybugDB WAL auto-checkpoint threshold in bytes (default 67108864 = 64 MiB; -1 keeps Ladybug stock ~16 MiB).\n GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES=N Worker job byte budget. Default 8388608.\n GITNEXUS_WORKER_POOL_SIZE=N Parse worker count override. Default cores-1 capped at 16.\n GITNEXUS_PARSE_CHUNK_CONCURRENCY=N Concurrent in-flight parse chunks. Default 2.\n GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT=N Max replacement spawns per slot before drop. Default 3.\n GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS=N Total retry wall-time per job. Default 5x sub-batch timeout.\n GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD=N Per-slot deaths to trip circuit breaker. Default max(3, poolSize).\n GITNEXUS_EMBEDDING_THREADS=N Limit local ONNX CPU threads for --embeddings.\n GITNEXUS_SEMANTIC_EXACT_SCAN_LIMIT=N Max embedding chunks for exact-scan fallback. Default 10000.\n\nFlags override the corresponding env vars when both are provided.\n\nTip: `.gitnexusignore` supports `.gitignore`-style negation. Add e.g.\n `!__tests__/` to index a directory that is auto-filtered by default (#771).',
'\nEnvironment variables:\n GITNEXUS_NO_GITIGNORE=1 Skip .gitignore parsing (still reads .gitnexusignore)\n GITNEXUS_MAX_FILE_SIZE=N Override large-file skip threshold (KB). Default 512, max 32768.\n GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS=N Worker idle timeout in milliseconds. Default 30000.\n GITNEXUS_WAL_CHECKPOINT_THRESHOLD=N LadybugDB WAL auto-checkpoint threshold in bytes (default 67108864 = 64 MiB; -1 keeps Ladybug stock ~16 MiB).\n GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES=N Worker job byte budget. Default 8388608.\n GITNEXUS_WORKER_POOL_SIZE=N Parse worker count override. Default cores-1 capped at 16.\n GITNEXUS_PARSE_CHUNK_CONCURRENCY=N Concurrent in-flight parse chunks. Default 2.\n GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT=N Max replacement spawns per slot before drop. Default 3.\n GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS=N Total retry wall-time per job. Default 5x sub-batch timeout.\n GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD=N Per-slot deaths to trip circuit breaker. Default max(3, poolSize).\n GITNEXUS_EMBEDDING_THREADS=N Limit local ONNX CPU threads for --embeddings.\n GITNEXUS_SEMANTIC_EXACT_SCAN_LIMIT=N Max embedding chunks for exact-scan fallback. Default 10000.\n GITNEXUS_VECTOR_MAX_DISTANCE=N Max accepted semantic/vector cosine distance (0 < N <= 2; higher values clamp to 2). Default 0.6 for MCP, 0.5 elsewhere.\n\nFlags override the corresponding env vars when both are provided.\n\nTip: `.gitnexusignore` supports `.gitignore`-style negation. Add e.g.\n `!__tests__/` to index a directory that is auto-filtered by default (#771).',
} as const;
+1 -1
View File
@@ -265,5 +265,5 @@ export const zhCN = {
'help.option.group.contracts.repo': '按仓库过滤',
'help.option.group.contracts.unmatched': '仅显示未匹配契约',
'help.analyze.environment':
'\n环境变量:\n GITNEXUS_NO_GITIGNORE=1 跳过 .gitignore 解析(仍读取 .gitnexusignore)\n GITNEXUS_MAX_FILE_SIZE=N 覆盖大文件跳过阈值(KB)。默认 512,最大 32768。\n GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS=N Worker 空闲超时(毫秒)。默认 30000。\n GITNEXUS_WAL_CHECKPOINT_THRESHOLD=N LadybugDB WAL 自动 checkpoint 阈值(字节,默认 67108864 = 64 MiB;-1 保持 Ladybug 默认约 16 MiB)。\n GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES=N Worker 作业字节预算。默认 8388608。\n GITNEXUS_WORKER_POOL_SIZE=N 解析 worker 数量覆盖值。默认 cores-1,最多 16。\n GITNEXUS_PARSE_CHUNK_CONCURRENCY=N 并发进行中的解析分块数。默认 2。\n GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT=N 每个 slot 丢弃前允许的最大替换进程数。默认 3。\n GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS=N 每个作业的总重试墙钟时间。默认 5 倍子批次超时。\n GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD=N 每个 slot 触发熔断的死亡次数。默认 max(3, poolSize)。\n GITNEXUS_EMBEDDING_THREADS=N 限制 --embeddings 的本地 ONNX CPU 线程数。\n GITNEXUS_SEMANTIC_EXACT_SCAN_LIMIT=N exact-scan 回退的最大嵌入分块数。默认 10000。\n\n当参数和对应环境变量同时提供时,参数优先。\n\n提示:`.gitnexusignore` 支持 `.gitignore` 风格的取反。比如添加\n `!__tests__/` 可以索引默认自动过滤的目录(#771)。',
'\n环境变量:\n GITNEXUS_NO_GITIGNORE=1 跳过 .gitignore 解析(仍读取 .gitnexusignore)\n GITNEXUS_MAX_FILE_SIZE=N 覆盖大文件跳过阈值(KB)。默认 512,最大 32768。\n GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS=N Worker 空闲超时(毫秒)。默认 30000。\n GITNEXUS_WAL_CHECKPOINT_THRESHOLD=N LadybugDB WAL 自动 checkpoint 阈值(字节,默认 67108864 = 64 MiB;-1 保持 Ladybug 默认约 16 MiB)。\n GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES=N Worker 作业字节预算。默认 8388608。\n GITNEXUS_WORKER_POOL_SIZE=N 解析 worker 数量覆盖值。默认 cores-1,最多 16。\n GITNEXUS_PARSE_CHUNK_CONCURRENCY=N 并发进行中的解析分块数。默认 2。\n GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT=N 每个 slot 丢弃前允许的最大替换进程数。默认 3。\n GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS=N 每个作业的总重试墙钟时间。默认 5 倍子批次超时。\n GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD=N 每个 slot 触发熔断的死亡次数。默认 max(3, poolSize)。\n GITNEXUS_EMBEDDING_THREADS=N 限制 --embeddings 的本地 ONNX CPU 线程数。\n GITNEXUS_SEMANTIC_EXACT_SCAN_LIMIT=N exact-scan 回退的最大嵌入分块数。默认 10000。\n GITNEXUS_VECTOR_MAX_DISTANCE=N 语义/向量搜索接受的最大余弦距离(0 < N <= 2;超出则钳制为 2)。MCP 默认 0.6,其他路径默认 0.5。\n\n当参数和对应环境变量同时提供时,参数优先。\n\n提示:`.gitnexusignore` 支持 `.gitignore` 风格的取反。比如添加\n `!__tests__/` 可以索引默认自动过滤的目录(#771)。',
} satisfies EnglishMessages;
+47
View File
@@ -1,6 +1,53 @@
import { defaultEmbeddingThreads } from '../platform/capabilities.js';
import { logger } from '../logger.js';
import { DEFAULT_EMBEDDING_CONFIG, type EmbeddingConfig } from './types.js';
export const DEFAULT_VECTOR_MAX_DISTANCE = 0.5;
export const DEFAULT_MCP_VECTOR_MAX_DISTANCE = 0.6;
/**
* Cosine distance over normalized embeddings is bounded to [0, 2], so any threshold
* above this accepts every row and silently disables the relevance filter. Values
* over the ceiling are clamped to it rather than passed through.
*/
export const VECTOR_MAX_DISTANCE_CEILING = 2;
const warned = new Set<string>();
const warnOnce = (key: string, message: string): void => {
if (warned.has(key)) return;
warned.add(key);
logger.warn(message);
};
/**
* Resolve the effective max accepted vector/semantic cosine distance.
* Reads `GITNEXUS_VECTOR_MAX_DISTANCE`. Unset/empty/whitespace → silent fallback.
* Invalid (non-numeric, <= 0, non-finite) → fallback plus a one-time warning.
* Values above the cosine ceiling (2) are clamped to it with a one-time warning.
*/
export const getVectorMaxDistance = (fallback: number = DEFAULT_VECTOR_MAX_DISTANCE): number => {
const raw = process.env.GITNEXUS_VECTOR_MAX_DISTANCE;
if (raw === undefined || raw.trim() === '') return fallback;
const parsed = Number(raw);
if (!Number.isFinite(parsed) || parsed <= 0) {
warnOnce(
`invalid:${raw}`,
` GITNEXUS_VECTOR_MAX_DISTANCE must be a positive number in (0, ${VECTOR_MAX_DISTANCE_CEILING}], got "${raw}" — using default ${fallback}`,
);
return fallback;
}
if (parsed > VECTOR_MAX_DISTANCE_CEILING) {
warnOnce(
`clamp:${raw}`,
` GITNEXUS_VECTOR_MAX_DISTANCE=${parsed} exceeds the cosine-distance ceiling (${VECTOR_MAX_DISTANCE_CEILING}) — clamping`,
);
return VECTOR_MAX_DISTANCE_CEILING;
}
return parsed;
};
const parsePositiveInt = (name: string, value: string | undefined, fallback: number): number => {
if (value === undefined) return fallback;
const parsed = Number(value);
@@ -26,7 +26,6 @@ import {
type EmbeddableNode,
type SemanticSearchResult,
type ModelProgress,
type EmbeddingContext,
EMBEDDABLE_LABELS,
isShortLabel,
LABEL_METHOD,
@@ -34,7 +33,11 @@ import {
STRUCTURAL_LABELS,
collectBestChunks,
} from './types.js';
import { resolveEmbeddingConfig } from './config.js';
import {
DEFAULT_VECTOR_MAX_DISTANCE,
getVectorMaxDistance,
resolveEmbeddingConfig,
} from './config.js';
import { rankExactEmbeddingRows, type ExactEmbeddingRow } from './exact-search.js';
import { EMBEDDING_TABLE_NAME, EMBEDDING_INDEX_NAME, STALE_HASH_SENTINEL } from '../lbug/schema.js';
import { loadVectorExtension, createVectorIndex } from '../lbug/lbug-adapter.js';
@@ -76,7 +79,7 @@ const ensureVectorExtensionAvailable = async (): Promise<boolean> => {
* invalidate existing vectors, such as metadata/header shape changes,
* structural container context changes, or preceding-context formatting rules.
*/
export const EMBEDDING_TEXT_VERSION = 'v2';
export const EMBEDDING_TEXT_VERSION = 'v4';
/**
* Compute a stable content fingerprint for an embeddable node.
@@ -251,6 +254,42 @@ export interface EmbeddingPipelineResult {
semanticMode: 'vector-index' | 'exact-scan';
}
/**
* DELETE stale embedding rows for the given nodeIds so they can be re-inserted.
*
* Kuzu forbids SET on vector-indexed properties; DELETE-then-INSERT is the
* sanctioned pattern. A `"does not exist"` error means the rows are already gone
* (safe to proceed); any other error risks vector-index corruption, so it
* propagates and aborts the pipeline.
*
* Called per-batch (just before each batch's INSERT), not once up front — see
* the caller comment / KTD7: an up-front bulk delete of every stale row leaves
* the whole index deleted-not-reinserted if the re-embed is interrupted. Per-batch
* interleaving bounds that window to a single batch.
*/
const deleteStaleEmbeddingRows = async (
executeWithReusedStatement: (
cypher: string,
paramsList: Array<Record<string, any>>,
) => Promise<void>,
nodeIds: string[],
): Promise<void> => {
if (nodeIds.length === 0) return;
try {
await executeWithReusedStatement(
`MATCH (e:${EMBEDDING_TABLE_NAME} {nodeId: $nodeId}) DELETE e`,
nodeIds.map((nodeId) => ({ nodeId })),
);
} catch (err) {
const msg = err instanceof Error ? err.message : String(err);
if (!msg.includes('does not exist')) {
throw new Error(
`[embed] Failed to delete stale embedding rows — aborting to prevent vector-index corruption: ${msg}`,
);
}
}
};
/**
* Run the embedding pipeline
*
@@ -259,11 +298,9 @@ export interface EmbeddingPipelineResult {
* @param onProgress - Callback for progress updates
* @param config - Optional configuration override
* @param skipNodeIds - Optional set of node IDs that already have embeddings (incremental mode)
* @param context - Optional repo/server context for metadata enrichment
* @param existingEmbeddings - Optional map of nodeId → contentHash for incremental mode.
* Nodes whose hash matches are skipped; nodes with a changed hash are DELETE'd
* and re-embedded; nodes not in the map are embedded fresh.
*/
export const runEmbeddingPipeline = async (
executeQuery: (cypher: string) => Promise<any[]>,
@@ -274,7 +311,6 @@ export const runEmbeddingPipeline = async (
onProgress: EmbeddingProgressCallback,
config: Partial<EmbeddingConfig> = {},
skipNodeIds?: Set<string>,
context?: EmbeddingContext,
existingEmbeddings?: Map<string, string>,
): Promise<EmbeddingPipelineResult> => {
const finalConfig = resolveEmbeddingConfig(config);
@@ -317,21 +353,16 @@ export const runEmbeddingPipeline = async (
// Phase 2: Query embeddable nodes
let nodes = await queryEmbeddableNodes(executeQuery);
// Apply context metadata
if (context?.repoName) {
for (const node of nodes) {
node.repoName = context.repoName;
node.serverName = context.serverName;
}
}
// Incremental mode: compare content hashes, delete stale rows, skip fresh ones.
// Computed hashes for stale nodes are cached so batchInsertEmbeddings can reuse them
// (avoids double computation).
const computedStaleHashes = new Map<string, string>();
// Stale rows are DELETE'd per-batch (just before each batch's INSERT) rather
// than all up front — see U6 / KTD7. `staleNodeIds` is consulted inside the
// batch loop; it stays empty in full (non-incremental) mode so no deletes fire.
const staleNodeIds = new Set<string>();
if (existingEmbeddings && existingEmbeddings.size > 0) {
const beforeCount = nodes.length;
const staleNodeIds: string[] = [];
nodes = nodes.filter((n) => {
const existingHash = existingEmbeddings.get(n.id);
if (existingHash === undefined) {
@@ -342,40 +373,16 @@ export const runEmbeddingPipeline = async (
if (currentHash !== existingHash) {
// Content changed — cache hash for reuse during insert, mark for DELETE + re-embed
computedStaleHashes.set(n.id, currentHash);
staleNodeIds.push(n.id);
staleNodeIds.add(n.id);
return true;
}
// Hash matches — skip (fresh); no need to cache hash for skipped nodes
return false;
});
// DELETE stale embedding rows so they can be re-inserted
// (Kuzu forbids SET on vector-indexed properties; DELETE-then-INSERT is the sanctioned pattern)
if (staleNodeIds.length > 0) {
if (isDev) {
logger.info(`🔄 Deleting ${staleNodeIds.length} stale embedding rows for re-embed`);
}
try {
await executeWithReusedStatement(
`MATCH (e:${EMBEDDING_TABLE_NAME} {nodeId: $nodeId}) DELETE e`,
staleNodeIds.map((nodeId) => ({ nodeId })),
);
} catch (err) {
// "does not exist" = rows already gone — safe to proceed.
// All other errors risk vector-index corruption (Kuzu requires DELETE-before-INSERT
// for vector-indexed properties) — propagate so the pipeline aborts cleanly.
const msg = err instanceof Error ? err.message : String(err);
if (!msg.includes('does not exist')) {
throw new Error(
`[embed] Failed to delete stale embedding rows — aborting to prevent vector-index corruption: ${msg}`,
);
}
}
}
if (isDev) {
logger.info(
`📦 Incremental embeddings: ${beforeCount} total, ${existingEmbeddings.size} cached, ${staleNodeIds.length} stale, ${nodes.length} to embed`,
`📦 Incremental embeddings: ${beforeCount} total, ${existingEmbeddings.size} cached, ${staleNodeIds.size} stale, ${nodes.length} to embed`,
);
}
}
@@ -500,6 +507,12 @@ export const runEmbeddingPipeline = async (
}
}
// U6 / KTD7: delete this batch's stale rows immediately before its inserts,
// so an interrupted re-embed loses at most one batch (not the whole index).
// Preserves Kuzu's required DELETE-before-INSERT for vector-indexed rows.
const batchStaleIds = batch.filter((n) => staleNodeIds.has(n.id)).map((n) => n.id);
await deleteStaleEmbeddingRows(executeWithReusedStatement, batchStaleIds);
// Embed chunk texts in sub-batches to control memory
const EMBED_SUB_BATCH = finalConfig.subBatchSize;
for (let si = 0; si < allTexts.length; si += EMBED_SUB_BATCH) {
@@ -595,7 +608,7 @@ export const semanticSearch = async (
executeQuery: (cypher: string) => Promise<any[]>,
query: string,
k: number = 10,
maxDistance: number = 0.5,
maxDistance: number = getVectorMaxDistance(DEFAULT_VECTOR_MAX_DISTANCE),
): Promise<SemanticSearchResult[]> => {
if (!isEmbedderReady()) {
throw new Error('Embedding model not initialized. Run embedding pipeline first.');
@@ -741,7 +754,7 @@ export const semanticSearchWithContext = async (
k: number = 5,
_hops: number = 1,
): Promise<any[]> => {
const results = await semanticSearch(executeQuery, query, k, 0.5);
const results = await semanticSearch(executeQuery, query, k);
return results.map((r) => ({
matchId: r.nodeId,
+53 -25
View File
@@ -1,7 +1,7 @@
/**
* Text Generator Module
*
* Generates enriched embedding text from code nodes with metadata.
* Generates compact, description-forward embedding text from code nodes.
* Supports chunkable labels (Function/Method with AST chunking),
* Class-specific structural text, and short-node direct embed.
*
@@ -58,33 +58,51 @@ const cleanContent = (content: string): string => {
};
/**
* Build metadata header for a node
* Compact location signal for the embedding header: the last 1-2 path segments
* (immediate parent dir + basename), never the full deep path.
*
* #2333 / PR #2334 tri-review: U1 dropped the location entirely, which regressed
* path/service-qualified semantic search (e.g. `billing/handler` vs
* `identity/handler` in a monorepo) — and FTS indexes only name/content/description,
* never `filePath`, so there is no keyword backfill. The bounded form restores the
* discriminating tokens (service dir + filename-concept) at a fraction of the
* dilution the full path caused.
*/
const buildMetadataHeader = (node: EmbeddableNode, config: Partial<EmbeddingConfig>): string => {
const boundedLocation = (filePath: string): string => {
const segments = filePath.replace(/\\/g, '/').split('/').filter(Boolean);
return segments.slice(-2).join('/');
};
/**
* Build a compact, description-forward header for embedding text.
*
* Issue #2333 (sub-issue of #2326), Option A: lead the embedding text with the
* symbol name + doc-comment description and drop the low-signal metadata lines
* (`Repo`/`Server`/`Export` and the verbose full `Path`). For short doc comments
* those lines used to be ~25-30% of the embedding text, diluting the description's
* semantic weight in the vector and weakening description-shaped search — worst
* for CJK, where a complete concept is often 4-20 characters.
*
* A *bounded* location signal (last 1-2 path segments) is kept after the
* description — see `boundedLocation` for why the full path drop was reversed.
*
* Full metadata is unaffected: it lives on the graph node properties, which is
* what display/context tools read. Only the embedding text changes here.
*
* Option B (reorder only, keep metadata) was rejected — mean-pooled embeddings
* weight by token proportion, not position, so reordering alone barely moves the
* signal. Option C (a separate description-only embedding + hybrid merge) is
* deferred to follow-up; build it only if Option A proves insufficient against
* real measurement. Any change to this template MUST bump EMBEDDING_TEXT_VERSION.
*/
const buildEmbeddingHeader = (node: EmbeddableNode, config: Partial<EmbeddingConfig>): string => {
const parts: string[] = [];
// Label + name
parts.push(`${node.label}: ${node.name}`);
// Repo name
if (node.repoName) {
parts.push(`Repo: ${node.repoName}`);
}
// Server name (optional)
if (node.serverName) {
parts.push(`Server: ${node.serverName}`);
}
// Full file path
parts.push(`Path: ${node.filePath}`);
// Export status
if (node.isExported !== undefined) {
parts.push(`Export: ${node.isExported}`);
}
// Description (truncated)
// Description hoisted above everything else so its semantic signal dominates
// the embedding vector and is never the part lost to token-limit truncation.
if (node.description) {
const maxLen = config.maxDescriptionLength ?? DEFAULT_EMBEDDING_CONFIG.maxDescriptionLength;
const truncated = truncateDescription(node.description, maxLen);
@@ -93,6 +111,16 @@ const buildMetadataHeader = (node: EmbeddableNode, config: Partial<EmbeddingConf
}
}
// Bounded location signal — placed after the description so the description
// still leads the vector. Restores path/service disambiguation lost when the
// full Path line was dropped (FTS does not index filePath to backfill it).
if (node.filePath) {
const loc = boundedLocation(node.filePath);
if (loc) {
parts.push(`Loc: ${loc}`);
}
}
return parts.join('\n');
};
@@ -102,7 +130,7 @@ const generateCodeBodyText = (
config: Partial<EmbeddingConfig>,
prevTail?: string,
): string => {
const header = buildMetadataHeader(node, config);
const header = buildEmbeddingHeader(node, config);
const parts = [header];
if (prevTail) {
parts.push(`[preceding context]: ...${cleanContent(prevTail)}`);
@@ -128,7 +156,7 @@ const generateStructuralTypeText = (
chunkIndex?: number,
prevTail?: string,
): string => {
const header = buildMetadataHeader(node, config);
const header = buildEmbeddingHeader(node, config);
const parts: string[] = [header];
const isFirstChunk = chunkIndex === undefined || chunkIndex === 0;
const cleanedContent = cleanContent(node.content);
@@ -253,7 +281,7 @@ export const generateEmbeddingText = (
prevTail?: string,
): string => {
if (isShortLabel(node.label)) {
const header = buildMetadataHeader(node, config);
const header = buildEmbeddingHeader(node, config);
const cleaned = cleanContent(node.content);
return `${header}\n\n${cleaned}`;
}
-8
View File
@@ -289,14 +289,6 @@ export interface CachedEmbedding {
contentHash?: string;
}
/**
* Context info for embedding pipeline (repo/server metadata enrichment)
*/
export interface EmbeddingContext {
repoName?: string;
serverName?: string;
}
/**
* Model download progress from transformers.js
*/
+33 -5
View File
@@ -112,11 +112,14 @@ flowchart TD
EMIT --> BRIDGE[(bridge.lbug<br/>#795)]
```
Label-scoped queries in `resolveSymbol` keep accidental cross-matches
out:
- `topic` → `(n:Function|Method|Class|Interface)`
- `grpc` method → `(n:Function|Method)`, service → `(n:Class|Interface)`
- `lib` → `(n:Package|Module)`
Label-scoped queries in `resolveSymbol` keep accidental cross-matches out.
They use the `MATCH (n) WHERE labels(n) IN [...]` allowlist form, NOT the
`MATCH (n:A|B)` disjunction — LadybugDB's parser rejects a disjunction that
names a reserved keyword (e.g. `Macro`, `Union`), which is what broke the
`custom` branch in #2325:
- `topic` → `labels(n) IN ['Function','Method','Class','Interface']`
- `grpc`/`thrift` method → `labels(n) IN ['Function','Method']`, service → `labels(n) IN ['Class','Interface']`
- `lib` → `labels(n) IN ['Module']`
## Cross-impact query (PR #606)
@@ -137,3 +140,28 @@ The bridge stores every extracted contract keyed by `symbolUid`.
Manifest-sourced contracts use the synthetic uid form so both sides
of the `(local impact) ↔ (bridge query)` join derive the same uid
without coordinating through any shared state.
## Cross-repo trace (`cross-trace.ts`)
A second consumer of the bridge. Where cross-impact fans a blast radius
*outward* from one symbol, cross-trace stitches a directed **path** between
two symbols that live in different repos:
```mermaid
flowchart TD
FT[from / to resolved<br/>across all members] --> SR{same repo?}
SR -- yes --> LT[single-repo trace<br/>no crossing]
SR -- no --> SEGA[trace: from → consumer symbol<br/>in home repo]
SEGA --> XB[Bridge pair query<br/>consumer.symbolUid → provider.symbolUid<br/>one ContractLink boundary]
XB --> SEGB[trace: provider symbol → to<br/>in target repo]
SEGB --> STITCH[stitched hops + CONTRACT_LINK edge<br/>+ optional REACHING_DEF data-flow]
```
It reuses the same `symbolUid` join as cross-impact, but issues its own
*pair* query (`listCrossingsBetween`) because a path needs BOTH endpoints of
a crossing — the uid-filtered neighbor join (`resolveBridgeNeighbors`, shared
with impact) returns only the far side. The crossing is clamped to one
boundary (`MAX_SUPPORTED_CROSS_DEPTH`). With `pdg: true` the boundary-adjacent
segments are enriched with intra-procedural REACHING_DEF data-flow (never
across the boundary). Full cross-program data flow across the boundary is a
deferred follow-up.
+437 -34
View File
@@ -13,7 +13,9 @@ import {
import { dedupeContracts, dedupeCrossLinks } from './normalization.js';
import { createLogger } from '../logger.js';
const bridgeLogger = createLogger('bridge-db', { debugEnvVar: 'GITNEXUS_DEBUG_BRIDGE' });
const bridgeLogger = createLogger('bridge-db', {
debugEnvVar: 'GITNEXUS_DEBUG_BRIDGE',
});
/**
* Sidecar files that LadybugDB creates next to a `bridge.lbug` file.
@@ -33,6 +35,358 @@ const bridgeLogger = createLogger('bridge-db', { debugEnvVar: 'GITNEXUS_DEBUG_BR
*/
const LBUG_SIDECAR_SUFFIXES = ['.wal', '.shadow'] as const;
/* ------------------------------------------------------------------ */
/* Read-only bridge handle cache */
/* ------------------------------------------------------------------ */
/**
* Cache of read-only bridge handles keyed by groupDir. Keeps one RO handle
* per groupDir alive across @group tool calls so a long-lived MCP server
* never reopens the same bridge.lbug in-process — reopening fails on Windows
* because the OS file handle isn't fully released before the next open races
* in (see PR #2269, #2274).
*
* deliberation: mtime-based invalidation was chosen over a simpler
* time-to-live or explicit-close model because:
* 1. TTL would force a reopen on a timer even when nothing changed.
* 2. Explicit-close requires every caller to know about the cache.
* 3. A cheap `fsp.stat` (uncached, but typically a single inode lookup on
* modern kernels) before each `ensureBridgeReady` call detects external
* writers (e.g. another process ran group sync) with zero false
* positives and no timer complication.
* 4. Same-process writes invalidate explicitly via `invalidateBridgeCache`
* before the atomic rename so the cached RO handle does not block it.
*/
interface CachedBridgeEntry {
handle: BridgeHandle;
mtime: number;
/**
* Active leases: callers between `getCachedBridgeReadOnly` (acquire, `refs++`)
* and `closeBridgeDb` (release, `refs--`). The native handle is never closed
* while `refs > 0` — a concurrent `@group` reader may still be querying it,
* and closing under a live query is a native use-after-free.
*/
refs: number;
/** Set once the entry leaves the cache; the native close is deferred to the last release. */
evicted: boolean;
/** Guards `finalizeBridgeClose` so the native close runs exactly once. */
closeStarted: boolean;
/**
* Per-handle FIFO serialization tail. The cached RO handle is shared across
* concurrent `@group` callers, but a LadybugDB `Connection` is NOT safe for
* concurrent query execution (see `lbug/conn-lock.ts` — two queries on one
* connection corrupt the native heap). `queryBridge` runs each op on this
* chain so no two ever overlap on one handle. Per-handle (not a single global
* lock) so different groups — separate connections — stay parallel.
*/
lockTail: Promise<void>;
/**
* Resolves when the native handle has actually been closed. `writeBridge` on
* Windows awaits this (bounded — see `WINDOWS_DRAIN_TIMEOUT_MS`) before its
* atomic rename, because Windows cannot rename over an open handle. On POSIX
* the rename succeeds over an open RO handle (the old inode survives for the
* in-flight reader), so the close stays fully non-blocking there.
*/
drained: Promise<void>;
/** Resolver for {@link CachedBridgeEntry.drained}; called once by `finalizeBridgeClose`. */
resolveDrained: () => void;
}
/**
* Windows-only bound on how long `invalidateBridgeCache` waits for in-flight
* readers to release before letting `writeBridge` rename. Past this, it falls
* through and `retryRename` (EBUSY ×3) copes — so a pathologically long reader
* can never wedge `group_sync`. ponytail: fixed 5s ceiling; make it
* configurable if a real workload shows reads routinely outlasting it.
*/
const WINDOWS_DRAIN_TIMEOUT_MS = 5000;
const cachedBridgeHandles = new Map<string, CachedBridgeEntry>();
/**
* Reverse lookup: cache entry by its `BridgeHandle`. Lets `queryBridge` and
* `closeBridgeDb` find an entry from just the handle — including an *evicted*
* entry that is no longer in `cachedBridgeHandles` but whose native handle a
* lease still holds open. Uncached/writable handles (the `writeBridge` temp DB)
* are absent here, which is how those paths opt out of the lock and refcount.
*/
const bridgeEntryByHandle = new WeakMap<BridgeHandle, CachedBridgeEntry>();
/**
* In-flight opens keyed by groupDir. Prevents the TOCTOU race where two
* concurrent cache-miss calls both open a fresh handle and the second
* overwrites the first in `cachedBridgeHandles` — leaking the first
* handle. Mirrors the `local-backend.ts:1293` reinitPromises pattern.
*/
const inFlightOpens = new Map<string, Promise<BridgeHandle | null>>();
function bridgeCacheKey(groupDir: string): string {
return path.resolve(groupDir);
}
/**
* Serialize an operation on a cached handle's per-handle FIFO chain. Mirrors the
* promise-chain mechanic of `lbug/conn-lock.ts` (install a fresh unresolved
* tail, await the prior holder, release in `finally` so a throw never wedges the
* chain) — but keyed per handle, not a single global lock. No re-entry guard:
* `queryBridge` is a leaf (it never calls another locked bridge helper), and the
* native close runs outside the lock gated on `refs === 0`.
*/
export async function withHandleLock<T>(
lock: { lockTail: Promise<void> },
fn: () => Promise<T>,
): Promise<T> {
const prior = lock.lockTail;
let release!: () => void;
lock.lockTail = new Promise<void>((resolve) => {
release = resolve;
});
await prior;
try {
return await fn();
} finally {
release();
}
}
/**
* Close a cached entry's native handle exactly once. Guarded by `closeStarted`
* so the mtime-evict path, `invalidateBridgeCache`, the last lease release, and
* `closeAllCachedBridges` can all reach here and only one native close runs.
*/
async function finalizeBridgeClose(entry: CachedBridgeEntry): Promise<void> {
if (entry.closeStarted) return;
entry.closeStarted = true;
bridgeEntryByHandle.delete(entry.handle);
try {
await closeBridgeHandle(entry.handle);
} finally {
entry.resolveDrained();
}
}
/**
* Remove an entry from the cache and release its native handle. The native
* close is DEFERRED until in-flight leases drain (`refs === 0`): closing a
* handle a concurrent `@group` reader is still querying is a native
* use-after-free (the `conn-lock.ts` hazard). When `refs === 0` (the common
* single-threaded case — e.g. `group_sync` with no concurrent read) the close
* runs now and the returned promise resolves when it completes, so
* `writeBridge`'s atomic rename never races a live RO handle on Windows.
*
* When `refs > 0` (a concurrent reader holds a lease), the native close is
* deferred to the last `closeBridgeDb` release — closing now would be a
* use-after-free. Platform split for the rename that follows:
* - POSIX: return immediately. The rename succeeds over the still-open RO
* handle (old inode survives for the reader); no wait, no starvation.
* - Windows: a rename over an open handle fails (EBUSY), so wait — bounded by
* `WINDOWS_DRAIN_TIMEOUT_MS` — for the reader to release and the deferred
* close to complete, then the rename is clean. On timeout, fall through and
* let `retryRename` cope, so a slow reader can never wedge `group_sync`.
*
* This is the single eviction path for BOTH the mtime-change branch and
* `invalidateBridgeCache`.
*/
async function evictBridgeEntry(key: string, entry: CachedBridgeEntry): Promise<void> {
if (!entry.evicted) {
entry.evicted = true;
if (cachedBridgeHandles.get(key) === entry) cachedBridgeHandles.delete(key);
}
if (entry.refs <= 0) {
await finalizeBridgeClose(entry);
return;
}
// refs > 0: close deferred to the last closeBridgeDb release.
if (process.platform === 'win32') {
// Windows needs the handle closed before writeBridge renames. Wait (bounded)
// for readers to drain; on timeout, retryRename handles the residual EBUSY.
let timer: ReturnType<typeof setTimeout>;
const timeout = new Promise<void>((resolve) => {
timer = setTimeout(resolve, WINDOWS_DRAIN_TIMEOUT_MS);
});
await Promise.race([entry.drained, timeout]).finally(() => clearTimeout(timer));
}
}
/**
* Close a BridgeHandle's native resources without touching the cache.
* Shared by `closeBridgeDb` (uncached handles) and the cache invalidation
* / shutdown paths so neither duplicates the close logic.
*/
async function closeBridgeHandle(handle: BridgeHandle): Promise<void> {
if (!handle._readOnly) {
try {
await (handle._conn as lbug.Connection).query('CHECKPOINT');
} catch {
/* ignore — older LadybugDB or schemaless DB may not accept it */
}
}
try {
await (handle._conn as lbug.Connection).close();
} catch {
/* ignore */
}
try {
await (handle._db as lbug.Database).close();
} catch {
/* ignore */
}
}
/**
* Get or create a cached read-only bridge handle for `groupDir`.
*
* - First call: delegates to `openBridgeDbReadOnly`, records the file's
* `mtimeMs`, and caches the handle.
* - Subsequent calls (mtime unchanged): returns the cached handle — no
* reopen, no OS file-handle churn.
* - After the file's mtime changes (external writer, e.g. another process
* ran `gitnexus group sync`): closes the stale handle, opens a fresh
* one, and updates the cache.
* - After the file disappears (ENOENT): invalidates cache, returns null.
*
* Returns `null` when the bridge file is missing, has an incompatible
* schema version, or cannot be opened even after the retry loop in
* `openBridgeDbReadOnly`.
*/
export async function getCachedBridgeReadOnly(groupDir: string): Promise<BridgeHandle | null> {
const key = bridgeCacheKey(groupDir);
const dbPath = path.join(groupDir, 'bridge.lbug');
// Fast path: cache hit, unchanged mtime → lease the cached handle.
const entry = cachedBridgeHandles.get(key);
if (entry) {
try {
const stat = await fsp.stat(dbPath);
// Re-check `evicted` AFTER the await: a concurrent writeBridge/invalidate
// may have evicted this entry while we awaited `stat`. Leasing an evicted
// (closing) handle would be a use-after-close. The `refs++` is the first
// synchronous statement after the check, so no evictor can slip between.
if (!entry.evicted && stat.mtimeMs === entry.mtime) {
entry.refs++;
return entry.handle;
}
} catch {
// File disappeared (ENOENT) — fall through to evict + reopen.
}
// mtime changed or file gone — evict (defers the native close if a
// concurrent reader still holds a lease; closes now otherwise).
if (!entry.evicted) await evictBridgeEntry(key, entry);
}
// TOCTOU guard: if another caller is already opening for this key, await
// their in-flight promise and take a lease on the result instead of opening
// a second handle.
const inFlight = inFlightOpens.get(key);
if (inFlight) {
const handle = await inFlight;
if (!handle) return null;
// Same post-await guard as the fast path: the opener's entry may have been
// evicted between caching and this awaiter resuming. Only lease a live,
// identity-matched entry; otherwise retry from the top for a fresh handle.
const opened = cachedBridgeHandles.get(key);
if (opened && !opened.evicted && opened.handle === handle) {
opened.refs++;
return handle;
}
return getCachedBridgeReadOnly(groupDir);
}
const openPromise: Promise<BridgeHandle | null> = (async () => {
try {
const handle = await openBridgeDbReadOnly(groupDir);
if (!handle) return null;
let mtime = 0;
try {
const stat = await fsp.stat(dbPath);
mtime = stat.mtimeMs;
} catch {
// bridge.lbug not stat-able right after open (rare race). Leaving
// mtime at 0 means the next call's fast-path comparison won't match
// (a real file's mtime is never 0), so it re-opens. Benign: the handle
// still works for this caller; we just don't cache-reuse it until a
// later open records a real mtime.
}
let resolveDrained!: () => void;
const drained = new Promise<void>((resolve) => {
resolveDrained = resolve;
});
const newEntry: CachedBridgeEntry = {
handle,
mtime,
refs: 0,
evicted: false,
closeStarted: false,
lockTail: Promise.resolve(),
drained,
resolveDrained,
};
cachedBridgeHandles.set(key, newEntry);
bridgeEntryByHandle.set(handle, newEntry);
return handle;
} finally {
inFlightOpens.delete(key);
}
})();
inFlightOpens.set(key, openPromise);
// Each caller (the opener and every awaiter) takes exactly one lease here, so
// refs counts callers correctly even under inFlightOpens coalescing.
const handle = await openPromise;
if (!handle) return null;
const opened = cachedBridgeHandles.get(key);
if (opened && !opened.evicted && opened.handle === handle) {
opened.refs++;
return handle;
}
return getCachedBridgeReadOnly(groupDir);
}
/**
* Invalidate the cached read-only handle for `groupDir`. Drops it from the
* cache immediately; the native close is deferred until any in-flight reader
* leases drain (see {@link evictBridgeEntry}). With no concurrent reader this
* resolves only after the handle is actually closed — which is why
* `writeBridge` awaits it before its atomic rename (Windows: a still-open RO
* handle would block the rename with EBUSY).
*/
export async function invalidateBridgeCache(groupDir: string): Promise<void> {
const key = bridgeCacheKey(groupDir);
const entry = cachedBridgeHandles.get(key);
if (entry) await evictBridgeEntry(key, entry);
}
/**
* Close ALL cached bridge handles. Call on process shutdown only — it force-
* closes regardless of refs (safe at `beforeExit`, which fires only at
* event-loop quiescence, so no query is in flight). Do NOT wire this to a
* SIGTERM/SIGINT handler that can fire mid-request: that would close a handle
* under a live query. Routes through `finalizeBridgeClose` for the close-once
* guarantee.
*/
export async function closeAllCachedBridges(): Promise<void> {
const entries = [...cachedBridgeHandles.values()];
cachedBridgeHandles.clear();
await Promise.all(entries.map((e) => finalizeBridgeClose(e)));
}
// Best-effort process-exit cleanup. 'beforeExit' fires before 'exit' and
// lets async work drain (unlike 'exit' which is synchronous-only). It does
// NOT fire on process.exit()/SIGTERM/SIGINT — but that is fine here: the OS
// reclaims all handles on any exit path, and for read-only handles there is
// no WAL to flush, so the only thing lost on signal death is a tidy close
// (cosmetic). We deliberately do NOT register a SIGTERM/SIGINT handler: a
// signal can fire mid-request, and closeAllCachedBridges force-closes
// regardless of refs, which would close a handle under a live query. Shutdown
// sequencing is the MCP server's responsibility — it should call
// closeAllCachedBridges() at a quiescent point (also how tests get a
// deterministic teardown).
process.once('beforeExit', () => {
void closeAllCachedBridges();
});
async function removeLbugFile(basePath: string): Promise<void> {
const candidates = [basePath, ...LBUG_SIDECAR_SUFFIXES.map((s) => `${basePath}${s}`)];
for (const f of candidates) {
@@ -195,20 +549,29 @@ export async function queryBridge<T>(
cypher: string,
params?: Record<string, LbugValue>,
): Promise<T[]> {
const conn = handle._conn as lbug.Connection;
if (params && Object.keys(params).length > 0) {
const stmt = await conn.prepare(cypher);
if (!stmt.isSuccess()) {
const errMsg = await stmt.getErrorMessage();
throw new Error(`Bridge query prepare failed: ${errMsg}`);
const run = async (): Promise<T[]> => {
const conn = handle._conn as lbug.Connection;
if (params && Object.keys(params).length > 0) {
const stmt = await conn.prepare(cypher);
if (!stmt.isSuccess()) {
const errMsg = await stmt.getErrorMessage();
throw new Error(`Bridge query prepare failed: ${errMsg}`);
}
const queryResult = await conn.execute(stmt, params);
const result = unwrapQueryResult(queryResult);
return (await result.getAll()) as T[];
}
const queryResult = await conn.execute(stmt, params);
const queryResult = await conn.query(cypher);
const result = unwrapQueryResult(queryResult);
return (await result.getAll()) as T[];
}
const queryResult = await conn.query(cypher);
const result = unwrapQueryResult(queryResult);
return (await result.getAll()) as T[];
};
// Cached RO handles are shared across concurrent @group callers, so serialize
// conn ops per handle (a LadybugDB Connection is not safe for concurrent
// queries — conn-lock.ts). Uncached/writable handles (the writeBridge temp DB)
// are single-threaded — they're absent from bridgeEntryByHandle and skip the
// lock at zero cost.
const entry = bridgeEntryByHandle.get(handle);
return entry ? withHandleLock(entry, run) : run();
}
/**
@@ -230,30 +593,54 @@ function unwrapQueryResult(queryResult: lbug.QueryResult | lbug.QueryResult[]):
return queryResult;
}
/**
* Release a caller's reference to a bridge handle.
*
* - **Cache-owned handle** (returned by `getCachedBridgeReadOnly`): this is the
* matching *release* for that acquire — it decrements the lease refcount, it
* does NOT close the native handle. The cache owns the lifetime; the handle
* closes on explicit `invalidateBridgeCache`, mtime-eviction, or process
* shutdown. If the entry was already evicted and this is the last lease, the
* deferred native close fires here (exactly once).
* - **Uncached/writable handle** (e.g. the `writeBridge` temp DB): closes the
* native handle for real (CHECKPOINT-flush for writable handles).
*
* Contract: before renaming or deleting `bridge.lbug`, call
* `invalidateBridgeCache` (not this) — `closeBridgeDb` on a cache-owned handle
* is a lease release, so the file may stay open under other readers.
*/
export async function closeBridgeDb(handle: BridgeHandle): Promise<void> {
// CHECKPOINT before close so the WAL/.shadow contents are flushed into
// the main database file. Without this, LadybugDB 0.16.0's non-blocking
// checkpoint thread can outlive the close call and leave sidecar pages
// pending on disk, which makes a subsequent read-side open either race
// with the WAL replay or trip the database-id check on the sidecars.
// CHECKPOINT is a no-op when there's nothing pending, so it's cheap.
try {
await (handle._conn as lbug.Connection).query('CHECKPOINT');
} catch {
/* ignore — older LadybugDB or schemaless DB may not accept it */
}
try {
await (handle._conn as lbug.Connection).close();
} catch {
/* ignore */
}
try {
await (handle._db as lbug.Database).close();
} catch {
/* ignore */
const entry = bridgeEntryByHandle.get(handle);
if (!entry) {
// Uncached or writable handle — close for real.
await closeBridgeHandle(handle);
return;
}
// Cache-owned handle: release this lease. Close only the evicted handle whose
// last lease just dropped (deferred-close completion); the live cached handle
// stays open for reuse.
if (entry.refs > 0) entry.refs--;
if (entry.evicted && entry.refs <= 0) await finalizeBridgeClose(entry);
}
// NOTE: Windows in-process write→read reopen of the SAME bridge.lbug is still a
// known limitation (the writable close's OS file handle is not released before
// the read open races; the existing open-side LBUG_OPEN_RETRY only retries
// lock-pattern errors, not the post-rename sidecar database-id mismatch). The
// bridge's close-then-reopen tests stay Windows-skipped. A close-side
// waitForWindowsHandleRelease + finalizeLbugSidecarsAfterClose probe (mirroring
// safeClose) was tried and did NOT close that gap on Windows CI, so it was
// removed rather than carry latency/duplication for no Windows benefit.
//
// Scope of the RO bridge-handle cache (getCachedBridgeReadOnly): it removes the
// PRODUCTION symptom — a long-lived MCP serve process reopening bridge.lbug on
// every @group call — by keeping one RO handle alive for read→READ reuse.
// It does NOT fix the write→READ reopen: the first @group read right after an
// in-process group_sync is a cache miss → openBridgeDbReadOnly, i.e. the same
// unfixed reopen, so on Windows that first post-sync read still returns null.
// The read-only CHECKPOINT skip above remains the load-bearing fix on
// Linux/macOS.
/* ------------------------------------------------------------------ */
/* retryRename — handles transient EBUSY/EPERM/EACCES on Windows */
/* ------------------------------------------------------------------ */
@@ -361,6 +748,13 @@ export async function writeBridge(
input: WriteBridgeInput,
): Promise<WriteBridgeReport> {
await fsp.mkdir(groupDir, { recursive: true });
// Invalidate the RO cache before writing. On Windows the cached handle
// would block the atomic rename (tmp → bridge.lbug) because the OS keeps
// a shared-mode lock on the open file. Closing it first guarantees the
// rename succeeds without EBUSY.
await invalidateBridgeCache(groupDir);
const contracts = dedupeContracts(input.contracts);
const crossLinks = dedupeCrossLinks(input.crossLinks);
@@ -713,7 +1107,12 @@ export async function openBridgeDbReadOnly(groupDir: string): Promise<BridgeHand
// (where we can retry) instead of on the first user query.
await handle.db.init();
await handle.conn.init();
return { _db: handle.db, _conn: handle.conn, groupDir } as BridgeHandle;
return {
_db: handle.db,
_conn: handle.conn,
groupDir,
_readOnly: true,
} as BridgeHandle;
} catch (err) {
lastErr = err;
if (handle) await closeLbugConnection(handle);
@@ -730,7 +1129,11 @@ export async function openBridgeDbReadOnly(groupDir: string): Promise<BridgeHand
const safeErrMsg =
lastErr instanceof Error ? String(lastErr.message).replace(/[\r\n]/g, ' ') : undefined;
bridgeLogger.debug(
{ groupDir: safeGroupDir, errMsg: safeErrMsg, attempts: LBUG_OPEN_RETRY_ATTEMPTS },
{
groupDir: safeGroupDir,
errMsg: safeErrMsg,
attempts: LBUG_OPEN_RETRY_ATTEMPTS,
},
'openBridgeDbReadOnly gave up',
);
return null;
+47 -13
View File
@@ -22,7 +22,12 @@ import {
repoInSubgroup,
} from './group-path-utils.js';
import { getGroupDir } from './storage.js';
import { closeBridgeDb, openBridgeDbReadOnly, queryBridge, readBridgeMeta } from './bridge-db.js';
import {
closeBridgeDb,
getCachedBridgeReadOnly,
queryBridge,
readBridgeMeta,
} from './bridge-db.js';
import { BRIDGE_SCHEMA_VERSION } from './bridge-schema.js';
// High limit for the local phase of group impact so collectImpactSymbolUids
@@ -63,7 +68,7 @@ RETURN provider.repo AS neighborRepo,
provider.type AS contractType
`;
type BridgeNeighborRow = {
export type BridgeNeighborRow = {
neighborRepo: string;
neighborUid: string;
neighborFilePath?: string;
@@ -352,7 +357,7 @@ export function mergeRisk(localRisk: string, cross: CrossRepoImpact[]): string {
return localRisk;
}
async function ensureBridgeReady(
export async function ensureBridgeReady(
groupDir: string,
): Promise<{ handle: BridgeHandle } | { error: string }> {
const meta = await readBridgeMeta(groupDir);
@@ -369,7 +374,10 @@ async function ensureBridgeReady(
error: `No bridge.lbug in this group directory. Run gitnexus group sync (schema ${BRIDGE_SCHEMA_VERSION}).`,
};
}
const handle = await openBridgeDbReadOnly(groupDir);
// Use the cached read-only handle if available — avoids reopening the same
// bridge.lbug in a long-lived MCP server, which fails on Windows because
// the OS handle isn't fully released before the next open races in.
const handle = await getCachedBridgeReadOnly(groupDir);
if (!handle) {
return {
error: `Could not open bridge.lbug read-only (schema ${BRIDGE_SCHEMA_VERSION}). Run gitnexus group sync.`,
@@ -394,6 +402,39 @@ function rowToNeighbor(r: Record<string, unknown>): BridgeNeighborRow | null {
};
}
/**
* Resolve cross-repo neighbors over `ContractLink` for a set of local symbol
* UIDs, in a single direction, sorted by descending confidence.
*
* This is the one shared consumer↔provider bridge join. `runGroupImpact`'s
* Phase-2 fan-out uses it directly; the cross-repo trace path (`cross-trace.ts`)
* reuses the same `queryBridge` + row-normalization primitives but issues a
* distinct *pair* query, because a trace must keep BOTH endpoints of a crossing
* (this neighbor join intentionally returns only the far side, which is lossy
* for stitching a path). Keeping this helper as the single uid-filtered join
* means impact never forks its own copy of the neighbor Cypher.
*
* Returns `[]` for an empty `uids` set without touching the DB.
*/
export async function resolveBridgeNeighbors(
handle: BridgeHandle,
opts: { localRepo: string; uids: string[]; direction: 'upstream' | 'downstream' },
): Promise<BridgeNeighborRow[]> {
if (opts.uids.length === 0) return [];
const cypher = opts.direction === 'upstream' ? CY_NEIGHBORS_UPSTREAM : CY_NEIGHBORS_DOWNSTREAM;
const rows = await queryBridge<Record<string, unknown>>(handle, cypher, {
localRepo: opts.localRepo,
uids: opts.uids,
});
const neighbors: BridgeNeighborRow[] = [];
for (const raw of rows) {
const n = rowToNeighbor(raw);
if (n) neighbors.push(n);
}
neighbors.sort((a, b) => b.confidence - a.confidence);
return neighbors;
}
export async function runGroupImpact(
deps: RunGroupImpactDeps,
params: Record<string, unknown>,
@@ -537,19 +578,12 @@ export async function runGroupImpact(
const truncatedRepos: string[] = [];
try {
const cypher = direction === 'upstream' ? CY_NEIGHBORS_UPSTREAM : CY_NEIGHBORS_DOWNSTREAM;
const rows = await queryBridge<Record<string, unknown>>(handle, cypher, {
const neighbors = await resolveBridgeNeighbors(handle, {
localRepo: repoPath,
uids,
direction,
});
const neighbors: BridgeNeighborRow[] = [];
for (const raw of rows) {
const n = rowToNeighbor(raw);
if (n) neighbors.push(n);
}
neighbors.sort((a, b) => b.confidence - a.confidence);
const seen = new Set<string>();
for (const n of neighbors) {
File diff suppressed because it is too large Load Diff
@@ -17,8 +17,11 @@ import type { HttpDetection, HttpLanguagePlugin } from './types.js';
// ─── Provider: framework routing ──────────────────────────────────────
// Matches `\w+\.GET(...)` etc. (gin, echo, chi all share this shape).
// Captures the HTTP method (field name), path literal, and handler
// identifier passed as the second argument.
// Captures the HTTP method (field name), path literal, and the handler —
// anchored to the LAST argument (`@handler .`) so a variadic middleware
// chain (`r.GET("/x", mw, handler)`, gin/echo/chi style) binds the real
// handler, not a middleware identifier (which would otherwise over-match
// and attach the route to the wrong symbol — see #2276 review).
const FRAMEWORK_ROUTE_PATTERNS = compilePatterns({
name: 'go-framework-route',
language: Go,
@@ -31,7 +34,8 @@ const FRAMEWORK_ROUTE_PATTERNS = compilePatterns({
field: (field_identifier) @http_method (#match? @http_method "^(GET|POST|PUT|DELETE|PATCH)$"))
arguments: (argument_list
(interpreted_string_literal) @path
(identifier) @handler))
[(identifier) (func_literal)] @handler
.))
`,
},
],
@@ -51,7 +55,8 @@ const HANDLE_FUNC_PATTERNS = compilePatterns({
field: (field_identifier) @fn (#eq? @fn "HandleFunc"))
arguments: (argument_list
(interpreted_string_literal) @path
(identifier) @handler))
[(identifier) (func_literal)] @handler
.))
`,
},
],
@@ -138,12 +143,18 @@ export const GO_HTTP_PLUGIN: HttpLanguagePlugin = {
if (!methodNode || !pathNode) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
// An inline `func(){…}` handler has no name → emit `name: null` and a
// `line` so it resolves to its containing/closure symbol by line-span
// containment (like a consumer). A named identifier handler keeps its
// name and resolves by name; `line` is harmless there.
const isInlineHandler = handlerNode?.type === 'func_literal';
out.push({
role: 'provider',
framework: 'go-framework',
method: methodNode.text.toUpperCase(),
path,
name: handlerNode?.text ?? null,
name: isInlineHandler ? null : (handlerNode?.text ?? null),
line: (handlerNode ?? pathNode).startPosition.row + 1,
confidence: 0.8,
});
}
@@ -155,12 +166,16 @@ export const GO_HTTP_PLUGIN: HttpLanguagePlugin = {
if (!pathNode) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
// Inline `func(){…}` handler → resolve by containment (see go-framework
// note above); a named handler resolves by name.
const isInlineHandler = handlerNode?.type === 'func_literal';
out.push({
role: 'provider',
framework: 'go-stdlib',
method: 'GET',
path,
name: handlerNode?.text ?? null,
name: isInlineHandler ? null : (handlerNode?.text ?? null),
line: (handlerNode ?? pathNode).startPosition.row + 1,
confidence: 0.8,
});
}
@@ -180,6 +195,7 @@ export const GO_HTTP_PLUGIN: HttpLanguagePlugin = {
method: httpMethod,
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -198,6 +214,7 @@ export const GO_HTTP_PLUGIN: HttpLanguagePlugin = {
method: method.toUpperCase(),
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -215,6 +232,7 @@ export const GO_HTTP_PLUGIN: HttpLanguagePlugin = {
method: methodNode.text.toUpperCase(),
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -10,6 +10,8 @@ import {
METHOD_ANNOTATION_TO_HTTP,
isRouteMemberKey,
findEnclosingClass,
joinPath,
type SharedSpringType,
} from '../../../ingestion/route-extractors/spring-shared.js';
import {
REST_TEMPLATE_TO_HTTP,
@@ -18,9 +20,7 @@ import {
EXCHANGE_ANNOTATION_TO_HTTP,
parseRequestLine,
pushPrefix,
joinPath,
scanSpringInheritanceProject,
type SharedSpringType,
OPENFEIGN_FRAMEWORK,
HTTP_INTERFACE_FRAMEWORK,
FEIGN_CONFIDENCE,
@@ -676,6 +676,30 @@ function scanSpringProject(files: readonly HttpScanInput[]): HttpFileDetections[
export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
name: 'java-http',
language: Java,
// routeCoverage intentionally LEFT at the default 'partial' (#2138 Part 2).
// The graph provider set is a strict *subset* of this scan()'s provider set —
// ingestion does NOT emit a Route node for a method-level array route nested
// under a class-level array-form `@RequestMapping` (ingestion suppresses it
// rather than drop the prefix; bare/scalar-prefixed array methods ARE now
// emitted — see #2280). Interface-inherited Spring routes ARE now emitted by
// ingestion (#2288), and same-URL multi-verb routes are now per-`(method,url)`
// Route nodes (#2289), so they are no longer coverage gaps. Declaring
// 'complete' here would let the parse-skip drop the remaining group-only
// providers (the array-prefix gap above). Java flips to 'complete' only once
// ingestion provider extraction matches this scan — class-level array-form
// prefix support is the final follow-up tracked in #2280.
// `hasConsumerSignals` below is kept ready for that flip.
// Consumer signals this plugin's scan() can detect: RestTemplate / WebClient /
// OkHttp / Java-HttpClient / Apache-HttpClient call sites, OpenFeign
// (`@FeignClient` + `@RequestLine`) interfaces, and Spring 6 HTTP Interface
// `@(Get|...)Exchange` / `@HttpExchange`. A provider-covered file containing
// any of these must still be parsed so its consumer contracts are not dropped
// (ingestion emits no FETCHES for Java). Conservative by design.
hasConsumerSignals(content) {
return /\brestTemplate\b|\bwebClient\b|Request\.Builder|HttpRequest|HttpMethod\.|new\s+Http(Get|Post|Put|Delete|Patch)\b|@RequestLine|@FeignClient|Exchange/.test(
content,
);
},
scan(tree) {
const out: HttpDetection[] = [];
@@ -708,6 +732,7 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
method: route.httpMethod,
path: joinPath(prefix, route.rawPath),
name: route.methodName,
line: route.methodNode.startPosition.row + 1,
confidence: FEIGN_CONFIDENCE,
});
}
@@ -725,6 +750,13 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
method: route.httpMethod,
path: joinPath(prefix, route.rawPath),
name: route.methodName,
// Spring providers are named controller methods resolved BY NAME, so
// `line` is inert — a named provider never falls through to line-span
// containment. Gate it on a present name so a (grammar-impossible)
// nameless provider degrades to file-level rather than resolving by
// containment to the enclosing class. Wired for consumer-emit parity
// and a future inline DSL.
line: route.methodName ? route.methodNode.startPosition.row + 1 : undefined,
confidence: 0.8,
});
}
@@ -751,6 +783,7 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
method: requestLine.parsed.method,
path: joinPath(prefix, requestLine.parsed.path),
name: requestLine.methodName,
line: requestLine.methodNode.startPosition.row + 1,
confidence: REQUEST_LINE_CONFIDENCE,
});
}
@@ -772,6 +805,7 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
method: route.httpMethod,
path: joinPath(prefix, route.rawPath),
name: route.methodName,
line: route.methodNode.startPosition.row + 1,
confidence: EXCHANGE_CONFIDENCE,
});
}
@@ -792,6 +826,7 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
method: httpMethod,
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -808,6 +843,7 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
method: httpMethodNode.text.toUpperCase(),
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -830,6 +866,7 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
method: httpMethod,
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -853,6 +890,7 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
method: verbText,
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -879,6 +917,7 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
method,
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -15,6 +15,8 @@ import type {
import {
METHOD_ANNOTATION_TO_HTTP,
findEnclosingClass,
joinPath,
type SharedSpringType,
} from '../../../ingestion/route-extractors/spring-shared.js';
import {
REST_TEMPLATE_TO_HTTP,
@@ -23,9 +25,7 @@ import {
EXCHANGE_ANNOTATION_TO_HTTP,
parseRequestLine,
pushPrefix,
joinPath,
scanSpringInheritanceProject,
type SharedSpringType,
OPENFEIGN_FRAMEWORK,
HTTP_INTERFACE_FRAMEWORK,
FEIGN_CONFIDENCE,
@@ -996,6 +996,7 @@ function buildKotlinPlugin(language: unknown): HttpLanguagePlugin {
method: httpMethod,
path: joinPath(prefix, rawPath),
name: nameNode?.text ?? null,
line: methodNode.startPosition.row + 1,
confidence: FEIGN_CONFIDENCE,
});
}
@@ -1018,6 +1019,13 @@ function buildKotlinPlugin(language: unknown): HttpLanguagePlugin {
method: httpMethod,
path: joinPath(prefix, rawPath),
name: nameNode?.text ?? null,
// Spring providers are named controller methods resolved BY NAME, so
// `line` is inert — a named provider never falls through to line-span
// containment. Gate it on a present name so a (grammar-impossible)
// nameless provider degrades to file-level rather than resolving by
// containment to the enclosing class. Wired for consumer-emit parity
// and a future inline DSL.
line: nameNode?.text ? methodNode.startPosition.row + 1 : undefined,
confidence: 0.8,
});
}
@@ -1038,6 +1046,7 @@ function buildKotlinPlugin(language: unknown): HttpLanguagePlugin {
method: httpMethod,
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -1057,6 +1066,7 @@ function buildKotlinPlugin(language: unknown): HttpLanguagePlugin {
method: httpMethod,
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -1084,6 +1094,7 @@ function buildKotlinPlugin(language: unknown): HttpLanguagePlugin {
method: verbText,
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -1108,6 +1119,7 @@ function buildKotlinPlugin(language: unknown): HttpLanguagePlugin {
method,
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -1134,6 +1146,7 @@ function buildKotlinPlugin(language: unknown): HttpLanguagePlugin {
method: httpMethod,
path: joinPath(prefix, rawPath),
name: nameNode?.text ?? null,
line: methodNode.startPosition.row + 1,
confidence: EXCHANGE_CONFIDENCE,
});
}
@@ -65,7 +65,7 @@ const EXPRESS_SPEC: PatternSpec<Record<string, never>> = {
function: (member_expression
object: (identifier) @obj (#match? @obj "^(router|app)$")
property: (property_identifier) @http_method (#match? @http_method "^(get|post|put|delete|patch)$"))
arguments: (arguments . [(string) (template_string)] @path))
arguments: (arguments . [(string) (template_string)] @path . (_)? @handler))
`,
};
@@ -295,8 +295,53 @@ function findDecoratedMethod(decoratorNode: Parser.SyntaxNode): Parser.SyntaxNod
return null;
}
/**
* Map each named import's LOCAL binding to its DECLARED export name and source
* module, by walking the file's `import { x as y } from 'm'` statements. Lets
* the express handler resolve through an alias (the local `y`) to the real
* symbol (`x` in `m`) instead of looking up the alias text. Only named imports
* are mapped — default and namespace imports are left to fall through as
* locally-scoped identifiers.
*/
function buildImportMap(tree: Parser.Tree): Map<string, { name: string; module: string }> {
const map = new Map<string, { name: string; module: string }>();
const walk = (node: Parser.SyntaxNode): void => {
if (node.type === 'import_statement') {
const sourceNode = node.childForFieldName('source');
const module = sourceNode ? unquoteLiteral(sourceNode.text) : null;
if (module !== null) {
const collect = (n: Parser.SyntaxNode): void => {
if (n.type === 'import_specifier') {
const nameNode = n.childForFieldName('name');
const aliasNode = n.childForFieldName('alias');
const local = aliasNode ?? nameNode;
if (nameNode && local && local.type === 'identifier') {
map.set(local.text, { name: nameNode.text, module });
}
}
for (let i = 0; i < n.namedChildCount; i++) {
const c = n.namedChild(i);
if (c) collect(c);
}
};
collect(node);
}
}
for (let i = 0; i < node.namedChildCount; i++) {
const c = node.namedChild(i);
if (c) walk(c);
}
};
walk(tree.rootNode);
return map;
}
function scanBundle(bundle: NodePatternBundle, tree: Parser.Tree): HttpDetection[] {
const out: HttpDetection[] = [];
// Local-binding → { declared export name, module } for the file's named
// imports, so an express handler that is an imported (possibly aliased)
// symbol resolves to the real definition rather than its local alias text.
const importMap = buildImportMap(tree);
// NestJS: collect `@Controller('prefix')` class decorators, keyed by
// the `class_declaration` they decorate.
@@ -348,6 +393,7 @@ function scanBundle(bundle: NodePatternBundle, tree: Parser.Tree): HttpDetection
method: httpMethod,
path: joinPath(prefix, rawPath),
name,
line: methodNode.startPosition.row + 1,
confidence: 0.8,
});
}
@@ -359,12 +405,24 @@ function scanBundle(bundle: NodePatternBundle, tree: Parser.Tree): HttpDetection
if (!methodNode || !pathNode) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
// Capture the handler argument identifier (`router.get('/x', listUsers)`
// → `listUsers`) so a named handler resolves by name. For an inline/anonymous
// handler emit `name: null` (NOT the sentinel `'handler'`) so the resolver
// does NOT match an unrelated function that happens to be named `handler` —
// it uses the registration line for containment instead. When the handler is
// an imported (possibly aliased) symbol, carry the resolved import so the
// extractor can pin it to the source module rather than the local alias text.
const handlerNode = match.captures.handler;
const localHandler = handlerNode?.type === 'identifier' ? handlerNode.text : null;
const imported = localHandler !== null ? importMap.get(localHandler) : undefined;
out.push({
role: 'provider',
framework: 'express',
method: methodNode.text.toUpperCase(),
path,
name: 'handler',
name: imported ? imported.name : localHandler,
handlerImport: imported,
line: (handlerNode ?? pathNode).startPosition.row + 1,
confidence: 0.8,
});
}
@@ -385,6 +443,7 @@ function scanBundle(bundle: NodePatternBundle, tree: Parser.Tree): HttpDetection
method: method.toUpperCase(),
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -403,6 +462,7 @@ function scanBundle(bundle: NodePatternBundle, tree: Parser.Tree): HttpDetection
method: 'GET',
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -420,6 +480,7 @@ function scanBundle(bundle: NodePatternBundle, tree: Parser.Tree): HttpDetection
method: methodNode.text.toUpperCase(),
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -437,6 +498,7 @@ function scanBundle(bundle: NodePatternBundle, tree: Parser.Tree): HttpDetection
method: methodNode.text.toUpperCase(),
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -456,6 +518,7 @@ function scanBundle(bundle: NodePatternBundle, tree: Parser.Tree): HttpDetection
method,
path,
name: null,
line: optionsNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -476,6 +539,7 @@ function scanBundle(bundle: NodePatternBundle, tree: Parser.Tree): HttpDetection
method,
path,
name: null,
line: optionsNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -37,7 +37,9 @@ const LARAVEL_ROUTE_SPEC: PatternSpec<Record<string, never>> = {
(scoped_call_expression
scope: (name) @scope (#eq? @scope "Route")
name: (name) @method (#match? @method "^(get|post|put|delete|patch)$")
arguments: (arguments . (argument (string) @path)))
arguments: (arguments
. (argument (string) @path)
(argument [(anonymous_function) (arrow_function)] @closure)?))
`,
};
@@ -130,6 +132,17 @@ function isHttpUrlLiteral(path: string): boolean {
export const PHP_HTTP_PLUGIN: HttpLanguagePlugin = {
name: 'php-http',
language: PHP.php_only,
// Laravel `Route::<verb>(...)` definitions are emitted as Route nodes by
// ingestion, so the graph is authoritative for PHP providers (#2138 Part 2).
routeCoverage: 'complete',
// Consumer signals scan() can detect: Laravel `Http::<verb>`, Guzzle client
// `->get/post/.../request(...)`, and `file_get_contents` of an HTTP URL. A
// provider-covered file with any of these must still be parsed (ingestion
// emits no FETCHES for PHP). Conservative — the `->verb(` shape over-matches
// ordinary method calls, which only costs a parse, never data.
hasConsumerSignals(content) {
return /Http::|file_get_contents|->\s*(get|post|put|delete|patch|request)\s*\(/i.test(content);
},
scan(tree) {
const out: HttpDetection[] = [];
@@ -139,12 +152,22 @@ export const PHP_HTTP_PLUGIN: HttpLanguagePlugin = {
if (!methodNode || !pathNode) continue;
const path = phpStringText(pathNode);
if (path === null) continue;
// A closure handler (`Route::get('/x', function(){…})` / `fn() => …`) has
// no name → emit `name: null` + the registration line so it resolves to
// its containing symbol (e.g. a service-provider `boot()` or controller
// method) by line-span containment. A named-controller route keeps the
// `'route'` label — resolving its array/string handler to a real method is
// a separate, graph-backed concern. NOTE: a closure at FILE scope
// (routes/web.php) has no enclosing function and PHP closures are not yet
// indexed as symbols, so it still degrades to file-level (see #2276).
const closureNode = match.captures.closure;
out.push({
role: 'provider',
framework: 'laravel',
method: methodNode.text.toUpperCase(),
path,
name: 'route',
name: closureNode ? null : 'route',
line: (closureNode ?? pathNode).startPosition.row + 1,
confidence: 0.8,
});
}
@@ -161,6 +184,7 @@ export const PHP_HTTP_PLUGIN: HttpLanguagePlugin = {
method: methodNode.text.toUpperCase(),
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -177,6 +201,7 @@ export const PHP_HTTP_PLUGIN: HttpLanguagePlugin = {
method: methodNode.text.toUpperCase(),
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -192,6 +217,7 @@ export const PHP_HTTP_PLUGIN: HttpLanguagePlugin = {
method: 'GET',
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -6,6 +6,7 @@ import {
unquoteLiteral,
type LanguagePatterns,
} from '../tree-sitter-scanner.js';
import { normalizeExtractedRoutePath } from '../../../ingestion/route-extractors/route-path.js';
import type { HttpDetection, HttpLanguagePlugin, RepoContext } from './types.js';
/**
@@ -79,6 +80,33 @@ const FASTAPI_ROUTER_PATTERNS = compilePatterns({
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Provider: Flask `app.add_url_rule('/path', view_func=handler)` ───
// The imperative Flask route registration: unlike `@app.route` (whose handler
// is the decorated function, same-file), `view_func` is frequently an IMPORTED
// (and sometimes aliased) view, so the handler resolves through the file's
// imports. `add_url_rule` + a `view_func=` keyword is highly Flask-specific, so
// the false-positive risk is low. Method(s) come from a `methods=[...]` keyword
// (default GET), extracted in code from the captured call.
const FLASK_ADD_URL_RULE_PATTERNS = compilePatterns({
name: 'python-flask-add-url-rule',
language: Python,
patterns: [
{
meta: {},
query: `
(call
function: (attribute
attribute: (identifier) @fn (#eq? @fn "add_url_rule"))
arguments: (argument_list
. (string) @path
(keyword_argument
name: (identifier) @kw (#eq? @kw "view_func")
value: (identifier) @handler))) @call
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── include_router(<router_obj>, prefix='/x') across the repo ────────
// Two shapes are common:
// app.include_router(assistant.router, prefix='/ai')
@@ -135,6 +163,26 @@ const INCLUDE_ROUTER_NAME_PATTERNS = compilePatterns({
],
} satisfies LanguagePatterns<Record<string, never>>);
const API_ROUTER_PREFIX_PATTERNS = compilePatterns({
name: 'python-fastapi-apirouter-prefix',
language: Python,
patterns: [
{
meta: {},
query: `
(assignment
left: (identifier) @router_name (#eq? @router_name "router")
right: (call
function: (identifier) @factory (#eq? @factory "APIRouter")
arguments: (argument_list
(keyword_argument
name: (identifier) @kw (#eq? @kw "prefix")
value: (string) @prefix))))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// `from .api.assistant import router` style — used together with
// INCLUDE_ROUTER_NAME so we can map a local name back to its module
// path, then back to the file the router was declared in.
@@ -331,6 +379,73 @@ const WRAPPER_URI_VAR_PATTERNS = compilePatterns({
],
} satisfies LanguagePatterns<Record<string, never>>);
/**
* Map each `from <module> import <name> [as <alias>]` binding to its declared
* name + raw module specifier (the spec keeps the leading dots for relative
* imports — `.users`, `..pkg.users` — which the extractor resolves to a target
* file). Lets a Flask `view_func` handler resolve through an alias to the real
* symbol in its module rather than the local alias text. `import x` / `import x
* as y` (module imports, not symbol imports) are left out — a route handler is a
* symbol, addressed via `from … import …`.
*/
function buildPythonImportMap(tree: Parser.Tree): Map<string, { name: string; module: string }> {
const map = new Map<string, { name: string; module: string }>();
const walk = (node: Parser.SyntaxNode): void => {
if (node.type === 'import_from_statement') {
const moduleNode = node.childForFieldName('module_name');
const module = moduleNode?.text ?? null;
if (module !== null) {
for (let i = 0; i < node.namedChildCount; i++) {
const c = node.namedChild(i);
if (!c || c.id === moduleNode?.id) continue;
if (c.type === 'dotted_name') {
map.set(c.text, { name: c.text, module });
} else if (c.type === 'aliased_import') {
const nameNode = c.childForFieldName('name');
const aliasNode = c.childForFieldName('alias');
if (nameNode && aliasNode) {
map.set(aliasNode.text, { name: nameNode.text, module });
}
}
}
}
}
for (let i = 0; i < node.namedChildCount; i++) {
const c = node.namedChild(i);
if (c) walk(c);
}
};
walk(tree.rootNode);
return map;
}
/**
* HTTP verbs declared on a Flask `add_url_rule(..., methods=[...])` call, upper-
* cased. Defaults to `['GET']` when no `methods` keyword is present (Flask's own
* default). Reads the captured call node directly since the list value is awkward
* to capture in a tree-sitter query.
*/
function extractFlaskMethods(callNode: Parser.SyntaxNode): string[] {
const args = callNode.childForFieldName('arguments');
if (args) {
for (let i = 0; i < args.namedChildCount; i++) {
const kw = args.namedChild(i);
if (!kw || kw.type !== 'keyword_argument') continue;
if (kw.childForFieldName('name')?.text !== 'methods') continue;
const list = kw.childForFieldName('value');
if (!list) continue;
const methods: string[] = [];
for (let j = 0; j < list.namedChildCount; j++) {
const el = list.namedChild(j);
const v = el && el.type === 'string' ? unquoteLiteral(el.text) : null;
if (v) methods.push(v.toUpperCase());
}
if (methods.length > 0) return methods;
}
}
return ['GET'];
}
// Pre-scan: collect local string assignments (uri = "api/v1/endpoint/")
function buildLocalStringMap(tree: Parser.Tree): Map<string, string> {
const map = new Map<string, string>();
@@ -811,10 +926,10 @@ function buildPythonRepoContext(
const prefixesByLongKey = new Map<string, Set<string>>();
const prefixesByShortKey = new Map<string, Set<string>>();
// Pre-pass over .py files. We deliberately run this even on files
// that don't contain `include_router` — the cost of an extra parse
// is bounded by the file count, and detecting `include_router`
// beforehand would require its own grep/scan.
// Cross-file pre-pass: only `include_router` sites need it — they bind a
// prefix declared in one file to a router defined in another. Same-file
// `APIRouter(prefix=...)` is resolved in scan() from the file's own tree, so
// APIRouter-only files are left out here and never parsed twice.
for (const rel of files) {
if (!rel.endsWith('.py')) continue;
const src = readFile(rel);
@@ -907,19 +1022,37 @@ function buildPythonRepoContext(
}
}
return { prefixesByLongKey, prefixesByShortKey };
return {
prefixesByLongKey,
prefixesByShortKey,
};
}
function joinPrefix(prefix: string, route: string): string {
// Mirror FastAPI's path joining: trim trailing slash off prefix,
// ensure exactly one leading slash on the result.
const p = prefix.replace(/\/+$/, '');
const r = route.startsWith('/') ? route : `/${route}`;
return `${p}${r}`;
// Delegate to the shared route-path normalizer so the group contract and the
// ingestion Route node join prefixes identically — one helper, no
// trailing-slash drift on empty routes (`APIRouter(prefix="/x")` + `@get("")`).
return normalizeExtractedRoutePath(route, prefix);
}
export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
name: 'python-http',
language: Python,
// routeCoverage intentionally LEFT at the default 'partial' (#2138 Part 2).
// It would be a no-op even if set to 'complete': FastAPI decorator routes set
// no handlerName (generic worker path) and Django sets methodName: null, so no
// Python file ever resolves a handlerSymbolId and none would be parse-skipped.
// Declaring 'complete' now is only a latent trap for the moment a follow-up
// gives FastAPI routes a handlerName. `hasConsumerSignals` is kept (and is a
// true superset of scan()'s consumer shapes) so the precondition already holds
// when Python is later flipped to 'complete'.
// Consumer signals scan() can detect: `requests.<verb>`/`requests.request`,
// `httpx` (sync/async client), the `uri=`/`url=` keyword/variable wrapper
// calls, plus aiohttp/urllib. Conservative — over-matching only costs a parse.
hasConsumerSignals(content) {
return /\brequests\s*\.|\bhttpx\b|\baiohttp\b|\burllib\b|\burlopen\b|\buri\s*=|\burl\s*=/.test(
content,
);
},
prepareRepo({ files, parser, readFile, parseSource }): RepoContext {
return buildPythonRepoContext(files, parser, readFile, parseSource);
},
@@ -927,6 +1060,10 @@ export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
const out: HttpDetection[] = [];
const httpxAsyncClients = collectHttpxAsyncClients(tree);
const ctx = repoContext as PythonRepoContext | undefined;
// Local-binding → { declared name, module } for the file's `from … import …`
// statements, so an imperatively-registered handler (Flask `view_func`) that
// is an imported (possibly aliased) symbol resolves to its real definition.
const importMap = buildPythonImportMap(tree);
// Providers: FastAPI @app.<verb>("/path") — already absolute path.
for (const match of runCompiledPatterns(FASTAPI_APP_PATTERNS, tree)) {
@@ -943,6 +1080,12 @@ export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
method: httpMethod,
path,
name: null,
// The decorated handler has no captured name → resolve by line-span
// containment. Best-effort fallback: FastAPI routes are graph-backed
// (ingestion decorator routes) and the function span starts at `def`
// (decorators excluded), so this lands the single-decorator case and
// degrades to file-level for multi-decorator stacks.
line: pathNode.startPosition.row + 1,
confidence: 0.8,
});
}
@@ -951,6 +1094,18 @@ export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
// the ingestion route extractor), not a per-file source scan — see the note
// at the top of this file.
// Same-file `router = APIRouter(prefix="/x")` (router-only). Read from this
// file's own tree, so there is no cross-file map and no prefix bleed across
// same-stem files; it stacks under any include_router(prefix=...) below.
let constructorPrefix: string | undefined;
for (const m of runCompiledPatterns(API_ROUTER_PREFIX_PATTERNS, tree)) {
const prefixNode = m.captures.prefix;
if (!prefixNode) continue;
const p = unquoteLiteral(prefixNode.text);
if (p === null) continue;
constructorPrefix = p;
}
// Providers: FastAPI @router.<verb>("/path") — must be joined
// with the prefix(es) declared at the include_router site. When
// no prefix is found we still emit the unprefixed path so this
@@ -975,10 +1130,13 @@ export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
const shortPrefixes =
longPrefixes || !shortKey ? undefined : ctx?.prefixesByShortKey.get(shortKey);
const prefixSet = longPrefixes ?? shortPrefixes;
// Stack the same-file APIRouter(prefix=...) under any cross-file
// include_router prefix.
const localPath = constructorPrefix ? joinPrefix(constructorPrefix, rawPath) : rawPath;
const paths =
prefixSet && prefixSet.size > 0
? [...prefixSet].map((p) => joinPrefix(p, rawPath))
: [rawPath];
? [...prefixSet].map((p) => joinPrefix(p, localPath))
: [localPath];
for (const p of paths) {
out.push({
@@ -987,6 +1145,34 @@ export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
method: httpMethod,
path: p,
name: null,
// Best-effort containment fallback — see the @app provider note above.
line: pathNode.startPosition.row + 1,
confidence: 0.8,
});
}
}
// Providers: Flask `app.add_url_rule('/path', view_func=handler, methods=[…])`.
// The handler is a `view_func` identifier, frequently an imported (possibly
// aliased) view, so resolve it through the file's imports to the declared
// symbol + its module for import-pinned resolution downstream.
for (const match of runCompiledPatterns(FLASK_ADD_URL_RULE_PATTERNS, tree)) {
const pathNode = match.captures.path;
const handlerNode = match.captures.handler;
const callNode = match.captures.call;
if (!pathNode || !handlerNode || !callNode) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
const imported = importMap.get(handlerNode.text);
for (const method of extractFlaskMethods(callNode)) {
out.push({
role: 'provider',
framework: 'flask',
method,
path,
name: imported ? imported.name : handlerNode.text,
handlerImport: imported,
line: (imported ? pathNode : handlerNode).startPosition.row + 1,
confidence: 0.8,
});
}
@@ -1005,6 +1191,7 @@ export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
method: methodNode.text.toUpperCase(),
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -1022,6 +1209,7 @@ export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
method: methodNode.text.toUpperCase(),
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -1040,6 +1228,7 @@ export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
method: methodRaw.toUpperCase(),
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -1059,6 +1248,7 @@ export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
method: methodNode.text.toUpperCase(),
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -1079,6 +1269,7 @@ export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
method: methodRaw.toUpperCase(),
path,
name: null,
line: pathNode.startPosition.row + 1,
confidence: 0.7,
});
}
@@ -1111,6 +1302,7 @@ export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
method: httpMethod,
path,
name: null,
line: methodNode.startPosition.row + 1,
confidence: 0.65,
});
}
@@ -1137,6 +1329,7 @@ export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
method: httpMethod,
path: normalized,
name: null,
line: methodNode.startPosition.row + 1,
confidence: 0.6,
});
}
@@ -16,6 +16,10 @@
*/
import type { HttpDetection, HttpFileDetections } from './types.js';
import {
resolveInheritedSpringRoutes,
type SharedSpringType,
} from '../../../ingestion/route-extractors/spring-shared.js';
/**
* RestTemplate method-name → HTTP verb. Source-scan only: the receiver must be
@@ -125,141 +129,28 @@ export function parseRequestLine(raw: string): { method: string; path: string }
}
/**
* Join a class/interface-level prefix and a method-level path into a single
* URL path: strip leading/trailing slashes on the prefix and leading slashes
* on the method path, then ensure exactly one slash between them.
*/
export function joinPath(prefix: string, methodPath: string): string {
const cleanPrefix = prefix.replace(/^\/+/, '').replace(/\/+$/, '');
const cleanSub = methodPath.replace(/^\/+/, '');
if (!cleanPrefix) return `/${cleanSub}`;
return `/${cleanPrefix}/${cleanSub}`;
}
/**
* Join a controller's own class prefix with a route inherited from an interface
* (interface-based controllers, #1743). The inherited path already has the
* interface's own class prefix (`inheritedOwnerPrefix`) baked in; when the
* controller repeats that same prefix we must NOT prepend it twice (#2057).
* Shared by both plugins' `scanProject` so Java and Kotlin agree.
*/
export function joinInheritedSpringPath(
controllerPrefix: string,
inheritedPath: string,
inheritedOwnerPrefix = '',
): string {
const joined = joinPath(controllerPrefix, inheritedPath);
const cleanPrefix = controllerPrefix.replace(/^\/+/, '').replace(/\/+$/, '');
const cleanOwnerPrefix = inheritedOwnerPrefix.replace(/^\/+/, '').replace(/\/+$/, '');
const cleanInherited = inheritedPath.replace(/^\/+/, '');
if (!cleanPrefix) return joined;
if (
cleanPrefix === cleanOwnerPrefix &&
(cleanInherited === cleanPrefix || cleanInherited.startsWith(`${cleanPrefix}/`))
) {
return `/${cleanInherited}`;
}
return joined;
}
/**
* Language-agnostic view of a Spring class/interface that each plugin's
* grammar-specific collector produces. The interface-based-controller
* inheritance algorithm (`scanSpringInheritanceProject`) operates only on this
* shape, so the Java and Kotlin plugins share one algorithm and cannot drift.
*
* `methods[].routes` carry only `{ method, path }` — the interface's own class
* prefix is applied *inside* `scanSpringInheritanceProject` (it is not part of
* the collector's output).
*/
export interface SharedSpringType {
filePath: string;
kind: 'class' | 'interface';
name: string;
/** Class-level `@RequestMapping` prefixes — one per array element. */
classPrefixes: string[];
implementedInterfaces: string[];
isController: boolean;
methods: Array<{ name: string; routes: Array<{ method: string; path: string }> }>;
}
/**
* Resolve interface-based-controller provider routes (#1743): a concrete
* Resolve interface-based-controller provider *detections* (#1743): a concrete
* `@RestController`/`@Controller` class inherits the `@(Get|...)Mapping` routes
* declared on the interface it implements. Shared by the Java and Kotlin plugins
* so both emit byte-identical provider contracts.
*
* An interface name that resolves to two distinct interfaces is ambiguous and
* its routes are dropped (the `null` marker). The controller's own class
* prefix(es) cross-product the inherited routes; `joinInheritedSpringPath`
* avoids doubling a prefix the interface already baked in (#2057).
* declared on the interface it implements. Thin group-layer adapter over the
* shared, language-agnostic `resolveInheritedSpringRoutes` (in
* `ingestion/route-extractors/spring-shared.ts`) — it maps each inherited route
* to a provider `HttpDetection`. Shared by the Java and Kotlin plugins so both
* emit byte-identical provider contracts; the ingestion route extractor calls
* the same underlying algorithm so all three stay in parity.
*/
export function scanSpringInheritanceProject(types: SharedSpringType[]): HttpFileDetections[] {
// interface name → (method name → routes). `ownerPrefix` records the
// interface's own class prefix so the controller side avoids doubling it
// (#2057). `null` marks an ambiguous (duplicated) interface name.
type InheritedRoute = { method: string; path: string; ownerPrefix: string };
const interfaceRoutes = new Map<string, Map<string, InheritedRoute[]> | null>();
for (const type of types) {
if (type.kind !== 'interface') continue;
if (interfaceRoutes.has(type.name)) {
interfaceRoutes.set(type.name, null);
continue;
}
const prefixes = type.classPrefixes.length ? type.classPrefixes : [''];
const methodMap = new Map<string, InheritedRoute[]>();
for (const method of type.methods) {
// Cross-product the interface's class prefixes with each method route, so a
// multi-element `@RequestMapping(["/a","/b"])` interface yields N bindings.
const routes = method.routes.flatMap((route) =>
prefixes.map((prefix) => ({
method: route.method,
path: prefix ? joinPath(prefix, route.path) : route.path,
ownerPrefix: prefix,
})),
);
if (routes.length > 0) methodMap.set(method.name, routes);
}
interfaceRoutes.set(type.name, methodMap);
}
const detectionsByFile = new Map<string, HttpDetection[]>();
for (const type of types) {
if (type.kind !== 'class' || !type.isController) continue;
// Cross-product the controller's own class prefixes with each inherited
// route; `['']` keeps the common no-prefix controller emitting the
// interface path unchanged.
const controllerPrefixes = type.classPrefixes.length ? type.classPrefixes : [''];
for (const method of type.methods) {
if (method.routes.length > 0) continue; // own @*Mapping → already a provider via scan()
const inherited = type.implementedInterfaces.flatMap((iface) => {
const routeMap = interfaceRoutes.get(iface);
if (!routeMap) return [];
const routes = routeMap.get(method.name) ?? [];
return routes.flatMap((route) =>
controllerPrefixes.map((controllerPrefix) => ({
method: route.method,
path: joinInheritedSpringPath(controllerPrefix, route.path, route.ownerPrefix),
})),
);
});
const seen = new Set<string>();
for (const route of inherited) {
const key = `${route.method} ${route.path}`;
if (seen.has(key)) continue;
seen.add(key);
const detections = detectionsByFile.get(type.filePath) ?? [];
detections.push({
role: 'provider',
framework: 'spring',
method: route.method,
path: route.path,
name: method.name,
confidence: 0.8,
});
detectionsByFile.set(type.filePath, detections);
}
}
for (const route of resolveInheritedSpringRoutes(types)) {
const detections = detectionsByFile.get(route.filePath) ?? [];
detections.push({
role: 'provider',
framework: 'spring',
method: route.method,
path: route.path,
name: route.methodName,
confidence: 0.8,
});
detectionsByFile.set(route.filePath, detections);
}
return [...detectionsByFile.entries()].map(([filePath, detections]) => ({
@@ -36,6 +36,26 @@ export interface HttpDetection {
* Null when no good candidate is available.
*/
name: string | null;
/**
* 1-based source line of the call/registration site (the `fetch(...)` for
* consumers, the `router.get(...)` / decorator for providers). Lets the
* extractor resolve the contract to the *containing* symbol (the function
* the call lives in) via line-span containment, so HTTP contracts carry a
* real `symbolUid` instead of an empty one. Optional — a plugin that does
* not set it falls back to file-level boundary resolution downstream.
*/
line?: number;
/**
* When the handler is an IMPORTED symbol, the import resolved to its declared
* (exported) `name` and the `module` specifier it came from. The extractor
* pins resolution to the import's target file, so an aliased import
* (`import { listUsers as handleUsers }`) or a name that collides with a local
* symbol resolves to the right handler instead of a same-named decoy. `name`
* here is the DECLARED export name (not the local alias); `module` is the raw
* specifier (e.g. `./handlers/users`). Set only for named imports; omitted for
* locally-defined or anonymous handlers.
*/
handlerImport?: { name: string; module: string };
/** Confidence in (0, 1]. Source-scan plugins typically use 0.7–0.8. */
confidence: number;
}
@@ -78,6 +98,43 @@ export interface HttpLanguagePlugin {
name: string;
/** tree-sitter grammar object (passed to the shared parser). */
language: unknown;
/**
* Whether ingestion is known to emit a `Route` graph node for EVERY
* provider route in this language (Spring/FastAPI/Laravel annotations are
* extracted into Route nodes during parse). When `'complete'`, the
* orchestrator may skip the source-scan + tree-sitter parse for a file whose
* graph provider routes all resolved a handler symbol (#2138 Part 2) — the
* graph is authoritative, the scan would only re-discover the same routes.
*
* Defaults to `'partial'` (the safe assumption): the source scan always runs,
* so a language whose ingestion coverage is incomplete never loses routes.
* This is a deliberate, per-language trust assertion — set it only for
* languages whose route ingestion is provably complete.
*/
routeCoverage?: 'complete' | 'partial';
/**
* Cheap, parse-free pre-check used by the parse-skip optimization (#2138
* Part 2). Given a file's raw source text, return `false` ONLY when the file
* provably contains no outbound-HTTP (consumer) call that this plugin's
* `scan()` would detect; return `true` on any doubt.
*
* Why it exists: `routeCoverage: 'complete'` asserts *provider* Route-node
* completeness only. A provider-covered file may ALSO be a consumer (e.g. a
* Spring `@RestController` that calls `restTemplate`/`webClient`, a Laravel
* controller using Guzzle, a FastAPI handler calling `requests`/`httpx`).
* Ingestion's `FETCHES` edges are JS/TS-only, so the graph cannot back up
* those server-side consumers — they come solely from the source scan. The
* orchestrator may therefore skip a provider-covered file's parse only when
* this returns `false`; otherwise the file is still scanned so its consumer
* contracts are not dropped.
*
* MUST be implemented by any plugin whose `scan()` can emit `'consumer'`
* detections AND that declares `routeCoverage: 'complete'`; otherwise that
* language's provider-covered files are never parse-skipped (safe, no win).
* The check is intentionally conservative — over-matching only costs a parse
* that could have been skipped; it never drops data.
*/
hasConsumerSignals?(content: string): boolean;
/**
* Optional pre-pass: walk the relevant files in the repo and produce
* an opaque context that `scan` can use to resolve cross-file facts.
@@ -49,6 +49,7 @@ MATCH (handlerFile:File)-[r:CodeRelation {type: 'HANDLES_ROUTE'}]->(route:Route)
RETURN handlerFile.id AS fileId, handlerFile.filePath AS filePath,
route.name AS routePath, route.id AS routeId,
route.method AS routeMethod,
route.handlerSymbolId AS handlerSymbolId,
route.responseKeys AS responseKeys,
r.reason AS routeSource`;
const FETCHES_QUERY = `
@@ -57,11 +58,163 @@ RETURN callerFile.id AS fileId, callerFile.filePath AS filePath,
route.name AS routePath, route.id AS routeId,
r.reason AS fetchReason`;
const CONTAINS_QUERY = `
MATCH (file:File {id: $fileId})<-[:CodeRelation {type: 'CONTAINS'}]-(sym)
WHERE sym.startLine IS NOT NULL
RETURN sym.id AS uid, sym.name AS name, sym.filePath AS filePath, labels(sym) AS labels
ORDER BY sym.startLine`;
// Function/Method/CodeElement symbols (with line spans) in a file, addressed by
// repo-relative path so the source-scan paths — which have a path but no graph
// `fileId` — can resolve the symbol CONTAINING an HTTP call by line-span
// containment. Matched by `filePath` rather than a File-[DEFINES]->sym edge so
// it also reaches methods nested in classes (Java/Kotlin), where the File
// defines the class and the class defines the method.
const CONTAINING_QUERY = `
MATCH (sym:Function)
WHERE sym.filePath = $filePath AND sym.startLine IS NOT NULL AND sym.endLine IS NOT NULL
RETURN sym.id AS uid, sym.name AS name, sym.filePath AS filePath,
sym.startLine AS startLine, sym.endLine AS endLine, labels(sym) AS labels
UNION ALL
MATCH (sym:Method)
WHERE sym.filePath = $filePath AND sym.startLine IS NOT NULL AND sym.endLine IS NOT NULL
RETURN sym.id AS uid, sym.name AS name, sym.filePath AS filePath,
sym.startLine AS startLine, sym.endLine AS endLine, labels(sym) AS labels
UNION ALL
MATCH (sym:CodeElement)
WHERE sym.filePath = $filePath AND sym.startLine IS NOT NULL AND sym.endLine IS NOT NULL
RETURN sym.id AS uid, sym.name AS name, sym.filePath AS filePath,
sym.startLine AS startLine, sym.endLine AS endLine, labels(sym) AS labels`;
// Repo-wide lookup of a symbol by exact name. Used to resolve a provider's
// named handler when it is defined in a file OTHER than its route registration —
// and only honored when the result is unique (see resolveSymbolByNameUnique).
//
// Label filtering uses `labels(n) IN [...]` rather than the openCypher
// disjunction `MATCH (n:A|B|C)`. NOTE: this 3-label set (Function/Method/
// CodeElement) actually PARSES — LadybugDB only rejects a disjunction that
// names a reserved keyword (e.g. `Macro`, `Union`) or a missing node table,
// neither of which applies here. So this query was NOT broken by #2325; it
// uses the `labels(n) IN` form for consistency with the manifest custom-branch
// fix (which WAS broken) and to stay immune if a reserved-keyword label is
// added later. `labels(n)` returns the node's single label as a string here, so
// `IN [...]` is an exact allowlist. (Exported so integration tests can run the
// exact production query against a real LadybugDB — the bug shipped because no
// test ran these strings against the real parser.)
//
// `n.filePath <> ''` excludes synthetic non-source `CodeElement` nodes that
// carry no real file — ORM model/table nodes (orm.ts emits `filePath: ''`) and
// similar — so a handler name colliding with an ORM model neither resolves to a
// degenerate edge-less node NOR inflates the uniqueness count and masks the real
// handler. `LIMIT 2` bounds materialization: distinguishing unique (1) from
// ambiguous (>=2) never needs more than two rows (the count guard stays exact).
export const RESOLVE_BY_NAME_QUERY = `
MATCH (n) WHERE labels(n) IN ['Function','Method','CodeElement']
AND n.name = $name AND n.filePath <> ''
RETURN n.id AS uid, n.name AS name, n.filePath AS filePath
LIMIT 2`;
// Resolve an IMPORTED handler by pinning it to the import's target module: the
// declared export `$name` whose file is the module the handler was imported from
// (`$fileDot` matches `mod.ext`, `$fileSlash` matches `mod/index.ext`). This is
// the precise rung — it survives aliases and local same-name collisions that a
// repo-wide name lookup cannot, and only resolves on a unique match within that
// module. `LIMIT 2` keeps the uniqueness count exact (see RESOLVE_BY_NAME_QUERY).
export const RESOLVE_IN_MODULE_QUERY = `
MATCH (n) WHERE labels(n) IN ['Function','Method','CodeElement']
AND n.name = $name AND (n.filePath STARTS WITH $fileDot OR n.filePath STARTS WITH $fileSlash)
RETURN n.id AS uid, n.name AS name, n.filePath AS filePath
LIMIT 2`;
// Source-file extensions an import specifier may resolve to (stripped before
// building the module file-prefix so `./h/users` and `./h/users.ts` agree).
const SOURCE_EXT_RE = /\.(?:m|c)?[jt]sx?$/;
/**
* Resolve an import specifier to a repo-relative FILE BASE (path without
* extension) so the target module can be matched by `filePath STARTS WITH`.
* Handles two relative-import dialects and returns null for bare/absolute
* imports (which fall back to a repo-wide name lookup):
* - path-style (JS/TS): `./handlers/users`, `../x` → joined against the
* importing file's directory.
* - dotted-relative (Python): `.users`, `..pkg.users` → leading dots are
* package levels (one dot = the file's own package), the rest dot→slash.
*/
function resolveModuleBase(fromFile: string, module: string): string | null {
const dir = path.posix.dirname(fromFile.replace(/\\/g, '/'));
if (module.includes('/')) {
// path-style relative import
if (!module.startsWith('.')) return null;
return path.posix.normalize(path.posix.join(dir, module)).replace(SOURCE_EXT_RE, '');
}
if (module.startsWith('.')) {
// Python dotted-relative import
const dots = module.length - module.replace(/^\.+/, '').length;
const rest = module.slice(dots).replace(/\./g, '/');
let base = dir;
for (let i = 1; i < dots; i++) base = path.posix.dirname(base);
return rest ? path.posix.normalize(path.posix.join(base, rest)) : base;
}
return null; // bare / absolute import — repo-wide fallback
}
interface ResolvedSymbol {
uid: string;
name: string;
filePath: string;
}
/**
* The innermost Function/Method whose `[startLine, endLine]` span contains
* `line` — i.e. the symbol the HTTP call lives inside. For a consumer this is
* the function making the `fetch`; for an inline-arrow provider it is the
* handler arrow itself. Returns null when nothing encloses the line (e.g. a
* route registered at module scope referencing a named handler defined
* elsewhere — that case resolves by name instead).
*/
function resolveContainingSymbol(
rows: Record<string, unknown>[],
line: number,
): ResolvedSymbol | null {
const norm = (x: unknown): string => String(x ?? '');
// Detection lines are 1-based; symbol spans are stored 0-based for the
// languages indexed today (parse-worker records `startPosition.row`). So the
// base-correct probe is `line - 1`. Pick the INNERMOST (smallest-span) symbol
// whose span contains the probe. Only if nothing contains `line - 1` do we
// retry with the raw `line` — a defensive fallback for any future language
// that stores 1-based spans. Probing `line - 1` first (rather than OR-ing both)
// avoids the +1 slack mis-picking a one-line sibling that sits on `line`.
const pick = (probe: number): ResolvedSymbol | null => {
let best: ResolvedSymbol | null = null;
let bestSpan = Number.POSITIVE_INFINITY;
for (const r of rows) {
const labels = JSON.stringify(r.labels ?? r[5] ?? '');
if (!['Function', 'Method', 'CodeElement'].some((l) => labels.includes(l))) continue;
const start = Number(r.startLine ?? r[3]);
const end = Number(r.endLine ?? r[4]);
if (!Number.isFinite(start) || !Number.isFinite(end)) continue;
if (probe < start || probe > end) continue;
const span = end - start;
if (span < bestSpan) {
bestSpan = span;
best = {
uid: norm(r.uid ?? r[0]),
name: norm(r.name ?? r[1]),
filePath: norm(r.filePath ?? r[2]),
};
}
}
return best && best.uid ? best : null;
};
return pick(line - 1) ?? pick(line);
}
/** A Function/Method in the file matching `name` exactly (for named handlers). */
function resolveSymbolByName(rows: Record<string, unknown>[], name: string): ResolvedSymbol | null {
const norm = (x: unknown): string => String(x ?? '');
for (const r of rows) {
const labels = JSON.stringify(r.labels ?? r[5] ?? '');
if (!['Function', 'Method', 'CodeElement'].some((l) => labels.includes(l))) continue;
if (norm(r.name ?? r[1]) !== name) continue;
const uid = norm(r.uid ?? r[0]);
if (uid) return { uid, name, filePath: norm(r.filePath ?? r[2]) };
}
return null;
}
// ─── Path normalization (shared between provider / consumer paths) ──
@@ -111,6 +264,10 @@ function contractIdFor(method: string, pathNorm: string): string {
return `http::${method.toUpperCase()}::${pathNorm}`;
}
export function normalizeRepoRelPath(filePath: string): string {
return filePath.replace(/\\/g, '/').replace(/^\.\//, '');
}
// ─── Graph row helpers ───────────────────────────────────────────────
function methodFromRouteReason(reason: string): string | null {
@@ -123,35 +280,6 @@ function methodFromRouteReason(reason: string): string | null {
return null;
}
function pickSymbolUid(
rows: Record<string, unknown>[],
preferredName: string | null,
): { uid: string; name: string; filePath: string } {
const norm = (x: unknown) => String(x ?? '');
const labeled = rows.filter((r) => {
const labels = r.labels ?? r[3];
const s = JSON.stringify(labels);
return s.includes('Method') || s.includes('Function');
});
const pool = labeled.length > 0 ? labeled : rows;
if (preferredName) {
const hit = pool.find((r) => norm(r.name ?? r[1]) === preferredName);
if (hit) {
return {
uid: norm(hit.uid ?? hit[0]),
name: norm(hit.name ?? hit[1]),
filePath: norm(hit.filePath ?? hit[2]),
};
}
}
const first = pool[0] || rows[0];
return {
uid: norm(first?.uid ?? first?.[0]),
name: norm(first?.name ?? first?.[1]),
filePath: norm(first?.filePath ?? first?.[2]),
};
}
// ─── Orchestrator ────────────────────────────────────────────────────
export class HttpRouteExtractor implements ContractExtractor {
@@ -282,22 +410,197 @@ export class HttpRouteExtractor implements ContractExtractor {
};
const files = await getScannedFiles();
await collectProjectDetections(files);
// Resolve an HTTP detection to the symbol it lives in — the containing
// function for a consumer / inline-arrow provider, or a named handler for
// a provider — addressed by repo-relative file path so the source-scan
// paths (which have no graph `fileId`) can resolve too. Per-file symbol
// lists are cached. Returns null without a DB or when nothing resolves (a
// named provider resolves by name even with no `line`; containment needs
// one); the contract then keeps an empty symbolUid and downstream falls
// back to file-level boundary matching.
const fileSymbolCache = new Map<string, Record<string, unknown>[]>();
const loadFileSymbols = async (filePath: string): Promise<Record<string, unknown>[]> => {
if (!dbExecutor) return [];
const cached = fileSymbolCache.get(filePath);
if (cached) return cached;
let rows: Record<string, unknown>[] = [];
try {
rows = await dbExecutor(CONTAINING_QUERY, { filePath });
} catch {
rows = [];
}
fileSymbolCache.set(filePath, rows);
return rows;
};
// Repo-wide UNAMBIGUOUS resolution for a provider handler defined in a file
// other than its route registration (e.g. `router.get('/x', listUsers)` with
// `listUsers` imported from another module). Returns the symbol ONLY when
// exactly one Function/Method/CodeElement carries that name across the repo.
// The strict uniqueness guard is intentionally conservative: when a name is
// shared across files (homonyms like `handler`/`index`), we prefer a
// false-negative (no attribution → file-level fallback) over a false-positive
// (wrong symbol).
//
// An IMPORTED handler (the common cross-file case) is pinned to its source
// module first by resolveImportedSymbol, so an alias or a name colliding with
// a local symbol resolves correctly; this repo-wide-by-name rung is the
// fallback for non-relative/bare imports and for plugins that supply only a
// name. Cached by name for the lifetime of this extract().
const globalNameCache = new Map<string, ResolvedSymbol | null>();
const toResolvedSymbol = (rows: Record<string, unknown>[]): ResolvedSymbol | null => {
const norm = (x: unknown): string => String(x ?? '');
const uid = rows.length === 1 ? norm(rows[0]!.uid ?? rows[0]![0]) : '';
const filePath = uid ? norm(rows[0]!.filePath ?? rows[0]![2]) : '';
// Reject a unique match that carries no real file (a synthetic ORM /
// non-source node) so it can never anchor a cross-trace on an edge-less
// node — defence in depth alongside the queries' filePath predicates.
return uid && filePath ? { uid, name: norm(rows[0]!.name ?? rows[0]![1]), filePath } : null;
};
const resolveSymbolByNameUnique = async (name: string): Promise<ResolvedSymbol | null> => {
if (!dbExecutor) return null;
const cached = globalNameCache.get(name);
if (cached !== undefined) return cached;
let rows: Record<string, unknown>[] = [];
try {
rows = await dbExecutor(RESOLVE_BY_NAME_QUERY, { name });
} catch {
rows = [];
}
const result = toResolvedSymbol(rows);
globalNameCache.set(name, result);
return result;
};
// Resolve a handler imported from a RELATIVE module to the unique declared
// symbol of that name inside the import's target file. Returns null for
// non-relative (bare/aliased-path) imports — those fall back to the repo-wide
// name lookup. Cached by (target-file-prefix, declared name).
const importedSymbolCache = new Map<string, ResolvedSymbol | null>();
const resolveImportedSymbol = async (
fromFile: string,
imp: { name: string; module: string },
): Promise<ResolvedSymbol | null> => {
if (!dbExecutor) return null;
const base = resolveModuleBase(fromFile, imp.module);
if (base === null) return null; // bare/absolute import → repo-wide fallback
const cacheKey = JSON.stringify([base, imp.name]);
const cached = importedSymbolCache.get(cacheKey);
if (cached !== undefined) return cached;
let rows: Record<string, unknown>[] = [];
try {
rows = await dbExecutor(RESOLVE_IN_MODULE_QUERY, {
name: imp.name,
fileDot: `${base}.`,
fileSlash: `${base}/`,
});
} catch {
rows = [];
}
const result = toResolvedSymbol(rows);
importedSymbolCache.set(cacheKey, result);
return result;
};
const resolveDetectionSymbol = async (
filePath: string,
d: HttpDetection,
): Promise<ResolvedSymbol | null> => {
if (!dbExecutor) return null;
const syms = await loadFileSymbols(filePath);
// Name resolution does NOT need a detection line — a named provider
// handler (Spring/Go/etc. method name) resolves by name even when the
// plugin didn't set `line`. Try the registration file FIRST; then, for a
// handler defined in another file, the unique repo-wide match. Only the
// containment fallback requires a line.
if (d.role === 'provider' && d.name) {
// IMPORTED handler: pin to the import's target module first. This is the
// precise rung — it survives aliases and names that collide with a local
// symbol. The handler is defined ELSEWHERE, so a file-scoped lookup of
// its (declared) name would be wrong; on a miss go straight to a unique
// repo-wide match on the declared name, never file-scoped.
if (d.handlerImport) {
const byImport = await resolveImportedSymbol(filePath, d.handlerImport);
if (byImport) return byImport;
const byGlobal = await resolveSymbolByNameUnique(d.handlerImport.name);
if (byGlobal) return byGlobal;
return null;
}
const byName = resolveSymbolByName(syms, d.name);
if (byName) return byName;
const byGlobal = await resolveSymbolByNameUnique(d.name);
if (byGlobal) return byGlobal;
// A NAMED handler we could not resolve by name (neither file-scoped nor
// the unique repo-wide match) must NOT fall through to line-span
// containment: `d.line` is the route REGISTRATION site, so containment
// would attach the route to the enclosing registrar (e.g. a
// `setupRoutes()` wrapper) rather than the handler. Leave it empty →
// file-level boundary fallback, upholding the invariant that a
// zero/ambiguous name match never yields a wrong-symbol attribution.
return null;
}
// Consumers (the function making the fetch) and inline-arrow providers
// (d.name === null) DO resolve by containment — there the enclosing symbol
// is the right one.
if (syms.length === 0 || d.line == null) return null;
return resolveContainingSymbol(syms, d.line);
};
// Run the graph provider pass FIRST. After #2138 Part 2 it reads handler
// symbols from the graph (no source parse for resolved routes), so it can
// report which files are fully graph-covered BEFORE we decide what to
// parse. Files fully covered by a `routeCoverage: 'complete'` language are
// candidates to skip the source scan + tree-sitter parse — but only their
// *providers* are graph-authoritative; the consumer-safety gate below
// removes any candidate that still needs scanning for outbound calls.
const coveredFiles = new Set<string>();
const graphProviders =
dbExecutor != null ? await this.extractProvidersGraph(dbExecutor, getDetections) : [];
// Source scan always runs to capture routes in languages/files not covered
// by graph edges; the glob and per-file parse results are cached above.
dbExecutor != null
? await this.extractProvidersGraph(
dbExecutor,
getDetections,
resolveDetectionSymbol,
coveredFiles,
)
: [];
// Consumer-safety gate (#2138 Part 2): `extractProvidersGraph` marks a file
// covered on *provider* grounds (all HANDLES_ROUTE rows resolved + a
// `routeCoverage: 'complete'` language). But a provider-covered file may also
// be a *consumer* (a controller that calls RestTemplate/WebClient/Guzzle/
// requests/...), and ingestion emits no FETCHES edges for those server-side
// languages — the graph can't back them up. So a covered file is only truly
// safe to skip (parse) when its plugin can PROVE, from a cheap parse-free
// text scan, that it holds no such consumer call. Anything else (a positive
// signal, no `hasConsumerSignals` hook, or an unreadable file) stays in the
// scan set so its consumer contracts are preserved.
for (const f of [...coveredFiles]) {
const plugin = getPluginForFile(f);
const content = readSafe(repoPath, f);
const provenNoConsumer =
content != null && typeof plugin?.hasConsumerSignals === 'function'
? plugin.hasConsumerSignals(content) === false
: false;
if (!provenNoConsumer) coveredFiles.delete(f);
}
// Everything the graph did not fully cover still gets a full source scan
// (fail-open: partial-coverage languages, unresolved routes, and graph-less
// runs all land here).
const scanFiles = files.filter((f) => !coveredFiles.has(f));
await collectProjectDetections(scanFiles);
const providers = this.mergeGraphAndSourceContracts(
graphProviders,
await this.extractProvidersSourceScan(files, getDetections),
await this.extractProvidersSourceScan(scanFiles, getDetections, resolveDetectionSymbol),
);
const graphConsumers =
dbExecutor != null ? await this.extractConsumersGraph(dbExecutor, getDetections) : [];
dbExecutor != null
? await this.extractConsumersGraph(dbExecutor, getDetections, resolveDetectionSymbol)
: [];
const consumers = this.mergeGraphAndSourceContracts(
graphConsumers,
await this.extractConsumersSourceScan(files, getDetections),
await this.extractConsumersSourceScan(scanFiles, getDetections, resolveDetectionSymbol),
);
return [...providers, ...consumers];
@@ -323,8 +626,15 @@ export class HttpRouteExtractor implements ContractExtractor {
private async extractProvidersGraph(
db: CypherExecutor,
getDetections: (rel: string) => Promise<HttpDetection[]>,
resolveSymbol: (filePath: string, d: HttpDetection) => Promise<ResolvedSymbol | null>,
coveredFiles?: Set<string>,
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
// Per-file coverage tracking (#2138 Part 2): a file is "fully graph-covered"
// when every one of its HANDLES_ROUTE rows resolved a handlerSymbolId AND its
// language plugin declares `routeCoverage: 'complete'`. Such files can skip
// the source scan + parse entirely — the graph is authoritative for them.
const fileAllResolved = new Map<string, boolean>();
let rows: Record<string, unknown>[];
try {
rows = await db(HANDLES_ROUTE_QUERY);
@@ -354,67 +664,79 @@ export class HttpRouteExtractor implements ContractExtractor {
.toUpperCase();
let method = (graphMethod || null) ?? methodFromRouteReason(routeSource);
// Look up handler name (and backfill method if missing) from the
// plugin's scan of the handler file. This replaces the old
// regex-based `inferMethodFromFileScan` and `pickJavaHandlerName`
// helpers — tree-sitter gives both pieces of information
// structurally. Always run the lookup: even when method is set by
// `methodFromRouteReason`, we still need the handler name.
const detections = filePath ? await getDetections(filePath) : [];
const providerDetections = detections.filter((d) => d.role === 'provider');
let handlerName: string | null = null;
const normalizedRoute = normalizeHttpPath(routePath);
// Candidates share the same normalized path. When multiple
// detections at the same path exist (e.g. GET + POST /api/orders
// in one router), a blind `.find()` silently returned the first
// verb — attaching the wrong handler and, when method was not
// already pinned by the route reason, the wrong method too.
// Disambiguate by method when we know it; refuse to guess when
// we don't.
const candidates = providerDetections.filter(
(d) => normalizeHttpPath(d.path) === normalizedRoute,
);
let match: (typeof candidates)[number] | undefined;
const ambiguousCandidates = !method && candidates.length > 1;
if (method) {
match = candidates.find((d) => d.method === method);
} else if (candidates.length === 1) {
match = candidates[0];
const handlerSymbolId = String(row.handlerSymbolId ?? '').trim();
const fileId = row.fileId ?? row[0];
// Track per-file resolution for the parse-skip coverage set: a file stays
// "all resolved" only while every one of its rows carries a handlerSymbolId.
if (filePath) {
const prev = fileAllResolved.get(filePath);
fileAllResolved.set(filePath, (prev ?? true) && handlerSymbolId.length > 0);
}
// else: multiple candidates + unknown method → leave match
// undefined so handlerName stays null and skip symbol
// enrichment below, keeping the file-basename fallback instead
// of letting pickSymbolUid silently pick the first Function /
// Method in the file (which reintroduces the mis-attribution
// we were trying to avoid). Method stays at the conservative
// 'GET' default set below.
if (match) {
if (!method) method = match.method;
handlerName = match.name;
}
if (!method) method = 'GET';
const pathNorm = normalizeHttpPath(routePath);
const cid = contractIdFor(method, pathNorm);
const pathNormEarly = normalizeHttpPath(routePath);
let symbolUid = '';
let symbolName = path.basename(filePath) || 'handler';
let symPath = filePath;
const fileId = row.fileId ?? row[0];
if (fileId && !ambiguousCandidates) {
try {
const syms = await db(CONTAINS_QUERY, { fileId });
if (syms.length > 0) {
const picked = pickSymbolUid(syms, handlerName);
symbolUid = picked.uid;
symbolName = picked.name;
symPath = picked.filePath || filePath;
if (handlerSymbolId) {
// Fast path (Part 2, #2138): the handler symbol was resolved during
// ingestion and persisted on the Route node, so the uid is authoritative
// and we SKIP the source-scan/parse the legacy path needed. Recover the
// display name from the file's symbols via CONTAINING_QUERY (the correct
// File-[DEFINES]->symbol edge — NOT CONTAINS, which is File->Folder).
if (!method) method = 'GET';
symbolUid = handlerSymbolId;
if (filePath) {
try {
const syms = await db(CONTAINING_QUERY, { filePath });
const hit = syms.find((s) => String(s.uid ?? s[0]) === handlerSymbolId);
if (hit) {
symbolName = String(hit.name ?? hit[1]) || symbolName;
symPath = String(hit.filePath ?? hit[2]) || filePath;
}
} catch {
/* keep the authoritative uid + basename fallback */
}
} catch {
/* ignore */
}
} else {
// Legacy fallback (old index / unresolved handler): recover the handler
// from the plugin's scan and resolve it to a real symbol by name (the
// handler/method name) or, for an inline handler, by line-span containment
// — both over File-[DEFINES]->symbol via resolveSymbol. No CONTAINS /
// pickSymbolUid: CONTAINS is File->Folder and the old first-symbol guess
// could win the contractId merge with a wrong uid.
const detections = filePath ? await getDetections(filePath) : [];
const providerDetections = detections.filter((d) => d.role === 'provider');
// Candidates share the same normalized path. When multiple detections at
// the same path exist (GET + POST /api/orders in one router), a blind
// `.find()` silently returned the first verb — attaching the wrong
// handler/method. Disambiguate by method when known; refuse to guess.
const candidates = providerDetections.filter(
(d) => normalizeHttpPath(d.path) === pathNormEarly,
);
let match: (typeof candidates)[number] | undefined;
const ambiguousCandidates = !method && candidates.length > 1;
if (method) {
match = candidates.find((d) => d.method === method);
} else if (candidates.length === 1) {
match = candidates[0];
}
// else: multiple candidates + unknown method → leave match undefined and
// skip symbol enrichment, keeping the file-basename fallback rather than
// guessing the wrong handler.
if (match && !method) method = match.method;
if (!method) method = 'GET';
const resolved =
match && !ambiguousCandidates ? await resolveSymbol(filePath, match) : null;
if (resolved) {
symbolUid = resolved.uid;
symbolName = resolved.name;
symPath = resolved.filePath || filePath;
}
}
const pathNorm = pathNormEarly;
const cid = contractIdFor(method, pathNorm);
out.push({
contractId: cid,
type: 'http',
@@ -432,6 +754,18 @@ export class HttpRouteExtractor implements ContractExtractor {
},
});
}
// Populate the parse-skip coverage set: files whose every provider route
// resolved a handler symbol AND whose language declares complete ingestion
// route coverage. Fail-open — any unresolved row or a 'partial' language
// leaves the file out, so it still gets a full source scan.
if (coveredFiles) {
for (const [fp, allResolved] of fileAllResolved) {
if (allResolved && getPluginForFile(fp)?.routeCoverage === 'complete') {
coveredFiles.add(fp);
}
}
}
return out;
}
@@ -440,26 +774,35 @@ export class HttpRouteExtractor implements ContractExtractor {
private async extractProvidersSourceScan(
files: string[],
getDetections: (rel: string) => Promise<HttpDetection[]>,
resolveSymbol: (filePath: string, d: HttpDetection) => Promise<ResolvedSymbol | null>,
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
for (const rel of files) {
const detections = await getDetections(rel);
const filePath = normalizeRepoRelPath(rel);
for (const d of detections) {
if (d.role !== 'provider') continue;
const pathNorm = normalizeHttpPath(d.path);
// Resolve the handler to a real symbol (named handler, or the inline
// arrow that encloses the registration line) so the contract carries a
// real symbolUid; fall back to the file + detection name otherwise.
const resolved = await resolveSymbol(filePath, d);
out.push({
contractId: contractIdFor(d.method, pathNorm),
type: 'http',
role: 'provider',
symbolUid: '',
symbolRef: { filePath: rel, name: d.name ?? 'handler' },
symbolName: d.name ?? 'handler',
symbolUid: resolved?.uid ?? '',
symbolRef: {
filePath: resolved?.filePath || filePath,
name: resolved?.name ?? d.name ?? 'handler',
},
symbolName: resolved?.name ?? d.name ?? 'handler',
confidence: d.confidence,
meta: {
method: d.method,
path: pathNorm,
pathSegments: pathNorm.split('/').filter(Boolean),
extractionStrategy: 'source_scan',
extractionStrategy: resolved ? 'source_scan_resolved' : 'source_scan',
framework: d.framework,
},
});
@@ -473,6 +816,7 @@ export class HttpRouteExtractor implements ContractExtractor {
private async extractConsumersGraph(
db: CypherExecutor,
getDetections: (rel: string) => Promise<HttpDetection[]>,
resolveSymbol: (filePath: string, d: HttpDetection) => Promise<ResolvedSymbol | null>,
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
let rows: Record<string, unknown>[];
@@ -512,19 +856,19 @@ export class HttpRouteExtractor implements ContractExtractor {
let symbolUid = '';
let symbolName = 'fetch';
let symPath = filePath;
const fileId = row.fileId ?? row[0];
if (fileId) {
try {
const syms = await db(CONTAINS_QUERY, { fileId });
if (syms.length > 0) {
const picked = pickSymbolUid(syms, null);
symbolUid = picked.uid;
symbolName = picked.name;
symPath = picked.filePath || filePath;
}
} catch {
/* ignore */
}
// Resolve the function CONTAINING the fetch by line-span. Do NOT fall back
// to the old `pickSymbolUid(syms, null)` first-symbol-in-file guess: an
// arbitrary wrong uid is worse than an empty one because it would win the
// contractId merge over a correctly-resolved source-scan contract (and the
// empty case degrades to the file-level boundary fallback downstream).
const resolved =
consumerCandidates.length === 1
? await resolveSymbol(filePath, consumerCandidates[0])
: null;
if (resolved) {
symbolUid = resolved.uid;
symbolName = resolved.name;
symPath = resolved.filePath || filePath;
}
out.push({
contractId: cid,
@@ -550,25 +894,31 @@ export class HttpRouteExtractor implements ContractExtractor {
private async extractConsumersSourceScan(
files: string[],
getDetections: (rel: string) => Promise<HttpDetection[]>,
resolveSymbol: (filePath: string, d: HttpDetection) => Promise<ResolvedSymbol | null>,
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
for (const rel of files) {
const detections = await getDetections(rel);
const filePath = normalizeRepoRelPath(rel);
for (const d of detections) {
if (d.role !== 'consumer') continue;
const pathNorm = normalizeConsumerPath(d.path);
// Resolve the function CONTAINING the fetch/axios call so the consumer
// contract carries a real symbolUid (was always '' — the gap that left
// cross-repo trace/impact unable to traverse HTTP links).
const resolved = await resolveSymbol(filePath, d);
out.push({
contractId: contractIdFor(d.method, pathNorm),
type: 'http',
role: 'consumer',
symbolUid: '',
symbolRef: { filePath: rel, name: 'fetch' },
symbolName: 'fetch',
symbolUid: resolved?.uid ?? '',
symbolRef: { filePath: resolved?.filePath || filePath, name: resolved?.name ?? 'fetch' },
symbolName: resolved?.name ?? 'fetch',
confidence: d.confidence,
meta: {
method: d.method,
path: pathNorm,
extractionStrategy: 'source_scan',
extractionStrategy: resolved ? 'source_scan_resolved' : 'source_scan',
framework: d.framework,
},
});
@@ -7,6 +7,21 @@ export interface ManifestExtractResult {
crossLinks: CrossLink[];
}
// Repo-wide symbol lookup for `custom` workspace contracts. Exported so the
// #2325 integration test can run the EXACT production query against a real
// LadybugDB — a hand-copied query string in the test would silently drift
// from this allowlist. Uses the `labels(n) IN [...]` allowlist form rather
// than a `MATCH (n:A|B)` disjunction: this 21-label list contains the
// reserved-keyword labels `Macro` and `Union`, and LadybugDB's parser rejects
// a disjunction that names a reserved keyword (#2325) — which the resolver's
// try/catch then swallowed. `labels(n) IN` has no such collision.
export const CUSTOM_CONTRACT_RESOLVE_QUERY = `MATCH (n)
WHERE labels(n) IN ['Function','Method','Class','Interface','Struct','Enum','Trait','Constructor','TypeAlias','Impl','Macro','Union','Typedef','Property','Record','Delegate','Annotation','Template','Const','Static','CodeElement']
AND n.name = $symbolName
RETURN n.id AS uid, n.name AS name, n.filePath AS filePath
ORDER BY n.filePath ASC
LIMIT 1`;
/**
* Canonicalize an HTTP path for matching against Route.name in the graph.
* Mirrors core/ingestion/pipeline.ts ensureSlash semantics:
@@ -189,13 +204,27 @@ export class ManifestExtractor {
// Cross-impact still works: the bridge query joins on the synthetic
// uid, and the local impact engine derives the same uid for the
// unresolved symbol — name-based hints are the additional safety net.
//
// Label filtering uses `MATCH (n) WHERE labels(n) IN [...]`, NOT the
// openCypher disjunction `MATCH (n:A|B|C)`. LadybugDB's parser rejects a
// disjunction that names a reserved keyword (`Macro` and `Union` both are)
// OR a label with no node table (e.g. the old `lib` branch's `Package`).
// The `custom` branch (reserved keywords in its list) and `lib` branch
// (missing `Package` table) genuinely threw (#2325) and the whole try/catch
// below swallowed it; the other branches parsed but use the same form for
// consistency and future-proofing. `labels(n)` returns the node's single
// label as a string here, so `IN [...]` is an exact allowlist that includes
// listed labels and excludes everything else — and is immune to both
// failure modes (no keyword collision; an unknown label is just a non-match).
try {
let rows: Record<string, unknown>[];
if (link.type === 'http') {
// Route.name is the canonicalized URL path (see
// core/ingestion/pipeline.ts ensureSlash + generateId('Route', ...)).
// Normalize the manifest contract the same way so a user-written
// "/api/orders" matches "api/orders" in the graph.
// Route.name is the canonicalized URL path. Since #2289 a Route node's
// *id* is `(method, url)`-composite (`routeNodeKey`), but `route.name`
// continues to carry the bare URL so URL-keyed group queries like this
// one keep working without a schema change. Normalize the manifest
// contract the same way so a user-written "/api/orders" matches
// "api/orders" in the graph.
//
// The contract may also use the explicit-method form "GET::/api/orders"
// recommended by buildContractId. Strip the METHOD:: prefix before
@@ -220,7 +249,7 @@ export class ManifestExtractor {
// avoid cross-matching Files/Variables/Imports that happen to
// share the topic name.
rows = await executor(
`MATCH (n:Function|Method|Class|Interface) WHERE n.name = $contract
`MATCH (n) WHERE labels(n) IN ['Function','Method','Class','Interface'] AND n.name = $contract
RETURN n.id AS uid, n.name AS name, n.filePath AS filePath
ORDER BY n.filePath ASC
LIMIT 1`,
@@ -244,7 +273,7 @@ export class ManifestExtractor {
const methodName = parts[1]?.trim() ?? '';
if (methodName) {
rows = await executor(
`MATCH (n:Function|Method) WHERE n.name = $methodName
`MATCH (n) WHERE labels(n) IN ['Function','Method'] AND n.name = $methodName
RETURN n.id AS uid, n.name AS name, n.filePath AS filePath
ORDER BY n.filePath ASC
LIMIT 1`,
@@ -252,7 +281,7 @@ export class ManifestExtractor {
);
} else if (serviceName) {
rows = await executor(
`MATCH (n:Class|Interface) WHERE n.name = $serviceName
`MATCH (n) WHERE labels(n) IN ['Class','Interface'] AND n.name = $serviceName
RETURN n.id AS uid, n.name AS name, n.filePath AS filePath
ORDER BY n.filePath ASC
LIMIT 1`,
@@ -264,11 +293,12 @@ export class ManifestExtractor {
} else if (link.type === 'lib') {
// Only exact match on the symbol's name. Previous fallback to
// CONTAINS on n.filePath would promote "react" to "react-native"
// or "@types/react" — silent wrong attribution. Restrict to
// package-level labels so we don't return arbitrary symbols
// named after a library.
// or "@types/react" — silent wrong attribution. Restrict to the
// package-level `Module` label so we don't return arbitrary symbols
// named after a library. (There is no `Package` node table — see
// NODE_TABLES — so a `Package` entry only ever matched nothing.)
rows = await executor(
`MATCH (n:Package|Module) WHERE n.name = $contract
`MATCH (n) WHERE labels(n) IN ['Module'] AND n.name = $contract
RETURN n.id AS uid, n.name AS name, n.filePath AS filePath
ORDER BY n.filePath ASC
LIMIT 1`,
@@ -289,14 +319,7 @@ export class ManifestExtractor {
const symbolName = link.contract.includes('::')
? link.contract.split('::').pop()!
: link.contract;
rows = await executor(
`MATCH (n:Function|Method|Class|Interface|Struct|Enum|Trait|Constructor|TypeAlias|Impl|Macro|Union|Typedef|Property|Record|Delegate|Annotation|Template|Const|Static|CodeElement)
WHERE n.name = $symbolName
RETURN n.id AS uid, n.name AS name, n.filePath AS filePath
ORDER BY n.filePath ASC
LIMIT 1`,
{ symbolName },
);
rows = await executor(CUSTOM_CONTRACT_RESOLVE_QUERY, { symbolName });
} else {
return null;
}
+78
View File
@@ -90,6 +90,79 @@ export interface GroupToolPort {
include_content?: boolean;
},
): Promise<unknown>;
// ── Cross-repo trace support (optional on the port) ────────────────
// These are optional so existing GroupToolPort test mocks (which predate
// the trace path and only stub impact/query/context/impactByUid) keep
// type-checking. The real LocalBackend port supplies all three; runGroupTrace
// guards on their presence and degrades to a clear error/note when absent.
//
// Single-repo directed-path trace over CALLS + HAS_METHOD. Returns the same
// shape as the `trace` MCP tool (`{ status, from, to, hopCount, hops, edges }`).
trace?(
repo: GroupRepoHandle,
params: {
from?: string;
to?: string;
from_uid?: string;
to_uid?: string;
from_file?: string;
to_file?: string;
maxDepth?: number;
includeTests?: boolean;
},
): Promise<unknown>;
// Resolve a symbol within one repo to its node id (== bridge symbolUid) and
// location, or report ambiguity / absence. Wraps the same resolver the
// context()/trace() tools use.
resolveSymbol?(
repo: GroupRepoHandle,
query: { name?: string; uid?: string; file_path?: string },
): Promise<GroupSymbolResolution>;
// Intra-procedural REACHING_DEF data-flow from an anchor symbol, used to
// enrich a boundary-adjacent trace segment. `available:false` signals the
// repo has no PDG `flows` layer (degraded, not an error).
pdgFlows?(
repo: GroupRepoHandle,
anchor: { name?: string; uid?: string; file_path?: string },
opts: { limit?: number },
): Promise<GroupPdgFlowResult>;
}
export type GroupSymbolResolution =
| {
kind: 'ok';
symbol: {
id: string;
name: string;
type: string;
filePath: string;
startLine: number;
endLine: number;
};
}
| {
kind: 'ambiguous';
candidates: Array<{
id: string;
name: string;
type: string;
filePath: string;
startLine: number;
}>;
}
| { kind: 'not_found' };
export interface GroupPdgFlowHop {
line: number;
text: string;
variable?: string;
}
export interface GroupPdgFlowResult {
available: boolean;
variable?: string;
hops: GroupPdgFlowHop[];
truncated?: boolean;
}
function isStoredContract(raw: unknown): raw is StoredContract {
@@ -313,6 +386,11 @@ export class GroupService {
return runGroupImpact({ port: this.port, gitnexusDir: getDefaultGitnexusDir() }, params);
}
async groupTrace(params: Record<string, unknown>): Promise<unknown> {
const { runGroupTrace } = await import('./cross-trace.js');
return runGroupTrace({ port: this.port, gitnexusDir: getDefaultGitnexusDir() }, params);
}
async groupContext(params: Record<string, unknown>): Promise<GroupContextResult> {
const name = String(params.name ?? '').trim();
const target = typeof params.target === 'string' ? params.target.trim() : '';
+7
View File
@@ -192,6 +192,13 @@ export interface BridgeHandle {
readonly _db: unknown;
readonly _conn: unknown;
readonly groupDir: string;
/**
* True when the handle was opened read-only. `closeBridgeDb` must NOT issue a
* CHECKPOINT on a read-only connection — doing so leaves a WAL/shadow lock
* artifact that makes the next read-only open of the same file fail in-process
* (repeated `@group` impact/trace calls in a long-lived server).
*/
readonly _readOnly?: boolean;
}
export interface BridgeMeta {
+119 -13
View File
@@ -21,7 +21,13 @@ import { generateId } from '../../lib/utils.js';
import type { SymbolDefinition } from 'gitnexus-shared';
import { yieldToEventLoop } from './utils/event-loop.js';
import type { ExtractedRoute, ExtractedFetchCall } from './workers/parse-worker.js';
import type { ExtractedDecoratorRoute } from './workers/parse-worker.js';
import { normalizeFetchURL, routeMatches } from './route-extractors/nextjs.js';
import {
normalizeExtractedRoutePath,
normalizeRouteMethod,
routeNodeKey,
} from './route-extractors/route-path.js';
import { extractReturnTypeName } from './type-extractors/shared.js';
const MAX_EXPORTS_PER_FILE = 500;
@@ -243,6 +249,93 @@ export const processRoutesFromExtracted = async (
onProgress?.(extractedRoutes.length, extractedRoutes.length);
};
/**
* Resolve each route's handler to a real symbol UID, keyed by the route's
* `(method, url)` identity (`routeNodeKey` — the same key the routes phase uses
* for the `Route` node). This is the Part 2 (#2138) groundwork that lets
* `HttpRouteExtractor.extractProvidersGraph` read the handler symbol from the
* graph instead of re-parsing source via `getDetections()`.
*
* Two route shapes, one resolution target — `(filePath, name) → nodeId`:
* - Laravel framework routes (`ExtractedRoute`) carry `controllerName` +
* `methodName`; resolve the controller (qualified-first) then the method in
* the controller's own file (mirrors `processRoutesFromExtracted`).
* - Decorator routes (`ExtractedDecoratorRoute`, e.g. Spring/FastAPI) carry
* `handlerName` (the decorated method, captured at extraction); resolve it
* directly in the route's own file.
*
* First-writer-wins per route identity, matching the routes phase's dedup (it
* keeps the first route registered for a `(method, url)` key and counts the rest
* as duplicates). The first route to claim a key reserves it **even when its
* handler is unresolvable**, so a later same-key route can never stamp its
* handler onto the first route's Route node (the routes phase made that first
* route the node-winner). Keying is `routeNodeKey(method, url)` (#2289): a
* same-URL multi-verb pair (`GET /x` + `POST /x`) resolves two handlers, one per
* node; method-less / wildcard routes key by URL alone, byte-identical to the
* pre-#2289 behavior. Routes whose handler cannot be *uniquely* resolved (no
* name, zero matches, or an ambiguous same-name match) carry no
* `handlerSymbolId`; the extractor then falls back to source scan for that route
* (fail-open, no regression, never a wrong handler).
*/
export function resolveRouteHandlerSymbols(
model: SemanticModel,
extractedRoutes: readonly ExtractedRoute[],
decoratorRoutes: readonly ExtractedDecoratorRoute[],
): Map<string, string> {
const out = new Map<string, string>();
// Route identities already claimed by an earlier route (resolved or not).
// Mirrors the routes phase `addRoute` first-writer-wins so the handler we
// stamp always belongs to the route that actually won the Route node.
const claimed = new Set<string>();
// Resolve a single same-file symbol by name, refusing to guess on ambiguity:
// exactly one match → its nodeId; zero or many → undefined (fail-open).
const uniqueSymbolId = (filePath: string, name: string): string | undefined => {
const defs = model.symbols.lookupExactAll(filePath, name);
return defs.length === 1 ? defs[0]?.nodeId : undefined;
};
const claim = (
routePath: string | null,
prefix: string | null,
httpMethod: string | null | undefined,
symbolId: string | undefined,
) => {
if (!routePath) return;
const url = normalizeExtractedRoutePath(routePath, prefix);
const key = routeNodeKey(normalizeRouteMethod(httpMethod), url);
if (claimed.has(key)) return; // first-writer-wins: later same-key routes can't override
claimed.add(key);
if (symbolId) out.set(key, symbolId);
};
// Laravel framework routes — controller class + method name.
for (const route of extractedRoutes) {
let methodId: string | undefined;
if (route.controllerName && route.methodName) {
let controllerDef: SymbolDefinition | undefined;
if (route.controllerQualifiedName) {
controllerDef = resolveControllerByQualifiedName(model, route.controllerQualifiedName);
}
if (!controllerDef) {
const controllerDefs = model.types.lookupClassByName(route.controllerName);
if (controllerDefs.length === 1) controllerDef = controllerDefs[0];
}
if (controllerDef) methodId = uniqueSymbolId(controllerDef.filePath, route.methodName);
}
claim(route.routePath, route.prefix ?? null, route.httpMethod, methodId);
}
// Decorator routes (Spring / FastAPI / generic) — the decorated handler in
// the route's own file.
for (const dr of decoratorRoutes) {
const handlerId = dr.handlerName ? uniqueSymbolId(dr.filePath, dr.handlerName) : undefined;
claim(dr.routePath, dr.prefix ?? null, dr.httpMethod, handlerId);
}
return out;
}
/** Common method names on response/data objects that are NOT property accesses */
// Properties/methods to ignore when extracting consumer accessed keys from `data.X` patterns.
// Avoids false positives from Fetch API, Array, Object, Promise, and DOM access on variables
@@ -386,19 +479,29 @@ export const extractConsumerAccessedKeys = (content: string): string[] => {
* Create FETCHES edges from extracted fetch() calls to matching Route nodes.
* When consumerContents is provided, extracts property access patterns from
* consumer files and encodes them in the edge reason field.
*
* Matching stays URL-only (#2289): a verb-less consumer (a `fetch()` call has
* no statically-known HTTP method) matches a route by URL and connects to
* **every** Route node sharing that URL — i.e. both the `GET /x` and `POST /x`
* nodes when a URL carries multiple verbs. `routeUrlToKeys` therefore maps each
* route URL to the list of `routeNodeKey` identities at that URL; a single-verb
* (or method-less) URL has a one-element list, keeping edges byte-identical to
* the pre-#2289 behavior.
*/
export const processNextjsFetchRoutes = (
graph: KnowledgeGraph,
fetchCalls: ExtractedFetchCall[],
routeRegistry: Map<string, string>, // routeURL → handlerFilePath
routeUrlToKeys: Map<string, string[]>, // routeURL → route node keys at that URL
consumerContents?: Map<string, string>, // filePath → file content
) => {
// Pre-count how many routes each consumer file matches (for confidence attribution)
// Pre-count how many route URLs each consumer file matches (for confidence
// attribution). Counts once per call that matches any URL — independent of how
// many verbs share that URL — so the multi-fetch heuristic is unchanged.
const routeCountByFile = new Map<string, number>();
for (const call of fetchCalls) {
const normalized = normalizeFetchURL(call.fetchURL);
if (!normalized) continue;
for (const [routeURL] of routeRegistry) {
for (const routeURL of routeUrlToKeys.keys()) {
if (routeMatches(normalized, routeURL)) {
routeCountByFile.set(call.filePath, (routeCountByFile.get(call.filePath) ?? 0) + 1);
break;
@@ -410,10 +513,9 @@ export const processNextjsFetchRoutes = (
const normalized = normalizeFetchURL(call.fetchURL);
if (!normalized) continue;
for (const [routeURL] of routeRegistry) {
for (const [routeURL, routeKeys] of routeUrlToKeys) {
if (routeMatches(normalized, routeURL)) {
const sourceId = generateId('File', call.filePath);
const routeNodeId = generateId('Route', routeURL);
// Extract consumer accessed keys if file content is available
let reason = 'fetch-url-match';
@@ -433,14 +535,18 @@ export const processNextjsFetchRoutes = (
reason = `${reason}|fetches:${fetchCount}`;
}
graph.addRelationship({
id: generateId('FETCHES', `${sourceId}->${routeNodeId}`),
sourceId,
targetId: routeNodeId,
type: 'FETCHES',
confidence: 0.9,
reason,
});
// Connect to every Route node at this URL (one per verb).
for (const routeKey of routeKeys) {
const routeNodeId = generateId('Route', routeKey);
graph.addRelationship({
id: generateId('FETCHES', `${sourceId}->${routeNodeId}`),
sourceId,
targetId: routeNodeId,
type: 'FETCHES',
confidence: 0.9,
reason,
});
}
break;
}
}
@@ -38,6 +38,7 @@ import type { SyntaxNode } from './utils/ast-helpers.js';
import type { CfgVisitor } from './cfg/types.js';
import type { NodeLabel } from 'gitnexus-shared';
import type { ExtractedRoute } from './route-extractors/laravel.js';
import type { SharedSpringType } from './route-extractors/spring-shared.js';
import type Parser from 'tree-sitter';
import type { ExtractedDecoratorRoute } from './workers/parse-worker.js';
@@ -288,6 +289,23 @@ interface LanguageProviderConfig {
lineOffset: number,
) => ExtractedDecoratorRoute[];
/**
* Collect a project-wide, language-agnostic view of route-defining
* class/interface declarations (`SharedSpringType`) from a parsed file.
*
* When defined, the parse worker calls this per file and the parse phase
* aggregates the results, then runs a cross-file pass that resolves
* interface-inherited routes (a concrete controller inherits the `@*Mapping`s
* its interfaces declare) and appends them to `decoratorRoutes`. Separate from
* `extractDecoratorRoutes` because inheritance needs all files, not one.
*
* Default: undefined (no interface-inheritance route resolution).
*/
readonly extractRouteInheritanceTypes?: (
tree: Parser.Tree,
filePath: string,
) => SharedSpringType[];
// ── Noise filtering ────────────────────────────────────────────────
/** Built-in/stdlib names that should be filtered from the call graph for this language.
* Default: undefined (no language-specific filtering). */
@@ -31,6 +31,7 @@ const FUNCTION_DECLARATION_TYPES = new Set([
'function_item',
]);
import type { SyntaxNode } from '../utils/ast-helpers.js';
import { createLeadingDocDescriptionExtractor } from '../utils/ast-helpers.js';
import type { NodeLabel } from 'gitnexus-shared';
import type { LanguageProvider } from '../language-provider.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
@@ -397,6 +398,8 @@ export const cProvider = defineLanguage({
}),
variableExtractor: createVariableExtractor(cVariableConfig),
classExtractor: cClassExtractor,
// ── Doxygen doc comment → description (issue #2270) ──
descriptionExtractor: createLeadingDocDescriptionExtractor(),
labelOverride: cppLabelOverride,
builtInNames: C_BUILT_INS,
@@ -482,6 +485,8 @@ export const cppProvider = defineLanguage({
}),
variableExtractor: createVariableExtractor(cppVariableConfig),
classExtractor: cppClassExtractor,
// ── Doxygen doc comment → description (issue #2270) ──
descriptionExtractor: createLeadingDocDescriptionExtractor(),
labelOverride: cppLabelOverride,
builtInNames: C_BUILT_INS,
extractTemplateConstraints: extractCppTemplateConstraintsForProvider,
@@ -15,6 +15,7 @@ import { csharpExportChecker } from '../export-detection.js';
import { createImportResolver } from '../import-resolvers/resolver-factory.js';
import { csharpImportConfig } from '../import-resolvers/configs/csharp.js';
import { CSHARP_QUERIES } from '../tree-sitter-queries.js';
import { createLeadingDocDescriptionExtractor } from '../utils/ast-helpers.js';
import type { AstFrameworkPatternConfig } from '../language-provider.js';
import { createCallExtractor } from '../call-extractors/generic.js';
import { csharpCallConfig } from '../call-extractors/configs/csharp.js';
@@ -194,6 +195,8 @@ export const csharpProvider = defineLanguage({
methodExtractor: createMethodExtractor(csharpMethodConfig),
variableExtractor: createVariableExtractor(csharpVariableConfig),
classExtractor: createClassExtractor(csharpClassConfig),
// ── XML doc comments (`///`) → description (issue #2270) ──
descriptionExtractor: createLeadingDocDescriptionExtractor(),
builtInNames: BUILT_INS,
// ── RFC #909 Ring 3: scope-based resolution hooks (RFC §5) ──────────
@@ -9,9 +9,12 @@
* The hook resolves the enclosing function by inspecting the previous sibling.
*/
import type { SyntaxNode } from '../utils/ast-helpers.js';
import {
createLeadingDocDescriptionExtractor,
FUNCTION_NODE_TYPES,
type SyntaxNode,
} from '../utils/ast-helpers.js';
import type { NodeLabel } from 'gitnexus-shared';
import { FUNCTION_NODE_TYPES } from '../utils/ast-helpers.js';
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { dartClassConfig } from '../class-extractors/configs/dart.js';
@@ -124,6 +127,8 @@ export const dartProvider = defineLanguage({
methodExtractor: createMethodExtractor(dartMethodConfig),
variableExtractor: createVariableExtractor(dartVariableConfig),
classExtractor: createClassExtractor(dartClassConfig),
// ── Dartdoc (`///`) → description (issue #2270) ──
descriptionExtractor: createLeadingDocDescriptionExtractor(),
enclosingFunctionFinder: dartEnclosingFunctionFinder,
builtInNames: DART_BUILT_INS,
@@ -11,6 +11,7 @@
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { goClassConfig } from '../class-extractors/configs/go.js';
import { createLeadingDocDescriptionExtractor } from '../utils/ast-helpers.js';
import { createGoCfgVisitor } from '../cfg/visitors/go.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as goConfig } from '../type-extractors/go.js';
@@ -138,6 +139,12 @@ export const goProvider = defineLanguage({
methodExtractor: createMethodExtractor(goMethodConfig),
variableExtractor: createVariableExtractor(goVariableConfig),
classExtractor: createClassExtractor(goClassConfig),
// ── godoc (`//` leading comments) → description (issue #2270). Build/tool
// directives (//go:…, // +build, //nolint, //line) are not documentation. ──
descriptionExtractor: createLeadingDocDescriptionExtractor({
lineCommentPrefixes: ['//'],
lineDirectivePrefixes: ['//go:', '// +build', '//nolint', '//line'],
}),
builtInNames: GO_BUILT_INS,
// ── RFC #909 Ring 3: scope-based resolution hooks ──────────
@@ -12,8 +12,9 @@ import { createClassExtractor } from '../class-extractors/generic.js';
import { javaClassConfig } from '../class-extractors/configs/jvm.js';
import { defineLanguage } from '../language-provider.js';
import type { AstFrameworkPatternConfig } from '../language-provider.js';
import { createLeadingDocDescriptionExtractor } from '../utils/ast-helpers.js';
import { javaTypeConfig } from '../type-extractors/jvm.js';
import { extractSpringRoutes } from '../route-extractors/spring.js';
import { extractSpringRoutes, extractSpringTypes } from '../route-extractors/spring.js';
import { javaExportChecker } from '../export-detection.js';
import { createImportResolver } from '../import-resolvers/resolver-factory.js';
import { javaImportConfig } from '../import-resolvers/configs/jvm.js';
@@ -117,6 +118,9 @@ export const javaProvider = defineLanguage({
variableExtractor: createVariableExtractor(javaVariableConfig),
classExtractor: createClassExtractor(javaClassConfig),
// ── Javadoc → description (issue #2270) ──
descriptionExtractor: createLeadingDocDescriptionExtractor(),
// ── RFC #909 Ring 3: scope-based resolution hooks ──
emitScopeCaptures: emitJavaScopeCaptures,
@@ -134,4 +138,5 @@ export const javaProvider = defineLanguage({
// ── Route extraction ──
extractDecoratorRoutes: extractSpringRoutes,
extractRouteInheritanceTypes: extractSpringTypes,
});
@@ -11,6 +11,7 @@ import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { kotlinClassConfig } from '../class-extractors/configs/jvm.js';
import { defineLanguage } from '../language-provider.js';
import { createLeadingDocDescriptionExtractor } from '../utils/ast-helpers.js';
import { assertCloneable } from '../workers/clone-safety.js';
import { kotlinTypeConfig } from '../type-extractors/jvm.js';
import { kotlinExportChecker } from '../export-detection.js';
@@ -170,6 +171,10 @@ export const kotlinProvider = defineLanguage({
variableExtractor: createVariableExtractor(kotlinVariableConfig),
classExtractor: createClassExtractor(kotlinClassConfig),
builtInNames: BUILT_INS,
// ── KDoc → description (issue #2270) ──
descriptionExtractor: createLeadingDocDescriptionExtractor(),
labelOverride: (functionNode, defaultLabel) => {
if (defaultLabel !== 'Function') return defaultLabel;
if (isKotlinClassMethod(functionNode)) return 'Method';
+28 -9
View File
@@ -21,13 +21,22 @@ import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { phpClassConfig } from '../class-extractors/configs/php.js';
import { createPhpCfgVisitor } from '../cfg/visitors/php.js';
import { defineLanguage, type AstFrameworkPatternConfig } from '../language-provider.js';
import {
defineLanguage,
type AstFrameworkPatternConfig,
type CaptureMap,
} from '../language-provider.js';
import { typeConfig as phpConfig } from '../type-extractors/php.js';
import { phpExportChecker } from '../export-detection.js';
import { createImportResolver } from '../import-resolvers/resolver-factory.js';
import { phpImportConfig } from '../import-resolvers/configs/php.js';
import { PHP_QUERIES } from '../tree-sitter-queries.js';
import { findDescendant, extractStringContent, type SyntaxNode } from '../utils/ast-helpers.js';
import {
findDescendant,
extractStringContent,
createLeadingDocDescriptionExtractor,
type SyntaxNode,
} from '../utils/ast-helpers.js';
import type { NodeLabel } from 'gitnexus-shared';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { phpConfig as phpFieldConfig } from '../field-extractors/configs/php.js';
@@ -221,22 +230,32 @@ function extractEloquentRelationDescription(methodNode: SyntaxNode): string | nu
return null;
}
/** PHPDoc-docblock fallback, shared with the other leading-comment languages. */
const phpLeadingDocFallback = createLeadingDocDescriptionExtractor();
/**
* LanguageProvider.descriptionExtractor implementation for PHP.
* Extracts Eloquent model property metadata and relationship descriptions.
* Eloquent model property metadata and relationship descriptions take
* precedence (they are richer than prose); otherwise documentable symbols fall
* back to their leading PHPDoc docblock (issue #2270), mirroring the other
* leading-comment languages.
*/
function phpDescriptionExtractor(
nodeLabel: NodeLabel,
nodeName: string,
captureMap: Record<string, SyntaxNode>,
captureMap: CaptureMap,
): string | undefined {
if (nodeLabel === 'Property' && captureMap['definition.property']) {
return extractPhpPropertyDescription(nodeName, captureMap['definition.property']) ?? undefined;
const propertyNode = captureMap['definition.property'];
if (nodeLabel === 'Property' && propertyNode) {
const eloquentProperty = extractPhpPropertyDescription(nodeName, propertyNode);
if (eloquentProperty) return eloquentProperty;
}
if (nodeLabel === 'Method' && captureMap['definition.method']) {
return extractEloquentRelationDescription(captureMap['definition.method']) ?? undefined;
const methodNode = captureMap['definition.method'];
if (nodeLabel === 'Method' && methodNode) {
const eloquentRelation = extractEloquentRelationDescription(methodNode);
if (eloquentRelation) return eloquentRelation;
}
return undefined;
return phpLeadingDocFallback(nodeLabel, nodeName, captureMap);
}
/** Detect Laravel route files by path convention. */
+15 -1
View File
@@ -13,7 +13,7 @@ import { createClassExtractor } from '../class-extractors/generic.js';
import { rubyClassConfig } from '../class-extractors/configs/ruby.js';
import { defineLanguage } from '../language-provider.js';
import type { AstFrameworkPatternConfig } from '../language-provider.js';
import type { SyntaxNode } from '../utils/ast-helpers.js';
import { createLeadingDocDescriptionExtractor, type SyntaxNode } from '../utils/ast-helpers.js';
import { typeConfig as rubyConfig } from '../type-extractors/ruby.js';
import { routeRubyCall } from '../call-routing.js';
import { rubyExportChecker } from '../export-detection.js';
@@ -197,6 +197,20 @@ export const rubyProvider = defineLanguage({
}),
variableExtractor: createVariableExtractor(rubyVariableConfig),
classExtractor: createClassExtractor(rubyClassConfig),
// ── Leading `#` comments (RDoc/YARD) → description (issue #2270). Magic
// comments and the shebang are not documentation. ──
descriptionExtractor: createLeadingDocDescriptionExtractor({
lineCommentPrefixes: ['#'],
lineDirectivePrefixes: [
'# frozen_string_literal:',
'# encoding:',
'# coding:',
'# -*-',
'#!',
'# rubocop:',
'# typed:',
],
}),
labelOverride: rubyLabelOverride,
// Ruby MRO is kind-aware: prepend providers beat the class's own method,
// which in turn beats include providers. The graph-level MRO phase
@@ -13,7 +13,7 @@ import type { NodeLabel } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { rustClassConfig } from '../class-extractors/configs/rust.js';
import { defineLanguage } from '../language-provider.js';
import type { SyntaxNode } from '../utils/ast-helpers.js';
import { createLeadingDocDescriptionExtractor, type SyntaxNode } from '../utils/ast-helpers.js';
import { typeConfig as rustConfig } from '../type-extractors/rust.js';
import { rustExportChecker } from '../export-detection.js';
import { createImportResolver } from '../import-resolvers/resolver-factory.js';
@@ -176,6 +176,13 @@ export const rustProvider = defineLanguage({
}),
variableExtractor: createVariableExtractor(rustVariableConfig),
classExtractor: createClassExtractor(rustClassConfig),
// ── Rust outer doc comments (`///`, `/** */`) → description (issue #2270).
// `//!` / `/*!` are INNER docs (document the enclosing item), so they must
// not attach to the following item — opt out of both. ──
descriptionExtractor: createLeadingDocDescriptionExtractor({
lineCommentPrefixes: ['///'],
blockDocPrefixes: ['/**'],
}),
builtInNames: BUILT_INS,
// ── RFC #909 Ring 3: scope-based resolution hooks ──────────
emitScopeCaptures: emitRustScopeCaptures,
@@ -17,6 +17,7 @@ import { createImportResolver } from '../import-resolvers/resolver-factory.js';
import { swiftImportConfig } from '../import-resolvers/configs/swift.js';
import { SWIFT_QUERIES } from '../tree-sitter-queries.js';
import type { SyntaxNode } from '../utils/ast-helpers.js';
import { createLeadingDocDescriptionExtractor } from '../utils/ast-helpers.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { swiftConfig as swiftFieldConfig } from '../field-extractors/configs/swift.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
@@ -243,6 +244,8 @@ export const swiftProvider = defineLanguage({
}),
variableExtractor: createVariableExtractor(swiftVariableConfig),
classExtractor: createClassExtractor(swiftClassConfig),
// ── Swift doc comments (`///`, `/** */`) → description (issue #2270) ──
descriptionExtractor: createLeadingDocDescriptionExtractor(),
orderSameNameTypeCandidates: orderSwiftSameNameTypeCandidates,
builtInNames: BUILT_INS,
// ── Scope-based resolution hooks (RFC #909 Ring 3, issue #937). See
@@ -16,6 +16,7 @@ import {
javascriptClassConfig,
} from '../class-extractors/configs/typescript-javascript.js';
import type { SyntaxNode } from '../utils/ast-helpers.js';
import { createLeadingDocDescriptionExtractor } from '../utils/ast-helpers.js';
import { createTypeScriptCfgVisitor } from '../cfg/visitors/typescript.js';
import { typeConfig as typescriptConfig } from '../type-extractors/typescript.js';
import { tsExportChecker } from '../export-detection.js';
@@ -344,6 +345,11 @@ export const typescriptProvider = defineLanguage({
}),
variableExtractor: createVariableExtractor(typescriptVariableConfig),
classExtractor: createClassExtractor(typescriptClassConfig),
// ── JSDoc → description (issue #2270). An exported decl is captured as the
// inner declaration; its JSDoc precedes the wrapping `export_statement`. ──
descriptionExtractor: createLeadingDocDescriptionExtractor({
wrapperNodeTypes: ['export_statement'],
}),
builtInNames: BUILT_INS,
// ── RFC #909 Ring 3: scope-based resolution hooks (RFC §5) ──────────
@@ -406,6 +412,11 @@ export const javascriptProvider = defineLanguage({
}),
variableExtractor: createVariableExtractor(javascriptVariableConfig),
classExtractor: createClassExtractor(javascriptClassConfig),
// ── JSDoc → description (issue #2270). An exported decl is captured as the
// inner declaration; its JSDoc precedes the wrapping `export_statement`. ──
descriptionExtractor: createLeadingDocDescriptionExtractor({
wrapperNodeTypes: ['export_statement'],
}),
builtInNames: BUILT_INS,
// ── RFC #909 Ring 3: scope-based resolution hooks (RFC §5) ──────────
@@ -0,0 +1,377 @@
/**
* Nuxt v4 auto-import resolution for the TypeScript scope-resolver.
*
* Nuxt and Nitro both make symbols available project-wide without explicit
* import statements. Two sources cover the common cases:
*
* 1. `.nuxt/imports.d.ts` - client/shared composables generated by Nuxt at
* build time. Only entries whose source is a project-local relative path
* (i.e. NOT a package in node_modules or a Nuxt runtime alias) are
* indexed. These are project utils and composables the developer wrote.
*
* 2. `server/utils/**` - Nitro auto-imports all exports from this directory
* tree into every `server/api/`, `server/routes/`, and
* `server/middleware/` file. These are not included in `imports.d.ts`
* because they are server-only.
*
* The two sources are kept in separate maps because their visibility scopes
* differ: app/client files do not see `server/utils`, while Nitro server
* handlers prefer same-named server utilities over client composables.
*
* Detection: if `.nuxt/imports.d.ts` is absent the repo is not a Nuxt project
* and this module returns null immediately, adding zero overhead to non-Nuxt
* analysis runs. The `server/utils` scan is only attempted when the Nuxt
* detection succeeds, so it also costs nothing for non-Nuxt repos.
*
* Limitations:
* - Auto-import edges are still heuristic rather than type-checked. The
* emitter tags them with confidence 0.75.
* - Re-exports with non-matching aliases (e.g. `flatUnwrap as unwrapSlot`
* where `flatUnwrap` is the graph node name) try both the original export
* name and the local alias when looking up the graph node.
* - Only direct exports are indexed; barrel re-exports that chain through
* multiple files are not followed.
*/
import fs from 'fs/promises';
import path from 'path';
// ---- types ------------------------------------------------------------------
export type NuxtAutoImportScope = 'client' | 'server';
/** A single auto-imported symbol and where it lives in the repo. */
export interface NuxtAutoImportEntry {
/** The name used at call sites (the alias when `export { X as Y }` form). */
readonly localName: string;
/** The original export name in the source file (before `as`). */
readonly exportName: string;
/** Repo-relative POSIX path to the source file (with extension). */
readonly sourceFile: string;
/** Which Nuxt/Nitro visibility scope supplied this entry. */
readonly scope: NuxtAutoImportScope;
}
/**
* Aggregated auto-import maps for one Nuxt workspace, keyed by the local
* name used in calling code. Client/shared entries come from
* `.nuxt/imports.d.ts`; server entries come from `server/utils`.
*/
export interface NuxtAutoImportConfig {
readonly clientByLocalName: ReadonlyMap<string, NuxtAutoImportEntry>;
readonly serverByLocalName: ReadonlyMap<string, NuxtAutoImportEntry>;
}
const FILE_EXTENSIONS = ['.ts', '.tsx', '.js', '.jsx'] as const;
const GENERATED_DIR_NAMES = new Set([
'.git',
'.nuxt',
'.output',
'.next',
'build',
'coverage',
'dist',
'node_modules',
]);
const IMPORTS_DTS_EXPORT_RE = /^export\s*\{([^}]+)\}\s*from\s*['"]([^'"]+)['"]/gm;
const NITRO_DECLARATION_EXPORT_RE =
/^export\s+(?:default\s+)?(?:async\s+)?function\s+([A-Za-z_$][A-Za-z0-9_$]*)|^export\s+(?:default\s+)?class\s+([A-Za-z_$][A-Za-z0-9_$]*)/gm;
const NITRO_VARIABLE_EXPORT_RE = /^export\s+(?:const|let|var)\s+([^;\n]+)/gm;
// Matches the binding name at the head of a single declarator (`name`, `name:`, `name =`).
// A leading `{`/`[` (destructuring) does not match, so destructured exports are skipped.
const DECLARATOR_NAME_RE = /^\s*([A-Za-z_$][A-Za-z0-9_$]*)/;
// ---- loader -----------------------------------------------------------------
/**
* Load the Nuxt auto-import map for `repoRoot`.
*
* Returns null when:
* - `.nuxt/imports.d.ts` does not exist (non-Nuxt project or pre-build), or
* - the file exists but yields zero project-local entries and `server/utils`
* is absent or empty.
*
* The `server/utils` scan is only attempted when `.nuxt/imports.d.ts` was
* successfully read, confirming this is an initialized Nuxt project. This
* avoids partial results from repos that have a `server/utils` directory but
* are not Nuxt projects.
*/
export async function loadNuxtAutoImports(repoRoot: string): Promise<NuxtAutoImportConfig | null> {
const clientByLocalName = new Map<string, NuxtAutoImportEntry>();
const serverByLocalName = new Map<string, NuxtAutoImportEntry>();
const nuxtInitialized = await collectImportsDts(repoRoot, clientByLocalName);
// Only scan server/utils when imports.d.ts was present, confirming this is
// an initialized Nuxt project. Without this gate, a non-Nuxt repo with a
// server/utils directory would get spurious Nitro auto-import edges.
if (nuxtInitialized) {
await collectNitroServerUtils(repoRoot, serverByLocalName);
}
const config = { clientByLocalName, serverByLocalName };
return hasNuxtAutoImports(config) ? config : null;
}
export function hasNuxtAutoImports(config: NuxtAutoImportConfig): boolean {
return config.clientByLocalName.size > 0 || config.serverByLocalName.size > 0;
}
export function getNuxtAutoImportEntry(
config: NuxtAutoImportConfig,
localName: string,
callerFile: string,
): NuxtAutoImportEntry | undefined {
if (isNitroServerRuntimeFile(callerFile)) {
// Nitro auto-imports only `server/utils/**` into the server context; app
// `composables/` are Vue-app-only. Server callers therefore resolve the
// server map only — no client fallback, which would emit cross-context
// CALLS edges Nitro never actually creates.
return config.serverByLocalName.get(localName);
}
return config.clientByLocalName.get(localName);
}
/**
* Directories whose files run in the Nitro server context and so receive
* `server/utils` auto-imports without an explicit import: API routes, server
* routes, middleware, plugins, and (Nitro 2.6+) tasks.
*/
const NITRO_RUNTIME_DIR_PREFIXES = [
'server/api/',
'server/routes/',
'server/middleware/',
'server/plugins/',
'server/tasks/',
] as const;
export function isNitroServerRuntimeFile(filePath: string): boolean {
const normalized = filePath.replace(/\\/g, '/');
return NITRO_RUNTIME_DIR_PREFIXES.some((prefix) => normalized.startsWith(prefix));
}
// ---- .nuxt/imports.d.ts -----------------------------------------------------
/**
* Parse `.nuxt/imports.d.ts` and add project-local entries to `byLocalName`.
*
* The file contains lines of the form:
* export { name1, name2, origName as alias } from 'source'
*
* Only entries with a source that is a project-local relative path
* (starts with `./` or `../` AND does not contain `node_modules`) are
* included. Nuxt runtime paths (`#app/...`) and third-party packages are
* intentionally skipped because they have no graph nodes in the repo.
*
* @returns true when `.nuxt/imports.d.ts` was successfully read.
*/
async function collectImportsDts(
repoRoot: string,
byLocalName: Map<string, NuxtAutoImportEntry>,
): Promise<boolean> {
const importsPath = path.join(repoRoot, '.nuxt', 'imports.d.ts');
let content: string;
try {
content = await fs.readFile(importsPath, 'utf-8');
} catch {
return false;
}
const nuxtDir = path.join(repoRoot, '.nuxt');
IMPORTS_DTS_EXPORT_RE.lastIndex = 0;
let m: RegExpExecArray | null;
while ((m = IMPORTS_DTS_EXPORT_RE.exec(content)) !== null) {
const symbolsRaw = m[1]!;
const source = m[2]!;
if (!isProjectLocalPath(source)) continue;
// Containment guard: a crafted source (`from '../../../../etc/passwd'`)
// passes the relative-path check but escapes the repo. Skip anything that
// resolves outside repoRoot before touching the filesystem.
const resolvedBase = path.resolve(nuxtDir, source);
if (!isWithinRepo(repoRoot, resolvedBase)) continue;
const resolvedFile = await resolveExtension(resolvedBase);
if (resolvedFile === null) continue;
const sourceFile = toRepoPosix(repoRoot, resolvedFile);
for (const raw of symbolsRaw.split(',')) {
const trimmed = raw.trim();
if (!trimmed) continue;
const asIdx = trimmed.indexOf(' as ');
const exportName = asIdx >= 0 ? trimmed.slice(0, asIdx).trim() : trimmed;
const localName = asIdx >= 0 ? trimmed.slice(asIdx + 4).trim() : trimmed;
if (!localName || !exportName) continue;
if (!byLocalName.has(localName)) {
byLocalName.set(localName, { localName, exportName, sourceFile, scope: 'client' });
}
}
}
return true;
}
// ---- server/utils (Nitro auto-imports) --------------------------------------
/**
* Scan `server/utils/**` and register every exported symbol as a Nitro
* auto-import. Nitro makes all exports from this directory tree available
* without an import statement in server routes and middleware.
*/
async function collectNitroServerUtils(
repoRoot: string,
byLocalName: Map<string, NuxtAutoImportEntry>,
): Promise<void> {
const serverUtilsDir = path.join(repoRoot, 'server', 'utils');
if (!(await dirExists(serverUtilsDir))) return;
const tsFiles = await collectTsFiles(serverUtilsDir);
for (const absPath of tsFiles) {
let content: string;
try {
content = await fs.readFile(absPath, 'utf-8');
} catch {
continue;
}
const sourceFile = toRepoPosix(repoRoot, absPath);
for (const name of extractNitroExportNames(content)) {
if (!byLocalName.has(name)) {
byLocalName.set(name, { localName: name, exportName: name, sourceFile, scope: 'server' });
}
}
}
}
// ---- helpers ----------------------------------------------------------------
function isProjectLocalPath(source: string): boolean {
if (!source.startsWith('./') && !source.startsWith('../')) return false;
if (source.includes('node_modules')) return false;
return true;
}
/** True when `absPath` is `repoRoot` itself or lives beneath it. */
function isWithinRepo(repoRoot: string, absPath: string): boolean {
const root = path.resolve(repoRoot);
return absPath === root || absPath.startsWith(root + path.sep);
}
async function resolveExtension(base: string): Promise<string | null> {
const directFile = await firstExistingFile([...FILE_EXTENSIONS.map((ext) => base + ext), base]);
if (directFile !== null) return directFile;
return firstExistingFile(FILE_EXTENSIONS.map((ext) => path.join(base, `index${ext}`)));
}
async function firstExistingFile(candidates: readonly string[]): Promise<string | null> {
const matches = await Promise.all(
candidates.map(async (candidate) => ((await isFile(candidate)) ? candidate : null)),
);
return matches.find((candidate): candidate is string => candidate !== null) ?? null;
}
async function isFile(filePath: string): Promise<boolean> {
try {
const stat = await fs.stat(filePath);
return stat.isFile();
} catch {
return false;
}
}
function toRepoPosix(repoRoot: string, absPath: string): string {
return path.relative(repoRoot, absPath).replace(/\\/g, '/');
}
async function dirExists(dirPath: string): Promise<boolean> {
try {
const stat = await fs.stat(dirPath);
return stat.isDirectory();
} catch {
return false;
}
}
/** Recursively collect supported TypeScript/JavaScript files under a directory. */
async function collectTsFiles(dir: string): Promise<string[]> {
const results: string[] = [];
let entries: import('fs').Dirent[];
try {
entries = await fs.readdir(dir, { withFileTypes: true });
} catch {
return results;
}
for (const entry of entries) {
if (entry.isDirectory() && GENERATED_DIR_NAMES.has(entry.name)) continue;
const full = path.join(dir, entry.name);
if (entry.isDirectory()) {
results.push(...(await collectTsFiles(full)));
} else if (entry.isFile() && hasSupportedSourceExtension(entry.name)) {
results.push(full);
}
}
return results;
}
function hasSupportedSourceExtension(fileName: string): boolean {
return FILE_EXTENSIONS.some((ext) => fileName.endsWith(ext));
}
function extractNitroExportNames(content: string): string[] {
const names = new Set<string>();
NITRO_DECLARATION_EXPORT_RE.lastIndex = 0;
let declaration: RegExpExecArray | null;
while ((declaration = NITRO_DECLARATION_EXPORT_RE.exec(content)) !== null) {
const name = declaration[1] ?? declaration[2];
if (name) names.add(name);
}
NITRO_VARIABLE_EXPORT_RE.lastIndex = 0;
let variableDeclaration: RegExpExecArray | null;
while ((variableDeclaration = NITRO_VARIABLE_EXPORT_RE.exec(content)) !== null) {
// Only the LHS binding name of each top-level declarator is a Nitro export.
// Splitting on top-level commas and reading the leading identifier avoids
// capturing RHS tokens (arrow params, object keys, operands) as export names,
// and tolerates commas inside generic type annotations (`x: Map<a, b> = …`).
for (const declarator of splitTopLevelDeclarators(variableDeclaration[1]!)) {
const name = DECLARATOR_NAME_RE.exec(declarator);
if (name) names.add(name[1]!);
}
}
return [...names];
}
/**
* Split a `const`/`let`/`var` declarator list on top-level commas, tracking
* `()`, `[]`, `{}`, and `<>` nesting so commas inside call args, object/array
* literals, and generic type arguments do not split a declarator. Errs toward
* under-splitting on pathological RHS (a missed binding name, never a spurious
* one) — Nitro only auto-imports real top-level binding names.
*/
function splitTopLevelDeclarators(text: string): string[] {
const parts: string[] = [];
let depth = 0;
let start = 0;
for (let i = 0; i < text.length; i++) {
const ch = text[i];
if (ch === '(' || ch === '[' || ch === '{' || ch === '<') {
depth++;
} else if (ch === ')' || ch === ']' || ch === '}' || ch === '>') {
if (depth > 0) depth--;
} else if (ch === ',' && depth === 0) {
parts.push(text.slice(start, i));
start = i + 1;
}
}
parts.push(text.slice(start));
return parts;
}
@@ -12,11 +12,14 @@
* ./query.ts (TYPESCRIPT_SCOPE_QUERY constant).
*/
import type { ParsedFile } from 'gitnexus-shared';
import type { NodeLabel, ParsedFile, ScopeId } from 'gitnexus-shared';
import { SupportedLanguages } from 'gitnexus-shared';
import { generateId } from '../../../../lib/utils.js';
import { buildMro, defaultLinearize } from '../../scope-resolution/passes/mro.js';
import { populateClassOwnedMembers } from '../../scope-resolution/scope/walkers.js';
import type { ScopeResolver } from '../../scope-resolution/contract/scope-resolver.js';
import { simpleKey } from '../../scope-resolution/graph-bridge/node-lookup.js';
import type { ScopeResolutionIndexes } from '../../model/scope-resolution-indexes.js';
import { typescriptProvider } from '../typescript.js';
import { loadTsconfigPaths, type TsconfigPaths } from '../../language-config.js';
import { buildSuffixIndex, type SuffixIndex } from '../../import-resolvers/utils.js';
@@ -26,12 +29,30 @@ import {
resolveTsTarget,
type TsResolveContext,
} from './index.js';
import {
getNuxtAutoImportEntry,
hasNuxtAutoImports,
loadNuxtAutoImports,
type NuxtAutoImportConfig,
} from './nuxt-auto-imports.js';
/** Shape the orchestrator threads in via `RunScopeResolutionInput.resolutionConfig`. */
interface TypescriptResolutionConfig {
readonly tsconfigPaths: TsconfigPaths | null;
/** Nuxt/Nitro auto-import map. Null for non-Nuxt projects. */
readonly nuxtAutoImports: NuxtAutoImportConfig | null;
}
const TYPESCRIPT_TYPE_ONLY_BINDING_TYPES = new Set<NodeLabel>([
'Interface',
'Type',
'TypeAlias',
'Typedef',
'Trait',
'Annotation',
'Decorator',
]);
/**
* Build a `resolveImportTarget` adapter that memoizes the workspace
* file list, the lower-cased file list, and the per-pass `resolveCache`
@@ -92,11 +113,14 @@ const typescriptScopeResolver: ScopeResolver = {
resolveImportTarget: makeTsResolveImportTarget(),
// Threaded into `resolveImportTarget` so tsconfig path aliases
// (`@/services/user`, `~/x`, …) resolve through the same standard
// (`@/services/user`, `~/x`, ...) resolve through the same standard
// resolver branch the legacy DAG uses. One I/O round-trip per
// workspace pass; the orchestrator awaits this once.
// `nuxtAutoImports` is null for non-Nuxt projects (no .nuxt/imports.d.ts),
// so this adds zero overhead to ordinary TypeScript repos.
loadResolutionConfig: async (repoPath: string) => ({
tsconfigPaths: await loadTsconfigPaths(repoPath),
nuxtAutoImports: await loadNuxtAutoImports(repoPath),
}),
// TypeScript declaration merging + LEGB: local > import > wildcard,
@@ -130,24 +154,198 @@ const typescriptScopeResolver: ScopeResolver = {
propagatesReturnTypesAcrossImports: true,
// TypeScript uses `.values()` / `.keys()` method-call syntax for
// collection views — no property-style accessors like C#'s
// collection views -- no property-style accessors like C#'s
// `Dictionary<K,V>.Values`. Leave `unwrapCollectionAccessor`
// undefined and let the regular member-call branch handle them.
//
// `collapseMemberCallsByCallerTarget` left undefined (= false) —
// `collapseMemberCallsByCallerTarget` left undefined (= false) --
// TypeScript legacy DAG emits one edge per call site, so
// per-site dedup is the parity target.
//
// `populateNamespaceSiblings` left undefined — TypeScript requires
// `populateNamespaceSiblings` left undefined -- TypeScript requires
// an explicit `import` / namespace augmentation for cross-file
// visibility; there's no implicit same-namespace sibling rule
// like C#'s.
//
// `hoistTypeBindingsToModule` — `tsBindingScopeFor` DOES hoist
// `hoistTypeBindingsToModule` -- `tsBindingScopeFor` DOES hoist
// method return-type bindings to the enclosing Module scope
// (mirrors C#), so enable the walk-up that lets the compound-
// receiver resolver find them.
hoistTypeBindingsToModule: true,
/**
* Emit CALLS edges for Nuxt/Nitro auto-imported symbols that are used
* without an explicit import statement.
*
* Nuxt makes composables and server utils available project-wide via its
* auto-import system. Because no `import` statement exists, the standard
* scope-resolution passes cannot create call-graph edges for these symbols.
* This hook recovers those edges after all normal resolution has run.
*
* For each TypeScript file the hook:
* 1. Builds the set of files already explicitly imported (to avoid
* creating duplicate edges for symbols imported conventionally).
* 2. Iterates parsed free-call reference sites and checks each against the
* auto-import map selected for the caller's Nuxt/Nitro scope.
* 3. For each hit that is not shadowed by a local binding or explicit
* import, emits a CALLS edge from the file's File node to the target
* function node, and an IMPORTS edge from the caller file to the source
* file (once per pair).
*
* Confidence is 0.75 (below the 0.9 used for fully resolved edges) to
* signal that these edges are heuristic rather than type-checked.
*/
emitPostResolutionEdges(graph, parsedFiles, nodeLookup, indexes, ctx) {
const cfg = ctx.resolutionConfig as TypescriptResolutionConfig | undefined;
const autoImports = cfg?.nuxtAutoImports;
if (!autoImports || !hasNuxtAutoImports(autoImports)) return;
// Pre-build a file -> explicit imported local names index so importing one
// symbol from a source does not suppress other auto-imported symbols from it.
const explicitImportNamesByFile = new Map<string, Set<string>>();
for (const [scopeId, edges] of indexes.imports) {
const scope = indexes.scopeTree.getScope(scopeId);
if (!scope?.filePath) continue;
let names = explicitImportNamesByFile.get(scope.filePath);
if (!names) {
names = new Set<string>();
explicitImportNamesByFile.set(scope.filePath, names);
}
for (const edge of edges) {
// Record the local name whether or not the import resolved to a file.
// An explicit import of a name — even from an unresolved external
// package (`import { useAuto } from '@vueuse/core'`) — is authoritative
// shadowing intent and must suppress the auto-import for that name.
if (edge.localName) names.add(edge.localName);
}
}
for (const parsedFile of parsedFiles) {
const { filePath } = parsedFile;
if (!filePath.endsWith('.ts') && !filePath.endsWith('.tsx')) continue;
const fileId = generateId('File', filePath);
const explicitImports = explicitImportNamesByFile.get(filePath) ?? new Set<string>();
// Track (sourceFile) pairs already handled for this caller to avoid
// emitting duplicate IMPORTS edges and duplicate CALLS edges per symbol.
const emittedImports = new Set<string>();
const emittedCalls = new Set<string>();
for (const site of parsedFile.referenceSites) {
if (site.kind !== 'call' || site.callForm !== 'free') continue;
const localName = site.name;
const entry = getNuxtAutoImportEntry(autoImports, localName, filePath);
if (!entry) continue;
const { exportName, sourceFile } = entry;
// Skip when the file already binds this name explicitly, when the file
// IS the source, or when a lexical same-file binding shadows it.
if (
explicitImports.has(localName) ||
sourceFile === filePath ||
hasLocalBindingInScopeChain(site.inScope, localName, filePath, indexes)
) {
continue;
}
// Emit one IMPORTS edge per (caller, sourceFile) pair.
if (!emittedImports.has(sourceFile)) {
emittedImports.add(sourceFile);
const targetFileId = generateId('File', sourceFile);
if (graph.getNode(targetFileId)) {
graph.addRelationship({
id: generateId('IMPORTS', `${fileId}->nuxt-auto-import->${targetFileId}`),
sourceId: fileId,
targetId: targetFileId,
type: 'IMPORTS',
confidence: 0.75,
reason: 'nuxt-auto-import-file',
});
}
}
// Emit one CALLS edge per (caller, symbol) pair.
const callKey = `${sourceFile}::${localName}`;
if (emittedCalls.has(callKey)) continue;
emittedCalls.add(callKey);
// Look up the graph node by export name first, fall back to local name.
// The fallback handles `default as X` where the function is named X.
const targetNodeId =
nodeLookup.get(simpleKey(sourceFile, exportName)) ??
(exportName !== localName ? nodeLookup.get(simpleKey(sourceFile, localName)) : undefined);
if (!targetNodeId || !graph.getNode(targetNodeId)) continue;
graph.addRelationship({
id: generateId('CALLS', `${fileId}:nuxt-auto-import:${localName}->${targetNodeId}`),
sourceId: fileId,
targetId: targetNodeId,
type: 'CALLS',
confidence: 0.75,
reason: 'nuxt-auto-import',
});
}
}
},
};
function occupiesTypeScriptValueSpace(type: NodeLabel): boolean {
return !TYPESCRIPT_TYPE_ONLY_BINDING_TYPES.has(type);
}
function hasLocalBindingInScopeChain(
scopeId: ScopeId,
name: string,
filePath: string,
indexes: ScopeResolutionIndexes,
): boolean {
const visited = new Set<ScopeId>();
let cursor: ScopeId | null | undefined = scopeId;
while (cursor !== null && cursor !== undefined && !visited.has(cursor)) {
visited.add(cursor);
const scope = indexes.scopeTree.getScope(cursor);
if (!scope) return false;
const localBindings = scope.bindings.get(name);
if (
localBindings?.some(
(binding) =>
binding.origin === 'local' &&
binding.def.filePath === filePath &&
occupiesTypeScriptValueSpace(binding.def.type),
)
) {
return true;
}
// Type-annotated function parameters — and other value-space type facts
// such as `self` and variable annotations — live in `scope.typeBindings`,
// not `scope.bindings`. A parameter named like a composable genuinely
// shadows the auto-import, and typeBindings never holds a pure type that
// belongs to callable space, so a same-file presence check here cannot
// over-suppress a legitimate auto-import.
//
// Residual (known limitation): this catches parameters whose annotation the
// TS scope query records as a type-binding (`p: Named`, generics, unions,
// predefined, arrays). Function-typed params (`p: () => void`), untyped
// params, destructured locals (`const { x } = …`), and catch-clause vars
// are captured by NEITHER map — the scope query emits no `@declaration` /
// `@type-binding` for them — so those shadow forms still leak an edge.
// Closing that needs shared TS scope-query/extractor changes that alter call
// resolution beyond Nuxt, so it is deferred to a follow-up rather than fixed
// here.
if (scope.filePath === filePath && scope.typeBindings.has(name)) {
return true;
}
cursor = scope.parent;
}
return false;
}
export { typescriptScopeResolver };
@@ -362,13 +362,24 @@ export const kotlinMethodConfig: MethodExtractionConfig = {
},
extractReceiverType(node) {
// Extension function: user_type appears before the simple_identifier (name)
// e.g., fun String.format(template: String) → receiver is "String"
// Extension function receiver. Newer tree-sitter-kotlin exposes it as a
// `receiver` field wrapping a `receiver_type` (which wraps the user_type);
// older grammars emitted a bare user_type/nullable_type child before the
// name (e.g. fun String.format(...) → receiver is "String").
const receiverField = node.childForFieldName('receiver');
if (receiverField) {
const inner = receiverField.namedChild(0) ?? receiverField;
return extractSimpleTypeName(inner) ?? inner.text?.trim();
}
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (!child) continue;
if (child.type === 'simple_identifier') break; // past the name — no receiver
if (child.type === 'user_type' || child.type === 'nullable_type') {
if (
child.type === 'receiver_type' ||
child.type === 'user_type' ||
child.type === 'nullable_type'
) {
return extractSimpleTypeName(child) ?? child.text?.trim();
}
}
@@ -22,10 +22,12 @@ import type {
FetchWrapperDef,
} from './workers/parse-worker.js';
import type {
ExtractedRouterConstructorPrefix,
ExtractedRouterImport,
ExtractedRouterInclude,
ExtractedRouterModuleAlias,
} from './route-extractors/fastapi-router-bindings.js';
import type { SharedSpringType } from './route-extractors/spring-shared.js';
export type FileProgressCallback = (current: number, total: number, filePath: string) => void;
@@ -36,9 +38,12 @@ export interface WorkerExtractedData {
decoratorRoutes: ExtractedDecoratorRoute[];
routerIncludes: ExtractedRouterInclude[];
routerImports: ExtractedRouterImport[];
routerConstructorPrefixes: ExtractedRouterConstructorPrefix[];
routerModuleAliases: ExtractedRouterModuleAlias[];
toolDefs: ExtractedToolDef[];
ormQueries: ExtractedORMQuery[];
/** Project-wide Spring class/interface views for the #2288 inheritance pass. */
springTypes: SharedSpringType[];
fileScopeBindings: FileScopeBindings[];
/**
* Per-file `ParsedFile` artifacts from the new scope-based resolution
@@ -77,7 +82,9 @@ export const mergeChunkResults = (
const allDecoratorRoutes: ExtractedDecoratorRoute[] = [];
const allRouterIncludes: ExtractedRouterInclude[] = [];
const allRouterImports: ExtractedRouterImport[] = [];
const allRouterConstructorPrefixes: ExtractedRouterConstructorPrefix[] = [];
const allRouterModuleAliases: ExtractedRouterModuleAlias[] = [];
const allSpringTypes: SharedSpringType[] = [];
const allToolDefs: ExtractedToolDef[] = [];
const allORMQueries: ExtractedORMQuery[] = [];
const fileScopeBindingsByFile: FileScopeBindings[] = [];
@@ -119,7 +126,11 @@ export const mergeChunkResults = (
for (const item of result.decoratorRoutes) allDecoratorRoutes.push(item);
for (const item of result.routerIncludes ?? []) allRouterIncludes.push(item);
for (const item of result.routerImports ?? []) allRouterImports.push(item);
for (const item of result.routerConstructorPrefixes ?? []) {
allRouterConstructorPrefixes.push(item);
}
for (const item of result.routerModuleAliases ?? []) allRouterModuleAliases.push(item);
for (const item of result.springTypes ?? []) allSpringTypes.push(item);
for (const item of result.toolDefs) allToolDefs.push(item);
if (result.ormQueries) for (const item of result.ormQueries) allORMQueries.push(item);
if (result.fileScopeBindings)
@@ -134,9 +145,11 @@ export const mergeChunkResults = (
decoratorRoutes: allDecoratorRoutes,
routerIncludes: allRouterIncludes,
routerImports: allRouterImports,
routerConstructorPrefixes: allRouterConstructorPrefixes,
routerModuleAliases: allRouterModuleAliases,
toolDefs: allToolDefs,
ormQueries: allORMQueries,
springTypes: allSpringTypes,
fileScopeBindings: fileScopeBindingsByFile,
parsedFiles: allParsedFiles,
};
@@ -40,6 +40,7 @@ import { DEFAULT_PDG_MAX_FUNCTION_LINES } from '../cfg/collect.js';
import type { WorkerExtractedData } from '../parsing-processor.js';
import {
processRoutesFromExtracted,
resolveRouteHandlerSymbols,
buildExportedTypeMapFromGraph,
type ExportedTypeMap,
} from '../call-processor.js';
@@ -75,10 +76,16 @@ import type {
FetchWrapperDef,
} from '../workers/parse-worker.js';
import type {
ExtractedRouterConstructorPrefix,
ExtractedRouterImport,
ExtractedRouterInclude,
ExtractedRouterModuleAlias,
} from '../route-extractors/fastapi-router-bindings.js';
import { normalizeExtractedRoutePath } from '../route-extractors/route-path.js';
import {
resolveInheritedSpringRoutes,
type SharedSpringType,
} from '../route-extractors/spring-shared.js';
import type { KnowledgeGraph } from '../../graph/types.js';
import type { PipelineOptions } from '../pipeline.js';
import fs from 'node:fs';
@@ -370,6 +377,10 @@ export async function runChunkedParseAndResolve(
allToolDefs: ExtractedToolDef[];
allORMQueries: ExtractedORMQuery[];
bindingAccumulator: BindingAccumulator;
/** Route URL → resolved handler symbol UID (Part 2, #2138). Lets the routes
* phase stamp `handlerSymbolId` on Route nodes so contract extraction can
* read the handler from the graph instead of re-parsing source. */
routeHandlerSymbols: ReadonlyMap<string, string>;
/** SemanticModel populated during parse — scope-resolution reads its
* TypeRegistry / MethodRegistry / SymbolTable indexes. */
model: MutableSemanticModel;
@@ -614,7 +625,9 @@ export async function runChunkedParseAndResolve(
const allDecoratorRoutes: ExtractedDecoratorRoute[] = [];
const allRouterIncludes: ExtractedRouterInclude[] = [];
const allRouterImports: ExtractedRouterImport[] = [];
const allRouterConstructorPrefixes: ExtractedRouterConstructorPrefix[] = [];
const allRouterModuleAliases: ExtractedRouterModuleAlias[] = [];
const allSpringTypes: SharedSpringType[] = [];
const allToolDefs: ExtractedToolDef[] = [];
const allORMQueries: ExtractedORMQuery[] = [];
// Aggregated per-file ParsedFile artifacts produced by workers' calls
@@ -769,9 +782,17 @@ export async function runChunkedParseAndResolve(
if (chunkWorkerData.routerImports?.length) {
for (const item of chunkWorkerData.routerImports) allRouterImports.push(item);
}
if (chunkWorkerData.routerConstructorPrefixes?.length) {
for (const item of chunkWorkerData.routerConstructorPrefixes) {
allRouterConstructorPrefixes.push(item);
}
}
if (chunkWorkerData.routerModuleAliases?.length) {
for (const item of chunkWorkerData.routerModuleAliases) allRouterModuleAliases.push(item);
}
if (chunkWorkerData.springTypes?.length) {
for (const item of chunkWorkerData.springTypes) allSpringTypes.push(item);
}
if (chunkWorkerData.toolDefs?.length) {
for (const item of chunkWorkerData.toolDefs) allToolDefs.push(item);
}
@@ -1143,7 +1164,10 @@ export async function runChunkedParseAndResolve(
// decorator inherits its file-basename's prefix. When a router is mounted
// under multiple prefixes we duplicate the route entry, mirroring FastAPI's
// runtime behaviour.
if (allRouterIncludes.length > 0 && allDecoratorRoutes.length > 0) {
if (
(allRouterIncludes.length > 0 || allRouterConstructorPrefixes.length > 0) &&
allDecoratorRoutes.length > 0
) {
// Group `routerImports` by file so we can resolve Shape-B locals against
// imports declared in the SAME file as the include_router call. We carry
// both the short module key (file basename) and, when available, the long
@@ -1188,6 +1212,11 @@ export async function runChunkedParseAndResolve(
// without a corresponding import statement).
const prefixesByLongKey = new Map<string, Set<string>>();
const prefixesByShortKey = new Map<string, Set<string>>();
// Constructor prefixes are `router`-only (the apply gate below and the
// group-layer tree-sitter both pin to the literal name `router`), so a
// flat file-key → prefix map suffices — mirrors the group layer's shape.
const constructorPrefixesByLongKey = new Map<string, string>();
const constructorPrefixesByShortKey = new Map<string, string>();
const recordPrefix = (target: Map<string, Set<string>>, key: string, prefix: string): void => {
let set = target.get(key);
@@ -1228,7 +1257,11 @@ export async function runChunkedParseAndResolve(
}
}
if (prefixesByLongKey.size > 0 || prefixesByShortKey.size > 0) {
if (
prefixesByLongKey.size > 0 ||
prefixesByShortKey.size > 0 ||
allRouterConstructorPrefixes.length > 0
) {
const fileLongKey = (rel: string): string => {
// Strip `.py`, then take the last two path segments. `api/users.py`
// → `api/users`. Files at the repo root return the empty string,
@@ -1250,6 +1283,15 @@ export async function runChunkedParseAndResolve(
return file.endsWith('.py') ? file.slice(0, -3) : file;
};
for (const ctor of allRouterConstructorPrefixes) {
const longKey = fileLongKey(ctor.filePath);
if (longKey) {
constructorPrefixesByLongKey.set(longKey, ctor.prefix);
} else {
constructorPrefixesByShortKey.set(fileShortKey(ctor.filePath), ctor.prefix);
}
}
const expanded: ExtractedDecoratorRoute[] = [];
for (const dr of allDecoratorRoutes) {
if (dr.decoratorReceiver !== 'router' || !dr.filePath.endsWith('.py')) {
@@ -1265,12 +1307,24 @@ export async function runChunkedParseAndResolve(
? undefined
: prefixesByShortKey.get(fileShortKey(dr.filePath));
const prefixes = longPrefixes ?? shortPrefixes;
// Constructor prefixes are keyed like include_router prefixes:
// long-key entries are precise, while short-key entries are only
// valid for repo-root/single-segment files where `fileLongKey`
// returns ''. Do not fall back from a missing long-key match to the
// short key or a root `users.py` prefix can leak onto
// `admin/users.py`.
const constructorPrefix = longKey
? constructorPrefixesByLongKey.get(longKey)
: constructorPrefixesByShortKey.get(fileShortKey(dr.filePath));
const routePath = constructorPrefix
? normalizeExtractedRoutePath(dr.routePath, constructorPrefix)
: dr.routePath;
if (!prefixes || prefixes.size === 0) {
expanded.push(dr);
expanded.push(routePath === dr.routePath ? dr : { ...dr, routePath });
continue;
}
for (const prefix of prefixes) {
expanded.push({ ...dr, prefix });
expanded.push({ ...dr, routePath, prefix });
}
}
allDecoratorRoutes.length = 0;
@@ -1278,10 +1332,37 @@ export async function runChunkedParseAndResolve(
}
}
// Cross-file Spring interface-inheritance pass (#2288): a concrete
// `@RestController` inherits the `@*Mapping`s declared on the interfaces it
// implements. The per-file `SharedSpringType` views collected by the Java
// provider's `extractRouteInheritanceTypes` hook are resolved here, project-
// wide, into decorator routes attributed to the implementing controller (the
// interface's own per-file routes were suppressed at extraction). Mirrors the
// group layer via the shared `resolveInheritedSpringRoutes` so both agree.
if (allSpringTypes.length > 0) {
for (const inherited of resolveInheritedSpringRoutes(allSpringTypes)) {
allDecoratorRoutes.push({
filePath: inherited.filePath,
routePath: inherited.path,
httpMethod: inherited.method,
decoratorName: 'inherited-mapping',
lineNumber: 0,
handlerName: inherited.methodName,
});
}
}
logHeapProbe(
'parse-impl-return',
`exportedTypeMap=${exportedTypeMap.size} parsedFiles=${allParsedFiles.length} nodes=${graph.nodeCount}`,
);
// Part 2 (#2138): resolve each route's handler to a real symbol UID now that
// the model is fully populated and decorator-route prefixes are finalized.
const routeHandlerSymbols = resolveRouteHandlerSymbols(
model,
allExtractedRoutes,
allDecoratorRoutes,
);
return {
exportedTypeMap,
allFetchCalls,
@@ -1291,6 +1372,7 @@ export async function runChunkedParseAndResolve(
allToolDefs,
allORMQueries,
bindingAccumulator,
routeHandlerSymbols,
model,
// Whether a worker pool was actually constructed for this run. False means
// no pool was needed: a warm all-cache-hit run replays cached worker output
@@ -51,6 +51,9 @@ export interface ParseOutput {
readonly allDecoratorRoutes: readonly ExtractedDecoratorRoute[];
readonly allToolDefs: readonly ExtractedToolDef[];
readonly allORMQueries: readonly ExtractedORMQuery[];
/** Route URL → resolved handler symbol UID (Part 2, #2138). Consumed by the
* routes phase to stamp `handlerSymbolId` on Route nodes. */
readonly routeHandlerSymbols: ReadonlyMap<string, string>;
bindingAccumulator: BindingAccumulator;
/** SemanticModel populated during parse — scope-resolution reads its
* TypeRegistry / MethodRegistry / SymbolTable indexes. */
@@ -17,6 +17,7 @@ import type { ToolsOutput } from './tools.js';
import type { StructureOutput } from './structure.js';
import { processProcesses, type ProcessDetectionResult } from '../process-processor.js';
import { generateId } from '../../../lib/utils.js';
import { routeNodeKey } from '../route-extractors/route-path.js';
import { isDev } from '../utils/env.js';
import { logger } from '../../logger.js';
@@ -106,14 +107,39 @@ export const processesPhase: PipelinePhase<ProcessesOutput> = {
// Link Route and Tool nodes to Processes
if (routeRegistry.size > 0 || toolDefs.length > 0) {
const routesByFile = new Map<string, string[]>();
for (const [url, entry] of routeRegistry) {
let list = routesByFile.get(entry.filePath);
// Two-tier route lookup, mirroring the tool tables 10 lines below.
// Routes whose handler resolved key by `handlerSymbolId` (read from
// the Route node's graph properties — routes.ts stamps it there) and
// link ONLY to the process whose entryPoint matches; routes without a
// resolved handler fall back to a per-file bucket so we still attach
// the Route node to a same-file process (best-effort).
//
// Pre-#2289-review-P2 this was a single per-file bucket: every verb
// on a file's controller was linked to every process in that file,
// cross-wiring same-file `GET /items` and `POST /items` to each
// other's handler processes. The per-verb `handlerSymbolId` the
// routes phase stamps on the Route node was never consulted.
const routesByHandlerId = new Map<string, string[]>();
const routesWithoutHandlerByFile = new Map<string, string[]>();
for (const [, entry] of routeRegistry) {
// Push the Route node identity (`routeNodeKey`), not the bare URL, so the
// ENTRY_POINT_OF edge targets the same node id the routes phase created
// (#2289: a same-URL GET/POST pair is two distinct Route nodes).
const routeKey = routeNodeKey(entry.method, entry.url);
// Source of truth for handlerSymbolId is the Route node in the
// graph (routes.ts populates it from `routeHandlerSymbols`); the
// routes phase runs before processes (see `deps`), so the node is
// always present here.
const routeNode = ctx.graph.getNode(generateId('Route', routeKey));
const handlerSymbolId = routeNode?.properties.handlerSymbolId as string | undefined;
const targetMap = handlerSymbolId ? routesByHandlerId : routesWithoutHandlerByFile;
const bucketKey = handlerSymbolId ?? entry.filePath;
let list = targetMap.get(bucketKey);
if (!list) {
list = [];
routesByFile.set(entry.filePath, list);
targetMap.set(bucketKey, list);
}
list.push(url);
list.push(routeKey);
}
const toolsByHandlerId = new Map<string, string[]>();
const toolsWithoutHandlerByFile = new Map<string, string[]>();
@@ -136,10 +162,12 @@ export const processesPhase: PipelinePhase<ProcessesOutput> = {
const entryFile = entryNode.properties.filePath;
if (!entryFile) continue;
const routeURLs = routesByFile.get(entryFile);
if (routeURLs) {
for (const routeURL of routeURLs) {
const routeNodeId = generateId('Route', routeURL);
const exactRouteKeys = routesByHandlerId.get(proc.entryPointId);
const fallbackRouteKeys = routesWithoutHandlerByFile.get(entryFile);
const routeKeys = exactRouteKeys ?? fallbackRouteKeys;
if (routeKeys) {
for (const routeKey of routeKeys) {
const routeNodeId = generateId('Route', routeKey);
ctx.graph.addRelationship({
id: generateId('ENTRY_POINT_OF', `${routeNodeId}->${proc.id}`),
sourceId: routeNodeId,
@@ -29,6 +29,11 @@ import {
compiledMatcherMatchesRoute,
} from '../route-extractors/middleware.js';
import { processNextjsFetchRoutes } from '../call-processor.js';
import {
normalizeExtractedRoutePath,
normalizeRouteMethod,
routeNodeKey,
} from '../route-extractors/route-path.js';
import { generateId } from '../../../lib/utils.js';
import { readFileContents } from '../filesystem-walker.js';
import { isDev } from '../utils/env.js';
@@ -42,6 +47,13 @@ const EXPO_NAV_PATTERNS = [
export interface RouteEntry {
filePath: string;
source: string;
/**
* The route's URL path (leading-slash, prefix-joined). This is the Route
* node's `name`. Stored explicitly because the registry is keyed by the
* `(method, url)` identity (`routeNodeKey`), so the key is no longer the URL
* — downstream URL consumers (middleware/fetch matching) read this instead.
*/
url: string;
/**
* HTTP verb for this route when ingestion knows it structurally
* (Spring/Laravel framework routes and decorator routes carry
@@ -133,48 +145,10 @@ export function extractTemplateStaticFetchCalls(
return calls;
}
export function normalizeExtractedRoutePath(routePath: string, prefix: string | null): string {
const pathPart = routePath.trim().replace(/^\/+/, '').replace(/\/+$/g, '');
const prefixPart = prefix?.trim().replace(/^\/+/, '').replace(/\/+$/g, '');
const joined = prefixPart ? `/${prefixPart}${pathPart ? `/${pathPart}` : ''}` : `/${pathPart}`;
return joined.replace(/\/+/g, '/') || '/';
}
function escapeRegex(s: string): string {
return s.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
}
/**
* Canonicalize a route's HTTP verb for persistence on the Route node.
* Returns an upper-cased standard method, or `undefined` when the value
* is not a real HTTP verb. Laravel `Route::resource` / `apiResource`
* surface `httpMethod` values like `resource` / `apiResource` (they
* expand to several verbs at runtime), so they must not be stored as a
* method — leaving them `undefined` keeps the column clean and lets the
* contract extractor fall back to its source-scan path for those routes.
*/
const VALID_HTTP_METHODS = new Set([
'GET',
'POST',
'PUT',
'PATCH',
'DELETE',
'HEAD',
'OPTIONS',
'TRACE',
'CONNECT',
]);
export function normalizeRouteMethod(raw: string | null | undefined): string | undefined {
if (typeof raw !== 'string') return undefined;
const verb = raw.trim().toUpperCase();
// '*' marks a method-agnostic route (e.g. a Django function view handles any
// verb). Preserve it so the contract layer emits a wildcard provider that
// matches consumers of any method, instead of silently narrowing to GET.
if (verb === '*') return '*';
return VALID_HTTP_METHODS.has(verb) ? verb : undefined;
}
export const routesPhase: PipelinePhase<RoutesOutput> = {
name: 'routes',
deps: ['parse'],
@@ -189,6 +163,7 @@ export const routesPhase: PipelinePhase<RoutesOutput> = {
allFetchWrapperDefs,
allExtractedRoutes,
allDecoratorRoutes,
routeHandlerSymbols,
} = getPhaseOutput<ParseOutput>(deps, 'parse');
// Local copy — routes phase must not mutate upstream ParseOutput
@@ -221,31 +196,45 @@ export const routesPhase: PipelinePhase<RoutesOutput> = {
if (expoAppPaths.has(p)) {
const expoURL = expoFileToRouteURL(p);
if (expoURL && !routeRegistry.has(expoURL)) {
routeRegistry.set(expoURL, { filePath: p, source: 'expo-filesystem-route' });
routeRegistry.set(expoURL, {
filePath: p,
source: 'expo-filesystem-route',
url: expoURL,
});
continue;
}
}
const nextjsURL = nextjsFileToRouteURL(p);
if (nextjsURL && !routeRegistry.has(nextjsURL)) {
routeRegistry.set(nextjsURL, { filePath: p, source: 'nextjs-filesystem-route' });
routeRegistry.set(nextjsURL, {
filePath: p,
source: 'nextjs-filesystem-route',
url: nextjsURL,
});
continue;
}
if (p.endsWith('.php')) {
const phpURL = phpFileToRouteURL(p);
if (phpURL && !routeRegistry.has(phpURL)) {
routeRegistry.set(phpURL, { filePath: p, source: 'php-file-route' });
routeRegistry.set(phpURL, { filePath: p, source: 'php-file-route', url: phpURL });
}
}
}
let duplicateRoutes = 0;
const namedRouteRegistry = new Map<string, string>();
const addRoute = (url: string, entry: RouteEntry) => {
if (routeRegistry.has(url)) {
// Routes are keyed by their `(method, url)` identity (#2289): a same-URL
// multi-verb pair (`GET /x` + `POST /x`) is two entries, not one. Method-less
// / wildcard routes key by URL (see `routeNodeKey`), so filesystem/resource
// routes stay byte-identical. A true duplicate (same method AND url) is still
// dropped.
const addRoute = (url: string, entry: Omit<RouteEntry, 'url'>) => {
const key = routeNodeKey(entry.method, url);
if (routeRegistry.has(key)) {
duplicateRoutes++;
return;
}
routeRegistry.set(url, entry);
routeRegistry.set(key, { ...entry, url });
};
for (const route of allExtractedRoutes) {
if (!route.routePath) continue;
@@ -273,8 +262,8 @@ export const routesPhase: PipelinePhase<RoutesOutput> = {
const handlerPaths = [...routeRegistry.values()].map((e) => e.filePath);
handlerContents = await readFileContents(ctx.repoPath, handlerPaths);
for (const [routeURL, entry] of routeRegistry) {
const { filePath: handlerPath, source: routeSource, method: routeMethod } = entry;
for (const [routeKey, entry] of routeRegistry) {
const { filePath: handlerPath, source: routeSource, method: routeMethod, url } = entry;
const content = handlerContents.get(handlerPath);
const { responseKeys, errorKeys } = content
@@ -286,14 +275,16 @@ export const routesPhase: PipelinePhase<RoutesOutput> = {
const mwResult = content ? extractMiddlewareChain(content) : undefined;
const middleware = mwResult?.chain;
const routeNodeId = generateId('Route', routeURL);
const routeNodeId = generateId('Route', routeKey);
const handlerSymbolId = routeHandlerSymbols.get(routeKey);
ctx.graph.addNode({
id: routeNodeId,
label: 'Route',
properties: {
name: routeURL,
name: url,
filePath: handlerPath,
...(routeMethod ? { method: routeMethod } : {}),
...(handlerSymbolId ? { handlerSymbolId } : {}),
...(responseKeys ? { responseKeys } : {}),
...(errorKeys ? { errorKeys } : {}),
...(middleware && middleware.length > 0 ? { middleware } : {}),
@@ -344,13 +335,13 @@ export const routesPhase: PipelinePhase<RoutesOutput> = {
.filter((m): m is NonNullable<typeof m> => m !== null);
let linkedCount = 0;
for (const [routeURL] of routeRegistry) {
for (const [routeKey, entry] of routeRegistry) {
const matches =
compiled.length === 0 ||
compiled.some((cm) => compiledMatcherMatchesRoute(cm, routeURL));
compiled.some((cm) => compiledMatcherMatchesRoute(cm, entry.url));
if (!matches) continue;
const routeNodeId = generateId('Route', routeURL);
const routeNodeId = generateId('Route', routeKey);
const existing = ctx.graph.getNode(routeNodeId);
if (!existing) continue;
@@ -487,13 +478,19 @@ export const routesPhase: PipelinePhase<RoutesOutput> = {
}
if (routeRegistry.size > 0 && allFetchCalls.length > 0) {
const routeURLToFile = new Map<string, string>();
for (const [url, entry] of routeRegistry) routeURLToFile.set(url, entry.filePath);
// url → [route node keys at that url] (one per verb). A verb-less fetch()
// consumer matches by URL and connects to every Route node at that URL.
const routeUrlToKeys = new Map<string, string[]>();
for (const [routeKey, entry] of routeRegistry) {
const existing = routeUrlToKeys.get(entry.url);
if (existing) existing.push(routeKey);
else routeUrlToKeys.set(entry.url, [routeKey]);
}
const consumerPaths = [...new Set(allFetchCalls.map((c) => c.filePath))];
const consumerContents = await readFileContents(ctx.repoPath, consumerPaths);
processNextjsFetchRoutes(ctx.graph, allFetchCalls, routeURLToFile, consumerContents);
processNextjsFetchRoutes(ctx.graph, allFetchCalls, routeUrlToKeys, consumerContents);
if (isDev) {
logger.info(
`🔗 Processed ${allFetchCalls.length} fetch() calls against ${routeRegistry.size} routes`,
@@ -105,6 +105,11 @@ export interface ExtractedRouterModuleAlias {
moduleKeyLong: string;
}
export interface ExtractedRouterConstructorPrefix {
filePath: string;
prefix: string;
}
// `<host>.include_router(<module>.router, ..., prefix='/x')` (Shape A).
// `<host>` is left unrestricted — common production names include
// `app`, `api`, `application`, `asgi_app`. Pinning to the literal
@@ -122,6 +127,8 @@ const INCLUDE_ROUTER_NAME_RE =
// The latter is the common case and the only one we can map back to
// a module stem.
const FROM_IMPORT_ROUTER_RE = /^\s*from\s+(\.+|\.*[A-Za-z_][\w.]*)\s+import\s+([^#\n]+)/gm;
const API_ROUTER_ASSIGN_RE = /\brouter\s*=\s*APIRouter\s*\(/g;
const API_ROUTER_PREFIX_ARG_RE = /\bprefix\s*=\s*(['"])([^'"]*)\1/;
/**
* Last `.`-separated segment of a (possibly relative) Python module
@@ -157,6 +164,35 @@ export function lastTwoSegmentsAsPath(text: string): string {
return `${parent}/${stem}`;
}
function findMatchingParen(content: string, openIndex: number): number {
let depth = 0;
let quote: string | null = null;
for (let i = openIndex; i < content.length; i++) {
const ch = content[i];
if (quote) {
if (ch === '\\') {
i++;
continue;
}
if (ch === quote) quote = null;
continue;
}
if (ch === '"' || ch === "'") {
quote = ch;
continue;
}
if (ch === '(' || ch === '[' || ch === '{') {
depth++;
continue;
}
if (ch === ')' || ch === ']' || ch === '}') {
depth--;
if (depth === 0 && ch === ')') return i;
}
}
return -1;
}
/**
* Scan a single Python file's source text for FastAPI router
* `include_router` sites and `from <module> import router` imports,
@@ -176,6 +212,7 @@ export function extractFastAPIRouterBindings(
outIncludes: ExtractedRouterInclude[],
outImports: ExtractedRouterImport[],
outModuleAliases?: ExtractedRouterModuleAlias[],
outConstructorPrefixes?: ExtractedRouterConstructorPrefix[],
): void {
if (!content.includes('include_router') && !content.includes('router')) return;
@@ -239,6 +276,24 @@ export function extractFastAPIRouterBindings(
}
}
if (outConstructorPrefixes && content.includes('APIRouter') && content.includes('prefix')) {
API_ROUTER_ASSIGN_RE.lastIndex = 0;
// Only `router = APIRouter(...)` is captured (the apply gate and the
// group-layer tree-sitter both pin to the literal name `router`).
while (API_ROUTER_ASSIGN_RE.exec(content) !== null) {
const openParen = API_ROUTER_ASSIGN_RE.lastIndex - 1;
const closeParen = findMatchingParen(content, openParen);
if (closeParen < 0) continue;
const args = content.slice(openParen + 1, closeParen);
const prefixMatch = API_ROUTER_PREFIX_ARG_RE.exec(args);
if (!prefixMatch) continue;
outConstructorPrefixes.push({
filePath,
prefix: prefixMatch[2],
});
}
}
if (!content.includes('include_router')) return;
// Shape A: `<host>.include_router(<module>.router, prefix='/x')`.
@@ -0,0 +1,68 @@
/**
* Shared route-path normalization.
*
* Extracted from the routes phase so both the routes phase (which creates the
* `Route` graph node, keyed by `(method, url)` via `routeNodeKey` — #2289) and
* the parse phase (which resolves each route's handler symbol and needs the
* SAME key to associate the resolved id back to the route) can compute an
* identical route identity without a phase-to-phase import cycle. Pure string
* logic, no dependencies.
*/
/**
* Join a route's path with its (optional) prefix into a normalized,
* leading-slash URL used as the Route node identity. Collapses duplicate
* slashes and strips trailing ones; an empty result degrades to `/`.
*/
export function normalizeExtractedRoutePath(routePath: string, prefix: string | null): string {
const pathPart = routePath.trim().replace(/^\/+/, '').replace(/\/+$/g, '');
const prefixPart = prefix?.trim().replace(/^\/+/, '').replace(/\/+$/g, '');
const joined = prefixPart ? `/${prefixPart}${pathPart ? `/${pathPart}` : ''}` : `/${pathPart}`;
return joined.replace(/\/+/g, '/') || '/';
}
const VALID_HTTP_METHODS = new Set([
'GET',
'POST',
'PUT',
'PATCH',
'DELETE',
'HEAD',
'OPTIONS',
'TRACE',
'CONNECT',
]);
/**
* Canonicalize a route's HTTP verb for persistence on the Route node and for
* the route identity key. Returns an upper-cased standard method, `'*'` for a
* method-agnostic route (e.g. a Django function view), or `undefined` when the
* value is not a real HTTP verb. Laravel `Route::resource` / `apiResource`
* surface values like `resource` / `apiResource` (they expand to several verbs
* at runtime), so they come back `undefined` — keeping the column clean and
* letting the contract extractor fall back to its source-scan path.
*/
export function normalizeRouteMethod(raw: string | null | undefined): string | undefined {
if (typeof raw !== 'string') return undefined;
const verb = raw.trim().toUpperCase();
// '*' marks a method-agnostic route. Preserve it so the contract layer emits a
// wildcard provider that matches consumers of any method.
if (verb === '*') return '*';
return VALID_HTTP_METHODS.has(verb) ? verb : undefined;
}
/**
* The Route node identity (#2289): `(method, path)` when the verb is known and
* specific, falling back to URL-only when the method is `undefined` (filesystem
* routes, Laravel `resource`/`apiResource`) or `'*'` (method-agnostic routes,
* e.g. Django function views). The URL-fallback keeps those byte-identical to
* the pre-#2289 URL-only ids, so only genuine declaration-style multi-verb
* routes (`GET /x` + `POST /x`) split into separate nodes.
*
* Used by the routes phase (node id + registry key), the processes phase
* (ENTRY_POINT_OF), and the handler-symbol resolver — all three must key
* identically.
*/
export function routeNodeKey(method: string | undefined, url: string): string {
return method && method !== '*' ? `${method} ${url}` : url;
}
@@ -59,6 +59,24 @@ export function findEnclosingClass(node: Parser.SyntaxNode): Parser.SyntaxNode |
return null;
}
/**
* Find the nearest enclosing Java type declaration (class OR interface) for a
* node, reporting its kind. Used by the ingestion route extractor to tell an
* interface-declared `@*Mapping` (handled by the cross-file inheritance pass,
* #2288) apart from a concrete class route, and by the type collector.
*/
export function findEnclosingType(
node: Parser.SyntaxNode,
): { node: Parser.SyntaxNode; kind: 'class' | 'interface' } | null {
let cur: Parser.SyntaxNode | null = node.parent;
while (cur) {
if (cur.type === 'class_declaration') return { node: cur, kind: 'class' };
if (cur.type === 'interface_declaration') return { node: cur, kind: 'interface' };
cur = cur.parent;
}
return null;
}
/**
* Strip enclosing quotes from a tree-sitter string-literal node's text.
* Handles single / double / template (backtick) quotes and triple-quoted
@@ -86,3 +104,164 @@ export function unquoteSpringLiteral(raw: string): string | null {
return raw;
}
/**
* Join a class/interface-level prefix and a method-level path into a single
* URL path: strip leading/trailing slashes on the prefix and leading slashes
* on the method path, then ensure exactly one slash between them.
*
* Lives here (the lower shared layer) so the ingestion route extractor and the
* group-layer Spring/Kotlin plugins join prefixes identically; the group
* `spring-consumer-shared.ts` re-exports it for its existing importers.
*/
export function joinPath(prefix: string, methodPath: string): string {
const cleanPrefix = prefix.replace(/^\/+/, '').replace(/\/+$/, '');
const cleanSub = methodPath.replace(/^\/+/, '');
if (!cleanPrefix) return `/${cleanSub}`;
return `/${cleanPrefix}/${cleanSub}`;
}
/**
* Join a controller's own class prefix with a route inherited from an interface
* (interface-based controllers, #1743). The inherited path already has the
* interface's own class prefix (`inheritedOwnerPrefix`) baked in; when the
* controller repeats that same prefix we must NOT prepend it twice (#2057).
*/
export function joinInheritedSpringPath(
controllerPrefix: string,
inheritedPath: string,
inheritedOwnerPrefix = '',
): string {
const joined = joinPath(controllerPrefix, inheritedPath);
const cleanPrefix = controllerPrefix.replace(/^\/+/, '').replace(/\/+$/, '');
const cleanOwnerPrefix = inheritedOwnerPrefix.replace(/^\/+/, '').replace(/\/+$/, '');
const cleanInherited = inheritedPath.replace(/^\/+/, '');
if (!cleanPrefix) return joined;
if (
cleanPrefix === cleanOwnerPrefix &&
(cleanInherited === cleanPrefix || cleanInherited.startsWith(`${cleanPrefix}/`))
) {
return `/${cleanInherited}`;
}
return joined;
}
/**
* Language-agnostic view of a Spring class/interface that each extractor's
* grammar-specific collector produces. The interface-based-controller
* inheritance algorithm (`resolveInheritedSpringRoutes`) operates only on this
* shape, so the ingestion route extractor and the group Java/Kotlin plugins
* share one algorithm and cannot drift.
*
* `methods[].routes` carry only `{ method, path }` — the interface's own class
* prefix is applied *inside* `resolveInheritedSpringRoutes` (it is not part of
* the collector's output).
*/
export interface SharedSpringType {
filePath: string;
kind: 'class' | 'interface';
name: string;
/** Class-level `@RequestMapping` prefixes — one per array element. */
classPrefixes: string[];
implementedInterfaces: string[];
isController: boolean;
methods: Array<{ name: string; routes: Array<{ method: string; path: string }> }>;
}
/** One provider route a concrete controller inherits from an interface. */
interface InheritedSpringRoute {
/** File of the implementing controller (where the route should be attributed). */
filePath: string;
/** Name of the controller method that inherits the route. */
methodName: string;
method: string;
path: string;
}
/**
* An interface route with its own class prefix already baked in. `ownerPrefix`
* records that prefix so the controller side avoids doubling it (#2057).
*/
interface IntermediateRoute {
method: string;
path: string;
ownerPrefix: string;
}
/**
* Resolve interface-based-controller provider routes (#1743): a concrete
* `@RestController`/`@Controller` class inherits the `@(Get|...)Mapping` routes
* declared on the interface it implements. Pure and language-agnostic — shared
* by the ingestion route extractor and the group Java/Kotlin plugins so all
* three emit the same inherited routes.
*
* An interface name that resolves to two distinct interfaces is ambiguous and
* its routes are dropped (the `null` marker). The controller's own class
* prefix(es) cross-product the inherited routes; `joinInheritedSpringPath`
* avoids doubling a prefix the interface already baked in (#2057). Duplicate
* `(method, path)` results per controller method are de-duped.
*/
export function resolveInheritedSpringRoutes(types: SharedSpringType[]): InheritedSpringRoute[] {
// interface name → (method name → routes). `null` marks an ambiguous
// (duplicated) interface name. `IntermediateRoute` is declared at module scope.
const interfaceRoutes = new Map<string, Map<string, IntermediateRoute[]> | null>();
for (const type of types) {
if (type.kind !== 'interface') continue;
if (interfaceRoutes.has(type.name)) {
interfaceRoutes.set(type.name, null);
continue;
}
const prefixes = type.classPrefixes.length ? type.classPrefixes : [''];
const methodMap = new Map<string, IntermediateRoute[]>();
for (const method of type.methods) {
// Cross-product the interface's class prefixes with each method route, so a
// multi-element `@RequestMapping(["/a","/b"])` interface yields N bindings.
const routes = method.routes.flatMap((route) =>
prefixes.map((prefix) => ({
method: route.method,
path: prefix ? joinPath(prefix, route.path) : route.path,
ownerPrefix: prefix,
})),
);
if (routes.length > 0) methodMap.set(method.name, routes);
}
interfaceRoutes.set(type.name, methodMap);
}
const out: InheritedSpringRoute[] = [];
for (const type of types) {
if (type.kind !== 'class' || !type.isController) continue;
// Cross-product the controller's own class prefixes with each inherited
// route; `['']` keeps the common no-prefix controller emitting the
// interface path unchanged.
const controllerPrefixes = type.classPrefixes.length ? type.classPrefixes : [''];
for (const method of type.methods) {
if (method.routes.length > 0) continue; // own @*Mapping → already a provider
const inherited = type.implementedInterfaces.flatMap((iface) => {
const routeMap = interfaceRoutes.get(iface);
if (!routeMap) return [];
const routes = routeMap.get(method.name) ?? [];
return routes.flatMap((route) =>
controllerPrefixes.map((controllerPrefix) => ({
method: route.method,
path: joinInheritedSpringPath(controllerPrefix, route.path, route.ownerPrefix),
})),
);
});
const seen = new Set<string>();
for (const route of inherited) {
const key = `${route.method} ${route.path}`;
if (seen.has(key)) continue;
seen.add(key);
out.push({
filePath: type.filePath,
methodName: method.name,
method: route.method,
path: route.path,
});
}
}
}
return out;
}
@@ -25,8 +25,9 @@ import type { ExtractedDecoratorRoute } from '../workers/parse-worker.js';
import {
METHOD_ANNOTATION_TO_HTTP,
isRouteMemberKey,
findEnclosingClass,
findEnclosingType,
unquoteSpringLiteral,
type SharedSpringType,
} from './spring-shared.js';
/**
@@ -39,6 +40,18 @@ import {
* @node → enclosing declaration (class_declaration | method_declaration)
* @value → the string-literal argument
* @key → the named-argument member key (absent for positional form)
*
* Method-level routes accept both the bare string form `@GetMapping("/x")` and
* the array form `@GetMapping({"/a","/b"})` (positional or `path =`/`value =`):
* a multi-element array yields one match per element, so the Phase 2 loop emits
* one route per path with no special-casing. This mirrors the group-layer
* `java.ts` query so the two Spring extractors stay in parity (#2138 follow-up;
* the divergence here was the root of the #2265 array-form gap). The class-level
* `@RequestMapping` branches also match the array form, but only to *detect* it:
* an array-form class prefix can't be resolved to a single string, so Phase 2
* suppresses that class's method-level array routes rather than emit them with a
* dropped prefix (a wrong route). Full class-array cross-product support is left
* to a follow-up (#2280).
*/
const ROUTE_ANNOTATION_QUERY = new Parser.Query(
Java,
@@ -48,7 +61,9 @@ const ROUTE_ANNOTATION_QUERY = new Parser.Query(
(modifiers
(annotation
name: (identifier) @ann
arguments: (annotation_argument_list (string_literal) @value)))) @node
arguments: (annotation_argument_list
[(string_literal) @value
(element_value_array_initializer (string_literal) @value)])))) @node
(class_declaration
(modifiers
(annotation
@@ -56,12 +71,15 @@ const ROUTE_ANNOTATION_QUERY = new Parser.Query(
arguments: (annotation_argument_list
(element_value_pair
key: (identifier) @key
value: (string_literal) @value))))) @node
value: [(string_literal) @value
(element_value_array_initializer (string_literal) @value)]))))) @node
(method_declaration
(modifiers
(annotation
name: (identifier) @ann
arguments: (annotation_argument_list (string_literal) @value)))) @node
arguments: (annotation_argument_list
[(string_literal) @value
(element_value_array_initializer (string_literal) @value)])))) @node
(method_declaration
(modifiers
(annotation
@@ -69,7 +87,8 @@ const ROUTE_ANNOTATION_QUERY = new Parser.Query(
arguments: (annotation_argument_list
(element_value_pair
key: (identifier) @key
value: (string_literal) @value))))) @node
value: [(string_literal) @value
(element_value_array_initializer (string_literal) @value)]))))) @node
]
`,
);
@@ -93,8 +112,15 @@ export function extractSpringRoutes(
): ExtractedDecoratorRoute[] {
const matches = ROUTE_ANNOTATION_QUERY.matches(tree.rootNode);
// Phase 1: collect class-level @RequestMapping prefixes keyed by node id
// Phase 1: collect class-level @RequestMapping prefixes keyed by node id.
// A scalar prefix (`@RequestMapping("/base")`) is stored in prefixByClassId.
// A class whose @RequestMapping uses the array form (`@RequestMapping({...})`)
// is instead recorded in classesWithArrayPrefix: there is no single prefix to
// store, and Phase 2 uses this to suppress that class's method-level array
// routes rather than emit them unprefixed (a wrong route — see #2280). Full
// class-array cross-product support is out of scope here.
const prefixByClassId = new Map<number, string>();
const classesWithArrayPrefix = new Set<number>();
for (const match of matches) {
const caps: Record<string, Parser.SyntaxNode> = {};
@@ -109,6 +135,10 @@ export function extractSpringRoutes(
if (node.type === 'class_declaration' && annNode.text === 'RequestMapping') {
if (!isRouteMemberKey(keyNode)) continue;
if (valueNode.parent?.type === 'element_value_array_initializer') {
classesWithArrayPrefix.add(node.id);
continue;
}
const prefix = unquoteSpringLiteral(valueNode.text);
if (prefix !== null) prefixByClassId.set(node.id, prefix);
}
@@ -137,8 +167,33 @@ export function extractSpringRoutes(
const routePath = unquoteSpringLiteral(valueNode.text);
if (routePath === null) continue;
const enclosingClass = findEnclosingClass(node);
const enclosingType = findEnclosingType(node);
// Interface-declared `@*Mapping`s are not concrete routes on their own — the
// implementing controller inherits them. Skip here; the cross-file
// inheritance pass (#2288) re-emits them attributed to the controller, with
// both the interface's and the controller's class prefixes resolved. Emitting
// the interface route directly would be wrong (unprefixed, wrong owner).
if (enclosingType?.kind === 'interface') continue;
const enclosingClass = enclosingType?.kind === 'class' ? enclosingType.node : null;
// Suppress a method-level *array-form* route nested under a class-level
// array-form @RequestMapping. The class prefix is one of several values that
// cannot be resolved to a single string here, so emitting the route would
// drop the prefix and yield a wrong unprefixed Route (a false signal, worse
// than a missing one). Skipping keeps ingestion a strict subset of the group
// scan — safe under routeCoverage:'partial'. Full class-array cross-product
// support is tracked in #2280. (Scalar method paths under an array class
// prefix are left unchanged: that pre-existing divergence is out of scope.)
const isArrayElement = valueNode.parent?.type === 'element_value_array_initializer';
if (isArrayElement && enclosingClass && classesWithArrayPrefix.has(enclosingClass.id)) {
continue;
}
const classPrefix = enclosingClass ? (prefixByClassId.get(enclosingClass.id) ?? '') : '';
// `node` is the annotated `method_declaration`; its name field is the
// handler method name (resolved to a symbol UID later by the routes phase).
const handlerName = node.childForFieldName('name')?.text;
routes.push({
filePath,
@@ -147,8 +202,173 @@ export function extractSpringRoutes(
decoratorName: ann,
lineNumber: annNode.startPosition.row + lineOffset,
...(classPrefix ? { prefix: classPrefix } : {}),
...(handlerName ? { handlerName } : {}),
});
}
return routes;
}
/**
* Tree-sitter query capturing every Java type declaration (class + interface),
* used by `extractSpringTypes` to build the project-wide `SharedSpringType`
* view the cross-file interface-inheritance pass consumes (#2288).
*/
const TYPE_DECLARATION_QUERY = new Parser.Query(
Java,
`[(class_declaration) @type (interface_declaration) @type]`,
);
/** Direct annotations on a type/method declaration (reads its `modifiers` child). */
function declarationAnnotations(node: Parser.SyntaxNode): Parser.SyntaxNode[] {
// `modifiers` is a NAMED CHILD of the declaration (not a named field) in
// tree-sitter-java — matching the group layer's `hasAnnotation`.
const modifiers = node.namedChildren.find((c) => c.type === 'modifiers');
if (!modifiers) return [];
return modifiers.namedChildren.filter(
(c) => c.type === 'annotation' || c.type === 'marker_annotation',
);
}
const annotationName = (ann: Parser.SyntaxNode): string | undefined => {
// Trailing segment of a possibly fully-qualified annotation name
// (`org.springframework.web.bind.annotation.GetMapping` → `GetMapping`), so a
// FQN annotation is classified the same as its simple form — matching the
// group layer's `simpleName` normalization. A simple name maps to itself.
const text = ann.childForFieldName('name')?.text;
return text?.split('.').pop() ?? text;
};
/**
* Collect the route path(s) carried by an annotation's argument list, honoring
* both positional and `path =`/`value =` named arguments and both the bare
* string and array forms. Non-route named args (`consumes`, `produces`, …) are
* dropped via `isRouteMemberKey`. Returns one entry per string element.
*/
function annotationRoutePaths(ann: Parser.SyntaxNode): string[] {
const args = ann.childForFieldName('arguments');
if (!args) return [];
const out: string[] = [];
const pushLiteral = (lit: Parser.SyntaxNode): void => {
const v = unquoteSpringLiteral(lit.text);
if (v !== null) out.push(v);
};
const pushFromValue = (valueNode: Parser.SyntaxNode): void => {
if (valueNode.type === 'string_literal') pushLiteral(valueNode);
else if (valueNode.type === 'element_value_array_initializer') {
for (const el of valueNode.namedChildren) if (el.type === 'string_literal') pushLiteral(el);
}
};
for (const child of args.namedChildren) {
if (child.type === 'string_literal' || child.type === 'element_value_array_initializer') {
pushFromValue(child); // positional
} else if (child.type === 'element_value_pair') {
const key = child.childForFieldName('key');
if (!isRouteMemberKey(key ?? undefined)) continue;
const value = child.childForFieldName('value');
if (value) pushFromValue(value);
}
}
return out;
}
/** Class-level `@RequestMapping` prefixes for a type (array-aware; may be []). */
function typeClassPrefixes(typeNode: Parser.SyntaxNode): string[] {
const prefixes: string[] = [];
for (const ann of declarationAnnotations(typeNode)) {
if (annotationName(ann) === 'RequestMapping') prefixes.push(...annotationRoutePaths(ann));
}
return prefixes;
}
/** The simple names of the interfaces a class declares via `implements`. */
function implementedInterfaceNames(typeNode: Parser.SyntaxNode): string[] {
const interfacesNode = typeNode.childForFieldName('interfaces');
if (!interfacesNode) return [];
const out: string[] = [];
const visit = (node: Parser.SyntaxNode): void => {
if (node.type === 'type_identifier' || node.type === 'scoped_type_identifier') {
out.push(node.text.split('.').pop() ?? node.text);
return;
}
for (const child of node.namedChildren) visit(child);
};
visit(interfacesNode);
return out;
}
/** Direct `method_declaration` children of a type (not methods of nested types). */
function directMethods(typeNode: Parser.SyntaxNode): Parser.SyntaxNode[] {
const out: Parser.SyntaxNode[] = [];
const visit = (node: Parser.SyntaxNode): void => {
for (const child of node.namedChildren) {
if (child.type === 'method_declaration') {
out.push(child);
continue;
}
// Don't descend into a nested type — its methods aren't this type's.
if (
child !== typeNode &&
(child.type === 'class_declaration' || child.type === 'interface_declaration')
) {
continue;
}
visit(child);
}
};
visit(typeNode);
return out;
}
/**
* Build the project-wide `SharedSpringType` view for one Java file: every class
* and interface with its class prefixes, implemented interfaces, controller
* flag, and per-method route annotations. The cross-file inheritance pass
* (#2288) feeds these into the shared `resolveInheritedSpringRoutes` so a
* concrete controller inherits the `@*Mapping`s declared on its interfaces.
*
* This is the ingestion counterpart of the group layer's `collectSpringTypes`
* (`group/extractors/http-patterns/java.ts`); both produce the same neutral
* shape so the two layers resolve inheritance identically (#2078 parity).
*/
export function extractSpringTypes(tree: Parser.Tree, filePath: string): SharedSpringType[] {
const out: SharedSpringType[] = [];
for (const match of TYPE_DECLARATION_QUERY.matches(tree.rootNode)) {
const typeNode = match.captures.find((c) => c.name === 'type')?.node;
if (!typeNode) continue;
const name = typeNode.childForFieldName('name')?.text;
if (!name) continue;
const kind = typeNode.type === 'interface_declaration' ? 'interface' : 'class';
const annNames = declarationAnnotations(typeNode).map(annotationName);
const isController =
kind === 'class' && (annNames.includes('RestController') || annNames.includes('Controller'));
const methods = directMethods(typeNode)
.map((methodNode) => {
const methodName = methodNode.childForFieldName('name')?.text;
if (!methodName) return null;
const routes: Array<{ method: string; path: string }> = [];
for (const ann of declarationAnnotations(methodNode)) {
const verb = METHOD_ANNOTATION_TO_HTTP[annotationName(ann) ?? ''];
if (!verb) continue;
for (const path of annotationRoutePaths(ann)) routes.push({ method: verb, path });
}
return { name: methodName, routes };
})
.filter(
(m): m is { name: string; routes: Array<{ method: string; path: string }> } => m !== null,
);
out.push({
filePath,
kind,
name,
classPrefixes: typeClassPrefixes(typeNode),
implementedInterfaces: kind === 'class' ? implementedInterfaceNames(typeNode) : [],
isController,
methods,
});
}
return out;
}
@@ -27,6 +27,9 @@ export const TYPESCRIPT_QUERIES = `
(function_declaration
name: (identifier) @name) @definition.function
(generator_function_declaration
name: (identifier) @name) @definition.function
; TypeScript overload signatures (function_signature is a separate node type from function_declaration)
(function_signature
name: (identifier) @name) @definition.function
@@ -374,6 +377,9 @@ export const JAVASCRIPT_QUERIES = `
(function_declaration
name: (identifier) @name) @definition.function
(generator_function_declaration
name: (identifier) @name) @definition.function
(method_definition
name: (property_identifier) @name) @definition.method
@@ -87,7 +87,7 @@ export const DEFINITION_CAPTURE_KEYS = [
/** Extract the definition node from a tree-sitter query capture map. */
export const getDefinitionNodeFromCaptures = (
captureMap: Record<string, SyntaxNode>,
captureMap: Record<string, SyntaxNode | undefined>,
): SyntaxNode | null => {
for (const key of DEFINITION_CAPTURE_KEYS) {
if (captureMap[key]) return captureMap[key];
@@ -310,7 +310,7 @@ export function findAncestorBeforeBoundary(
* Returns null if the capture should be skipped (import, call, C/C++ duplicate, missing name).
*/
export function getLabelFromCaptures(
captureMap: Record<string, SyntaxNode>,
captureMap: Record<string, SyntaxNode | undefined>,
provider: LanguageProvider,
): NodeLabel | null {
if (captureMap['import'] || captureMap['call']) return null;
@@ -926,6 +926,227 @@ export function findChild(node: SyntaxNode, type: string): SyntaxNode | null {
return null;
}
/** Remove bidi-override and zero-width control characters. Doc text is
* attacker-influenced (any indexed repo) and is returned verbatim to MCP
* clients, so strip Trojan-Source-style hidden controls from the description
* before it leaves the extractor (#2286 review). Scoped to the doc-comment path
* only — global `sanitizeUTF8` is intentionally untouched. */
const stripBidiAndZeroWidth = (text: string): string =>
Array.from(text)
.filter((ch) => {
const c = ch.codePointAt(0) ?? 0;
// Bidi overrides/isolates (U+202A–202E, U+2066–2069), zero-width
// space/joiners (U+200B–200D), and BOM/zero-width-no-break (U+FEFF).
return !(
(c >= 0x202a && c <= 0x202e) ||
(c >= 0x2066 && c <= 0x2069) ||
(c >= 0x200b && c <= 0x200d) ||
c === 0xfeff
);
})
.join('');
/** Normalize a block doc comment body: strip the opening (double-star or
* bang) delimiter, the closing delimiter, and per-line gutter stars, then
* collapse whitespace so tag content stays as searchable words. */
const normalizeBlockDocComment = (text: string): string | undefined => {
const inner = stripBidiAndZeroWidth(
text
.replace(/^\/\*[*!]/, '')
// Close delimiter: tolerate the degenerate empty comment `/**/`, where the
// opening strip already consumed the shared `*`, leaving a lone `/`.
.replace(/\*?\/\s*$/, '')
.replace(/^[ \t]*\*[ \t]?/gm, ' ')
.replace(/\s+/g, ' ')
.trim(),
);
return inner.length > 0 ? inner : undefined;
};
/** Default line-comment prefixes treated as documentation: the universal
* triple-slash / bang-slash doc markers (Rust, C#, Dart, Swift, Doxygen).
* Go (`//`) and Ruby (`#`) opt into their conventional markers explicitly. */
const DEFAULT_LINE_DOC_PREFIXES: readonly string[] = ['///', '//!'];
/** Default block-comment doc openers: Javadoc/JSDoc-style `/**` and Doxygen
* `/*!`. Rust opts out of `/*!` (and `//!`) because those are *inner* docs that
* document the enclosing item, not the following one. */
const DEFAULT_BLOCK_DOC_PREFIXES: readonly string[] = ['/**', '/*!'];
/** A file-top `/** … *\/` license/copyright/file-overview block has no
* package/import sibling to shield it, so it would otherwise be absorbed as the
* first declaration's description (PR #2286 review). These markers identify such
* headers; they are specific enough not to fire on an ordinary symbol doc that
* merely mentions the word "copyright". `@file`/`@fileoverview` are explicitly
* file-level JSDoc tags, so a block carrying them is not a symbol doc. */
const FILE_HEADER_MARKER =
/SPDX-License-Identifier|@licen[sc]e\b|@fileoverview\b|@file\b|Licen[sc]ed under|copyright\s*(\(c\)|©|\d{4})/i;
/**
* Extract the normalized text of a leading doc comment immediately preceding a
* definition node — covering both block doc comments (Javadoc / KDoc / JSDoc /
* PHPDoc / Doxygen, opened by `/**` or `/*!`) and runs of line doc comments
* (`///`, `//!`, or the caller-supplied prefixes such as Go's `//` or Ruby's
* `#`). Returns `undefined` when there is no preceding doc comment or it is
* empty.
*
* Grammar-agnostic by design: matches on the comment text prefix rather than a
* grammar node type, because the comment node is named differently across
* grammars (`block_comment`, `multiline_comment`, `comment`, `line_comment`).
* Annotations and modifiers live inside the definition node, so the doc comment
* remains the definition's `previousNamedSibling` even on annotated/decorated
* declarations.
*
* Block comments are taken as the immediately-preceding sibling (intervening
* package/import/code siblings already shield a file-level license block from
* the first declaration). Line doc comments enforce row-adjacency: the first
* comment must sit on the line directly above the definition, and each comment
* walked further up must sit directly above the previous one — so a run stops
* at a blank line. This matches godoc/RDoc/rustdoc convention and prevents an
* unrelated comment block (a license header, a Ruby shebang + magic comment)
* separated by a blank line from being absorbed. Adjacency is checked on
* `startPosition.row` (reliable) rather than `endPosition.row`, since some
* grammars fold the trailing newline into the comment node.
*
* Normalization mirrors Python docstring handling: strip the comment delimiters
* / per-line markers, then collapse whitespace to single spaces so tag content
* (`@param`, `@deprecated since 2.0, use computeBalanceV2`) survives.
*
* When the captured definition is an inner node and its own preceding sibling
* carries no doc, the search retries from a wrapping node whose type is listed in
* `opts.wrapperNodeTypes` (e.g. an `export_statement` wrapping an exported
* function/class — the JSDoc precedes the wrapper, not the inner declaration).
*/
export interface LeadingDocCommentOptions {
/** Line-comment doc prefixes (defaults to {@link DEFAULT_LINE_DOC_PREFIXES};
* Go passes `['//']`, Ruby passes `['#']`). */
lineCommentPrefixes?: readonly string[];
/** Grammar node types that wrap a definition such that the doc comment is the
* wrapper's preceding sibling rather than the definition's. TS/JS pass
* `['export_statement']`. Empty by default → no wrapper retry. */
wrapperNodeTypes?: readonly string[];
/** Line-comment prefixes that are tool/build directives or magic comments
* rather than documentation (Go passes `['//go:', '// +build', …]`, Ruby
* passes `['# frozen_string_literal:', '#!', …]`). A matching line is skipped
* in the doc run rather than absorbed. Empty by default. */
lineDirectivePrefixes?: readonly string[];
/** Block-comment doc openers (defaults to `['/**', '/*!']`). Rust passes
* `['/**']` so its inner-doc `/*!` does not attach to the following item. */
blockDocPrefixes?: readonly string[];
}
export function extractLeadingDocComment(
node: SyntaxNode,
opts: LeadingDocCommentOptions = {},
): string | undefined {
const lineCommentPrefixes = opts.lineCommentPrefixes ?? DEFAULT_LINE_DOC_PREFIXES;
const wrapperNodeTypes = opts.wrapperNodeTypes ?? [];
const lineDirectivePrefixes = opts.lineDirectivePrefixes ?? [];
const blockDocPrefixes = opts.blockDocPrefixes ?? DEFAULT_BLOCK_DOC_PREFIXES;
const fromNode = (anchor: SyntaxNode): string | undefined => {
const prev = anchor.previousNamedSibling;
if (!prev) return undefined;
// Block doc comment: /** ... */ or /*! ... */
if (blockDocPrefixes.some((p) => prev.text.startsWith(p))) {
// Skip a file-top license/copyright/overview header (no package/import
// sibling shields it from the first declaration). A strict row-adjacency
// check is unreliable here — some grammars fold the trailing newline into
// the comment node — so match header markers instead.
if (FILE_HEADER_MARKER.test(prev.text)) return undefined;
return normalizeBlockDocComment(prev.text);
}
// Run of row-adjacent preceding line doc comments (e.g. `///` or `//`).
const matchedPrefix = (text: string): string | undefined =>
lineCommentPrefixes.find((prefix) => text.trimStart().startsWith(prefix));
const isDirective = (text: string): boolean =>
lineDirectivePrefixes.some((prefix) => text.trimStart().startsWith(prefix));
const lines: string[] = [];
let current: SyntaxNode | null = prev;
let expectedRow = anchor.startPosition.row - 1;
while (current) {
const text = current.text;
const prefix = matchedPrefix(text);
if (prefix === undefined || current.startPosition.row !== expectedRow) break;
// A build/tool directive or magic comment (e.g. `//go:build`,
// `# frozen_string_literal:`) is not documentation: skip it but keep
// walking the adjacent run, so a real doc above it is still collected.
if (!isDirective(text)) lines.unshift(text.trimStart().slice(prefix.length));
expectedRow = current.startPosition.row - 1;
current = current.previousNamedSibling;
}
const joined = stripBidiAndZeroWidth(lines.join(' ').replace(/\s+/g, ' ').trim());
return joined.length > 0 ? joined : undefined;
};
const direct = fromNode(node);
if (direct !== undefined) return direct;
const parent = node.parent;
if (parent && wrapperNodeTypes.includes(parent.type)) {
return fromNode(parent);
}
return undefined;
}
/** Node labels that can carry a leading doc comment — callables and type-like
* declarations. Field/property/variable/const doc is intentionally excluded
* (issue #2270 scopes this to method/type documentation). Language-neutral:
* a label a given grammar never emits simply never matches.
*
* Bounded to labels that are also in `embeddings/types.ts` `EMBEDDABLE_LABELS`:
* the description is only useful once it reaches the embedding metadata header,
* and the embedding pipeline only queries embeddable labels. Extracting docs
* for a non-embeddable label is a wasted write that never becomes searchable.
* A subset invariant in the unit tests guards against drift. Making currently-
* non-embeddable doc-bearing labels (Module, Delegate, Annotation, and C++
* `Template`) searchable is tracked as a follow-up — it needs an embedding-
* pipeline/schema change beyond this fix. */
export const DOC_BEARING_LABELS: ReadonlySet<NodeLabel> = new Set<NodeLabel>([
'Function',
'Method',
'Constructor',
'Class',
'Interface',
'Enum',
'Struct',
'Trait',
'Record',
'Union',
'Namespace',
'TypeAlias',
'Macro',
]);
/**
* Build a `LanguageProvider.descriptionExtractor` that surfaces a definition's
* leading doc comment as its `description` (issue #2270). For labels in
* {@link DOC_BEARING_LABELS} (which is bounded to embeddable labels) the text
* then reaches the embedding metadata header and becomes semantically searchable.
*
* Language-neutral factory (names no language): guards on
* {@link DOC_BEARING_LABELS}; callers pass per-language doc-comment behavior via
* {@link LeadingDocCommentOptions} (line prefixes, export-style wrappers, …)
* which is threaded straight through to {@link extractLeadingDocComment}.
*/
export const createLeadingDocDescriptionExtractor = (
opts: LeadingDocCommentOptions = {},
): ((
nodeLabel: NodeLabel,
nodeName: string,
captureMap: Record<string, SyntaxNode | undefined>,
) => string | undefined) => {
return (nodeLabel, _nodeName, captureMap) => {
if (!DOC_BEARING_LABELS.has(nodeLabel)) return undefined;
const definitionNode = getDefinitionNodeFromCaptures(captureMap);
return definitionNode ? extractLeadingDocComment(definitionNode, opts) : undefined;
};
};
// ============================================================================
// Capture + range helpers (formerly python/ast-utils.ts — language-agnostic)
// ============================================================================
@@ -30,6 +30,7 @@ import { postResultCloneSafe } from './post-result.js';
import { mergeResult } from './result-merge.js';
import type { SymbolTableReader } from '../model/symbol-table.js';
import type {
ExtractedRouterConstructorPrefix,
ExtractedRouterInclude,
ExtractedRouterImport,
ExtractedRouterModuleAlias,
@@ -130,6 +131,7 @@ import {
persistDurableParsedFileShardSync,
} from '../../../storage/parsedfile-store.js';
import { extractLaravelRoutes, type ExtractedRoute } from '../route-extractors/laravel.js';
import type { SharedSpringType } from '../route-extractors/spring-shared.js';
import {
collectFunctionCfgs,
DEFAULT_PDG_MAX_FUNCTION_LINES,
@@ -317,6 +319,16 @@ export interface ExtractedDecoratorRoute {
* absent ⇒ no prefix applies.
*/
prefix?: string | null;
/**
* Name of the handler the route decorator sits on (the decorated
* method/function — e.g. `create` for `@PostMapping("/orders") Order create()`).
* Captured at extraction where the decorated definition node is in hand, so
* the routes phase can resolve it to a real handler symbol UID via the
* SemanticModel (same `(filePath, name) → nodeId` lookup Laravel routes use).
* Absent when the extractor could not identify the decorated definition;
* resolution then falls back (the Route node simply carries no handlerSymbolId).
*/
handlerName?: string;
}
export interface ExtractedToolDef {
@@ -383,6 +395,16 @@ export interface ParseWorkerResult {
decoratorRoutes: ExtractedDecoratorRoute[];
routerIncludes: ExtractedRouterInclude[];
routerImports: ExtractedRouterImport[];
routerConstructorPrefixes?: ExtractedRouterConstructorPrefix[];
/**
* Optional. Project-wide `SharedSpringType` view of route-defining
* class/interface declarations, produced by the provider's
* `extractRouteInheritanceTypes` hook (Java/Spring). parse-impl aggregates
* these and runs a cross-file pass that resolves interface-inherited routes
* into additional `decoratorRoutes` (#2288). Optional for cache backward
* compatibility; consumers must guard with `?? []`.
*/
springTypes?: SharedSpringType[];
/**
* Optional. `from <pkg> import <module>` records from Python files
* where `<module>` is later used as a Shape-A include receiver
@@ -886,6 +908,7 @@ const processBatch = (
decoratorRoutes: [],
routerIncludes: [],
routerImports: [],
routerConstructorPrefixes: [],
routerModuleAliases: [],
toolDefs: [],
ormQueries: [],
@@ -2141,7 +2164,20 @@ const processFileGroup = (
`${file.path}:${qualifiedName}${classTemplateTag}${arityTag}${parameterShapeTag}${constraintsTag}`,
);
const description = provider.descriptionExtractor?.(nodeLabel, nodeName, captureMap);
let description: string | undefined;
try {
description = provider.descriptionExtractor?.(nodeLabel, nodeName, captureMap);
} catch (err) {
// A throw here (an unexpected tree-sitter node shape, a provider bug) must
// NOT propagate — it would escape processFileGroup to the language-group
// catch, which treats any throw as "parser unavailable" and silently drops
// every remaining file in the group. Mirrors the extractTemplateConstraints
// guard above (#2286 review).
reportWarning(
`Description extraction failed for ${file.path}: ${err instanceof Error ? err.message : String(err)}`,
);
description = undefined;
}
let frameworkHint = definitionNode
? detectFrameworkFromAST(language, (definitionNode.text || '').slice(0, 300))
@@ -2390,6 +2426,7 @@ const processFileGroup = (
result.routerIncludes,
result.routerImports,
(result.routerModuleAliases ??= []),
(result.routerConstructorPrefixes ??= []),
);
}
@@ -2402,6 +2439,14 @@ const processFileGroup = (
for (const r of frameworkRoutes) result.decoratorRoutes.push(r);
}
// Project-wide route-inheritance type collection via provider hook (#2288).
// The per-file SharedSpringType views are aggregated by the parse phase,
// which then resolves interface-inherited routes cross-file.
if (provider.extractRouteInheritanceTypes) {
const springTypes = provider.extractRouteInheritanceTypes(tree, file.path);
if (springTypes.length > 0) (result.springTypes ??= []).push(...springTypes);
}
// Vue: emit CALLS edges for components used in <template>
if (language === SupportedLanguages.Vue) {
const templateComponents = extractTemplateComponents(file.content);
@@ -2434,6 +2479,7 @@ let accumulated: ParseWorkerResult = {
decoratorRoutes: [],
routerIncludes: [],
routerImports: [],
routerConstructorPrefixes: [],
routerModuleAliases: [],
toolDefs: [],
ormQueries: [],
@@ -2574,6 +2620,7 @@ parentPort!.on('message', (msg: WorkerIncomingMessage) => {
decoratorRoutes: [],
routerIncludes: [],
routerImports: [],
routerConstructorPrefixes: [],
routerModuleAliases: [],
toolDefs: [],
ormQueries: [],
@@ -37,10 +37,18 @@ export const mergeResult = (target: ParseWorkerResult, src: ParseWorkerResult):
appendAll(target.decoratorRoutes, src.decoratorRoutes);
if (src.routerIncludes) appendAll(target.routerIncludes, src.routerIncludes);
if (src.routerImports) appendAll(target.routerImports, src.routerImports);
if (src.routerConstructorPrefixes) {
target.routerConstructorPrefixes ??= [];
appendAll(target.routerConstructorPrefixes, src.routerConstructorPrefixes);
}
if (src.routerModuleAliases) {
target.routerModuleAliases ??= [];
appendAll(target.routerModuleAliases, src.routerModuleAliases);
}
if (src.springTypes) {
target.springTypes ??= [];
appendAll(target.springTypes, src.springTypes);
}
appendAll(target.toolDefs, src.toolDefs);
appendAll(target.ormQueries, src.ormQueries);
appendAll(target.constructorBindings, src.constructorBindings);
+2 -1
View File
@@ -381,7 +381,7 @@ export const streamAllCSVsToDisk = async (
// Route nodes for API endpoint mapping
const routeWriter = new BufferedCSVWriter(
path.join(csvDir, 'route.csv'),
'id,name,filePath,responseKeys,errorKeys,middleware,method',
'id,name,filePath,responseKeys,errorKeys,middleware,method,handlerSymbolId',
);
// Tool nodes for MCP tool definitions
@@ -561,6 +561,7 @@ export const streamAllCSVsToDisk = async (
escapeCSVField(errorKeysStr),
escapeCSVField(middlewareStr),
escapeCSVField(String(node.properties.method ?? '')),
escapeCSVField(String(node.properties.handlerSymbolId ?? '')),
].join(','),
);
break;
+10 -3
View File
@@ -1337,7 +1337,7 @@ export const getCopyQuery = (table: NodeTableName, filePath: string): string =>
return `COPY ${t}(id, name, filePath, startLine, endLine, level, content, description) FROM "${filePath}" ${COPY_CSV_OPTS}`;
}
if (table === 'Route') {
return `COPY ${t}(id, name, filePath, responseKeys, errorKeys, middleware, method) FROM "${filePath}" ${COPY_CSV_OPTS}`;
return `COPY ${t}(id, name, filePath, responseKeys, errorKeys, middleware, method, handlerSymbolId) FROM "${filePath}" ${COPY_CSV_OPTS}`;
}
if (table === 'Tool') {
return `COPY ${t}(id, name, filePath, description) FROM "${filePath}" ${COPY_CSV_OPTS}`;
@@ -2334,6 +2334,13 @@ export const loadVectorExtension = async (
if (loaded && useModuleState) vectorExtensionLoaded = true;
return loaded;
};
/**
* Default stemmer for FTS indexes. Single source so the analyze path
* (`getSearchFTSStemmer`) and the read-only `createFTSIndex`/`ensureFTSIndex`
* defaults can never silently diverge.
*/
export const DEFAULT_FTS_STEMMER = 'porter';
/**
* Create a full-text search index on a table
* @param tableName - The node table name (e.g., 'File', 'CodeSymbol')
@@ -2345,7 +2352,7 @@ export const createFTSIndex = async (
tableName: string,
indexName: string,
properties: string[],
stemmer: string = 'porter',
stemmer: string = DEFAULT_FTS_STEMMER,
): Promise<void> => {
if (!conn) {
throw new Error('LadybugDB not initialized. Call initLbug first.');
@@ -2441,7 +2448,7 @@ export const ensureFTSIndex = async (
tableName: string,
indexName: string,
properties: string[],
stemmer: string = 'porter',
stemmer: string = DEFAULT_FTS_STEMMER,
): Promise<void> => {
const key = ftsIndexKey(tableName, indexName);
if (ensuredFTSIndexes.has(key)) return;
+1
View File
@@ -195,6 +195,7 @@ CREATE NODE TABLE Route (
errorKeys STRING[],
middleware STRING[],
method STRING,
handlerSymbolId STRING,
PRIMARY KEY (id)
)`;
+15 -3
View File
@@ -38,6 +38,12 @@ export interface CreateLoggerOptions {
debugEnvVar?: string;
/** Override destination stream — primarily for tests. */
destination?: DestinationStream;
/**
* Explicit level for the destination-override path — primarily for tests that
* need to capture below the default `info` (e.g. asserting a `debug` record).
* Ignored unless `destination` is set; `debugEnvVar` still wins when truthy.
*/
level?: string;
}
function isTruthyEnv(value: string | undefined): boolean {
@@ -219,7 +225,7 @@ export function createLogger(name: string, opts?: CreateLoggerOptions): Logger {
if (opts?.destination) {
return pino(
{ level: debugRequested ? 'debug' : 'info', base: undefined, name },
{ level: debugRequested ? 'debug' : (opts.level ?? 'info'), base: undefined, name },
opts.destination,
);
}
@@ -247,6 +253,7 @@ export function createLogger(name: string, opts?: CreateLoggerOptions): Logger {
/* ------------------------------------------------------------------ */
let _activeDestination: DestinationStream | undefined;
let _activeLevel: string | undefined;
let _cached: Logger | undefined;
function _getInner(): Logger {
@@ -256,7 +263,7 @@ function _getInner(): Logger {
// by `_captureLogger()` below.
_cached = createLogger(
'gitnexus',
_activeDestination ? { destination: _activeDestination } : undefined,
_activeDestination ? { destination: _activeDestination, level: _activeLevel } : undefined,
);
return _cached;
}
@@ -342,10 +349,13 @@ export interface LoggerCapture {
* expect(cap.records().some(r => r.msg?.includes('clamping'))).toBe(true);
* });
*
* Pass `level` (e.g. 'debug') to capture below the default 'info' — needed to
* assert that a record was emitted at debug rather than merely absent.
*
* Not a public API; underscore-prefixed and called only from test code.
* Throws if a previous capture is still active — see the body for context.
*/
export function _captureLogger(): LoggerCapture {
export function _captureLogger(level?: string): LoggerCapture {
// Guard against double-capture: forgetting `restore()` between two
// `_captureLogger()` calls silently abandoned the previous capture and
// corrupted logger state for the rest of the vitest worker. Throwing here
@@ -358,6 +368,7 @@ export function _captureLogger(): LoggerCapture {
}
const w = new MemoryWritable();
_activeDestination = w;
_activeLevel = level;
_cached = undefined;
return {
records: () =>
@@ -369,6 +380,7 @@ export function _captureLogger(): LoggerCapture {
text: () => w.chunks.join(''),
restore: () => {
_activeDestination = undefined;
_activeLevel = undefined;
_cached = undefined;
},
};
+50 -17
View File
@@ -30,7 +30,11 @@ import {
queryImporters,
loadFTSExtension,
} from './lbug/lbug-adapter.js';
import { createSearchFTSIndexes, verifySearchFTSIndexes } from './search/fts-indexes.js';
import {
createSearchFTSIndexes,
initialiseSearchFTSStemmer,
verifySearchFTSIndexes,
} from './search/fts-indexes.js';
import { resolveAnalyzeInstallPolicy } from './lbug/extension-loader.js';
import {
startWalCheckpointDriver,
@@ -547,6 +551,11 @@ export async function runFullAnalysis(
const progress = (phase: string, percent: number, message: string) =>
callbacks.onProgress(phase, percent, message);
// Resolve + validate operator-provided FTS config once, before the expensive
// parse/load phases. A typo fails here in ms; createSearchFTSIndexes reuses
// the cached value via getSearchFTSStemmer.
initialiseSearchFTSStemmer();
// Scope the degraded-parse log throttle to this run. On a reused process
// (e.g. tests, or any host that calls runFullAnalysis more than once) the
// module-level counter would otherwise stay saturated and suppress every
@@ -658,6 +667,22 @@ export async function runFullAnalysis(
}
try {
await initLbug(lbugPath);
// Gate on FTS availability BEFORE touching any index. createSearchFTSIndexes
// now DROPs each index before recreating it (so schema changes reach existing
// DBs); if the extension were unavailable, the drops would run and leave the
// DB index-less, only failing at the create step. Fail loudly first — mirrors
// the analyze path's `if (ftsAvailable)` gate below — so an unavailable
// extension never destroys the existing indexes.
const repairFtsAvailable = await loadFTSExtension(undefined, {
policy: resolveAnalyzeInstallPolicy(),
});
if (!repairFtsAvailable) {
throw new Error(
'Cannot repair FTS indexes: the LadybugDB FTS extension is unavailable ' +
'(not pre-installed and could not be installed on this machine). ' +
'Run `gitnexus doctor` to install it, then retry `--repair-fts`.',
);
}
progress('fts', 85, 'Repairing search indexes...');
await createSearchFTSIndexes({
onIndexStart: options.verbose
@@ -733,6 +758,30 @@ export async function runFullAnalysis(
options = { ...options, force: true };
}
// ── schema-version mismatch forces full rebuild (#2289 P1) ────────
// Mirrors the pdg-mode block above: a stamp from an older
// INCREMENTAL_SCHEMA_VERSION (e.g. pre-v5 URL-only Route ids) cannot be
// reconciled by an incremental top-up — same-commit re-analyze would
// strand stale rows next to new-schema writes. MUST sit before the
// alreadyUpToDate fast path below: an unchanged-commit clean tree would
// otherwise early-return without ever reaching the `isIncremental` gate
// that consults `schemaVersion`, defeating the bump's whole point.
//
// `schemaVersion === undefined` covers two cases that should still trip
// this guard: a non-git repo (which never stamps the field) and very old
// meta from before the field existed. Non-git repos take the
// `currentCommit === ''` rebuild branch below regardless, so the redundant
// force here is harmless; the friendlier `'pre-versioning'` log avoids a
// user-visible "stamped vundefined" line in that edge case.
if (existingMeta && existingMeta.schemaVersion !== INCREMENTAL_SCHEMA_VERSION) {
const stampedVersion = existingMeta.schemaVersion ?? 'pre-versioning';
log(
`index schema changed (stamped v${stampedVersion}, this build is v${INCREMENTAL_SCHEMA_VERSION}); ` +
`forcing a full rebuild so persisted rows match the current schema.`,
);
options = { ...options, force: true };
}
// ── Early-return: already up to date ──────────────────────────────
if (existingMeta && !options.force && existingMeta.lastCommit === currentCommit) {
// Non-git folders have currentCommit = '' — always rebuild since we can't detect changes
@@ -1320,21 +1369,6 @@ export async function runFullAnalysis(
}
}
const { readServerMapping } = await import('./embeddings/server-mapping.js');
// Mirror the registry's name-resolution chain so the server-mapping
// lookup key stays aligned with the final registry name (#1259):
// --name → remote-derived → canonical-root basename
// (preserved-alias is intentionally NOT consulted here — server
// mappings are addressed by the operationally-meaningful name the
// user configures, not by a sticky registry-only alias they may not
// know about. The previous canonical-only logic ignored both --name
// and remote-derived names, silently breaking server-mapping for
// anyone with a `--name` alias or remote-named repo.)
const projectName =
options.registryName ??
getInferredRepoName(repoPath) ??
path.basename(resolveRepoIdentityRoot(repoPath));
const serverName = await readServerMapping(projectName);
const embeddingResult = await runEmbeddingPipeline(
executeQuery,
executeWithReusedStatement,
@@ -1350,7 +1384,6 @@ export async function runFullAnalysis(
},
{},
cachedEmbeddingNodeIds.size > 0 ? cachedEmbeddingNodeIds : undefined,
{ repoName: projectName, serverName },
existingEmbeddings,
);
if (embeddingResult.semanticMode === 'exact-scan') {
+109 -19
View File
@@ -1,17 +1,100 @@
import { createFTSIndex } from '../lbug/lbug-adapter.js';
import { createFTSIndex, dropFTSIndex, DEFAULT_FTS_STEMMER } from '../lbug/lbug-adapter.js';
import { FTS_INDEXES } from './fts-schema.js';
// Stemmers shipped by the LadybugDB FTS extension. Mirrors the lowercase token
// set in the extension bundled with @ladybugdb/core 0.17.x (see package.json).
// Keep in sync on a LadybugDB minor bump — a value here that the installed
// extension rejects would pass validation but fail at CREATE_FTS_INDEX.
const SUPPORTED_FTS_STEMMERS = new Set<string>([
'arabic',
'basque',
'catalan',
'danish',
'dutch',
'english',
'finnish',
'french',
'german',
'greek',
'hindi',
'hungarian',
'indonesian',
'irish',
'italian',
'lithuanian',
'nepali',
'norwegian',
'none',
'porter',
'portuguese',
'romanian',
'russian',
'serbian',
'spanish',
'swedish',
'tamil',
'turkish',
]);
export interface CreateSearchFTSIndexesOptions {
onIndexStart?: (table: string, indexName: string) => void;
onIndexReady?: (table: string, indexName: string) => void;
}
let resolvedStemmer: string | undefined;
/** Read + validate `GITNEXUS_FTS_STEMMER`. Throws on an unsupported value. */
function resolveFTSStemmer(): string {
const raw = process.env.GITNEXUS_FTS_STEMMER?.trim().toLowerCase();
if (!raw) return DEFAULT_FTS_STEMMER;
if (SUPPORTED_FTS_STEMMERS.has(raw)) return raw;
throw new Error(
`Invalid GITNEXUS_FTS_STEMMER "${process.env.GITNEXUS_FTS_STEMMER}". ` +
`Expected one of: ${[...SUPPORTED_FTS_STEMMERS].sort().join(', ')}.`,
);
}
/**
* Resolve + validate `GITNEXUS_FTS_STEMMER` once, up front at analyze startup,
* and cache it. An invalid value throws here — in milliseconds — instead of
* ~85% into a run (after the expensive parse/scope-resolution work). The cached
* value is what {@link getSearchFTSStemmer} returns for the rest of the run, so
* config is read and validated in exactly one place.
*/
export function initialiseSearchFTSStemmer(): string {
resolvedStemmer = resolveFTSStemmer();
return resolvedStemmer;
}
/**
* Return the stemmer resolved by {@link initialiseSearchFTSStemmer}. Falls back
* to resolving on demand when init was never called (read-only hosts, unit
* tests) so validation always applies.
*/
export function getSearchFTSStemmer(): string {
return resolvedStemmer ?? resolveFTSStemmer();
}
export async function createSearchFTSIndexes(
options?: CreateSearchFTSIndexesOptions,
): Promise<void> {
const stemmer = getSearchFTSStemmer();
for (const { table, indexName, properties } of FTS_INDEXES) {
options?.onIndexStart?.(table, indexName);
await createFTSIndex(table, indexName, [...properties]);
// Drop first so the live `properties` always win. `createFTSIndex` is
// idempotent-by-name (skips when the index already exists), so without the
// drop a schema change — e.g. adding `description` (#2299) — would never
// reach an existing `.lbug` DB on an incremental re-analyze or `--repair-fts`;
// the old name+content index would silently persist. `dropFTSIndex` no-ops
// when the index is absent (first-ever analyze) and clears the per-connection
// memo so the create below actually runs.
// ponytail: this rebuilds every FTS index on every analyze instead of
// skipping when present; FTS build is proportional to symbol-table size and
// runs inside the existing FTS phase. Gate on a stored schema fingerprint if
// this rebuild cost ever shows up in analyze profiles.
await dropFTSIndex(table, indexName);
await createFTSIndex(table, indexName, [...properties], stemmer);
options?.onIndexReady?.(table, indexName);
}
}
@@ -19,25 +102,32 @@ export async function createSearchFTSIndexes(
export async function verifySearchFTSIndexes(
executeQuery: (cypher: string) => Promise<unknown[]>,
): Promise<string[]> {
const safeIdentifier = (value: string): string => {
if (!/^[A-Za-z_][A-Za-z0-9_]*$/.test(value)) {
throw new Error(`Invalid FTS identifier: ${value}`);
}
return value;
};
// Read the catalog once and check each configured index both EXISTS and
// covers its expected columns. A queryability-only probe (CALL QUERY_FTS_INDEX
// ... catch) is not enough: a stale `name+content`-only index left on a
// pre-#2299 DB stays queryable yet silently misses `description`, so the probe
// would pass while doc-comment search is still broken (#2299). SHOW_INDEXES
// exposes `property_names` (STRING[]) per index, so we assert coverage directly.
const rows = await executeQuery('CALL SHOW_INDEXES() RETURN *');
const propsByIndex = new Map<string, readonly string[]>();
for (const row of rows) {
if (typeof row !== 'object' || row === null) continue;
const record = row as Record<string, unknown>;
const indexName = record.index_name;
const propertyNames = record.property_names;
if (typeof indexName !== 'string' || !Array.isArray(propertyNames)) continue;
propsByIndex.set(
indexName,
propertyNames.filter((p): p is string => typeof p === 'string'),
);
}
const missing: string[] = [];
for (const { table, indexName } of FTS_INDEXES) {
const safeTable = safeIdentifier(table);
const safeIndex = safeIdentifier(indexName);
const probe = `
CALL QUERY_FTS_INDEX('${safeTable}', '${safeIndex}', '__gitnexus_fts_probe__', conjunctive := false)
RETURN score
LIMIT 1
`;
try {
await executeQuery(probe);
} catch {
for (const { table, indexName, properties } of FTS_INDEXES) {
const actual = propsByIndex.get(indexName);
// Absent from the catalog, or present but not covering every expected column.
if (!actual || !properties.every((p) => actual.includes(p))) {
missing.push(`${table}.${indexName}`);
}
}
+38 -4
View File
@@ -4,10 +4,44 @@ export interface FTSIndexDefinition {
readonly properties: readonly string[];
}
// Shared by both index creation (`createSearchFTSIndexes`) and querying
// (`searchFTSFromLbug` / `verifySearchFTSIndexes`) — the single source of truth
// for which tables/columns are full-text searchable. Adding `description` here
// makes doc comments (Javadoc/KDoc/JSDoc/Doxygen/godoc/RDoc) keyword-searchable
// once they are populated by `descriptionExtractor` (#2270/#2286, issue #2299).
//
// IMPORTANT: every property must be a real column on its table (see
// `core/lbug/schema.ts`). `File` has no `description` column, so it stays
// name+content. All other entries below carry a `description` column.
//
// Tables beyond the original 5 mirror `EMBEDDABLE_LABELS` (embeddings/types.ts):
// indexing the same set keeps a symbol's doc comment both keyword- and
// semantically-searchable.
const FTS_PROPERTIES = ['name', 'content', 'description'] as const;
export const FTS_INDEXES: readonly FTSIndexDefinition[] = [
// File has no `description` column — keep it name+content only.
{ table: 'File', indexName: 'file_fts', properties: ['name', 'content'] },
{ table: 'Function', indexName: 'function_fts', properties: ['name', 'content'] },
{ table: 'Class', indexName: 'class_fts', properties: ['name', 'content'] },
{ table: 'Method', indexName: 'method_fts', properties: ['name', 'content'] },
{ table: 'Interface', indexName: 'interface_fts', properties: ['name', 'content'] },
// Original 5 (minus File) gain `description`.
{ table: 'Function', indexName: 'function_fts', properties: FTS_PROPERTIES },
{ table: 'Class', indexName: 'class_fts', properties: FTS_PROPERTIES },
{ table: 'Method', indexName: 'method_fts', properties: FTS_PROPERTIES },
{ table: 'Interface', indexName: 'interface_fts', properties: FTS_PROPERTIES },
// Remaining EMBEDDABLE_LABELS symbol tables — all CODE_ELEMENT_BASE-shaped
// (or a superset), so all carry name + content + description columns.
{ table: 'Constructor', indexName: 'constructor_fts', properties: FTS_PROPERTIES },
{ table: 'Struct', indexName: 'struct_fts', properties: FTS_PROPERTIES },
{ table: 'Enum', indexName: 'enum_fts', properties: FTS_PROPERTIES },
{ table: 'Trait', indexName: 'trait_fts', properties: FTS_PROPERTIES },
{ table: 'Impl', indexName: 'impl_fts', properties: FTS_PROPERTIES },
{ table: 'Macro', indexName: 'macro_fts', properties: FTS_PROPERTIES },
{ table: 'Namespace', indexName: 'namespace_fts', properties: FTS_PROPERTIES },
{ table: 'TypeAlias', indexName: 'type_alias_fts', properties: FTS_PROPERTIES },
{ table: 'Typedef', indexName: 'typedef_fts', properties: FTS_PROPERTIES },
{ table: 'Const', indexName: 'const_fts', properties: FTS_PROPERTIES },
{ table: 'Property', indexName: 'property_fts', properties: FTS_PROPERTIES },
{ table: 'Record', indexName: 'record_fts', properties: FTS_PROPERTIES },
{ table: 'Union', indexName: 'union_fts', properties: FTS_PROPERTIES },
{ table: 'Static', indexName: 'static_fts', properties: FTS_PROPERTIES },
{ table: 'Variable', indexName: 'variable_fts', properties: FTS_PROPERTIES },
];
+362 -23
View File
@@ -39,9 +39,19 @@ import {
type RegistryEntry,
type BranchSummary,
} from '../../storage/repo-manager.js';
import { GroupService, type GroupToolPort } from '../../core/group/service.js';
import {
GroupService,
type GroupToolPort,
type GroupSymbolResolution,
type GroupPdgFlowResult,
type GroupPdgFlowHop,
} from '../../core/group/service.js';
import { resolveAtGroupMemberRepoPath } from '../../core/group/resolve-at-member.js';
import { collectBestChunks } from '../../core/embeddings/types.js';
import {
DEFAULT_MCP_VECTOR_MAX_DISTANCE,
getVectorMaxDistance,
} from '../../core/embeddings/config.js';
import {
rankExactEmbeddingRows,
type ExactEmbeddingRow,
@@ -260,10 +270,39 @@ export const IMPACT_RELATION_CONFIDENCE: Readonly<Record<string, number>> = {
const confidenceForRelType = (relType: string | undefined): number =>
IMPACT_RELATION_CONFIDENCE[relType ?? ''] ?? 0.5;
/** Structured error logging for query failures — replaces empty catch blocks */
/**
* Structured logging for *swallowed* query failures — replaces empty catch
* blocks. The level reflects telemetry severity, NOT a promise about the
* caller: most callers catch the failure and degrade to a genuinely safe
* fallback (a usable result, usually with a caller-visible `partial`/`ftsUsed`
* flag), so these are not operation-level errors and must not log at `error`:
*
* - A benign missing optional table/label/column — a repo analyzed without
* processes/communities, or a pre-v3 PDG index lacking the `calleeIds`
* column — is a normal configuration, not a failure. Logged at `debug`
* (suppressed at the default `info` level; surfaced only when troubleshooting).
* - Any other swallowed failure is an unexpected-but-handled degradation:
* logged at `warn` so it stays observable without raising a false `error`
* alarm that would drown genuine, operation-aborting failures.
*
* `error` is intentionally NOT used here — it is reserved for failures that
* actually abort an operation, which log directly rather than through this
* best-effort-degradation helper.
*
* Contract for callers (#2283 review): only route a failure here when the
* caller ALSO surfaces the degradation in its result (a `partial` flag,
* `failed_files`, `traversalComplete:false`, …). A mutating or safety-critical
* path that would otherwise report success/clean (e.g. `rename` apply, the
* `detect_changes` safety gate) MUST set that result-level signal — `warn`
* alone is not a substitute for an honest result.
*/
function logQueryError(context: string, err: unknown): void {
const msg = err instanceof Error ? err.message : String(err);
logger.error({ context, err: msg }, 'GitNexus query failed');
if (isBenignMissingTableError(err)) {
logger.debug({ context, err: msg }, 'GitNexus query skipped (missing optional data)');
return;
}
logger.warn({ context, err: msg }, 'GitNexus query failed (degraded)');
}
/**
@@ -276,7 +315,12 @@ function logQueryError(context: string, err: unknown): void {
*/
function isBenignMissingTableError(err: unknown): boolean {
const msg = err instanceof Error ? err.message : String(err ?? '');
return /does not exist|no such (table|label|rel)|unknown (table|label)|not (defined|found)/i.test(
// The `not (defined|found)` arm is scoped to a schema object (table/label/
// rel/column/property), mirroring lbug-adapter's isMissingColumnError
// (`/(table|column|property).*not found/i`): an unscoped "not found" matched
// operation failures like `rg: not found` (ripgrep absent) or `Symbol not
// found`, which this helper would then silently demote to `debug` (#2283).
return /does not exist|no such (table|label|rel)|unknown (table|label)|(table|label|rel|column|property)[^\n]*\bnot (defined|found)\b/i.test(
msg,
);
}
@@ -474,6 +518,41 @@ interface ImpactParams {
summaryOnly?: boolean;
}
/** One route in an `api_impact` result. `executionFlows` are process names. */
interface ApiImpactRoute {
route: string;
method: string | null;
handler: string;
responseShape: { success: string[]; error: string[] };
middleware: string[];
middlewareDetection?: 'partial';
middlewareNote?: string;
consumers: Array<{ name: string; file: string; accesses: string[]; attributionNote?: string }>;
mismatches?: Array<{
consumer: string;
field: string;
reason: string;
confidence: 'high' | 'low';
}>;
executionFlows: string[];
impactSummary: {
directConsumers: number;
affectedFlows: number;
riskLevel: 'LOW' | 'MEDIUM' | 'HIGH';
warning?: string;
};
}
/**
* `api_impact` is polymorphic by match count: a single matched route returns the
* route object directly; two or more return the wrapped `{ routes, total }`
* form; any guard failure returns `{ error }`.
*/
type ApiImpactResult =
| ApiImpactRoute
| { routes: ApiImpactRoute[]; total: number }
| { error: string };
/**
* One repository entry as returned by {@link LocalBackend.listRepos} and in each
* `list_repos` page. Named so the `listRepos`/`listReposPage` return types read
@@ -596,12 +675,164 @@ export class LocalBackend {
query: (r, p) => this.query(r as RepoHandle, p),
impactByUid: (id, uid, d, o) => this.impactByUid(id, uid, d, o),
context: (r, p) => this.context(r as RepoHandle, p),
trace: (r, p) => this.trace(r as RepoHandle, p),
resolveSymbol: (r, q) => this.resolveSymbolForGroup(r as RepoHandle, q),
pdgFlows: (r, anchor, opts) => this.pdgFlowsForGroup(r as RepoHandle, anchor, opts),
};
this.groupToolSvc = new GroupService(port);
}
return this.groupToolSvc;
}
/**
* Adapt the shared symbol resolver to the GroupToolPort contract. Used by the
* cross-repo trace path to locate which member repo an endpoint lives in and
* recover its node id (== bridge `Contract.symbolUid`).
*/
private async resolveSymbolForGroup(
repo: RepoHandle,
query: { name?: string; uid?: string; file_path?: string },
): Promise<GroupSymbolResolution> {
await this.ensureInitialized(repo);
const outcome = await this.resolveSymbolCandidates(
repo,
{ uid: query.uid, name: query.name },
{ file_path: query.file_path },
);
if (outcome.kind === 'ok') {
const s = outcome.symbol;
return {
kind: 'ok',
symbol: {
id: s.id,
name: s.name,
type: s.type,
filePath: s.filePath,
startLine: s.startLine,
endLine: s.endLine,
},
};
}
if (outcome.kind === 'ambiguous') {
return {
kind: 'ambiguous',
candidates: outcome.candidates.map((c) => ({
id: c.id,
name: c.name,
type: c.type,
filePath: c.filePath,
startLine: c.startLine,
})),
};
}
return { kind: 'not_found' };
}
/**
* Intra-procedural REACHING_DEF data-flow for a single anchor symbol, adapted
* to the GroupToolPort contract. Reuses the same anchor + `flows` query as the
* `pdg_query` tool. `available:false` (not an error) when the repo has no PDG
* `flows` layer, so the cross-repo trace degrades to call-level hops.
*/
private async pdgFlowsForGroup(
repo: RepoHandle,
anchor: { name?: string; uid?: string; file_path?: string },
opts: { limit?: number },
): Promise<GroupPdgFlowResult> {
try {
await this.ensureInitialized(repo);
return await this._pdgFlowsForGroupImpl(repo, anchor, opts);
} catch {
// Enrichment is auxiliary — never let a PDG query failure fail the trace.
return { available: false, hops: [] };
}
}
/**
* Intra-procedural REACHING_DEF data-flow within the anchor symbol's block
* span. Reuses the same anchored, bind-param-only `flows` query as
* `pdg_query` (no rel-property index ⇒ the BasicBlock id-prefix + line-span
* anchor IS the bound). The anchor is resolved by UID when available (the
* boundary symbol is known precisely), avoiding the name-ambiguity the
* by-name `resolveBlockAnchor` path can hit. Data flow never crosses the repo
* boundary — this only describes how values move toward the boundary call
* inside one function.
*/
private async _pdgFlowsForGroupImpl(
repo: RepoHandle,
anchor: { name?: string; uid?: string; file_path?: string },
opts: { limit?: number },
): Promise<GroupPdgFlowResult> {
const rawLimit = opts.limit ?? PDG_QUERY_DEFAULT_LIMIT;
const limit =
Number.isInteger(rawLimit) && rawLimit >= 1 && rawLimit <= PDG_QUERY_MAX_LIMIT
? rawLimit
: PDG_QUERY_DEFAULT_LIMIT;
// Meta probe: layer present iff the flows cap is stamped. `false` is a
// definitive absence (degrade to call-level); `undefined` is unreadable
// meta (fall through and infer presence from rows found).
const pdgStamped = await pdgStampForMode(repo.lbugPath, 'flows');
if (pdgStamped === false) return { available: false, hops: [] };
// Resolve the anchor symbol (UID is precise; fall back to name/file).
const resolved = await this.resolveSymbolCandidates(
repo,
{ uid: anchor.uid, name: anchor.name },
{ file_path: anchor.file_path },
);
if (resolved.kind !== 'ok') {
// Layer may exist but we couldn't anchor — report availability from the
// stamp so the caller's note reflects the layer, not the miss.
return { available: pdgStamped === true, hops: [] };
}
const sym = resolved.symbol;
// Same span-anchored clause as resolveBlockAnchor's symbol branch: the
// BasicBlock startLine is 1-based vs the 0-based symbol span, so shift both
// bounds +1. `idPrefix`/`symStart`/`symEnd` are bind params; the edge type
// is a hardcoded literal — no user string is ever interpolated.
const hasSpan =
typeof sym.startLine === 'number' &&
typeof sym.endLine === 'number' &&
sym.endLine >= sym.startLine;
const idPrefix = `BasicBlock:${sym.filePath}:`;
const anchorClause = hasSpan
? 'a.id STARTS WITH $idPrefix AND a.startLine >= $symStart AND a.startLine <= $symEnd'
: 'a.id STARTS WITH $idPrefix';
const queryParams: Record<string, unknown> = hasSpan
? { idPrefix, symStart: sym.startLine + 1, symEnd: sym.endLine + 1 }
: { idPrefix };
const rows = await executeParameterized(
repo.lbugPath,
`MATCH (a:BasicBlock)-[r:CodeRelation]->(b:BasicBlock)
WHERE r.type = 'REACHING_DEF' AND ${anchorClause}
RETURN a.startLine AS defLine, b.startLine AS useLine, b.text AS useText, r.reason AS reason
ORDER BY useLine, defLine, reason
LIMIT ${limit + 1}`,
queryParams,
);
const truncated = rows.length > limit;
const capped = truncated ? rows.slice(0, limit) : rows;
const hops: GroupPdgFlowHop[] = capped.map((r: Record<string, unknown>) => ({
// Number()/String() coerce the LadybugDB object/tuple cell; a bare
// `as number` cast on a nullish cell would surface NaN downstream.
line: Number(r.useLine ?? r[1] ?? 0),
text: String(r.useText ?? r[2] ?? '').trim(),
variable: decodeReachingDefReason(String(r.reason ?? r[3] ?? '')).name || undefined,
}));
const available = pdgStamped === true || hops.length > 0;
return {
available,
...(hops[0]?.variable ? { variable: hops[0].variable } : {}),
hops,
...(truncated ? { truncated: true } : {}),
};
}
/** Close all pooled LadybugDB connections (CLI one-shot; optional for long-lived MCP). */
async dispose(): Promise<void> {
await closeLbug();
@@ -1323,7 +1554,7 @@ export class LocalBackend {
// — third-party MCP clients may legitimately send "query", so the alias is not slated
// for removal even if Claude Code's argument handling later changes.
if (
(method === 'impact' || method === 'query' || method === 'context') &&
(method === 'impact' || method === 'query' || method === 'context' || method === 'trace') &&
typeof p.repo === 'string' &&
p.repo.startsWith('@')
) {
@@ -1797,9 +2028,14 @@ export class LocalBackend {
try {
ftsResponse = await searchFTSFromLbug(query, limit, repo.lbugPath);
} catch (err: any) {
logger.error(
// Swallowed, gracefully-degraded failure: the search falls back to
// semantic-only (a valid result), and the most common cause is simply an
// un-indexed FTS extension — a normal configuration, not an operation
// error. Logged at warn (matching the sibling import-failure fallback
// above), never error, so it does not raise a false alarm.
logger.warn(
{ err: err.message },
'GitNexus: BM25/FTS search failed (FTS indexes may not exist) -',
'GitNexus: BM25/FTS search failed (FTS indexes may not exist) — falling back to semantic-only',
);
return { results: [], ftsUsed: false };
}
@@ -1895,6 +2131,7 @@ export class LocalBackend {
const queryVec = await embedQuery(query);
const dims = getEmbeddingDims();
const queryVecStr = `[${queryVec.join(',')}]`;
const maxDistance = getVectorMaxDistance(DEFAULT_MCP_VECTOR_MAX_DISTANCE);
let bestChunks = new Map<
string,
@@ -1908,7 +2145,7 @@ export class LocalBackend {
CAST(${queryVecStr} AS FLOAT[${dims}]), ${fetchLimit})
YIELD node AS emb, distance
WITH emb, distance
WHERE distance < 0.6
WHERE distance < ${maxDistance}
RETURN emb.nodeId AS nodeId, emb.chunkIndex AS chunkIndex,
emb.startLine AS startLine, emb.endLine AS endLine, distance
ORDER BY distance
@@ -1958,7 +2195,7 @@ export class LocalBackend {
embedding: row.embedding ?? row[4] ?? [],
}));
bestChunks = new Map(
rankExactEmbeddingRows(exactRows, queryVec, limit, 0.6).map((row) => [
rankExactEmbeddingRows(exactRows, queryVec, limit, maxDistance).map((row) => [
row.nodeId,
{
distance: row.distance,
@@ -3667,6 +3904,9 @@ export class LocalBackend {
// Map diff hunks to indexed symbols via range overlap
const changedSymbols: any[] = [];
// Set if a swallowed graph query fails below — surfaces `partial:true` so a
// degraded run cannot report a false-clean `risk_level:'low'` (#2283).
let queryDegraded = false;
for (const fileDiff of fileDiffs) {
if (fileDiff.hunks.length === 0) continue;
@@ -3714,6 +3954,12 @@ export class LocalBackend {
}
} catch (e) {
logQueryError('detect-changes:file-symbols', e);
// The symbol query failed: changedSymbols stays empty and the result
// would otherwise look like a clean no-op (`changed_count:0`,
// `risk_level:'low'`). detect_changes is the pre-commit safety gate, so
// flag the result `partial` rather than let a swallowed failure
// masquerade as "nothing changed" (#2283).
queryDegraded = true;
}
}
@@ -3752,6 +3998,7 @@ export class LocalBackend {
}
} catch (e) {
logQueryError('detect-changes:process-lookup', e);
queryDegraded = true;
}
}
@@ -3774,6 +4021,9 @@ export class LocalBackend {
},
changed_symbols: changedSymbols,
affected_processes: Array.from(affectedProcesses.values()),
// A swallowed query failure makes the counts/risk above incomplete — tell
// the caller so the safety gate isn't trusted as a clean result (#2283).
...(queryDegraded && { partial: true }),
};
}
@@ -3974,6 +4224,7 @@ export class LocalBackend {
const allChanges = Array.from(changes.values());
const totalEdits = allChanges.reduce((sum, c) => sum + c.edits.length, 0);
const failedFiles: string[] = [];
if (!dry_run) {
// Apply edits to files
for (const change of allChanges) {
@@ -3984,13 +4235,17 @@ export class LocalBackend {
content = content.replace(regex, new_name);
await fs.writeFile(fullPath, content, 'utf-8');
} catch (e) {
// A swallowed write failure must not be reported as a full success
// (#2283): record the file so the result can degrade to 'partial'
// with the unwritten files listed, rather than masquerading as done.
logQueryError('rename:apply-edit', e);
failedFiles.push(change.file_path);
}
}
}
return {
status: 'success',
status: failedFiles.length > 0 ? 'partial' : 'success',
old_name: oldName,
new_name,
files_affected: allChanges.length,
@@ -3999,6 +4254,7 @@ export class LocalBackend {
text_search_edits: astSearchEdits,
changes: allChanges,
applied: !dry_run,
...(failedFiles.length > 0 && { failed_files: failedFiles }),
};
}
@@ -4039,6 +4295,22 @@ export class LocalBackend {
};
}
// A single-repo trace needs a target. Omitting `to` is the destination-trace
// shorthand, but that only exists for a cross-repo @group trace — reject a
// to-less single-repo call with an actionable error rather than the opaque
// "Target symbol 'undefined' not found".
const hasTo =
(typeof params.to === 'string' && params.to.trim() !== '') ||
(typeof params.to_uid === 'string' && params.to_uid.trim() !== '');
if (!hasTo) {
return {
status: 'error',
error: 'trace requires `to` (or `to_uid`) for a single-repo trace.',
suggestion:
'Pass a target symbol, or use repo:"@<group>" and omit `to` to trace `from` to its HTTP destination.',
};
}
const fromOutcome = await this.resolveSymbolCandidates(
repo,
{ uid: params.from_uid, name: params.from },
@@ -4316,9 +4588,21 @@ export class LocalBackend {
}
const mode = modeResult.mode;
// #2279: some MCP client/agent adapters serialize an *omitted* optional
// numeric field as `0` rather than dropping it, so callgraph calls arrive
// carrying a spurious `line: 0`. `line` is meaningless on the callgraph path
// (the symbol→symbol BFS has no statement notion), so treat a literal `0`
// there as omitted and let the normal traversal run. The coercion is
// deliberately narrow — only the literal `0`, only when mode !== 'pdg':
// a genuine positive `line` on callgraph still errors (real mode mistake),
// negative/fractional values still error, and pdg mode is untouched (the
// normalization is an identity there, so `line: 0` is still rejected below —
// there is no 1-based source line `0` to anchor on).
const effectiveLine = mode !== 'pdg' && params.line === 0 ? undefined : params.line;
// `line` is a PDG-only statement anchor. Reject it on the callgraph path
// rather than silently ignore (the symbol→symbol BFS has no statement notion).
if (params.line !== undefined && mode !== 'pdg') {
if (effectiveLine !== undefined && mode !== 'pdg') {
return {
error: `Parameter 'line' is only supported with mode:'pdg' (it anchors the dependence slice on a statement). Remove it or set mode:'pdg'.`,
target: { name: params.target },
@@ -4329,8 +4613,8 @@ export class LocalBackend {
}
// A provided `line` must be a positive integer.
if (
params.line !== undefined &&
(!Number.isInteger(params.line) || (params.line as number) < 1)
effectiveLine !== undefined &&
(!Number.isInteger(effectiveLine) || (effectiveLine as number) < 1)
) {
// Line param fails validation before target resolution → partial-but-typed
// target on the pdg path (typed PdgImpactTarget, not an inline literal).
@@ -4666,7 +4950,11 @@ export class LocalBackend {
symType,
direction,
maxDepth,
line: params.line,
// Use the normalized line, not raw params.line, so the gate and the
// engine share one source of truth (#2283). Identity in pdg mode today
// — effectiveLine === params.line when mode === 'pdg' — but this stays
// correct if the normalization ever stops being an identity here.
line: effectiveLine,
limit: Number.isFinite(params.limit) ? params.limit : 100,
// KTD2 extraction-seam discipline: hand the engine its DB dependency
// explicitly rather than `this.`-binding it. LocalBackend owns repo
@@ -5782,6 +6070,26 @@ export class LocalBackend {
if (resolved.ok === false) return { error: resolved.error };
const svc = this.getGroupService();
if (method === 'trace') {
// Cross-repo trace resolves `from`/`to` across ALL members (it does not
// anchor on a single member like impact/query/context), so the member
// path in `@group/path` is advisory here — `resolved` above still
// validates that the group exists. groupTrace owns cross-member
// resolution and the single-boundary bridge crossing.
const traceArgs: Record<string, unknown> = { name: groupName };
if (params.from !== undefined) traceArgs.from = params.from;
if (params.to !== undefined) traceArgs.to = params.to;
if (params.from_uid !== undefined) traceArgs.from_uid = params.from_uid;
if (params.to_uid !== undefined) traceArgs.to_uid = params.to_uid;
if (params.from_file !== undefined) traceArgs.from_file = params.from_file;
if (params.to_file !== undefined) traceArgs.to_file = params.to_file;
if (params.maxDepth !== undefined) traceArgs.maxDepth = params.maxDepth;
if (params.crossDepth !== undefined) traceArgs.crossDepth = params.crossDepth;
if (params.includeTests !== undefined) traceArgs.includeTests = params.includeTests;
if (params.pdg !== undefined) traceArgs.pdg = params.pdg;
if (params.limit !== undefined) traceArgs.limit = params.limit;
return svc.groupTrace(traceArgs);
}
if (method === 'impact') {
// KTD5/KTD12 — validate `mode` at the group-forward boundary too (the
// JSON-schema enum is advisory). An invalid mode errors; `mode:'pdg'` is
@@ -5938,6 +6246,7 @@ export class LocalBackend {
Array<{
id: string;
name: string;
method: string | null;
filePath: string;
responseKeys: string[] | null;
errorKeys: string[] | null;
@@ -5960,7 +6269,7 @@ export class LocalBackend {
RETURN n.id AS routeId, n.name AS routeName, n.filePath AS handlerFile,
n.responseKeys AS responseKeys, n.errorKeys AS errorKeys, n.middleware AS middleware,
consumer.name AS consumerName, consumer.filePath AS consumerFile,
r.reason AS fetchReason
r.reason AS fetchReason, n.method AS method
`,
params,
);
@@ -5975,6 +6284,7 @@ export class LocalBackend {
{
id: string;
name: string;
method: string | null;
filePath: string;
responseKeys: string[] | null;
errorKeys: string[] | null;
@@ -5997,11 +6307,17 @@ export class LocalBackend {
const consumerName = row.consumerName ?? row[6];
const consumerFile = row.consumerFile ?? row[7];
const fetchReason: string | null = row.fetchReason ?? row[8] ?? null;
// Verb is the literal '*' for method-agnostic routes (Django function
// views) and absent (null) for method-less routes (filesystem, Laravel
// resource). Appended last in RETURN so positional fallbacks for the
// consumer/reason columns above stay stable.
const method: string | null = row.method ?? row[9] ?? null;
if (!routeMap.has(id)) {
routeMap.set(id, {
id,
name,
method,
filePath,
responseKeys,
errorKeys,
@@ -6099,6 +6415,7 @@ export class LocalBackend {
return {
routes: routes.map((r) => ({
route: r.name,
method: r.method,
handler: r.filePath,
middleware: r.middleware || [],
consumers: r.consumers,
@@ -6166,6 +6483,7 @@ export class LocalBackend {
return {
route: r.name,
method: r.method,
handler: r.filePath,
...(responseKeys.length > 0 ? { responseKeys } : {}),
...(errorKeys.length > 0 ? { errorKeys } : {}),
@@ -6233,8 +6551,8 @@ export class LocalBackend {
private async apiImpact(
repo: RepoHandle,
params: { route?: string; file?: string },
): Promise<any> {
params: { route?: string; file?: string; method?: unknown },
): Promise<ApiImpactResult> {
await this.ensureInitialized(repo);
if (!params.route && !params.file) {
@@ -6253,11 +6571,30 @@ export class LocalBackend {
queryParams.file = params.file;
}
const routes = await this.fetchRoutesWithConsumers(repo.lbugPath, routeFilter, queryParams);
// After #2302 the same URL/handler can expose one Route node per HTTP verb.
// An optional `method` narrows to that one verb so the response collapses to
// the singular shape. A method-agnostic route (method `'*'`, e.g. a Django
// function view) matches any selector; verbless routes (null method) never do.
// `method` arrives unvalidated from the MCP envelope (the JSON schema is
// advisory), so reject a non-string verb with a structured error instead of
// throwing on `.toUpperCase()`; empty/whitespace collapses to no selector.
const rawMethod = params.method;
if (rawMethod !== undefined && typeof rawMethod !== 'string') {
return { error: '"method" must be a string (e.g. "GET", "POST").' };
}
const wantedMethod =
typeof rawMethod === 'string' ? rawMethod.trim().toUpperCase() || undefined : undefined;
const matched = await this.fetchRoutesWithConsumers(repo.lbugPath, routeFilter, queryParams);
const routes = matched.filter(
(r) => !wantedMethod || r.method === '*' || r.method?.toUpperCase() === wantedMethod,
);
if (routes.length === 0) {
const target = params.route || params.file;
return { error: `No routes found matching "${target}".` };
// Only append the verb when the URL/file matched routes but none used it;
// a non-existent URL/file gets the plain "no routes found" message.
const verb = wantedMethod && matched.length > 0 ? ` with method "${wantedMethod}"` : '';
return { error: `No routes found matching "${target}"${verb}.` };
}
const flowMap = await this.fetchLinkedFlowsBatch(
@@ -6265,15 +6602,16 @@ export class LocalBackend {
routes.map((r) => r.id),
);
// Count how many routes share the same handler file (for middleware partial detection)
// Count verbs per handler from the FULL match (before the method filter) so a
// method-scoped query still flags a multi-verb handler's partial middleware.
const routeCountByHandler = new Map<string, number>();
for (const r of routes) {
for (const r of matched) {
if (r.filePath) {
routeCountByHandler.set(r.filePath, (routeCountByHandler.get(r.filePath) ?? 0) + 1);
}
}
const results = routes.map((r) => {
const results: ApiImpactRoute[] = routes.map((r) => {
// Keys already normalized by fetchRoutesWithConsumers (quotes stripped)
const responseKeys = r.responseKeys ?? [];
const errorKeys = r.errorKeys ?? [];
@@ -6346,6 +6684,7 @@ export class LocalBackend {
return {
route: r.name,
method: r.method,
handler: r.filePath,
responseShape: {
success: responseKeys,
@@ -6356,7 +6695,7 @@ export class LocalBackend {
? {
middlewareDetection: 'partial' as const,
middlewareNote:
'Middleware captured from first HTTP method export only — other methods in this handler may use different middleware chains.',
'Middleware captured from the first route export only — other route exports in this handler may use different middleware chains.',
}
: {}),
consumers,
+58 -8
View File
@@ -27,6 +27,12 @@ export interface ToolDefinition {
}
>;
required: string[];
/**
* JSON-Schema `anyOf` for cross-property constraints `required` cannot express
* — e.g. "at least one of route/file". Forwarded verbatim to clients by the
* server's ListTools handler, so MCP clients see the constraint.
*/
anyOf?: Array<{ required: string[] }>;
};
}
@@ -469,9 +475,14 @@ SERVICE: optional monorepo path prefix (case-sensitive path segments). When "rep
},
line: {
type: 'integer',
minimum: 1,
// `minimum: 0` (not 1) so strict client/agent adapters that materialize
// an omitted optional numeric field as `0` do not reject the request
// before sending (#2279). A positive line is still required for a real
// pdg anchor — the backend enforces that — but `0`/omitted means "no
// statement anchor" and is tolerated on the callgraph path.
minimum: 0,
description:
"1-based source line — PDG statement anchor (mode:'pdg'). Seeds affectedStatements on the statement at this line; inter-procedural symbols are still returned in interproceduralByDepth/pdgInterprocedural and the compatibility byDepth bucket.",
"1-based source line — PDG statement anchor (mode:'pdg'). Seeds affectedStatements on the statement at this line; inter-procedural symbols are still returned in interproceduralByDepth/pdgInterprocedural and the compatibility byDepth bucket. Omit line for whole-symbol pdg (whole-symbol reach + diagnostics); a positive line anchors a statement slice. Literal 0 is tolerated only as an omitted-line compatibility sentinel on the callgraph path and is rejected for mode:'pdg'.",
},
file_path: {
type: 'string',
@@ -670,7 +681,7 @@ CONTRACT CAVEATS:
WHEN TO USE: Understanding API consumption patterns, finding orphaned routes. For pre-change analysis, prefer \`api_impact\` which combines this data with mismatch detection and risk assessment.
AFTER THIS: Use impact() on specific route handlers to see full blast radius.
Returns: route nodes with their handlers, middleware wrapper chains (e.g., withAuth, withRateLimit), and consumers.`,
Returns: route nodes with their handlers, middleware wrapper chains (e.g., withAuth, withRateLimit), and consumers. Each route object includes its "method" (the HTTP verb, "*" for method-agnostic routes, or null for method-less routes).`,
annotations: READ_ONLY_TOOL_ANNOTATIONS,
inputSchema: {
type: 'object',
@@ -711,7 +722,7 @@ Returns: tool nodes with their handler files and descriptions.`,
WHEN TO USE: Detecting mismatches between what an API route returns and what consumers expect. Finding shape drift. For pre-change analysis, prefer \`api_impact\` which combines this data with mismatch detection and risk assessment.
REQUIRES: Route nodes with responseKeys (extracted from .json({...}) calls during indexing).
Returns routes that have both detected response keys AND consumers. Shows top-level keys each endpoint returns (e.g., data, pagination, error) and what keys each consumer accesses. Reports MISMATCH status when a consumer accesses keys not present in the route's response shape.`,
Returns routes that have both detected response keys AND consumers. Shows top-level keys each endpoint returns (e.g., data, pagination, error) and what keys each consumer accesses. Reports MISMATCH status when a consumer accesses keys not present in the route's response shape. Each route object includes its "method" (the HTTP verb, "*" for method-agnostic routes, or null for method-less routes).`,
annotations: READ_ONLY_TOOL_ANNOTATIONS,
inputSchema: {
type: 'object',
@@ -736,16 +747,24 @@ WHEN TO USE: BEFORE modifying any API route handler. Shows what consumers depend
Risk levels: LOW (0-3 consumers), MEDIUM (4-9 or any mismatches), HIGH (10+ consumers or mismatches with 4+ consumers). Mismatches with confidence "low" indicate the consumer file fetches multiple routes — property attribution is approximate.
Returns: single route object when one match, or { routes: [...], total: N } for multiple matches. Combines route_map, shape_check, and impact data.`,
Response shape is keyed on how many routes match, not on the data: exactly one match returns a single route object; two or more return { routes: [...], total: N }. The same URL can expose multiple HTTP verbs (e.g. GET and POST /api/orders are distinct routes that share the URL), so a bare-URL lookup may return the wrapped form — every route object carries its own "method" so verbs are distinguishable. Pass "method" to narrow to one verb; the single-object shape is returned only when exactly one route remains after filtering — a substring route/file match spanning several URLs can still return the wrapped form. A URL/file that exists but has no route for the given verb returns an error. Each route's "method" is the literal "*" for method-agnostic routes (e.g. Django function views), which match any "method" selector, or null for method-less routes (filesystem, Laravel resource), which never match a selector. Combines route_map, shape_check, and impact data.`,
annotations: READ_ONLY_TOOL_ANNOTATIONS,
inputSchema: {
type: 'object',
properties: {
route: { type: 'string', description: 'Route path (e.g., "/api/grants")' },
file: { type: 'string', description: 'Handler file path (alternative to route)' },
method: {
type: 'string',
description:
'Optional HTTP verb — GET, POST, PUT, PATCH, DELETE, etc. — to narrow a multi-verb route or file lookup to a single method. Returns an error if no matched route uses that verb.',
},
repo: { type: 'string', description: 'Repository name or path.' },
},
required: [],
// Exactly one lookup key is needed, but either works (route wins if both
// are passed) — so the structural constraint is "at least one of route/file".
anyOf: [{ required: ['route'] }, { required: ['file'] }],
},
},
{
@@ -791,7 +810,11 @@ WHEN TO USE: Debugging "how does A reach B?" — answers in one call what would
Traverses CALLS edges plus HAS_METHOD (class → member) edges, so a trace can descend from a class into its methods. Each hop's edge type is reported in edges[], so call hops and containment hops remain distinguishable.
Returns: ordered hops with file:line, and an aligned edges[] of edge type + confidence. When no path exists, reports the furthest reachable node so you know where the chain breaks (and truncated: true if a traversal cap was hit first).`,
Returns: ordered hops with file:line, and an aligned edges[] of edge type + confidence. When no path exists, reports the furthest reachable node so you know where the chain breaks (and truncated: true if a traversal cap was hit first).
CROSS-REPO (experimental): pass repo as "@groupName" to trace across repositories in a group. When from/to live in different member repos, the trace stitches the two repo-local segments across a single ContractLink boundary (e.g. an HTTP consumer→provider link), clamped to one crossing. The result adds crossings[] (the bridged contract with matchType/confidence), tags each hop with its member repo, and a notes[] channel for degraded states. The boundary hop is reported with edge type CONTRACT_LINK. Pass pdg:true to also attach the intra-procedural data-flow (REACHING_DEF) for boundary-adjacent segments when those repos were indexed with --pdg; absent a PDG layer it degrades to call-level hops with a note.
DESTINATION TRACE (cross-repo): for an "@groupName" trace, OMIT to/to_uid/to_file to trace 'from' to wherever its outgoing HTTP call lands. The result ends at the provider endpoint (reported by route + file even when the handler is an anonymous function with no nameable symbol). This is the way to follow a client call to a backend handler you cannot name.`,
annotations: READ_ONLY_TOOL_ANNOTATIONS,
inputSchema: {
type: 'object',
@@ -799,7 +822,11 @@ Returns: ordered hops with file:line, and an aligned edges[] of edge type + conf
from: { type: 'string', description: 'Source symbol name' },
from_uid: { type: 'string', description: 'Source symbol UID (zero-ambiguity)' },
from_file: { type: 'string', description: 'Source file path hint for disambiguation' },
to: { type: 'string', description: 'Target symbol name' },
to: {
type: 'string',
description:
"Target symbol name. Omit (with to_uid/to_file) on an @group trace to trace 'from' to its HTTP destination.",
},
to_uid: { type: 'string', description: 'Target symbol UID (zero-ambiguity)' },
to_file: { type: 'string', description: 'Target file path hint for disambiguation' },
maxDepth: {
@@ -814,9 +841,32 @@ Returns: ordered hops with file:line, and an aligned edges[] of edge type + conf
description: 'Include test-file symbols in traversal (default: false)',
default: false,
},
pdg: {
type: 'boolean',
description:
'Cross-repo only (experimental): attach intra-procedural REACHING_DEF data-flow for boundary-adjacent segments when the repo has a --pdg layer. Default false.',
default: false,
},
crossDepth: {
type: 'number',
description:
'Cross-repo only: number of ContractLink boundaries to cross. Only 1 is supported today (multi-hop deferred); a direct caller that passes a higher value gets it clamped to 1 with a notes[] entry.',
default: 1,
minimum: 1,
maximum: 1,
},
limit: {
type: 'number',
description:
'Cross-repo + pdg:true only: max REACHING_DEF data-flow hops attached per boundary-adjacent segment (default 50, max 200). When a segment dataFlow is truncated, re-issue with a higher limit.',
default: 50,
minimum: 1,
maximum: 200,
},
repo: {
type: 'string',
description: 'Repository name or path. Omit if only one repo is indexed.',
description:
'Repository name or path, or "@groupName" / "@groupName/memberPath" for a cross-repo trace over a group. Omit if only one repo is indexed.',
},
},
required: [],
-1
View File
@@ -1767,7 +1767,6 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
},
{}, // config: use defaults
undefined, // skipNodeIds
undefined, // context
existingEmbeddings,
);
+6
View File
@@ -103,6 +103,12 @@ export const getRemoteUrl = (repoPath: string): string | undefined => {
* Find the git repository root from any path inside the repo
*/
export const getGitRoot = (fromPath: string): string | null => {
const resolved = path.resolve(fromPath);
// Avoid git rev-parse --show-toplevel trimming trailing spaces from the
// repository root on Windows; callers that need identity keys canonicalize
// this value with realpath before comparing it.
if (hasGitDir(resolved)) return resolved;
try {
const raw = chompGitOutput(
execSync('git rev-parse --show-toplevel', {

Some files were not shown because too many files have changed in this diff Show More