Compare commits

...
Author SHA1 Message Date
Gergo Magyar 3768e3bfd0 fix(server): log skip-embedding count and table-not-found swallow path
Addresses review feedback on PR #823:
- Log count of already-embedded nodes when skipNodeIds is populated
  (aids debugging if Kuzu driver row shape changes).
- Log when the 'table does not exist' swallow path fires so ops can
  catch it if Kuzu ever changes error wording.
- Document the {} config positional argument with an inline comment
  referencing the runEmbeddingPipeline signature.
2026-04-15 07:40:58 +01:00
Gergo Magyar 3384575ac6 style: prettier format gitnexus/src/server/api.ts 2026-04-15 07:38:16 +01:00
jonasvanderhaegen-xve dd194d56b1 fix(server): narrow catch to table-not-exist errors only in POST /api/embed
Bare catch{} would silently swallow connection errors and proceed to
re-embed all nodes, hiding infrastructure issues. Now only swallows
errors where the CodeEmbedding table does not yet exist.
2026-04-14 14:27:03 +02:00
jonasvanderhaegen-xve 80a6fde2ba fix(server): skip already-embedded nodes in POST /api/embed to avoid vector-index SET error
Kuzu/LadybugDB forbids SET on a property that is part of a vector index.
The /api/embed endpoint was calling runEmbeddingPipeline without skipNodeIds,
causing it to attempt MERGE+SET on every node including those already embedded.

Fix: query existing CodeEmbedding nodeIds before running the pipeline and pass
them as skipNodeIds so only new (unembedded) nodes are processed.
2026-04-14 13:33:48 +02:00
jonasvanderhaegen-xve 8d38cc99fa fix(embeddings): use MERGE instead of CREATE for CodeEmbedding inserts
CREATE fails with duplicate PK when a CodeEmbedding node already exists,
which happens when:
- A PostToolUse hook triggers a concurrent gitnexus analyze during an
  active analyze run (git commits fire the hook)
- A partial prior run left some embeddings in the DB before a crash

Switching to MERGE makes the insert idempotent: existing embeddings are
updated in place, new ones are created, no PK violations.

Fixes: #822
2026-04-14 12:59:14 +02:00
jonasvanderhaegen-xve 41844edf88 fix(csv-generator): deduplicate all node types, not just File nodes
The pipeline can produce duplicate node IDs across all symbol types
(Class, Method, Function, etc.). Only File nodes were guarded by a
seenFileIds Set, leaving every other type unprotected. When the CSV
was COPY'd into LadybugDB, duplicate PKs caused mass "Batch execution
error: Found duplicated primary key value" warnings on gitnexus serve.

Replace the per-type seenFileIds with a single seenNodeIds Set checked
at the top of the iteration loop, before the switch, so every label is
covered by the same O(1) deduplication guard.

Fixes: #822
2026-04-14 12:32:08 +02:00
Filipe Oliveira (Redis)andGergo Magyar 9ad1984b17 fix: resolve C/C++ cross-file calls through transitive #include chains (#816)
* fix: resolve C/C++ cross-file calls through transitive #include chains

In C/C++, #include is transitive: if a.c includes b.h and b.h includes
c.h, then a.c can call any function declared in c.h. The wildcard import
synthesis only walked direct imports (1 hop), missing symbols reachable
through transitive header chains.

This is the dominant pattern in large C codebases — Redis's db.c includes
server.h which includes dict.h, so db.c should resolve calls to dictFind()
declared in dict.h and defined in dict.c. Before this fix, those cross-file
call edges were missing entirely.

The fix expands the import closure transitively for C/C++ files before
synthesizing wildcard bindings. A BFS walks ctx.importMap and graphImports
to collect all transitively reachable headers, then passes the full closure
to synthesizeForFile.

Tested on Redis (github.com/redis/redis):
- Before: dictFetchValue had 0 cross-file callers, processCommand had 0
- After: dictFetchValue has 9 callers, processCommand has 1, +1946 edges total

Fixes #813

* refactor(ingestion): dispatch wildcard synthesis by import-semantics strategy

Generalize PR #816's C/C++ transitive #include fix into a language-agnostic
strategy pattern. The `wildcard-synthesis.ts` pipeline phase no longer
references `SupportedLanguages.C` / `SupportedLanguages.CPlusPlus` — it
dispatches on `provider.importSemantics` via an exhaustive `switch`.

Also fixes a correctness bug the original BFS introduced: `queue.pop()`
(LIFO/DFS) reversed the iteration order of `#include` directives, which —
combined with first-seen-wins dedup in `synthesizeForFile` — silently
bound overloaded symbols to the wrong header. For the `cpp-calls`
fixture, `write_audit("hello")` was being resolved to `zero.h`'s arity-0
overload instead of `one.h`'s arity-1 overload, breaking arity
narrowing. Switched to FIFO (`queue.shift()`) with direct imports seeded
in declaration order.

Taxonomy (researched across 20+ languages + stack-graphs / SCIP prior art):

  | Tag                 | Traversal       | Languages                          |
  |---------------------|-----------------|------------------------------------|
  | named               | none            | TS, JS, Java, C#, Rust, PHP, Kotlin|
  | wildcard-transitive | BFS closure     | C, C++                             |
  | wildcard-leaf       | single hop      | Go, Ruby, Swift, Dart              |
  | namespace           | none at import  | Python                             |
  | explicit-reexport   | topological DAG | (scaffold; TS `export *` future)   |

Changes:
- Widen `ImportSemantics` union from 3 to 5 tags with full taxonomy JSDoc
- Retag 5 providers: c-cpp (x2) → wildcard-transitive; dart, go, ruby,
  swift → wildcard-leaf
- Move BFS closure into `wildcard-synthesis.ts` as `expandTransitiveIncludeClosure`
  (pipeline-owned; providers stay pure declarations)
- Replace `if (lang === C || CPP)` with `dispatchSynthesis` helper called
  by both Loop 1 (ctx.importMap) and Loop 2 (graphImports) so a future
  transitive language whose edges arrive via graphImports gets closure
  expansion consistently
- `never`-assertion default arm forces compile-time exhaustiveness
- `explicit-reexport` arm falls through to leaf behavior (scaffold;
  TODO: implement re-export DAG walk for TS `export *` / Rust `pub use`)
- New unit tests covering circular includes, deep chains, diamond dedup,
  graphImports-only paths, and order-preservation (the regression fix)

Verification:
- All existing C/C++ transitive tests pass unchanged
- Previously failing `cpp.test.ts > resolves run → write_audit to one.h
  via arity narrowing` now passes
- `tsc --noEmit` clean
- 225/225 tests pass across wildcard-synthesis, cross-file-binding,
  cpp resolver, and new closure unit tests

* fix(ingestion): bound closure size, O(1) dequeue, track Strategy 4 (#816 review)

Address @xkonjin's review feedback on the import-resolution strategy refactor:

1. **DoS guard**: cap transitive closures at 5,000 files via
   `MAX_TRANSITIVE_CLOSURE_SIZE`. Pathological codebases (boost-style headers,
   monoheader kernels) could previously produce closures with tens of thousands
   of entries per translation unit. BFS now stops early and returns a partial
   closure rather than risking OOM. The closest-headers-first BFS ordering
   means the partial closure still contains the files overload resolution
   cares about.

2. **Perf**: replace `Array.prototype.shift()` (O(n)) with a head-index queue
   (O(1) dequeue). Deep chains previously had quadratic BFS behavior; now
   linear in closure size.

3. **Strategy 4 tracking**: change TODO in `dispatchSynthesis` to
   `TODO(#821)` referencing the filed issue for TS `export *` / Rust
   `pub use` DAG-walk implementation, and clarify that today's leaf
   fallthrough preserves correctness for direct imports — only the extra
   re-export traversal is missing.

4. **Test**: new unit test exercising the 5,000-file cap on a 10k-file
   synthetic chain, verifying partial-closure invariants (starts from
   importer side, bounded, deep nodes excluded).

Not addressed in this commit (followups):
- Review point 3 (graphImports-only deep-chain *integration* fixture):
  unit tests already exercise the `graphImports` traversal path directly
  in isolation and combined with `importMap`. A fixture that stresses
  graphImports-only transitive resolution is valuable but requires
  understanding when the pipeline populates graphImports distinctly from
  ctx.importMap — tracking as a followup rather than blocking this PR.

---------

Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-14 09:39:17 +01:00
Copilotandmagyargergo 1a597f3cc6 Fix npm arborist crash caused by tree-sitter-dart tarball URL format (#820)
* Initial plan

* fix: change tree-sitter-dart from tarball URL to git URL to fix npm arborist crash, add error handling and troubleshooting docs

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/382c76c6-89c3-463a-8631-2a5d6510be4c

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refine error handler patterns and troubleshooting docs for arborist crash

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/909b319b-c367-40aa-8033-32dfb6231d4e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* style: run prettier on changed files

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/50eb6b94-9300-4bf2-9b61-c2d78f637fc6

* fix: use github: shorthand for tree-sitter-dart to avoid SSH in CI

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2ce3f4b3-4c1e-4c39-b824-c25cfe145529

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* revert: use git+https:// for tree-sitter-dart instead of github: shorthand (fixes arborist crash from PR #811)

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b786b68d-6c76-4054-88eb-ad46ea9f5b81

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-04-14 08:57:29 +01:00
Gergő Magyar 759c983dce fix(extractors): resolve 3 silent contract mis-resolution bugs (#793) (#817)
* fix(extractors): resolve 3 silent contract mis-resolution bugs (#793)

Addresses Codex adversarial review findings for extractor contract
resolution on the new group extractor surface.

F1 (manifest-extractor): resolveSymbol passed the full "METHOD::path"
contract string through normalizeRoutePath, producing "/GET::/api/orders"
which never matches Route.name. Adds parseHttpContract() helper that
strips the METHOD:: prefix before path normalization. Contract ID
construction (buildContractId) is unchanged.

F2 (http-route-extractor): graph-assisted backfill used path-only
detections.find(), so multi-verb same-URL files attached the wrong
verb/handler to provider rows and inferred the wrong verb on FETCHES
consumer edges. Now requires path+method match when method is known,
and skips backfill when method is unknown and multiple detections tie
on path.

F3 (grpc-extractor): resolveProtoConflict seeded bestScore=-1 and only
replaced on strict >, so all-zero-score ties silently selected
candidates[0]. Now computes all scores, counts ties at the top score,
and returns null on ambiguity (caller skips contract emission and
warns with service name + candidate paths).

All three fixes are test-first; 73 tests pass across the three suites.
No schema changes, no new dependencies, contract ID wire format
(http::METHOD::path, grpc::pkg.Service/Method, http::*::path) preserved.

* fix(extractors): address PR #817 review — ambiguous symbol pick + contract id casing

Copilot + Claude review on PR #817 flagged two follow-up bugs on top of
the F1/F2/F3 fixes:

1. http-route-extractor: ambiguous multi-verb case left handlerName null
   but still ran the CONTAINS DB query. pickSymbolUid(syms, null) then
   silently picked pool[0] — reintroducing handler mis-attribution via
   a different route than the .find() bug F2 fixed. Now gates symbol
   enrichment on an ambiguousCandidates flag so the file-basename
   fallback wins instead.

2. manifest-extractor: buildContractId passed raw user casing through
   for the explicit-method form, so get::/api/orders and
   GET::/api/orders produced different contract ids even though
   parseHttpContract upper-cases during lookup. Now reuses
   parseHttpContract + normalizeRoutePath to canonicalize both method
   and path, so logically equivalent manifest inputs share a contract
   id (and share a manifestSymbolUid fallback).

Adds one regression test per bug: lowercase vs uppercase manifest
contract ids must match, and ambiguous multi-verb with CONTAINS rows
must not silently attach a real handler or call the CONTAINS query
at all. 75 tests pass across the three extractor suites.

* chore: prettier formatting
2026-04-14 08:03:02 +01:00
Abhigyan Patwari 988b905abe Merge pull request #767 from noCharger/feat/chat-scroll-pause
feat(web): add smart chat scroll
2026-04-14 01:25:09 +05:30
Gergő Magyar 3fbee2d3d2 chore: release v1.6.1 (#815) 2026-04-13 20:38:09 +01:00
26ff700e37 refactor(pipeline): DAG-based phase architecture + container-logic extraction to LanguageProvider (#809)
* Initial plan

* refactor: move language-specific container node logic into LanguageProvider

- Add resolveEnclosingOwner hook to LanguageProviderConfig
- Add staticOwnerTypes to MethodExtractionConfig
- Implement Ruby resolveEnclosingOwner (singleton_class → class/module)
- Replace hardcoded STATIC_OWNER_TYPES with config.staticOwnerTypes
- Move Ruby static types to rubyMethodConfig
- Move Kotlin static types to kotlinMethodConfig
- Remove Ruby singleton_class branch from findEnclosingClassInfo
- Collapse seqFindEnclosingClassNode/seqFindRawEnclosingContainerNode
  into single provider-aware seqFindEnclosingOwnerNode
- Update worker path to pass provider.resolveEnclosingOwner

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/bc9f9d4d-f749-4872-9ff2-17fc86e08787

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: add regression tests for config-driven staticOwnerTypes and resolveEnclosingOwner hook

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/bc9f9d4d-f749-4872-9ff2-17fc86e08787

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor: implement DAG-based pipeline architecture with phase extraction

Restructure the ingestion pipeline from a ~1800-line monolithic orchestrator
into a DAG (Directed Acyclic Graph) of named phases with explicit dependencies.

New files under pipeline-phases/:
- types.ts: PipelinePhase, PipelineContext, PhaseResult contracts
- runner.ts: DAG runner with topological sort validation
- scan.ts, structure.ts, markdown.ts, cobol.ts: early phases
- parse.ts + parse-impl.ts: chunked parse + resolve (the core)
- routes.ts, tools.ts, orm.ts: post-parse enrichment phases
- cross-file.ts + cross-file-impl.ts: cross-file binding propagation
- mro.ts, communities.ts, processes.ts: graph analysis phases
- index.ts: barrel export

pipeline.ts reduced from ~1960 lines to ~184 lines:
- DAG phase array declaration
- runPipelineFromRepo as thin orchestrator
- topologicalLevelSort retained for backward compat

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/136bf9c3-2f4f-449b-9fff-001332c8371c

* test: add DAG runner unit tests, update ARCHITECTURE.md with phase DAG docs

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/136bf9c3-2f4f-449b-9fff-001332c8371c

* fix: address code review - pass resolutionContext through parse output, fix worker URL path

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/136bf9c3-2f4f-449b-9fff-001332c8371c

* fix: declare transitive parse dependency explicitly in mro/communities/processes phases

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/136bf9c3-2f4f-449b-9fff-001332c8371c

* refactor: improve pipeline-phases clean code and folder structure

- Extract synthesizeWildcardImportBindings to wildcard-synthesis.ts
- Extract extractORMQueriesInline to orm-extraction.ts
- Create shared constants.ts for AST_CACHE_CAP
- Fix inline type import in orm.ts (use proper top-level import)
- Add comprehensive JSDoc to getPhaseOutput explaining type safety
- Move isDev to module level in cross-file.ts (consistency)
- Improve module-level documentation across files
- Organize barrel exports in index.ts with section comments

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2bd6d4aa-6271-4009-8dd2-332ea8ec73ab

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* address review feedback: fix circular dep, allFetchCalls mutation, progress bugs, remove DAG naming, extract isDev, fix _item naming, fix O(n²) line calc

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6cf53c9b-d55d-4c6f-bf3d-7bfb82d512b6

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* improve JSDoc on lineNumberAtOffset binary search

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6cf53c9b-d55d-4c6f-bf3d-7bfb82d512b6

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* address review: filter deps in runner, move totalFiles to ctx, fix cycle JSDoc, centralize isDev, remove DAG naming

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b388424f-b939-4a94-97de-3855f9465564

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix doc consistency in graph-sort.ts module-level and function-level JSDoc

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b388424f-b939-4a94-97de-3855f9465564

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(pipeline): wrap phase errors with phase name and emit terminal error progress event

Restores phase diagnostics at CLI/MCP boundary. runPipeline now wraps
phase.execute() in try/catch and rethrows with 'Phase <name> failed: ...'
preserving the original via { cause }. Also emits a terminal
{ phase: 'error' } progress event so subscribers see the failure before
the rejection propagates. Handler errors during error reporting are
swallowed to keep the original cause authoritative.

Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U1)

* fix(pipeline): move bindingAccumulator dispose into crossFile try/finally; make single-use

crossFile.execute() now wraps its body in try/finally so the accumulator
is released on both the happy path and when runCrossFileBindingPropagation
throws. Dev-mode telemetry stays inside the try block before dispose (all
three counters return 0 after dispose clears internal maps).

BindingAccumulator becomes single-use: appendFile after dispose now throws
'BindingAccumulator: use after dispose' instead of silently re-animating
via the old _disposed auto-clear. Docs updated; the only production
construction site (parse-impl) always creates a fresh instance per run,
so no caller relied on the re-use contract.

Residual risk documented in crossFile module JSDoc: a future phase
inserted between parse and crossFile that throws would still leak the
accumulator. Any such phase must manage accumulator lifetime explicitly.

Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U2)

* docs(pipeline): explain why importCtx teardown is safe before crossFile

Investigation (plan U3) confirms: `importCtx` (ImportResolutionContext)
is a scratch workspace with no downstream consumer after parse.
`resolutionContext` (returned to crossFile) is a distinct object that
owns importMap / namedImportMap / packageMap / moduleAliasMap / model,
and never closes over importCtx. cross-file-impl consumes only that
ctx via processCalls. The two confusingly-similar "context" names
were the root of the adversarial reviewer's concern — comment locks
in the invariant so the next reader sees it.

No behavioral change.

Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U3)

* refactor(pipeline): remove ctx.totalFiles side-channel; promote to ParseOutput

totalFiles was a hidden mutable field on PipelineContext written by
parse and read by mro/communities/processes — five reviewers flagged
this as a violation of the immutable-context invariant. Removed from
PipelineContext, which is now fully readonly, and made the implicit
temporal dep explicit: mro/communities/processes now declare 'parse'
as a dep and read totalFiles via getPhaseOutput<ParseOutput>(...).

No behavior change. Topo-sort unchanged because parse was already a
transitive dep through crossFile.

Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U4)

* feat(method-extractor): runtime staticOwnerTypes guard at factory chokepoint

createMethodExtractor now rejects MethodExtractionConfigs that list
companion_object / singleton_class / object_declaration in
typeDeclarationNodes but omit the matching entry from staticOwnerTypes.
Fails loudly at provider construction time instead of producing
silent isStatic=false on the 50000th file analyzed.

Opt-out convention preserved: an explicit `new Set()` (empty Set)
signals intentional exclusion and passes the guard (memory obs #30588).

All 13 existing language configs pass the guard; the new negative test
fails without it. Test-first.

Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U5)

* fix(pipeline): wrap sequential-fallback in try/finally so cleanup survives throws

The sequential-fallback block in runChunkedParseAndResolve now runs
inside a try/finally that guarantees astCache.clear(), accumulator
finalize, and enrichExportedTypeMap execute even if readFileContents
or processCalls throws mid-fallback. Cleanup failures are caught
inside the finally so they can't mask the original error.

Accumulator disposal ownership remains with crossFile (U2) — U6 only
adds astCache cleanup and preserves finalize ordering on the error
path.

Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U6)

* test(pipeline): direct unit coverage for wildcard-synthesis and cross-file-impl

Both modules previously had zero direct unit coverage — branches were
exercised only through integration tests' happy paths.

wildcard-synthesis.test.ts covers: Go graph-IMPORTS fallback, Python
moduleAliasMap build, MAX_SYNTHETIC_BINDINGS_PER_FILE cap, dedup
against existing namedImportMap entries, and empty-exportedSymbols
early return.

cross-file-impl.test.ts covers: gapRatio below threshold no-op,
MAX_CROSS_FILE_REPROCESS cap, graph-only exportedTypeMap fallback,
and empty namedImportMap short-circuit.

Tests assert current behavior — any future regression flips them.

Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U7)

* test(pipeline): golden-file graph-parity regression guard on mini-repo fixture

Pins the current post-P1/P2 graph output (57 symbols, 92 relationships,
4 processes, deterministic edge digest) so future silent refactors
cannot drift behavior unnoticed. If any count changes or any edge
rewires, the test fails with a readable diff listing what changed
and a copy-pasteable UPDATE_GOLDEN=1 regen command.

Edge digest keyed by symbolic (label, name, filePath) triples rather
than raw generateId output — stays meaningful across id-encoding
refactors while still catching real semantic rewiring.

Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U8)

* fix(pipeline): minimal cycle reporting + resolveEnclosingOwner loop safeguards

U9: runner cycle detection now reports only the SCC members via DFS
back-edge trace ('Cycle detected: A -> B -> C -> A') rather than
everything with inDegree > 0 (which mixed cycle members with blocked
dependents). Also emits the 'error' progress event for graph-
validation failures, symmetric with U1's runtime-error path.

U16: findEnclosingClassInfo now defends against language-provider
hooks that return non-container nodes — visitedContainers Set breaks
repeat-visit loops, MAX_ENCLOSING_WALK_ITERATIONS is belt-and-braces.
Documented the hook contract invariant so future provider authors
know the walk-continues-upward expectation.

Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U9, U16)

* refactor(pipeline): type hygiene, dead code cleanup, shared allPathSet, graph-sort naming

Bundles plan units U10, U11, U12, U14, U15:

U10 — Type hygiene: readonly ParseOutput arrays (allExtractedRoutes,
allDecoratorRoutes, allToolDefs, allORMQueries, allPaths); removed
redundant 'as string[] | undefined' cast in routes.ts and 'as URL' in
parse-impl.ts; WorkerPool is now 'import type'. Readonly contract
propagated into processORMQueries (only iterates).

U11 — Dead code & shims: deleted constants.ts shim (AST_CACHE_CAP
inlined into its sole real consumer cross-file-impl.ts; isDev
consumers now import directly from ../utils/env.js). Removed internal
utility re-exports from pipeline-phases/index.ts (no external
consumers). Removed topologicalLevelSort re-export from pipeline.ts;
updated topological-sort.test.ts to import from the canonical
utils/graph-sort.js. Stripped 'Phase 3+4:' stale JSDoc from
parse-impl.ts.

U12 — Perf: StructureOutput now carries allPathSet (ReadonlySet<string>)
built once; cobol, markdown, and cross-file-impl consume the shared
set instead of allocating their own. Parse forwards it via
ParseOutput.allPathSet; processCobol/processMarkdown widened to
ReadonlySet<string>.

U14 — graph-sort.ts: renamed local 'inDegree' to
'pendingImportsPerFile' with expanded JSDoc explaining the reverse-
graph Kahn's formulation and warning future maintainers not to
'correct' it to standard in-degree semantics. Added self-edge test.

U15 — Unconditional worker-fallback logging: removed isDev guard on
the worker-pool-creation-failure console.warn so operators can
diagnose perf degradations in production.

No behavior change. U8 golden-file test confirms pipeline output is
byte-identical.

Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U10, U11, U12, U14, U15)

* docs: fix ARCHITECTURE.md table integrity; bump AGENTS.md/CLAUDE.md to 1.3.0

U13 — documentation fixes:

ARCHITECTURE.md: the prior insertion of the 'Pipeline Phase DAG'
section orphaned 7 rows from the 'Where to change what' header.
Moved those 7 rows back up under their header so the table reads
contiguously; DAG section now follows the completed table.

AGENTS.md + CLAUDE.md: bumped version 1.2.0 -> 1.3.0, updated Last
reviewed to 2026-04-13, added matching Changelog row documenting
the GitNexus index stats refresh after the DAG refactor. Stat
bumps (symbols/relationships/execution flows) that were sitting
uncommitted in the working tree are now landed under a proper
changelog entry per each file's own documented schema.

Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U13)

* refactor(pipeline): drop spurious parse deps, true-readonly ParseOutput.exportedTypeMap, skip redundant wildcard synth

- mro/communities/processes: switch redundant `parse` dep to `structure` —
  totalFiles originates in structure, so depending on parse for it was a
  spurious data dep that obscured the real DAG.
- ParseOutput.exportedTypeMap: typed as truly ReadonlyMap<...,ReadonlyMap>>;
  graph→exports enrichment moved into parse-impl so the snapshot is
  fully populated at parse return. crossFile builds its own local mutable
  working copy for per-file re-resolution writes — no cast at the boundary.
- parse-impl: hasSynthesized flag guards the unconditional final
  synthesizeWildcardImportBindings call when per-chunk/fallback synthesis
  already ran (graph-global + idempotent across chunks).
- cross-file-impl: documented the intentional `phase: 'parsing'` progress
  label so telemetry bucketing stays consistent with the parse phase.
- cross-file-impl test: replaced the now-moved fallback-enrichment
  assertion with a stronger one — crossFile must not mutate the
  parse-supplied map.

Addresses PR #809 review pass 5 carry-overs.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-13 20:31:05 +01:00
Copilotandmagyargergo 6388113e10 fix: prevent stack overflow and memory exhaustion on large repo analysis (#814)
* Initial plan

* fix: prevent stack overflow and memory issues on large repo analysis

- Convert c3Linearize from recursive to iterative (explicit work stack)
  to handle deep class hierarchies without stack overflow
- Replace push(...arr) spread patterns with safe loops in parse-worker.ts
  and lbug-adapter.ts to prevent stack overflow on large arrays
- Stream relationship CSV lines directly to per-pair temp files in
  lbug-adapter.ts instead of accumulating millions of lines in memory
- Add test for deep 500-level inheritance chain

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9cb2eed2-adc7-4fa4-9216-e7ac3facb9b5

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: add stack size flag and enhanced error messages for large repos

- Auto-set --stack-size=4096 alongside --max-old-space-size in analyze
  command to prevent stack overflow on deep class hierarchies
- Add helpful error guidance for known large-repo failure modes
  (stack overflow, heap OOM, Map size limits)

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9cb2eed2-adc7-4fa4-9216-e7ac3facb9b5

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: address code review feedback

- Add error handling for write stream close in lbug-adapter.ts
- Handle backpressure when writing relationship CSV lines to disk
- Clarify ENTER/MERGE phase transition comment in resolve.ts
- Fix inconsistent stack size in error message (4096 not 8192)

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9cb2eed2-adc7-4fa4-9216-e7ac3facb9b5

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: address PR review — resource leak, redundant guard, Set, test depth

- Fix write-stream resource leak on readline error by destroying all
  open WriteStreams before rejecting (lbug-adapter.ts)
- Switch failedPairCsvPaths from array to Set for O(1) lookup
- Remove redundant MERGE-phase empty-parents guard in resolve.ts
  (unreachable — ENTER phase already handles that case)
- Increase deep inheritance test DEPTH from 500 to 2000 for
  reliable regression coverage across platforms

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/cf1f3e22-3864-454a-a3a5-2bded9ebfdba

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* style: fix prettier formatting in lbug-adapter.ts

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b5ca33c4-bb03-402f-a206-21ea7e1e310e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore: revert unintended package.json/lock changes

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b5ca33c4-bb03-402f-a206-21ea7e1e310e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: strip NODE_OPTIONS in skip-git-cli test child processes

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/fc59cd11-348b-4e22-b9ea-98787300de48

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: don't put --stack-size in NODE_OPTIONS (rejected by Node 24)

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/fc59cd11-348b-4e22-b9ea-98787300de48

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: pass --stack-size as CLI arg only, not in NODE_OPTIONS (Node 24 compat)

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/fc59cd11-348b-4e22-b9ea-98787300de48

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-04-13 20:08:49 +01:00
Copilotandmagyargergo c672697012 fix: replace tree-sitter-dart git URL with tarball to fix npm install crash (#811)
* Initial plan

* fix: replace tree-sitter-dart git URL with tarball URL to fix npm install crash

The `github:` git URL for tree-sitter-dart caused npm's arborist to
create a dependency node with a null target during the rebuild phase,
crashing global installs with:
  Cannot destructure property 'package' of 'node.target' as it is null.

Using a GitHub archive tarball URL instead avoids the arborist bug while
still installing from the exact same commit (80e23c0).

Fixes #805

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c15109ae-0865-4d69-bd08-9972dcfe18f9

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-04-13 18:57:56 +01:00
Deepak Chauhan d786e692af [cli] Preserve Ruby singleton_class context in sequential parsing (#774)
* fix(parsing): preserve ruby singleton class context

* refactor(parsing): clarify singleton class helpers
2026-04-13 13:12:52 +01:00
Arkh74278andarkh a6421b3b1b [dart] Add call patterns for await, cascade, lambda, and widget-tree contexts (#801)
* feat(dart): add call patterns for await, cascade, lambda, and widget-tree contexts

* fix(dart): address review feedback — await member-chain, cascade comment, static_final comment, add to query-compilation smoke test

* test(dart): add integration tests for await and widget-tree call patterns

* style: apply prettier formatting to dart integration tests

---------

Co-authored-by: arkh <local@localhost>
2026-04-13 11:21:11 +01:00
Copilotandmagyargergo 9f4109a33f fix: remove file:../gitnexus-shared from runtime dependencies (#803)
* Initial plan

* fix: remove file:../gitnexus-shared from dependencies to fix npm install outside monorepo

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2ea9c1da-1b0c-4ab0-b370-f3970cc54ffa

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-04-13 10:17:30 +01:00
Copilotandmagyargergo 79e1d933fa fix: resolve generic TypeScript awaited function calls missing from call graph (#804)
* Initial plan

* fix: resolve generic TypeScript function callers missed by impact analysis

When a generic function call is combined with `await` (e.g. `await fn<T>(args)`),
tree-sitter-typescript parses it as a `call_expression` whose `function` field is
an `await_expression` rather than a bare `identifier`. The existing queries only
matched `call_expression { function: identifier }`, so these calls produced no
`@call.name` capture and were silently dropped from the call graph.

Fix: add two new tree-sitter query patterns to `TYPESCRIPT_QUERIES` that handle:
1. `await fn<T>(args)` — awaited generic free call
2. `await obj.fn<T>(args)` — awaited generic member call

Both patterns require the `(type_arguments)` child to be present (which is what
causes tree-sitter to parse the `function` field as an `await_expression`).
Non-generic awaited calls (`await fn(args)`) are unaffected: tree-sitter parses
them as `await_expression { call_expression { identifier } }`, which is still
captured by the existing first pattern.

Also adds a new test fixture `typescript-generic-calls` with two callers of a
generic `verifyToken<T>` function using `await` and three new integration tests.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/4cf75290-900b-4cea-8a65-2a245ff86970

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: clean up test fixture interface ordering and imports

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/4cf75290-900b-4cea-8a65-2a245ff86970

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: add coverage for awaited generic member-call form (await obj.fn<T>())

Address review feedback: the member-call query pattern was untested.

Adds service.ts (TokenService with generic verify<T> method) and guest.ts
(calls await svc.verify<GuestPayload>()) to the typescript-generic-calls
fixture, plus a new integration test asserting the CALLS edge resolves.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/fcbf8d99-8dbc-40ce-b2a3-60b8d63c095a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* revert: undo accidental ladybugdb version bump in package files

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/fcbf8d99-8dbc-40ce-b2a3-60b8d63c095a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* style: run prettier on changed files

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/8c7d8291-74bb-4a86-ae47-7c79e2cbb57e

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-04-13 10:16:42 +01:00
ivkondandClaude 4d4756fe86 feat(group): extractor expansion + manifest extractor (2/4 of #606 split) (#796)
* feat(group): extractor expansion + manifest extractor

Part 2 of 4 in the split of #606 (ticket: #792). Follows #795
(bridge.lbug storage foundation, already merged), but this PR has no
code-level dependency on #795 — it only imports types and the
ContractExtractor interface that existed on upstream main before
either PR. It could have been reviewed in parallel with #795.

## What changed

Expands the 3 existing contract extractors with substantially more
language/framework coverage, and adds a new `manifest-extractor`
that resolves `group.yaml`-declared cross-links against the per-repo
graph via exact-name lookups.

### New file (228 LOC)

- `gitnexus/src/core/group/extractors/manifest-extractor.ts` —
  exact graph lookup for `group.yaml`-declared cross-links. HTTP
  paths are canonicalized before Route.name matching; gRPC is
  resolved by service/method name (NO `.proto`-filename fallback);
  topic and lib use exact-name match. Falls back to a synthetic
  `manifest::<repo>::<contractId>` uid when the graph has no
  matching symbol, so cross-impact traversal still has a stable
  anchor for the contract.

### Modified extractors (+958 LOC prod)

- `extractors/grpc-extractor.ts` (+522) — `.proto` parser with
  comment and string-literal sanitization (braces inside strings no
  longer truncate service bodies); package/service/method canonical
  IDs; server/client detection across Go (`grpc.NewServer`,
  `RegisterXxxServer`, `XxxGrpc.XxxImplBase`), Java (`@GrpcService`,
  `BlockingStub`), Python (`servicer_to_server`, `XxxStub`), and
  TypeScript/Node (`@GrpcMethod`, `ClientGrpc`, `loadPackageDefinition`).
- `extractors/http-route-extractor.ts` (+174) — Go gin/echo/stdlib
  `HandleFunc`, NestJS `@Controller`+`@Get`/etc, Python FastAPI
  decorators, Java Spring `@RequestMapping`/`@GetMapping`,
  restTemplate / WebClient / OkHttp consumers.
- `extractors/topic-extractor.ts` (+98) — sarama `ProducerMessage{}`
  struct literal detection (replaces a constructor-anchored regex
  that missed topics inside producer loops), kafka-go Writer/Reader,
  Python NATS (`await nc.subscribe`/`await nc.publish`), JetStream
  helpers.

### Modified and new tests (+1264 LOC)

- `grpc-extractor.test.ts` (+539) — full coverage of the new proto
  parser (strings-with-braces regression, comments-with-braces
  regression), per-language server/client detection
- `http-route-extractor.test.ts` (+240) — per-framework route
  extraction + normalization edge cases
- `topic-extractor.test.ts` (+177) — the sarama in-loop regression,
  JetStream, Python NATS, kafka-go Writer/Reader
- `manifest-extractor.test.ts` (+308 NEW) — HTTP path normalization,
  gRPC exact lookup with proto-fallback regression, lib and topic
  exact matching, synthetic-uid fallback behavior

### Self-review fixes folded in

Carried forward from the #606 self-review (commit `d15b8cb`):

- **HIGH #1** — `manifest-extractor.resolveSymbol` was too fuzzy.
  Previously used `CONTAINS` on route/name fields plus an
  unconditional `filePath ENDS WITH '.proto'` fallback for gRPC.
  Consequences: `/orders` matched `/suborders`, and any repo with
  any `.proto` file returned a random proto symbol for a gRPC
  manifest entry. Replaced with exact equality + deterministic
  `ORDER BY` + synthetic-uid fallback for unresolved manifests.
  Regression tests included.
- **MED #3** — gRPC proto parser brace-depth counting now sanitizes
  strings and comments first (`stripProtoCommentsAndStrings`). A
  valid proto with `option deprecated_reason = "use NewService {
  instead"` used to have its service body closed early by the `"{"`
  inside the literal, silently dropping methods after the offending
  string. Regression tests for both string-with-brace and
  comment-with-brace cases.
- **MED #4** — sarama Kafka regex changed from
  `sarama.NewSyncProducer[\s\S]{0,300}?Topic:` (anchored on
  constructor, caught only first topic in a loop) to
  `sarama.ProducerMessage{...Topic:}` (matches every struct literal
  directly). Regression test with a for-loop that constructs
  multiple `ProducerMessage`s.
- **MED #7** — `manifest-extractor.resolveSymbol` no longer has a
  silent `catch { /* fall through */ }`. Errors from the graph
  executor are logged via `console.warn` with link type, contract
  name, repo key, and error message before falling through to the
  synthetic-uid path.

## Why

Reviewer focus here is pure regex / parser correctness — no
storage, no Cypher queries, no algorithmic changes to the cross-link
algorithm. Separating this from the bridge foundation PR (#795)
meant reviewers could stay in a single mental mode (parsing logic)
instead of context-switching between DDL, Cypher, and regex.

## How to verify

- `cd gitnexus && npx tsc --noEmit`
- `cd gitnexus && npx vitest run test/unit/group/grpc-extractor.test.ts --pool=forks`
- `cd gitnexus && npx vitest run test/unit/group/http-route-extractor.test.ts --pool=forks`
- `cd gitnexus && npx vitest run test/unit/group/topic-extractor.test.ts --pool=forks`
- `cd gitnexus && npx vitest run test/unit/group/manifest-extractor.test.ts --pool=forks`

Local pre-push: typecheck clean, all 99 extractor unit tests pass
(grpc 43, http 18, topic 30, manifest 8).

## Risk / rollback

**Low.** Extractors have no user-facing surface in this PR — they
produce `ExtractedContract[]` that is consumed by `sync.ts` in the
next split (#793). No existing behavior changes for users who don't
run a `group sync`. Rollback = `git revert` of the merge commit;
the modifications to `grpc-extractor.ts` / `http-route-extractor.ts`
/ `topic-extractor.ts` revert to the pre-PR versions that still
work (they're subsets of the new functionality).

## Scope discipline (per GUARDRAILS.md)

- Only the 8 files above are touched; no drive-by refactors
- No CI/release/security config changes
- No secrets or machine-specific paths
- Content lifted from #606 (CI 11/11 green on `d15b8cb`)

## Dependencies

- **Base:** `main` (upstream already includes #795 as `1ff324c`)
- **Blocks:** sync pipeline (#793) and the cross-impact feature (#794)
- **Tracker issue:** #792
- **Parent PR:** #606

Co-authored-by: Claude <noreply@anthropic.com>

* refactor(group): migrate topic-extractor from regex to tree-sitter queries

Addresses @magyargergo's feedback on #796 that regex-based lookups
should use tree-sitter nodes instead, and that the top-level
extractors must NOT carry language dependencies. This is phase 1 of
a multi-step migration — topic-extractor first because its patterns
are the most uniform (16 "call/annotation with first-arg string
literal" variants), which makes it a clean proof of the approach
before grpc-extractor and http-route-extractor get the same treatment.

## Architecture: language-agnostic orchestrator + per-language plugins

The top-level extractor is a thin orchestrator that never imports a
tree-sitter grammar or a query string. Per-language knowledge lives
in a new `topic-patterns/` folder with one file per language plus a
registry that maps file extensions to compiled plugins:

```
src/core/group/extractors/
├── tree-sitter-scanner.ts         # shared, language-agnostic scanning utilities
├── topic-extractor.ts              # thin orchestrator (no grammar imports)
└── topic-patterns/
    ├── types.ts                    # TopicMeta, Broker
    ├── index.ts                    # registry: extension → compiled provider
    ├── java.ts                     # tree-sitter-java + JAVA_TOPIC_PROVIDER
    ├── go.ts                       # tree-sitter-go + GO_TOPIC_PROVIDER
    ├── python.ts                   # tree-sitter-python + PYTHON_TOPIC_PROVIDER
    └── node.ts                     # tree-sitter-javascript + tree-sitter-typescript
                                    # → JAVASCRIPT_/TYPESCRIPT_/TSX_TOPIC_PROVIDER
```

**Shared scanner (`tree-sitter-scanner.ts`)** — defines
`PatternSpec<TMeta>`, `LanguagePatterns<TMeta>`, `CompiledPatterns<TMeta>`
and the `scanFile(parser, plugin, content)` helper. Plugins compile their
queries eagerly at module load via `compilePatterns()`, so a broken
pattern fails loudly at import time instead of silently at scan time.
`unquoteLiteral()` handles single/double/template quotes, Python
triple-quoted strings, and Go raw backtick strings.

**Per-language plugins** own:
- the tree-sitter grammar import (this is the ONLY place in
  `src/core/group/` where tree-sitter grammars are imported),
- the query S-expressions,
- the `TopicMeta` payload (role, broker, confidence, symbolName) that
  the orchestrator receives back on every match.

Each plugin uses a `@value` capture name to bind the topic literal node.
The JavaScript and TypeScript grammars share AST node names for every
construct we query, so `node.ts` defines the pattern sources once and
compiles them against `JavaScript`, `TypeScript.typescript`, and
`TypeScript.tsx` — exporting three providers because `Parser.Query`
objects are NOT portable across grammar instances.

**Registry (`topic-patterns/index.ts`)** — maps `.java` → Java provider,
`.go` → Go, `.py` → Python, `.js`/`.jsx` → JS, `.ts` → TS, `.tsx` → TSX.
Also exports `TOPIC_SCAN_GLOB` so adding a new language is a single
file-level edit (drop `topic-patterns/<lang>.ts`, import + register it
here — zero edits required in `topic-extractor.ts`).

**Orchestrator (`topic-extractor.ts`)** — ~110 lines, no grammar or
query imports. Per file: `getProviderForFile(rel)` → `scanFile(parser,
provider, content)` → `unquoteLiteral(valueText)` → `makeContract(...)`.
Reuses one `Parser` instance across files; the scanner calls
`setLanguage` per plugin.

## Why this is better than regex

1. **Comments and strings are respected for free.** The old regex
   would match `// kafkaTemplate.send("fake.topic")` as a real
   producer; tree-sitter never visits comments or string literals as
   code nodes, so false positives from commented-out code are
   eliminated.
2. **Struct/object literal patterns are structural, not textual.**
   `sarama.ProducerMessage{Topic: "..."}` no longer needs a 300-char
   lookahead (which was a known cross-match bug partly mitigated by a
   loop regression test in the self-review). The new query matches a
   specific `composite_literal` with a specific `qualified_type` and
   `keyed_element` — exactly one struct literal per match.
3. **No order-of-operations fragility.** Regex for
   `channel.publish` vs `channel.consume` was independent and
   file-wide; the AST scopes matches to the specific `call_expression`.
4. **Language-agnostic extension.** Adding Ruby, Rust, or C# topic
   detection later means dropping one file in `topic-patterns/` — no
   changes to shared scanner or orchestrator, and no tree-sitter
   imports leak into top-level code.

## Per-file fault tolerance

- Malformed files that tree-sitter can't parse are silently skipped
  (`parser.parse` is wrapped by `scanFile`). The ingestion pipeline
  already logs unparseable files at index time.
- A syntactically invalid query is caught at `compilePatterns` time,
  not scan time — broken plugins fail loudly at import.
- Per-pattern `matches()` failures are swallowed so one broken query
  in a plugin doesn't block the rest.

## Tests

All 30 existing `topic-extractor.test.ts` tests pass **without any
changes to the test file** — they were written as input/output contract
tests (given this source file, expect these `ExtractedContract` objects)
and that contract is unchanged. Regression coverage includes:

- Kafka: Java `@KafkaListener` + `kafkaTemplate.send`; Node
  `producer.send` + `consumer.subscribe`; Go sarama producer/consumer
  (sync and async); kafka-go Writer/Reader; Python `KafkaConsumer` +
  `producer.send/produce`
- RabbitMQ: Java `@RabbitListener` + `rabbitTemplate.convertAndSend`;
  Node `channel.consume/publish/sendToQueue`; Python `basic_consume/
  basic_publish` with keyword args
- NATS: Go and Node `nc.Subscribe/Publish`; Go and Node JetStream
  `js.Subscribe/Publish`; Python `await nc.subscribe/publish`

Including the regression test for the sarama `ProducerMessage`
in-loop case — the AST-based query captures every literal in the
file independently, not just the first one after `NewSyncProducer`.

## Neighbor regression check

- `topic-extractor.test.ts` — 30/30 pass (rewritten extractor)
- `http-route-extractor.test.ts` — 18/18 pass (untouched)
- `grpc-extractor.test.ts` — 43/43 pass (untouched)
- `manifest-extractor.test.ts` — 8/8 pass (untouched)
- Full `npx tsc --noEmit` clean

## Scope discipline (per GUARDRAILS.md)

- Only files under `src/core/group/extractors/` are touched; no
  changes to other extractors, tests, MCP surface, or pipeline.ts.
- No CI/release/security config changes, no secrets.
- New tree-sitter imports all reference grammars that are already
  installed as dependencies (`tree-sitter`, `tree-sitter-javascript`,
  `tree-sitter-typescript`, `tree-sitter-python`, `tree-sitter-java`,
  `tree-sitter-go` — all in `package.json` for the existing pipeline).

## Phase 2 / phase 3 plan

- **Phase 2 (next commit):** rewrite `http-route-extractor.ts`
  Strategy B (regex fallback) on the same plugin pattern. Graph-assisted
  Strategy A stays as-is (already uses pipeline-built tree-sitter data
  via `HANDLES_ROUTE` Cypher queries).
- **Phase 3 (commit after):** rewrite `grpc-extractor.ts` for Java /
  Go / Python / TypeScript detection. `.proto` files are the one
  outstanding question — there is no `tree-sitter-proto` grammar
  installed; the in-tree string-sanitizing parser stays as a pragmatic
  exception with a comment, alternative being to add
  `tree-sitter-proto` as a dep (open for the maintainer).

Co-authored-by: Claude <noreply@anthropic.com>

* refactor(group): migrate http-route-extractor Strategy B to tree-sitter plugins

Phase 2 of the extractor refactor requested by @magyargergo on #796.
Same architecture as the phase 1 topic-extractor rewrite: a thin,
language-agnostic orchestrator plus per-language plugins that own
tree-sitter grammars and query sources. The top-level extractor file
no longer imports any tree-sitter grammar or query string.

## Architecture

```
src/core/group/extractors/
├── tree-sitter-scanner.ts          # shared, language-agnostic primitives
├── http-route-extractor.ts         # thin orchestrator (no grammar imports)
└── http-patterns/
    ├── types.ts                    # HttpDetection, HttpLanguagePlugin, HttpRole
    ├── index.ts                    # registry: ext → plugin + HTTP_SCAN_GLOB
    ├── java.ts                     # tree-sitter-java: Spring + RestTemplate/WebClient/OkHttp
    ├── go.ts                       # tree-sitter-go: gin/echo/HandleFunc + http/resty consumers
    ├── python.ts                   # tree-sitter-python: FastAPI + requests
    ├── php.ts                      # tree-sitter-php: Laravel Route::get/...
    └── node.ts                     # tree-sitter-javascript + tree-sitter-typescript:
                                    #   NestJS controllers, Express, fetch, axios
```

**Shared scanner (`tree-sitter-scanner.ts`)** — generalised from phase 1:
- `ScanMatch<TMeta>.captures` is now a full `CaptureMap` (every named
  capture the query binds, not just a single `@value`). Topic extractor
  updated to read `match.captures.value` accordingly.
- New `runCompiledPatterns(plugin, tree)` helper lets plugins run
  multiple query bundles against the same pre-parsed tree. This is
  needed for HTTP plugins that combine a class-prefix query with a
  method-route query (Spring, NestJS).
- `scanFile` becomes a thin wrapper over `parser.parse + runCompiledPatterns`.

**HTTP plugin shape** — unlike topic plugins, HTTP plugins expose a
`scan(tree)` function rather than a flat pattern list. This reflects
HTTP's more complex extraction: each detection needs method + path +
handler name, and framework patterns like Spring `@RequestMapping` /
NestJS `@Controller` require cross-referencing a class-level prefix
with method-level annotations. Plugins internally use
`compilePatterns` + `runCompiledPatterns` and walk the AST to resolve
the class/method relationships.

**Per-framework coverage:**

- **Java (`java.ts`)**
  - Spring: `@RequestMapping("/api/v2")` class prefix + `@(Get|Post|Put|
    Delete|Patch)Mapping("/sub")` method routes, joined via the
    enclosing `class_declaration` node id.
  - `RestTemplate.getForObject/postForEntity/put/delete/patchForObject` →
    method derived from API name.
  - `WebClient.method(HttpMethod.X, "/path")` → method from
    `HttpMethod.X` capture.
  - `new Request.Builder().url("/path")` → OkHttp consumer.

- **Go (`go.ts`)**
  - gin / echo / chi frameworks: `\w+.GET("/path", handler)` captures
    upper-case verb + handler identifier.
  - `net/http.HandleFunc("/path", handler)` → provider (default GET).
  - `http.Get/Post/Head` consumer, `http.NewRequest("METHOD", ...)`,
    resty `client.R().Get/Post/...`.

- **Python (`python.ts`)**
  - `@app.get("/path")` FastAPI decorators.
  - `requests.get/post/...` and `requests.request("METHOD", "url")`.

- **PHP (`php.ts`)**
  - Laravel `Route::get/post/.../patch('/path', ...)` via
    `scoped_call_expression`. Uses `PHP.php_only` to match the
    existing ingestion pipeline's grammar selection.

- **Node (`node.ts`) — JS + TS + TSX**
  - Pattern sources defined once, compiled against three grammar
    variants (`JavaScript`, `TypeScript.typescript`, `TypeScript.tsx`)
    because `Parser.Query` objects are not portable across grammars.
    Exports three plugins sharing the same `scan` logic.
  - NestJS: `@Controller('prefix')` decorators are siblings of the
    class in `export_statement` / `program`; `@Get(':id')` decorators
    are siblings of the method in `class_body`. The plugin walks
    decorator → next named sibling to find the decorated class /
    method, then combines the class prefix with the method path.
    Only emits NestJS detections when the enclosing class has a real
    `@Controller` decorator — prevents false positives from generic
    classes that happen to use `@Get` from another library.
  - Express: `(router|app).<verb>('/path', ...)`.
  - `fetch(url)` (default GET) + `fetch(url, { method: 'X' })`
    (uses two queries + a SyntaxNode-id dedupe set so URL literals
    aren't double-emitted by the options variant).
  - `axios.get/post/...`.

## Orchestrator changes

`http-route-extractor.ts` drops every `scanXxxProviders` / `scanXxxConsumers`
regex method and replaces them with a single source-scan loop that
delegates to `getPluginForFile(rel).scan(tree)`. The orchestrator
still owns:

- **Path normalization** (`normalizeHttpPath`, `normalizeConsumerPath`)
  — language-agnostic string processing shared by both strategies.
- **Graph-assisted Strategy A** (`HANDLES_ROUTE` / `FETCHES` / `CONTAINS`
  Cypher queries) — unchanged in spirit. The only regex helpers it
  used (`inferMethodFromFileScan`, `pickJavaHandlerName`) are now
  replaced by a lookup against the plugin's detections for the same
  file: for each route row, find the detection whose normalized path
  matches, and pull the HTTP method + handler name from it.
- **Per-file parse cache** — the orchestrator parses each relevant
  file at most once per `extract()` call. Both the graph-assisted
  enrichment loop and the source-scan fallback share the same
  `cachedDetections` map, so we never run the plugin twice for the
  same file.

## Why this is better than the regex version

1. **Comments and strings for free.** The old regex would match
   `// router.get('/fake')` as a real Express route; tree-sitter
   never visits string/comment nodes.
2. **Structural controller-prefix.** Spring and NestJS class-prefix
   joining is now scoped to the enclosing class via `class_declaration`
   node ids, eliminating file-wide state that broke when a file had
   multiple controllers.
3. **Precise NestJS disambiguation.** The plugin only emits a NestJS
   detection when the enclosing class has a real `@Controller`
   decorator — the old regex would fire on any `@Get(...)` in the
   file regardless of surrounding context.
4. **Language-agnostic extension.** Adding Ruby / Rust / Kotlin HTTP
   detection later means dropping one file in `http-patterns/` — no
   changes to the shared scanner, the orchestrator, or the Strategy A
   Cypher queries.

## Tests

- `http-route-extractor.test.ts` — **18/18 pass** (tests unchanged;
  they're contract-style input/output tests and the contract shape is
  unchanged). Covers Spring class prefix, Express, gin/echo, stdlib
  HandleFunc, NestJS, Laravel, FastAPI for providers and
  fetch/axios/python-requests/rest-template/webClient/okhttp/go-stdlib/
  resty for consumers, plus graph-first Strategy A for both.
- `topic-extractor.test.ts` — **30/30 pass** after the `captures.value`
  API migration.
- `grpc-extractor.test.ts` — 43/43 pass (untouched; phase 3).
- `manifest-extractor.test.ts` — 8/8 pass (untouched).
- `service.test.ts`, `sync.test.ts`, `storage.test.ts` — 41/41 pass.
- `npx tsc -p tsconfig.json --noEmit` clean.

## Scope discipline (per GUARDRAILS.md)

- Only files under `src/core/group/extractors/` are touched.
- No changes to pipeline.ts, MCP surface, ingestion, or tests.
- No CI / release / security / secrets changes.
- Tree-sitter grammars imported by plugins (`tree-sitter-java`,
  `tree-sitter-go`, `tree-sitter-python`, `tree-sitter-php`,
  `tree-sitter-javascript`, `tree-sitter-typescript`) are all already
  in `package.json` for the existing ingestion pipeline.

## Phase 3 plan

- **grpc-extractor** gets the same treatment: plugin-per-language under
  `grpc-patterns/` for Java / Go / Python / TS detection. `.proto`
  files remain an open question — no `tree-sitter-proto` grammar is
  installed, so the in-tree string-sanitizing parser from PR #796's
  self-review stays as a pragmatic exception unless the maintainer
  wants us to add `tree-sitter-proto` as a new dep.

Co-authored-by: Claude <noreply@anthropic.com>

* refactor(group): migrate grpc-extractor source scans to tree-sitter plugins

Phase 3 (final) of the extractor refactor requested by @magyargergo on
#796. Same architecture as phase 1 (topic) and phase 2 (http): thin
language-agnostic orchestrator + per-language plugins that own
tree-sitter grammars and query sources. With this commit the top-level
extractors under `src/core/group/extractors/` import ZERO tree-sitter
grammars and ZERO query strings — every grammar import lives in a
`*-patterns/<lang>.ts` plugin file, and the orchestrators go through
the registry indirection.

## Architecture

```
src/core/group/extractors/
├── tree-sitter-scanner.ts         # shared primitives (unchanged)
├── grpc-extractor.ts               # orchestrator (only `.proto` parser left)
└── grpc-patterns/
    ├── types.ts                    # GrpcDetection, GrpcLanguagePlugin, GrpcRole
    ├── index.ts                    # registry: ext → plugin + GRPC_SCAN_GLOB
    ├── go.ts                       # tree-sitter-go: RegisterXxxServer, Unimplemented, NewXxxClient
    ├── java.ts                     # tree-sitter-java: @GrpcService + XxxImplBase + newBlockingStub
    ├── python.ts                   # tree-sitter-python: add_XxxServicer_to_server + XxxStub
    └── node.ts                     # tree-sitter-javascript + tree-sitter-typescript:
                                    #   @GrpcMethod, @GrpcClient field type,
                                    #   .getService<X>('Svc'), new XxxServiceClient,
                                    #   loadPackageDefinition dynamic constructors
```

## Per-language coverage

**Go (`go.ts`)**
- Provider: `\w+.RegisterXxxServer(...)` via `call_expression →
  selector_expression → field_identifier` + JS regex filter
  `^Register(\w+)Server$`.
- Provider: `pb.UnimplementedXxxServer` embedded in a struct via
  `struct_type → field_declaration_list → field_declaration →
  qualified_type → type_identifier` + JS filter.
- Consumer: `\w+.NewXxxClient(...)` via the same call_expression
  query + JS filter `^New(\w+)Client$`.

**Java (`java.ts`)**
- Provider: `class X extends YyyGrpc.YyyImplBase` — two queries
  handle the scoped and plain forms. `scoped_type_identifier`'s
  children are positional (no `scope:`/`name:` fields), so the
  query matches the two `type_identifier` children by position.
- `#match? @inner "ImplBase$"` restricts matches at query time.
- Whether the class has `@GrpcService` or not controls only the
  `source` metadata label — the plugin walks the class_declaration's
  `modifiers` child in JS to detect the marker_annotation.
- Consumer: `YyyGrpc.newStub(ch)` / `newBlockingStub(ch)` via a
  `method_invocation` query with `#match? @method
  "^new(Blocking)?Stub$"`, service name extracted via
  `^(\w+)Grpc$` on the object identifier.

**Python (`python.ts`)**
- Single call-expression query covers both bare identifier and
  `obj.method` attribute forms:
  `(call function: [(identifier) @fn (attribute attribute: (identifier) @fn)])`.
- Plugin filters `@fn.text` against two JS regexes:
  `^add_(\w+)Servicer_to_server$` (provider) and `^(\w+)Stub$`
  (consumer), with a reserved-names ignore list for the Stub case
  (Mock / Test / Fake / Stub).

**Node — JavaScript + TypeScript + TSX (`node.ts`)**
- Pattern sources defined once, compiled three times (one per grammar)
  because `Parser.Query` objects are not portable across grammars.
  Exports three `GrpcLanguagePlugin`s sharing the same `scan`.
- `@GrpcMethod('Service', 'Method')`: decorator query captures the
  two string literals. Confidence is hard-coded 0.8 regardless of
  proto map resolution (matches the original regex version's
  behaviour).
- `@GrpcClient(...) field: XxxServiceClient`: decorator query
  captures the decorator node, plugin walks up to find the enclosing
  `public_field_definition` (decorators on fields are CHILDREN of
  the field definition in tree-sitter-typescript, not siblings) and
  reads its first `type_annotation → type_identifier`, then runs the
  `^(\w+Service)Client$` JS filter.
- `client.getService<X>('AuthService')`: call-expression query on
  `member_expression.property = "getService"` + string literal arg.
- `new XxxServiceClient(...)`: `new_expression` with a bare
  identifier constructor, filtered by `^(\w+Service)Client$` so
  generic `new AuthClient(...)` (missing the `Service` infix) does
  NOT falsely register as a consumer. Preserves the regression test
  `test_extract_ts_non_service_client_constructor_is_ignored`.
- `loadPackageDefinition` dynamic loader: gated on
  `tree.rootNode.text.includes('loadPackageDefinition')`. When set,
  `new foo.bar.Xxx(...)` qualified constructors with a capitalised
  property name register as consumers.

## Orchestrator changes

`grpc-extractor.ts` loses every `scanGoProviders` / `scanJavaProviders`
/ ... helper and replaces them with a single source-scan loop that:

1. Parses each file with the plugin's grammar (one shared `Parser`
   instance across all files, `setLanguage` called per plugin).
2. Calls `plugin.scan(tree)` to get `GrpcDetection[]`.
3. Converts each detection to an `ExtractedContract` via the private
   `detectionToContract` helper, which:
   - Looks the short service name up in the proto map (filled by
     the `.proto` parser).
   - Picks confidence = `confidenceWithProto` if resolved, else
     `confidenceWithoutProto`.
   - Builds a method-level contract id (`grpc::pkg.Svc/Method`) when
     the detection carries a `methodName` (TS `@GrpcMethod` only),
     otherwise a service-level id (`grpc::pkg.Svc/*`).

Everything else — the `.proto` parser, `buildProtoContext`,
`buildProtoMap`, `resolveProtoConflict`, `serviceContractId`,
`stripProtoCommentsAndStrings`, `extractServiceBlocks`, the dedupe
function — stays exactly as before. The `.proto` parser is kept as a
pragmatic exception to the "no regex in extractors" rule because no
`tree-sitter-proto` grammar is installed in the repo; a comment at the
top of the file explains this and flags the maintainer option of
adding `tree-sitter-proto` as a dependency.

## Why this is better than the regex version

1. **Comments and strings are respected for free.** Matched node types
   are only code constructs, never text inside comments or string
   literals.
2. **No false positives on partial names.** The old `(\w+?)Grpc`-style
   regexes would cross-match unrelated identifiers; structural queries
   restrict matches to the exact AST shape (`scoped_type_identifier →
   type_identifier` pairs, `method_invocation → identifier` etc.).
3. **NestJS `@GrpcClient` is structural, not regex-based.** The old
   regex required a specific textual layout
   (`@GrpcClient(...) private readonly foo!: XxxServiceClient`); the
   plugin now walks the AST, so modifier order / optional modifiers /
   multi-line formatting don't break it.
4. **Language-agnostic extension.** Adding Kotlin / Rust / C# gRPC
   detection later is a one-file edit in `grpc-patterns/index.ts` —
   no touches to the shared scanner, the orchestrator, or the proto
   parser.

## Tests

- `grpc-extractor.test.ts` — **43/43 pass** (tests unchanged; the
  contract shape is identical). Covers .proto parsing (including the
  brace-inside-string regression), Go provider/consumer,
  Java @GrpcService / plain ImplBase provider + newBlockingStub
  consumer, Python servicer + stub, TS @GrpcMethod + @GrpcClient +
  .getService + new XxxServiceClient + loadPackageDefinition + the
  `AuthClient` vs `AuthServiceClient` discrimination, dedupe across
  multiple patterns in one file, proto-aware confidence, and the
  inherited-package resolution for split proto definitions.
- `topic-extractor.test.ts` — 30/30 pass.
- `http-route-extractor.test.ts` — 18/18 pass.
- `manifest-extractor.test.ts` — 8/8 pass.
- `service.test.ts`, `sync.test.ts`, `storage.test.ts` — 41/41 pass.
- `npx tsc -p tsconfig.json --noEmit` clean.

## Scope discipline (per GUARDRAILS.md)

- Only files under `src/core/group/extractors/` are touched.
- No pipeline.ts, MCP surface, ingestion, CI / release / security, or
  test changes.
- New tree-sitter grammar imports (`tree-sitter-go`, `tree-sitter-java`,
  `tree-sitter-python`, `tree-sitter-javascript`, `tree-sitter-typescript`)
  are all already installed for the ingestion pipeline.

## End of phase series

This commit completes the three-phase extractor refactor:
  - **Phase 1** (`ea06d11`): topic-extractor → `topic-patterns/`
  - **Phase 2** (`b6015f6`): http-route-extractor → `http-patterns/`
  - **Phase 3** (this commit): grpc-extractor → `grpc-patterns/`

Every remaining regex-based extractor helper under the `src/core/group/
extractors/` directory is either (a) language-agnostic string
processing (path normalization, dedupe keys) or (b) the `.proto`
parser, which is documented as an explicit exception.

Co-authored-by: Claude <noreply@anthropic.com>

* feat(group): add tree-sitter-proto for .proto file parsing

Addresses @magyargergo's suggestion on #796 to replace the manual
string-sanitizing .proto parser with a tree-sitter grammar.

- **Vendored `tree-sitter-proto`** in `vendor/tree-sitter-proto/`.
  Grammar source from [coder3101/tree-sitter-proto](https://github.com/coder3101/tree-sitter-proto)
  (latest `grammar.js`), parser.c regenerated with `tree-sitter-cli
  0.24` to produce ABI version 14 — compatible with the project's
  `tree-sitter 0.25` runtime (which supports ABI ≤ 14). Added as
  `optionalDependency` with `file:./vendor/tree-sitter-proto`.

- **New `grpc-patterns/proto.ts` plugin** — uses the same
  `compilePatterns` + `runCompiledPatterns` infrastructure as every
  other plugin. Two queries:
  - `(package (full_ident) @pkg)` — package declaration
  - `(service (service_name) @service_name (rpc (rpc_name) @rpc_name))`
    — one match per (service, rpc) pair

- **Graceful fallback** — `tree-sitter-proto` is an optional
  dependency. If it fails to install (platform incompatibility) or
  fails the runtime smoke-test (`setLanguage` + `parse` on a trivial
  proto), `PROTO_GRPC_PLUGIN` stays `null` and the orchestrator
  uses the existing manual parser. The smoke-test catches the
  `SyntaxNode` TDZ error that occurs in vitest's fork-based test
  runner.

- **Orchestrator updated** — when `hasProtoPlugin` is true, `.proto`
  files are handled by the plugin loop (they're included in
  `GRPC_SCAN_GLOB`), and the manual `parseProtoFile` loop is
  skipped. `buildProtoContext` still runs to build the proto map
  for cross-referencing source-file detections.

1. **No manual comment/string stripping.** The old parser needed
   `stripProtoCommentsAndStrings` (110 lines) to avoid counting
   braces inside comments and string literals. tree-sitter handles
   this natively.
2. **No brace-depth tracking.** `extractServiceBlocks` used a manual
   depth counter to find service boundaries. tree-sitter's AST gives
   us `service` → `service_name` + `rpc` → `rpc_name` directly.
3. **Performance.** tree-sitter's C-based parser is faster than
   character-by-character JS scanning + regex on large proto files.

- `grpc-extractor.test.ts` — **43/43 pass** (unchanged)
- All other extractor tests — 99/99 pass
- `npx tsc -p tsconfig.json --noEmit` clean

Co-authored-by: Claude <noreply@anthropic.com>

* chore: add .gitignore for vendored tree-sitter-proto build artifacts

https://claude.ai/code/session_01SFUCxgKMMQ8EgRHYw91xPU

* fix: correct .gitignore paths for vendored tree-sitter-proto

Patterns should be relative to the .gitignore file's directory.

https://claude.ai/code/session_01SFUCxgKMMQ8EgRHYw91xPU

* refactor(group): address Copilot review feedback on #796

Six fixes suggested by the Copilot AI review:

1. **`normalizeHttpPath` root-path edge case** — stripping trailing
   slashes on the input `/` produced an empty string, yielding
   malformed contract ids like `http::GET::`. Now preserves `/` for
   the root handler/fetch case.

2. **Dedupe `scanFiles` call** — `extract()` was globbing the
   source-scan file list twice (once for the provider fallback, once
   for the consumer fallback). Moved to a single lazy call that
   memoizes the result for the rest of the method.

3. **HTTP `scanFiles` now ignores `**/vendor/**`** — every other
   extractor's glob already ignored vendored sources; the HTTP one
   didn't. Fixed for consistency.

4. **`loadPackageDefinition` check is now structural** — was calling
   `tree.rootNode.text.includes('loadPackageDefinition')` which forces
   materialization of the entire file text from the parse tree
   (expensive on large files). Replaced with a dedicated compiled
   query on `(call_expression function: [(identifier) | (member_expression)])`
   so the check stays in the AST domain.

5. **`grpc-extractor.ts` header docstring updated** — still claimed
   ".proto parsing is not tree-sitter-based because no grammar is
   installed". Now describes the actual behaviour: tree-sitter when
   `tree-sitter-proto` is available (optionalDependency), manual
   fallback otherwise.

6. **Eliminated the double proto file parse on the fallback path** —
   `buildProtoContext` already globs + parses every `.proto` file to
   build `servicesByName`. On the `!hasProtoPlugin` branch the
   extractor was globbing + parsing again via the now-removed
   `parseProtoFile` helper. The fallback branch now iterates the map
   that `buildProtoContext` already produced to emit provider
   contracts directly — single pass per proto file.

## Tests

- `topic-extractor.test.ts` — 30/30 pass
- `http-route-extractor.test.ts` — 18/18 pass
- `grpc-extractor.test.ts` — 43/43 pass
- `manifest-extractor.test.ts` — 8/8 pass
- `npx tsc -p tsconfig.json --noEmit` clean

Co-authored-by: Claude <noreply@anthropic.com>

* refactor(group): address Claude review feedback (bugs + dedup + hygiene) on #796

Follows up `2f28bfc` with the remaining items from the Claude AI review:

## Bugs

**Bug 2 — Label-unaware Cypher queries in `resolveSymbol`.**
The manifest-extractor's lookup queries were `MATCH (n) WHERE n.name = $x`
with no label filter, so a topic/service/package name could silently match
any node type (File, Variable, Import, Folder, …). Added label filters:
- `topic` → `(n:Function|Method|Class|Interface)` (topics are best-effort
  symbol-name matches against listener/publisher symbols)
- `grpc` method → `(n:Function|Method)`
- `grpc` service → `(n:Class|Interface)`
- `lib` → `(n:Package|Module)`

All 8 manifest-extractor tests still pass (mock executor is
label-agnostic, but the production LadybugDB graph now gets correctly
scoped queries).

**Bug 8 — Tautological `!handlerName` condition.**
`http-route-extractor.ts:extractProvidersGraph` had
`let handlerName = null; if (!method || !handlerName) { ... }` — the
`!handlerName` clause was always true since there was no intervening
assignment. Simplified to always run the plugin-scan lookup (we need
the handler name even when `methodFromRouteReason` already resolved
the method).

## Clean code / dedup

**Design 7 — `readSafe` was copy-pasted in all three orchestrators.**
Extracted to `extractors/fs-utils.ts` as the single source of truth
for the path-traversal guard. Dropped the three local copies and the
now-unused `fs`/`path` imports from topic-extractor.

**Style 10 — Language-specific `_test.go` skip in the topic orchestrator.**
Was `if (rel.endsWith('_test.go')) continue;` inside the language-
agnostic extraction loop. Pushed into the glob's ignore list
(`'**/*_test.go'`) alongside the existing `node_modules`, `vendor`,
`dist`, `build` entries, with a comment explaining that other
languages' test file conventions either live in separate directories
(Python `tests/`, Java `src/test/`) or are already covered by the
existing ignores.

## Already addressed in `2f28bfc` (mentioned again in Claude review)

- Bug 3: `normalizeHttpPath('/')` returns `''` — fixed
- Bug 4: double glob + double parse of `.proto` — fixed
- Bug 5: `scanFiles` called twice in HTTP — fixed
- Bug 6: missing `**/vendor/**` in HTTP glob — fixed
- Design 9 partially: `tree.rootNode.text.includes('loadPackageDefinition')`
  replaced with a dedicated structural query

## Deferred

- Bug 1 (`http::*::path` vs `http::GET::path` matching) — out of scope;
  sync.ts matching logic lands in #793, manifest extractor already
  emits correct synthetic uids for unresolved HTTP contracts.
- Design 9 full (change plugin `scan(tree)` → `scan(tree, source)`) —
  the only real use case (`loadPackageDefinition` gate) is already
  fixed via a structural query, so the interface change would be
  cosmetic churn without a concrete consumer.

## Tests

- `topic-extractor.test.ts` — 30/30 pass
- `http-route-extractor.test.ts` — 18/18 pass
- `grpc-extractor.test.ts` — 43/43 pass
- `manifest-extractor.test.ts` — 8/8 pass
- `npx tsc -p tsconfig.json --noEmit` clean

Co-authored-by: Claude <noreply@anthropic.com>

* docs+fix(group): address remaining Claude review items + add pipeline flow chart

## Fixes

**Remaining 🔴 — HTTP contract id wildcard format.** Documented the
`http::*::<path>` format as an intentional wildcard for manifest links
that omit the HTTP method, alongside the explicit-method form
(`GET::/path` → `http::GET::/path`). The docblock on `buildContractId`
now states both forms, notes that wildcard-aware matching is the
responsibility of the sync / cross-impact layer (#793), and
recommends the explicit-method form whenever the author knows the
method (it round-trips through exact equality without needing
wildcard logic downstream). Tests unchanged — the wildcard format is
what they've always asserted.

**Minor 1 — stale comment at `manifest-extractor.ts:124-126`.** The
comment claimed "creates a contract with an empty symbolUid/ref" but
the code switched to `manifestSymbolUid(repo, contractId)` a few
commits back. Updated to describe the actual synthetic-uid fallback
semantics and the cross-impact path that relies on both sides of the
join deriving the same uid.

**Minor 2 — exhaustiveness guard on `buildContractId`.** The
`switch(type)` covered all five current `ContractType` variants but
silently returned `undefined` if a new variant was added. Added a
`default: const _exhaustive: never = type; throw new Error(...)`
clause so the build fails loudly on an unhandled variant.

**Minor 3 — `tree.rootNode.text` in `grpc-patterns/node.ts`.** Already
fixed in `2f28bfc` via a dedicated structural query
(`LOAD_PACKAGE_DEFINITION_SPEC`). No action needed.

## New: pipeline flow chart (per @magyargergo's request)

Added `src/core/group/PIPELINE.md` with four Mermaid diagrams:
1. **High-level overview** — `group.yaml` → extractors + manifest →
   contract matching → `bridge.lbug` → `runGroupImpact`.
2. **Per-repo extractor two-strategy shape** — graph-assisted
   Strategy A vs. source-scan Strategy B.
3. **Plugin architecture** — orchestrator → registry →
   per-language `*-patterns/<lang>.ts` → `tree-sitter-scanner.ts` →
   `ExtractedContract`.
4. **Manifest extraction** — label-scoped `resolveSymbol` with the
   synthetic-uid fallback.
5. **Cross-impact query (#606)** — local impact → bridge join →
   cross-repo fan-out.

Each diagram is annotated with which PRs own which stage (this PR:
extractors + manifest; #795: bridge storage; #606: cross-impact
runtime) and points at the concrete files/functions involved.

## Tests

- 99/99 extractor tests pass
- `npx tsc -p tsconfig.json --noEmit` clean

Co-authored-by: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-04-13 08:49:30 +01:00
Gergő Magyar b10d25bbca chore: release v1.6.0 — update CHANGELOG and package-lock (#798) 2026-04-12 12:31:21 +01:00
CopilotandGergo Magyar a94d6ef80b Extract registries into model/ module with SemanticModel interface (#786)
* Initial plan

* feat(SM-20): extract registries into model/ module with SemanticModel interface

- Create model/type-registry.ts — TypeRegistry interface + factory
- Create model/method-registry.ts — MethodRegistry interface + factory
- Create model/field-registry.ts — FieldRegistry interface + factory
- Create model/semantic-model.ts — SemanticModel interface + factory
- Create model/heritage-map.ts — re-export HeritageMap types
- Create model/binding-accumulator.ts — re-export BindingAccumulator types
- Create model/resolve.ts — move lookupMethodByOwnerWithMRO from call-processor
- Update symbol-table.ts — delegate to SemanticModel for registry ops
- Update call-processor.ts — re-export lookupMethodByOwnerWithMRO from model/resolve

No circular dependencies: model/resolve.ts does NOT import resolution-context.ts.
All 775 related unit tests pass with no regressions.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/27ad2975-1a31-4f50-815b-178ee8a95277

* fix: clarify re-export comment per code review feedback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/27ad2975-1a31-4f50-815b-178ee8a95277

* refactor(SM-20): wire up SemanticModel as first-class resolution input

PR #786 extracted TypeRegistry/MethodRegistry/FieldRegistry into model/
behind SemanticModel, but consumers still routed through SymbolTable
delegates. This change completes Phase 6 of the fuzzy-lookup elimination
roadmap by making call-processor, resolution-context, type-env, and
heritage-map query the model directly via `table.model.{types,methods,fields}`.

Also absorbs the open PR #786 review findings so the branch lands clean:
- Removed duplicate JSDoc block on lookupMethodByOwner (symbol-table.ts)
- Added model/index.ts barrel for the public model/ surface
- Fixed O(n) buildParentMapFromHeritage BFS via head-pointer queue
- Clarified re-export facade framing on binding-accumulator.ts and
  heritage-map.ts inside model/
- Refined @internal JSDoc on lookupMethodByOwnerWithMRO

Changes:
- symbol-table.ts: expose `readonly model: SemanticModel` on the
  SymbolTable interface. SymbolTable delegate wrappers (lookupClassByName
  etc.) stay as thin pass-throughs for backward compat; deletion is a
  follow-up once all internal callers are migrated.
- model/resolve.ts: lookupMethodByOwnerWithMRO now takes SemanticModel
  instead of SymbolTable, removing the last SymbolTable import from the
  model/ module. Preserves circular-dependency firewall.
- call-processor.ts: 6 call sites in D0 member resolution, field
  resolution, ctor override, and ctor disambiguation migrated to
  model.types/methods/fields.
- resolution-context.ts: tier 3 class+impl lookup migrated.
- type-env.ts: 5 sites across lookupClassDefsByName, resolveFieldType,
  and resolveMethodReturnType migrated.
- heritage-map.ts: parent/child class-name resolution migrated.

Tests:
- symbol-table.test.ts: +10 parity and feeding-audit tests covering
  every model.{types,methods,fields} path (Class, Method, Property,
  Impl, Function-with-ownerId, Property-without-ownerId skip, arity
  filtering, clear cascade).
- call-processor.test.ts: classLookupSpy now targets
  ctx.symbols.model.types since the wrapper is bypassed.
- type-env.test.ts: createMockSymbolTable and the destructured-call
  makeSymbolTable helpers gained a model shim that forwards to the
  (possibly overridden) top-level lookup stubs.

Validation: full suite 5603 passed / 159 skipped, resolver integration
suite (19 files, 1766 tests) clean, tsc --noEmit clean.

* refactor(SM-21): invert ownership — SemanticModel contains SymbolTable

Follow-up to SM-20. Previously SymbolTable owned a `model` subfield;
this commit turns the ownership direction around so the SemanticModel
is the top-level container and SymbolTable is nested as `.symbols`:

    SemanticModel (top-level, passed everywhere)
      ├── types   (TypeRegistry)
      ├── methods (MethodRegistry)
      ├── fields  (FieldRegistry)
      └── symbols (SymbolTable — file-indexed + callable-name index)

The owner-scoped registries live directly on the model; file and
callable-name lookups go through `.symbols`. Consumers receive a
`SemanticModel` and reach into the appropriate field — no more
`table.model.types.X` double-hop.

Core changes:
- symbol-table.ts: createSymbolTable now takes injected
  TypeRegistry/MethodRegistry/FieldRegistry via a SymbolTableDeps
  argument. When omitted (test fallback), it creates standalone
  registries locally and clears them in clear() — production callers
  always inject. The five registry convenience delegates
  (lookupClassByName, lookupMethodByOwner, lookupFieldByOwner,
  lookupClassByQualifiedName, lookupImplByName) remain as thin
  forwards to the injected registries so standalone SymbolTable use
  (chiefly tests) stays ergonomic.
- model/semantic-model.ts: createSemanticModel() now creates the
  three registries AND a SymbolTable wired to them, exposing the
  SymbolTable as `.symbols`. clear() cascades through all four.
- resolution-context.ts: `readonly symbols: SymbolTable` field is
  replaced with `readonly model: SemanticModel`. Internal factory
  builds a SemanticModel and keeps a local `symbols` alias for
  backward-compatible inner body.

Consumer migrations (src/):
- call-processor.ts: ctx.symbols.add/.lookupExactAll/
  .lookupCallableByName → ctx.model.symbols.*; ctx.symbols.model.X →
  ctx.model.X. buildTypeEnv option key renamed symbolTable → model.
- type-env.ts: symbolTable parameter renamed model (type
  SemanticModel), all internal call sites rewritten to use
  model.types.*, model.methods.*, model.fields.*,
  model.symbols.lookupExactAll / .lookupCallableByName.
- heritage-map.ts: 2 class-lookup sites migrated.
- pipeline.ts: ctx.symbols → ctx.model.symbols throughout.

Test migrations:
- symbol-table.test.ts: parity tests (which validated the old
  table.model.X hop) replaced with direct SemanticModel coverage via
  createSemanticModel(). New tests exercise types/methods/fields/
  symbols feeding end-to-end.
- type-env.test.ts: createMockSymbolTable rebuilt as a
  SemanticModel-shaped mock that still accepts the legacy flat
  override bag for backward compat; inline `makeSymbolTable` helpers
  for destructured-call and importedReturnTypes suites rewritten to
  match the new shape; buildTypeEnv options `symbolTable: X` and
  `{ symbolTable }` shorthand renamed to `model:`; one real
  createSymbolTable-based test rewritten to use createSemanticModel.
- call-processor.test.ts, heritage-map.test.ts,
  heritage-processor.test.ts, symbol-resolver.test.ts: bulk sed
  `ctx.symbols.` → `ctx.model.symbols.`. call-processor.test.ts spy
  updated to target `ctx.model.types.lookupClassByName`.

Validation: full test suite 5589 passed / 169 skipped / 0 failed;
tsc --noEmit clean; pre-commit eslint + prettier + typecheck all
green. CLAUDE.md / AGENTS.md stats bumped from an earlier `npx
gitnexus analyze` refresh (3965 symbols / 10012 edges / 243 flows).

* refactor(SM-22/SM-23): dispatch table + DAG rearchitecture

SM-22: Extract registration dispatch table into model/registration-table.ts.
Replaces the if/else ladder inside SymbolTable.add() with an O(1)
Map<NodeLabel, RoutingDecision> fan-out. SemanticModel wires the table
per-instance so hooks close over the correct registries.

SM-23: DAG rearchitecture. symbol-table.ts is now a pure 2-index leaf
(fileIndex + callableByName) with zero imports from model/. All
type/method/field routing lives in the model/ layer. Tests migrated to
createSemanticModel() + model.symbols access pattern.

Tests: 5632 passed, 0 failures.

* refactor: delete dead code (skipCallableIndex + model/ facades)

Removes the unused skipCallableIndex flag from the registration dispatch
table and deletes two facade files that had zero consumers.

skipCallableIndex was declared on RoutingDecision and populated for all
10 entries but never read at runtime — semantic-model.ts explicitly
documented that the flag was NOT consulted. The callable-index gate
lives inside SymbolTable.add() via CALLABLE_TYPES.has(type), which is
the single source of truth. Deleting the flag keeps SymbolTable as the
sole decision point and removes documentation-as-data.

model/binding-accumulator.ts and model/heritage-map.ts were facade
pass-throughs of their parent-directory counterparts. Grep confirms no
consumer imports either from the model/ path — all usage goes through
../binding-accumulator.js and ../heritage-map.js directly. model/index.ts
was the only "user" and re-exported them with a note about unifying the
import boundary, but that boundary has no actual consumers today.

Resolves review findings M-01 and M-03 from
.context/compound-engineering/ce-review/20260411-144641-59605d93/maintainability.json

Tests: 5631 passed, 0 failures (1 less than pre-Unit-1: the
skipCallableIndex-specific assertion was removed).

* refactor: remove lookupMethodByOwnerWithMRO backward-compat shim

call-processor.ts re-exported lookupMethodByOwnerWithMRO from
./model/resolve.js as a backward-compat shim for symbol-table.test.ts.
The function already lives in model/resolve.ts and is re-exported
properly from model/index.ts (the barrel) — the call-processor shim
was a duplicate export path with no durable reason to exist.

Migrated the test import from call-processor.js to model/index.js
(the canonical barrel). Deleted the re-export statement and the stale
"re-exported for backward compatibility" comment block. Hoisted the
remaining import to the top of the file with the other imports; the
bottom-of-file position was a relic of the shim pattern.

Resolves review finding M-02 from
.context/compound-engineering/ce-review/20260411-144641-59605d93/maintainability.json

Tests: 5631 passed, 0 failures.

* refactor: harden registration dispatch runtime safety

Two hardening changes in semantic-model.ts, both closing silent-failure
paths in the SM-series dispatcher-bypass failure mode.

1. model.symbols.clear() now cascades to the owner-scoped registries.
   Previously, the SymbolTable facade exposed rawSymbols.clear directly,
   which only emptied fileIndex + callableByName — the types/methods/
   fields registries stayed populated. Any caller holding a SymbolTable
   reference that invoked .clear() left the model in a split state where
   subsequent .add() calls double-registered in the registries. No
   current caller exercises this path, but it was a latent phantom-
   resolution risk that didn't belong in a public API. Extracted the
   cascade into a single cascadeClear closure wired into both
   model.clear() and the facade's clear field.

2. runExhaustivenessGuard now throws instead of console.warn on drift.
   The production short-circuit via NODE_ENV === 'production' is
   preserved, so real users never see the throw — but CI and dev runs
   now fail loudly if a NodeLabel is added to gitnexus-shared without
   being placed in one of the three registration-table allowlists. The
   previous warn-only behavior was silent in test output volume; SM-19
   already documented dispatcher-bypass as the dominant silent-failure
   mode in this codebase.

Test-first: added test/unit/model/semantic-model.test.ts covering
model.symbols.clear() cascade (4 registries × clear = 4 tests), the
existing model.clear() cascade (regression guard), and a happy-path
construction test that verifies the current allowlists have zero drift.

Resolves correctness P2 finding (symbols.clear() partial clear),
correctness P3 (exhaustiveness warn-only), and kieran-typescript KT-03
(same exhaustiveness finding, agreement boost).

Tests: 5638 passed (+7 new), 0 failures.

* docs: fix stale JSDoc references in resolveStaticCall

call-processor.ts:2215-2216 referenced SymbolTable.lookupClassByName
and SymbolTable.lookupMethodByOwner via {@link}. Both methods were
removed from SymbolTable during SM-20 — they now live on TypeRegistry
and MethodRegistry respectively, accessible via model.types and
model.methods.

Other SymbolTable.* references in the codebase (lookupExactFull, add,
lookupCallableByName in call-processor.ts:593, symbol-table.ts:86,
type-extractors/types.ts:57) target methods that are still on
SymbolTable and remain valid.

Resolves correctness P3 and kieran-typescript KT-02 (same finding,
agreement boost).

* refactor: deduplicate ALL_NODE_LABELS constant

ALL_NODE_LABELS was private in semantic-model.ts and duplicated
verbatim in registration-table.test.ts. Two hardcoded lists meant a
new NodeLabel added to gitnexus-shared could land in one copy but not
the other, silently drifting the exhaustiveness invariant.

Exported ALL_NODE_LABELS from semantic-model.ts, re-exported through
model/index.ts for barrel consistency, and switched the test to import
it instead of redeclaring. The explanatory comment now describes the
single-source-of-truth contract.

Resolves maintainability M-04.

Tests: 5638 passed, 0 failures.

* refactor: add compile-time NodeLabel exhaustiveness check

The runtime exhaustiveness guard in semantic-model.ts caught drift at
test time. Added a type-level check in registration-table.ts that
catches drift at BUILD time — if a new NodeLabel is added to
gitnexus-shared without being classified into one of the three
allowlists, TypeScript fails the _exhaustiveCheck assignment and
names the missing label.

The runtime guard stays as belt-and-suspenders: if a future contributor
bypasses the type check with @ts-ignore, the runtime guard still fires
in dev/test.

Implementation: converted the three allowlist Set<NodeLabel> initializers
to use `as const` tuples, then derived a union type from the tuples and
asserted `Exclude<NodeLabel, union> extends never`. Zero runtime impact
— the exported Sets are unchanged, Map.get hot-path performance is
unchanged, the test API is unchanged.

Resolves kieran-typescript KT-04.

Tests: 21/21 registration-table tests pass with zero modifications.

* refactor(test): restore type safety to createMockSymbolTable

createMockSymbolTable was widened to (overrides: any = {}): any with an
eslint-disable-next-line, and every buildTypeEnv call site passed the
mock as `model: mockSymbolTable as any`. The widening masked silent
false-green tests: buildTypeEnv accesses model.types/methods/fields,
and a flat any-typed override could silently return undefined from a
path that TypeScript should have caught at compile time.

Defined LegacyMockOverrides interface with typed stubs for each method
the mock can override (SymbolTable reads + TypeRegistry/MethodRegistry/
FieldRegistry lookups). Return type is now SemanticModel, so the mock
object is compile-checked against the real interface — a missing
registry method is a type error, not a silent runtime undefined.

Removed the eslint-disable and all 9 `as any` casts at call sites
(lines 1287, 1300, 1307, 2124, 2138, 5823, 5835, 5850, 5870). The
mock's return value now flows through buildTypeEnv's typed `model`
option without coercion.

Resolves kieran-typescript KT-01 and testing gap TG-02. This was the
highest-value cleanup in the plan — the only finding representing real
hidden test weakness.

Tests: 360 passed | 7 skipped (type-env.test.ts), typecheck clean.

* test: close coverage gaps in model/ registries

Added direct unit tests for the three owner-scoped registries that
previously had only transitive coverage via symbol-table.test.ts and
registration-table.test.ts. These new tests pin behaviors that were
flagged by the testing reviewer as untested or undertested.

method-registry.test.ts (14 tests):
- T-01: arity-fallback branch — when argCount matches no overload,
  fall back to the full pool so fuzzy resolution still has candidates.
  Previously untested and would have returned undefined instead of
  a valid candidate if the branch regressed.
- T-02: requiredParameterCount range filtering — methods with default
  parameters accept any argCount in [requiredParameterCount,
  parameterCount]. Previously untested at the registry level.
- Variadic fallback (parameterCount=undefined is retained during arity
  narrowing, bypassing range check).
- Return-type dedup paths: shared returnType → first wins, differing
  returnTypes → undefined, firstReturnType=undefined → undefined,
  single-overload skips dedup entirely.

type-registry.test.ts (9 tests):
- classByName homonym accumulation (two User classes in different
  packages both returned).
- classByQualifiedName disambiguation — same simple name, different
  FQNs resolve independently.
- Partial classes with identical simple + qualified name accumulate
  in both indexes.
- registerImpl stores Rust impl blocks separately from classes.
- Multiple impl blocks per type accumulate.

field-registry.test.ts (6 tests):
- register/lookup round-trip, owner-scope isolation, last-wins on
  duplicate key (flat map, not overload list).
- clear + re-register round-trip.

Extended symbol-table.test.ts cascade test (renamed from "both
registries" to "all three registries and the nested symbol table") to
also assert model.methods and model.fields are cleared — the test
name previously implied full coverage but only asserted types + symbols.

Resolves testing findings T-01, T-02, T-03, T-05.

Tests: 5667 passed (+29 new), 0 failures.

* refactor(test): replace brittle reference-equality tests + add intent comments

Two cleanups flagged as low-severity P3 by the testing reviewer:

1. registration-table.test.ts: Replaced three reference-equality tests
   (hook identity via toBe) with behavioral tests that survive a future
   refactor to per-label closures. The new "class-like behavior group"
   describe iterates Class/Struct/Interface/Enum/Record/Trait and
   verifies each one writes to types.registerClass. Same pattern for
   Method/Constructor. A separate "behavior group isolation" describe
   verifies class-like hooks don't leak into methods/fields and Impl
   never pollutes registerClass. Strictly more coverage than the
   reference-equality tests provided and implementation-independent.

2. symbol-resolver.test.ts: Added a comment above the lookupExactFull
   and SM-16: getFiles() describes explaining why they intentionally
   use createSymbolTable() directly instead of createSemanticModel().
   The DAG leaf-only behaviors they test do not involve registries, so
   testing the bare SymbolTable keeps the unit isolated. Prevents a
   future reader from "fixing" the inconsistency.

3. qualified-class-lookups.test.ts: Added a comment above
   `const symbolTable = model.symbols` explaining that processParsing
   writes still reach the owner-scoped registries via SemanticModel's
   fan-out — the alias is convenience, not a leaf in isolation.

Resolves testing T-04, kieran-typescript KT-05, kieran-typescript KT-06.

Tests: affected files all green (112 passed in registration-table +
symbol-resolver + qualified-class-lookups).

* refactor(model): collapse RoutingDecision wrapper and trim barrel surface

Two cleanups against the advanced-review findings on post-Unit-9 state:

S2 (cross-reviewer agreement — architecture-strategist + code-simplicity):
Delete the RoutingDecision single-field wrapper interface. Post-Unit-1
it held exactly one field (hook: RegistrationHook) and added pure
ceremony at every call site — `dispatchTable.get(key)!.hook(name, def)`
vs the now-direct `dispatchTable.get(key)!(name, def)`. Change the Map
type from Map<NodeLabel, RoutingDecision> to Map<NodeLabel,
RegistrationHook>, drop the interface, and update 17 test call sites.

A3 (architecture-strategist): Trim model/index.ts barrel surface.
createRegistrationTable, RegistrationHook, and RegistrationTableDeps
were re-exported from the barrel despite having zero legitimate
consumers outside model/ itself. The only callers (semantic-model.ts
and registration-table.test.ts) import directly from
./registration-table.js. Barrel exposure invited external callers to
construct orphan dispatch tables with independent registries,
weakening the SM-21 ownership inversion where SemanticModel is the
composition root. Kept CALLABLE_ONLY_LABELS, INERT_LABELS,
DISPATCH_LABELS exported since those remain useful for downstream
resolution logic and have no construction risk.

Resolves review findings:
- S2 (code-simplicity P3, 0.85) + architecture-strategist residual
- A3 (architecture-strategist P3, 0.82)

Tests: 5674 passed, 0 failures. Typecheck clean.

* refactor(model): replace runtime exhaustiveness guard with compile-time bijection

Replace the three-layer drift protection (hardcoded ALL_NODE_LABELS
array + 3 tuple consts + _ExhaustiveLabelCheck type + runExhaustivenessGuard
runtime + CI taxonomy test) with a single Record<NodeLabel, LabelBehavior>
map that structurally proves every invariant at compile time.

## Before

- ALL_NODE_LABELS hardcoded in semantic-model.ts (36 entries, could drift)
- DISPATCH_LABELS_TUPLE / CALLABLE_ONLY_LABELS_TUPLE / INERT_LABELS_TUPLE
  private tuples (36 more entries total, could overlap or miss)
- _ClassifiedLabel / _UncoveredLabel type-level check (caught missing
  labels but NOT duplicates across tuples)
- runExhaustivenessGuard runtime throw (only defense against duplicates)
- NodeLabel taxonomy coverage test in CI (same check as runtime guard)

Four defenses for invariants that the type system can express directly.

## After

```ts
type LabelBehavior = 'dispatch' | 'callable-only' | 'inert';

const LABEL_BEHAVIOR = {
  Class: 'dispatch',
  // ...36 entries...
  Tool: 'inert',
} as const satisfies Record<NodeLabel, LabelBehavior>;
```

The `as const satisfies Record<NodeLabel, LabelBehavior>` combo enforces:

1. **Every NodeLabel must be a key** — Record requires all K keys.
   Adding a NodeLabel to gitnexus-shared without classifying it here
   fails with "Property 'X' is missing in type ..." naming the drifted label.
2. **No non-NodeLabel keys allowed** — `satisfies` with object literals
   triggers excess-property checking. A typo'd key fails to compile.
3. **No duplicate classification** — impossible by construction; object
   keys are unique at the source level.
4. **Valid category** — LabelBehavior is a narrow union, typos caught.

`ALL_NODE_LABELS`, `DISPATCH_LABELS`, `CALLABLE_ONLY_LABELS`, and
`INERT_LABELS` are now derived via `Object.keys(LABEL_BEHAVIOR)` and
`filter(l => LABEL_BEHAVIOR[l] === ...)` — single source of truth,
structurally impossible to drift.

## Deleted

- runExhaustivenessGuard() function in semantic-model.ts (~18 lines)
- ALL_NODE_LABELS hardcoded array in semantic-model.ts (~38 lines)
- DISPATCH_LABELS_TUPLE / CALLABLE_ONLY_LABELS_TUPLE / INERT_LABELS_TUPLE
  private consts in registration-table.ts (~30 lines)
- _ClassifiedLabel / _UncoveredLabel / _exhaustiveCheck type machinery
  (~20 lines)

## Kept named proofs: none

The `as const satisfies` on the object literal already catches all four
drift modes. Named type-level proofs (_MissingFromMap / _ExtraKeysInMap)
are pure duplication and were removed per review.

## Also in this commit

- S6: trim wrappedAdd narration comments in semantic-model.ts
  (Step 1/2/3 block comments removed; kept the Function+ownerId WHY note)
- A3: tighten model/index.ts barrel — createRegistrationTable,
  RegistrationHook, RegistrationTableDeps remain direct-imports only;
  ALL_NODE_LABELS and LabelBehavior re-exported from the new home in
  registration-table.ts

## Resolves

- Advanced-review S4 (runtime guard per-call cost) — guard no longer exists
- Advanced-review S1 (tuple three-defenses indirection) — single Record replaces all tuples
- Correctness P3 (exhaustiveness warns-only) — structurally impossible to drift
- Unit 6 type-level check — subsumed by the Record type
- Unit 3 runtime throw — no longer needed

Tests: 5674 passed, 0 failures. Typecheck clean.

* test(model): delete duplicate closure-isolation spy tests

S5 (code-simplicity P3): The 'closure isolation — each hook can only
write to its registry' describe block duplicated the 'behavior group
isolation' block's coverage via a different mechanism.

Behavioral tests (lines 151-174, kept):
  table.get('Class')!('User', def);
  expect(deps.methods.lookupMethodByOwner('unrelated', 'User')).toBeUndefined();
  expect(deps.fields.lookupFieldByOwner('unrelated', 'User')).toBeUndefined();

Spy tests (deleted, ~55 lines):
  vi.spyOn(deps.methods, 'register')
  table.get('Class')!('User', def);
  expect(methodsSpy).not.toHaveBeenCalled();

Both assert the same invariant — classHook does not touch the methods or
fields registries. The behavioral form observes the END STATE of the
registry (lookup returns undefined), which is the actual contract.
The spy form asserts the IMPLEMENTATION (a specific method was not
called), which couples to internal wiring — a refactor to a different
register function name would break the spy test while the behavioral
test would still pass.

Also dropped the now-unused `vi` import from vitest.

Tests: 24/24 registration-table.test.ts pass (-4 from spy deletion).

* refactor(model): compile-time cross-invariant between CLASS_TYPES and dispatch classHook

A1 (architecture-strategist P2, 0.90): CLASS_TYPES in symbol-table.ts
and the class-like entries of the dispatch table were two independent
hardcoded sets. Adding a new class-like label (e.g. Swift 'Extension')
to one but not the other would silently degrade qualifiedName
population — the symptom is subtle (partial qualified-name lookups)
and no test asserted the co-extensive invariant.

Fixed with a single source of truth and a two-layer compile-time
enforcement:

## symbol-table.ts

- Add `CLASS_TYPES_TUPLE` as `readonly [...] as const satisfies
  readonly NodeLabel[]`. The `satisfies` forces every tuple entry to
  be a valid NodeLabel at compile time.
- Export derived type `ClassLikeLabel = typeof CLASS_TYPES_TUPLE[number]`.
- Derive `CLASS_TYPES` Set from the tuple — same runtime shape as
  before, now typed `ReadonlySet<NodeLabel>`.

## registration-table.ts

- Import `CLASS_TYPES_TUPLE` and `ClassLikeLabel` from symbol-table.ts.
- Narrow the `satisfies` on `LABEL_BEHAVIOR` via intersection:
      Record<NodeLabel, LabelBehavior> & Record<ClassLikeLabel, 'dispatch'>
  This forces every class-like label to have value 'dispatch' at
  compile time. Adding a label to CLASS_TYPES_TUPLE without
  classifying it as dispatch in LABEL_BEHAVIOR fails to compile with
  a type error naming the drifted label.
- Build the class-like entries of the dispatch Map by iterating
  `CLASS_TYPES_TUPLE` at factory time. Adding a label to the tuple
  automatically wires it to classHook — no second place to update.

## What the design prevents

1. Drift scenario A (A1 original): 'Extension' added to CLASS_TYPES_TUPLE
   but not to LABEL_BEHAVIOR → compile error on LABEL_BEHAVIOR's
   satisfies.
2. Drift scenario B: 'Extension' added to CLASS_TYPES_TUPLE but not
   wired to classHook → impossible because the Map is derived from the
   tuple.
3. Drift scenario C: class-like label classified as something other
   than 'dispatch' in LABEL_BEHAVIOR → compile error on the narrowed
   intersection.

Runtime behavior unchanged: same 6 labels in CLASS_TYPES, same 6
class-like entries in the dispatch Map. Tests pin the behavior via
the existing behavior-group tests in registration-table.test.ts.

DAG unchanged: registration-table.ts already imported from symbol-table.ts
(the allowed upward direction). symbol-table.ts still imports nothing
from model/.

Tests: 5670 passed, 0 failures. Typecheck clean.

* test(field-extraction): use SemanticModel facade instead of raw SymbolTable

A6 (architecture-strategist P3, 0.85): field-extraction.test.ts created
its FieldExtractorContext fixture with `symbolTable: createSymbolTable()` —
a raw SymbolTable leaf, not the facade. In production, the context's
symbolTable field is always `model.symbols` (the SemanticModel-wrapped
facade where .add() dispatches through the owner-scoped registries).

The current field extractors don't call symbolTable.add() at all, so
this change is behavior-neutral today. The value is architectural
consistency — matching the test fixture to the production shape
prevents silent drift if a future field extractor starts registering
dynamically-discovered properties via the context. Without the fix,
such writes would hit the raw leaf and skip the fan-out, and tests
would pass even though the symptom (empty FieldRegistry) would
manifest in production.

Tests: 50/50 field-extraction.test.ts pass. Production tsc --noEmit
clean. Test-tsconfig error count unchanged (634 pre-existing errors
in unrelated test files, out of scope).

* refactor(A5): decouple model/resolve.ts from language registry

Move the MroStrategy type into gitnexus-shared and replace the
language: SupportedLanguages parameter on lookupMethodByOwnerWithMRO
with a direct mroStrategy: MroStrategy literal. Callers derive the
strategy from their language provider before invoking the resolver.

model/resolve.ts no longer imports from ../languages/index.js, so the
model/ layer is free of cross-layer coupling with the language
registry — this closes finding A5 from the SM-20/21/22/23 advanced
review (plan 006).

* feat(A4): add MethodRegistry.lookupMethodByName flat-by-name index

Add a secondary `methodsByName: Map<string, SymbolDefinition[]>` index
on MethodRegistry that returns every method with a given unqualified
name, accumulated across owners and overloads. The new index shares
SymbolDefinition references with methodByOwner — no duplication.

This is step 1 of the A4 double-index removal (plan 006). Tier 3
global resolution will switch to this index in Unit 3 so Method and
Constructor can be removed from CALLABLE_TYPES in Unit 4.

* refactor(A4): extend Tier 3 + memberCallByFile to consult method registry

Add model.methods.lookupMethodByName to Tier 3 global resolution in
resolution-context.ts and to the callable-pool build in
call-processor.ts (resolveMemberCallByFile + D2 widen path).

Intentionally behavior-preserving: Method and Constructor are still
in CALLABLE_TYPES so the new lookup returns identical candidates that
already reach Tier 3 through callableByName. Both paths dedup by
nodeId during this intermediate state — Unit 4 shrinks CALLABLE_TYPES
and the dedup is removed.

Part of plan 006 A4 step 2.

* refactor(A4): shrink CALLABLE_TYPES to free callables only

CALLABLE_TYPES = {Function, Macro, Delegate}. Method and Constructor
are no longer double-indexed in callableByName — they reach resolvers
through model.methods.lookupMethodByName instead.

Companion changes:
- Introduce CALL_TARGET_TYPES = CALLABLE_TYPES ∪ {Method, Constructor}
  for the resolver's kind filter (filterCallableCandidates,
  countCallableCandidates). Separates registration semantics (narrow)
  from the resolver's acceptable-target set (wide).
- type-env.ts for-loop return-type inference consults both indexes,
  treating the union as the authoritative call pool.
- resolveMemberCallByFile + D2 widen path keep the nodeId dedup in
  place: Python/Rust/Kotlin class methods emitted as Function+ownerId
  still land in both indexes until Unit 5 unblocks the normalization.
- Tier 3 global resolution (resolution-context.ts) keeps the same
  dedup for the same reason.

Test updates reflect the new contract: Method/Constructor live in
methodsByName, not callableByName. Orphan Method-without-ownerId now
lives only in the file index (no registry coverage).

Part of plan 006 — closes A4 for strictly-labeled methods. Python/
Rust/Kotlin Function+ownerId normalization is tracked as Unit 5
(blocked).

* refactor: rename CALLABLE_TYPES → FREE_CALLABLE_TYPES

Pure rename. The constant's meaning changed in Unit 4 (free callables
only — no methods, no constructors) so the name now reflects that
scope: "callables that have no owner scope". Updates the constant
declaration and every consumer in src/ and test/.

Closes plan 006 Unit 6.

* refactor(A2): strict SymbolTableReader (pure reads) + SymbolTableWriter (+add)

Split the SymbolTable interface into three strictly layered surfaces:

- SymbolTableReader: lookups + iteration. NO add, NO clear. Holders
  cannot mutate the table in any way.
- SymbolTableWriter extends Reader: + add. NO clear. Holders can
  register new symbols but cannot trigger a leaf-index reset.
- InternalSymbolTable (private, not exported): + clear. The cascading
  reset capability is reachable only through createSymbolTable's
  return type, held exclusively by SemanticModel.rawSymbols.

SemanticModel.symbols is now typed as SymbolTableWriter — external
consumers (workers, processors, pipelines) can register symbols and
query them, but cannot reach .clear(). The A2 LSP fix holds: callers
holding any public reference cannot desync the leaf indexes from the
owner-scoped registries.

Delete the transitional `type SymbolTable = SymbolTableReader` alias
and migrate every consumer (src + test) to the explicit names:
- Field and parameter annotations use SymbolTableReader by default;
  only code that calls .add() uses SymbolTableWriter.
- parsing-processor (workers + sequential paths) takes
  SymbolTableWriter so it can register extracted symbols.
- field-types, call-processor, named-binding-processor,
  workers/parse-worker: use SymbolTableReader (query-only).
- Tests: drop the stale `clear` fields from mock factories and
  migrate the semantic-model cascade tests from the removed
  model.symbols.clear() path to model.clear().

Closes plan 006 Unit 7. Industry sources: TypeScript compiler API
builder pattern, Salsa ParallelDatabase, .NET IReadOnlyList. See the
a2-lsp-clear-contract-research artifact for full citations.

* feat(A2): add SemanticModel.resetFileIndex() partial-reset entry point

Add a named method that clears only the leaf file and callable
indexes without cascading to the three owner-scoped registries
(types, methods, fields). Replaces the rare partial-reset use case
that was previously reachable via the now-removed symbols.clear()
path from A2 (plan 006 Unit 7).

JSDoc makes the semantic difference with model.clear() explicit so
future readers don't have to guess which method to call for a given
reingestion scenario.

Test-first: three scenarios cover the partial-vs-full semantics,
re-add after reset, and idempotency.

Closes plan 006 Unit 8.

* docs(S7): trim registration-table module JSDoc

Remove the ~24 lines of design-provenance citations from the module
JSDoc. The rust-analyzer, TypeScript-compiler, and Fowler references
are preserved in git history via the original SM-22 commits and in
plan 006 Unit 9.

Keep the ownership diagram, behavior-group table, and the
'How to add a new NodeLabel' checklist — those are load-bearing for
future contributors.

Closes plan 006 Unit 9 (S7 advanced-review finding).

* test(S3): migrate type-env.test.ts off LegacyMockOverrides

Replace the createMockSymbolTable bridge and LegacyMockOverrides
interface with real createSemanticModel() + add() calls across all
14 call sites. Where a test needs a specific registry lookup that
can't be pre-populated cleanly, use vi.spyOn on the real registry
instead.

Pattern breakdown:
- Pattern A (pre-populate via model.symbols.add): 13 sites
- Pattern B (vi.spyOn on registry lookup): 1 site

Deletes LegacyMockOverrides + createMockSymbolTable entirely. The
real MethodRegistry arity/returnType semantics match the hand-rolled
mock behavior in every migrated case, and no 'as any' casts remain
in the file.

Closes plan 006 Unit 10 (S3 advanced-review finding).

* refactor: remove unused MroStrategy type exports from language-provider and resolve modules

* refactor: relocate symbol-table, heritage-map, resolution-context into model/

Use git mv so blame and history follow each file:
- gitnexus/src/core/ingestion/symbol-table.ts → model/symbol-table.ts
- gitnexus/src/core/ingestion/heritage-map.ts → model/heritage-map.ts
- gitnexus/src/core/ingestion/resolution-context.ts → model/resolution-context.ts

These three files are part of the SemanticModel layer (file/callable
indexes, heritage parent map, tiered resolver) and now sit alongside
the registries they collaborate with. Updates every consumer import
path across src/ and test/ to the new locations.

* refactor(model): enforce pure-leaf DAG + delete legacy re-exports

model/ is now a pure leaf: zero upward imports and zero compat
shims in its parent processors. Completes the DAG cleanup started
in the previous commit.

1. walkBindingChain — moved into model/resolution-context.ts;
   named-binding-processor.ts deleted.

2. NamedImportMap + NamedImportBinding + isFileInPackageDir —
   moved into model/resolution-context.ts. Every consumer now
   imports from the canonical location directly. Legacy re-exports
   in import-processor.ts deleted.

3. c3Linearize + gatherAncestors — moved into model/resolve.ts.
   mro-processor.ts imports them back for computeMRO. Legacy
   c3Linearize re-export from mro-processor.ts deleted.

4. ExtractedHeritage type — moved into model/heritage-map.ts.
   call-processor.ts, parsing-processor.ts, pipeline.ts,
   heritage-processor.ts, and the test files now import it from
   the canonical location. Legacy re-exports in parse-worker.ts
   and heritage-processor.ts deleted.

5. resolveExtendsType — rewritten in model/heritage-map.ts to
   take an explicit HeritageResolutionStrategy (A5-style DI).
   buildHeritageMap accepts an optional getHeritageStrategy
   callback; production uses getHeritageStrategyForLanguage from
   heritage-processor.ts. Legacy resolveExtendsType re-export
   from heritage-processor.ts deleted.

Verified:
- grep 'from "..' gitnexus/src/core/ingestion/model → empty
- grep 'Re-export for legacy' gitnexus/src/core/ingestion → empty
- npx tsc --noEmit → clean
- npx vitest run → 5686 passing

* docs(model): strip phase/plan references from module comments

Remove SM-20/21/22/23, A2/A4/A5, plan 006, Unit N labels and historical
phrasing ("previously", "legacy", "model-leaf DAG cleanup") from all 10
files in src/core/ingestion/model/. Preserve domain vocabulary (Tier
1/2/3), invariants, and caveats — only the plan archaeology is gone.

* refactor(model): tighten interface segregation + compile-time invariants

Apply four gated findings from branch-wide code review:

- SemanticModel.symbols now typed as SymbolTableReader; MutableSemanticModel
  widens it back to SymbolTableWriter. ResolutionContext.model is typed as
  MutableSemanticModel since it owns the lifecycle. Resolvers that only
  query symbols can annotate their own fields as SemanticModel to drop
  write access at the type level.

- Lookup methods (lookupExactAll, lookupCallableByName, lookupClassByName,
  lookupClassByQualifiedName, lookupImplByName) now return
  readonly SymbolDefinition[]. The returned arrays are live views into
  the internal indexes; the readonly marker prevents accidental caller
  mutation. walkBindingChain return type narrowed to match.

- FREE_CALLABLE_TUPLE + FreeCallableLabel exported from symbol-table.ts
  as the single source of truth for free-callable labels. LABEL_BEHAVIOR
  now satisfies Record<FreeCallableLabel, 'callable-only'> as a second
  cross-invariant alongside Record<ClassLikeLabel, 'dispatch'>. Adding a
  label to the tuple without classifying it as 'callable-only' fails at
  build time. CALLABLE_ONLY_LABELS is now a re-export alias of
  FREE_CALLABLE_TYPES so the two sets cannot drift.

- walkBindingChain fast-exits before allocating its cycle-detection Set
  when the caller's file has no named bindings. Skips ~200k transient
  Set allocations per large-repo resolution pass.

Also fixes five stale comments flagged by the review: duplicate JSDoc
block on RegistrationHook merged; resolve.ts "delegates to mro-processor"
direction corrected; RegistrationTableDeps JSDoc names
createRegistrationTable (not createSymbolTable); mro-processor.ts
"re-exported at top" stale comment removed; gatherAncestors export
comment matches reality.

tsc --noEmit clean, full test suite green (5786 tests).

* refactor(model): resolve four deferred P2 review findings

Address the four gated items from the branch-wide review that needed
design decisions before applying:

F#3 — Method/Constructor without ownerId fallback to callable index.
The dispatch hook silently skips owner-scoped labels that lack an owner
(an extractor contract violation — AST-degraded parse, or a buggy
language extractor). Pre-dispatch-table code let such defs fall through
to callableByName and stay reachable at Tier 3 global resolution. This
restores that fallback in SymbolTable.add so orphaned Methods and
Constructors don't silently vanish. Property deliberately does NOT
participate in the fallback to avoid polluting common names like
id / name / type.

F#4 — Delete MutableSemanticModel.resetFileIndex. The method had zero
production callers (only three tests), documented a "rare partial-
reingestion flow" that was never implemented, and contained the
adversarial-reviewer's double-populate trap: calling resetFileIndex
followed by re-adding the same class symbol would push a duplicate
SymbolDefinition into TypeRegistry.classByName without ever clearing
the first one. If incremental reingestion is ever needed, it can be
designed properly with per-file TypeRegistry invalidation. For now,
deleting the footgun is safer than documenting it.

F#5 — Compile-time dispatch-table completeness check. `LABEL_BEHAVIOR`
already enforces "every NodeLabel is classified" via
`Record<NodeLabel, LabelBehavior>`, but the dispatch-table factory
populated its Map with manual `table.set(...)` calls that TypeScript
could not correlate back to the `'dispatch'` classification. Add a
type-level `DispatchLabel` extracted from `LABEL_BEHAVIOR` via a
conditional mapped type, and build the table from an object literal
that satisfies `Record<DispatchLabel, RegistrationHook>`. Adding a new
dispatch-classified label without wiring it to a hook now fails the
build with a named-key error — no more silent no-op hooks.

F#7 — Tier 3 dedup fast-path via MethodRegistry.hasFunctionMethods.
The Set-based dedup between callableDefs and methodDefs is only needed
when a Python/Rust/Kotlin class method (emitted as Function+ownerId by
the worker) lands in both indexes. For TS/Java/C#/C++/Ruby-only repos
— where the two indexes are disjoint by construction — the dedup was
pure overhead on every global-tier hit. MethodRegistry now tracks
whether any Function-typed def was ever registered, and resolution-
context branches Tier 3 into a concat-only fast path when that flag
is false. Slow path with dedup survives unchanged for mixed-language
repos.

New tests pin the invariants: hasFunctionMethods flag transitions,
Method/Constructor orphan fallback, Property non-fallback, and the
MethodRegistry clear() reset. Full test suite green (5756 tests).

* refactor(model): close remaining P3 review findings + coverage gaps

Address the remaining review items in one batch.

Production refactors:

- Rename classHook → classLikeHook (M05). The hook handles Class /
  Struct / Interface / Enum / Record / Trait; the vocabulary used in
  surrounding docs and the behavior-group table is "class-like". The
  rename makes the code match the taxonomy without forcing readers
  through a mental glossary.

- Extract MAX_BINDING_CHAIN_DEPTH constant in resolution-context.ts
  and document it as a known silent false-negative source (ADV-003).
  Five hops cover the common TypeScript monorepo pattern; raising the
  cap is a one-line change if a real repo exceeds it. walkBindingChain
  consumes the constant so the 5 magic number no longer floats free.

- Replace defs.filter() allocation in MethodRegistry.lookupMethodByOwner
  with a two-pass streaming count + conditional materialization
  (PERF-04). Pure-match and pure-reject arity paths now skip the
  filtered-array allocation entirely; only the discriminating case
  (at least one match AND at least one rejection) pays it.

- Rewrite NOOP_SYMBOL_TABLE in parse-worker.ts and NOOP_SYMBOL_TABLE_SEQ
  in parsing-processor.ts to implement all six SymbolTableReader
  methods (ADV-005). The `as unknown as SymbolTableReader` cast is
  removed in favor of a direct SymbolTableReader annotation, so future
  additions to the interface surface as compile errors on the stubs
  instead of silently falling through.

- type-env.ts getCallableUnionCount and getFirstCallable now take
  `model: SemanticModel` as an explicit argument instead of reaching
  into the enclosing `model!` non-null assertion (KT-003). Callers
  enter via an `if (model)` guard and pass the narrowed reference, so
  the non-null precondition is visible at the type level and the
  closures cannot be accidentally extracted into a context without
  the guard.

- Tier 3 dedup in resolution-context.ts now covers all four index reads
  (classDefs, implDefs, callableDefs, methodDefs) via a pushUnique
  helper (C-03). Previously classDefs and implDefs were spread directly
  without dedup; any theoretical nodeId collision would have produced
  duplicates in globalDefs.

Test infrastructure:

- Extract makeDef / makeMethod factory helpers into
  test/unit/model/helpers.ts (T-07). The four registry/table test
  files now import the shared helper and specialize with overrides,
  removing ~25 lines of duplicated boilerplate and creating a single
  point of maintenance.

New test coverage:

- T-01: c3 BFS fallback — cyclic Python hierarchy that fails c3
  linearization and must fall back to heritageMap.getAncestors() BFS
  order. Added to the lookupMethodByOwnerWithMRO describe block.

- T-02: Tier 2a-named precedence — verifies the binding chain walker
  fires before Tier 2a import-scoped when an aliased import
  `import { User as U } from B` competes with a raw same-name Tier 2a
  hit. Also pins Tier 1 same-file precedence over Tier 2a-named.

- T-03: Tier 3 Function+ownerId dedup — end-to-end test that a Python
  class method emitted as `Function + ownerId` yields exactly ONE Tier
  3 candidate (not two). Companion test pins the fast-path branch for
  hasFunctionMethods === false repos.

- T-06: walkBindingChain guards — circular re-export detection,
  depth-cap exceeded drop, and boundary case at exactly
  MAX_BINDING_CHAIN_DEPTH hops resolving successfully.

All tests added to a new test/unit/model/resolution-context.test.ts
dedicated to ResolutionContext.resolve() tier-precedence invariants.

Full suite: 5708 passing (minus the known Windows LBUG lock flake
that passes in isolation).

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-12 01:06:55 +01:00
ivkondandClaude 1ff324ca16 feat(group): bridge.lbug storage + contract matching expansion (1/4 of #606 split) (#795)
* feat(group): bridge.lbug storage + contract matching expansion

Part 1 of 4 in the split of #606 (ticket: #791, closes #790 with a
revised plan per @magyargergo's request).

## What changed

Adds the LadybugDB-backed bridge storage infrastructure and extends
the contract matching algorithm with wildcard support. All changes are
additive: storage.ts, sync.ts, service.ts, cli/group.ts, mcp/tools.ts
are left on their upstream main versions and will migrate to the new
bridge in follow-up PRs (#792, #793, #794).

### Files

**New (844 LOC prod):**
- `gitnexus/src/core/group/bridge-db.ts` — atomic write-to-temp with
  `retryRename` for Windows EBUSY/EPERM, per-item write tolerance via
  `WriteBridgeReport`, `findContractNode` with three-tier symbol
  lookup (uid → filePath+name → filePath)
- `gitnexus/src/core/group/bridge-schema.ts` — schema DDL
- `gitnexus/src/core/group/normalization.ts` — contract ID
  canonicalization + `dedupeContracts` / `dedupeCrossLinks` helpers
  used by both matching and bridge write

**Modified (+137 LOC prod):**
- `gitnexus/src/core/group/matching.ts` — adds `runWildcardMatch` for
  `grpc::Service/*` wildcard consumers, `buildProviderIndex` helper,
  and canonical gRPC ID handling in `normalizeContractId`
- `gitnexus/src/core/group/types.ts` — `MatchType` gains `'wildcard'`;
  new `BridgeHandle` and `BridgeMeta` interfaces

**New tests (658 LOC):**
- `gitnexus/test/unit/group/bridge-db.test.ts` — core write/read round
  trip, `WriteBridgeReport` shape, dropped-links counter, retryRename
  behavior on EBUSY/ENOENT/EPERM/EACCES
- `gitnexus/test/unit/group/bridge-db-edge.test.ts` — edge cases
  (malformed meta, missing contract nodes, concurrent access)

**Modified tests (+225 LOC):**
- `gitnexus/test/unit/group/matching.test.ts` — wildcard consumer
  matching, gRPC canonical ID handling, same-service guard

### Self-review fixes folded in

Carried forward from the original #606 self-review:
- `writeBridge` try/finally handle lifecycle + `handleClosed` sentinel
- `openBridgeDbReadOnly` partial-handle cleanup
- `writeBridgeMeta` uses `retryRename` for Windows consistency
- `retryRename` unit tests (was zero coverage)
- Per-item try/catch around every CREATE loop so one malformed contract
  doesn't abort the whole write
- Dropped cross-link counter (`linksDroppedMissingNode`)

### Why now

magyargergo asked for the #606 PR to be split so we can iterate with
confidence (https://github.com/abhigyanpatwari/GitNexus/pull/606#issuecomment-4229612271).
This is the foundational layer — pure infra, no user-facing surface,
no callers of the new APIs in this PR. Later PRs wire it in.

### How to verify

- `cd gitnexus && npx tsc --noEmit`
- `cd gitnexus && npx vitest run test/unit/group/bridge-db.test.ts --pool=forks`
- `cd gitnexus && npx vitest run test/unit/group/bridge-db-edge.test.ts --pool=forks`
- `cd gitnexus && npx vitest run test/unit/group/matching.test.ts --pool=forks`
- Pre-commit hook runs clean

### Risk / rollback

**Low.** All new code sits under `src/core/group/` in new files plus a
minimal `+16/-1` diff to `types.ts` and a `+136/-0` diff to `matching.ts`
(both purely additive). No existing callers reference the new APIs
(bridge-db, openBridgeOrFallback, runWildcardMatch) — the PRs that wire
them in come later in the split chain. Rollback = `git revert` of the
merge commit; no state introduced, no schema migration triggered.

### Scope discipline (per GUARDRAILS.md)

- Only the 8 files listed above are touched; no drive-by refactors
- No CI/release/security config changes
- No secrets, tokens, or machine-specific paths
- Content is lifted from the #606 branch which already passed CI 11/11
  green on `d15b8cb` (before the split)

### Dependencies

- **Base:** `main` (no dependencies on other split PRs)
- **Blocks:** extractor expansion (#792), sync pipeline (#793),
  cross-impact feature (#794)
- **Related ticket:** #791

Co-authored-by: Claude <noreply@anthropic.com>

* fix(group): address @claude review on #795

Addresses the findings from the automated review on PR #795
(https://github.com/abhigyanpatwari/GitNexus/pull/795#issuecomment-4229770000
— posted by @magyargergo / claude-code Action run).

### Medium severity (reviewer flagged as blockers)

- **bridge-db.ts `openBridgeDbReadOnly` bak recovery** — the `.bak`
  recovery path used bare `fsp.rename(bakPath, dbPath)`, which is
  exactly the scenario most likely to hit Windows EBUSY/EPERM (an
  interrupted writer still holding the handle for a few ms). Switched
  to `retryRename` for consistency with the rest of the file's
  Windows-safe rename path.
- **bridge-db.ts `ensureBridgeSchema` error detection** — the inline
  `msg.includes('already exists')` substring match has been lifted
  into a named constant `LBUG_ALREADY_EXISTS_MSG` with a comment
  documenting the coupling to LadybugDB's error message wording and
  why we can't use `IF NOT EXISTS` (LadybugDB DDL doesn't support it)
  or typed errors (LadybugDB's JS driver doesn't expose error codes).
  Also tightened the `catch (err: any)` to `catch (err: unknown)`.
- **bridge-db.ts `findContractNode` — extracted out of writeBridge**
  — the 35-line async closure living inside `writeBridge` has been
  lifted to three module-level functions: `createContractLookupIndex`,
  `indexContract`, and `findContractNode`. `findContractNode` is now
  a pure synchronous function taking a prebuilt index instead of
  doing its own DB queries. The `writeBridge` cross-link loop is now
  ~25 lines instead of ~100.
- **bridge-db.ts `findContractNode` — N+1 query elimination** — the
  old inner-closure version issued up to 6 DB round-trips per
  cross-link (2 endpoints × up to 3 tiers of fallback queries). For a
  group with 1000 cross-links, that's up to 6000 DB queries just to
  resolve endpoints. The new version consults an in-memory
  `ContractLookupIndex` built incrementally as contracts are inserted
  (`indexContract` called AFTER each successful insert so failed
  inserts don't poison the index). Cross-link resolution is now
  O(1) per link instead of O(3) DB queries per link, with zero DB
  round-trips during the cross-link loop.

### Minor severity

- **bridge-db.ts `queryBridge` empty-array guard** — if LadybugDB
  ever returns an empty `QueryResult[]` at the top level (shouldn't
  happen with single-statement calls, but driver contract isn't
  explicit), the old code would call `.getAll()` on `undefined` and
  crash with a confusing stack. Added an `unwrapQueryResult` helper
  that throws an explicit `'empty QueryResult array'` error instead,
  making a potential driver regression visible immediately.
- **normalization.ts `contractRichness` weights** — added a
  block-level comment documenting the weight ordering (+3 for
  symbolUid, +2 for each symbol-identifying field, +1 for service
  tag or non-manifest origin) and explicitly noting that the
  absolute numbers don't matter, only the relative ordering. Matches
  the "comment for contributors" suggestion in the review.
- **bridge-schema.ts `BRIDGE_SCHEMA_VERSION` migration comment** —
  added a 4-point contract explaining what bumping the constant
  means ("discard and re-sync" strategy for V1, no in-place
  migration yet, new migration logic should live in a separate
  `bridge-migrations.ts` module when it becomes necessary).
- **test/unit/group/fixtures.ts** — extracted the `makeContract`
  helper previously copy-pasted between `bridge-db.test.ts` and
  `bridge-db-edge.test.ts` into a shared fixtures module. Both test
  files now import from `./fixtures.js`. Kept the scope minimal:
  fixtures is NOT a general-purpose factory module, just the shared
  baseline contract builder.

### New tests

Added 9 pure-function unit tests for the now-extracted
`findContractNode` in `bridge-db.test.ts`:
  - returns null on empty index
  - tier 1 (symbolUid) match, including repo-scope and role-scope
    isolation
  - tier 2 (filePath + symbolName) fallback when symbolUid is empty
    or mismatches
  - tier 3 (filePath only) when exactly one contract lives in the
    file, and refusal when multiple do
  - priority ordering when multiple tiers could resolve

These are fully isolated — no DB, no temp directories, no native
LadybugDB binding — so they run in &lt;10ms total and are
immediately trustworthy as a regression safety net.

### Deliberately deferred (reviewer marked as "fine for now")

- `BridgeHandle._db` / `._conn` typing to `unknown` with casts in
  `bridge-db.ts` — reviewer's note: "The typing is fine for now."
- Batch inserts via `UNWIND` — needs LadybugDB support confirmation,
  tracked as a follow-up; the per-item pattern remains.
- `queryBridge` prepared-statement lifecycle — the current pattern
  (prepare → execute → GC) relies on LadybugDB's internals, worth
  verifying against their docs in a separate audit.

### Scope discipline (per `GUARDRAILS.md`)

- Only files touched by this PR (`bridge-db.ts`, `bridge-schema.ts`,
  `normalization.ts`, both bridge test files, new `fixtures.ts`) —
  no drive-by refactors
- No CI/release/security config changes
- No secrets

### Test + typecheck status

- `npx tsc --noEmit` clean
- `bridge-db.test.ts`: added 9 `findContractNode` tests, all pass in
  isolation. The full-file run still hits the pre-existing native
  LadybugDB cleanup segfault that flakes the reported count — same
  as every prior commit on this branch, not a regression.
- `bridge-db-edge.test.ts`: 4/4 pass
- `matching.test.ts`: 28/28 pass
- `types.test.ts`: 5/5 pass
- `retryRename` tests (4/4) and `findContractNode` tests (9/9)
  verified in isolation via `-t` filter

Co-authored-by: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-04-11 19:46:12 +01:00
Dave Brophy 9364739fb4 fix: restore tree-sitter-swift postinstall patch for macOS ARM64 (#788)
* fix: restore tree-sitter-swift postinstall patch for macOS ARM64

PR #516 (77dcb06) deleted `scripts/patch-tree-sitter-swift.cjs` and
the `postinstall` script entry when bumping to `tree-sitter-swift@0.7.1`,
since 0.7.1 ships prebuilt darwin-arm64 binaries and no longer needs the
patch. PR #538 (01ddc3e) then had to revert `tree-sitter-swift` back to
`^0.6.0` (and `tree-sitter` back to `^0.21.1`) because `npm overrides`
doesn't apply when gitnexus is installed via `npx -y` (gitnexus isn't the
root project, so overrides are silently ignored, producing ERESOLVE errors).

PR #538 reverted the grammar package changes but did not restore the patch
script, leaving `tree-sitter-swift@0.6.0` unable to build its native
binding on macOS ARM64. The symptom is `gitnexus analyze` printing
"Skipping swift" or "swift parser not available".

`Dockerfile.test` still references `node scripts/patch-tree-sitter-swift.cjs`
(added in the same PR #516), confirming the regression — the test image
build is also broken.

This commit restores the patch script from commit `0c8ec95` (the last
revision before it was deleted) and re-adds the `postinstall` entry to
`package.json`. No logic changes — it is an exact restoration.

The TODO comment in the script ("Remove this script when tree-sitter is
upgraded to ^0.22.x") still applies.

* style: run prettier on patch-tree-sitter-swift.cjs
2026-04-11 18:24:02 +01:00
smTheApexandProta100 75635638b1 feat(csharp): capture interface-to-interface heritage (#789)
The C# tree-sitter query set only matched `base_list` on
`class_declaration`, so interfaces extending other interfaces
(`interface IFoo : IBar`) were never captured as heritage edges.

This broke transitive interface implementation chains. For example,
given:

    interface IBase { }
    interface IFoo : IBase { }
    class MyClass : IFoo { }

only `MyClass -> IFoo` was emitted, and the `IFoo -> IBase` edge was
silently dropped. Any analysis that relies on walking the full
interface inheritance chain (e.g. "which classes implement IBase?")
therefore returned incomplete results.

This patch adds two new query patterns mirroring the existing
class_declaration heritage patterns, but targeting
`interface_declaration`:

    (interface_declaration name: (identifier) @heritage.class
      (base_list (identifier) @heritage.extends)) @heritage
    (interface_declaration name: (identifier) @heritage.class
      (base_list (generic_name (identifier) @heritage.extends))) @heritage

The existing heritage-processor pipeline already handles these
captures correctly once the query emits them, so no changes are
needed outside of tree-sitter-queries.ts.

Testing:
- New fixture `csharp-interface-heritage/` covering:
    * interface : interface  (single base)
    * interface : interface, interface  (multiple bases)
    * class : interface (where that interface derives from others)
- 6 new test cases in test/integration/resolvers/csharp.test.ts
  asserting exactly 4 IMPLEMENTS edges and 0 EXTENDS edges for the
  fixture.
- Full C# resolver suite: 175/175 passing, no regressions.

Co-authored-by: Prota100 <Prota100@users.noreply.github.com>
2026-04-11 17:41:51 +01:00
JWWD | ModusOpandClaude Opus 4.6 5be0537ce4 Fix stack overflow on large PHP files — iterative AST traversal (#783)
* Fix: replace recursive AST traversal with iterative stack to prevent stack overflow on large files

Fixes #752

Large PHP files (2000+ lines) with deeply nested AST structures (closures,
array literals, chained method calls) cause "Maximum call stack size exceeded"
during analysis. This converts three recursive tree traversal functions to
iterative loops using explicit stacks:

1. `walk()` in type-env.ts — the main AST walker that processes every node.
   On a 2,462-line PHP controller, this recurses through 5,000-10,000+ nodes.

2. `findRelationCall()` in languages/php.ts — recursive search for Eloquent
   relationship calls within method bodies.

3. `findDescendant()` in utils/ast-helpers.ts — generic recursive utility
   used by PHP property extraction and other parsers.

All three now use a while loop with an array-based stack instead of function
call recursion, eliminating V8's ~10K frame call stack limit as a constraint.

Tested against a production Laravel codebase with 373 PHP files (87,723 lines
total, largest file 2,462 lines) — indexes successfully in 17.4s with zero
errors, where the recursive version would crash with stack overflow.

* Fix: reverse child push order in findRelationCall iterative traversal

The iterative stack-based traversal pushed children in forward order,
causing the last child to be processed first (LIFO). This reversed the
original recursive left-to-right DFS order. Push children in reverse
so the first child ends up on top of the stack.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Fix: move stack declaration before processNode, rename walk to processNode

Move the stack initialization above the function that pushes onto it,
making the data-flow order match the code order. Rename walk to
processNode since it now processes a single node rather than
recursively traversing the tree.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 15:32:06 +01:00
Louis Chu a162f66254 fix(web): keep chat pinned on async content growth 2026-04-11 05:51:29 -07:00
Mr. WorldwideBrown 6d9ec1009e fix: load VECTOR extension during DB init for semantic search (#782)
* fix: load VECTOR extension during DB init for semantic search

The VECTOR extension was only loaded inside the embedding generation

pipeline (createVectorIndex). On a fresh gitnexus serve session,

semantic and hybrid search failed because QUERY_VECTOR_INDEX was

unknown.

Now loads the VECTOR extension alongside FTS during database

initialization in both the single-connection and pool-based paths.

Fixes #766

* fix: reset vectorExtensionLoaded on DB close and retry paths

The vectorExtensionLoaded flag was not being reset in closeLbug() or
the busy-retry cleanup path in withLbugDb(). This caused the VECTOR
extension to not be re-loaded after a close+re-init cycle, breaking
semantic search on reconnection.

Also resets shared.ftsLoaded and shared.vectorLoaded in the pool
adapter closeOne() for external DB entries, preventing stale extension
state when the pool is re-opened.

Adds integration tests covering vector extension loading, idempotency,
and state reset on both close and busy-retry paths.

* fix: set ftsLoaded flag in initLbugWithDb to avoid redundant extension reloads

* fix: set shared.vectorLoaded flag in initLbugWithDb to avoid redundant reloads
2026-04-11 12:17:40 +01:00
Mr. WorldwideBrown 4911201664 fix: map diff hunks to symbol line ranges in detect_changes (#779)
* fix: map diff hunks to symbol line ranges in detect_changes

The detect_changes tool previously used `git diff --name-only` and
picked the first 20 arbitrary symbols from each changed file. This
produced false positives (unchanged symbols reported as modified) and
false negatives (actually changed symbols dropped by the LIMIT).

Now uses `git diff -U0` to get unified diff with hunk headers, parses
the @@ line ranges, and queries for symbols whose [startLine, endLine]
range overlaps the diff hunks. Only truly touched symbols are reported.

Also fixed the CONTAINS path match to ENDS WITH to prevent cross-file
false positives from substring matching.

Fixes #758

* fix: address review feedback - variable shadowing, batch queries, tests

- Rename `params` to `queryParams` in detectChanges hunk-mapping loop
  to avoid shadowing the outer method parameter
- Replace N+1 per-symbol process lookup with a single batched query
  using WHERE n.id IN $ids (same pattern as impact BFS traversal)
- Add unit tests for parseDiffHunks covering single/multi file,
  single/multi hunk, omitted count, pure-deletion, and empty input

* style: fix prettier formatting in parse-diff-hunks test
2026-04-11 11:29:52 +01:00
Mr. WorldwideBrown 08541e2857 Fix HTTP client vs Express route detection and Spring interface attribution (#780)
* fix: correctly identify HTTP client calls vs Express routes in receiver extraction

* fix: skip Spring route extraction for Feign client interfaces

* fix: address review feedback - receiver walk edge case, regex anchoring, add tests

* style: fix prettier formatting in route extractor and test files
2026-04-11 11:24:47 +01:00
d9960c62bf SM-19: Delete resolveCallTarget — replace with thin dispatcher (#770)
* Initial plan

* SM-19: Replace resolveCallTarget with thin dispatcher

Delete the monolithic resolveCallTarget function (~200 lines) and replace it
with a 15-line thin dispatcher that routes to resolveMemberCall,
resolveStaticCall, or resolveFreeCall. Extract module-alias resolution and
file-based member-call fallback into dedicated helper functions.

- resolveCallTarget body reduced from ~200 lines to ~15 lines
- Extract resolveModuleAliasedCall helper (Python/Ruby module imports)
- Extract resolveMemberCallByFile helper (trait dispatch, overload disambiguation)
- Extract singleCandidate helper (constructor alias fallback, name-based fallback)
- Update unit tests for new dispatcher semantics
- Update doc comments referencing deleted D0-D4 paths

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/469eac38-b0c0-4a26-a2ff-3eb06299730b

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* SM-19: Add singleCandidate tail fallback for member calls with unresolvable receiver type

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/469eac38-b0c0-4a26-a2ff-3eb06299730b

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(SM-19): address all PR #770 review findings + fix CI

Fixes all 5 test failures (2 unit + 3 integration) and addresses 10
review findings from comment 4225312416.

Critical fix — singleCandidate null-route guard
The SM-19 dispatcher chained singleCandidate as an unconditional tail
fallback for member calls with receiverTypeName. This bypassed the
SM-10 R3 null-route contract: when the receiver type IS in the index
but file/owner filtering produced zero matches, the old code returned
null (genuine miss), but the new code fell through to singleCandidate
(false-positive CALLS edge).

Root cause: resolveMemberCallByFile returns null for two semantically
different reasons — (1) type not found in the index at all, and
(2) type found but no candidate matched after narrowing. The dispatcher
treated both as "try the next fallback." The old resolveCallTarget
exited the entire function on case 2.

Fix: after the scoped resolvers both return null, check whether the
receiver type resolves in the index. If it does (case 2), null-route
— the scoped resolvers made the right decision. If it doesn't (case 1,
e.g. PHP 'mixed', dynamic types), singleCandidate is the correct last
resort. ctx.resolve is cached so the check is free.

This fixes:
- Unit: no heritageMap null-route test (was getting 1 edge, expects 0)
- Integration: Rust c.trait_only() negative test
- Integration: 3 PHP heritage + alias tests (singleCandidate correctly
  fires when the receiver type is not in the index)

Performance (findings #1, #2, #3)
- Thread pre-computed tiered result into resolveModuleAliasedCall via
  new tieredOverride parameter — eliminates the duplicate ctx.resolve
  call on every module-alias path.
- Add countCallableCandidates helper that short-circuits at threshold
  without allocating an intermediate array — replaces the
  filterCallableCandidates(...).length > 1 allocation in skipMember.
- resolveMemberCallByFile lookupCallableByName caching deferred to a
  follow-up (finding #2) — the fix requires threading widenCache
  through the file-scoped resolver which is a larger change.

Code quality (findings #4, #5)
- Remove dead code: redundant conditional in resolveMemberCallByFile
  where both branches returned null.
- Move WidenCache type declaration from mid-file (between JSDoc blocks)
  to adjacent to CONSTRUCTOR_TARGET_TYPES with other type declarations.

Formatting
- Applied prettier to call-processor.ts (CI format check was failing).

Verification
- tsc --noEmit clean
- 3188 unit tests pass (0 skipped real tests)
- 1766 resolver integration tests pass
- Zero regressions — all PHP, Rust, and no-heritageMap tests green

Review: https://github.com/abhigyanpatwari/GitNexus/pull/770#issuecomment-4225312416

* fix(SM-19): restore module-alias narrowing and constructor disambiguation

Codex adversarial review on PR #770 surfaced two silent regressions in the
SM-19 thin dispatcher:

Finding 1 [high] — Typed member calls bypassed module-alias narrowing.
When two homonym receiver types are both imported by the caller, the
import-scoped tier no longer narrows and the owner/file resolvers see
genuine ambiguity. The dispatcher null-routed silently, dropping valid
CALLS edges. Fix: consult `resolveModuleAliasedCall` at the top of the
typed-member branch so an active alias on `call.receiverName` picks the
aliased file before the generic resolvers run.

Finding 2 [medium] — Constructor dispatch lost overload disambiguation.
When `resolveStaticCall` bails (ambiguous or ownerless Constructor pool)
and the caller supplied `overloadHints` / `preComputedArgTypes`, the
branch fell straight through to `singleCandidate` — which also bails on
multiple same-arity survivors. Fix: between `resolveStaticCall` and
`singleCandidate`, run constructor-filtered overload disambiguation on
the tiered pool. Only engages when a narrowing signal is present;
preserves SM-10 R3 null-route for genuinely ambiguous cases.

Tests:
- call-processor.test.ts: 3 new dispatcher-level regression tests
  covering real-homonym alias narrowing, constructor overload
  disambiguation with `argTypes`, and null-route control
- symbol-table.test.ts: update `module alias homonyms` test which
  previously codified the Finding 1 regression; now asserts resolution
  to the aliased file's method

Verification: 3191 unit + 2398 integration tests pass; tsc --noEmit
clean; prettier clean.

* refactor(SM-19): address code review findings with clean-code pass

Code review on commit f424685e surfaced one P1 correctness regression and
two P2 maintainability concerns. This commit closes all ten findings:

P1 — Alias helper placement regression
  - resolveModuleAliasedCall now runs as a FALLBACK in the typed-member
    branch, after resolveMemberCall/resolveMemberCallByFile return null.
    Previously it short-circuited BEFORE scoped resolvers, leaking unrelated
    homonyms from the aliased file when a local var coincidentally matched
    a module alias.
  - Added type-file verification guard: alias narrowing only fires when the
    alias target file is among the receiver type's defining files. Prevents
    cross-type false positives and hardens SM-10 R3.

P2 — Thin-dispatcher drift (roadmap Phase 3)
  - Extracted disambiguateByOverloadOrArgTypes shared helper. Centralizes
    the overloadHints → preComputedArgTypes precedence rule used by both
    member and constructor resolvers.
  - Folded constructor overload disambiguation into resolveStaticCall as
    step 4.5 (between the ambiguous-pool bail and the instantiable-class
    fallback). resolveStaticCall now accepts optional overloadHints /
    preComputedArgTypes symmetric with resolveMemberCallByFile.
  - Dispatcher's constructor branch returns to a 2-line delegation.
  - resolveMemberCallByFile now calls the shared helper instead of inlining
    the ternary.

P2 — Missing test coverage
  - owner-scoped wins over alias narrowing (alias with unrelated target
    class must not override unique owner-scoped answer)
  - alias narrowing rejects unrelated target type (type-file guard)
  - alias fallthrough: receiverName not in alias map
  - alias fallthrough: alias target file has no matching method
  (overloadHints-for-constructor variant transitively covered via the
   extracted helper's member-path tests; direct dispatcher test deferred
   as it requires real OverloadHints fixture parsing)

P3 — Clarity and durability
  - Stripped "Codex SM-19 Finding N" prefixes from comments. Replaced with
    durable explanations of WHY each guarded branch exists.
  - Added cross-reference comment at the tail-branch resolveModuleAliasedCall
    call site pointing to the typed-member branch usage.

Verification: 3195 unit + 1766 resolver integration + 2398 full integration
tests pass. tsc --noEmit clean. prettier clean.

Plan: docs/plans/2026-04-11-002-fix-sm19-code-review-findings-plan.md

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-11 10:20:57 +01:00
Mr. WorldwideBrown 7c983d798f Fix OpenCode config path, FTS extension load order, error messages, and CLAUDE.md stats (#781) 2026-04-11 06:18:14 +01:00
Yogesh Singh d87744fffc fix: resolve false 404 errors and stale repo context during multi-repo switching on Windows (#633)
* fix: resolve false 404s and stale repo context during multi-repo switching on Windows

* test(e2e): add repo-switching tests — hold-queue 503, ?project= URL, Windows path normalization

* test(e2e): fix repo-switching specs — use live backend with ?server= param
2026-04-10 19:59:16 +01:00
100858f8c8 feat(SM-18): Delete lookupFuzzy, lookupFuzzyCallable, globalIndex, callableIndex (#769)
* Initial plan

* Update test files for SymbolTable interface changes

Remove lookupFuzzy, lookupFuzzyCallable, globalIndex, and callableIndex
references from all test files. Replace lookupFuzzyCallable with
lookupCallableByName. Update getStats assertions to only expect
{ fileCount }. Remove tests that exclusively tested removed methods.

Files updated:
- symbol-table.test.ts: Remove lookupFuzzy describe block and all
  globalIndex/callableIndex tests, update callable method references
- symbol-resolver.test.ts: Remove SM-16 lookupFuzzy test block,
  update Tier 3 describe title
- type-env.test.ts: Update all mock SymbolTable objects and spy
  variable names
- call-form.test.ts: Update ownerId propagation test

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(SM-18): Remove lookupFuzzy, lookupFuzzyCallable, globalIndex, callableIndex

Remove from SymbolTable interface and implementation:
- lookupFuzzy method
- lookupFuzzyCallable method
- globalIndex Map
- callableIndex Map (renamed to callableByName, backing lookupCallableByName)

Add lookupCallableByName as the targeted replacement for fuzzy callable
lookups. Migrate all production callers:
- resolution-context.ts: lookupFuzzyCallable → lookupCallableByName
- type-env.ts: lookupFuzzyCallable → lookupCallableByName
- call-processor.ts: lookupFuzzy → lookupCallableByName (D2 widen paths)

Remove fuzzyCallCount/fuzzyCallableCallCount stats and globalSymbolCount
from getStats(). Update pipeline.ts logging accordingly.

Memory savings: globalIndex stored every non-Property symbol (typically
the largest index by entry count). Removing it eliminates one Map plus
all its per-name arrays — net savings proportional to unique symbol
count in the project.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/4a658c69-41a9-4d57-8527-50ca544ca967

* fix(SM-18): address all PR #769 review findings

1. type-env.test.ts mock: add missing lookupImplByName + getFiles methods.

2. Macro/Delegate tests: 2 new tests confirm C/C++ Macro and C# Delegate
   are indexed in callableByName.

3. D2 widen path test: module-alias scenario verifying lookupCallableByName
   resolves methods in aliased files that shadow same-file definitions.

4. CALLABLE_TYPES unified: exported from symbol-table.ts (single source of
   truth), imported in call-processor.ts. Removed duplicate
   CALLABLE_SYMBOL_TYPES constant.

5. getStats() observability restored: tier hit counters (tierSameFile,
   tierImportScoped, tierGlobal, tierMiss) replace the removed
   fuzzyCallCount diagnostic.

* chore(SM-18): remove unnecessary `as any` casts on valid NodeLabel types

Macro, Delegate, TypeAlias, Const, and Variable are all valid NodeLabel
values in gitnexus-shared. The casts suppressed type checking without
purpose and signaled false uncertainty.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-10 14:56:24 +01:00
e5dafce9f2 feat(SM-16): Restructure resolveUncached — replace lookupFuzzy data source for all tiers (#764)
* Initial plan

* chore: initial plan for SM-16 resolveUncached refactor

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0f505332-25be-46a7-b78e-fde58c1fc6fd

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* feat(SM-16): restructure resolveUncached — replace lookupFuzzy with targeted index lookups

- Remove single lookupFuzzy call that fed all Tier 2a/2b/3 in resolveUncached
- Tier 2a: iterate importedFiles with lookupExactAll per file (O(imports) × O(1))
- Tier 2b: iterate symbols.getFiles() filtered by isFileInPackageDir + lookupExactAll
  (O(files) × O(1), avoids global name scan)
- Tier 3: replace with lookupClassByName + lookupImplByName + lookupFuzzyCallable
  (three O(1) index lookups covering class-like, Rust impl blocks, and callables)
- Add getFiles() to SymbolTable interface (exposes fileIndex.keys() for Tier 2b)
- Add lookupImplByName() to SymbolTable — dedicated Rust Impl index kept separate
  from classByName to preserve correct heritage-map resolution
- Remove allDefs parameter from walkBindingChain; always use lookupExactAll directly
- Add 29 new unit tests covering SM-16 changes and per-language fixtures
- fuzzyCallCount in getStats() is now 0 for all resolve() calls (acceptance criterion)

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0f505332-25be-46a7-b78e-fde58c1fc6fd

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(SM-16): clean up — readable Tier 3 if-else, correct doc comment, remove unused import

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0f505332-25be-46a7-b78e-fde58c1fc6fd

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(SM-16): address all PR #764 review findings

1. Eager callableIndex — maintained on add() like classByName/implByName,
   removing the O(globalIndex) lazy rebuild on the Tier 3 hot path.

2. Tier 2b inverted index — packageDirSuffix→Set<filePath> built lazily on
   first Tier 2b hit. Changes O(allFiles×packages) per resolution to
   O(packages×filesInPackage).

3. Tier 3 type exclusion documented — TypeAlias, Const, Variable are
   intentionally not reachable at Tier 3. 4 negative/positive tests added.

4. Tier 3 allocation guard simplified — single spread replaces 4-way if-else.

5. getFiles() live iterator documented with safety contract.

6. Remaining lookupFuzzy callers in call-processor.ts documented in the
   Tier 3 comment block.

7. Tier 2b language fixtures — added Rust, Kotlin, PHP tests (3 new).

Also merges origin/main (SM-15 accumulator fixes).

* fix(SM-16): address Codex adversarial review — Tier 2b cache lifecycle + Macro/Delegate at Tier 3

1. Tier 2b packageDirIndex now invalidated in clearCache() and clear(),
   preventing stale snapshots when symbols/packages are added between
   chunk processing phases.

2. Macro (C/C++) and Delegate (C#) added to CALLABLE_TYPES in the eager
   callableIndex, restoring Tier 3 reachability for these call targets
   that the old lookupFuzzy returned.

* fix(SM-16): address ce:review findings — Tier 2b cache lifecycle + Tier 3 perf + test gaps

1. packageDirIndex no longer invalidated in clearCache() — the index
   persists across file boundaries since packageMap and symbols are
   append-only during the calls phase. Only clear() (pipeline reset)
   invalidates. Prevents O(files×dirs) rebuild per-file.

2. Tier 3 short-circuit: return null before spread when all three
   indexes are empty, avoiding allocation on the common miss path.

3. Add Macro (C/C++) and Delegate (C#) Tier 3 regression tests —
   the only newly-added CALLABLE_TYPES were completely untested.

4. Add packageDirIndex invalidation regression test — verifies clear()
   resets the index and newly-added symbols are visible.

* fix(SM-16): address final review — deduplicate NamedImportMap + doc fixes

1. NamedImportMap: removed duplicate definition from resolution-context.ts,
   now imported directly from import-processor.ts (no re-export needed —
   no consumers imported it from resolution-context).

2. packageDirIndex build cost documented accurately in comment.

3. fuzzyCallCount scope documented in test comment.

4. Tier 2a test suite: added comment about Go/Kotlin/PHP coverage.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-10 13:00:33 +01:00
CopilotandGergo Magyar ab956f113c feat(SM-15): Wire BindingAccumulator into processCallsFromExtracted for cross-file return type propagation (#763)
* Initial plan

* Initial setup - Phase 9 BindingAccumulator cross-file return type wiring

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/7cee6490-090d-4714-8cb5-a704168ff47a

* feat(SM-15): wire BindingAccumulator into processCallsFromExtracted for Phase 9 cross-file return type propagation

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/7cee6490-090d-4714-8cb5-a704168ff47a

* fix(SM-15): address all PR #763 review findings

Performance (R1)
- Changed _fileScopeByFile from Map<string, [string,string][]> to
  Map<string, Map<string,string>>. fileScopeGet(filePath, name) is
  now O(1) — replaces the O(n) linear scan + defensive-copy alloc
  that ran once per ConstructorBinding entry. fileScopeEntries()
  reconstructs tuples from Map.entries() for backward compat.
- Updated finalize() dev-mode invariant to compare deduplicated Map
  size rather than raw array length (Map.set deduplicates same-name).

Lifecycle (R2)
- Documented that Phase 9 intentionally reads pre-finalize because
  finalize() cannot move before both the worker consumer (line 984)
  AND the sequential-path writer (line 1061). Pre-finalize reads are
  safe because finalize() is write-lock-only with no side effects.
  Replaced the ambiguous "populated but not yet finalized" comment
  with the full lifecycle ordering explanation.

Sequential-path parity (R3)
- Wired bindingAccumulator into processCalls at line 797 (sequential
  path) so verifyConstructorBindings gets the Phase 9 fallback.
- Added bindingAccumulator parameter to processAssignmentsFromExtracted
  signature and wired it at the pipeline.ts call site (line 1026).
- Both paths now produce identical Phase 9 behavior for the same code.

Tracking comments (R4)
- Added "Overlapping mechanism (N of 3)" cross-references at:
  1. buildImportedReturnTypes (~line 109)
  2. collectExportedBindings (~line 168)
  3. Phase 9 fallback in verifyConstructorBindings (~line 563)
  Each links to the other two and notes future unification.

Language coverage (R5)
- Added 5 new Phase 9 integration test suites in cross-file-binding.test.ts:
  JavaScript, C++, C#, PHP, Ruby. Each uses the existing fixture
  directories and asserts getUser() → User → user.save() resolves.
  Total cross-file binding tests: 52 (was 37).

Quality asymmetry (R6)
- Added inline comment at the Phase 9 fallback noting worker-path
  entries are Tier 0/1 only and that binding accuracy is structurally
  lower for large repos where the worker path dominates.

Tests (+21 new)
- 6 fileScopeGet unit tests (happy path, unknown file/name, mixed
  scopes, post-dispose, duplicate varName last-write-wins)
- 15 integration tests across 5 new language suites

Verification
- tsc --noEmit clean
- 3147 unit tests pass (+6 new)
- 52 cross-file binding integration tests pass (+15 new)
- 1766 resolver integration tests pass
- Zero regressions

Plan: docs/plans/2026-04-10-001-fix-sm15-review-findings-plan.md
Review: https://github.com/abhigyanpatwari/GitNexus/pull/763#issuecomment-4220354242

* fix(SM-15): gate accumulator fallback on resolution tier and fix sequential file-order dependency

Two Codex adversarial reviews identified medium-severity bugs in the Phase 9
BindingAccumulator fallback:

1. Local-first violation: the fallback fired regardless of whether ctx.resolve()
   found same-file candidates, letting an imported callee shadow a local one
   and produce false CALLS edges. Fixed by gating on tiered.tier !== 'same-file'
   and callableDefs.length <= 1.

2. Sequential file-order dependency: processCalls flushed and verified per-file,
   so consumer files processed before their providers missed accumulator bindings.
   Fixed by splitting into a flush pre-pass (all files) then a resolution loop,
   mirroring the worker path's "all appends before any reads" pattern.

Also adds 11 consumer-before-provider integration test fixtures (one per
supported language) and 4 unit tests for tier gating edge cases.

* refactor(SM-15): eliminate duplicated prepare logic in processCalls two-pass split

Replace the duplicated pre-pass + legacy-path code (parse → query → heritage
→ TypeEnv → exports) with a single preparation loop followed by a resolution
loop. Both paths now share the same preparation code — the only conditional
is the accumulator flush.

Side benefit: globalParentMap is now fully populated before any resolution
runs, improving cross-file isSubclassOf accuracy regardless of file order.

Net -118 lines (226 removed, 108 added).

* fix(SM-15): address PR #763 third-pass review findings

1. Update stale dispose() JSDoc — remove forward-reference to Phase 9
   wiring that is now complete; document actual consumers.

2. Add processAssignmentsFromExtracted Phase 9 unit test — verifies the
   accumulator fallback produces ACCESSES write edges when the SymbolTable
   has no returnType for the callee.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-10 10:29:31 +01:00
Louis Chu ad2a397137 feat(web): add smart chat scroll 2026-04-10 00:06:24 -07:00
Mr. WorldwideBrown 6147579e54 Fix security issues and critical bugs found in code review (#709) 2026-04-10 05:20:29 +01:00
Abhigyan Patwari 4a1f912aee feat(sm-14): add BindingAccumulator — collect TypeEnv outputs across files (#743) 2026-04-09 21:03:25 +01:00
d09078925e Extract resolveFreeCall from resolveCallTarget (SM-13) (#756)
* Initial plan

* feat(SM-13): extract resolveFreeCall from resolveCallTarget

Extract the free-function call resolution path into a dedicated
`resolveFreeCall(calledName, filePath, ctx)` function that uses
`lookupExact` + import-scoped resolution via `ctx.resolve()`.

- Free function calls (foo()) now route through `resolveFreeCall`
- Swift/Kotlin implicit constructors (User()) delegate to
  `resolveStaticCall` within `resolveFreeCall`
- `resolveCallTarget` dispatches `callForm === 'free'` early,
  removing the inline freeFormHasClassTarget logic
- S0 block simplified to only handle `callForm === 'constructor'`
- Global (Tier 3) fallthrough preserved via ctx.resolve() until Phase 5
- 9 new unit tests for resolveFreeCall
- All 163 unit tests pass, all 1199 integration resolver tests pass

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c5f2e73a-259a-438c-b5c8-286b82e3c215

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore: revert unrelated package-lock.json change

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c5f2e73a-259a-438c-b5c8-286b82e3c215

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(SM-13): address PR #756 review findings on resolveFreeCall

Addresses all 7 findings from the PR #756 review comment.

Code (R1, finding #1)
- Replace the literal `'Class' | 'Struct' | 'Record'` check in
  `hasClassTarget` with `INSTANTIABLE_CLASS_TYPES.has(c.type)`. Converts
  an invariant that was previously comment-enforced ("keep this list
  aligned with INSTANTIABLE_CLASS_TYPES") into one enforced structurally.
  Any future extension of the set propagates here automatically. The
  narrower Swift extension dedup block below still uses literal
  `'Class' | 'Struct'` by design — Swift extensions only produce Class
  duplicates in practice, Record is deliberately excluded there, and
  the inline comment now documents that asymmetry.

Tests (+12 regression scenarios)

Finding #2 — language coverage
- Go free function (doStuff())
- Python free function (def helper(): ... helper())
- Rust free function outside any impl block
- Java statically-imported function
- JavaScript module-level function
Each exercises `_resolveCallTargetForTesting` with `callForm='free'`
and the language-specific file extension. `resolveFreeCall` has no
file-extension branching, so these guard the dispatch chain per
language without assuming extractor-specific symbol shapes.

Finding #3 — argCount threading
- 2-arg overload selected when argCount=2
- 0-arg overload selected when argCount=0

Finding #5 — Tier 3 (global) resolution
- Function globally visible but not imported. Asserts exact
  `TIER_CONFIDENCE.global === 0.5` and `reason === 'global'` to catch
  silent drift if the tier table is ever refactored.

Finding #6 — preComputedArgTypes worker path
- String overload matched via preComputedArgTypes=['String']
- Int overload matched via preComputedArgTypes=['int'] (lowercase,
  mirroring the parse-worker's inferred-literal shape; stored 'Int' is
  normalized via normalizeJvmTypeName at comparison time)

Finding #7 — Enum null-route documentation
- Enum-only free call asserts `toBeNull()` with an explanatory comment
  linking to the INSTANTIABLE_CLASS_TYPES rationale. NOT marked skipped
  — current behavior is intentional, not broken.

Finding #4 — Swift extension dedup guard
- Two same-name Class entries at different path lengths; exercises the
  full dispatch chain:
    1. filterCallableCandidates with 'free' strips Class → length 0
    2. hasClassTarget triggers resolveStaticCall
    3. Homonym ambiguity null-routes per SM-12 round-1 contract
    4. Constructor-form retry repopulates with both Classes
    5. Dedup block sorts by filePath.length → shortest path wins

Verification
- `tsc --noEmit` clean
- 3064 unit tests pass (+12)
- 1766 integration tests pass
- Zero regressions

Plan: docs/plans/2026-04-09-003-fix-sm13-resolve-free-call-review-findings-plan.md
Review: https://github.com/abhigyanpatwari/GitNexus/pull/756#issuecomment-4213879002

* refactor(SM-13): extract dedupSwiftExtensionCandidates shared helper

Follow-up to the PR #756 review fix. SM-13 duplicated the Swift
extension same-name collision dedup block between `resolveCallTarget`
and `resolveFreeCall` — two copies of identical 15-line logic with the
same heuristic (`filePath.length` sort, Class/Struct-only, `length > 1`
guard). Extract a single shared helper so the two sites cannot drift.

Changes
- New `dedupSwiftExtensionCandidates(candidates, tier)` helper defined
  alongside `tryOverloadDisambiguation`, with JSDoc documenting:
  - The Swift extension scenario it addresses
  - Why it is intentionally narrower than INSTANTIABLE_CLASS_TYPES
    (Class/Struct only, not Record — C#/Kotlin records don't exhibit
    the multi-file definition pattern, widening risks accidental
    dedup of legitimately distinct record types)
  - The return-null-on-no-match contract so callers can fall through
- `resolveCallTarget` tail dedup (was lines 1593-1610): replaced with
  a single `dedupSwiftExtensionCandidates` call
- `resolveFreeCall` tail dedup (was lines 1994-2012): same replacement
- Net line count: -32 insertions, -9 deletions in the consumer sites,
  +36 for the shared helper + JSDoc

Verification
- `tsc --noEmit` clean
- 3064 unit tests pass (including the R7 Swift dedup guard test added
  in the previous commit that exercises the full free-form retry
  chain through this helper)
- 1766 integration tests pass
- Zero regressions

Follows-up on: https://github.com/abhigyanpatwari/GitNexus/pull/756

* docs(SM-13): address PR #756 final review — comment cleanup only

Three documentation-only findings from the approval review. No
behavior change, no new tests, no code path modifications.

Finding #1 — stale line-number comment
- The comment inside `resolveFreeCall` at the `hasClassTarget` site
  referenced "lines ~1994-2008" for the Swift extension dedup block.
  Those lines were the inlined pre-SM-13 version; the block has since
  been extracted to `dedupSwiftExtensionCandidates`. Replaced the line
  reference with the helper name so future readers don't chase dead
  line numbers.

Finding #2 — fuzzy-widening asymmetry undocumented
- `resolveFreeCall` intentionally has no `widenCache` parameter and no
  D2 fuzzy-widening pass (unlike `resolveCallTarget`'s member-call
  path). Added an explicit "Asymmetry vs `resolveCallTarget`" paragraph
  to the JSDoc so a caller comparing the two signatures knows the
  skipped pass is deliberate and tied to Phase 5.

Finding #3 — constructor-form retry reasons undocumented
- `resolveStaticCall` can return null for three distinct reasons
  (empty instantiable pool, homonym ambiguity, ownerless Constructor
  nodes). The retry below it unconditionally re-filters with
  `'constructor'` form, which is correct for all three but not
  obvious. Added a structured three-case comment enumerating each
  reason and linking (a) to the SM-12 null-route contract, (b) to
  the R7 dedup test, and (c) to the currently-uncovered ownerless-
  Constructor path (noted as a future test candidate).

Verification
- `tsc --noEmit` clean
- 175 `resolveFreeCall` + `resolveStaticCall` + sibling tests pass
  (sanity check — no behavior change expected)
- No regressions

Follows-up on: https://github.com/abhigyanpatwari/GitNexus/pull/756#issuecomment-4215739052

* test(SM-13): cover ownerless-Constructor retry + PHP free function

Two low-severity test gaps from PR #756 review comment 4215739052 —
previously addressed doc-only, now have concrete test coverage.

Finding #3 low — ownerless-Constructor retry path (previously comment-only)
- The retry after resolveStaticCall returns null handles three distinct
  null-return reasons. Cases (a) and (b) were already tested (Interface/
  Trait null-route from SM-12, Swift shadowing dedup from R7). Case (c) —
  resolveStaticCall step-4 bailout when the tiered pool contains
  ownerless Constructor nodes — was only covered by a comment.
- New test: Class + ownerless Constructor in tiered pool, callForm='free'.
  Exercises the full chain:
    1. resolveStaticCall step 3 walks classCandidates via
       lookupMethodByOwner — ownerless Constructor not in methodByOwner,
       nothing found.
    2. Step 4 detects Constructor in tiered pool, bails with null.
    3. resolveFreeCall retry re-runs filterCallableCandidates with
       'constructor' form, which prefers Constructor over Class per
       CONSTRUCTOR_TARGET_TYPES ordering.
    4. Single survivor returned.
- Asserts the Constructor node (not the Class) is the resolved target.

Low — PHP free function coverage gap
- The language coverage table in the same review flagged PHP free
  functions (top-level `function helper()` outside any class) as
  uncovered. Added a test mirroring the existing Go/Python/Rust/Java/
  JS language tests — exercises the `.php` dispatch path for free
  calls. Ruby and C/C++ remain uncovered; deferred to a future round
  since those languages also have other gaps in the broader test file.

Verification
- `tsc --noEmit` clean
- 3066 unit tests pass (+2 new regression tests)
- 1766 integration tests pass
- Zero regressions

Follows-up on: https://github.com/abhigyanpatwari/GitNexus/pull/756#issuecomment-4215739052

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-09 17:41:28 +01:00
JaysonAlbertandgfwangjie 338cb01ee0 [codex] fix large repository graph loading (#732)
* fix(web): stream large graph responses

* fix(server): harden graph streaming

* fix(ci): stabilize graph loading coverage

---------

Co-authored-by: gfwangjie <gfwangjie@gf.com.cn>
2026-04-09 17:40:24 +01:00
4450a14b98 feat(SM-12): Extract resolveStaticCall from resolveCallTarget (#754)
* Initial plan

* feat(SM-12): extract resolveStaticCall from resolveCallTarget

- Add resolveStaticCall(className, methodName, currentFile, ctx, argCount?) using
  lookupClassByName + lookupMethodByOwner for O(1) constructor/static resolution
- Add S0 fast path in resolveCallTarget for constructor/free-form class calls
- Export resolveStaticCall from call-processor.ts
- Add 11 unit tests covering constructor resolution, confidence tiers,
  arity disambiguation, and resolveCallTarget delegation

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c9471ca9-57ff-4dae-956e-e7ffdc326bc4

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore: revert unrelated package-lock.json change

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c9471ca9-57ff-4dae-956e-e7ffdc326bc4

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor: shorten verbose test name per code review feedback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c9471ca9-57ff-4dae-956e-e7ffdc326bc4

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(SM-12): address PR #754 review findings

Addresses Claude's review comments on PR #754:

Performance
- Pass pre-computed `tiered` result into `resolveStaticCall` as optional
  `tieredOverride` parameter, eliminating the duplicate `ctx.resolve(className,
  currentFile)` on every constructor call path.
- Cache `freeFormHasClassTarget` in `resolveCallTarget` so the S0 fast path
  and the free-form constructor retry share a single `.some()` scan.

Architecture
- Reconcile `CLASS_LIKE_TYPES` (call-processor) with `CLASS_TYPES`
  (symbol-table): `CLASS_LIKE_TYPES = [...CLASS_TYPES, 'Impl']`. This makes
  the relationship explicit — the call resolver's set is a strict superset
  of the heritage-index set, guaranteeing anything reachable via
  `lookupClassByName` also passes the resolver filter. Trait is now included
  (harmless: traits have no Constructor nodes, so step-3 returns undefined
  and step-5 still returns the class-like node when unique). Documented
  the Interface inclusion rationale (static methods + MRO walker).
- Collapse `resolveStaticCall`'s `methodName` parameter into `className` —
  all call sites passed identical values. Named constructors (Dart
  `User.fromJson()`) arrive as member calls and go through
  `resolveMemberCall`. Documented the reserved path for when a language
  surfaces a static-method-shaped call with a distinct member name.
- Document the known gap: `callForm === 'member'` constructor patterns
  (e.g. Python `models.User()`) are handled by the tail fallback, not S0.

Tests
- Add tiered-override test asserting `ctx.resolve` is not re-invoked when
  a pre-computed result is passed in.
- Add language-specific `_resolveCallTargetForTesting` integration tests
  for Java (`new User()`), Python (`User()`), and Kotlin (`User()`).

Verification: 3031 unit + 1766 integration tests pass, zero regressions.

* fix(SM-12): restrict resolveStaticCall fallback to instantiable kinds

Addresses the high-severity finding from the Codex adversarial review of
PR #754: `resolveStaticCall`'s step-5 "return the class itself when no
Constructor node is found" fallback reused `CLASS_LIKE_TYPES`, which —
after SM-11 and PR #754's reconciliation — now includes `Interface`,
`Trait`, and `Impl`. That is the method-dispatch set, not the
instantiable set, so constructor-shaped calls could resolve to
non-instantiable nodes and emit false `CALLS` edges.

Concrete failure: Rust same-file `impl User { ... }` alongside
`struct User { ... }` — both land at same-file tier, the Impl is not
filtered out, and the step-5 fallback produces a `CALLS` edge to the
`Impl` block instead of the `Struct`. The same widening exposed
Interface / Trait targets in Java / C# / PHP / Scala.

Fix
- Introduce `INSTANTIABLE_CLASS_TYPES = {'Class', 'Struct', 'Record'}`
  as a sibling to `CLASS_LIKE_TYPES`, documenting the contract
  explicitly and cross-referencing `CONSTRUCTOR_TARGET_TYPES`.
- Update `CLASS_LIKE_TYPES` JSDoc to clarify it is the method-dispatch
  set and add an anti-pattern warning against reusing it for
  constructor-fallback filtering.
- Tighten `resolveStaticCall` step 5: filter `classCandidates` through
  `INSTANTIABLE_CLASS_TYPES` before the `length === 1` check. This
  strips `Impl` from the Rust shadowing scenario (leaving `Struct` as
  the sole instantiable target) and null-routes Interface / Trait /
  `Impl`-alone scenarios, matching the SM-10 R3 null-route precedent.
- Step 3 (explicit Constructor lookup via `lookupMethodByOwner`) is
  intentionally unchanged — its `def.type === 'Constructor'` check is
  the correct contract, and legitimate Constructor nodes attached to
  `Impl` owners still resolve correctly.

Tests (+10 regression scenarios)
- Positive guards: Struct, Record fallback paths.
- Null-route: Interface (Java/C#/TS), PHP Trait, Rust Trait.
- Rust same-file shadowing: Struct wins over Impl.
- Rust Impl-alone: null-routes (no Struct present).
- Step-3 preservation: Constructor owned by Impl still resolves to the
  Constructor node, proving step-5 tightening doesn't leak into step 3.
- Full cascade via `_resolveCallTargetForTesting` for Interface and
  Trait — confirms no downstream path silently re-introduces the edge.

Verification
- `tsc --noEmit` clean
- 3041 unit tests pass (+10)
- 1766 integration tests pass
- Zero regressions

Plan: docs/plans/2026-04-09-002-fix-sm12-constructor-fallback-instantiable-only-plan.md
Codex review job: review-mnrao7fr-nv9y0e

* fix(SM-12): address PR #754 second review round

Addresses the 9 findings from the follow-up review on PR #754.

Performance
- Align `freeFormHasClassTarget` with `INSTANTIABLE_CLASS_TYPES`: drop
  `Enum` (S0 would always return null for it — wasted lookup work) and
  add `Record` (C# records and Kotlin data classes were bypassing S0
  entirely). The trigger set and the fallback filter set now agree by
  construction, documented inline.

Documentation
- Remove stale single-line JSDoc on `CLASS_LIKE_TYPES` (line 57) that
  duplicated the full multi-line block immediately below it — tooling
  picks up the first block so the old one-liner was shadowing the
  current explanation.
- Rewrite the `resolveStaticCall` JSDoc step list to match the actual
  step boundaries in the implementation (steps 3, 4, 5 were blurred in
  the old description).
- Add inline comment on step 3 documenting the same-name lookup
  assumption (`${candidate.nodeId}\0${className}`) and the symmetric
  miss case for Python `__init__`-style constructors.
- Add inline comment on step 4 documenting that it also catches the
  ambiguous-step-3 case, and warning against removing the check
  without handling that path explicitly.
- Add inline comment on step 5 enumerating the three length outcomes
  (0 / 1 / >1) so future readers see the dominant null-route case.
- Document Ruby `User.new` as a known gap alongside Python
  `models.User()` in the S0 header comment.

Tests (+2 scenarios)
- Record free-form constructor call via `_resolveCallTargetForTesting`
  exercises the aligned `freeFormHasClassTarget` trigger end-to-end,
  closing the gap where the direct `resolveStaticCall` test passed
  but the integration path was silently bypassing S0.
- Arity threading via `_resolveCallTargetForTesting` asserts that
  `call.argCount` flows through resolveCallTarget → S0 →
  resolveStaticCall → lookupMethodByOwner, catching any future
  regression where the argCount is dropped at the S0 call site.

Verification
- `tsc --noEmit` clean
- 3043 unit tests pass (+2)
- 1766 integration tests pass
- Zero regressions

Plan: docs/plans/2026-04-09-002-fix-sm12-constructor-fallback-instantiable-only-plan.md
Review: https://github.com/abhigyanpatwari/GitNexus/pull/754#issuecomment-4213536094

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-09 12:09:07 +01:00
Kunal Hemnani 3f0b8c1a5b fix(ingestion): replace lookupExact with lookupExactAll in named-binding-processor (#755) 2026-04-09 11:57:36 +01:00
bb68cc1eb0 Extract resolveMemberCall from resolveCallTarget (SM-11) (#744)
* Initial plan

* feat(SM-11): extract resolveMemberCall from resolveCallTarget

- Create resolveMemberCall(ownerType, methodName, currentFile, ctx, heritageMap?)
  that uses owner-scoped + MRO resolution only (no fuzzy lookup)
- resolveCallTarget delegates member calls (D0 path) to resolveMemberCall
- walkMixedChain uses resolveMemberCall for owner-scoped member-call resolution
- Add 7 unit tests for resolveMemberCall covering direct, inherited, MRO,
  null cases, and confidence tier assertions
- Export resolveMemberCall for external use

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/3b7889a9-5f2f-4572-8904-45084210f10d

* fix(SM-11): address PR #744 review

Blocking fixes:

- B1: Revert unrelated package-lock.json gitnexus-shared addition

- B2: Document confidence-tier semantic change on resolveMemberCall

Performance / coupling fixes:

- S1: walkMixedChain now calls resolveMethodByOwner directly (hot path) to avoid throwaway ResolveResult allocation per chain step

- S2: Thread tier from resolveMethodByOwner via { def, tier } tuple; eliminates double ctx.resolve

Alignment with semantic-model plan (Phase 3 target):

- resolveMethodByOwner now iterates ALL class-like candidates from ctx.resolve, deduplicating matches by nodeId. Absorbs D4's ownerId-filtering into the owner-scoped path.

- Handles homonym classes (two Users in different files) without falling through to D1-D4 fuzzy widening

- Shared-ancestor MRO walks automatically dedup (both homonyms walk to same base method)

- Unified direct-vs-MRO lookup under a single canWalkMRO check

Tests added:

- T1: Three D0 skip-condition tests via new _resolveCallTargetForTesting internal export (overloadHints, preComputedArgTypes, hasActiveModuleAlias)

- T2: Rust qualified-syntax null test (trait-inherited method) + direct impl control

- T3: C++ leftmost-base diamond inheritance test

- B2 lock-in: cross-file class tier assertion

- Homonym disambiguation: only-one-owns-method, both-own-method ambiguity, shared-ancestor MRO convergence

Verification:

- tsc --noEmit: clean

- vitest run test/unit/: 3014 passed

- vitest run test/integration/resolvers/: 1746 passed

* test(SM-11): address second PR #744 review round + per-language integration tests

Review fixes (https://github.com/abhigyanpatwari/GitNexus/pull/744#issuecomment-4211877593):

P1 (Performance): Replace Map allocation in resolveMethodByOwner with a firstDef+ambiguous flag pattern. Zero allocation for the common single-candidate case on the hot path — the previous Map approach allocated on every member call regardless of whether deduplication was needed.

P2 (Test gap): Strengthen the module-alias D0 skip test with a homonym fixture (two Users in different files). Previously the test passed whether or not D0 was actually bypassed; the new version proves D0 must be skipped by showing that resolveMemberCall directly returns null (ambiguous) but D1-D4 with alias narrowing picks the right one. Also fixes the underlying D2-vs-alias widening interaction: when filteredCandidates was narrowed by module-alias disambiguation, D2 no longer widens back to the full fuzzy pool (introduces aliasNarrowed boolean flag).

L1 (Language coverage): Add C# and Kotlin implements-split tests at the resolveMemberCall layer.

L2 (Maintainability): Export OverloadHints as @internal so the test can use a direct cast instead of fragile Parameters<...> type inference.

Per-language integration tests:

- rust-child-extends-parent: Direct impl method resolution via D0 (with honest documentation of the trait-method-as-Function gap that is Phase 5 / SM-16 scope)

- java-interface-default-method: User implements Validator with default method resolved via implements-split MRO

- csharp-interface-default-method: Same pattern for C# 8.0+ default interface methods

- kotlin-interface-default-method: Same pattern for Kotlin interfaces with default implementations

- python-multi-level-mro: 3-level C3 linearization (Grandparent ← Parent ← Child)

- cpp-diamond-inheritance: Classic diamond (Base ← A, B ← Derived) via leftmost-base MRO

Verification:

- tsc --noEmit: clean

- vitest run test/unit/: 3015 passed

- vitest run test/integration/resolvers/: 1763 passed (+17 new per-language tests)

* fix(SM-11): Codex adversarial review corrections + deeper D0 fixes

Addresses the three high-severity findings from the Codex adversarial review of PR #744 (https://github.com/abhigyanpatwari/GitNexus/pull/744#issuecomment-4212075120), plus four deeper fixes discovered during regression triage. All discovered issues are now addressed end-to-end rather than papered over with tail-return fallbacks.

Codex review findings:

R1 (C++ diamond): The cpp-diamond-inheritance fixture used non-virtual inheritance, which is genuinely ambiguous in real C++ (two Base subobjects). Changed A and B to use 'virtual public Base' so there's a single shared Base subobject and d.method() is an unambiguous call that the leftmost-base MRO walk correctly resolves.

R2 (C# default-interface): The csharp-interface-default-method fixture called user.Validate() via a User-typed variable, but C# does not inherit default interface methods as callable class members — the call is only valid through an interface-typed variable. Changed App.cs to 'IValidator user = new User(...)' which is the idiomatic dispatch pattern.

R3 (resolveCallTarget tail-return): When D1-D4 receiver filtering produced zero file-matched and zero owner-matched candidates for a member call, the function fell through to the permissive single-candidate tail return — silently emitting CALLS edges for methods that don't belong to the receiver. Added an explicit null-route inside the D1-D4 block that fires only when both filters yielded 0.

R4 (Rust negative assertion): Added the c.trait_only() negative integration test in rust.test.ts demonstrating that direct member calls on Rust structs do not walk trait ancestry. The test now passes because of R3 (previously fell through to the tail return).

Regression triage discoveries:

1. D0 was dead code on the sequential pipeline. The sequential path sets overloadHints for every call regardless of whether the method is overloaded, and the original D0 skip condition '!overloadHints && !preComputedArgTypes' was therefore always false. The Java/C#/C++ SM-9/SM-10 inheritance tests were passing ONLY via the tail-return fallback. Fix: narrow the skip to 'overloadHints && filteredCandidates.length > 1' — skip D0 only when there are actually multiple candidates that need overload disambiguation.

2. lookupMethodByOwner couldn't disambiguate arity-differing overloads (e.g. C++ greet() vs greet(string)). With D0 now firing on the sequential path, same-name/different-arity overloads would collapse to an arbitrary first pick. Fix: added an optional argCount parameter to lookupMethodByOwner + lookupMethodByOwnerWithMRO that filters the overload set by parameterCount/requiredParameterCount before the returnType dedup.

3. Python and Rust class methods are captured as Function nodes (not Method) with ownerId set to the class. The methodByOwner index only accepted 'Method' and 'Constructor' types, so Python class methods and Rust trait methods were invisible to D0. Fix: extended the methodByOwner indexing condition to include 'Function' when ownerId is set. This also unlocks the Rust trait-method negative assertion by ensuring the qualified-syntax MRO strategy has something to return null for.

4. D0 was being skipped when a local variable shadowed an imported module name (Python 'from models.c import C; c = C()' creates both a module alias 'c → models/c.py' AND a typed local 'c'). Fix: the D0 skip now gates on 'aliasNarrowed' (a new boolean tracking whether the alias block actually narrowed filteredCandidates) instead of 'hasActiveModuleAlias'. If the method isn't in the aliased module, the receiver is a typed local variable and D0 should run.

5. PHP trait walk missed the HasTimestamps trait because lookupClassByName did not include 'Trait' type. buildHeritageMap uses lookupClassByName to resolve parent names, so 'BaseModel use HasTimestamps' was failing to register an ancestor edge for BaseModel → HasTimestamps. Fix: added 'Trait' to CLASS_TYPES. The trait is now a valid class-like type for heritage resolution (PHP use, Rust impl Trait for Struct, Scala traits).

Test updates:

- Updated the 'no heritageMap' unit test in call-processor.test.ts to assert the correct null-route behavior instead of the old tail-return fallback.

- Added a new unit test asserting Trait inclusion in the class set.

- Updated the 'does NOT include other type-like labels' test to remove Trait from its rejection set.

Verification:

- tsc --noEmit: clean

- vitest run test/unit/: 3016 passed (+1 new Trait inclusion test)

- vitest run test/integration/resolvers/: 1764 passed (+1 new Rust negative assertion)

- Zero regressions

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergo Magyar <magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-09 09:52:12 +01:00
Roshan Warrierandtxhno d6debf3324 fix(symbol-table): index constructors in methodByOwner (#753)
Co-authored-by: txhno <198242577+txhno@users.noreply.github.com>
2026-04-09 08:26:15 +01:00
Pratyush Sharma 9ab92a97d0 fix(deps): pin tree-sitter-c override to resolve peer dep conflict (#720) (#723) 2026-04-09 06:40:17 +01:00
Murat Çelik 4fde5f241b feat: print skipped large file paths in verbose analyze output (#745) 2026-04-09 06:18:32 +01:00
evolution 9f9bbcd744 feat: support GITNEXUS_HOME env var to customize global directory (#746) 2026-04-09 06:18:17 +01:00
Cocoon-Break fd67cfd5a7 docs: fix web UI install link spacing in README (#731) 2026-04-09 06:16:14 +01:00
Pratyush Sharma 3f28f7ead5 fix(web): correct dev-mode serve command in OnboardingGuide (#725) 2026-04-09 06:15:14 +01:00
d9ba9aa998 SM-10: Add MRO fast path before D2 fuzzy widening in resolveCallTarget (#741)
* Initial plan

* Add MRO fast path before D2 fuzzy widening in resolveCallTarget

When receiverTypeName is known, try resolveMethodByOwner (owner-scoped
+ MRO lookup) before falling back to the expensive lookupFuzzy in D2.
This short-circuits cross-file member call resolution for the common
non-overloaded case.

The fast path is skipped when overload disambiguation hints are
available (overloadHints or preComputedArgTypes) to avoid picking the
wrong overload for same-return-type overloaded methods.

Passes heritageMap to resolveCallTarget from all 4 call sites:
- Language seed path (processCalls)
- Sequential path (processCalls)
- walkMixedChain fallback
- Worker path (processCallsFromExtracted)

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9e49521f-2472-47bc-96e9-be4a46b073f0

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(SM-10): address PR #741 review

Correctness:
- Module-alias guard for D0. When call.receiverName matches an active
  entry in ctx.moduleAliasMap for the current file, D0 is now skipped
  and resolution falls through to D1-D4 which respects the
  alias-narrowed candidate pool. Prevents a homonymous class in a
  different file from being picked by ctx.resolve(receiverTypeName)
  inside resolveMethodByOwner. New unit test pins the contract.

Unit tests (call-processor.test.ts — 3 new):
- D0 hit: child.parentMethod() resolves via MRO walk when
  heritageMap is provided.
- D0 skipped: same scenario still resolves via D1-D4 when heritageMap
  is undefined (backward-compat guard).
- Module-alias guard: two files both define class User with a save()
  method; 'import auth_mod as auth' in app.py must resolve
  auth.user.save() to auth_mod.py, not user_mod.py.

Integration language coverage (+3 fixtures/tests):
- swift-child-extends-parent — first-wins, gated on swiftAvailable.
- ruby-child-extends-parent   — first-wins.
- php-child-extends-parent    — first-wins (uses ParentClass since
  'Parent' is a PHP reserved word).

* test(SM-10): address second PR #741 review round

Unit tests (call-processor.test.ts, +2 new):
- overloadHints guard: Java source with two same-return-type overloads
  method(int) and method(String), int added first so lookupMethodByOwner
  would return it. processCalls auto-generates overloadHints for Java,
  forcing D0 to be skipped. o.method("hello") must resolve to
  method(String) via literal-inferred disambiguation.
- preComputedArgTypes guard: worker-path equivalent via
  processCallsFromExtracted with ExtractedCall.argTypes=['String'].
  Same two overloads, same correctness guarantee.

Integration tests (+2 fixtures + test blocks):
- go-child-extends-parent    — struct embedding, first-wins
  (Go structs are labeled 'Struct' not 'Class' in GitNexus).
- dart-child-extends-parent  — extends, first-wins, gated on
  dartAvailable like other Dart tests.

Documentation:
- Expanded the fallthrough comment in resolveMethodByOwner to clarify
  that unknown-extension paths land on plain lookupMethodByOwner
  without an ancestor walk, and that D1-D4 still runs on D0 miss.

* test(SM-10): D0 miss with heritageMap present falls through to D1-D4

Closes the last remaining gap from PR #741 review round 3. The existing
'D0 skipped' test only covered the heritageMap=undefined case, leaving
the miss-with-heritageMap path implicitly covered by integration tests
only. This adds a focused unit test where:

- Class Obj has a method doWork findable via tiered resolution
  (import-scoped) but intentionally NOT registered in methodByOwner
  (no ownerId), so lookupMethodByOwner misses.
- heritageMap is provided but built from an empty heritage array, so
  getAncestors(class:Obj) returns []. The MRO walk yields no parents.
- lookupMethodByOwnerWithMRO therefore returns undefined → D0 miss.
- D1 resolves the receiver type; D2 widens via lookupFuzzy;
  D3 file-filter picks the single matching candidate.
- A CALLS edge must still be emitted — D0 miss must not swallow
  the call.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-08 23:24:07 +01:00
c19e76a4a3 feat(SM-9): Add lookupMethodByOwnerWithMRO using HeritageMap (#740)
* Initial plan

* feat(SM-9): add lookupMethodByOwnerWithMRO with HeritageMap parent chain walking

- Export c3Linearize from mro-processor.ts for reuse
- Add lookupMethodByOwnerWithMRO in call-processor.ts with MRO strategy support
- Update resolveMethodByOwner to fall back to MRO walk when HeritageMap available
- Thread heritageMap through walkMixedChain for chain resolution
- Add 10 unit tests covering all acceptance criteria

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/cc58249b-42f1-45a9-89fb-e3917e4d0171

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* feat(SM-9): add Java integration test with class Child extends Parent fixture

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/cc58249b-42f1-45a9-89fb-e3917e4d0171

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docs: address code review comments on MRO strategy documentation

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/cc58249b-42f1-45a9-89fb-e3917e4d0171

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* perf(SM-9): address PR #740 review comments

- Eliminate double direct lookup in resolveMethodByOwner: delegate
  straight to lookupMethodByOwnerWithMRO when a HeritageMap is
  available (the MRO helper already does the direct lookup before
  walking ancestors). Fallback path handles the no-HeritageMap case.
- Memoize C3 linearization per HeritageMap via a WeakMap keyed cache.
  HeritageMap is immutable after build, so C3 results are stable for
  its lifetime; WeakMap lets the cache auto-drain when the HeritageMap
  is GC'd. Null sentinel caches linearization failures so cyclic
  hierarchies are not reprocessed. Eliminates per-call buildParentMap +
  c3Linearize on Python codebases.
- ancestors variable typed as readonly to accept the cached result
  without copying.
- Add four missing MRO unit tests: Kotlin implements-split, C#
  implements-split, JavaScript first-wins (separate provider from TS),
  and C++ leftmost-base diamond (first diamond test for C++).

* fix(SM-9): CI prettier + address PR #740 follow-up review

- Fix CI prettier failure in test/integration/resolvers/java.test.ts
  (auto-formatted — was introduced in 37563a31 before my first fix
  commit but had not been caught locally).
- Pin caller on the SM-9 Java integration test (parentMethodCall.source
  === 'run') so a regression that misattributes the CALLS edge fails.
- Add two implements-split unit tests:
  * Ambiguous default from two interfaces → BFS first-wins. Pins the
    contract that lookupMethodByOwnerWithMRO returns a defined result
    (full ambiguity detection is deferred to computeMRO graph pass).
  * Class method precedence over interface default: Child extends Base
    implements IFoo where both define handle() — documents that BFS
    visits the extends edge first, matching Java's class-wins rule.
- Add @internal JSDoc on lookupMethodByOwnerWithMRO clarifying it is
  exported only for testing; resolveMethodByOwner is the proper entry
  point for callers.

* test(SM-9): per-language integration fixtures and tests for inherited method resolution

Extends the SM-9 integration coverage beyond Java with six new
child-extends-parent fixtures, one per MRO strategy:

- python-child-extends-parent       → C3 strategy
- typescript-child-extends-parent   → first-wins
- javascript-child-extends-parent   → first-wins (separate provider)
- kotlin-child-extends-parent       → implements-split
- csharp-child-extends-parent       → implements-split
- cpp-child-extends-parent          → leftmost-base

Each fixture follows the java-child-extends-parent pattern:
- Parent class with a single method
- Child class extending Parent, no override
- App class/function that instantiates Child and calls
  the parent method — exercises the full ingestion pipeline,
  HeritageMap construction, and lookupMethodByOwnerWithMRO walk.

For every fixture the matching integration test asserts:
- Parent and Child classes are detected
- Child → Parent EXTENDS edge is emitted
- The parent-method call resolves to the correct target file
- The caller is pinned (source === 'run' / 'Run') to catch
  edge misattribution regressions

Rust is intentionally omitted — its qualified-syntax strategy
returns undefined from lookupMethodByOwnerWithMRO by design, so
there is no inherited-method resolution to assert against.

All 1739 integration resolver tests pass (+18 new SM-9 tests
across 6 languages).

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-08 20:46:45 +01:00
b75e76d44a feat(SM-8): Build HeritageMap from accumulated ExtractedHeritage[] (#739)
* Initial plan

* feat(SM-8): add HeritageMap with MRO-aware parent/ancestor lookup

- New heritage-map.ts: HeritageMap interface with getParents() and getAncestors()
- buildHeritageMap() consumes ExtractedHeritage[], resolves names via lookupClassByName
- Cycle protection and bounded depth (MAX_ANCESTOR_DEPTH=32) in getAncestors
- Worker path: HeritageMap built from deferredWorkerHeritage, threaded into processCallsFromExtracted
- Sequential path: Heritage accumulated across chunks, HeritageMap built after all chunks, passed to processCalls
- 18 unit tests covering parent lookup, multi-level, diamond, cycles, missing parent, bounded depth

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c413e0a3-5d63-4ddb-8ece-02fe6ed99efd

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: rename cycle test for clarity per code review

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c413e0a3-5d63-4ddb-8ece-02fe6ed99efd

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor(SM-8): merge implementor map into heritage map

- Add `getImplementorFiles(interfaceName)` to HeritageMap interface
- Build implementor index (interface name → file paths) alongside parent
  lookup in `buildHeritageMap`, using same `resolveExtendsType` logic
- Remove `ImplementorMap` type, `buildImplementorMap`, `mergeImplementorMaps`
  from call-processor.ts
- Update `findInterfaceDispatchTargets`, `processCalls`, and
  `processCallsFromExtracted` to use HeritageMap for both parent
  lookup and implementor dispatch
- Pipeline: single `buildHeritageMap` call replaces separate
  buildImplementorMap + buildHeritageMap for both worker and
  sequential paths
- Migrate implementor tests from call-processor.test.ts to
  heritage-map.test.ts (4 new getImplementorFiles tests)
- Update interface dispatch test to use buildHeritageMap instead
  of hand-constructed ImplementorMap

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/085dffb4-b31e-4aa5-9aa3-4314bc0010e7

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: rename implementor test for clarity per code review

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/085dffb4-b31e-4aa5-9aa3-4314bc0010e7

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(SM-8): address PR #739 review comments

- pipeline.ts: cache chunk file contents from Pass 1 to eliminate
  double-read of sequential chunks in Pass 2. Peak memory drains
  incrementally as Pass 2 processes each chunk.
- heritage-map.ts: document Rust trait-impl omission from implementor
  index and the interface-name collision limitation.
- heritage-map.test.ts: add six tests covering the extends->IMPLEMENTS
  path across C# (interfaceNamePattern), Swift (heritageDefaultEdge),
  Java (symbol-table Interface lookup), Kotlin, PHP, and the Rust
  trait-impl omission.
- pipeline.ts: comment why the heritage accumulation uses a manual
  push loop instead of spread (ref #650).

* test(SM-8): address second PR #739 review pass

- Add TypeScript implements test to getImplementorFiles (closes
  the .ts coverage gap flagged by the bot reviewer).
- Tighten deep-chain boundary assertion from toBeLessThanOrEqual(32)
  to toBe(32) so a future regression returning fewer ancestors
  fails loudly. Added an ancestors[31] === 'class:Level32' check
  to pin the upper boundary.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-08 19:00:08 +01:00
MyShiningand许恩宁 83b5bec293 [cli] Replace owner-filtered method lookups in type-env (#736)
* refactor(type-env): use owner method lookup

* test(type-env): cover owner lookup edge cases

* test(type-env): cover inherited overload ambiguity

---------

Co-authored-by: 许恩宁 <xuenning@qiyi.com>
2026-04-08 17:41:09 +01:00
MyShiningand许恩宁 3388ae16d7 [cli] Replace Phase P class checks with class lookup index (#734)
* refactor(call-processor): use class lookup index in phase p

* test(call-processor): cover class lookup fallback

---------

Co-authored-by: 许恩宁 <xuenning@qiyi.com>
2026-04-08 14:48:54 +01:00
MyShiningand许恩宁 d784f591b2 [cli] Replace class-type fuzzy lookups in type-env.ts (#733)
* refactor(type-env): use class lookup index for type resolution

* test(type-env): add lookupClassByName regression coverage

* test(type-env): expand class lookup regression coverage

---------

Co-authored-by: 许恩宁 <xuenning@qiyi.com>
2026-04-08 14:12:01 +01:00
Kunal Hemnani 0f43190543 feat(symbol-table): add fuzzy lookup counters (#708) 2026-04-08 07:53:30 +01:00
Roshan Warrier fe87ff8f74 fix(symbol-table): index constructors in methodByOwner (#694) 2026-04-08 06:32:02 +01:00
Deepak Chauhan be2401061e [cli] Add qualified class lookups to SymbolTable (#716) 2026-04-07 22:57:18 +01:00
Deepak Chauhan b73233d232 feat(symbol-table): add class name lookup index (#707) 2026-04-07 13:29:50 +01:00
Tushar Dhawas (Kyo) 1c8ae5eb46 refactor: extract CLASS_LIKE_TYPES constant (#693)
* refactor: extract CLASS_LIKE_TYPES constant

* chore: apply prettier formatting
2026-04-07 11:56:03 +01:00
Zander Raycraft b73928f732 scarf (#688) 2026-04-06 18:39:43 -05:00
Dmytro Semchuk 19faf3b326 fix(docs): fix codex duplicate typo in main readme file (#687) 2026-04-06 22:20:35 +01:00
Gergő Magyar cb772b9e29 feat: lookupMethodByOwner index for O(1) cross-class chain resolution (#665)
Add eagerly-populated methodByOwner index to SymbolTable, keyed by
ownerNodeId\0methodName. Used by walkMixedChain as a fast path for
resolving intermediate method calls in cross-class chains like
user.getAddress().getCity().getZipCode(), avoiding expensive fuzzy
lookups when the owner type is already known.

Handles overloaded methods: returns the first match when all overloads
share the same returnType, undefined when return types differ (ambiguous).

- Add lookupMethodByOwner to SymbolTable interface + implementation
- Add resolveMethodByOwner helper in call-processor.ts
- Add fast path in walkMixedChain before resolveCallTarget fallback
- Add Java cross-class chain fixture + 6 integration tests
- Add 148 unit tests for methodByOwner index behavior
2026-04-06 10:20:04 +01:00
ivkondandClaude Opus 4.6 10f8815639 fix(ignore): respect negation patterns in .gitnexusignore (#654)
* fix(ignore): respect negation patterns in .gitnexusignore childrenIgnored

childrenIgnored checked `ig.ignores(rel) || ig.ignores(rel + '/')` which
short-circuited on the bare path — directory-only negation patterns like
`!iOS/` were missed because `ig.ignores('iOS')` treats the path as a file.
Now only checks with trailing slash since childrenIgnored is only called
for directories. Bare-name patterns (e.g. `local`) still match per gitignore spec.

Fixes #596

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test(ignore): add edge-case for bare `!dir` negation pattern

Verifies that `!iOS` (without trailing slash) also un-ignores the iOS/
directory — confirms the `ignore` package normalizes both `!dir` and
`!dir/` forms consistently when tested with a trailing-slash path.

Addresses non-blocking review suggestion on #654.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs(ignore): link ignore package docs for bare-name normalization

Adds references to the `ignore` package documentation in both the
childrenIgnored comment and the bare-negation test, explaining why
`!iOS` (without trailing slash) also re-includes the iOS/ directory.

Addresses non-blocking review suggestion on #654.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-06 08:32:47 +01:00
Abhigyan Patwari 6ead5e5986 fix(setup): prefer global gitnexus binary over npx for MCP config (#653) 2026-04-06 07:46:10 +01:00
Abhigyan Patwari 14791ded4a fix(server): return clean CORS rejection instead of 500 error (#646) 2026-04-06 07:45:19 +01:00
tantkandtian kian tan 9eeb20bb04 fix: replace Array.push(...spread) with loop to prevent stack overflow (#650)
* fix: replace Array.push(...spread) with loop to prevent stack overflow

On large codebases (78K+ C files), deferred arrays in
runChunkedParseAndResolve accumulate 100K+ entries. The spread
operator in push(...array) puts every element on the call stack
as a function argument, exceeding the maximum call stack size.

Replace all 11 occurrences of `arr.push(...other)` with
`for (const _item of other) arr.push(_item)` which uses
constant stack space regardless of array size.

Fixes #649

* fix: also replace push(...spread) in parsing-processor.ts

* style: format pipeline.ts to match prettier config

---------

Co-authored-by: tian kian tan <tan@example.com>
2026-04-05 21:54:09 +01:00
Gergő Magyar 5a7c0fdbb1 feat: same-arity overload disambiguation via type-hash suffix (#651) (#658)
* feat: same-arity overload disambiguation via type-hash suffix (#651)

Add ~type1,type2 suffix to Method/Constructor node IDs when same-arity
overloads with different parameter types exist in the same class. Also add
$const suffix for C++ const-qualified method overloads via new isConst field.

Key changes:
- typeTagForId() detects same-arity collisions and appends ~typeTag
- constTagForId() detects const/non-const collisions and appends $const
- TS/JS excluded from type-hashing (overload signatures collapse to impl body)
- Sequential findEnclosingFunction fixed: falls through on ambiguous same-class
  candidates instead of picking first; fallback path includes typeTag + constTag
- Per-call-site integration tests across Java, C#, Kotlin, C++, TypeScript
- Cross-file + chain resolution tests for all 5 languages
- C++ isConst extraction via tree-sitter type_qualifier in function_declarator

1710 integration + 18 unit tests pass.

* fix: preserve generic/template args in type-hash, perf + type safety fixes

- Add rawType field to ParameterInfo preserving full type text (vector<int>)
  while type stays simplified (vector). typeTagForId uses rawType for tags.
- Populate rawType in all 11 language method extractors
- Add buildCollisionGroups() to pre-group methods by name#arity (O(N) once
  per class instead of O(N) per method call)
- Cache method extraction in call-processor findEnclosingFunction fallback
- Fix null guards on getLanguageFromFilename in all findEnclosing paths
- Tighten SKIP_TYPE_HASH_LANGUAGES to ReadonlySet<SupportedLanguages>
- Document ID stability invariant on first overload introduction
- C++ integration tests: template overloads (vector<int> vs vector<string>),
  cross-file template + chain resolution, out-of-class method definitions

1718 integration + 20 unit tests pass.

* fix: add rawType to method-extraction unit test assertions

All 26 parameter .toEqual() assertions in method-extraction.test.ts
needed the new rawType field added to match ParameterInfo schema change.

* perf: cache tempMap/groups per class, consolidate extractFromNode

- Cache derived method map + collision groups per classNode.id in
  parsing-processor (avoids rebuild per method in same class)
- Replace per-call extractFromNode with cached class extraction +
  funcName:line lookup in call-processor fallback (avoids AST walk
  per call site)
- Remove dead clearEnclosingFunctionCache export, fix JSDoc

* test: add sequential-path integration test for same-arity overloads

Add skipWorkers option to PipelineOptions to force sequential parsing.
New test suite verifies type-hash disambiguation produces identical
results through the sequential path (parsing-processor + call-processor
findEnclosingFunction) as the worker path.
2026-04-05 21:51:55 +01:00
Gergő Magyar 0561d24efd feat: METHOD_IMPLEMENTS edges, overload disambiguation, MethodExtractor unification (#574) (#642) 2026-04-04 18:41:47 +01:00
Abhigyan Patwari 153262304c fix(mcp): unify stdout silencing to prevent embedder/pool-adapter conflicts (#645) 2026-04-04 11:56:49 +01:00
Abhigyan Patwari 16cf4c503e fix(web): replace aggressive heartbeat disconnect with graceful reconnection (#643) 2026-04-04 11:56:35 +01:00
Abhigyan Patwari 57951a197b fix(web): scope all backend calls to the active repo, not always the first (#644) 2026-04-04 11:55:34 +01:00
Gergő MagyarandClaude Opus 4.6 63fc4c795f feat: MethodExtractor configs for Python, PHP, Swift, Dart, Rust, Ruby (#624)
* feat: MethodExtractor configs for Python, PHP, Swift, Dart, Rust, Ruby with exhaustive integration tests

Add per-language MethodExtractionConfig for all remaining tree-sitter languages
(RFC #568 PR 2). Each config follows the established createMethodExtractor()
factory pattern — no new types, no parse-worker changes.

Configs:
- Python: @abstractmethod, @staticmethod/@classmethod, *args/**kwargs, type hints, _/__ visibility
- PHP: abstract/final/static keywords, PHP 8 #[] attributes, __construct/__destruct
- Swift: 5-level visibility, protocol-as-abstract, static/class methods, @ attributes
- Dart: _ convention visibility, abstract (no body), method_signature unwrapping
- Rust: pub visibility, &self receiver, trait_item + impl_item, #[] attributes
- Ruby: positional visibility via sibling-walk, singleton_method as static

Integration fixtures (18 directories) covering 3 resolution patterns:
- Method enrichment: parameterTypes, isAbstract, isFinal, annotations on graph nodes
- Overload dispatch: arity-based CALLS resolution via parameterTypes
- Abstract dispatch: abstract/concrete method distinction (Python, PHP, Rust, Swift)

Go deferred — requires factory changes for receiver-based method extraction.

Closes #571

* fix: address code review findings across 6 MethodExtractor configs

Fix all actionable items from the PR #624 deep-dive review:

Dart (critical — fixes 6 CI failures):
- isDartStatic: check children first, siblings as fallback
- isDartAbstract: handle declaration nodes for abstract methods
- extractSingleParam: detect required keyword as sibling token
- Add declaration to methodNodeTypes, mixin_declaration to typeDeclarationNodes
- Add member call query for variable assignments in tree-sitter-queries

Python:
- hasDecorator now matches dotted paths (e.g. @abc.abstractmethod)
- Fix version comment from ^0.23.6 to 0.23.4

PHP:
- Add enum_declaration to typeDeclarationNodes (PHP 8.1+)
- Add version comment for 0.23.12

Swift:
- Add isOverride using hasKeyword/hasModifier pattern

Rust:
- Fix version comment from ^0.23.2 to 0.23.1

Also: identifier fallback in generic.ts for mixin owner names,
Dart integration test label fix (Method vs Function), version
comment for tree-sitter-dart 1.0.0.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: Dart extension_declaration and Ruby module_function support

Dart:
- Add extension_declaration to typeDeclarationNodes and extension_body
  to bodyNodeTypes — extension methods are now extracted into the graph
- Add extension_declaration and mixin_declaration to CLASS_CONTAINER_TYPES
  for HAS_METHOD edge resolution

Ruby:
- module_function now maps to visibility 'private' in extractRubyVisibility
- module_function methods marked isStatic via backward-walk in isStatic
- Override semantics: private/public after module_function resets isStatic

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(go): Go MethodExtractor config with receiver-based extraction

Add Go as the 13th language with a per-language MethodExtractor config.
Go methods are top-level (not nested in struct bodies), so this adds
extractFromNode() to the MethodExtractor interface for direct method
node extraction without an enclosing class.

Config extracts:
- Name from field_identifier (methods) / identifier (functions)
- Return type including multi-return (first type from parameter_list)
- Parameters with variadic support
- Visibility via uppercase/lowercase convention
- Receiver type with pointer unwrapping (*User → User)
- isStatic for functions (no receiver)

Infrastructure:
- extractOwnerName optional hook on MethodExtractionConfig
- extractFromNode on MethodExtractor (factory auto-implements)
- Parse-worker uses extractFromNode when no enclosing class found
- method_declaration added to CLASS_CONTAINER_TYPES

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test: method enrichment integration tests for 7 languages + TS abstract class fix

Add method-enrichment integration test fixtures and test blocks for
Go, C++, Java, Kotlin, TypeScript, JavaScript, and C#. Each fixture
tests: class detection, HAS_METHOD edges, EXTENDS edges, isAbstract,
isStatic, annotations, parameterTypes, and CALLS edge resolution.

Fixes found during testing:
- Remove method_declaration from CLASS_CONTAINER_TYPES (added for Go
  but broke Java/C# HAS_METHOD edge resolution — method_declaration
  is also Java's method node type)
- Add abstract_class_declaration query to TypeScript tree-sitter
  queries (was missing, so abstract classes were invisible to pipeline)

1699 integration tests pass across 20 test files, 0 regressions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: format typeDeclarationNodes array for better readability in PHP config

* fix: Go interface methods + Rust impl-for-Struct owner resolution

Go:
- Add method_elem to methodNodeTypes so interface method signatures
  are extractable as abstract methods
- Integration test: Animal interface detected, Speak isAbstract,
  CALLS edges from app.go

Rust:
- Add extractOwnerName to resolve impl Trait for Struct to the
  concrete Struct (not the Trait) — fixes method misattribution
- Fix findEnclosingClassId to generate Struct: label (not Impl:)
  for impl blocks so HAS_METHOD edges resolve to struct nodes
- Tighten abstract-dispatch test: assert SqlRepo owns find/save

generic.ts:
- Fix extractOwnerName fallback: when hook returns a value, skip
  both name-field and type_identifier scan (was overwriting result)

1703 integration tests pass, 0 regressions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: code review response — Rust impl label, Swift params, Dart async, sequential methodExtractor

Address code review findings from PR #624:

- ast-helpers: Rust `impl Trait for Struct` uses Struct label (matches existing
  graph node), plain `impl Struct` uses Impl label (matches definition.impl)
- swift: fix parameter type extraction (user_type not type_annotation), detect
  default values as function_declaration siblings, add version comment
- dart: isDartAsync now detects async*/sync* generators, add clarifying comment
  for declaration nodes in extension bodies
- python: correct isFinal comment (PEP 591 @typing.final exists, just not modeled)
- parsing-processor: port methodExtractor enrichment to sequential path so
  isAbstract/isStatic/visibility/annotations/isFinal populate on <15-file repos
- tests: remove silent `if (prop !== undefined)` guards, assert properties
  directly, fix label queries (Dart Method vs Function, Swift Method for protocol
  methods), add Rust HAS_METHOD sourceLabel tests, Swift parameterTypes tests,
  and Dart async/sync* integration tests with fixture

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: Rust grammar gap + qualified method IDs to resolve same-file collisions

Phase 1 — Rust grammar:
- Add function_signature_item query to RUST_QUERIES so abstract trait methods
  (fn speak(&self) -> String;) become graph nodes with isAbstract=true

Phase 2 — Qualified method IDs:
- findEnclosingClassInfo returns {classId, className} for AST-based class lookup
- Both parsing paths (sequential + worker) qualify method/property IDs with
  enclosing class: Method:file:ClassName.method instead of Method:file:method
- extractFuncNameFromSourceId handles ClassName.method format
- Fixes silent data loss when same-name methods in different classes shared a
  file (e.g., Animal.speak and Dog.speak both now exist as distinct graph nodes)

Test updates:
- Rust: abstract+concrete trait methods both verified, function count adjusted
- Python: static method disambiguation now emits 2 CALLS edges (correct — no
  more ID collision masking the second call)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: owner-aware resolution for qualified method IDs

Address Codex adversarial review findings after qualified ID change:

- findEnclosingFunction: disambiguate candidates by ownerId when multiple
  same-name methods exist in file; qualify fallback-generated IDs
- findEnclosingFunctionId (worker): qualify sourceIds with enclosing class
  name so CALLS source attribution matches definition-phase node IDs
- buildExportedTypeMapFromGraph: use lookupExactAll + nodeId match instead
  of lookupExactFull which returns first definition for bare name

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: methodExtractor variadic arity, return type preservation, PHP abstract dispatch

Three bugs in the methodExtractor enrichment path broke 17 integration tests:

1. Variadic parameterCount: buildMethodProps and parse-worker set
   parameterCount = info.parameters.length even for variadic functions,
   causing arity filtering to reject valid calls. Now checks isVariadic
   and sets parameterCount = undefined (matching extractMethodSignature).

2. C++ bare `...` token: extractCppParameters only iterated named
   children, missing the unnamed `...` token in C-style variadics like
   log_entry(const char* fmt, ...). Added fallback scan of all children.

3. Return type stripping: All 11 language extractReturnType functions
   used extractSimpleTypeName() which strips generic parameters
   (List<User> → "List", Task<User> → "Task"). Changed to .text?.trim()
   to preserve full generic types needed for for-loop iterable resolution,
   async-await binding, and return-type inference.

Also fixes PHP abstract dispatch test that matched SqlRepository instead
of the interface due to ambiguous filePath.includes('Repository') filter,
and adds parent-walk fallback in PHP isAbstract for extractFromNode path.

* chore: remove plan and review artifacts from PR

* fix: address Round 4 review findings + infrastructure improvements

- Ruby: add singleton_class support for class << self methods (4 new tests)
- PHP: add enum_declaration to CLASS_CONTAINER_TYPES
- Dart: add mixin/extension labels to CONTAINER_TYPE_TO_LABEL
- Swift: add TODO for unverifiable struct/enum node types on Node 22
- C#: add grammar version comment (0.23.1)
- Ruby: fix version comment range to pin (0.23.1)
- Rust/ast-helpers: add cross-reference comments for impl_item duplication
- ast-helpers: document CLASS_CONTAINER_TYPES ↔ typeDeclarationNodes invariant
- generic.ts: replace Array.includes with Set for O(1) dedup in addNestedBodies
- Go/Python/Ruby: align isAbstract signature with 2-param interface contract
- CLAUDE.md: fix malformed backtick around gitnexus:start HTML comment
- parsing-processor: add per-class method extraction cache (eliminates O(N*M))
- ast-helpers: add scoped_type_identifier to impl_item resolution
- call-processor: add dev-mode warnings at silent candidates[0] fallbacks
- MCP context(): surface methodMetadata for Method/Function/Constructor nodes
- resources.ts: update schema to list all stored Method properties

* fix: singleton_class HAS_METHOD edge regression in findEnclosingClassInfo

singleton_class (class << self) was added to CLASS_CONTAINER_TYPES but
has no name field — its receiver `self` has node type 'self', not
'identifier'. findEnclosingClassInfo now walks up to the enclosing
class/module to inherit its name, matching ruby.ts:extractOwnerName.

Also fixes findEnclosingClassNode in parse-worker.ts to skip
singleton_class and return the actual class/module node.

Adds integration test assertions for from_habitat (class << self method):
HAS_METHOD edge from Animal, isStatic=true, parameterCount=1.

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 16:11:31 +01:00
489 changed files with 57111 additions and 6182 deletions
+4 -3
View File
@@ -1,10 +1,10 @@
<!-- version: 1.2.0 -->
<!-- version: 1.3.0 -->
<!--
Metadata: version, last reviewed, scope, model policy, reference docs, changelog.
Last updated: 2026-03-22
-->
Last reviewed: 2026-03-24
Last reviewed: 2026-04-13
**Project:** GitNexus · **Environment:** dev · **Maintainer:** repository maintainers (see GitHub)
@@ -54,6 +54,7 @@ Generic “core standards” playbooks are often long and stack-specific. For th
| Date | Version | Change |
|------|---------|--------|
| 2026-04-13 | 1.3.0 | Updated GitNexus index stats after DAG refactor. |
| 2026-03-24 | 1.2.0 | Fixed gitnexus:start block duplication (was inlined in Reference Docs bullet). |
| 2026-03-23 | 1.1.0 | Updated agent instructions (sections, references, Cursor layout). |
| 2026-03-22 | 1.0.0 | Added structured agent header and changelog. |
@@ -63,7 +64,7 @@ Generic “core standards” playbooks are often long and stack-specific. For th
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus**. Use the GitNexus MCP tools to understand code, assess impact, and navigate safely. For current symbol stats, run `npx gitnexus analyze` and inspect `.gitnexus/meta.json`.
This project is indexed by GitNexus as **GitNexus** (4325 symbols, 10556 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
+117 -2
View File
@@ -16,7 +16,7 @@ This repository is a **monorepo** with two main products: the **CLI / MCP packag
1. **Ingestion** (`gitnexus analyze`)
- Entry: `gitnexus/src/cli/analyze.ts` → `runPipelineFromRepo` in `gitnexus/src/core/ingestion/pipeline.ts`.
- Walks the git working tree, parses supported languages via **Tree-sitter**, resolves imports/calls/inheritance, detects **communities** and **processes** (execution flows), and builds an in-memory **knowledge graph** (`gitnexus/src/core/graph/`).
- The pipeline is structured as a **DAG (Directed Acyclic Graph)** of named phases (see [Pipeline Phase DAG](#pipeline-phase-dag) below).
- Output is loaded into **LadybugDB** under **`.gitnexus/`** at the repo root (`lbug/`, `meta.json`, etc.). Optional **FTS** indexes and **embeddings** attach to the same store.
- The repo is registered in **`~/.gitnexus/registry.json`** so MCP can find it from any working directory.
@@ -49,7 +49,7 @@ This repository is a **monorepo** with two main products: the **CLI / MCP packag
| If you are changing… | Start in… |
|----------------------|-----------|
| CLI commands / flags | `gitnexus/src/cli/` (`index.ts`, per-command modules). |
| Parsing or graph construction | `gitnexus/src/core/ingestion/` (pipeline, processors, resolvers, type-extractors). |
| Parsing or graph construction | `gitnexus/src/core/ingestion/pipeline-phases/` (individual phase files), `pipeline.ts` (orchestrator). |
| Graph schema / DB access | `gitnexus/src/core/lbug/` (`schema.ts`, `lbug-adapter.ts`), `gitnexus/src/mcp/core/lbug-adapter.ts` if MCP-specific. |
| MCP protocol, tools, resources | `gitnexus/src/mcp/server.ts`, `tools.ts`, `resources.ts`. |
| Search ranking | `gitnexus/src/core/search/` (BM25, hybrid fusion). |
@@ -58,8 +58,123 @@ This repository is a **monorepo** with two main products: the **CLI / MCP packag
| Web UI behavior | `gitnexus-web/src/` (components, workers, graph client). |
| CI | `.github/workflows/*.yml`, `.github/actions/setup-gitnexus/`. |
## Pipeline Phase DAG
The ingestion pipeline is a DAG of named phases. Each phase is defined in its own file under `gitnexus/src/core/ingestion/pipeline-phases/` with explicit dependencies, typed inputs, and typed outputs.
```
scan → structure → [markdown, cobol] → parse → [routes, tools, orm]
→ crossFile → mro → communities → processes
```
### Phase files
| Phase | File | Dependencies | What it does |
|-------|------|-------------|--------------|
| `scan` | `scan.ts` | (root) | Walk repo filesystem, collect paths + sizes |
| `structure` | `structure.ts` | `scan` | Build File/Folder nodes + CONTAINS edges |
| `markdown` | `markdown.ts` | `structure` | Extract headings and cross-links from .md/.mdx |
| `cobol` | `cobol.ts` | `structure` | Regex-based COBOL/JCL extraction |
| `parse` | `parse.ts` + `parse-impl.ts` | `structure`, `markdown`, `cobol` | Chunked tree-sitter parse, import/call/heritage resolution |
| `routes` | `routes.ts` | `parse` | Route registry (Next.js, Expo, PHP, decorator-based) |
| `tools` | `tools.ts` | `parse` | MCP/RPC tool detection |
| `orm` | `orm.ts` | `parse` | Prisma/Supabase ORM query edges |
| `crossFile` | `cross-file.ts` + `cross-file-impl.ts` | `parse`, `routes`, `tools`, `orm` | Cross-file type propagation in topological order |
| `mro` | `mro.ts` | `crossFile` | Method Resolution Order, METHOD_OVERRIDES edges |
| `communities` | `communities.ts` | `mro` | Leiden community detection |
| `processes` | `processes.ts` | `communities`, `routes`, `tools` | Execution flow detection, Route/Tool → Process links |
### How to add a new phase
1. Create a new file in `pipeline-phases/` (e.g. `my-phase.ts`)
2. Define a `PipelinePhase<MyOutput>` object with `name`, `deps`, and `execute(ctx, deps)`
3. Export it from `pipeline-phases/index.ts`
4. Add it to the `buildPhaseList()` function in `pipeline.ts`
```typescript
// pipeline-phases/my-phase.ts
import type { PipelinePhase, PipelineContext, PhaseResult } from './types.js';
import { getPhaseOutput } from './types.js';
import type { ParseOutput } from './parse.js';
export interface MyPhaseOutput { /* ... */ }
export const myPhase: PipelinePhase<MyPhaseOutput> = {
name: 'myPhase',
deps: ['parse'], // runs after parse completes
async execute(ctx, deps) {
const { allPaths } = getPhaseOutput<ParseOutput>(deps, 'parse');
// ... do work, write to ctx.graph ...
return { /* typed output */ };
},
};
```
### DAG runner
The runner (`pipeline-phases/runner.ts`) validates the DAG at startup (detects cycles and missing deps via topological sort), then executes phases in dependency order. Each phase receives:
- `ctx: PipelineContext` — shared graph, repoPath, progress callback
- `deps: Map<string, PhaseResult>` — outputs from all upstream phases
## Known limitations
### Overloaded method resolution
Method and Constructor node IDs include an arity suffix (`#<paramCount>`) to
disambiguate overloaded methods. Two overloads with different parameter counts
produce distinct graph nodes: `Method:file:Class.method#1` vs
`Method:file:Class.method#2`.
**Same-arity overload disambiguation:** When two overloads share the same
parameter count but differ in types (e.g. `save(int)` vs `save(String)`), a
type-hash suffix `~type1,type2` is appended to produce distinct node IDs:
`Method:file:Class.save#1~int` vs `Method:file:Class.save#1~String`. The suffix
is only added when a same-arity collision is detected within a class and all
parameters have non-null type annotations. Languages without type info (Python,
Ruby, JS) fall back to arity-only IDs. TypeScript/JavaScript overload signatures
are intentionally excluded from type-hashing because they are declaration-only
contracts that should collapse to the implementation body's node ID. See issue
\#651.
**C++ const-qualified overload disambiguation:** Methods overloaded by const
qualification (e.g. `begin()` vs `begin() const`) are disambiguated via an
`isConst` property and a `$const` ID suffix appended to the const-qualified
variant when a non-const collision exists. The `$const` suffix appears after the
type-hash suffix: e.g. `Method:file:Container.begin#0$const`.
**Generic/template type preservation in type-hash:** The type-hash suffix uses
`rawType` (full AST text including generic/template args) rather than the
simplified `type` from `extractSimpleTypeName`. This means C++ template overloads
like `process(vector<int>)` vs `process(vector<string>)` produce distinct IDs:
`~vector<int>` vs `~vector<std::string>`. Java generic overloads like
`process(List<String>)` vs `process(List<Integer>)` are a compile error due to
type erasure, so this gap is theoretical for Java.
**ID stability on first overload:** Type and const tags are collision-only. When
a class has `save(int)` as its only `save` method, the ID is `save#1` (no tag).
Adding `save(String)` changes the original to `save#1~int`. This is correct for
fresh analysis but means IDs are not stable across overload additions. Future
incremental re-analysis should account for this.
**Variadic method matching:** When one side is variadic (`parameterCount`
undefined) and the other has a fixed count, `METHOD_IMPLEMENTS` edges are
emitted with confidence 0.7 instead of 1.0. Variadic methods like
`foo(String... args)` may superficially match `foo(String s)` by type but
are not guaranteed to be interchangeable across all languages (Java/Kotlin
accept this via varargs sugar; TypeScript, C#, Rust do not).
**Confidence tiering** for `METHOD_IMPLEMENTS` edges:
| Match quality | Confidence | When |
|---|---|---|
| Exact parameter types match | 1.0 | Both sides have `parameterTypes` arrays and they match |
| Arity (count) matches | 1.0 | Both sides have `parameterCount`, types unavailable |
| Variadic vs fixed | 0.7 | One side is variadic, other has fixed count |
| Lenient (insufficient info) | 0.7 | One or both sides lack type and count data |
## Related docs
- [MIGRATION.md](MIGRATION.md) — breaking changes and migration guidance.
- [RUNBOOK.md](RUNBOOK.md) — operational commands and recovery.
- [GUARDRAILS.md](GUARDRAILS.md) — safety boundaries for humans and agents.
- [TESTING.md](TESTING.md) — how to run tests.
+206 -3
View File
@@ -1,10 +1,10 @@
<!-- version: 1.2.0 -->
<!-- version: 1.3.0 -->
<!--
Metadata: version, last reviewed, scope, model policy, reference docs, changelog.
Last updated: 2026-03-22
-->
Last reviewed: 2026-03-24
Last reviewed: 2026-04-13
**Project:** GitNexus · **Environment:** dev · **Maintainer:** repository maintainers (see GitHub)
@@ -41,6 +41,7 @@ If always-on instructions grow, load deep conventions via conditional reads (e.g
| Date | Version | Change |
|------|---------|--------|
| 2026-04-13 | 1.3.0 | Updated GitNexus index stats after DAG refactor. |
| 2026-03-24 | 1.2.0 | Removed duplicated gitnexus:start block and scope table; replaced with pointers to AGENTS.md. |
| 2026-03-23 | 1.1.0 | Updated agent instructions to match AGENTS.md. |
| 2026-03-22 | 1.0.0 | Added structured header and changelog. |
@@ -49,4 +50,206 @@ If always-on instructions grow, load deep conventions via conditional reads (e.g
## GitNexus rules
GitNexus MCP rules are in the `<!-- gitnexus:start -->` … `<!-- gitnexus:end -->` block in **[AGENTS.md](AGENTS.md)** — load that section when working with MCP tools or the graph index.
GitNexus MCP rules are in the `<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (4325 symbols, 10556 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## When Debugging
1. `gitnexus_query({query: "<error or symptom>"})` — find execution flows related to the issue
2. `gitnexus_context({name: "<suspect function>"})` — see all callers, callees, and process participation
3. `READ gitnexus://repo/GitNexus/process/{processName}` — trace the full execution flow step by step
4. For regressions: `gitnexus_detect_changes({scope: "compare", base_ref: "main"})` — see what your branch changed
## When Refactoring
- **Renaming**: MUST use `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` first. Review the preview — graph edits are safe, text_search edits need manual review. Then run with `dry_run: false`.
- **Extracting/Splitting**: MUST run `gitnexus_context({name: "target"})` to see all incoming/outgoing refs, then `gitnexus_impact({target: "target", direction: "upstream"})` to find all external callers before moving code.
- After any refactor: run `gitnexus_detect_changes({scope: "all"})` to verify only expected files changed.
## Never Do
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Tools Quick Reference
| Tool | When to use | Command |
|------|-------------|---------|
| `query` | Find code by concept | `gitnexus_query({query: "auth validation"})` |
| `context` | 360-degree view of one symbol | `gitnexus_context({name: "validateUser"})` |
| `impact` | Blast radius before editing | `gitnexus_impact({target: "X", direction: "upstream"})` |
| `detect_changes` | Pre-commit scope check | `gitnexus_detect_changes({scope: "staged"})` |
| `rename` | Safe multi-file rename | `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` |
| `cypher` | Custom graph queries | `gitnexus_cypher({query: "MATCH ..."})` |
## Impact Risk Levels
| Depth | Meaning | Action |
|-------|---------|--------|
| d=1 | WILL BREAK — direct callers/importers | MUST update these |
| d=2 | LIKELY AFFECTED — indirect deps | Should test |
| d=3 | MAY NEED TESTING — transitive | Test if critical path |
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
## Self-Check Before Finishing
Before completing any code modification task, verify:
1. `gitnexus_impact` was run for all modified symbols
2. No HIGH/CRITICAL risk warnings were ignored
3. `gitnexus_detect_changes()` confirms changes match expected scope
4. All d=1 (WILL BREAK) dependents were updated
## Keeping the Index Fresh
After committing code changes, the GitNexus index becomes stale. Re-run analyze to update it:
```bash
npx gitnexus analyze
```
If the index previously included embeddings, preserve them by adding `--embeddings`:
```bash
npx gitnexus analyze --embeddings
```
To check whether embeddings exist, inspect `.gitnexus/meta.json` — the `stats.embeddings` field shows the count (0 means no embeddings). **Running analyze without `--embeddings` will delete any previously generated embeddings.**
> Claude Code users: A PostToolUse hook handles this automatically after `git commit` and `git merge`.
## CLI
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
<!-- gitnexus:end -->` block in **[AGENTS.md](AGENTS.md)** — load that section when working with MCP tools or the graph index.
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (3298 symbols, 7954 relationships, 185 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## When Debugging
1. `gitnexus_query({query: "<error or symptom>"})` — find execution flows related to the issue
2. `gitnexus_context({name: "<suspect function>"})` — see all callers, callees, and process participation
3. `READ gitnexus://repo/GitNexus/process/{processName}` — trace the full execution flow step by step
4. For regressions: `gitnexus_detect_changes({scope: "compare", base_ref: "main"})` — see what your branch changed
## When Refactoring
- **Renaming**: MUST use `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` first. Review the preview — graph edits are safe, text_search edits need manual review. Then run with `dry_run: false`.
- **Extracting/Splitting**: MUST run `gitnexus_context({name: "target"})` to see all incoming/outgoing refs, then `gitnexus_impact({target: "target", direction: "upstream"})` to find all external callers before moving code.
- After any refactor: run `gitnexus_detect_changes({scope: "all"})` to verify only expected files changed.
## Never Do
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Tools Quick Reference
| Tool | When to use | Command |
|------|-------------|---------|
| `query` | Find code by concept | `gitnexus_query({query: "auth validation"})` |
| `context` | 360-degree view of one symbol | `gitnexus_context({name: "validateUser"})` |
| `impact` | Blast radius before editing | `gitnexus_impact({target: "X", direction: "upstream"})` |
| `detect_changes` | Pre-commit scope check | `gitnexus_detect_changes({scope: "staged"})` |
| `rename` | Safe multi-file rename | `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` |
| `cypher` | Custom graph queries | `gitnexus_cypher({query: "MATCH ..."})` |
## Impact Risk Levels
| Depth | Meaning | Action |
|-------|---------|--------|
| d=1 | WILL BREAK — direct callers/importers | MUST update these |
| d=2 | LIKELY AFFECTED — indirect deps | Should test |
| d=3 | MAY NEED TESTING — transitive | Test if critical path |
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
## Self-Check Before Finishing
Before completing any code modification task, verify:
1. `gitnexus_impact` was run for all modified symbols
2. No HIGH/CRITICAL risk warnings were ignored
3. `gitnexus_detect_changes()` confirms changes match expected scope
4. All d=1 (WILL BREAK) dependents were updated
## Keeping the Index Fresh
After committing code changes, the GitNexus index becomes stale. Re-run analyze to update it:
```bash
npx gitnexus analyze
```
If the index previously included embeddings, preserve them by adding `--embeddings`:
```bash
npx gitnexus analyze --embeddings
```
To check whether embeddings exist, inspect `.gitnexus/meta.json` — the `stats.embeddings` field shows the count (0 means no embeddings). **Running analyze without `--embeddings` will delete any previously generated embeddings.**
> Claude Code users: A PostToolUse hook handles this automatically after `git commit` and `git merge`.
## CLI
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
<!-- gitnexus:end -->
+27
View File
@@ -0,0 +1,27 @@
# Migration Guide
## OVERRIDES → METHOD_OVERRIDES (PR #642)
The `OVERRIDES` relationship type has been renamed to `METHOD_OVERRIDES` for
consistency with the new `METHOD_IMPLEMENTS` edge type.
### Do I need to migrate?
**No.** Backward compatibility is handled automatically at runtime:
- `local-backend.ts` dual-reads both `OVERRIDES` and `METHOD_OVERRIDES` in all
impact-analysis and context queries. Existing stored graphs with `OVERRIDES`
edges continue to return correct results without any manual intervention.
- The `REL_TYPES` array in `schema-constants.ts` includes both names so Cypher
queries that reference either will work.
### What happens on re-index?
Running `npx gitnexus analyze` on a repository produces `METHOD_OVERRIDES`
edges going forward. The old `OVERRIDES` edges are replaced as part of the
normal full re-index.
### When will the legacy alias be removed?
The `OVERRIDES` compat alias will remain until a future major version. Removal
will be announced in this file and in the changelog before it happens.
+1 -2
View File
@@ -52,7 +52,7 @@ https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
| **What** | Index repos locally, connect AI agents via MCP | Visual graph explorer + AI chat in browser |
| **For** | Daily development with Cursor, Claude Code, Codex, Windsurf, OpenCode | Quick exploration, demos, one-off analysis |
| **Scale** | Full repos, any size | Limited by browser memory (~5k files), or unlimited via backend mode |
| **Install** | `npm install -g gitnexus` | No install —[gitnexus.vercel.app](https://gitnexus.vercel.app) |
| **Install** | `npm install -g gitnexus` | No install — [gitnexus.vercel.app](https://gitnexus.vercel.app) |
| **Storage** | LadybugDB native (fast, persistent) | LadybugDB WASM (in-memory, per session) |
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
| **Privacy** | Everything local, no network | Everything in-browser, no server |
@@ -119,7 +119,6 @@ To configure MCP for your editor, run `npx gitnexus setup` once — or set it up
| **Codex** | Yes | Yes | — | MCP + Skills |
| **Windsurf** | Yes | — | — | MCP |
| **OpenCode** | Yes | Yes | — | MCP + Skills |
| **Codex** | Yes | — | — | MCP |
> **Claude Code** gets the deepest integration: MCP tools + agent skills + PreToolUse hooks that enrich searches with graph context + PostToolUse hooks that auto-reindex after commits.
+2 -1
View File
@@ -97,7 +97,8 @@ export type RelationshipType =
| 'CONTAINS'
| 'CALLS'
| 'INHERITS'
| 'OVERRIDES'
| 'METHOD_OVERRIDES'
| 'METHOD_IMPLEMENTS'
| 'IMPORTS'
| 'USES'
| 'DEFINES'
+1
View File
@@ -19,6 +19,7 @@ export type { NodeTableName, RelType } from './lbug/schema-constants.js';
// Language support
export { SupportedLanguages } from './languages.js';
export { getLanguageFromFilename, getSyntaxLanguageFromFilename } from './language-detection.js';
export type { MroStrategy } from './mro-strategy.js';
// Pipeline progress
export type { PipelinePhase, PipelineProgress } from './pipeline.js';
+3 -1
View File
@@ -55,7 +55,9 @@ export const REL_TYPES = [
'HAS_METHOD',
'HAS_PROPERTY',
'ACCESSES',
'OVERRIDES',
'METHOD_OVERRIDES',
'OVERRIDES', // Legacy compat alias — kept until all stored indexes are migrated
'METHOD_IMPLEMENTS',
'MEMBER_OF',
'STEP_IN_PROCESS',
'HANDLES_ROUTE',
+23
View File
@@ -0,0 +1,23 @@
/**
* MRO (Method Resolution Order) strategy — shared between CLI and any
* future consumer that reasons about multiple-inheritance semantics.
*
* Lives in `gitnexus-shared` so the low-level resolution module
* (`core/ingestion/model/resolve.ts`) does not need to import from
* `languages/` — keeping the `model/` layer free of language-registry
* coupling.
*
* Strategy semantics:
* - `first-wins`: BFS ancestor walk, first match wins (default).
* - `leftmost-base`: BFS ancestor walk, leftmost base wins (C++).
* - `c3`: C3-linearized ancestor order, first match wins (Python).
* - `implements-split`: BFS walk, first match wins (Java/C#/Kotlin) — full
* interface-default ambiguity is handled at graph level.
* - `qualified-syntax`: No auto-resolution (Rust — requires `<T as Trait>::m`).
*/
export type MroStrategy =
| 'first-wins'
| 'c3'
| 'leftmost-base'
| 'implements-split'
| 'qualified-syntax';
@@ -0,0 +1,109 @@
import { test, expect } from '@playwright/test';
/**
* E2E tests for heartbeat disconnect/reconnect behavior.
*
* Verifies the key regression: when the heartbeat fails, the UI shows a
* "reconnecting" banner instead of resetting to the onboarding screen.
*
* Strategy: block /api/heartbeat via route interception BEFORE loading the
* graph. The heartbeat EventSource can never connect, so onReconnecting
* fires on the first retry attempt. This reliably tests the banner behavior
* without depending on setOffline timing (which varies across CI environments).
*/
const BACKEND_URL = process.env.BACKEND_URL ?? 'http://localhost:4747';
const FRONTEND_URL = process.env.FRONTEND_URL ?? 'http://localhost:5173';
test.beforeAll(async () => {
if (process.env.E2E) return;
try {
const [backendRes, frontendRes] = await Promise.allSettled([
fetch(`${BACKEND_URL}/api/repos`),
fetch(FRONTEND_URL),
]);
if (
backendRes.status === 'rejected' ||
(backendRes.status === 'fulfilled' && !backendRes.value.ok)
) {
test.skip(true, 'gitnexus serve not available');
return;
}
if (
frontendRes.status === 'rejected' ||
(frontendRes.status === 'fulfilled' && !frontendRes.value.ok)
) {
test.skip(true, 'Vite dev server not available');
return;
}
if (backendRes.status === 'fulfilled') {
const repos = await backendRes.value.json();
if (!repos.length) {
test.skip(true, 'No indexed repos');
return;
}
}
} catch {
test.skip(true, 'servers not available');
}
});
test.describe('Heartbeat Reconnect', () => {
test('shows reconnecting banner instead of onboarding reset when heartbeat is unavailable', async ({
page,
}) => {
// Block the heartbeat BEFORE navigating — the EventSource will fail
// immediately on every connection attempt, triggering onReconnecting.
await page.route('**/api/heartbeat', (route) => route.abort('connectionrefused'));
// Load the app and connect to a repo (all other endpoints work normally)
await page.goto('/');
const landingCard = page.locator('[data-testid="landing-repo-card"]').first();
try {
await landingCard.waitFor({ state: 'visible', timeout: 15_000 });
await landingCard.click();
} catch {
// auto-connect may skip the landing screen
}
// Wait for graph to load (heartbeat is blocked, but graph loads fine)
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// The reconnecting banner should appear (heartbeat is failing)
const banner = page.getByText('Server connection lost');
await expect(banner).toBeVisible({ timeout: 15_000 });
// The graph canvas should STILL be visible — NOT reset to onboarding
await expect(page.locator('canvas').first()).toBeVisible();
});
test('banner clears when heartbeat becomes available', async ({ page }) => {
// Start with heartbeat blocked
await page.route('**/api/heartbeat', (route) => route.abort('connectionrefused'));
await page.goto('/');
const landingCard = page.locator('[data-testid="landing-repo-card"]').first();
try {
await landingCard.waitFor({ state: 'visible', timeout: 15_000 });
await landingCard.click();
} catch {
// auto-connect may skip the landing screen
}
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// Verify banner appears
const banner = page.getByText('Server connection lost');
await expect(banner).toBeVisible({ timeout: 15_000 });
// Unblock heartbeat — the real server is running, so reconnect will succeed
await page.unroute('**/api/heartbeat');
// Banner should disappear as heartbeat reconnects
await expect(banner).not.toBeVisible({ timeout: 30_000 });
// Graph should still be there
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible();
});
});
+112
View File
@@ -0,0 +1,112 @@
import { test, expect } from '@playwright/test';
/**
* E2E tests for multi-repo scoping and URL persistence.
*
* Verifies that:
* - Connecting via ?server= loads data and sets ?project= in the URL
* - The repo name appears in the UI after connecting
* - F5 with ?server=&project= reconnects to the correct repo
*
* Runs against the single indexed repo in CI — validates the plumbing
* works end-to-end even with one repo.
*/
const BACKEND_URL = process.env.BACKEND_URL ?? 'http://localhost:4747';
const FRONTEND_URL = process.env.FRONTEND_URL ?? 'http://localhost:5173';
let firstRepoName: string;
test.beforeAll(async () => {
if (process.env.E2E) {
// Still need to fetch the repo name for assertions
try {
const res = await fetch(`${BACKEND_URL}/api/repos`);
const repos = await res.json();
firstRepoName = repos[0]?.name ?? '';
} catch {
firstRepoName = '';
}
return;
}
try {
const [backendRes, frontendRes] = await Promise.allSettled([
fetch(`${BACKEND_URL}/api/repos`),
fetch(FRONTEND_URL),
]);
if (
backendRes.status === 'rejected' ||
(backendRes.status === 'fulfilled' && !backendRes.value.ok)
) {
test.skip(true, 'gitnexus serve not available');
return;
}
if (
frontendRes.status === 'rejected' ||
(frontendRes.status === 'fulfilled' && !frontendRes.value.ok)
) {
test.skip(true, 'Vite dev server not available');
return;
}
if (backendRes.status === 'fulfilled') {
const repos = await backendRes.value.json();
if (!repos.length) {
test.skip(true, 'No indexed repos');
return;
}
firstRepoName = repos[0].name;
}
} catch {
test.skip(true, 'servers not available');
}
});
test.describe('Multi-Repo Scoping', () => {
test('auto-connect via ?server= sets ?project= in URL', async ({ page }) => {
// Navigate with ?server= param (the bookmarkable shortcut)
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
// Wait for graph to load
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// URL should now contain ?project= with the repo name
const url = new URL(page.url());
const project = url.searchParams.get('project');
expect(project).toBeTruthy();
expect(project).toBe(firstRepoName);
});
test('?server= is preserved in URL for F5 recovery', async ({ page }) => {
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// URL should still have ?server=
const url = new URL(page.url());
expect(url.searchParams.get('server')).toBeTruthy();
// F5 should reconnect (not show onboarding)
await page.reload();
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
});
test('node count in status bar matches backend data', async ({ page }) => {
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// Fetch expected node count from backend
const res = await fetch(`${BACKEND_URL}/api/repo?repo=${encodeURIComponent(firstRepoName)}`);
const repoInfo = await res.json();
const expectedNodes = repoInfo.stats?.nodes;
if (expectedNodes) {
// Status bar shows node count — use the status-ready area to avoid
// matching multiple elements (file tree, header may also show counts)
const statusBar = page.locator('footer');
const nodeText = statusBar.getByText(/\d+ nodes/).first();
await expect(nodeText).toBeVisible({ timeout: 10_000 });
const text = await nodeText.textContent();
const displayedNodes = parseInt(text?.match(/(\d+)\s*nodes/)?.[1] ?? '0', 10);
expect(displayedNodes).toBeGreaterThan(0);
}
});
});
+168
View File
@@ -0,0 +1,168 @@
import { test, expect } from '@playwright/test';
/**
* E2E tests for the repo-switching and false-404 fixes.
*
* Most tests use the live backend (same pattern as multi-repo-scoping.spec.ts).
* The 503 hold-queue test uses route interception to simulate a slow analysis.
*/
const BACKEND_URL = process.env.BACKEND_URL ?? 'http://localhost:4747';
const FRONTEND_URL = process.env.FRONTEND_URL ?? 'http://localhost:5173';
let firstRepoName: string;
test.beforeAll(async () => {
if (process.env.E2E) {
try {
const res = await fetch(`${BACKEND_URL}/api/repos`);
const repos = await res.json();
firstRepoName = repos[0]?.name ?? '';
} catch {
firstRepoName = '';
}
return;
}
try {
const [backendRes, frontendRes] = await Promise.allSettled([
fetch(`${BACKEND_URL}/api/repos`),
fetch(FRONTEND_URL),
]);
if (
backendRes.status === 'rejected' ||
(backendRes.status === 'fulfilled' && !backendRes.value.ok)
) {
test.skip(true, 'gitnexus serve not available');
return;
}
if (
frontendRes.status === 'rejected' ||
(frontendRes.status === 'fulfilled' && !frontendRes.value.ok)
) {
test.skip(true, 'Vite dev server not available');
return;
}
if (backendRes.status === 'fulfilled') {
const repos = await backendRes.value.json();
if (!repos.length) {
test.skip(true, 'No indexed repos');
return;
}
firstRepoName = repos[0].name;
}
} catch {
test.skip(true, 'servers not available');
}
});
// ── 1. Hold-queue: 503 → descriptive user message ────────────────────────────
test.describe('Hold-queue timeout error', () => {
test('shows descriptive message when /api/repo returns 503', async ({ page }, testInfo) => {
// Intercept only /api/repo (singular) — not /api/repos — to return a 503
// regex: /api/repo followed by end, ?, or # — NOT /api/repos
await page.route(/\/api\/repo(?!s)(\?.*)?$/, (route) =>
route.fulfill({
status: 503,
contentType: 'application/json',
body: JSON.stringify({
error: `Repository analysis for "${firstRepoName}" is taking longer than expected. Please try again in a moment.`,
}),
}),
);
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
// UI should show the 503 error message
await expect(page.getByText(/taking longer than expected/i)).toBeVisible({
timeout: 20_000,
});
await page.screenshot({ path: testInfo.outputPath('hold-queue-503.png') });
});
});
// ── 2. ?project= URL persistence ─────────────────────────────────────────────
test.describe('?project= URL persistence', () => {
test('?project= is set in URL after connecting via ?server=', async ({ page }) => {
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
const url = new URL(page.url());
const project = url.searchParams.get('project');
expect(project).toBeTruthy();
// first repo returned by the live backend
if (firstRepoName) expect(project).toBe(firstRepoName);
});
test('?project= is still present after F5 reload', async ({ page }) => {
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// After connect, URL has ?server=&project= — F5 re-uses both params
await page.reload();
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
const url = new URL(page.url());
expect(url.searchParams.get('project')).toBeTruthy();
});
});
// ── 3. ?project= + ?server= combined auto-connect ────────────────────────────
test.describe('?project= auto-connect', () => {
test('navigating with ?server=&project= connects to the correct repo', async ({
page,
}, testInfo) => {
if (!firstRepoName) test.skip(true, 'no repo name available');
await page.goto(
`/?server=${encodeURIComponent(BACKEND_URL)}&project=${encodeURIComponent(firstRepoName)}`,
);
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// ?project= in URL should match what we passed in
const url = new URL(page.url());
expect(url.searchParams.get('project')).toBe(firstRepoName);
await page.screenshot({ path: testInfo.outputPath('project-param-connect.png') });
});
});
// ── 4. Windows path normalization ─────────────────────────────────────────────
test.describe('Windows path normalization', () => {
test('project name uses basename when /api/repo returns a Windows-style repoPath', async ({
page,
}) => {
const repoName = firstRepoName || 'test-repo';
const windowsPath = `C:\\Users\\LENOVO\\.gitnexus\\repos\\${repoName}`;
// Mock /api/repo to return a Windows backslash path while keeping name correct
await page.route(/\/api\/repo(?!s)(\?.*)?$/, (route) =>
route.fulfill({
contentType: 'application/json',
body: JSON.stringify({
// intentionally omit `name` to force path-based extraction
path: windowsPath,
repoPath: windowsPath,
}),
}),
);
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// URL ?project= must be the short basename, NOT the full Windows path
const url = new URL(page.url());
const project = url.searchParams.get('project');
expect(project).toBeTruthy();
expect(project).not.toContain('\\');
expect(project).not.toContain('LENOVO');
expect(project).toBe(repoName);
});
});
+24 -26
View File
@@ -1,4 +1,4 @@
import { test, expect, type TestInfo } from '@playwright/test';
import { test, expect } from '@playwright/test';
/**
* E2E tests for the GitNexus web UI — exploring view features.
@@ -58,36 +58,41 @@ test.beforeAll(async () => {
* For these tests we require at least one indexed repo, so pick the first
* landing card when present and then wait for the exploring view.
*/
async function waitForGraphLoaded(page: import('@playwright/test').Page, testInfo: TestInfo) {
async function waitForGraphLoaded(page: import('@playwright/test').Page) {
await page.goto('/');
const landingCard = page.locator('[data-testid="landing-repo-card"]').first();
const landingCards = page.locator('[data-testid="landing-repo-card"]');
const preferredLandingCard = landingCards
.filter({ hasText: /GitNexus|local-integration/ })
.first();
try {
await landingCard.waitFor({ state: 'visible', timeout: 15_000 });
await landingCards.first().waitFor({ state: 'visible', timeout: 15_000 });
const landingCard =
(await preferredLandingCard.count()) > 0 ? preferredLandingCard : landingCards.first();
await landingCard.click();
} catch {
// Landing screen may not appear (e.g. ?server auto-connect)
}
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
await expect(page.getByText(/\d+ nodes/).first()).toBeVisible();
await page.screenshot({ path: testInfo.outputPath('graph-loaded.png') });
const statusBar = page.getByRole('contentinfo');
await expect(statusBar.getByText('Ready', { exact: true })).toBeVisible({ timeout: 45_000 });
await expect(statusBar).toContainText(/nodes/, {
timeout: 20_000,
});
}
test.describe('Server Connection & Graph Loading', () => {
test('selects a repo from landing and loads graph', async ({ page }, testInfo) => {
await waitForGraphLoaded(page, testInfo);
await page.screenshot({ path: testInfo.outputPath('graph-loaded-full.png'), fullPage: true });
test('selects a repo from landing and loads graph', async ({ page }) => {
await waitForGraphLoaded(page);
});
});
test.describe('Nexus AI', () => {
test('panel opens and agent initializes without error', async ({ page }, testInfo) => {
await waitForGraphLoaded(page, testInfo);
test('panel opens and agent initializes without error', async ({ page }) => {
await waitForGraphLoaded(page);
await page.getByRole('button', { name: 'Nexus AI' }).click();
await expect(page.getByText('Ask me anything')).toBeVisible({ timeout: 15_000 });
await page.screenshot({ path: testInfo.outputPath('nexus-ai-panel.png'), fullPage: true });
const errorBanner = page.getByText('Database not ready');
expect(await errorBanner.isVisible().catch(() => false)).toBe(false);
@@ -95,8 +100,8 @@ test.describe('Nexus AI', () => {
});
test.describe('Processes Panel', () => {
test('shows process list and View button works', async ({ page }, testInfo) => {
await waitForGraphLoaded(page, testInfo);
test('shows process list and View button works', async ({ page }) => {
await waitForGraphLoaded(page);
await page.getByRole('button', { name: 'Nexus AI' }).click();
await page.getByText('Processes').click();
@@ -104,7 +109,6 @@ test.describe('Processes Panel', () => {
await expect(page.locator('[data-testid="process-list-loaded"]')).toBeVisible({
timeout: 15_000,
});
await page.screenshot({ path: testInfo.outputPath('processes-panel.png'), fullPage: true });
const processRow = page.locator('[data-testid="process-row"]').first();
await expect(processRow).toBeVisible({ timeout: 10_000 });
@@ -114,14 +118,10 @@ test.describe('Processes Panel', () => {
await viewBtn.waitFor({ state: 'visible', timeout: 5_000 });
await viewBtn.click();
await expect(page.locator('[data-testid="process-modal"]')).toBeVisible({ timeout: 5_000 });
await page.screenshot({
path: testInfo.outputPath('process-view-clicked.png'),
fullPage: true,
});
});
test('lightbulb highlights nodes in graph', async ({ page }, testInfo) => {
await waitForGraphLoaded(page, testInfo);
test('lightbulb highlights nodes in graph', async ({ page }) => {
await waitForGraphLoaded(page);
await page.getByRole('button', { name: 'Nexus AI' }).click();
await page.getByText('Processes').click();
@@ -137,13 +137,12 @@ test.describe('Processes Panel', () => {
await lightbulb.waitFor({ state: 'visible', timeout: 5_000 });
await lightbulb.click();
await expect(processRow).toHaveClass(/bg-amber-950/, { timeout: 5_000 });
await page.screenshot({ path: testInfo.outputPath('after-highlight.png'), fullPage: true });
});
});
test.describe('Turn Off All Highlights', () => {
test('selecting a node dims others, button clears it', async ({ page }, testInfo) => {
await waitForGraphLoaded(page, testInfo);
test('selecting a node dims others, button clears it', async ({ page }) => {
await waitForGraphLoaded(page);
await expect(page.locator('canvas').first()).toBeVisible({ timeout: 10_000 });
@@ -160,6 +159,5 @@ test.describe('Turn Off All Highlights', () => {
await expect(highlightToggle).toHaveAttribute('title', 'Turn on AI highlights', {
timeout: 5_000,
});
await page.screenshot({ path: testInfo.outputPath('highlights-cleared.png'), fullPage: true });
});
});
+84 -49
View File
@@ -1,4 +1,4 @@
import { useCallback, useEffect, useRef } from 'react';
import { useCallback, useEffect, useRef, useState } from 'react';
import { AppStateProvider, useAppState } from './hooks/useAppState';
import { DropZone } from './components/DropZone';
import { LoadingOverlay } from './components/LoadingOverlay';
@@ -36,7 +36,6 @@ const AppContent = () => {
refreshLLMSettings,
initializeAgent,
startEmbeddingsWithFallback,
embeddingStatus,
codeReferences,
selectedNode,
isCodePanelOpen,
@@ -45,17 +44,25 @@ const AppContent = () => {
availableRepos,
setAvailableRepos,
switchRepo,
setCurrentRepo,
} = useAppState();
const graphCanvasRef = useRef<GraphCanvasHandle>(null);
const [serverDisconnected, setServerDisconnected] = useState(false);
const handleServerConnect = useCallback(
async (result: ConnectResult): Promise<void> => {
// Extract project name from repoPath
// Use the canonical repo name from the server response so all subsequent
// backend calls (queries, search, grep, readFile) scope to this repo.
const repoName = result.repoInfo.name;
const repoPath = result.repoInfo.repoPath ?? result.repoInfo.path;
const parts = (repoPath || '').split('/').filter((p) => p && !p.startsWith('.'));
const projectName = parts[parts.length - 1] || parts[0] || 'server-project';
// Normalize both Windows (\) and Unix (/) path separators before splitting
const projectName =
result.repoInfo.name ||
(repoPath || '').replace(/\\/g, '/').split('/').filter(Boolean).pop() ||
'server-project';
setProjectName(projectName);
setCurrentRepo(projectName);
// Build KnowledgeGraph from server data for visualization
const graph = createKnowledgeGraph();
@@ -67,6 +74,11 @@ const AppContent = () => {
}
setGraph(graph);
// Persist the active project in the URL for bookmarkability and F5 refresh resilience
const urlObj = new URL(window.location.href);
urlObj.searchParams.set('project', projectName);
window.history.replaceState(null, '', urlObj.toString());
// Transition directly to exploring view
setViewMode('exploring');
@@ -80,20 +92,26 @@ const AppContent = () => {
console.warn('Failed to initialize agent:', err);
}
},
[setViewMode, setGraph, setProjectName, initializeAgent, startEmbeddingsWithFallback],
[
setViewMode,
setGraph,
setProjectName,
setCurrentRepo,
initializeAgent,
startEmbeddingsWithFallback,
],
);
// Auto-connect when ?server query param is present (bookmarkable shortcut)
// Auto-connect when ?server or ?project query param is present (bookmarkable shortcut)
const autoConnectRan = useRef(false);
useEffect(() => {
if (autoConnectRan.current) return;
const params = new URLSearchParams(window.location.search);
if (!params.has('server')) return;
autoConnectRan.current = true;
const serverUrlParam = params.get('server');
const projectParam = params.get('project');
// Clean the URL so a refresh won't re-trigger
const cleanUrl = window.location.pathname + window.location.hash;
window.history.replaceState(null, '', cleanUrl);
if (!serverUrlParam && !projectParam) return;
autoConnectRan.current = true;
setProgress({
phase: 'extracting',
@@ -103,36 +121,45 @@ const AppContent = () => {
});
setViewMode('loading');
const serverUrl = params.get('server') || window.location.origin;
const serverUrl = serverUrlParam || window.location.origin;
const baseUrl = normalizeServerUrl(serverUrl);
connectToServer(serverUrl, (phase, downloaded, total) => {
if (phase === 'validating') {
setProgress({
phase: 'extracting',
percent: 5,
message: 'Connecting to server...',
detail: 'Validating server',
});
} else if (phase === 'downloading') {
const pct = total ? Math.round((downloaded / total) * 90) + 5 : 50;
const mb = (downloaded / (1024 * 1024)).toFixed(1);
setProgress({
phase: 'extracting',
percent: pct,
message: 'Downloading graph...',
detail: `${mb} MB downloaded`,
});
} else if (phase === 'extracting') {
setProgress({
phase: 'extracting',
percent: 97,
message: 'Processing...',
detail: 'Extracting file contents',
});
}
})
const tryConnect = async () => {
return await connectToServer(
serverUrl,
(phase, downloaded, total) => {
if (phase === 'validating') {
setProgress({
phase: 'extracting',
percent: 5,
message: 'Connecting to server...',
detail: 'Validating server',
});
} else if (phase === 'downloading') {
const pct = total ? Math.round((downloaded / total) * 90) + 5 : 50;
const mb = (downloaded / (1024 * 1024)).toFixed(1);
setProgress({
phase: 'extracting',
percent: pct,
message: 'Downloading graph...',
detail: `${mb} MB downloaded`,
});
} else if (phase === 'extracting') {
setProgress({
phase: 'extracting',
percent: 97,
message: 'Processing...',
detail: 'Extracting file contents',
});
}
},
undefined,
projectParam || undefined,
{ awaitAnalysis: true }, // enable backend hold-queue for repos still being analyzed
);
};
tryConnect()
.then(async (result) => {
await handleServerConnect(result);
setProgress(null);
@@ -169,21 +196,18 @@ const AppContent = () => {
// ── Server heartbeat: detect when server goes down while exploring ────────
// Uses SSE (EventSource) for instant detection — no polling delay.
// On disconnect: show a reconnecting banner instead of resetting to onboarding.
// The heartbeat retries indefinitely with capped backoff and recovers automatically.
useEffect(() => {
if (viewMode !== 'exploring') return;
const cleanup = connectHeartbeat(
() => {}, // onConnect — already connected, no action needed
() => {
// Server went down — return to onboarding
setViewMode('onboarding');
setGraph(null);
setProgress(null);
},
() => setServerDisconnected(false),
() => setServerDisconnected(true),
);
return cleanup;
}, [viewMode, setViewMode, setGraph, setProgress]);
}, [viewMode]);
// Render based on view mode
if (viewMode === 'onboarding') {
@@ -196,7 +220,12 @@ const AppContent = () => {
await handleServerConnect(result);
setProgress(null);
if (serverUrl) {
setServerBaseUrl(normalizeServerUrl(serverUrl));
const base = normalizeServerUrl(serverUrl);
setServerBaseUrl(base);
// Add ?server= so F5 reconnects to this server
const url = new URL(window.location.href);
url.searchParams.set('server', base);
window.history.replaceState(null, '', url.toString());
}
}}
/>
@@ -268,6 +297,12 @@ const AppContent = () => {
<StatusBar />
{serverDisconnected && (
<div className="fixed bottom-12 left-1/2 z-50 -translate-x-1/2 rounded-lg border border-yellow-500/30 bg-yellow-900/80 px-4 py-2 text-sm text-yellow-200 shadow-lg backdrop-blur">
Server connection lost — reconnecting&hellip;
</div>
)}
{/* Settings Panel (modal) */}
<SettingsPanel
isOpen={isSettingsPanelOpen}
@@ -54,6 +54,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
clearCodeReferences,
setSelectedNode,
codeReferenceFocus,
projectName,
} = useAppState();
const nodeById = useMemo(() => {
@@ -174,7 +175,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
return () => {
rafIds.forEach((id) => cancelAnimationFrame(id));
};
}, [codeReferenceFocus?.ts, aiReferences]);
}, [codeReferenceFocus, aiReferences]);
const refsWithSnippets = useMemo(() => {
return aiReferences.map((ref) => {
@@ -223,13 +224,14 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
const isWholeFile = selectedIsFile || startLine === undefined;
const options = isWholeFile
? undefined
? { repo: projectName }
: {
startLine: Math.max(0, startLine - CONTEXT_LINES),
endLine: (endLine ?? startLine) + CONTEXT_LINES,
repo: projectName,
};
readFile(selectedFilePath, options)
readFile(selectedFilePath, { ...options, repo: projectName || undefined })
.then((result) => {
if (!cancelled) {
setFileResult(result);
@@ -251,6 +253,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
selectedNode?.properties?.startLine,
selectedNode?.properties?.endLine,
selectedIsFile,
projectName,
]);
// Scroll to the selected node's startLine after content loads
@@ -201,7 +201,7 @@ interface OnboardingGuideProps {
}
export const OnboardingGuide = ({ isPolling }: OnboardingGuideProps) => {
const primary = isDev ? 'cd gitnexus && npm run serve' : 'npx gitnexus@latest serve';
const primary = isDev ? 'npm run --prefix gitnexus serve' : 'npx gitnexus@latest serve';
const termLabel = isDev ? 'Start backend' : 'Terminal';
// Step states: step 1 = copy command, step 2 = run/wait, step 3 = auto-connect
@@ -277,7 +277,9 @@ export const OnboardingGuide = ({ isPolling }: OnboardingGuideProps) => {
state={step2State}
number={2}
title={isPolling ? 'Waiting for server to start' : 'Paste and run in your terminal'}
description={isPolling ? undefined : 'Open a new terminal window, paste, and hit Enter.'}
description={
isPolling ? undefined : 'Open a terminal at the project root, paste, and hit Enter.'
}
>
{isPolling && <PollingBar />}
</StepRow>
+24 -13
View File
@@ -8,8 +8,10 @@ import {
Loader2,
AlertTriangle,
GitBranch,
ArrowDown,
} from '@/lib/lucide-icons';
import { useAppState } from '../hooks/useAppState';
import { useAutoScroll } from '../hooks/useAutoScroll';
import { ToolCallCard } from './ToolCallCard';
import { isProviderConfigured } from '../core/llm/settings-service';
import { MarkdownRenderer } from './MarkdownRenderer';
@@ -35,14 +37,11 @@ export const RightPanel = () => {
const [chatInput, setChatInput] = useState('');
const [activeTab, setActiveTab] = useState<'chat' | 'processes'>('chat');
const textareaRef = useRef<HTMLTextAreaElement>(null);
const messagesEndRef = useRef<HTMLDivElement>(null);
// Auto-scroll to bottom when messages update or while streaming
useEffect(() => {
if (messagesEndRef.current) {
messagesEndRef.current.scrollIntoView({ behavior: 'smooth' });
}
}, [chatMessages, isChatLoading]);
// Keep streamed replies pinned unless the user intentionally scrolls away from the bottom.
const { scrollContainerRef, messagesContainerRef, isAtBottom, scrollToBottom } = useAutoScroll(
chatMessages,
isChatLoading,
);
const resolveFilePathForUI = useCallback((_requestedPath: string): string | null => {
return null;
@@ -265,7 +264,7 @@ export const RightPanel = () => {
{/* Chat Content - only show when chat tab is active */}
{activeTab === 'chat' && (
<div className="flex flex-1 flex-col overflow-hidden">
<div className="relative flex flex-1 flex-col overflow-hidden">
{/* Status bar */}
<div className="flex items-center gap-2.5 border-b border-border-subtle bg-elevated/50 px-4 py-3">
<div className="ml-auto flex items-center gap-2">
@@ -291,7 +290,7 @@ export const RightPanel = () => {
)}
{/* Messages */}
<div className="scrollbar-thin flex-1 overflow-y-auto p-4">
<div ref={scrollContainerRef} className="scrollbar-thin flex-1 overflow-y-auto p-4">
{chatMessages.length === 0 ? (
<div className="flex h-full flex-col items-center justify-center px-4 text-center">
<div className="mb-4 flex h-14 w-14 items-center justify-center rounded-xl bg-gradient-to-br from-accent to-node-interface text-2xl shadow-glow">
@@ -315,7 +314,7 @@ export const RightPanel = () => {
</div>
</div>
) : (
<div className="flex flex-col gap-6">
<div ref={messagesContainerRef} className="flex flex-col gap-6">
{chatMessages.map((message) => (
<div key={message.id} className="animate-fade-in">
{/* User message - compact label style */}
@@ -391,10 +390,22 @@ export const RightPanel = () => {
))}
</div>
)}
{/* Scroll anchor for auto-scroll */}
<div ref={messagesEndRef} />
</div>
{/* Scroll to bottom */}
<button
aria-label="Scroll to bottom"
onClick={() => scrollToBottom()}
className={`absolute bottom-20 left-1/2 z-10 -translate-x-1/2 rounded-full border border-border-subtle bg-elevated px-3 py-1.5 text-xs text-text-secondary shadow-lg transition-all duration-200 hover:border-accent hover:text-accent ${
!isAtBottom && chatMessages.length > 0
? 'translate-y-0 opacity-100'
: 'pointer-events-none translate-y-2 opacity-0'
}`}
>
<ArrowDown className="mr-1 inline h-3.5 w-3.5" />
Scroll to bottom
</button>
{/* Input */}
<div className="border-t border-border-subtle bg-surface p-3">
<div className="flex items-end gap-2 rounded-xl border border-border-subtle bg-elevated px-3 py-2 transition-all focus-within:border-accent focus-within:ring-2 focus-within:ring-accent/20">
+1 -1
View File
@@ -64,7 +64,7 @@ export const StatusBar = () => {
</a>
{/* Right - Stats */}
<div className="flex items-center gap-3">
<div className="flex items-center gap-3" data-testid="graph-stats">
{graph && (
<>
<span>{nodeCount} nodes</span>
+59 -22
View File
@@ -145,6 +145,7 @@ interface AppState {
availableRepos: BackendRepo[];
setAvailableRepos: (repos: BackendRepo[]) => void;
switchRepo: (repoName: string) => Promise<void>;
setCurrentRepo: (repoName: string) => void;
// Worker API (shared across app)
runQuery: (cypher: string) => Promise<any[]>;
@@ -456,6 +457,10 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
// Backend client — direct HTTP calls (no Worker/Comlink)
const repoRef = useRef<string | undefined>(undefined);
const setCurrentRepo = useCallback((repoName: string) => {
repoRef.current = repoName;
}, []);
const runQuery = useCallback(async (cypher: string): Promise<any[]> => {
return backendRunQuery(cypher, repoRef.current);
}, []);
@@ -574,6 +579,13 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
try {
const effectiveProjectName = overrideProjectName || projectName || 'project';
// Sync repoRef so all agent backend calls target the correct repo.
// initializeAgent can be called from App.tsx (handleServerConnect) which
// never sets repoRef.current directly — without this, queries default to repo[0].
if (overrideProjectName) {
repoRef.current = overrideProjectName;
}
const repo = repoRef.current;
// Build backend interface for Graph RAG tools
@@ -605,7 +617,8 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
setIsAgentInitializing(false);
}
},
[projectName],
// eslint-disable-next-line react-hooks/exhaustive-deps
[], // repoRef is a stable ref — we sync it explicitly on entry; no state deps needed
);
const sendChatMessage = useCallback(
@@ -1037,6 +1050,9 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
setCodePanelOpen(false);
setCodeReferenceFocus(null);
let connectedRepo: BackendRepo | undefined;
let pNameStr = repoName || 'server-project';
try {
const result: ConnectResult = await connectToServer(
serverBaseUrl,
@@ -1068,39 +1084,28 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
},
undefined,
repoName,
{ awaitAnalysis: true }, // enable backend hold-queue for repos still being analyzed
);
// Build graph for visualization
const repoPath = result.repoInfo.repoPath ?? result.repoInfo.path;
// Prefer the registry name, then normalize Windows \ and Unix / paths
const pName =
repoName || result.repoInfo.name || repoPath?.split('/').pop() || 'server-project';
repoName ||
result.repoInfo.name ||
(repoPath || '').replace(/\\/g, '/').split('/').filter(Boolean).pop() ||
'server-project';
setProjectName(pName);
repoRef.current = pName;
connectedRepo = result.repoInfo;
pNameStr = pName;
const newGraph = createKnowledgeGraph();
for (const node of result.nodes) newGraph.addNode(node);
for (const rel of result.relationships) newGraph.addRelationship(rel);
setGraph(newGraph);
// No fileContents needed — grep/read tools use backend HTTP
// Initialize agent with backend queries, then start embeddings
try {
if (getActiveProviderConfig()) {
await initializeAgent(pName);
}
setViewMode('exploring');
startEmbeddingsWithFallback();
setProgress(null);
} catch (err) {
console.warn('Failed to initialize agent:', err);
setIsAgentReady(false);
agentRef.current = null;
setAgentError('Failed to initialize agent');
setViewMode('exploring');
setProgress(null);
}
} catch (err) {
} catch (err: unknown) {
console.error('Repo switch failed:', err);
setProgress({
phase: 'error',
@@ -1114,6 +1119,36 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
setViewMode('exploring');
setProgress(null);
}, ERROR_RESET_DELAY_MS);
return; // Abort the whole switchRepo process
}
if (pNameStr) {
// Persist the selected project in the URL so a refresh re-opens it
const urlObj = new URL(window.location.href);
urlObj.searchParams.set('project', pNameStr);
window.history.replaceState(null, '', urlObj.toString());
}
// Reset the agent and clear chat history so the AI starts fresh for the new repo
agentRef.current = null;
setIsAgentReady(false);
setChatMessages([]);
// Re-initialize agent with the new repo's graph context
try {
if (getActiveProviderConfig()) {
await initializeAgent(pNameStr);
}
setViewMode('exploring');
startEmbeddingsWithFallback();
setProgress(null);
} catch (err) {
console.warn('Failed to initialize agent:', err);
setIsAgentReady(false);
agentRef.current = null;
setAgentError('Failed to initialize agent');
setViewMode('exploring');
setProgress(null);
}
},
[
@@ -1133,6 +1168,7 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
setCodeReferences,
setCodePanelOpen,
setCodeReferenceFocus,
setChatMessages,
],
);
@@ -1219,6 +1255,7 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
availableRepos,
setAvailableRepos,
switchRepo,
setCurrentRepo,
runQuery,
isDatabaseReady,
// Embedding state and methods
+145
View File
@@ -0,0 +1,145 @@
import { useCallback, useEffect, useLayoutEffect, useRef, useState } from 'react';
const DEFAULT_BOTTOM_THRESHOLD = 100;
const USER_SCROLL_EPSILON = 5;
export interface UseAutoScrollResult {
scrollContainerRef: React.RefObject<HTMLDivElement>;
messagesContainerRef: React.RefObject<HTMLDivElement>;
isAtBottom: boolean;
scrollToBottom: (behavior?: ScrollBehavior) => void;
}
function isNearBottom(element: HTMLElement, threshold: number): boolean {
return element.scrollHeight - element.scrollTop - element.clientHeight <= threshold;
}
export function useAutoScroll<T>(
chatMessages: T[],
isChatLoading: boolean,
bottomThreshold = DEFAULT_BOTTOM_THRESHOLD,
): UseAutoScrollResult {
const scrollContainerRef = useRef<HTMLDivElement>(null);
const messagesContainerRef = useRef<HTMLDivElement>(null);
const [isAtBottom, setIsAtBottom] = useState(true);
const shouldStickToBottomRef = useRef(true);
const lastScrollTopRef = useRef(0);
const scrollFrameIdRef = useRef<number | null>(null);
const syncScrollState = useCallback(() => {
const element = scrollContainerRef.current;
if (!element) return;
const currentScrollTop = element.scrollTop;
const nearBottom = isNearBottom(element, bottomThreshold);
if (nearBottom) {
shouldStickToBottomRef.current = true;
} else if (currentScrollTop < lastScrollTopRef.current - USER_SCROLL_EPSILON) {
shouldStickToBottomRef.current = false;
}
lastScrollTopRef.current = currentScrollTop;
setIsAtBottom(nearBottom);
}, [bottomThreshold]);
const scrollToBottom = useCallback(
(behavior: ScrollBehavior = 'smooth') => {
const element = scrollContainerRef.current;
if (!element) return;
shouldStickToBottomRef.current = true;
if (behavior === 'auto') {
element.scrollTop = element.scrollHeight;
lastScrollTopRef.current = element.scrollTop;
setIsAtBottom(isNearBottom(element, bottomThreshold));
return;
}
element.scrollTo({
top: element.scrollHeight,
behavior,
});
},
[bottomThreshold],
);
useEffect(() => {
const element = scrollContainerRef.current;
if (!element) return;
lastScrollTopRef.current = element.scrollTop;
const handleScroll = () => {
if (scrollFrameIdRef.current !== null) {
cancelAnimationFrame(scrollFrameIdRef.current);
}
scrollFrameIdRef.current = requestAnimationFrame(() => {
scrollFrameIdRef.current = null;
syncScrollState();
});
};
element.addEventListener('scroll', handleScroll, { passive: true });
syncScrollState();
return () => {
element.removeEventListener('scroll', handleScroll);
if (scrollFrameIdRef.current !== null) {
cancelAnimationFrame(scrollFrameIdRef.current);
scrollFrameIdRef.current = null;
}
};
}, [syncScrollState]);
useEffect(() => {
const content = messagesContainerRef.current;
const scrollEl = scrollContainerRef.current;
if (!content || !scrollEl || typeof ResizeObserver === 'undefined') return;
let resizeFrameId: number | null = null;
const observer = new ResizeObserver(() => {
if (shouldStickToBottomRef.current) {
if (resizeFrameId !== null) {
cancelAnimationFrame(resizeFrameId);
}
resizeFrameId = requestAnimationFrame(() => {
resizeFrameId = null;
scrollToBottom('auto');
});
} else {
syncScrollState();
}
});
observer.observe(content);
return () => {
observer.disconnect();
if (resizeFrameId !== null) {
cancelAnimationFrame(resizeFrameId);
resizeFrameId = null;
}
};
}, [chatMessages.length, scrollToBottom, syncScrollState]);
useLayoutEffect(() => {
if (!shouldStickToBottomRef.current) return;
scrollToBottom('auto');
}, [chatMessages.length, isChatLoading, scrollToBottom]);
return {
scrollContainerRef,
messagesContainerRef,
isAtBottom,
scrollToBottom,
};
}
+1
View File
@@ -9,6 +9,7 @@
export {
AlertCircle,
AlertTriangle,
ArrowDown,
ArrowRight,
AtSign,
Brain,
+115 -19
View File
@@ -222,7 +222,7 @@ export function normalizeServerUrl(input: string): string {
// ── Internal Helpers ───────────────────────────────────────────────────────
const DEFAULT_TIMEOUT_MS = 10_000;
const DEFAULT_TIMEOUT_MS = 30_000;
const PROBE_TIMEOUT_MS = 2_000;
const fetchWithTimeout = async (
@@ -264,11 +264,13 @@ const fetchWithTimeout = async (
const assertOk = async (response: Response): Promise<void> => {
if (response.ok) return;
let message = `Backend returned ${response.status} ${response.statusText}`;
let message = response.statusText;
try {
const body = await response.json();
if (body && typeof body.error === 'string') {
message = body.error;
} else if (body && typeof body.message === 'string') {
message = body.message;
}
} catch {
// Response body was not JSON
@@ -302,16 +304,26 @@ export const fetchServerInfo = async (): Promise<ServerInfo> => {
};
/**
* Connect an SSE heartbeat to the backend. Fires `onDisconnect` when the
* server goes down (after one retry to avoid false positives from transient
* network hiccups). Returns a cleanup function.
* Connect an SSE heartbeat to the backend. Retries indefinitely with capped
* exponential backoff so transient hiccups don't reset the UI.
*
* - `onConnect` fires on every successful (re)connection.
* - `onReconnecting` fires on the first retry after a drop — use it to show
* a "reconnecting" banner while keeping the current view intact.
*
* Returns a cleanup function that tears down the EventSource and timers.
*/
export const connectHeartbeat = (onConnect: () => void, onDisconnect: () => void): (() => void) => {
export const connectHeartbeat = (
onConnect: () => void,
onReconnecting: () => void,
): (() => void) => {
let closed = false;
let retryTimer: ReturnType<typeof setTimeout> | null = null;
let es: EventSource | null = null;
let attempt = 0;
const MAX_RETRIES = 3;
/** Whether we've already fired onReconnecting for the current drop. */
let notifiedReconnecting = false;
const MAX_BACKOFF_MS = 15_000;
const connect = () => {
if (closed) return;
@@ -319,6 +331,7 @@ export const connectHeartbeat = (onConnect: () => void, onDisconnect: () => void
es.onopen = () => {
if (!closed) {
attempt = 0;
notifiedReconnecting = false;
onConnect();
}
};
@@ -326,13 +339,15 @@ export const connectHeartbeat = (onConnect: () => void, onDisconnect: () => void
es?.close();
es = null;
if (closed) return;
if (attempt < MAX_RETRIES) {
const delay = 1_000 * Math.pow(2, attempt);
attempt++;
retryTimer = setTimeout(connect, delay);
} else {
onDisconnect();
if (!notifiedReconnecting) {
notifiedReconnecting = true;
onReconnecting();
}
const delay = Math.min(1_000 * Math.pow(2, attempt), MAX_BACKOFF_MS);
attempt++;
retryTimer = setTimeout(connect, delay);
};
};
@@ -373,10 +388,22 @@ export const fetchRepos = async (): Promise<BackendRepo[]> => {
return response.json() as Promise<BackendRepo[]>;
};
/** Fetch repo metadata. */
export const fetchRepoInfo = async (repo?: string): Promise<BackendRepo> => {
/** Fetch repo metadata.
* Pass `awaitAnalysis: true` when connecting to a repo that may still be cloning/analyzing —
* this enables the backend's hold-queue and uses a 5-minute timeout to match.
* Normal calls (e.g. repo switching between already-indexed repos) use the default 10s timeout.
*
* Must stay in sync with HOLD_QUEUE_TIMEOUT_SECS in gitnexus/src/server/api.ts.
*/
const HOLD_QUEUE_TIMEOUT_MS = 300_000; // 5 minutes — matches backend HOLD_QUEUE_TIMEOUT_SECS
export const fetchRepoInfo = async (
repo?: string,
opts?: { awaitAnalysis?: boolean },
): Promise<BackendRepo> => {
const url = `${_backendUrl}/api/repo${repo ? `?${repoParam(repo)}` : ''}`;
const response = await fetchWithTimeout(url);
const timeout = opts?.awaitAnalysis ? HOLD_QUEUE_TIMEOUT_MS : undefined;
const response = await fetchWithTimeout(url, {}, timeout);
await assertOk(response);
const data = await response.json();
return { ...data, repoPath: data.repoPath ?? data.path };
@@ -391,13 +418,19 @@ export const fetchGraph = async (
onProgress?: (downloaded: number, total: number | null) => void;
},
): Promise<{ nodes: GraphNode[]; relationships: GraphRelationship[] }> => {
const params = [repoParam(repo), opts?.includeContent ? 'includeContent=true' : '']
const params = [repoParam(repo), opts?.includeContent ? 'includeContent=true' : '', 'stream=true']
.filter(Boolean)
.join('&');
const url = `${_backendUrl}/api/graph${params ? `?${params}` : ''}`;
const response = await fetchWithTimeout(url, { signal: opts?.signal }, 60_000);
// Large repos can take a while to serialize the graph — use an elevated timeout
const response = await fetchWithTimeout(url, { signal: opts?.signal }, 120_000);
await assertOk(response);
const contentType = response.headers.get('Content-Type') || '';
if (contentType.includes('application/x-ndjson')) {
return parseNdjsonGraphResponse(response, opts?.onProgress);
}
if (!opts?.onProgress || !response.body) {
return response.json() as Promise<{ nodes: GraphNode[]; relationships: GraphRelationship[] }>;
}
@@ -426,6 +459,66 @@ export const fetchGraph = async (
return JSON.parse(new TextDecoder().decode(combined));
};
const parseNdjsonGraphResponse = async (
response: Response,
onProgress?: (downloaded: number, total: number | null) => void,
): Promise<{ nodes: GraphNode[]; relationships: GraphRelationship[] }> => {
if (!response.body) {
throw new BackendError('No response body', response.status, 'server');
}
const contentLength = response.headers.get('Content-Length');
const total = contentLength ? parseInt(contentLength, 10) : null;
const reader = response.body.getReader();
const decoder = new TextDecoder();
const nodes: GraphNode[] = [];
const relationships: GraphRelationship[] = [];
let buffer = '';
let downloaded = 0;
const parseLine = (line: string) => {
const trimmed = line.trim();
if (!trimmed) return;
const record = JSON.parse(trimmed) as
| { type: 'node'; data: GraphNode }
| { type: 'relationship'; data: GraphRelationship }
| { type: 'error'; error: string };
if (record.type === 'node') {
nodes.push(record.data);
return;
}
if (record.type === 'relationship') {
relationships.push(record.data);
return;
}
if (record.type === 'error') {
throw new BackendError(record.error, response.status || 500, 'server');
}
};
while (true) {
const { done, value } = await reader.read();
if (done) break;
downloaded += value.length;
onProgress?.(downloaded, total);
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split('\n');
buffer = lines.pop() || '';
for (const line of lines) {
parseLine(line);
}
}
buffer += decoder.decode();
parseLine(buffer);
return { nodes, relationships };
};
/** Execute a Cypher query. Returns rows. */
export const runQuery = async (
cypher: string,
@@ -657,18 +750,21 @@ export interface ConnectResult {
/**
* Connect to a server: validate, fetch repo info, download graph.
* Content is NOT included (use readFile/grep for file access).
* Pass `awaitAnalysis: true` when the repo may still be cloning/analyzing —
* this enables the backend hold-queue and a 5-minute fetch timeout.
*/
export async function connectToServer(
url: string,
onProgress?: (phase: string, downloaded: number, total: number | null) => void,
signal?: AbortSignal,
repoName?: string,
opts?: { awaitAnalysis?: boolean },
): Promise<ConnectResult> {
const baseUrl = normalizeServerUrl(url);
setBackendUrl(baseUrl);
onProgress?.('validating', 0, null);
const repoInfo = await fetchRepoInfo(repoName);
const repoInfo = await fetchRepoInfo(repoName, { awaitAnalysis: opts?.awaitAnalysis });
onProgress?.('downloading', 0, null);
const { nodes, relationships } = await fetchGraph(repoName, {
+147
View File
@@ -0,0 +1,147 @@
import { describe, it, expect, vi, beforeEach, afterEach } from 'vitest';
import { connectHeartbeat } from '../../src/services/backend-client';
// Mock EventSource to simulate SSE behavior
class MockEventSource {
onopen: (() => void) | null = null;
onerror: (() => void) | null = null;
closed = false;
close() {
this.closed = true;
}
}
let lastEventSource: MockEventSource | null = null;
beforeEach(() => {
lastEventSource = null;
vi.stubGlobal(
'EventSource',
vi.fn().mockImplementation(() => {
lastEventSource = new MockEventSource();
return lastEventSource;
}),
);
vi.useFakeTimers();
});
afterEach(() => {
vi.useRealTimers();
vi.unstubAllGlobals();
});
describe('connectHeartbeat', () => {
it('calls onConnect when EventSource opens', () => {
const onConnect = vi.fn();
const onReconnecting = vi.fn();
connectHeartbeat(onConnect, onReconnecting);
lastEventSource!.onopen!();
expect(onConnect).toHaveBeenCalledOnce();
expect(onReconnecting).not.toHaveBeenCalled();
});
it('calls onReconnecting on first error, then retries', () => {
const onConnect = vi.fn();
const onReconnecting = vi.fn();
connectHeartbeat(onConnect, onReconnecting);
// Simulate connection drop
lastEventSource!.onerror!();
expect(onReconnecting).toHaveBeenCalledOnce();
expect(lastEventSource!.closed).toBe(true);
// Advance past first retry delay (1s)
vi.advanceTimersByTime(1_000);
// A new EventSource should have been created
expect(EventSource).toHaveBeenCalledTimes(2);
});
it('fires onReconnecting only once per disconnect', () => {
const onConnect = vi.fn();
const onReconnecting = vi.fn();
connectHeartbeat(onConnect, onReconnecting);
// First error
lastEventSource!.onerror!();
expect(onReconnecting).toHaveBeenCalledOnce();
// Second retry fires error again
vi.advanceTimersByTime(1_000);
lastEventSource!.onerror!();
expect(onReconnecting).toHaveBeenCalledOnce(); // still 1
// Third retry fires error
vi.advanceTimersByTime(2_000);
lastEventSource!.onerror!();
expect(onReconnecting).toHaveBeenCalledOnce(); // still 1
});
it('retries indefinitely instead of giving up after 3 attempts', () => {
const onConnect = vi.fn();
const onReconnecting = vi.fn();
connectHeartbeat(onConnect, onReconnecting);
// Simulate 10 consecutive failures — should never stop retrying
for (let i = 0; i < 10; i++) {
lastEventSource!.onerror!();
// Advance past the max backoff (15s) to ensure the next retry fires
vi.advanceTimersByTime(16_000);
}
// Should have created 11 EventSources (1 initial + 10 retries)
expect(EventSource).toHaveBeenCalledTimes(11);
});
it('resets reconnecting state when connection recovers', () => {
const onConnect = vi.fn();
const onReconnecting = vi.fn();
connectHeartbeat(onConnect, onReconnecting);
// Drop
lastEventSource!.onerror!();
expect(onReconnecting).toHaveBeenCalledOnce();
// Retry succeeds
vi.advanceTimersByTime(1_000);
lastEventSource!.onopen!();
expect(onConnect).toHaveBeenCalledOnce();
// Drop again — should fire onReconnecting again (reset after recovery)
lastEventSource!.onerror!();
expect(onReconnecting).toHaveBeenCalledTimes(2);
});
it('caps backoff at 15 seconds', () => {
const onConnect = vi.fn();
const onReconnecting = vi.fn();
connectHeartbeat(onConnect, onReconnecting);
// Fail many times to push backoff past the cap
for (let i = 0; i < 6; i++) {
lastEventSource!.onerror!();
// The delay for attempt i is min(1000 * 2^i, 15000)
// i=0: 1s, i=1: 2s, i=2: 4s, i=3: 8s, i=4: 15s (capped), i=5: 15s (capped)
vi.advanceTimersByTime(16_000);
}
// All retries should have fired — 7 EventSources total
expect(EventSource).toHaveBeenCalledTimes(7);
});
it('stops retrying when cleanup is called', () => {
const onConnect = vi.fn();
const onReconnecting = vi.fn();
const cleanup = connectHeartbeat(onConnect, onReconnecting);
lastEventSource!.onerror!();
cleanup();
// Advance time — no new EventSource should be created
vi.advanceTimersByTime(30_000);
expect(EventSource).toHaveBeenCalledTimes(1);
});
});
@@ -1,5 +1,5 @@
import { describe, expect, it } from 'vitest';
import { normalizeServerUrl } from '../../src/services/backend-client';
import { afterEach, describe, expect, it, vi } from 'vitest';
import { fetchGraph, normalizeServerUrl, setBackendUrl } from '../../src/services/backend-client';
describe('normalizeServerUrl', () => {
it('adds http:// to localhost', () => {
@@ -31,3 +31,137 @@ describe('normalizeServerUrl', () => {
expect(normalizeServerUrl('https://gitnexus.example.com')).toBe('https://gitnexus.example.com');
});
});
afterEach(() => {
vi.restoreAllMocks();
});
describe('fetchGraph', () => {
it('requests streamed graph responses from the backend', async () => {
setBackendUrl('http://localhost:4747');
const fetchMock = vi.fn().mockResolvedValue(
new Response('{"nodes":[],"relationships":[]}', {
status: 200,
headers: {
'Content-Type': 'application/json',
},
}),
);
vi.stubGlobal('fetch', fetchMock);
await fetchGraph('big-repo');
expect(fetchMock).toHaveBeenCalledWith(
expect.stringContaining('/api/graph?repo=big-repo&stream=true'),
expect.any(Object),
);
});
it('parses NDJSON graph streams incrementally', async () => {
setBackendUrl('http://localhost:4747');
const encoder = new TextEncoder();
const stream = new ReadableStream<Uint8Array>({
start(controller) {
controller.enqueue(
encoder.encode(
[
'{"type":"node","data":{"id":"File:src/app.ts","label":"File","properties":{"name":"app.ts","filePath":"src/app.ts"}}}\n',
'{"type":"relationship","data":{"id":"File:src/app.ts_CONTAINS_Function:src/app.ts:main","type":"CONTAINS","sourceId":"File:src/app.ts","targetId":"Function:src/app.ts:main"}}\n',
].join(''),
),
);
controller.close();
},
});
vi.stubGlobal(
'fetch',
vi.fn().mockResolvedValue(
new Response(stream, {
status: 200,
headers: {
'Content-Type': 'application/x-ndjson',
},
}),
),
);
const progress = vi.fn();
const result = await fetchGraph('big-repo', { onProgress: progress });
expect(result.nodes).toHaveLength(1);
expect(result.relationships).toHaveLength(1);
expect(result.nodes[0].id).toBe('File:src/app.ts');
expect(result.relationships[0].type).toBe('CONTAINS');
expect(progress).toHaveBeenCalled();
});
it('parses NDJSON graph lines split across chunks', async () => {
setBackendUrl('http://localhost:4747');
const encoder = new TextEncoder();
const stream = new ReadableStream<Uint8Array>({
start(controller) {
controller.enqueue(
encoder.encode(
'{"type":"node","data":{"id":"File:src/app.ts","label":"File","properties":{"name":"app.ts"',
),
);
controller.enqueue(
encoder.encode(
',"filePath":"src/app.ts"}}}\n{"type":"relationship","data":{"id":"File:src/app.ts_CONTAINS_Function:src/app.ts:main","type":"CONTAINS","sourceId":"File:src/app.ts","targetId":"Function:src/app.ts:main"}}\n',
),
);
controller.close();
},
});
vi.stubGlobal(
'fetch',
vi.fn().mockResolvedValue(
new Response(stream, {
status: 200,
headers: {
'Content-Type': 'application/x-ndjson',
},
}),
),
);
const result = await fetchGraph('big-repo');
expect(result.nodes).toHaveLength(1);
expect(result.relationships).toHaveLength(1);
expect(result.nodes[0].properties.filePath).toBe('src/app.ts');
});
it('throws backend errors emitted in the NDJSON stream', async () => {
setBackendUrl('http://localhost:4747');
const encoder = new TextEncoder();
const stream = new ReadableStream<Uint8Array>({
start(controller) {
controller.enqueue(encoder.encode('{"type":"error","error":"stream failed"}\n'));
controller.close();
},
});
vi.stubGlobal(
'fetch',
vi.fn().mockResolvedValue(
new Response(stream, {
status: 200,
headers: {
'Content-Type': 'application/x-ndjson',
},
}),
),
);
await expect(fetchGraph('big-repo')).rejects.toMatchObject({
message: 'stream failed',
});
});
});
@@ -0,0 +1,287 @@
import { act, fireEvent, render, screen } from '@testing-library/react';
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest';
import { useAutoScroll } from '../../src/hooks/useAutoScroll';
interface HarnessProps {
messages: unknown[];
isChatLoading: boolean;
}
function AutoScrollHarness({ messages, isChatLoading }: HarnessProps) {
const { scrollContainerRef, messagesContainerRef, isAtBottom, scrollToBottom } = useAutoScroll(
messages,
isChatLoading,
);
return (
<>
<div data-testid="is-at-bottom">{String(isAtBottom)}</div>
<div data-testid="container" ref={scrollContainerRef}>
{messages.length > 0 ? (
<div data-testid="messages-container" ref={messagesContainerRef}>
{messages.map((message, index) => (
<div key={index}>{String(message)}</div>
))}
</div>
) : null}
</div>
<button type="button" onClick={() => scrollToBottom()}>
Scroll to bottom
</button>
</>
);
}
function setScrollMetrics(
element: HTMLDivElement,
metrics: { scrollTop?: number; scrollHeight?: number; clientHeight?: number },
) {
if (metrics.scrollTop !== undefined) {
Object.defineProperty(element, 'scrollTop', {
configurable: true,
writable: true,
value: metrics.scrollTop,
});
}
if (metrics.scrollHeight !== undefined) {
Object.defineProperty(element, 'scrollHeight', {
configurable: true,
value: metrics.scrollHeight,
});
}
if (metrics.clientHeight !== undefined) {
Object.defineProperty(element, 'clientHeight', {
configurable: true,
value: metrics.clientHeight,
});
}
}
async function flushAnimationFrame() {
await act(async () => {
vi.runAllTimers();
});
}
async function scrollContainer(element: HTMLDivElement, scrollTop: number) {
setScrollMetrics(element, { scrollTop });
fireEvent.scroll(element);
await flushAnimationFrame();
}
const resizeObserverInstances: ResizeObserverMock[] = [];
class ResizeObserverMock {
callback: ResizeObserverCallback;
observedElements: Element[] = [];
observe = vi.fn((element: Element) => {
this.observedElements.push(element);
});
unobserve = vi.fn();
disconnect = vi.fn();
constructor(callback: ResizeObserverCallback) {
this.callback = callback;
resizeObserverInstances.push(this);
}
}
async function triggerResize(instance: ResizeObserverMock) {
await act(async () => {
instance.callback([], instance as unknown as ResizeObserver);
});
await flushAnimationFrame();
}
describe('useAutoScroll', () => {
beforeEach(() => {
vi.useFakeTimers();
resizeObserverInstances.length = 0;
vi.stubGlobal(
'requestAnimationFrame',
vi.fn((callback: FrameRequestCallback) => {
return window.setTimeout(() => callback(performance.now()), 0);
}),
);
vi.stubGlobal(
'cancelAnimationFrame',
vi.fn((frameId: number) => {
clearTimeout(frameId);
}),
);
Object.defineProperty(HTMLElement.prototype, 'scrollTo', {
configurable: true,
value: function (options: ScrollToOptions) {
if (options.top !== undefined) {
Object.defineProperty(this, 'scrollTop', {
configurable: true,
writable: true,
value: options.top,
});
}
},
});
vi.stubGlobal('ResizeObserver', ResizeObserverMock);
});
afterEach(() => {
vi.useRealTimers();
vi.unstubAllGlobals();
});
it('starts with isAtBottom true and auto-scrolls the very first message', () => {
const { rerender } = render(<AutoScrollHarness messages={[]} isChatLoading={false} />);
expect(screen.getByTestId('is-at-bottom')).toHaveTextContent('true');
const container = screen.getByTestId('container') as HTMLDivElement;
setScrollMetrics(container, { scrollTop: 0, scrollHeight: 500, clientHeight: 200 });
rerender(<AutoScrollHarness messages={[{ id: 1 }]} isChatLoading={false} />);
expect(container.scrollTop).toBe(500);
expect(screen.getByTestId('is-at-bottom')).toHaveTextContent('true');
});
it('follows streaming updates while the view stays pinned to the bottom', () => {
const { rerender } = render(<AutoScrollHarness messages={[{ id: 1 }]} isChatLoading={false} />);
const container = screen.getByTestId('container') as HTMLDivElement;
setScrollMetrics(container, { scrollTop: 700, scrollHeight: 1000, clientHeight: 200 });
rerender(<AutoScrollHarness messages={[{ id: 1 }]} isChatLoading={true} />);
expect(container.scrollTop).toBe(1000);
expect(screen.getByTestId('is-at-bottom')).toHaveTextContent('true');
});
it('stops auto-scroll after the user scrolls up', async () => {
const { rerender } = render(<AutoScrollHarness messages={[{ id: 1 }]} isChatLoading={false} />);
const container = screen.getByTestId('container') as HTMLDivElement;
setScrollMetrics(container, { scrollTop: 700, scrollHeight: 1000, clientHeight: 200 });
await scrollContainer(container, 700);
await scrollContainer(container, 250);
expect(screen.getByTestId('is-at-bottom')).toHaveTextContent('false');
setScrollMetrics(container, { scrollTop: 250, scrollHeight: 1400, clientHeight: 200 });
rerender(<AutoScrollHarness messages={[{ id: 1 }, { id: 2 }]} isChatLoading={true} />);
expect(container.scrollTop).toBe(250);
});
it('re-enables auto-scroll once the user returns near the bottom', async () => {
const { rerender } = render(<AutoScrollHarness messages={[{ id: 1 }]} isChatLoading={false} />);
const container = screen.getByTestId('container') as HTMLDivElement;
setScrollMetrics(container, { scrollTop: 700, scrollHeight: 1000, clientHeight: 200 });
await scrollContainer(container, 700);
await scrollContainer(container, 250);
setScrollMetrics(container, { scrollTop: 1120, scrollHeight: 1400, clientHeight: 200 });
await scrollContainer(container, 1120);
expect(screen.getByTestId('is-at-bottom')).toHaveTextContent('true');
setScrollMetrics(container, { scrollTop: 1120, scrollHeight: 1800, clientHeight: 200 });
rerender(<AutoScrollHarness messages={[{ id: 1 }, { id: 2 }]} isChatLoading={true} />);
expect(container.scrollTop).toBe(1800);
});
it('scrollToBottom re-engages auto-scroll and scrolls to the container bottom', async () => {
const { rerender } = render(<AutoScrollHarness messages={[{ id: 1 }]} isChatLoading={false} />);
const container = screen.getByTestId('container') as HTMLDivElement;
const scrollTo = vi.spyOn(container, 'scrollTo');
setScrollMetrics(container, { scrollTop: 700, scrollHeight: 1000, clientHeight: 200 });
await scrollContainer(container, 700);
await scrollContainer(container, 250);
fireEvent.click(screen.getByRole('button', { name: 'Scroll to bottom' }));
expect(scrollTo).toHaveBeenCalledWith({ top: 1000, behavior: 'smooth' });
expect(screen.getByTestId('is-at-bottom')).toHaveTextContent('false');
setScrollMetrics(container, { scrollTop: 250, scrollHeight: 1600, clientHeight: 200 });
rerender(<AutoScrollHarness messages={[{ id: 1 }, { id: 2 }]} isChatLoading={true} />);
expect(container.scrollTop).toBe(1600);
expect(screen.getByTestId('is-at-bottom')).toHaveTextContent('true');
});
it('re-pins to the latest bottom when inner content grows asynchronously', async () => {
render(<AutoScrollHarness messages={[{ id: 1 }]} isChatLoading={false} />);
const container = screen.getByTestId('container') as HTMLDivElement;
setScrollMetrics(container, { scrollTop: 700, scrollHeight: 1000, clientHeight: 200 });
await scrollContainer(container, 700);
setScrollMetrics(container, { scrollTop: 1000, scrollHeight: 1450, clientHeight: 200 });
await triggerResize(resizeObserverInstances[0]);
expect(container.scrollTop).toBe(1450);
expect(screen.getByTestId('is-at-bottom')).toHaveTextContent('true');
});
it('does not auto-scroll on async growth after user intentionally scrolls away', async () => {
render(<AutoScrollHarness messages={[{ id: 1 }]} isChatLoading={false} />);
const container = screen.getByTestId('container') as HTMLDivElement;
setScrollMetrics(container, { scrollTop: 700, scrollHeight: 1000, clientHeight: 200 });
await scrollContainer(container, 700);
await scrollContainer(container, 250);
setScrollMetrics(container, { scrollTop: 250, scrollHeight: 1400, clientHeight: 200 });
await triggerResize(resizeObserverInstances[0]);
expect(container.scrollTop).toBe(250);
expect(screen.getByTestId('is-at-bottom')).toHaveTextContent('false');
});
it('cancels the pending ResizeObserver rAF when the component unmounts', () => {
const cancelRAF = vi.mocked(cancelAnimationFrame);
const { unmount } = render(<AutoScrollHarness messages={[{ id: 1 }]} isChatLoading={false} />);
const container = screen.getByTestId('container') as HTMLDivElement;
setScrollMetrics(container, { scrollTop: 950, scrollHeight: 1000, clientHeight: 200 });
const callsBefore = cancelRAF.mock.calls.length;
act(() => {
resizeObserverInstances[0].callback(
[],
resizeObserverInstances[0] as unknown as ResizeObserver,
);
});
unmount();
expect(cancelRAF.mock.calls.length).toBeGreaterThan(callsBefore);
expect(() => vi.runAllTimers()).not.toThrow();
});
it('attaches the observer when the messages wrapper first appears and disconnects on unmount', () => {
const { rerender, unmount } = render(<AutoScrollHarness messages={[]} isChatLoading={false} />);
expect(screen.queryByTestId('messages-container')).toBeNull();
expect(resizeObserverInstances).toHaveLength(0);
rerender(<AutoScrollHarness messages={[{ id: 1 }]} isChatLoading={false} />);
const messagesContainer = screen.getByTestId('messages-container');
const resizeObserver = resizeObserverInstances[0];
expect(resizeObserverInstances).toHaveLength(1);
expect(resizeObserver.observe).toHaveBeenCalledWith(messagesContainer);
unmount();
expect(resizeObserver.disconnect).toHaveBeenCalledTimes(1);
});
});
+75
View File
@@ -2,6 +2,81 @@
All notable changes to GitNexus will be documented in this file.
## [1.6.1] - 2026-04-13
### Added
- **Service group extractor expansion** — manifest extractor and broader extractor coverage (2/4 of #606 split) (#796)
- **Dart call patterns** for `await`, cascade, lambda, and widget-tree contexts (#801)
### Fixed
- **Stack overflow and memory exhaustion** on large repository analysis (#814)
- **`tree-sitter-dart` install crash** — switched from git URL to npm tarball (#811)
- **Generic TypeScript awaited function calls** missing from the call graph (#804)
- **Runtime dependency on `file:../gitnexus-shared`** removed from the published package (#803)
- **Ruby `singleton_class` context** preserved during sequential parsing (#774)
### Changed
- **DAG-based ingestion pipeline architecture** — pipeline phases now declare typed dependencies and run via a topologically sorted DAG; container-node logic extracted to `LanguageProvider`. Includes hardened lifecycle (try/finally cleanup, error wrapping, cycle reporting), tightened `ParseOutput.exportedTypeMap` immutability, and corrected phase dependencies (#809)
## [1.6.0] - 2026-04-12
### Added
- **SemanticModel architecture refactor (SM-8 through SM-19)** — extracted registries into `model/` module with ISP-compliant interfaces: TypeRegistry, MethodRegistry, FieldRegistry, RegistrationTable, ResolutionContext (#786)
- HeritageMap built from accumulated `ExtractedHeritage[]` for MRO-aware resolution (#739)
- `lookupMethodByOwnerWithMRO` using HeritageMap for cross-class method dispatch (#740)
- MRO fast path before D2 fuzzy widening in call resolution (#741)
- BindingAccumulator for cross-file return type propagation (#743, #763)
- Restructured `resolveUncached` replacing `lookupFuzzy` data source for all tiers (#764)
- Deleted `lookupFuzzy`, `lookupFuzzyCallable`, `globalIndex`, `callableIndex` — replaced with structured lookups (#769)
- Deleted `resolveCallTarget` god-method — replaced with thin dispatcher delegating to `resolveMemberCall` (#744), `resolveStaticCall` (#754), `resolveFreeCall` (#756) (#770)
- **Service group infrastructure** — service boundary detection, contract extractors, sync pipeline, CLI/MCP tools, monorepo fixture; bridge.lbug storage and contract matching expansion (#795)
- **C# interface-to-interface heritage** capture (#789)
- **Vue SFC support** with destructured call result tracking (#604)
- **Java method reference** resolution — `obj::method` as call sites (#622)
- **C/C++ MethodExtractor** config with pure virtual detection (#617)
- **MethodExtractor configs** for Python, PHP, Swift, Dart, Rust, Ruby (#624)
- **METHOD_IMPLEMENTS edges** with overload disambiguation and MethodExtractor unification (#642)
- **Same-arity overload disambiguation** via type-hash suffix (#658)
- **`GITNEXUS_HOME` env var** to customize global directory (#746)
- **Verbose analyze output** prints skipped large file paths (#745)
- **Class name lookup index** for O(1) qualified lookups (#707, #716)
- **`lookupMethodByOwner` index** for O(1) cross-class chain resolution (#665)
- **Fuzzy lookup counters** for performance visibility (#708)
### Fixed
- **Stack overflow on large PHP files** — iterative AST traversal (#783)
- **Large repository graph loading** failure (#732)
- **Windows multi-repo switching** — false 404 errors and stale repo context (#633)
- **`detect_changes` diff mapping** — map diff hunks to symbol line ranges (#779)
- **HTTP client vs Express route detection** and Spring interface attribution (#780)
- **VECTOR extension** not loaded during DB init for semantic search (#782)
- **tree-sitter-swift** postinstall patch for macOS ARM64 (#788)
- **tree-sitter-c** peer dependency conflict pinned (#723)
- **Constructor indexing** in methodByOwner (#694, #753)
- **Named binding processor** — `lookupExact` replaced with `lookupExactAll` (#755)
- **`.gitnexusignore` negation patterns** now respected (#654)
- **MCP setup** prefers global gitnexus binary over npx (#653)
- **CORS rejection** returns clean error instead of 500 (#646)
- **Array.push stack overflow** — replaced spread with loop (#650)
- **MCP stdout silencing** prevents embedder/pool-adapter conflicts (#645)
- **Web heartbeat** — graceful reconnection replaces aggressive disconnect (#643)
- **Web repo scoping** — backend calls scoped to active repo (#644)
- **OpenCode config path** and FTS extension load order (#781)
- **OnboardingGuide** dev-mode serve command corrected (#725)
- **Security issues** and critical bugs from code review (#709)
### Changed
- Replaced class-type fuzzy lookups with structured indices in type-env (#733, #734, #736)
- Extracted `CLASS_LIKE_TYPES` constant (#693)
## [1.5.3] - 2026-04-01
### Added
- **TypeScript/JavaScript MethodExtractor** config (#588)
### Fixed
- **Wiki Azure OpenAI** compat and HTML viewer script injection (#618)
## [1.5.2] - 2026-04-01
### Fixed
+50
View File
@@ -234,6 +234,56 @@ Installed automatically by both `gitnexus analyze` (per-repo) and `gitnexus setu
- Node.js >= 18
- Git repository (uses git for commit tracking)
## Troubleshooting
### `Cannot destructure property 'package' of 'node.target' as it is null`
This crash was caused by a dependency URL format that is incompatible with
certain npm/arborist versions ([npm/cli#8126](https://github.com/npm/cli/issues/8126)).
It is fixed in **gitnexus v1.6.2+**. Upgrade to the latest version:
```bash
npx gitnexus@latest analyze # always uses the newest release
# — or —
npm install -g gitnexus@latest # upgrade a global install
```
If you still hit npm install issues after upgrading, these generic workarounds
may help:
```bash
npm install -g npm@latest # update npm itself
npm cache clean --force # clear a possibly corrupt cache
```
### Installation fails with native module errors
Some optional language grammars (Dart, Kotlin, Swift) require native compilation. If they fail, GitNexus still works — those languages will be skipped.
If `npm install -g gitnexus` fails on native modules:
```bash
# Ensure build tools are available (Linux/macOS)
# Ubuntu/Debian: sudo apt install python3 make g++
# macOS: xcode-select --install
# Retry installation
npm install -g gitnexus
```
### Analysis runs out of memory
For very large repositories:
```bash
# Increase Node.js heap size
NODE_OPTIONS="--max-old-space-size=16384" npx gitnexus analyze
# Exclude large directories
echo "vendor/" >> .gitnexusignore
echo "dist/" >> .gitnexusignore
```
## Privacy
- All processing happens locally on your machine
+41 -4
View File
@@ -1,17 +1,19 @@
{
"name": "gitnexus",
"version": "1.5.3",
"version": "1.6.1",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "gitnexus",
"version": "1.5.3",
"version": "1.6.1",
"hasInstallScript": true,
"license": "PolyForm-Noncommercial-1.0.0",
"dependencies": {
"@huggingface/transformers": "^3.0.0",
"@ladybugdb/core": "^0.15.2",
"@modelcontextprotocol/sdk": "^1.0.0",
"@scarf/scarf": "^1.4.0",
"cli-progress": "^3.12.0",
"commander": "^12.0.0",
"cors": "^2.8.5",
@@ -60,8 +62,9 @@
"node": ">=20.0.0"
},
"optionalDependencies": {
"tree-sitter-dart": "github:UserNobody14/tree-sitter-dart#80e23c07b64494f7e21090bb3450223ef0b192f4",
"tree-sitter-dart": "git+https://github.com/UserNobody14/tree-sitter-dart.git#80e23c07b64494f7e21090bb3450223ef0b192f4",
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-proto": "file:./vendor/tree-sitter-proto",
"tree-sitter-swift": "^0.6.0"
}
},
@@ -1878,6 +1881,13 @@
"dev": true,
"license": "MIT"
},
"node_modules/@scarf/scarf": {
"version": "1.4.0",
"resolved": "https://registry.npmjs.org/@scarf/scarf/-/scarf-1.4.0.tgz",
"integrity": "sha512-xxeapPiUXdZAE3che6f3xogoJPeZgig6omHEy1rIY5WVsB3H2BHNnZH+gHG6x91SCWyQCzWGsuL2Hh3ClO5/qQ==",
"hasInstallScript": true,
"license": "Apache-2.0"
},
"node_modules/@standard-schema/spec": {
"version": "1.1.0",
"resolved": "https://registry.npmjs.org/@standard-schema/spec/-/spec-1.1.0.tgz",
@@ -5122,7 +5132,7 @@
},
"node_modules/tree-sitter-dart": {
"version": "1.0.0",
"resolved": "git+ssh://git@github.com/UserNobody14/tree-sitter-dart.git#80e23c07b64494f7e21090bb3450223ef0b192f4",
"resolved": "git+https://github.com/UserNobody14/tree-sitter-dart.git#80e23c07b64494f7e21090bb3450223ef0b192f4",
"integrity": "sha512-Bs/1wAOIJ2akPEXlE/XVpuES19Oo3NqoSJRJ/0N2r38qAd9nTXdqmaGHQ44/JXnA6QHcbgD2YzCCc4wUc98cyQ==",
"hasInstallScript": true,
"license": "ISC",
@@ -5286,6 +5296,10 @@
"node": "^18 || ^20 || >= 21"
}
},
"node_modules/tree-sitter-proto": {
"resolved": "vendor/tree-sitter-proto",
"link": true
},
"node_modules/tree-sitter-python": {
"version": "0.23.4",
"resolved": "https://registry.npmjs.org/tree-sitter-python/-/tree-sitter-python-0.23.4.tgz",
@@ -5867,6 +5881,29 @@
"peerDependencies": {
"zod": "^3.25.28 || ^4"
}
},
"vendor/tree-sitter-proto": {
"version": "0.4.1",
"hasInstallScript": true,
"license": "MIT",
"optional": true,
"dependencies": {
"node-addon-api": "^8.0.0",
"node-gyp-build": "^4.8.0"
},
"peerDependencies": {
"tree-sitter": ">=0.21.0"
}
},
"vendor/tree-sitter-proto/node_modules/node-addon-api": {
"version": "8.7.0",
"resolved": "https://registry.npmjs.org/node-addon-api/-/node-addon-api-8.7.0.tgz",
"integrity": "sha512-9MdFxmkKaOYVTV+XVRG8ArDwwQ77XIgIPyKASB1k3JPq3M8fGQQQE3YpMOrKm6g//Ktx8ivZr8xo1Qmtqub+GA==",
"license": "MIT",
"optional": true,
"engines": {
"node": "^18 || ^20 || >= 21"
}
}
}
}
+7 -4
View File
@@ -1,6 +1,6 @@
{
"name": "gitnexus",
"version": "1.5.3",
"version": "1.6.1",
"description": "Graph-powered code intelligence for AI agents. Index any codebase, query via MCP or CLI.",
"author": "Abhigyan Patwari",
"license": "PolyForm-Noncommercial-1.0.0",
@@ -46,6 +46,7 @@
"test:integration": "vitest run test/integration",
"test:watch": "vitest",
"test:coverage": "vitest run --coverage",
"postinstall": "node scripts/patch-tree-sitter-swift.cjs",
"prepare": "node scripts/build.js",
"prepack": "node scripts/build.js"
},
@@ -53,11 +54,11 @@
"@huggingface/transformers": "^3.0.0",
"@ladybugdb/core": "^0.15.2",
"@modelcontextprotocol/sdk": "^1.0.0",
"@scarf/scarf": "^1.4.0",
"cli-progress": "^3.12.0",
"commander": "^12.0.0",
"cors": "^2.8.5",
"express": "^4.19.2",
"gitnexus-shared": "file:../gitnexus-shared",
"glob": "^11.0.0",
"graphology": "^0.25.4",
"graphology-indices": "^0.17.0",
@@ -83,8 +84,9 @@
"uuid": "^13.0.0"
},
"optionalDependencies": {
"tree-sitter-dart": "github:UserNobody14/tree-sitter-dart#80e23c07b64494f7e21090bb3450223ef0b192f4",
"tree-sitter-dart": "git+https://github.com/UserNobody14/tree-sitter-dart.git#80e23c07b64494f7e21090bb3450223ef0b192f4",
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-proto": "file:./vendor/tree-sitter-proto",
"tree-sitter-swift": "^0.6.0"
},
"devDependencies": {
@@ -103,7 +105,8 @@
"overrides": {
"@huggingface/transformers": {
"onnxruntime-node": "$onnxruntime-node"
}
},
"tree-sitter-c": "0.23.2"
},
"engines": {
"node": ">=20.0.0"
@@ -0,0 +1,78 @@
#!/usr/bin/env node
/**
* WORKAROUND: tree-sitter-swift@0.6.0 binding.gyp build failure
*
* Background:
* tree-sitter-swift@0.6.0's binding.gyp contains an "actions" array that
* invokes `tree-sitter generate` to regenerate parser.c from grammar.js.
* This is intended for grammar developers, but the published npm package
* already ships pre-generated parser files (parser.c, scanner.c), so the
* actions are unnecessary for consumers. Since consumers don't have
* tree-sitter-cli installed, the actions always fail during `npm install`.
*
* Why we can't just upgrade:
* tree-sitter-swift@0.7.1 fixes this (removes postinstall, ships prebuilds),
* but it requires tree-sitter@^0.22.1. The upstream project pins tree-sitter
* to ^0.21.0 and all other grammar packages depend on that version.
* Upgrading tree-sitter would be a separate breaking change.
*
* How this workaround works:
* 1. tree-sitter-swift's own postinstall fails (npm warns but continues)
* 2. This script runs as gitnexus's postinstall
* 3. It removes the "actions" array from binding.gyp
* 4. It rebuilds the native binding with the cleaned binding.gyp
*
* TODO: Remove this script when tree-sitter is upgraded to ^0.22.x,
* which allows using tree-sitter-swift@0.7.1+ directly.
*/
const fs = require('fs');
const path = require('path');
const { execSync } = require('child_process');
const swiftDir = path.join(__dirname, '..', 'node_modules', 'tree-sitter-swift');
const bindingPath = path.join(swiftDir, 'binding.gyp');
try {
if (!fs.existsSync(bindingPath)) {
process.exit(0);
}
const content = fs.readFileSync(bindingPath, 'utf8');
let needsRebuild = false;
if (content.includes('"actions"')) {
// Strip Python-style comments (#) and trailing commas before JSON parsing
const cleaned = content
.replace(/#[^\n]*/g, '') // Remove # comments
.replace(/,(\s*[\]}])/g, '$1'); // Remove trailing commas before ] or }
const gyp = JSON.parse(cleaned);
if (gyp.targets && gyp.targets[0] && gyp.targets[0].actions) {
delete gyp.targets[0].actions;
fs.writeFileSync(bindingPath, JSON.stringify(gyp, null, 2) + '\n');
console.log('[tree-sitter-swift] Patched binding.gyp (removed actions array)');
needsRebuild = true;
}
}
// Check if native binding exists
const bindingNode = path.join(swiftDir, 'build', 'Release', 'tree_sitter_swift_binding.node');
if (!fs.existsSync(bindingNode)) {
needsRebuild = true;
}
if (needsRebuild) {
console.log('[tree-sitter-swift] Rebuilding native binding...');
execSync('npx node-gyp rebuild', {
cwd: swiftDir,
stdio: 'pipe',
timeout: 120000,
});
console.log('[tree-sitter-swift] Native binding built successfully');
}
} catch (err) {
console.warn('[tree-sitter-swift] Could not build native binding:', err.message);
console.warn(
'[tree-sitter-swift] You may need to manually run: cd node_modules/tree-sitter-swift && npx node-gyp rebuild',
);
}
+10 -2
View File
@@ -26,6 +26,7 @@ interface RepoStats {
export interface AIContextOptions {
skipAgentsMd?: boolean;
noStats?: boolean;
}
const GITNEXUS_START_MARKER = '<!-- gitnexus:start -->';
@@ -64,6 +65,7 @@ function generateGitNexusContent(
stats: RepoStats,
generatedSkills?: GeneratedSkillInfo[],
groupNames?: string[],
noStats?: boolean,
): string {
const generatedRows =
generatedSkills && generatedSkills.length > 0
@@ -87,7 +89,7 @@ function generateGitNexusContent(
return `${GITNEXUS_START_MARKER}
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **${projectName}** (${stats.nodes || 0} symbols, ${stats.edges || 0} relationships, ${stats.processes || 0} execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **${projectName}**${noStats ? '' : ` (${stats.nodes || 0} symbols, ${stats.edges || 0} relationships, ${stats.processes || 0} execution flows)`}. Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run \`npx gitnexus analyze\` in terminal first.
@@ -332,7 +334,13 @@ export async function generateAIContextFiles(
options?: AIContextOptions,
): Promise<{ files: string[] }> {
const groupNames = await findGroupsContainingRegistryName(projectName);
const content = generateGitNexusContent(projectName, stats, generatedSkills, groupNames);
const content = generateGitNexusContent(
projectName,
stats,
generatedSkills,
groupNames,
options?.noStats,
);
const createdFiles: string[] = [];
if (!options?.skipAgentsMd) {
+59 -4
View File
@@ -20,8 +20,11 @@ import fs from 'fs/promises';
const HEAP_MB = 8192;
const HEAP_FLAG = `--max-old-space-size=${HEAP_MB}`;
/** Increase default stack size (KB) to prevent stack overflow on deep class hierarchies. */
const STACK_KB = 4096;
const STACK_FLAG = `--stack-size=${STACK_KB}`;
/** Re-exec the process with an 8GB heap if we're currently below that. */
/** Re-exec the process with an 8GB heap and larger stack if we're currently below that. */
function ensureHeap(): boolean {
const nodeOpts = process.env.NODE_OPTIONS || '';
if (nodeOpts.includes('--max-old-space-size')) return false;
@@ -29,8 +32,13 @@ function ensureHeap(): boolean {
const v8Heap = v8.getHeapStatistics().heap_size_limit;
if (v8Heap >= HEAP_MB * 1024 * 1024 * 0.9) return false;
// --stack-size is a V8 flag not allowed in NODE_OPTIONS on Node 24+,
// so pass it only as a direct CLI argument, not via the environment.
const cliFlags = [HEAP_FLAG];
if (!nodeOpts.includes('--stack-size')) cliFlags.push(STACK_FLAG);
try {
execFileSync(process.execPath, [HEAP_FLAG, ...process.argv.slice(1)], {
execFileSync(process.execPath, [...cliFlags, ...process.argv.slice(1)], {
stdio: 'inherit',
env: { ...process.env, NODE_OPTIONS: `${nodeOpts} ${HEAP_FLAG}`.trim() },
});
@@ -47,6 +55,8 @@ export interface AnalyzeOptions {
verbose?: boolean;
/** Skip AGENTS.md and CLAUDE.md gitnexus block updates. */
skipAgentsMd?: boolean;
/** Omit volatile symbol/relationship counts from AGENTS.md and CLAUDE.md. */
noStats?: boolean;
/** Index the folder even when no .git directory is present. */
skipGit?: boolean;
}
@@ -177,6 +187,7 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
embeddings: options?.embeddings,
skipGit: options?.skipGit,
skipAgentsMd: options?.skipAgentsMd,
noStats: options?.noStats,
},
{
onProgress: (_phase, percent, message) => {
@@ -240,7 +251,7 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
processes: s.processes,
},
skillResult.skills,
{ skipAgentsMd: options?.skipAgentsMd },
{ skipAgentsMd: options?.skipAgentsMd, noStats: options?.noStats },
);
}
} catch {
@@ -282,7 +293,51 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
console.warn = origWarn;
console.error = origError;
bar.stop();
console.error(`\n Analysis failed: ${err.message}\n`);
const msg = err.message || String(err);
console.error(`\n Analysis failed: ${msg}\n`);
// Provide helpful guidance for known failure modes
if (
msg.includes('Maximum call stack size exceeded') ||
msg.includes('call stack') ||
msg.includes('Map maximum size') ||
msg.includes('Invalid array length') ||
msg.includes('Invalid string length') ||
msg.includes('allocation failed') ||
msg.includes('heap out of memory') ||
msg.includes('JavaScript heap')
) {
console.error(' This error typically occurs on very large repositories.');
console.error(' Suggestions:');
console.error(' 1. Add large vendored/generated directories to .gitnexusignore');
console.error(' 2. Increase Node.js heap: NODE_OPTIONS="--max-old-space-size=16384"');
console.error(' 3. Increase stack size: NODE_OPTIONS="--stack-size=4096"');
console.error('');
} else if (msg.includes('ERESOLVE') || msg.includes('Could not resolve dependency')) {
// Note: the original arborist "Cannot destructure property 'package' of
// 'node.target'" crash happens inside npm *before* gitnexus code runs,
// so it can't be caught here. This branch handles dependency-resolution
// errors that surface at runtime (e.g. dynamic require failures).
console.error(' This looks like an npm dependency resolution issue.');
console.error(' Suggestions:');
console.error(' 1. Clear the npm cache: npm cache clean --force');
console.error(' 2. Update npm: npm install -g npm@latest');
console.error(' 3. Reinstall gitnexus: npm install -g gitnexus@latest');
console.error(' 4. Or try npx directly: npx gitnexus@latest analyze');
console.error('');
} else if (
msg.includes('MODULE_NOT_FOUND') ||
msg.includes('Cannot find module') ||
msg.includes('ERR_MODULE_NOT_FOUND')
) {
console.error(' A required module could not be loaded. The installation may be corrupt.');
console.error(' Suggestions:');
console.error(' 1. Reinstall: npm install -g gitnexus@latest');
console.error(' 2. Clear cache: npm cache clean --force && npx gitnexus@latest analyze');
console.error('');
}
process.exitCode = 1;
return;
}
+1
View File
@@ -26,6 +26,7 @@ program
.option('--embeddings', 'Enable embedding generation for semantic search (off by default)')
.option('--skills', 'Generate repo-specific skill files from detected communities')
.option('--skip-agents-md', 'Skip updating the gitnexus section in AGENTS.md and CLAUDE.md')
.option('--no-stats', 'Omit volatile file/symbol counts from AGENTS.md and CLAUDE.md')
.option('--skip-git', 'Index a folder without requiring a .git directory')
.option('-v, --verbose', 'Enable verbose ingestion warnings (default: false)')
.addHelpText(
+4 -1
View File
@@ -14,7 +14,10 @@ process.on('unhandledRejection', (reason: any) => {
export const serveCommand = async (options?: { port?: string; host?: string }) => {
const port = Number(options?.port ?? 4747);
const host = options?.host ?? '127.0.0.1';
// Default to 'localhost' so the OS decides whether to bind to 127.0.0.1 or
// ::1 based on system configuration, avoiding spurious CORS errors when the
// hosted frontend at gitnexus.vercel.app connects to localhost.
const host = options?.host ?? 'localhost';
try {
await createServer(port, host);
+36 -3
View File
@@ -9,7 +9,7 @@
import fs from 'fs/promises';
import path from 'path';
import os from 'os';
import { execFile } from 'child_process';
import { execFile, execFileSync } from 'child_process';
import { promisify } from 'util';
import { fileURLToPath } from 'url';
import { glob } from 'glob';
@@ -25,11 +25,44 @@ interface SetupResult {
errors: string[];
}
/**
* Resolve the absolute path to the `gitnexus` binary if it's installed
* globally (or via npm -g / yarn global). Returns null when not found.
*/
function resolveGitnexusBin(): string | null {
try {
const cmd = process.platform === 'win32' ? 'where' : 'which';
const resolved = execFileSync(cmd, ['gitnexus'], {
encoding: 'utf-8',
timeout: 5000,
stdio: ['ignore', 'pipe', 'ignore'],
})
.split('\n')[0]
.trim();
return resolved || null;
} catch {
return null;
}
}
/**
* The MCP server entry for all editors.
* On Windows, npx must be invoked via cmd /c since it's a .cmd script.
*
* Prefers the globally-installed `gitnexus` binary (starts in ~1 s) over
* `npx -y gitnexus@latest` (cold-cache install of native deps can take
* >60 s, exceeding Claude Code's 30 s MCP connection timeout).
*
* Falls back to npx when the binary isn't on PATH — e.g. first-time
* users who ran `npx gitnexus analyze` but haven't done `npm i -g`.
*/
function getMcpEntry() {
const bin = resolveGitnexusBin();
if (bin) {
return { command: bin, args: ['mcp'] };
}
// Fallback: npx (works without a global install, but slow cold-start)
if (process.platform === 'win32') {
return {
command: 'cmd',
@@ -232,7 +265,7 @@ async function setupOpenCode(result: SetupResult): Promise<void> {
return;
}
const configPath = path.join(opencodeDir, 'config.json');
const configPath = path.join(opencodeDir, 'opencode.json');
try {
const existing = await readJsonFile(configPath);
const config = existing || {};
+8 -3
View File
@@ -398,11 +398,16 @@ export const createIgnoreFilter = async (repoPath: string, options?: IgnoreOptio
// defense-in-depth — do not remove `dot: false` assuming this covers it.
if (DEFAULT_IGNORE_LIST.has(p.name)) return true;
// Check against .gitignore / .gitnexusignore patterns.
// Test both bare path and path with trailing slash to handle
// bare-name patterns (e.g. `local`) and dir-only patterns (e.g. `local/`).
// Since childrenIgnored is only called for directories, always test with
// a trailing slash. This ensures directory-only negation patterns (e.g.
// `!iOS/`) are applied correctly — without the slash, `ig.ignores('iOS')`
// treats the path as a file and misses the negation.
// Bare-name patterns (e.g. `local`) still match `local/` per gitignore spec:
// the `ignore` package normalizes `dir` and `dir/` to match directories.
// See: https://github.com/kaelzhang/node-ignore#2-filenames-and-dirnames
if (ig) {
const rel = p.relative();
if (rel && (ig.ignores(rel) || ig.ignores(rel + '/'))) return true;
if (rel && ig.ignores(rel + '/')) return true;
}
return false;
},
@@ -100,8 +100,8 @@ const batchInsertEmbeddings = async (
) => Promise<void>,
updates: Array<{ id: string; embedding: number[] }>,
): Promise<void> => {
// INSERT into separate embedding table - much more memory efficient!
const cypher = `CREATE (e:CodeEmbedding {nodeId: $nodeId, embedding: $embedding})`;
// MERGE instead of CREATE — idempotent, handles concurrent analyzes and partial prior runs
const cypher = `MERGE (e:CodeEmbedding {nodeId: $nodeId}) SET e.embedding = $embedding`;
const paramsList = updates.map((u) => ({ nodeId: u.id, embedding: u.embedding }));
await executeWithReusedStatement(cypher, paramsList);
};
+139
View File
@@ -0,0 +1,139 @@
# Group Analysis Pipeline
Flow chart of the cross-repo contract extraction + matching pipeline.
This covers what runs **inside this PR** (extractors + manifest) and
the downstream handoff to the bridge storage (PR #795) and
cross-impact query (PR #606).
## High-level overview
```mermaid
flowchart TD
A[group.yaml] --> B[GroupConfig parser]
B --> C{For each repo<br/>in group}
C --> D[Per-repo LadybugDB<br/>indexed by main pipeline]
D --> E1[TopicExtractor]
D --> E2[HttpRouteExtractor]
D --> E3[GrpcExtractor]
E1 --> F[ExtractedContract array<br/>per repo]
E2 --> F
E3 --> F
B --> M[ManifestExtractor]
M --> G[Manifest contracts<br/>+ cross-links]
F --> H[Contract matching<br/>exact + wildcard]
G --> H
H --> I[(bridge.lbug<br/>#795)]
I --> J[runGroupImpact<br/>#606]
J --> K[CrossRepoImpact]
```
## Per-repo extractor pipeline
Each extractor under `src/core/group/extractors/` follows the same
two-strategy shape:
```mermaid
flowchart TD
R[RepoHandle + CypherExecutor<br/>for this repo] --> S{Graph-assisted<br/>Strategy A<br/>available?}
S -->|yes| A1[Cypher query against<br/>per-repo LadybugDB]
A1 --> A2{non-empty<br/>result?}
A2 -->|yes| OUT[ExtractedContract array]
A2 -->|no| B1
S -->|no| B1[Source-scan Strategy B]
B1 --> B2[glob repo source files]
B2 --> B3{ext in registry?}
B3 -->|yes| B4[Per-language plugin<br/>scan parsed tree]
B3 -->|no| SKIP[skip file]
B4 --> OUT
SKIP --> B2
```
**Strategy A** (graph-assisted) uses Cypher over edges already produced
by the main ingestion pipeline:
- HTTP: `HANDLES_ROUTE` / `FETCHES` edges from `(File)-[]->(Route)`
- topic: none (pipeline doesn't yet produce topic nodes — Strategy B only)
- gRPC: none (Strategy B + proto map only)
**Strategy B** (source-scan) is 100% tree-sitter based after this PR.
Each `*-patterns/<lang>.ts` plugin owns its grammar + S-expression
queries; the top-level orchestrator imports neither.
## Plugin architecture
```mermaid
flowchart LR
O[Orchestrator<br/>topic|http|grpc-extractor.ts] --> REG[REGISTRY<br/>*-patterns/index.ts]
REG --> P1[java.ts<br/>tree-sitter-java]
REG --> P2[go.ts<br/>tree-sitter-go]
REG --> P3[python.ts<br/>tree-sitter-python]
REG --> P4[node.ts<br/>JS + TS + TSX]
REG --> P5[php.ts<br/>tree-sitter-php<br/>HTTP only]
REG --> P6[proto.ts<br/>tree-sitter-proto<br/>gRPC only, optional]
P1 --> SCAN[tree-sitter-scanner.ts<br/>compilePatterns + runCompiledPatterns]
P2 --> SCAN
P3 --> SCAN
P4 --> SCAN
P5 --> SCAN
P6 --> SCAN
SCAN --> DET[Detection objects<br/>TopicMeta / HttpDetection / GrpcDetection]
DET --> O
O --> CT[ExtractedContract array]
```
The orchestrator never imports a grammar. Adding a new language /
framework = drop one file in `*-patterns/`, register it in
`index.ts`. No orchestrator edits required.
## Manifest extraction
```mermaid
flowchart TD
Y[group.yaml links] --> ME[ManifestExtractor]
ME --> LOOP{for each link}
LOOP --> RES[resolveSymbol<br/>label-scoped Cypher]
RES --> OK{found?}
OK -->|yes| REF[real symbol uid + ref]
OK -->|no| SYN[synthetic uid<br/>manifest::repo::cid]
REF --> EMIT[emit provider + consumer<br/>Contract objects<br/>+ CrossLink]
SYN --> EMIT
EMIT --> BRIDGE[(bridge.lbug<br/>#795)]
```
Label-scoped queries in `resolveSymbol` keep accidental cross-matches
out:
- `topic` → `(n:Function|Method|Class|Interface)`
- `grpc` method → `(n:Function|Method)`, service → `(n:Class|Interface)`
- `lib` → `(n:Package|Module)`
## Cross-impact query (PR #606)
```mermaid
flowchart TD
U[User changes symbol S<br/>in repo R] --> LI[Local impact engine<br/>per-repo uid expansion]
LI --> IDS[Affected uid set]
IDS --> BR[Bridge query<br/>MATCH Contract WHERE uid IN ids]
BR --> CL[CrossLink traversal]
CL --> OTHER[Matching contract in<br/>other repo]
OTHER --> FE[Fan-out impact<br/>to consuming repo]
FE --> OUT[CrossRepoImpact<br/>per affected repo]
```
The bridge stores every extracted contract keyed by `symbolUid`.
Manifest-sourced contracts use the synthetic uid form so both sides
of the `(local impact) ↔ (bridge query)` join derive the same uid
without coordinating through any shared state.
+588
View File
@@ -0,0 +1,588 @@
import fsp from 'node:fs/promises';
import path from 'node:path';
import { createHash } from 'node:crypto';
import lbug from '@ladybugdb/core';
import type { LbugValue } from '@ladybugdb/core';
import type { BridgeHandle, BridgeMeta, StoredContract, CrossLink, RepoSnapshot } from './types.js';
import { BRIDGE_SCHEMA_QUERIES, BRIDGE_SCHEMA_VERSION } from './bridge-schema.js';
import { dedupeContracts, dedupeCrossLinks } from './normalization.js';
export function contractNodeId(
repo: string,
contractId: string,
role: string,
filePath: string,
): string {
return createHash('sha256').update(`${repo}\0${contractId}\0${role}\0${filePath}`).digest('hex');
}
/* ------------------------------------------------------------------ */
/* ContractLookupIndex — in-memory lookup for findContractNode */
/* ------------------------------------------------------------------ */
/**
* In-memory index of contract node IDs keyed three ways, mirroring the
* three-tier fallback lookup in {@link findContractNode}. Built once per
* `writeBridge` call after all contracts are successfully inserted, then
* consulted for every cross-link — which eliminates the former N+1 query
* pattern (up to `6 × cross-links` DB round-trips) and turns cross-link
* resolution into constant-time per link.
*
* Keys are deliberately flat strings (not tuples) so `Map<string, ...>`
* works; the separator `\0` can't occur in any legal repo path / file
* path / symbol identifier, which makes the encoding injection-safe.
*/
export interface ContractLookupIndex {
/** tier 1: `repo + role + symbolUid` → contract node id */
byUid: Map<string, string>;
/** tier 2: `repo + role + filePath + symbolName` → contract node id */
byRef: Map<string, string>;
/** tier 3: `repo + role + filePath` → list of contract node ids in that file */
byFile: Map<string, string[]>;
}
export function createContractLookupIndex(): ContractLookupIndex {
return {
byUid: new Map(),
byRef: new Map(),
byFile: new Map(),
};
}
function uidKey(repo: string, role: string, symbolUid: string): string {
return `${repo}\0${role}\0${symbolUid}`;
}
function refKey(repo: string, role: string, filePath: string, symbolName: string): string {
return `${repo}\0${role}\0${filePath}\0${symbolName}`;
}
function fileKey(repo: string, role: string, filePath: string): string {
return `${repo}\0${role}\0${filePath}`;
}
/**
* Add a successfully-inserted contract to the lookup index. Must be called
* AFTER the DB insert succeeds (not before) so failed inserts don't poison
* the index and cause cross-links to point at non-existent rows.
*/
export function indexContract(
index: ContractLookupIndex,
contract: StoredContract,
nodeId: string,
): void {
if (contract.symbolUid) {
index.byUid.set(uidKey(contract.repo, contract.role, contract.symbolUid), nodeId);
}
index.byRef.set(
refKey(contract.repo, contract.role, contract.symbolRef.filePath, contract.symbolRef.name),
nodeId,
);
const fk = fileKey(contract.repo, contract.role, contract.symbolRef.filePath);
const existing = index.byFile.get(fk);
if (existing) {
existing.push(nodeId);
} else {
index.byFile.set(fk, [nodeId]);
}
}
/**
* Resolve a cross-link endpoint (consumer or provider reference) to an
* already-inserted contract node id. Returns `null` if no match — the
* caller is expected to count that as a dropped link in `WriteBridgeReport`.
*
* The resolution order matches the pre-cache DB-query behavior:
* 1. exact `symbolUid` match in the same `(repo, role)` scope
* 2. exact `(filePath, symbolName)` match
* 3. if exactly one contract lives in the file → that one (fallback for
* legacy graph-assisted extractors that couldn't resolve a symbol name)
*
* This is a pure function — no I/O, no DB — so it's trivial to unit-test
* in isolation (which was the reviewer's main clean-code concern on the
* original 35-line inner closure in `writeBridge`).
*/
export function findContractNode(
index: ContractLookupIndex,
repo: string,
role: 'consumer' | 'provider',
symbolUid: string,
filePath: string,
symbolName: string,
): string | null {
if (symbolUid) {
const uidHit = index.byUid.get(uidKey(repo, role, symbolUid));
if (uidHit !== undefined) return uidHit;
}
const refHit = index.byRef.get(refKey(repo, role, filePath, symbolName));
if (refHit !== undefined) return refHit;
const fileCandidates = index.byFile.get(fileKey(repo, role, filePath));
if (fileCandidates && fileCandidates.length === 1) return fileCandidates[0];
return null;
}
export async function openBridgeDb(dbPath: string): Promise<BridgeHandle> {
const parentDir = path.dirname(dbPath);
await fsp.mkdir(parentDir, { recursive: true });
const db = new lbug.Database(dbPath, 0, false, false); // writable
const conn = new lbug.Connection(db);
return { _db: db, _conn: conn, groupDir: parentDir } as BridgeHandle;
}
/**
* LadybugDB returns an error whose message contains this substring when a
* CREATE NODE TABLE or CREATE REL TABLE statement hits an already-existing
* table. LadybugDB DDL doesn't support IF NOT EXISTS, and its JS driver
* doesn't expose typed error codes, so we match on the message substring —
* the same pattern used by `core/lbug/lbug-adapter.ts`. If a future
* LadybugDB release changes the wording, update this constant.
*/
const LBUG_ALREADY_EXISTS_MSG = 'already exists';
export async function ensureBridgeSchema(handle: BridgeHandle): Promise<void> {
const conn = handle._conn as lbug.Connection;
for (const q of BRIDGE_SCHEMA_QUERIES) {
try {
await conn.query(q);
} catch (err: unknown) {
const msg = err instanceof Error ? err.message : String(err);
if (!msg.includes(LBUG_ALREADY_EXISTS_MSG)) throw err;
}
}
}
export async function queryBridge<T>(
handle: BridgeHandle,
cypher: string,
params?: Record<string, LbugValue>,
): Promise<T[]> {
const conn = handle._conn as lbug.Connection;
if (params && Object.keys(params).length > 0) {
const stmt = await conn.prepare(cypher);
if (!stmt.isSuccess()) {
const errMsg = await stmt.getErrorMessage();
throw new Error(`Bridge query prepare failed: ${errMsg}`);
}
const queryResult = await conn.execute(stmt, params);
const result = unwrapQueryResult(queryResult);
return (await result.getAll()) as T[];
}
const queryResult = await conn.query(cypher);
const result = unwrapQueryResult(queryResult);
return (await result.getAll()) as T[];
}
/**
* LadybugDB's `conn.query` / `conn.execute` can return either a single
* `QueryResult` (for a single statement) or an array of them (when a
* multi-statement script is dispatched). We always pass a single statement,
* so the array form is a wrapper we unwrap here — but an empty top-level
* array would cause `.getAll()` on `undefined` and crash with a confusing
* stack. Throwing an explicit error makes a driver-contract regression
* visible immediately instead of masking it.
*/
function unwrapQueryResult(queryResult: lbug.QueryResult | lbug.QueryResult[]): lbug.QueryResult {
if (Array.isArray(queryResult)) {
if (queryResult.length === 0) {
throw new Error('Bridge query returned an empty QueryResult array');
}
return queryResult[0];
}
return queryResult;
}
export async function closeBridgeDb(handle: BridgeHandle): Promise<void> {
try {
await (handle._conn as lbug.Connection).close();
} catch {
/* ignore */
}
try {
await (handle._db as lbug.Database).close();
} catch {
/* ignore */
}
}
/* ------------------------------------------------------------------ */
/* retryRename — handles transient EBUSY/EPERM/EACCES on Windows */
/* ------------------------------------------------------------------ */
const RETRY_CODES = new Set(['EBUSY', 'EPERM', 'EACCES']);
export async function retryRename(src: string, dst: string, attempts = 3): Promise<void> {
for (let i = 1; i <= attempts; i++) {
try {
await fsp.rename(src, dst);
return;
} catch (err: unknown) {
const code = (err as NodeJS.ErrnoException).code;
if (!code || !RETRY_CODES.has(code) || i === attempts) throw err;
await new Promise((r) => setTimeout(r, 100 * Math.pow(2, i - 1)));
}
}
}
/* ------------------------------------------------------------------ */
/* writeBridgeMeta / readBridgeMeta */
/* ------------------------------------------------------------------ */
export async function writeBridgeMeta(groupDir: string, meta: BridgeMeta): Promise<void> {
const target = path.join(groupDir, 'meta.json');
const tmp = `${target}.tmp.${Date.now()}`;
await fsp.writeFile(tmp, JSON.stringify(meta, null, 2), 'utf-8');
// Use retryRename for consistency with writeBridge's atomic swap — on
// Windows a concurrent reader can cause EBUSY/EPERM even on a tiny
// meta.json, and we don't want meta write to be less robust than the
// bridge.lbug swap it accompanies.
await retryRename(tmp, target);
}
export async function readBridgeMeta(groupDir: string): Promise<BridgeMeta> {
try {
const content = await fsp.readFile(path.join(groupDir, 'meta.json'), 'utf-8');
return JSON.parse(content) as BridgeMeta;
} catch {
return { version: 0, generatedAt: '', missingRepos: [] };
}
}
/* ------------------------------------------------------------------ */
/* writeBridge — atomic write-to-temp-then-rename */
/* ------------------------------------------------------------------ */
export interface WriteBridgeInput {
contracts: StoredContract[];
crossLinks: CrossLink[];
repoSnapshots: Record<string, RepoSnapshot>;
missingRepos: string[];
}
/**
* Non-fatal issues encountered during writeBridge. Callers can log these to
* surface partial-success state without aborting the whole sync.
* `sampleErrors` is capped at MAX_SAMPLE_ERRORS per category to bound memory.
*/
export interface WriteBridgeReport {
contractsInserted: number;
contractsFailed: number;
snapshotsInserted: number;
snapshotsFailed: number;
linksInserted: number;
linksFailed: number;
/** Cross-links skipped because their from/to contract nodes weren't found. */
linksDroppedMissingNode: number;
sampleErrors: Array<{
kind: 'contract' | 'snapshot' | 'link';
id: string;
message: string;
}>;
}
const MAX_SAMPLE_ERRORS = 10;
function errMessage(err: unknown): string {
if (err instanceof Error) return err.message;
try {
return String(err);
} catch {
return 'unknown error';
}
}
export async function writeBridge(
groupDir: string,
input: WriteBridgeInput,
): Promise<WriteBridgeReport> {
await fsp.mkdir(groupDir, { recursive: true });
const contracts = dedupeContracts(input.contracts);
const crossLinks = dedupeCrossLinks(input.crossLinks);
const finalPath = path.join(groupDir, 'bridge.lbug');
const tmpPath = path.join(groupDir, 'bridge.lbug.tmp');
const bakPath = path.join(groupDir, 'bridge.lbug.bak');
const report: WriteBridgeReport = {
contractsInserted: 0,
contractsFailed: 0,
snapshotsInserted: 0,
snapshotsFailed: 0,
linksInserted: 0,
linksFailed: 0,
linksDroppedMissingNode: 0,
sampleErrors: [],
};
const recordError = (kind: 'contract' | 'snapshot' | 'link', id: string, err: unknown) => {
if (report.sampleErrors.length < MAX_SAMPLE_ERRORS) {
report.sampleErrors.push({ kind, id, message: errMessage(err) });
}
};
// Clean up any leftover tmp
try {
await fsp.rm(tmpPath, { recursive: true, force: true });
} catch {
/* ignore */
}
// 1. Create temp DB, insert all data.
//
// Everything after `openBridgeDb` must run inside a try/finally so that
// if ANY step before the explicit `closeBridgeDb` throws — schema
// creation, a contract insert loop that rethrows, a snapshot write, the
// cross-link loop, or anything else — the handle is still released. A
// leaked handle holds the native LadybugDB file lock on tmpPath, which
// (a) leaks a FD and (b) prevents the next writeBridge call from
// reusing the same tmp slot.
const handle = await openBridgeDb(tmpPath);
let handleClosed = false;
try {
await ensureBridgeSchema(handle);
// Build the lookup index incrementally as contracts are inserted, so
// failed inserts are never in the index (and therefore never resolved
// by the cross-link loop below). This replaces a previous N+1 query
// pattern where each link made up to 6 DB round-trips to find its
// endpoints — see ContractLookupIndex.
const lookupIndex = createContractLookupIndex();
// Insert contracts — tolerate individual failures (e.g., a corrupt meta
// that can't be serialized). The whole sync must not fail because one
// contract is broken.
for (const c of contracts) {
const id = contractNodeId(c.repo, c.contractId, c.role, c.symbolRef.filePath);
try {
await queryBridge(
handle,
`CREATE (n:Contract {
id: $id,
contractId: $contractId,
type: $type,
role: $role,
repo: $repo,
service: $service,
symbolUid: $symbolUid,
filePath: $filePath,
symbolName: $symbolName,
confidence: $confidence,
meta: $meta
})`,
{
id,
contractId: c.contractId,
type: c.type,
role: c.role,
repo: c.repo,
service: c.service ?? '',
symbolUid: c.symbolUid,
filePath: c.symbolRef.filePath,
symbolName: c.symbolName,
confidence: c.confidence,
meta: JSON.stringify(c.meta),
},
);
report.contractsInserted++;
// Only index on successful insert — the cross-link loop must never
// resolve to a row that isn't actually in the DB.
indexContract(lookupIndex, c, id);
} catch (err) {
report.contractsFailed++;
recordError('contract', id, err);
}
}
// Insert repo snapshots
for (const [repoId, snap] of Object.entries(input.repoSnapshots)) {
try {
await queryBridge(
handle,
`CREATE (s:RepoSnapshot {
id: $id,
indexedAt: $indexedAt,
lastCommit: $lastCommit
})`,
{
id: repoId,
indexedAt: snap.indexedAt,
lastCommit: snap.lastCommit,
},
);
report.snapshotsInserted++;
} catch (err) {
report.snapshotsFailed++;
recordError('snapshot', repoId, err);
}
}
// Insert cross-links (tolerating missing nodes).
//
// `findContractNode` consults the in-memory lookup index built above,
// not the DB — that's an O(1) pure-function lookup per endpoint instead
// of the previous 2-3 DB queries. For M cross-links, the previous code
// issued up to 6M round-trips; this version issues zero.
//
// `link.contractId` may differ between the consumer and provider sides
// (e.g. wildcard consumer `grpc::Service/*` → method-level provider
// `grpc::Service/Method`) — that's why we resolve each endpoint
// independently via its own `(repo, role, symbolUid, filePath, symbolName)`
// tuple rather than matching on contractId.
for (const link of crossLinks) {
const linkId = `${link.from.repo}::${link.contractId}->${link.to.repo}::${link.contractId}`;
try {
const fromId = findContractNode(
lookupIndex,
link.from.repo,
'consumer',
link.from.symbolUid,
link.from.symbolRef.filePath,
link.from.symbolRef.name,
);
const toId = findContractNode(
lookupIndex,
link.to.repo,
'provider',
link.to.symbolUid,
link.to.symbolRef.filePath,
link.to.symbolRef.name,
);
if (!fromId || !toId) {
report.linksDroppedMissingNode++;
continue;
}
await queryBridge(
handle,
`
MATCH (a:Contract), (b:Contract)
WHERE a.id = $fromId AND b.id = $toId
CREATE (a)-[:ContractLink {
matchType: $matchType,
confidence: $confidence,
contractId: $contractId,
fromRepo: $fromRepo,
toRepo: $toRepo
}]->(b)
`,
{
fromId,
toId,
matchType: link.matchType,
confidence: link.confidence,
contractId: link.contractId,
fromRepo: link.from.repo,
toRepo: link.to.repo,
},
);
report.linksInserted++;
} catch (err) {
report.linksFailed++;
recordError('link', linkId, err);
}
}
// 2. Close temp DB (happy path). The finally block also calls
// closeBridgeDb if we threw above; `handleClosed` prevents a
// double-close on the native handle.
await closeBridgeDb(handle);
handleClosed = true;
} finally {
if (!handleClosed) {
await closeBridgeDb(handle).catch(() => {
/* ignore: cleanup path, best effort */
});
}
}
// 3. Atomic swap: old→.bak, tmp→final, rm .bak
try {
await fsp.access(finalPath);
await retryRename(finalPath, bakPath);
} catch {
/* no existing db */
}
await retryRename(tmpPath, finalPath);
try {
await fsp.rm(bakPath, { recursive: true, force: true });
} catch {
/* ignore */
}
// 4. Write meta.json
await writeBridgeMeta(groupDir, {
version: BRIDGE_SCHEMA_VERSION,
generatedAt: new Date().toISOString(),
missingRepos: input.missingRepos,
});
return report;
}
/* ------------------------------------------------------------------ */
/* openBridgeDbReadOnly */
/* ------------------------------------------------------------------ */
export async function openBridgeDbReadOnly(groupDir: string): Promise<BridgeHandle | null> {
const dbPath = path.join(groupDir, 'bridge.lbug');
try {
await fsp.access(dbPath);
} catch {
// Check for .bak recovery. Use `retryRename` (not `fsp.rename`) for the
// exact same reason the rest of this file does: the scenario that
// triggers bak recovery is an interrupted writer, which on Windows may
// still be holding an open handle on `.bak` for a few milliseconds when
// a reader races in. EBUSY/EPERM retries recover that case silently.
const bakPath = path.join(groupDir, 'bridge.lbug.bak');
try {
await fsp.access(bakPath);
await retryRename(bakPath, dbPath);
} catch {
return null;
}
}
// Version gate: check meta.json version compatibility
const meta = await readBridgeMeta(groupDir);
if (meta.version > 0 && meta.version !== BRIDGE_SCHEMA_VERSION) {
return null; // incompatible schema version — fallback to JSON or re-sync
}
// Open the native handle. If Connection construction throws AFTER
// Database was successfully allocated, we'd leak the native Database
// object. Wrap each step separately and tear down the partial handle.
let db: lbug.Database | undefined;
let conn: lbug.Connection | undefined;
try {
db = new lbug.Database(dbPath, 0, false, true); // readOnly
conn = new lbug.Connection(db);
return { _db: db, _conn: conn, groupDir } as BridgeHandle;
} catch {
if (conn) {
try {
await conn.close();
} catch {
/* ignore */
}
}
if (db) {
try {
await db.close();
} catch {
/* ignore */
}
}
return null;
}
}
/* ------------------------------------------------------------------ */
/* bridgeExists */
/* ------------------------------------------------------------------ */
export async function bridgeExists(groupDir: string): Promise<boolean> {
const handle = await openBridgeDbReadOnly(groupDir);
if (!handle) return false;
await closeBridgeDb(handle);
return true;
}
+60
View File
@@ -0,0 +1,60 @@
/**
* Bridge LadybugDB schema for cross-repo Contract Registry.
* Separate from per-repo schema in lbug/schema.ts.
*/
/**
* Version of the bridge.lbug schema below. `openBridgeDbReadOnly` compares
* this against `meta.json`'s version field and returns `null` on mismatch,
* which trips the caller into either the JSON fallback path or a fresh
* `group sync` that rebuilds `bridge.lbug` from scratch.
*
* Migration contract for contributors bumping this constant:
* 1. Bump the number (e.g. `1` → `2`).
* 2. Update the DDL below to match the new schema.
* 3. DO NOT attempt an online migration in this file — the version gate
* is intentionally a "discard and re-sync" strategy for V1. An old
* bridge.lbug whose version doesn't match is treated as opaque and
* rebuilt by the next `group sync`.
* 4. If online migration becomes necessary (e.g. when groups accumulate
* large amounts of embedding data), add a migration path as a
* separate `bridge-migrations.ts` module rather than bloating this
* file — keep schema and migration concerns separate.
*/
export const BRIDGE_SCHEMA_VERSION = 1;
export const CONTRACT_SCHEMA = `
CREATE NODE TABLE Contract (
id STRING,
contractId STRING,
type STRING,
role STRING,
repo STRING,
service STRING DEFAULT '',
symbolUid STRING DEFAULT '',
filePath STRING DEFAULT '',
symbolName STRING DEFAULT '',
confidence DOUBLE DEFAULT 0.0,
meta STRING DEFAULT '{}',
PRIMARY KEY (id)
)`;
export const REPO_SNAPSHOT_SCHEMA = `
CREATE NODE TABLE RepoSnapshot (
id STRING,
indexedAt STRING DEFAULT '',
lastCommit STRING DEFAULT '',
PRIMARY KEY (id)
)`;
export const CONTRACT_LINK_SCHEMA = `
CREATE REL TABLE ContractLink (
FROM Contract TO Contract,
matchType STRING,
confidence DOUBLE,
contractId STRING,
fromRepo STRING,
toRepo STRING
)`;
export const BRIDGE_SCHEMA_QUERIES = [CONTRACT_SCHEMA, REPO_SNAPSHOT_SCHEMA, CONTRACT_LINK_SCHEMA];
@@ -0,0 +1,23 @@
import * as fs from 'node:fs';
import * as path from 'node:path';
/**
* Safely read a file inside a repo, rejecting any path that escapes
* `repoPath` via `..` traversal or absolute segments. Returns `null` if
* the path is outside the repo or the file can't be read.
*
* Used by every source-scan extractor under this directory. Kept as a
* single shared implementation so the path-traversal guard (security-
* sensitive) lives in exactly one place.
*/
export function readSafe(repoPath: string, rel: string): string | null {
const abs = path.resolve(repoPath, rel);
const base = path.resolve(repoPath);
const relToBase = path.relative(base, abs);
if (relToBase.startsWith('..') || path.isAbsolute(relToBase)) return null;
try {
return fs.readFileSync(abs, 'utf-8');
} catch {
return null;
}
}
@@ -1,20 +1,38 @@
import * as fs from 'node:fs';
import * as path from 'node:path';
import { glob } from 'glob';
import Parser from 'tree-sitter';
import type { ContractExtractor, CypherExecutor } from '../contract-extractor.js';
import type { ExtractedContract, RepoHandle } from '../types.js';
import { readSafe } from './fs-utils.js';
import {
GRPC_SCAN_GLOB,
getPluginForFile,
hasProtoPlugin,
type GrpcDetection,
} from './grpc-patterns/index.js';
function readSafe(repoPath: string, rel: string): string | null {
const abs = path.resolve(repoPath, rel);
const base = path.resolve(repoPath);
const relToBase = path.relative(base, abs);
if (relToBase.startsWith('..') || path.isAbsolute(relToBase)) return null;
try {
return fs.readFileSync(abs, 'utf-8');
} catch {
return null;
}
}
/**
* Language-agnostic orchestrator for gRPC (provider + consumer) contract
* extraction.
*
* Two parts:
*
* 1. **`.proto` parsing** — tree-sitter when `tree-sitter-proto` is
* installed (optionalDependency vendored in `vendor/tree-sitter-proto/`),
* via the `.proto` entry in `grpc-patterns/` and `hasProtoPlugin`.
* When the grammar isn't available (platform incompatibility, native
* build failure) the orchestrator falls back to the in-process
* string-sanitizing parser defined below (`stripProtoCommentsAndStrings`
* + `extractServiceBlocks`). The fallback preserves offsets so any
* downstream regex scans run against a sanitized copy without
* affecting line numbers of the original.
*
* 2. **Source-scan providers / consumers** — delegated to per-language
* plugins in `./grpc-patterns/`. The orchestrator imports NO
* tree-sitter grammars or query strings — each plugin owns its own.
*/
// ─── .proto fallback parser (used only when tree-sitter-proto is absent) ───
function contractId(pkg: string, service: string, method: string): string {
const prefix = pkg ? `${pkg}.${service}` : service;
@@ -25,20 +43,110 @@ function serviceOnlyContractId(serviceName: string): string {
return `grpc::${serviceName}/*`;
}
/**
* Replace all .proto comments and string literals with spaces, preserving the
* original length and character offsets of the input. This lets downstream
* regex / brace-depth parsers run on a "sanitized" copy without having to
* understand proto syntax, while any RegExp.exec/index-based lookups that
* were already positional against `content` continue to work against the
* original string.
*
* Supported comment forms: `// line comment`, `/* block comment * /`.
* Supported strings: double-quoted ("…") and single-quoted ('…') with `\`
* escape handling. Raw/unterminated strings are not supported — we stop
* on a line break for line-style comments and on EOF for unterminated
* strings/blocks, which matches how most real proto files parse.
*/
function stripProtoCommentsAndStrings(content: string): string {
const out = new Array<string>(content.length);
let i = 0;
while (i < content.length) {
const ch = content[i];
const next = content[i + 1];
// Line comment: // ... \n
if (ch === '/' && next === '/') {
out[i] = ' ';
out[i + 1] = ' ';
i += 2;
while (i < content.length && content[i] !== '\n') {
out[i] = content[i] === '\r' ? '\r' : ' ';
i++;
}
continue;
}
// Block comment: /* ... */
if (ch === '/' && next === '*') {
out[i] = ' ';
out[i + 1] = ' ';
i += 2;
while (i < content.length) {
if (content[i] === '*' && content[i + 1] === '/') {
out[i] = ' ';
out[i + 1] = ' ';
i += 2;
break;
}
// Preserve newlines so line numbers stay stable for downstream code.
out[i] = content[i] === '\n' || content[i] === '\r' ? content[i] : ' ';
i++;
}
continue;
}
// String literal: "..." or '...'
if (ch === '"' || ch === "'") {
const quote = ch;
out[i] = ' '; // replace opening quote
i++;
while (i < content.length) {
const c = content[i];
if (c === '\\' && i + 1 < content.length) {
// Skip escaped pair (e.g. \" \n \\)
out[i] = ' ';
out[i + 1] = ' ';
i += 2;
continue;
}
if (c === quote) {
out[i] = ' ';
i++;
break;
}
// Preserve newlines; proto technically disallows unescaped newlines
// inside strings, but real files occasionally have them.
out[i] = c === '\n' || c === '\r' ? c : ' ';
i++;
}
continue;
}
out[i] = ch;
i++;
}
return out.join('');
}
function extractServiceBlocks(content: string): Array<{ name: string; body: string }> {
const results: Array<{ name: string; body: string }> = [];
// v1: brace-depth only — braces inside comments or string literals are not filtered (see spec Fix 2)
// Sanitize comments and string literals so braces inside them don't
// throw off the depth counter. The sanitized copy has the same length
// and offsets as the original, so we use it ONLY to scan for service
// headers and braces; the service body we return is sliced from the
// ORIGINAL content to preserve exact source text for downstream use.
const sanitized = stripProtoCommentsAndStrings(content);
const headerRe = /service\s+(\w+)\s*\{/g;
let headerMatch: RegExpExecArray | null;
while ((headerMatch = headerRe.exec(content)) !== null) {
while ((headerMatch = headerRe.exec(sanitized)) !== null) {
const serviceName = headerMatch[1];
const bodyStart = headerMatch.index + headerMatch[0].length;
let depth = 1;
let pos = bodyStart;
while (pos < content.length && depth > 0) {
const ch = content[pos];
while (pos < sanitized.length && depth > 0) {
const ch = sanitized[pos];
if (ch === '{') depth++;
else if (ch === '}') depth--;
pos++;
@@ -75,6 +183,177 @@ function makeContract(
};
}
export interface ProtoServiceInfo {
package: string;
serviceName: string;
methods: string[];
protoPath: string;
}
function normalizeProtoPath(rel: string): string {
return rel.replace(/\\/g, '/');
}
function extractProtoImports(content: string): string[] {
const imports: string[] = [];
const re = /^\s*import\s+"([^"]+)"\s*;/gm;
let match: RegExpExecArray | null;
while ((match = re.exec(content)) !== null) {
imports.push(match[1]);
}
return imports;
}
function longestSharedSegmentRun(aPath: string, bPath: string): number {
const a = aPath.split('/').filter(Boolean);
const b = bPath.split('/').filter(Boolean);
let best = 0;
for (let i = 0; i < a.length; i++) {
for (let j = 0; j < b.length; j++) {
let run = 0;
while (a[i + run] && b[j + run] && a[i + run] === b[j + run]) {
run++;
}
if (run > best) best = run;
}
}
return best;
}
async function buildProtoContext(repoPath: string): Promise<{
packagesByProto: Map<string, string>;
servicesByName: Map<string, ProtoServiceInfo[]>;
}> {
const servicesByName = new Map<string, ProtoServiceInfo[]>();
const protoFiles = await glob('**/*.proto', {
cwd: repoPath,
absolute: false,
nodir: true,
ignore: ['**/node_modules/**', '**/.git/**', '**/vendor/**'],
});
const contents = new Map<string, string>();
for (const rel of protoFiles) {
const content = readSafe(repoPath, rel);
if (!content) continue;
contents.set(normalizeProtoPath(rel), content);
}
const packagesByProto = new Map<string, string>();
const resolvePackage = (protoPath: string, seen = new Set<string>()): string => {
if (packagesByProto.has(protoPath)) return packagesByProto.get(protoPath) ?? '';
if (seen.has(protoPath)) return '';
const content = contents.get(protoPath);
if (!content) return '';
seen.add(protoPath);
const pkgMatch = content.match(/^\s*package\s+([\w.]+)\s*;/m);
if (pkgMatch?.[1]) {
packagesByProto.set(protoPath, pkgMatch[1]);
return pkgMatch[1];
}
for (const importPath of extractProtoImports(content)) {
const normalizedImport = normalizeProtoPath(importPath);
const candidates = [
normalizeProtoPath(
path.posix.normalize(path.posix.join(path.posix.dirname(protoPath), normalizedImport)),
),
normalizedImport,
];
for (const candidate of candidates) {
if (!contents.has(candidate)) continue;
const inheritedPackage = resolvePackage(candidate, seen);
if (inheritedPackage) {
packagesByProto.set(protoPath, inheritedPackage);
return inheritedPackage;
}
}
}
packagesByProto.set(protoPath, '');
return '';
};
for (const rel of protoFiles) {
const normalizedRel = normalizeProtoPath(rel);
const content = contents.get(normalizedRel);
if (!content) continue;
const pkg = resolvePackage(normalizedRel);
const serviceBlocks = extractServiceBlocks(content);
for (const block of serviceBlocks) {
const rpcRe = /rpc\s+(\w+)\s*\(/g;
const methods: string[] = [];
let m: RegExpExecArray | null;
while ((m = rpcRe.exec(block.body)) !== null) {
methods.push(m[1]);
}
const info: ProtoServiceInfo = {
package: pkg,
serviceName: block.name,
methods,
protoPath: normalizedRel,
};
const existing = servicesByName.get(block.name) ?? [];
existing.push(info);
servicesByName.set(block.name, existing);
}
}
return { packagesByProto, servicesByName };
}
export async function buildProtoMap(repoPath: string): Promise<Map<string, ProtoServiceInfo[]>> {
const { servicesByName } = await buildProtoContext(repoPath);
return servicesByName;
}
export function resolveProtoConflict(
serviceName: string,
sourceFilePath: string,
candidates: ProtoServiceInfo[],
): ProtoServiceInfo | null {
if (candidates.length === 0) return null;
if (candidates.length === 1) return candidates[0];
const sourceDir = normalizeProtoPath(path.dirname(sourceFilePath));
const scored = candidates.map((c) => {
const protoDir = normalizeProtoPath(path.dirname(c.protoPath));
return { candidate: c, score: longestSharedSegmentRun(sourceDir, protoDir) };
});
let maxScore = -1;
for (const s of scored) {
if (s.score > maxScore) maxScore = s.score;
}
const winners = scored.filter((s) => s.score === maxScore);
// Path heuristic cannot uniquely identify a winner — refuse to guess.
// Ties (including all-zero ties) would otherwise silently merge unrelated
// services under a fabricated package-qualified contract id.
if (winners.length !== 1) {
const paths = candidates.map((c) => c.protoPath).join(', ');
console.warn(
`[grpc-extractor] Ambiguous proto resolution for service "${serviceName}" from ${sourceFilePath}: ${winners.length} candidates tied at score ${maxScore} among [${paths}] — skipping canonical contract`,
);
return null;
}
return winners[0].candidate;
}
export function serviceContractId(pkg: string, serviceName: string): string {
const prefix = pkg ? `${pkg}.${serviceName}` : serviceName;
return `grpc::${prefix}/*`;
}
// ─── Orchestrator ────────────────────────────────────────────────────
export class GrpcExtractor implements ContractExtractor {
type = 'grpc' as const;
@@ -88,270 +367,116 @@ export class GrpcExtractor implements ContractExtractor {
_repo: RepoHandle,
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
const protoContext = await buildProtoContext(repoPath);
const protoMap = protoContext.servicesByName;
// Proto files — definitive provider source
const protoFiles = await glob('**/*.proto', {
cwd: repoPath,
ignore: ['**/node_modules/**', '**/.git/**', '**/vendor/**'],
nodir: true,
});
for (const rel of protoFiles) {
const content = readSafe(repoPath, rel);
if (content) out.push(...this.parseProtoFile(content, rel));
// ─── Proto files — definitive provider source ─────────────────
// When tree-sitter-proto is available, .proto files are handled by
// the plugin loop below (they're in GRPC_SCAN_GLOB). Otherwise
// emit provider contracts directly from the proto map that
// `buildProtoContext` already built — no second glob / parse pass.
if (!hasProtoPlugin) {
for (const infos of protoMap.values()) {
for (const info of infos) {
for (const methodName of info.methods) {
const cid = contractId(info.package, info.serviceName, methodName);
out.push(
makeContract(
cid,
'provider',
info.protoPath,
`${info.serviceName}.${methodName}`,
0.85,
{
package: info.package,
service: info.serviceName,
method: methodName,
source: 'proto',
},
),
);
}
}
}
}
// Source files — server/client detection
const sourceFiles = await glob('**/*.{go,java,py,ts,tsx,js,jsx}', {
// ─── Source files (+ .proto when plugin available) ────────────
const sourceFiles = await glob(GRPC_SCAN_GLOB, {
cwd: repoPath,
ignore: ['**/node_modules/**', '**/.git/**', '**/vendor/**', '**/dist/**', '**/build/**'],
nodir: true,
});
const parser = new Parser();
for (const rel of sourceFiles) {
const plugin = getPluginForFile(rel);
if (!plugin) continue;
const content = readSafe(repoPath, rel);
if (!content) continue;
const ext = path.extname(rel).toLowerCase();
if (ext === '.go') {
out.push(...this.scanGoProviders(content, rel));
out.push(...this.scanGoConsumers(content, rel));
} else if (ext === '.java') {
out.push(...this.scanJavaProviders(content, rel));
out.push(...this.scanJavaConsumers(content, rel));
} else if (ext === '.py') {
out.push(...this.scanPythonProviders(content, rel));
out.push(...this.scanPythonConsumers(content, rel));
} else if (['.ts', '.tsx', '.js', '.jsx'].includes(ext)) {
out.push(...this.scanTsProviders(content, rel));
let detections: GrpcDetection[] = [];
try {
parser.setLanguage(plugin.language);
const tree = parser.parse(content);
detections = plugin.scan(tree);
} catch {
continue;
}
for (const d of detections) {
const contract = this.detectionToContract(d, rel, protoMap);
if (contract) out.push(contract);
}
}
return this.dedupe(out);
}
private parseProtoFile(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
const pkgMatch = content.match(/^package\s+([\w.]+)\s*;/m);
const pkg = pkgMatch ? pkgMatch[1] : '';
for (const { name: serviceName, body } of extractServiceBlocks(content)) {
const rpcRe = /rpc\s+(\w+)\s*\(/g;
let rpcMatch: RegExpExecArray | null;
while ((rpcMatch = rpcRe.exec(body)) !== null) {
const methodName = rpcMatch[1];
const cid = contractId(pkg, serviceName, methodName);
out.push(
makeContract(cid, 'provider', filePath, `${serviceName}.${methodName}`, 0.85, {
package: pkg,
service: serviceName,
method: methodName,
source: 'proto',
}),
);
}
}
return out;
}
private scanGoProviders(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
// pb.RegisterXxxServer(
const registerRe = /\w+\.Register(\w+)Server\s*\(/g;
let m: RegExpExecArray | null;
while ((m = registerRe.exec(content)) !== null) {
const serviceName = m[1];
out.push(
makeContract(
serviceOnlyContractId(serviceName),
'provider',
filePath,
`Register${serviceName}Server`,
0.8,
{ service: serviceName, source: 'go_register' },
),
);
}
// pb.UnimplementedXxxServer
const unimplRe = /\w+\.Unimplemented(\w+)Server\b/g;
while ((m = unimplRe.exec(content)) !== null) {
const serviceName = m[1];
out.push(
makeContract(
serviceOnlyContractId(serviceName),
'provider',
filePath,
`Unimplemented${serviceName}Server`,
0.8,
{ service: serviceName, source: 'go_unimplemented' },
),
);
}
return out;
}
private scanGoConsumers(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
const re = /\w+\.New(\w+)Client\s*\(/g;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const serviceName = m[1];
out.push(
makeContract(
serviceOnlyContractId(serviceName),
'consumer',
filePath,
`New${serviceName}Client`,
0.7,
{ service: serviceName, source: 'go_client' },
),
);
}
return out;
}
private scanJavaProviders(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
// @GrpcService
if (content.includes('@GrpcService')) {
const implBaseRe = /extends\s+(\w+)Grpc\.(\w+)ImplBase/;
const m = content.match(implBaseRe);
if (m) {
out.push(
makeContract(serviceOnlyContractId(m[1]), 'provider', filePath, m[2], 0.8, {
service: m[1],
source: 'java_grpc_service',
}),
);
} else {
// Try extracting service name from class name
const classRe =
/class\s+(\w*?)(?:Grpc)?(?:Service)?\s+extends\s+(\w+)(?:Grpc\.(\w+))?ImplBase/;
const cm = content.match(classRe);
if (cm) {
const svcName = cm[2].replace(/Grpc$/, '');
out.push(
makeContract(serviceOnlyContractId(svcName), 'provider', filePath, cm[1], 0.8, {
service: svcName,
source: 'java_grpc_service',
}),
);
}
}
}
// extends XxxImplBase (without @GrpcService)
if (!content.includes('@GrpcService')) {
const implRe = /extends\s+(\w+?)(?:Grpc\.(\w+))?ImplBase/;
const m = content.match(implRe);
if (m) {
const svcName = m[2] || m[1].replace(/Grpc$/, '');
out.push(
makeContract(serviceOnlyContractId(svcName), 'provider', filePath, svcName, 0.8, {
service: svcName,
source: 'java_impl_base',
}),
);
}
}
return out;
}
private scanJavaConsumers(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
// XxxGrpc.newBlockingStub( or XxxGrpc.newStub(
const re = /(\w+)Grpc\.new(?:Blocking)?Stub\s*\(/g;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const serviceName = m[1];
out.push(
makeContract(
serviceOnlyContractId(serviceName),
'consumer',
filePath,
`${serviceName}Stub`,
0.7,
{ service: serviceName, source: 'java_stub' },
),
);
}
return out;
}
private scanPythonProviders(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
// add_XxxServicer_to_server(
const re = /add_(\w+?)Servicer_to_server\s*\(/g;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const serviceName = m[1];
out.push(
makeContract(
serviceOnlyContractId(serviceName),
'provider',
filePath,
`add_${serviceName}Servicer_to_server`,
0.8,
{ service: serviceName, source: 'python_servicer' },
),
);
}
return out;
}
private scanPythonConsumers(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
// XxxStub(
const re = /(\w+)Stub\s*\(/g;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const name = m[1];
// Filter out common false positives
if (['Mock', 'Test', 'Fake', 'Stub'].includes(name)) continue;
out.push(
makeContract(serviceOnlyContractId(name), 'consumer', filePath, `${name}Stub`, 0.7, {
service: name,
source: 'python_stub',
}),
);
}
return out;
}
private scanTsProviders(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
// @GrpcMethod('ServiceName', 'MethodName')
const re = /@GrpcMethod\s*\(\s*['"](\w+)['"]\s*,\s*['"](\w+)['"]\s*\)/g;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const serviceName = m[1];
const methodName = m[2];
const cid = contractId('', serviceName, methodName);
out.push(
makeContract(cid, 'provider', filePath, `${serviceName}.${methodName}`, 0.8, {
service: serviceName,
method: methodName,
source: 'ts_grpc_method',
}),
);
}
return out;
/**
* Convert a plugin `GrpcDetection` into a concrete `ExtractedContract`
* by resolving the short service name against the proto map, building
* either a service-level (`grpc::pkg.Svc/*`) or method-level
* (`grpc::pkg.Svc/Method`) contract id, and selecting confidence
* based on whether the proto map had an entry.
*/
private detectionToContract(
d: GrpcDetection,
filePath: string,
protoMap: Map<string, ProtoServiceInfo[]>,
): ExtractedContract | null {
const candidates = protoMap.get(d.serviceName) ?? [];
const proto = resolveProtoConflict(d.serviceName, filePath, candidates);
// If there were proto candidates but resolution was ambiguous, skip
// contract emission rather than fabricating a package-qualified id from
// an arbitrary candidate. resolveProtoConflict already warned.
if (candidates.length > 0 && proto === null) return null;
const pkg = proto?.package ?? '';
const cid = d.methodName
? contractId(pkg, d.serviceName, d.methodName)
: proto
? serviceContractId(pkg, d.serviceName)
: serviceOnlyContractId(d.serviceName);
const confidence = proto ? d.confidenceWithProto : d.confidenceWithoutProto;
const meta: Record<string, unknown> = {
service: d.serviceName,
source: d.source,
};
if (d.methodName) meta.method = d.methodName;
return makeContract(cid, d.role, filePath, d.symbolName, confidence, meta);
}
private dedupe(items: ExtractedContract[]): ExtractedContract[] {
const seen = new Set<string>();
const out: ExtractedContract[] = [];
const byKey = new Map<string, ExtractedContract>();
for (const c of items) {
const k = `${c.contractId}|${c.role}|${c.symbolRef.filePath}`;
if (seen.has(k)) continue;
seen.add(k);
out.push(c);
const existing = byKey.get(k);
if (
!existing ||
c.confidence > existing.confidence ||
(c.confidence === existing.confidence &&
String(c.meta.source) < String(existing.meta.source))
) {
byKey.set(k, c);
}
}
return out;
return Array.from(byKey.values());
}
}
@@ -0,0 +1,109 @@
import Go from 'tree-sitter-go';
import {
compilePatterns,
runCompiledPatterns,
type LanguagePatterns,
} from '../tree-sitter-scanner.js';
import type { GrpcDetection, GrpcLanguagePlugin } from './types.js';
/**
* Go gRPC plugin. Detects:
* - Provider: `pb.RegisterXxxServer(...)` calls
* - Provider: `pb.UnimplementedXxxServer` embedded in a struct
* - Consumer: `pb.NewXxxClient(conn)` calls
*/
const REGISTER_RE = /^Register(\w+)Server$/;
const UNIMPLEMENTED_RE = /^Unimplemented(\w+)Server$/;
const NEW_CLIENT_RE = /^New(\w+)Client$/;
// Any `xxx.<fn>(...)` call — plugin filters the field identifier text.
const SELECTOR_CALL_PATTERNS = compilePatterns({
name: 'go-grpc-selector-call',
language: Go,
patterns: [
{
meta: {},
query: `
(call_expression
function: (selector_expression
field: (field_identifier) @fn))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// Any `qualified_type` used as a struct field — for `pb.UnimplementedXxxServer`.
const STRUCT_EMBEDDING_PATTERNS = compilePatterns({
name: 'go-grpc-struct-embedding',
language: Go,
patterns: [
{
meta: {},
query: `
(struct_type
(field_declaration_list
(field_declaration
type: (qualified_type
name: (type_identifier) @field_type))))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
export const GO_GRPC_PLUGIN: GrpcLanguagePlugin = {
name: 'go-grpc',
language: Go,
scan(tree) {
const out: GrpcDetection[] = [];
for (const match of runCompiledPatterns(SELECTOR_CALL_PATTERNS, tree)) {
const fnNode = match.captures.fn;
if (!fnNode) continue;
const fnText = fnNode.text;
const registerMatch = REGISTER_RE.exec(fnText);
if (registerMatch) {
out.push({
role: 'provider',
serviceName: registerMatch[1],
symbolName: fnText,
source: 'go_register',
confidenceWithProto: 0.8,
confidenceWithoutProto: 0.65,
});
continue;
}
const newClientMatch = NEW_CLIENT_RE.exec(fnText);
if (newClientMatch) {
out.push({
role: 'consumer',
serviceName: newClientMatch[1],
symbolName: fnText,
source: 'go_client',
confidenceWithProto: 0.75,
confidenceWithoutProto: 0.55,
});
continue;
}
}
for (const match of runCompiledPatterns(STRUCT_EMBEDDING_PATTERNS, tree)) {
const fieldNode = match.captures.field_type;
if (!fieldNode) continue;
const unimpl = UNIMPLEMENTED_RE.exec(fieldNode.text);
if (!unimpl) continue;
out.push({
role: 'provider',
serviceName: unimpl[1],
symbolName: fieldNode.text,
source: 'go_unimplemented',
confidenceWithProto: 0.8,
confidenceWithoutProto: 0.65,
});
}
return out;
},
};
@@ -0,0 +1,53 @@
import * as path from 'node:path';
import type { GrpcLanguagePlugin } from './types.js';
import { GO_GRPC_PLUGIN } from './go.js';
import { JAVA_GRPC_PLUGIN } from './java.js';
import { PYTHON_GRPC_PLUGIN } from './python.js';
import { JAVASCRIPT_GRPC_PLUGIN, TYPESCRIPT_GRPC_PLUGIN, TSX_GRPC_PLUGIN } from './node.js';
import { PROTO_GRPC_PLUGIN } from './proto.js';
export type { GrpcDetection, GrpcLanguagePlugin, GrpcRole } from './types.js';
export { PROTO_GRPC_PLUGIN, extractPackageFromTree } from './proto.js';
/**
* File-extension → gRPC language plugin registry. Mirrors the shape
* of `http-patterns/index.ts` and `topic-patterns/index.ts`.
*
* `.proto` files are registered only when `tree-sitter-proto` is
* available (it's an optionalDependency). When absent, the orchestrator
* falls back to the built-in manual proto parser.
*/
const REGISTRY: Record<string, GrpcLanguagePlugin> = {
'.go': GO_GRPC_PLUGIN,
'.java': JAVA_GRPC_PLUGIN,
'.py': PYTHON_GRPC_PLUGIN,
'.js': JAVASCRIPT_GRPC_PLUGIN,
'.jsx': JAVASCRIPT_GRPC_PLUGIN,
'.ts': TYPESCRIPT_GRPC_PLUGIN,
'.tsx': TSX_GRPC_PLUGIN,
...(PROTO_GRPC_PLUGIN ? { '.proto': PROTO_GRPC_PLUGIN } : {}),
};
/**
* Glob for source files worth scanning for gRPC server/client patterns.
* Includes `.proto` when the grammar is available.
*/
export const GRPC_SCAN_GLOB = PROTO_GRPC_PLUGIN
? '**/*.{go,java,py,ts,tsx,js,jsx,proto}'
: '**/*.{go,java,py,ts,tsx,js,jsx}';
/**
* Whether the tree-sitter proto plugin is available. The orchestrator
* uses this to decide between the tree-sitter path and the fallback
* manual parser for `.proto` files.
*/
export const hasProtoPlugin = PROTO_GRPC_PLUGIN !== null;
/**
* Return the gRPC plugin registered for the given file's extension,
* or `undefined` if the extension is not registered.
*/
export function getPluginForFile(rel: string): GrpcLanguagePlugin | undefined {
const ext = path.extname(rel).toLowerCase();
return REGISTRY[ext];
}
@@ -0,0 +1,179 @@
import Parser from 'tree-sitter';
import Java from 'tree-sitter-java';
import {
compilePatterns,
runCompiledPatterns,
type LanguagePatterns,
} from '../tree-sitter-scanner.js';
import type { GrpcDetection, GrpcLanguagePlugin } from './types.js';
/**
* Java gRPC plugin. Detects:
* - Provider: classes extending `XxxServiceGrpc.XxxServiceImplBase`
* (with or without a `@GrpcService` annotation; the annotation
* only affects confidence labelling in the original regex version
* — here we emit a single detection per class and pick the source
* label based on whether the annotation is present).
* - Consumer: `XxxServiceGrpc.newBlockingStub(ch)` /
* `XxxServiceGrpc.newStub(ch)` calls.
*/
const IMPL_BASE_RE = /^(\w+)ImplBase$/;
const GRPC_SUFFIX_RE = /^(\w+)Grpc$/;
// Classes extending `ScopedType.ScopedType` where the inner name ends
// in ImplBase. Covers `XxxServiceGrpc.XxxServiceImplBase`.
// Note: tree-sitter-java's `scoped_type_identifier` exposes its two
// segments as positional `type_identifier` children, NOT as named
// `scope:`/`name:` fields. We match positionally here and rely on the
// grammar's left-to-right ordering: first child = outer, second = inner.
const SCOPED_IMPL_BASE_PATTERNS = compilePatterns({
name: 'java-grpc-scoped-impl-base',
language: Java,
patterns: [
{
meta: {},
query: `
(class_declaration
name: (identifier) @class_name
superclass: (superclass
(scoped_type_identifier
(type_identifier) @outer
(type_identifier) @inner (#match? @inner "ImplBase$")))) @class
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// Classes extending a simple `XxxImplBase` identifier (no scope).
const PLAIN_IMPL_BASE_PATTERNS = compilePatterns({
name: 'java-grpc-plain-impl-base',
language: Java,
patterns: [
{
meta: {},
query: `
(class_declaration
name: (identifier) @class_name
superclass: (superclass
(type_identifier) @plain_type (#match? @plain_type "ImplBase$"))) @class
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// gRPC stub factories: `XxxGrpc.newStub(ch)` / `XxxGrpc.newBlockingStub(ch)`.
const STUB_PATTERNS = compilePatterns({
name: 'java-grpc-stub',
language: Java,
patterns: [
{
meta: {},
query: `
(method_invocation
object: (identifier) @grpc_cls
name: (identifier) @method (#match? @method "^new(Blocking)?Stub$"))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
/**
* Check whether a `class_declaration` node has a `@GrpcService`
* annotation in its modifiers list. In tree-sitter-java, class-level
* annotations live under `(class_declaration (modifiers (marker_annotation|annotation)))`.
*/
function hasGrpcServiceAnnotation(classNode: Parser.SyntaxNode): boolean {
for (let i = 0; i < classNode.namedChildCount; i++) {
const child = classNode.namedChild(i);
if (!child || child.type !== 'modifiers') continue;
for (let j = 0; j < child.namedChildCount; j++) {
const mod = child.namedChild(j);
if (!mod) continue;
if (mod.type !== 'marker_annotation' && mod.type !== 'annotation') continue;
const nameNode = mod.childForFieldName('name');
if (nameNode?.text === 'GrpcService') return true;
}
}
return false;
}
/**
* Given the inner type_identifier text like `AuthServiceImplBase`,
* return the service name (`AuthService`), or null if the text
* doesn't end in `ImplBase`.
*/
function extractServiceFromImplBase(text: string): string | null {
const m = IMPL_BASE_RE.exec(text);
if (!m) return null;
// Strip a trailing `Grpc` on the service name too — the original
// regex replaces `Grpc$` on the extracted prefix.
return m[1].replace(/Grpc$/, '');
}
export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
name: 'java-grpc',
language: Java,
scan(tree) {
const out: GrpcDetection[] = [];
const emittedClassIds = new Set<number>();
// ─── Providers: scoped form (`...Grpc.XxxImplBase`) ─────────────
for (const match of runCompiledPatterns(SCOPED_IMPL_BASE_PATTERNS, tree)) {
const classNode = match.captures.class;
const innerNode = match.captures.inner;
if (!classNode || !innerNode) continue;
const serviceName = extractServiceFromImplBase(innerNode.text);
if (!serviceName) continue;
emittedClassIds.add(classNode.id);
const annotated = hasGrpcServiceAnnotation(classNode);
out.push({
role: 'provider',
serviceName,
symbolName: serviceName,
source: annotated ? 'java_grpc_service' : 'java_impl_base',
confidenceWithProto: 0.8,
confidenceWithoutProto: 0.65,
});
}
// ─── Providers: plain form (`XxxImplBase`) ──────────────────────
for (const match of runCompiledPatterns(PLAIN_IMPL_BASE_PATTERNS, tree)) {
const classNode = match.captures.class;
const plainNode = match.captures.plain_type;
if (!classNode || !plainNode) continue;
if (emittedClassIds.has(classNode.id)) continue;
const serviceName = extractServiceFromImplBase(plainNode.text);
if (!serviceName) continue;
emittedClassIds.add(classNode.id);
const annotated = hasGrpcServiceAnnotation(classNode);
out.push({
role: 'provider',
serviceName,
symbolName: serviceName,
source: annotated ? 'java_grpc_service' : 'java_impl_base',
confidenceWithProto: 0.8,
confidenceWithoutProto: 0.65,
});
}
// ─── Consumers: `XxxGrpc.newBlockingStub(...)` / `newStub(...)` ─
for (const match of runCompiledPatterns(STUB_PATTERNS, tree)) {
const grpcClsNode = match.captures.grpc_cls;
if (!grpcClsNode) continue;
const grpcMatch = GRPC_SUFFIX_RE.exec(grpcClsNode.text);
if (!grpcMatch) continue;
const serviceName = grpcMatch[1];
out.push({
role: 'consumer',
serviceName,
symbolName: `${serviceName}Stub`,
source: 'java_stub',
confidenceWithProto: 0.75,
confidenceWithoutProto: 0.55,
});
}
return out;
},
};
@@ -0,0 +1,314 @@
import Parser from 'tree-sitter';
import JavaScript from 'tree-sitter-javascript';
import TypeScript from 'tree-sitter-typescript';
import {
compilePatterns,
runCompiledPatterns,
unquoteLiteral,
type CompiledPatterns,
type LanguagePatterns,
type PatternSpec,
} from '../tree-sitter-scanner.js';
import type { GrpcDetection, GrpcLanguagePlugin } from './types.js';
/**
* Node.js / TypeScript gRPC plugin family. Detects:
* - Provider: NestJS `@GrpcMethod('Service', 'Method')` decorators
* - Consumer: NestJS `@GrpcClient(...) readonly x!: XxxServiceClient`
* - Consumer: `client.getService<X>('AuthService')`
* - Consumer: `new XxxServiceClient(...)` (generated client constructor)
* - Consumer: `new foo.bar.Xxx(...)` when the file uses
* `loadPackageDefinition` (gRPC dynamic proto loader)
*
* As with the HTTP `node.ts`, pattern sources are defined once and
* compiled against three grammar variants (JS / TS / TSX) because
* `Parser.Query` is not portable across grammar objects.
*/
const SERVICE_CLIENT_RE = /^(\w+Service)Client$/;
const CAPITALIZED_SERVICE_RE = /^[A-Z]\w+$/;
// @GrpcMethod('Service', 'Method')
const GRPC_METHOD_SPEC: PatternSpec<Record<string, never>> = {
meta: {},
query: `
(decorator
(call_expression
function: (identifier) @dec (#eq? @dec "GrpcMethod")
arguments: (arguments
. [(string) (template_string)] @service
. [(string) (template_string)] @method)))
`,
};
// @GrpcClient(...) standalone decorator — the plugin walks to the next
// sibling (a field definition) to read its type annotation.
const GRPC_CLIENT_SPEC: PatternSpec<Record<string, never>> = {
meta: {},
query: `
(decorator
(call_expression
function: (identifier) @dec (#eq? @dec "GrpcClient"))) @grpc_client_decorator
`,
};
// `.getService<X>('AuthService')` / `.getService('AuthService')`
const GET_SERVICE_SPEC: PatternSpec<Record<string, never>> = {
meta: {},
query: `
(call_expression
function: (member_expression
property: (property_identifier) @method (#eq? @method "getService"))
arguments: (arguments . [(string) (template_string)] @service))
`,
};
// `new XxxServiceClient(...)` — bare identifier constructor.
const NEW_SIMPLE_CTOR_SPEC: PatternSpec<Record<string, never>> = {
meta: {},
query: `
(new_expression
constructor: (identifier) @ctor)
`,
};
// `new foo.bar.XxxService(...)` — qualified constructor.
const NEW_QUALIFIED_CTOR_SPEC: PatternSpec<Record<string, never>> = {
meta: {},
query: `
(new_expression
constructor: (member_expression
property: (property_identifier) @ctor))
`,
};
// Detect whether the file uses `loadPackageDefinition` (gRPC dynamic
// proto loader). Matches either a bare call or an `obj.loadPackageDefinition(...)`
// call. Plugin gates the qualified-constructor consumer on this —
// structural check avoids materializing `tree.rootNode.text` for every file.
const LOAD_PACKAGE_DEFINITION_SPEC: PatternSpec<Record<string, never>> = {
meta: {},
query: `
(call_expression
function: [
(identifier) @fn (#eq? @fn "loadPackageDefinition")
(member_expression property: (property_identifier) @fn (#eq? @fn "loadPackageDefinition"))
])
`,
};
interface NodeGrpcPatternBundle {
grpcMethod: CompiledPatterns<Record<string, never>>;
grpcClient: CompiledPatterns<Record<string, never>>;
getService: CompiledPatterns<Record<string, never>>;
newSimpleCtor: CompiledPatterns<Record<string, never>>;
newQualifiedCtor: CompiledPatterns<Record<string, never>>;
loadPackageDefinition: CompiledPatterns<Record<string, never>>;
}
function compileBundle(language: unknown, name: string): NodeGrpcPatternBundle {
const mk = (spec: PatternSpec<Record<string, never>>, suffix: string) =>
compilePatterns({
name: `${name}-${suffix}`,
language,
patterns: [spec],
} satisfies LanguagePatterns<Record<string, never>>);
return {
grpcMethod: mk(GRPC_METHOD_SPEC, 'grpc-method'),
grpcClient: mk(GRPC_CLIENT_SPEC, 'grpc-client'),
getService: mk(GET_SERVICE_SPEC, 'get-service'),
newSimpleCtor: mk(NEW_SIMPLE_CTOR_SPEC, 'new-simple-ctor'),
newQualifiedCtor: mk(NEW_QUALIFIED_CTOR_SPEC, 'new-qualified-ctor'),
loadPackageDefinition: mk(LOAD_PACKAGE_DEFINITION_SPEC, 'load-package-definition'),
};
}
const JAVASCRIPT_BUNDLE = compileBundle(JavaScript, 'javascript-grpc');
const TYPESCRIPT_BUNDLE = compileBundle(TypeScript.typescript, 'typescript-grpc');
const TSX_BUNDLE = compileBundle(TypeScript.tsx, 'tsx-grpc');
/**
* Given a `@GrpcClient(...)` decorator node, find the type annotation
* text of the field it decorates (e.g. `AuthServiceClient`).
*
* In tree-sitter-typescript, decorators on class fields can appear in
* two configurations:
* - As a CHILD of `public_field_definition` alongside the field's
* type annotation (the common case for NestJS `@GrpcClient`).
* - As a SIBLING of the field in `class_body` (for method
* decorators, but kept for resilience against grammar variants).
* We walk the parent container and search for a type annotation.
*/
function resolveGrpcClientFieldType(decoratorNode: Parser.SyntaxNode): string | null {
const parent = decoratorNode.parent;
if (!parent) return null;
// Case 1: decorator is a child of the field definition — search
// the parent itself (which is the field definition) for a
// type_annotation child.
if (parent.type === 'public_field_definition' || parent.type.endsWith('field_definition')) {
return findFirstTypeAnnotationText(parent);
}
// Case 2: decorator is a sibling of the field in a class_body — walk
// forward through subsequent siblings until we find a node containing
// a type annotation.
for (let i = 0; i < parent.namedChildCount; i++) {
const child = parent.namedChild(i);
if (child && child.id === decoratorNode.id) {
for (let j = i + 1; j < parent.namedChildCount; j++) {
const next = parent.namedChild(j);
if (!next) continue;
if (next.type === 'decorator') continue;
const typeText = findFirstTypeAnnotationText(next);
if (typeText) return typeText;
return null;
}
return null;
}
}
return null;
}
/**
* Recursively search `node` for the first `type_annotation` child and
* return the text of its inner `type_identifier`, or null. Handles
* both `public_field_definition` and its variants.
*/
function findFirstTypeAnnotationText(node: Parser.SyntaxNode): string | null {
if (node.type === 'type_annotation') {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (!child) continue;
if (child.type === 'type_identifier') return child.text;
}
return null;
}
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (!child) continue;
const found = findFirstTypeAnnotationText(child);
if (found) return found;
}
return null;
}
function scanBundle(bundle: NodeGrpcPatternBundle, tree: Parser.Tree): GrpcDetection[] {
const out: GrpcDetection[] = [];
// ─── Provider: @GrpcMethod('Service', 'Method') ──────────────────
for (const match of runCompiledPatterns(bundle.grpcMethod, tree)) {
const svcNode = match.captures.service;
const methodNode = match.captures.method;
if (!svcNode || !methodNode) continue;
const svc = unquoteLiteral(svcNode.text);
const mth = unquoteLiteral(methodNode.text);
if (!svc || !mth) continue;
out.push({
role: 'provider',
serviceName: svc,
symbolName: `${svc}.${mth}`,
source: 'ts_grpc_method',
methodName: mth,
// @GrpcMethod hard-coded confidence 0.8 in the original code
// regardless of whether the proto map has a match.
confidenceWithProto: 0.8,
confidenceWithoutProto: 0.8,
});
}
// ─── Consumer: @GrpcClient() field with XxxServiceClient type ────
for (const match of runCompiledPatterns(bundle.grpcClient, tree)) {
const decoratorNode = match.captures.grpc_client_decorator;
if (!decoratorNode) continue;
const typeText = resolveGrpcClientFieldType(decoratorNode);
if (!typeText) continue;
const svcMatch = SERVICE_CLIENT_RE.exec(typeText);
if (!svcMatch) continue;
const serviceName = svcMatch[1];
out.push({
role: 'consumer',
serviceName,
symbolName: `${serviceName}Client`,
source: 'ts_grpc_client_decorator',
confidenceWithProto: 0.75,
confidenceWithoutProto: 0.55,
});
}
// ─── Consumer: client.getService<X>('Service') ───────────────────
for (const match of runCompiledPatterns(bundle.getService, tree)) {
const svcNode = match.captures.service;
if (!svcNode) continue;
const svc = unquoteLiteral(svcNode.text);
if (!svc) continue;
out.push({
role: 'consumer',
serviceName: svc,
symbolName: `${svc}Client`,
source: 'ts_client_grpc_get_service',
confidenceWithProto: 0.75,
confidenceWithoutProto: 0.55,
});
}
// ─── Consumer: new XxxServiceClient(...) ─────────────────────────
for (const match of runCompiledPatterns(bundle.newSimpleCtor, tree)) {
const ctorNode = match.captures.ctor;
if (!ctorNode) continue;
const svcMatch = SERVICE_CLIENT_RE.exec(ctorNode.text);
if (!svcMatch) continue;
const serviceName = svcMatch[1];
out.push({
role: 'consumer',
serviceName,
symbolName: `${serviceName}Client`,
source: 'ts_generated_client',
confidenceWithProto: 0.75,
confidenceWithoutProto: 0.55,
});
}
// ─── Consumer: loadPackageDefinition dynamic proto loader ────────
// Only emit when the file uses loadPackageDefinition, otherwise a
// generic `new foo.bar.Something()` in unrelated code would falsely
// register as a gRPC consumer. Check structurally via a dedicated
// query — avoids materializing `tree.rootNode.text` for the whole
// file (expensive on large files).
const usesLoadPackage = runCompiledPatterns(bundle.loadPackageDefinition, tree).length > 0;
if (usesLoadPackage) {
for (const match of runCompiledPatterns(bundle.newQualifiedCtor, tree)) {
const ctorNode = match.captures.ctor;
if (!ctorNode) continue;
if (!CAPITALIZED_SERVICE_RE.test(ctorNode.text)) continue;
out.push({
role: 'consumer',
serviceName: ctorNode.text,
symbolName: `${ctorNode.text}Client`,
source: 'ts_load_package_definition',
confidenceWithProto: 0.75,
confidenceWithoutProto: 0.55,
});
}
}
return out;
}
export const JAVASCRIPT_GRPC_PLUGIN: GrpcLanguagePlugin = {
name: 'javascript-grpc',
language: JavaScript,
scan: (tree) => scanBundle(JAVASCRIPT_BUNDLE, tree),
};
export const TYPESCRIPT_GRPC_PLUGIN: GrpcLanguagePlugin = {
name: 'typescript-grpc',
language: TypeScript.typescript,
scan: (tree) => scanBundle(TYPESCRIPT_BUNDLE, tree),
};
export const TSX_GRPC_PLUGIN: GrpcLanguagePlugin = {
name: 'tsx-grpc',
language: TypeScript.tsx,
scan: (tree) => scanBundle(TSX_BUNDLE, tree),
};
@@ -0,0 +1,147 @@
import { createRequire } from 'node:module';
import {
compilePatterns,
runCompiledPatterns,
type CompiledPatterns,
type LanguagePatterns,
} from '../tree-sitter-scanner.js';
import type { GrpcDetection, GrpcLanguagePlugin } from './types.js';
/**
* Protobuf (.proto) tree-sitter plugin for gRPC contract extraction.
*
* Uses `tree-sitter-proto` (coder3101/tree-sitter-proto) as an
* optionalDependency — if the grammar is not installed (e.g. native
* compilation failed on an unusual platform), the plugin exports
* `null` and the orchestrator falls back to the existing manual
* string-sanitizing parser.
*
* The grammar is vendored in `vendor/tree-sitter-proto/` with
* parser.c regenerated against tree-sitter-cli 0.24 (ABI version 14)
* so it is compatible with the project's tree-sitter 0.25 runtime.
*/
const _require = createRequire(import.meta.url);
let ProtoGrammar: unknown = null;
try {
ProtoGrammar = _require('tree-sitter-proto');
} catch {
// Grammar not installed — PROTO_GRPC_PLUGIN will be null.
}
let PACKAGE_PATTERNS: CompiledPatterns<Record<string, never>> | null = null;
let SERVICE_PATTERNS: CompiledPatterns<Record<string, never>> | null = null;
if (ProtoGrammar) {
try {
// Validate that the grammar actually loads end-to-end: compile queries
// AND parse + walk a trivial proto file. tree-sitter's internal
// `initializeLanguageNodeClasses` can fail with a TDZ error in some
// test runners (vitest forks) when SyntaxNode isn't fully initialized
// yet. Catching that here ensures `PROTO_GRPC_PLUGIN` stays null and
// the orchestrator falls back to the manual parser.
// eslint-disable-next-line @typescript-eslint/no-explicit-any
const _Parser = _require('tree-sitter') as any;
// Smoke-test: parse + setLanguage to verify the grammar is
// end-to-end compatible with this tree-sitter runtime.
const _testParser = new _Parser();
_testParser.setLanguage(ProtoGrammar);
_testParser.parse('service X { rpc Y (R) returns (R); }');
PACKAGE_PATTERNS = compilePatterns({
name: 'proto-package',
language: ProtoGrammar,
patterns: [
{
meta: {},
query: `(package (full_ident) @pkg)`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
SERVICE_PATTERNS = compilePatterns({
name: 'proto-service',
language: ProtoGrammar,
patterns: [
{
meta: {},
query: `
(service
(service_name) @service_name
(rpc
(rpc_name) @rpc_name))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
} catch {
// Compilation failed (grammar ABI mismatch?) — fall back to null.
PACKAGE_PATTERNS = null;
SERVICE_PATTERNS = null;
ProtoGrammar = null;
}
}
function buildPlugin(): GrpcLanguagePlugin | null {
if (!ProtoGrammar || !PACKAGE_PATTERNS || !SERVICE_PATTERNS) return null;
const pkgPatterns = PACKAGE_PATTERNS;
const svcPatterns = SERVICE_PATTERNS;
return {
name: 'proto-grpc',
language: ProtoGrammar,
scan(tree) {
const out: GrpcDetection[] = [];
// Extract `package` declaration (first match wins).
let pkg = '';
for (const match of runCompiledPatterns(pkgPatterns, tree)) {
const pkgNode = match.captures.pkg;
if (pkgNode) {
pkg = pkgNode.text;
break;
}
}
// Extract `service → rpc` pairs. The query returns one match per
// (service, rpc) combination thanks to the nested structure.
for (const match of runCompiledPatterns(svcPatterns, tree)) {
const serviceNode = match.captures.service_name;
const rpcNode = match.captures.rpc_name;
if (!serviceNode || !rpcNode) continue;
const serviceName = serviceNode.text;
const methodName = rpcNode.text;
out.push({
role: 'provider',
serviceName,
symbolName: `${serviceName}.${methodName}`,
source: 'proto',
methodName,
// Proto definitions are the canonical source of truth — always
// high confidence regardless of cross-referencing.
confidenceWithProto: 0.85,
confidenceWithoutProto: 0.85,
});
}
return out;
},
};
}
/**
* The proto plugin, or `null` if tree-sitter-proto is not available.
* The orchestrator checks this at import time and decides whether to
* use the tree-sitter path or the fallback manual parser.
*/
export const PROTO_GRPC_PLUGIN: GrpcLanguagePlugin | null = buildPlugin();
/** The package declaration text from a proto file's tree. */
export function extractPackageFromTree(tree: import('tree-sitter').Tree): string {
if (!PACKAGE_PATTERNS) return '';
for (const match of runCompiledPatterns(PACKAGE_PATTERNS, tree)) {
const pkgNode = match.captures.pkg;
if (pkgNode) return pkgNode.text;
}
return '';
}
@@ -0,0 +1,77 @@
import Python from 'tree-sitter-python';
import {
compilePatterns,
runCompiledPatterns,
type LanguagePatterns,
} from '../tree-sitter-scanner.js';
import type { GrpcDetection, GrpcLanguagePlugin } from './types.js';
/**
* Python gRPC plugin. Detects:
* - Provider: `add_XxxServicer_to_server(...)` calls (bare identifier
* or qualified attribute form `auth_pb2_grpc.add_XxxServicer_to_server`)
* - Consumer: `XxxStub(channel)` calls (bare or `auth_pb2_grpc.XxxStub`)
*/
const ADD_SERVICER_RE = /^add_(\w+)Servicer_to_server$/;
const STUB_RE = /^(\w+)Stub$/;
/** Reserved names that would produce garbage service names. */
const STUB_IGNORE = new Set(['Mock', 'Test', 'Fake', 'Stub']);
// Any call whose target is either a bare identifier or an attribute
// access (`obj.method`). The plugin filters the function name in JS.
const CALL_PATTERNS = compilePatterns({
name: 'python-grpc-call',
language: Python,
patterns: [
{
meta: {},
query: `
(call
function: [
(identifier) @fn
(attribute attribute: (identifier) @fn)
])
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
export const PYTHON_GRPC_PLUGIN: GrpcLanguagePlugin = {
name: 'python-grpc',
language: Python,
scan(tree) {
const out: GrpcDetection[] = [];
for (const match of runCompiledPatterns(CALL_PATTERNS, tree)) {
const fnNode = match.captures.fn;
if (!fnNode) continue;
const fnText = fnNode.text;
const addServicer = ADD_SERVICER_RE.exec(fnText);
if (addServicer) {
out.push({
role: 'provider',
serviceName: addServicer[1],
symbolName: fnText,
source: 'python_servicer',
confidenceWithProto: 0.8,
confidenceWithoutProto: 0.65,
});
continue;
}
const stubMatch = STUB_RE.exec(fnText);
if (stubMatch && !STUB_IGNORE.has(stubMatch[1])) {
out.push({
role: 'consumer',
serviceName: stubMatch[1],
symbolName: fnText,
source: 'python_stub',
confidenceWithProto: 0.75,
confidenceWithoutProto: 0.55,
});
}
}
return out;
},
};
@@ -0,0 +1,54 @@
import type Parser from 'tree-sitter';
/**
* Shared types for the grpc-extractor language plugins.
*
* Each plugin lives in its own file (java.ts, go.ts, ...) and owns the
* tree-sitter grammar import + query sources. The top-level
* `grpc-extractor.ts` orchestrator only knows about this type module
* and the plugin registry (`./index.ts`). It MUST NOT import any
* grammar or query text directly.
*/
export type GrpcRole = 'provider' | 'consumer';
/**
* One raw gRPC detection produced by a plugin's `scan()` function. The
* orchestrator uses the proto map to resolve the full package-qualified
* contract id and choose a confidence based on whether the proto was
* found.
*
* Most patterns produce service-level detections; `TS @GrpcMethod` is
* the only pattern that captures an explicit `methodName`, producing
* a method-level contract (`grpc::pkg.Service/Method`).
*/
export interface GrpcDetection {
role: GrpcRole;
/** Short service name, e.g. `"AuthService"`. */
serviceName: string;
/** Symbol name emitted into the contract's symbolRef. */
symbolName: string;
/** Metadata source label (goes into `meta.source`). */
source: string;
/** Explicit method name; set only by TS `@GrpcMethod`. */
methodName?: string;
/** Confidence when the proto map resolves the service. */
confidenceWithProto: number;
/** Confidence when the proto map has no entry. */
confidenceWithoutProto: number;
}
/**
* One language-scoped gRPC plugin. Plugins own the tree-sitter grammar
* and a `scan(tree)` function that returns zero or more
* `GrpcDetection`s. The plugin is free to run multiple compiled query
* bundles and walk the AST to cross-reference captures.
*
* `language` is typed `unknown` for the same reason as in
* `tree-sitter-scanner.ts`.
*/
export interface GrpcLanguagePlugin {
name: string;
language: unknown;
scan(tree: Parser.Tree): GrpcDetection[];
}
@@ -0,0 +1,224 @@
import Go from 'tree-sitter-go';
import {
compilePatterns,
runCompiledPatterns,
unquoteLiteral,
type LanguagePatterns,
} from '../tree-sitter-scanner.js';
import type { HttpDetection, HttpLanguagePlugin } from './types.js';
/**
* Go HTTP plugin. Handles:
* - gin / echo / chi framework routing — `r.GET("/path", handler)`
* - net/http stdlib — `http.HandleFunc("/path", handler)`
* - net/http consumer — `http.Get(...)`, `http.NewRequest("METHOD", ...)`
* - resty consumer — `client.R().Delete("/path")`
*/
// ─── Provider: framework routing ──────────────────────────────────────
// Matches `\w+\.GET(...)` etc. (gin, echo, chi all share this shape).
// Captures the HTTP method (field name), path literal, and handler
// identifier passed as the second argument.
const FRAMEWORK_ROUTE_PATTERNS = compilePatterns({
name: 'go-framework-route',
language: Go,
patterns: [
{
meta: {},
query: `
(call_expression
function: (selector_expression
field: (field_identifier) @http_method (#match? @http_method "^(GET|POST|PUT|DELETE|PATCH)$"))
arguments: (argument_list
(interpreted_string_literal) @path
(identifier) @handler))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Provider: net/http `http.HandleFunc("/p", handler)` ─────────────
const HANDLE_FUNC_PATTERNS = compilePatterns({
name: 'go-handle-func',
language: Go,
patterns: [
{
meta: {},
query: `
(call_expression
function: (selector_expression
operand: (identifier) @pkg (#eq? @pkg "http")
field: (field_identifier) @fn (#eq? @fn "HandleFunc"))
arguments: (argument_list
(interpreted_string_literal) @path
(identifier) @handler))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: net/http stdlib Get / Post / Head ─────────────────────
const HTTP_CLIENT_METHOD_TO_HTTP: Record<string, string> = {
Get: 'GET',
Post: 'POST',
Head: 'GET', // HEAD has no body semantics we care about — treat as GET for contract matching
};
const HTTP_CLIENT_PATTERNS = compilePatterns({
name: 'go-http-client',
language: Go,
patterns: [
{
meta: {},
query: `
(call_expression
function: (selector_expression
operand: (identifier) @pkg (#eq? @pkg "http")
field: (field_identifier) @fn (#match? @fn "^(Get|Post|Head)$"))
arguments: (argument_list . (interpreted_string_literal) @path))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: net/http `http.NewRequest("METHOD", "/path", ...)` ────
const NEW_REQUEST_PATTERNS = compilePatterns({
name: 'go-new-request',
language: Go,
patterns: [
{
meta: {},
query: `
(call_expression
function: (selector_expression
operand: (identifier) @pkg (#eq? @pkg "http")
field: (field_identifier) @fn (#eq? @fn "NewRequest"))
arguments: (argument_list
.
(interpreted_string_literal) @http_method
(interpreted_string_literal) @path))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: resty `client.R().Delete("/path")` ─────────────────────
// Matches any chained call whose receiver is `something.R()` and whose
// method name is an HTTP verb. This is how go-resty's fluent API looks.
const RESTY_PATTERNS = compilePatterns({
name: 'go-resty',
language: Go,
patterns: [
{
meta: {},
query: `
(call_expression
function: (selector_expression
operand: (call_expression
function: (selector_expression
field: (field_identifier) @r (#eq? @r "R")))
field: (field_identifier) @http_method (#match? @http_method "^(Get|Post|Put|Delete|Patch)$"))
arguments: (argument_list . (interpreted_string_literal) @path))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
export const GO_HTTP_PLUGIN: HttpLanguagePlugin = {
name: 'go-http',
language: Go,
scan(tree) {
const out: HttpDetection[] = [];
// Framework providers: r.GET/POST/... with handler identifier
for (const match of runCompiledPatterns(FRAMEWORK_ROUTE_PATTERNS, tree)) {
const methodNode = match.captures.http_method;
const pathNode = match.captures.path;
const handlerNode = match.captures.handler;
if (!methodNode || !pathNode) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'provider',
framework: 'go-framework',
method: methodNode.text.toUpperCase(),
path,
name: handlerNode?.text ?? null,
confidence: 0.8,
});
}
// net/http HandleFunc: default method GET
for (const match of runCompiledPatterns(HANDLE_FUNC_PATTERNS, tree)) {
const pathNode = match.captures.path;
const handlerNode = match.captures.handler;
if (!pathNode) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'provider',
framework: 'go-stdlib',
method: 'GET',
path,
name: handlerNode?.text ?? null,
confidence: 0.8,
});
}
// net/http client: http.Get/Post/Head
for (const match of runCompiledPatterns(HTTP_CLIENT_PATTERNS, tree)) {
const fnNode = match.captures.fn;
const pathNode = match.captures.path;
if (!fnNode || !pathNode) continue;
const httpMethod = HTTP_CLIENT_METHOD_TO_HTTP[fnNode.text];
if (!httpMethod) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'go-stdlib',
method: httpMethod,
path,
name: null,
confidence: 0.7,
});
}
// net/http NewRequest
for (const match of runCompiledPatterns(NEW_REQUEST_PATTERNS, tree)) {
const methodNode = match.captures.http_method;
const pathNode = match.captures.path;
if (!methodNode || !pathNode) continue;
const method = unquoteLiteral(methodNode.text);
const path = unquoteLiteral(pathNode.text);
if (method === null || path === null) continue;
out.push({
role: 'consumer',
framework: 'go-stdlib',
method: method.toUpperCase(),
path,
name: null,
confidence: 0.7,
});
}
// resty
for (const match of runCompiledPatterns(RESTY_PATTERNS, tree)) {
const methodNode = match.captures.http_method;
const pathNode = match.captures.path;
if (!methodNode || !pathNode) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'go-resty',
method: methodNode.text.toUpperCase(),
path,
name: null,
confidence: 0.7,
});
}
return out;
},
};
@@ -0,0 +1,50 @@
import * as path from 'node:path';
import type { HttpLanguagePlugin } from './types.js';
import { JAVA_HTTP_PLUGIN } from './java.js';
import { GO_HTTP_PLUGIN } from './go.js';
import { PYTHON_HTTP_PLUGIN } from './python.js';
import { PHP_HTTP_PLUGIN } from './php.js';
import { JAVASCRIPT_HTTP_PLUGIN, TYPESCRIPT_HTTP_PLUGIN, TSX_HTTP_PLUGIN } from './node.js';
export type { HttpDetection, HttpLanguagePlugin, HttpRole } from './types.js';
/**
* File-extension → HTTP language plugin registry. The top-level
* orchestrator (`http-route-extractor.ts`) looks up the plugin for each
* file it visits and delegates the tree-sitter scanning to the plugin.
*
* Keys are lowercase extensions including the leading dot. To add a
* new language, drop a `http-patterns/<lang>.ts` that exports a
* `HttpLanguagePlugin`, import it here and register the extension(s).
* No edits to `http-route-extractor.ts` are required.
*/
const REGISTRY: Record<string, HttpLanguagePlugin> = {
'.java': JAVA_HTTP_PLUGIN,
'.go': GO_HTTP_PLUGIN,
'.py': PYTHON_HTTP_PLUGIN,
'.php': PHP_HTTP_PLUGIN,
'.js': JAVASCRIPT_HTTP_PLUGIN,
'.jsx': JAVASCRIPT_HTTP_PLUGIN,
'.ts': TYPESCRIPT_HTTP_PLUGIN,
'.tsx': TSX_HTTP_PLUGIN,
};
/**
* Glob for files worth scanning for HTTP routes. Kept alongside the
* registry so adding a new language widens the glob in one edit.
*
* `.vue` / `.svelte` files are intentionally omitted for the source-scan
* path — they need their own grammar-aware extraction and the existing
* regex fallback for them was never very accurate. The graph-assisted
* Strategy A still handles them via the ingestion pipeline.
*/
export const HTTP_SCAN_GLOB = '**/*.{ts,tsx,js,jsx,java,go,py,php}';
/**
* Return the HTTP plugin registered for the given file's extension,
* or `undefined` if the extension is not registered.
*/
export function getPluginForFile(rel: string): HttpLanguagePlugin | undefined {
const ext = path.extname(rel).toLowerCase();
return REGISTRY[ext];
}
@@ -0,0 +1,267 @@
import Parser from 'tree-sitter';
import Java from 'tree-sitter-java';
import {
compilePatterns,
runCompiledPatterns,
unquoteLiteral,
type LanguagePatterns,
} from '../tree-sitter-scanner.js';
import type { HttpDetection, HttpLanguagePlugin } from './types.js';
/**
* Java HTTP plugin. Handles:
* - Spring `@RequestMapping` class prefixes + `@(Get|Post|...)Mapping` method annotations
* - Spring `RestTemplate.getForObject/...`, `WebClient.method(HttpMethod.X, ...)`
* - OkHttp `new Request.Builder().url("...")`
*
* The plugin runs two pattern bundles: one to collect class-level
* `@RequestMapping` prefixes keyed by the enclosing class node, and a
* second to match method-level annotations. The `scan` function walks
* up from each matched annotation to find its enclosing class and
* combines the prefix with the method path.
*/
const METHOD_ANNOTATION_TO_HTTP: Record<string, string> = {
GetMapping: 'GET',
PostMapping: 'POST',
PutMapping: 'PUT',
DeleteMapping: 'DELETE',
PatchMapping: 'PATCH',
};
// ─── Provider: Spring class-level @RequestMapping prefix ──────────────
const SPRING_CLASS_PREFIX_PATTERNS = compilePatterns({
name: 'java-spring-class-prefix',
language: Java,
patterns: [
{
meta: {},
query: `
(class_declaration
(modifiers
(annotation
name: (identifier) @ann (#eq? @ann "RequestMapping")
arguments: (annotation_argument_list (string_literal) @prefix)))) @class
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Provider: Spring @(Get|Post|...)Mapping method annotations ───────
const SPRING_METHOD_ROUTE_PATTERNS = compilePatterns({
name: 'java-spring-method-route',
language: Java,
patterns: [
{
meta: {},
query: `
(method_declaration
(modifiers
(annotation
name: (identifier) @ann (#match? @ann "^(Get|Post|Put|Delete|Patch)Mapping$")
arguments: (annotation_argument_list (string_literal) @path)))
name: (identifier) @method_name) @method
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: Spring RestTemplate (object-named + method-named) ──────
// RestTemplate.getForObject / getForEntity → GET
// RestTemplate.postForObject / postForEntity → POST
// RestTemplate.put → PUT
// RestTemplate.delete → DELETE
// RestTemplate.patchForObject → PATCH
const REST_TEMPLATE_TO_HTTP: Record<string, string> = {
getForObject: 'GET',
getForEntity: 'GET',
postForObject: 'POST',
postForEntity: 'POST',
put: 'PUT',
delete: 'DELETE',
patchForObject: 'PATCH',
};
interface RestTemplateMeta {
framework: 'spring-rest-template';
}
const REST_TEMPLATE_PATTERNS = compilePatterns({
name: 'java-rest-template',
language: Java,
patterns: [
{
meta: { framework: 'spring-rest-template' },
query: `
(method_invocation
object: (identifier) @obj (#eq? @obj "restTemplate")
name: (identifier) @method
arguments: (argument_list . (string_literal) @path))
`,
},
],
} satisfies LanguagePatterns<RestTemplateMeta>);
// ─── Consumer: Spring WebClient — webClient.method(HttpMethod.X, "path") ─
const WEB_CLIENT_PATTERNS = compilePatterns({
name: 'java-web-client',
language: Java,
patterns: [
{
meta: {},
query: `
(method_invocation
object: (identifier) @obj (#eq? @obj "webClient")
name: (identifier) @method (#eq? @method "method")
arguments: (argument_list
(field_access
object: (identifier) @httpMethodCls (#eq? @httpMethodCls "HttpMethod")
field: (identifier) @http_method)
(string_literal) @path))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: OkHttp `new Request.Builder().url("path")` ─────────────
// Note: `Request.Builder` is a `scoped_type_identifier` whose text includes
// the dot, so `#eq?` against the literal string matches cleanly (no need
// to escape a regex dot).
const OK_HTTP_PATTERNS = compilePatterns({
name: 'java-okhttp',
language: Java,
patterns: [
{
meta: {},
query: `
(method_invocation
object: (object_creation_expression
type: (scoped_type_identifier) @type (#eq? @type "Request.Builder"))
name: (identifier) @method (#eq? @method "url")
arguments: (argument_list . (string_literal) @path))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
/**
* Find the nearest enclosing class_declaration ancestor for a node, or
* null if the node is top-level. Tree-sitter's SyntaxNode.parent walks
* one level at a time.
*/
function findEnclosingClass(node: Parser.SyntaxNode): Parser.SyntaxNode | null {
let cur: Parser.SyntaxNode | null = node.parent;
while (cur) {
if (cur.type === 'class_declaration') return cur;
cur = cur.parent;
}
return null;
}
/**
* Join a class-level prefix and a method-level path into a single URL
* path. Mirrors the semantics of the original regex implementation:
* strip trailing slashes on the prefix, then ensure a single slash
* between prefix and method path.
*/
function joinPath(prefix: string, methodPath: string): string {
const cleanPrefix = prefix.replace(/^\/+/, '').replace(/\/+$/, '');
const cleanSub = methodPath.replace(/^\/+/, '');
if (!cleanPrefix) return `/${cleanSub}`;
return `/${cleanPrefix}/${cleanSub}`;
}
export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
name: 'java-http',
language: Java,
scan(tree) {
const out: HttpDetection[] = [];
// ─── Providers: Spring class prefix + method annotations ────────
const prefixByClassId = new Map<number, string>();
for (const match of runCompiledPatterns(SPRING_CLASS_PREFIX_PATTERNS, tree)) {
const prefixNode = match.captures.prefix;
const classNode = match.captures.class;
if (!prefixNode || !classNode) continue;
const prefix = unquoteLiteral(prefixNode.text);
if (prefix !== null) prefixByClassId.set(classNode.id, prefix);
}
for (const match of runCompiledPatterns(SPRING_METHOD_ROUTE_PATTERNS, tree)) {
const annNode = match.captures.ann;
const pathNode = match.captures.path;
const nameNode = match.captures.method_name;
const methodNode = match.captures.method;
if (!annNode || !pathNode || !methodNode) continue;
const httpMethod = METHOD_ANNOTATION_TO_HTTP[annNode.text];
if (!httpMethod) continue;
const rawPath = unquoteLiteral(pathNode.text);
if (rawPath === null) continue;
const enclosingClass = findEnclosingClass(methodNode);
const prefix = enclosingClass ? (prefixByClassId.get(enclosingClass.id) ?? '') : '';
const fullPath = joinPath(prefix, rawPath);
out.push({
role: 'provider',
framework: 'spring',
method: httpMethod,
path: fullPath,
name: nameNode?.text ?? null,
confidence: 0.8,
});
}
// ─── Consumers: RestTemplate ────────────────────────────────────
for (const match of runCompiledPatterns(REST_TEMPLATE_PATTERNS, tree)) {
const methodNode = match.captures.method;
const pathNode = match.captures.path;
if (!methodNode || !pathNode) continue;
const httpMethod = REST_TEMPLATE_TO_HTTP[methodNode.text];
if (!httpMethod) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'spring-rest-template',
method: httpMethod,
path,
name: null,
confidence: 0.7,
});
}
// ─── Consumers: WebClient.method(HttpMethod.X, "path") ──────────
for (const match of runCompiledPatterns(WEB_CLIENT_PATTERNS, tree)) {
const httpMethodNode = match.captures.http_method;
const pathNode = match.captures.path;
if (!httpMethodNode || !pathNode) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'spring-web-client',
method: httpMethodNode.text.toUpperCase(),
path,
name: null,
confidence: 0.7,
});
}
// ─── Consumers: OkHttp Request.Builder().url("path") ────────────
for (const match of runCompiledPatterns(OK_HTTP_PATTERNS, tree)) {
const pathNode = match.captures.path;
if (!pathNode) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'okhttp',
method: 'GET',
path,
name: null,
confidence: 0.7,
});
}
return out;
},
};
@@ -0,0 +1,373 @@
import Parser from 'tree-sitter';
import JavaScript from 'tree-sitter-javascript';
import TypeScript from 'tree-sitter-typescript';
import {
compilePatterns,
runCompiledPatterns,
unquoteLiteral,
type CompiledPatterns,
type LanguagePatterns,
type PatternSpec,
} from '../tree-sitter-scanner.js';
import type { HttpDetection, HttpLanguagePlugin } from './types.js';
/**
* Node.js / TypeScript HTTP plugin family. Handles:
* - NestJS `@Controller('prefix')` classes with `@Get(':id')` methods
* - Express `router.get(...)` / `app.post(...)` providers
* - `fetch(url)` / `fetch(url, { method: 'POST' })` consumers
* - `axios.get(url)` / `axios.delete(url)` consumers
*
* Because the JavaScript and TypeScript tree-sitter grammars share
* node type names for every construct we query, pattern sources are
* defined once and compiled against each grammar variant. The plugin
* exports three `HttpLanguagePlugin`s (JS, TS, TSX) that share the
* same `scan` function but bind to different grammars.
*/
// ─── Provider: NestJS — class-level @Controller('prefix') ────────────
// In tree-sitter-typescript decorators are NOT children of
// class_declaration / method_definition — they're siblings in the
// surrounding class_body / program node. We therefore match the
// decorator standalone and walk to its related class/method in JS.
const NEST_CONTROLLER_SPEC: PatternSpec<Record<string, never>> = {
meta: {},
query: `
(decorator
(call_expression
function: (identifier) @dec (#eq? @dec "Controller")
arguments: (arguments . [(string) (template_string)] @prefix))) @ctrl_decorator
`,
};
// ─── Provider: NestJS — method-level @Get/@Post/... decorators ───────
// Matches either `@Get('path')` or `@Get()`. The `@path` capture is
// optional — when the first argument isn't a string, the plugin falls
// back to '/' for the method-level path.
const NEST_METHOD_SPEC: PatternSpec<Record<string, never>> = {
meta: {},
query: `
(decorator
(call_expression
function: (identifier) @dec (#match? @dec "^(Get|Post|Put|Delete|Patch)$")
arguments: (arguments) @args)) @method_decorator
`,
};
// ─── Provider: Express — router.get/app.post/... ─────────────────────
const EXPRESS_SPEC: PatternSpec<Record<string, never>> = {
meta: {},
query: `
(call_expression
function: (member_expression
object: (identifier) @obj (#match? @obj "^(router|app)$")
property: (property_identifier) @http_method (#match? @http_method "^(get|post|put|delete|patch)$"))
arguments: (arguments . [(string) (template_string)] @path))
`,
};
// ─── Consumer: fetch(url) with NO options ─────────────────────────────
const FETCH_NO_OPTIONS_SPEC: PatternSpec<Record<string, never>> = {
meta: {},
query: `
(call_expression
function: (identifier) @fn (#eq? @fn "fetch")
arguments: (arguments . [(string) (template_string)] @path .))
`,
};
// ─── Consumer: fetch(url, { method: 'X', ... }) ──────────────────────
const FETCH_WITH_OPTIONS_SPEC: PatternSpec<Record<string, never>> = {
meta: {},
query: `
(call_expression
function: (identifier) @fn (#eq? @fn "fetch")
arguments: (arguments
. [(string) (template_string)] @path
(object
(pair
key: (property_identifier) @key (#eq? @key "method")
value: (string) @http_method))))
`,
};
// ─── Consumer: axios.get/post/... ────────────────────────────────────
const AXIOS_SPEC: PatternSpec<Record<string, never>> = {
meta: {},
query: `
(call_expression
function: (member_expression
object: (identifier) @obj (#eq? @obj "axios")
property: (property_identifier) @http_method (#match? @http_method "^(get|post|put|delete|patch)$"))
arguments: (arguments . [(string) (template_string)] @path))
`,
};
interface NodePatternBundle {
controller: CompiledPatterns<Record<string, never>>;
methodDecorator: CompiledPatterns<Record<string, never>>;
express: CompiledPatterns<Record<string, never>>;
fetchNoOptions: CompiledPatterns<Record<string, never>>;
fetchWithOptions: CompiledPatterns<Record<string, never>>;
axios: CompiledPatterns<Record<string, never>>;
}
function compileBundle(language: unknown, name: string): NodePatternBundle {
const mk = (spec: PatternSpec<Record<string, never>>, suffix: string) =>
compilePatterns({
name: `${name}-${suffix}`,
language,
patterns: [spec],
} satisfies LanguagePatterns<Record<string, never>>);
return {
controller: mk(NEST_CONTROLLER_SPEC, 'nest-controller'),
methodDecorator: mk(NEST_METHOD_SPEC, 'nest-method-decorator'),
express: mk(EXPRESS_SPEC, 'express'),
fetchNoOptions: mk(FETCH_NO_OPTIONS_SPEC, 'fetch-no-options'),
fetchWithOptions: mk(FETCH_WITH_OPTIONS_SPEC, 'fetch-with-options'),
axios: mk(AXIOS_SPEC, 'axios'),
};
}
const JAVASCRIPT_BUNDLE = compileBundle(JavaScript, 'javascript-http');
const TYPESCRIPT_BUNDLE = compileBundle(TypeScript.typescript, 'typescript-http');
const TSX_BUNDLE = compileBundle(TypeScript.tsx, 'tsx-http');
const NEST_DECORATOR_TO_HTTP: Record<string, string> = {
Get: 'GET',
Post: 'POST',
Put: 'PUT',
Delete: 'DELETE',
Patch: 'PATCH',
};
/**
* Find the nearest enclosing class_declaration for a node, or null.
*/
function findEnclosingClass(node: Parser.SyntaxNode): Parser.SyntaxNode | null {
let cur: Parser.SyntaxNode | null = node.parent;
while (cur) {
if (cur.type === 'class_declaration') return cur;
cur = cur.parent;
}
return null;
}
function joinPath(prefix: string, sub: string): string {
const cleanPrefix = prefix.replace(/^\/+/, '').replace(/\/+$/, '');
const cleanSub = sub.replace(/^\/+/, '');
if (!cleanPrefix) return `/${cleanSub}`;
return `/${cleanPrefix}/${cleanSub}`;
}
/**
* For a standalone `decorator` node (child of class_body / program),
* find the related `class_declaration` node that it decorates. In
* tree-sitter-typescript the decorator is placed before the class
* declaration as a sibling (when decorating a class) or inside the
* class_body before a method_definition (when decorating a method);
* we walk the parent chain until we find the enclosing class.
*/
function findDecoratedClass(decoratorNode: Parser.SyntaxNode): Parser.SyntaxNode | null {
const parent = decoratorNode.parent;
if (!parent) return null;
// Case 1: decorator is a sibling of the class_declaration at program /
// export_statement level. Walk forward through siblings until we find
// the class_declaration this decorator belongs to.
for (let i = 0; i < parent.namedChildCount; i++) {
const child = parent.namedChild(i);
if (child && child.id === decoratorNode.id) {
for (let j = i + 1; j < parent.namedChildCount; j++) {
const next = parent.namedChild(j);
if (!next) continue;
if (next.type === 'decorator') continue; // adjacent decorators stack
if (next.type === 'class_declaration') return next;
if (next.type === 'export_statement') {
// `export class Foo { ... }` wraps the declaration.
for (let k = 0; k < next.namedChildCount; k++) {
const inner = next.namedChild(k);
if (inner?.type === 'class_declaration') return inner;
}
}
break;
}
break;
}
}
// Case 2: decorator is inside a class_body (decorating a method) —
// walk up to the enclosing class_declaration.
return findEnclosingClass(decoratorNode);
}
/**
* For a method-level decorator node (child of class_body before a
* method_definition), find the method_definition it decorates.
*/
function findDecoratedMethod(decoratorNode: Parser.SyntaxNode): Parser.SyntaxNode | null {
const parent = decoratorNode.parent;
if (!parent || parent.type !== 'class_body') return null;
for (let i = 0; i < parent.namedChildCount; i++) {
const child = parent.namedChild(i);
if (child && child.id === decoratorNode.id) {
for (let j = i + 1; j < parent.namedChildCount; j++) {
const next = parent.namedChild(j);
if (!next) continue;
if (next.type === 'decorator') continue;
if (next.type === 'method_definition') return next;
return null;
}
return null;
}
}
return null;
}
function scanBundle(bundle: NodePatternBundle, tree: Parser.Tree): HttpDetection[] {
const out: HttpDetection[] = [];
// NestJS: collect `@Controller('prefix')` class decorators, keyed by
// the `class_declaration` they decorate.
const prefixByClassId = new Map<number, string>();
for (const match of runCompiledPatterns(bundle.controller, tree)) {
const prefixNode = match.captures.prefix;
const decoratorNode = match.captures.ctrl_decorator;
if (!prefixNode || !decoratorNode) continue;
const prefix = unquoteLiteral(prefixNode.text);
if (prefix === null) continue;
const classNode = findDecoratedClass(decoratorNode);
if (!classNode) continue;
prefixByClassId.set(classNode.id, prefix);
}
// NestJS: method-level @Get/@Post/... decorators. The decorator's
// arguments list may be empty (`@Get()`), a string (`@Get('path')`),
// or something else (which we skip).
for (const match of runCompiledPatterns(bundle.methodDecorator, tree)) {
const decNode = match.captures.dec;
const argsNode = match.captures.args;
const decoratorNode = match.captures.method_decorator;
if (!decNode || !argsNode || !decoratorNode) continue;
const httpMethod = NEST_DECORATOR_TO_HTTP[decNode.text];
if (!httpMethod) continue;
const methodNode = findDecoratedMethod(decoratorNode);
if (!methodNode) continue;
const enclosingClass = findEnclosingClass(methodNode);
// Only emit NestJS detections when the class actually has a
// @Controller decorator — without it, the match is almost certainly
// something else (e.g. an unrelated library using similar names).
if (!enclosingClass || !prefixByClassId.has(enclosingClass.id)) continue;
const prefix = prefixByClassId.get(enclosingClass.id) ?? '';
let rawPath = '/';
const firstArg = argsNode.namedChild(0);
if (firstArg && (firstArg.type === 'string' || firstArg.type === 'template_string')) {
const unquoted = unquoteLiteral(firstArg.text);
if (unquoted !== null) rawPath = unquoted;
}
// Get the method name from the decorated method_definition.
const methodNameNode = methodNode.childForFieldName('name');
const name = methodNameNode?.text ?? null;
out.push({
role: 'provider',
framework: 'nest',
method: httpMethod,
path: joinPath(prefix, rawPath),
name,
confidence: 0.8,
});
}
// Express: router/app.<verb>(...)
for (const match of runCompiledPatterns(bundle.express, tree)) {
const methodNode = match.captures.http_method;
const pathNode = match.captures.path;
if (!methodNode || !pathNode) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'provider',
framework: 'express',
method: methodNode.text.toUpperCase(),
path,
name: 'handler',
confidence: 0.8,
});
}
// Consumer: fetch with options { method: 'X' }
const fetchSeen = new Set<number>();
for (const match of runCompiledPatterns(bundle.fetchWithOptions, tree)) {
const pathNode = match.captures.path;
const methodNode = match.captures.http_method;
if (!pathNode || !methodNode) continue;
const path = unquoteLiteral(pathNode.text);
const method = unquoteLiteral(methodNode.text);
if (path === null || method === null) continue;
fetchSeen.add(pathNode.id);
out.push({
role: 'consumer',
framework: 'fetch',
method: method.toUpperCase(),
path,
name: null,
confidence: 0.7,
});
}
// Consumer: plain fetch(path) — default GET. Skip path nodes we already
// matched with the options variant so we don't double-emit.
for (const match of runCompiledPatterns(bundle.fetchNoOptions, tree)) {
const pathNode = match.captures.path;
if (!pathNode) continue;
if (fetchSeen.has(pathNode.id)) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'fetch',
method: 'GET',
path,
name: null,
confidence: 0.7,
});
}
// Consumer: axios.<verb>(url)
for (const match of runCompiledPatterns(bundle.axios, tree)) {
const methodNode = match.captures.http_method;
const pathNode = match.captures.path;
if (!methodNode || !pathNode) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'axios',
method: methodNode.text.toUpperCase(),
path,
name: null,
confidence: 0.7,
});
}
return out;
}
export const JAVASCRIPT_HTTP_PLUGIN: HttpLanguagePlugin = {
name: 'javascript-http',
language: JavaScript,
scan: (tree) => scanBundle(JAVASCRIPT_BUNDLE, tree),
};
export const TYPESCRIPT_HTTP_PLUGIN: HttpLanguagePlugin = {
name: 'typescript-http',
language: TypeScript.typescript,
scan: (tree) => scanBundle(TYPESCRIPT_BUNDLE, tree),
};
export const TSX_HTTP_PLUGIN: HttpLanguagePlugin = {
name: 'tsx-http',
language: TypeScript.tsx,
scan: (tree) => scanBundle(TSX_BUNDLE, tree),
};
@@ -0,0 +1,79 @@
import PHP from 'tree-sitter-php';
import {
compilePatterns,
runCompiledPatterns,
unquoteLiteral,
type LanguagePatterns,
} from '../tree-sitter-scanner.js';
import type { HttpDetection, HttpLanguagePlugin } from './types.js';
/**
* PHP HTTP plugin — Laravel `Route::get/post/...` declarations.
*
* The pipeline already uses `PHP.php_only` for ingesting plain `.php`
* files (see `core/tree-sitter/parser-loader.ts`), and we do the same
* here so Laravel route files are parsed with the right grammar dialect.
*/
const LARAVEL_PATTERNS = compilePatterns({
name: 'php-laravel',
language: PHP.php_only,
patterns: [
{
meta: {},
query: `
(scoped_call_expression
scope: (name) @scope (#eq? @scope "Route")
name: (name) @method (#match? @method "^(get|post|put|delete|patch)$")
arguments: (arguments . (argument (string) @path)))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
/**
* Extract the inner text of a PHP `string` node. The tree-sitter-php
* grammar wraps single / double-quoted literals differently depending
* on content; we try both the raw `text` (with quotes) through
* `unquoteLiteral`, and a fallback via the `string_value` / `string_content`
* child nodes.
*/
function phpStringText(node: import('tree-sitter').SyntaxNode): string | null {
// Most single-quoted strings expose their inner content through the
// full node text (including quotes), which unquoteLiteral strips.
const direct = unquoteLiteral(node.text);
if (direct !== null && direct !== node.text) return direct;
// Fall back to child string_content / string_value node if present.
for (const child of node.children) {
if (child.type === 'string_content' || child.type === 'string_value') {
return child.text;
}
}
return direct;
}
export const PHP_HTTP_PLUGIN: HttpLanguagePlugin = {
name: 'php-http',
language: PHP.php_only,
scan(tree) {
const out: HttpDetection[] = [];
for (const match of runCompiledPatterns(LARAVEL_PATTERNS, tree)) {
const methodNode = match.captures.method;
const pathNode = match.captures.path;
if (!methodNode || !pathNode) continue;
const path = phpStringText(pathNode);
if (path === null) continue;
out.push({
role: 'provider',
framework: 'laravel',
method: methodNode.text.toUpperCase(),
path,
name: 'route',
confidence: 0.8,
});
}
return out;
},
};
@@ -0,0 +1,142 @@
import Python from 'tree-sitter-python';
import {
compilePatterns,
runCompiledPatterns,
unquoteLiteral,
type LanguagePatterns,
} from '../tree-sitter-scanner.js';
import type { HttpDetection, HttpLanguagePlugin } from './types.js';
/**
* Python HTTP plugin. Handles:
* - FastAPI `@app.get("/path")` provider decorators
* - `requests.get/post/...("url")` consumer calls
* - Generic `requests.request("METHOD", "url")` consumer calls
*/
const FASTAPI_VERBS: Record<string, string> = {
get: 'GET',
post: 'POST',
put: 'PUT',
delete: 'DELETE',
patch: 'PATCH',
};
// ─── Provider: FastAPI @app.get/... ──────────────────────────────────
const FASTAPI_PATTERNS = compilePatterns({
name: 'python-fastapi',
language: Python,
patterns: [
{
meta: {},
query: `
(decorator
(call
function: (attribute
object: (identifier) @obj (#eq? @obj "app")
attribute: (identifier) @method (#match? @method "^(get|post|put|delete|patch)$"))
arguments: (argument_list . (string) @path)))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: requests.get/post/... ──────────────────────────────────
const REQUESTS_VERB_PATTERNS = compilePatterns({
name: 'python-requests-verb',
language: Python,
patterns: [
{
meta: {},
query: `
(call
function: (attribute
object: (identifier) @obj (#eq? @obj "requests")
attribute: (identifier) @method (#match? @method "^(get|post|put|delete|patch)$"))
arguments: (argument_list . (string) @path))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: requests.request("METHOD", "url") ─────────────────────
const REQUESTS_GENERIC_PATTERNS = compilePatterns({
name: 'python-requests-generic',
language: Python,
patterns: [
{
meta: {},
query: `
(call
function: (attribute
object: (identifier) @obj (#eq? @obj "requests")
attribute: (identifier) @method (#eq? @method "request"))
arguments: (argument_list . (string) @http_method (string) @path))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
name: 'python-http',
language: Python,
scan(tree) {
const out: HttpDetection[] = [];
// Providers: FastAPI
for (const match of runCompiledPatterns(FASTAPI_PATTERNS, tree)) {
const methodNode = match.captures.method;
const pathNode = match.captures.path;
if (!methodNode || !pathNode) continue;
const httpMethod = FASTAPI_VERBS[methodNode.text];
if (!httpMethod) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'provider',
framework: 'fastapi',
method: httpMethod,
path,
name: null,
confidence: 0.8,
});
}
// Consumers: requests.<verb>
for (const match of runCompiledPatterns(REQUESTS_VERB_PATTERNS, tree)) {
const methodNode = match.captures.method;
const pathNode = match.captures.path;
if (!methodNode || !pathNode) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'python-requests',
method: methodNode.text.toUpperCase(),
path,
name: null,
confidence: 0.7,
});
}
// Consumers: requests.request("METHOD", "url")
for (const match of runCompiledPatterns(REQUESTS_GENERIC_PATTERNS, tree)) {
const methodNode = match.captures.http_method;
const pathNode = match.captures.path;
if (!methodNode || !pathNode) continue;
const methodRaw = unquoteLiteral(methodNode.text);
const path = unquoteLiteral(pathNode.text);
if (methodRaw === null || path === null) continue;
out.push({
role: 'consumer',
framework: 'python-requests',
method: methodRaw.toUpperCase(),
path,
name: null,
confidence: 0.7,
});
}
return out;
},
};
@@ -0,0 +1,65 @@
import type Parser from 'tree-sitter';
/**
* Shared types for the http-route-extractor language plugins.
*
* Each plugin lives in its own file (java.ts, node.ts, ...) and owns
* the tree-sitter grammar import + queries. The top-level
* `http-route-extractor.ts` orchestrator only knows about this type
* module and the plugin registry (`./index.ts`). It MUST NOT import
* any grammar or query text directly — language-specific knowledge
* belongs in the plugins.
*/
export type HttpRole = 'provider' | 'consumer';
/**
* One raw HTTP detection produced by a plugin's `scan()` function. The
* orchestrator converts this into a full `ExtractedContract` by running
* path normalization and building the contract id.
*
* `path` is the raw literal string as it appeared in source (with
* `${...}` template placeholders still in place); the orchestrator
* runs the appropriate normalizer for provider vs. consumer paths.
*/
export interface HttpDetection {
role: HttpRole;
/** Short framework label, e.g. `'spring'`, `'nest'`, `'express'`. */
framework: string;
/** HTTP method in upper case (`'GET'`, `'POST'`, ...). */
method: string;
/** Raw path literal as seen in source (template placeholders intact). */
path: string;
/**
* Symbol name of the handler (for providers) or calling function
* (for consumers) when the plugin can determine it structurally.
* Null when no good candidate is available.
*/
name: string | null;
/** Confidence in (0, 1]. Source-scan plugins typically use 0.7–0.8. */
confidence: number;
}
/**
* One language-scoped HTTP plugin. The plugin owns the tree-sitter
* grammar and the `scan` function that translates a parsed tree into
* zero or more `HttpDetection`s. Plugins are free to run multiple
* compiled pattern bundles internally (see the shared scanner's
* `runCompiledPatterns` helper).
*
* `language` is typed as `unknown` for the same reason as
* `LanguagePatterns.language` in `tree-sitter-scanner.ts` — the
* grammar modules export different shapes.
*/
export interface HttpLanguagePlugin {
/** Human-readable plugin name for diagnostics. */
name: string;
/** tree-sitter grammar object (passed to the shared parser). */
language: unknown;
/**
* Scan a parsed tree and return zero or more HTTP detections. Plugins
* must not throw — they should swallow per-match errors so a single
* malformed construct does not abort the whole file.
*/
scan(tree: Parser.Tree): HttpDetection[];
}
@@ -1,8 +1,34 @@
import * as fs from 'node:fs';
import * as path from 'node:path';
import { glob } from 'glob';
import Parser from 'tree-sitter';
import type { ContractExtractor, CypherExecutor } from '../contract-extractor.js';
import type { ExtractedContract, RepoHandle } from '../types.js';
import { readSafe } from './fs-utils.js';
import { getPluginForFile, HTTP_SCAN_GLOB, type HttpDetection } from './http-patterns/index.js';
/**
* Language-agnostic orchestrator for HTTP route (provider + consumer)
* contract extraction. Two strategies, in order of preference per role:
*
* 1. **Graph-assisted (Strategy A)** — if a per-repo LadybugDB executor
* is available, read `HANDLES_ROUTE` / `FETCHES` Cypher edges that
* the ingestion pipeline already produced via tree-sitter. This is
* the preferred path because the graph has richer symbol metadata
* (real uids, class/method structure, etc.).
*
* 2. **Source-scan fallback (Strategy B)** — parse files directly with
* the per-language plugin registry in `./http-patterns/`. Used when
* the graph has no routes/fetches for this repo (e.g. a repo that
* hasn't been indexed yet, or whose indexer doesn't know the
* framework). Each plugin owns its tree-sitter grammar and query
* sources — this orchestrator imports NO grammars or query strings.
*
* Adding a new language for Strategy B is a one-file edit in
* `http-patterns/index.ts`: register a new `HttpLanguagePlugin` and
* widen `HTTP_SCAN_GLOB` if needed.
*/
// ─── Graph-assisted queries ──────────────────────────────────────────
const HANDLES_ROUTE_QUERY = `
MATCH (handlerFile:File)-[r:CodeRelation {type: 'HANDLES_ROUTE'}]->(route:Route)
@@ -23,14 +49,56 @@ WHERE sym.startLine IS NOT NULL
RETURN sym.id AS uid, sym.name AS name, sym.filePath AS filePath, labels(sym) AS labels
ORDER BY sym.startLine`;
// ─── Path normalization (shared between provider / consumer paths) ──
/**
* Canonicalize a provider-side HTTP path for contract-id generation:
* - strip query string
* - lower-case
* - drop trailing slash
* - collapse `:id`, `{id}`, `[id]` path params into a single `{param}`
*/
export function normalizeHttpPath(p: string): string {
let s = p.trim().split('?')[0].toLowerCase().replace(/\/+$/, '');
s = s.replace(/:\w+/g, '{param}');
s = s.replace(/\{[^}]+\}/g, '{param}');
s = s.replace(/\[[^\]]+\]/g, '{param}');
return s;
// Preserve root: after stripping trailing slashes, the root "/"
// collapses to "" which would produce malformed contract ids like
// `http::GET::`. Restore a single slash for the root case.
return s === '' ? '/' : s;
}
/**
* Consumer-side normalization is more aggressive:
* - template literals (`${x}`) → `{param}`
* - strip protocol + host if the URL is absolute
* - numeric segments → `{param}` (so `/api/orders/42` → `/api/orders/{param}`)
*/
function normalizeConsumerPath(url: string): string {
const templated = url.replace(/\$\{[^}]+\}/g, '{param}').trim();
let pathOnly = templated;
if (/^https?:\/\//i.test(templated)) {
try {
pathOnly = new URL(templated).pathname;
} catch {
pathOnly = templated.replace(/^https?:\/\/[^/]+/i, '');
}
}
const normalized = normalizeHttpPath(pathOnly || '/');
const segments = normalized
.split('/')
.filter(Boolean)
.map((segment) => (/^\d+$/.test(segment) ? '{param}' : segment));
return `/${segments.join('/')}`.replace(/\/+$/, '') || '/';
}
function contractIdFor(method: string, pathNorm: string): string {
return `http::${method.toUpperCase()}::${pathNorm}`;
}
// ─── Graph row helpers ───────────────────────────────────────────────
function methodFromRouteReason(reason: string): string | null {
const r = reason || '';
if (/GetMapping|decorator-Get/i.test(r)) return 'GET';
@@ -41,50 +109,6 @@ function methodFromRouteReason(reason: string): string | null {
return null;
}
function contractIdFor(method: string, pathNorm: string): string {
return `http::${method.toUpperCase()}::${pathNorm}`;
}
function readSafe(repoPath: string, rel: string): string | null {
const abs = path.resolve(repoPath, rel);
const base = path.resolve(repoPath);
const relToBase = path.relative(base, abs);
if (relToBase.startsWith('..') || path.isAbsolute(relToBase)) return null;
try {
return fs.readFileSync(abs, 'utf-8');
} catch {
return null;
}
}
function pickJavaHandlerName(
content: string,
routePath: string,
httpMethod: string,
): string | null {
const tail = routePath.split('/').filter(Boolean).pop() || '';
const mapNames: Record<string, string> = {
GET: 'GetMapping',
POST: 'PostMapping',
PUT: 'PutMapping',
DELETE: 'DeleteMapping',
PATCH: 'PatchMapping',
};
const ann = mapNames[httpMethod] || 'GetMapping';
const lines = content.split(/\r?\n/);
for (let i = 0; i < lines.length; i++) {
const line = lines[i];
if (!line.includes(`@${ann}`)) continue;
if (!line.includes(`"${tail}"`) && !line.includes(`'${tail}'`) && tail && !line.includes(tail))
continue;
for (let j = i + 1; j < Math.min(i + 8, lines.length); j++) {
const m = lines[j].match(/(?:public|protected|private)\s+[\w<>,\s\[\]]+\s+(\w+)\s*\(/);
if (m) return m[1];
}
}
return null;
}
function pickSymbolUid(
rows: Record<string, unknown>[],
preferredName: string | null,
@@ -114,6 +138,8 @@ function pickSymbolUid(
};
}
// ─── Orchestrator ────────────────────────────────────────────────────
export class HttpRouteExtractor implements ContractExtractor {
type = 'http' as const;
@@ -124,20 +150,76 @@ export class HttpRouteExtractor implements ContractExtractor {
async extract(
dbExecutor: CypherExecutor | null,
repoPath: string,
repo: RepoHandle,
_repo: RepoHandle,
): Promise<ExtractedContract[]> {
const graphP = dbExecutor != null ? await this.extractProvidersGraph(dbExecutor, repoPath) : [];
const providers = graphP.length > 0 ? graphP : await this.extractProvidersSourceScan(repoPath);
// Parse each file at most once and reuse the plugin results across
// both graph-assisted enrichment and source-scan emission.
const parser = new Parser();
const cachedDetections = new Map<string, HttpDetection[]>();
const getDetections = (rel: string): HttpDetection[] => {
const cached = cachedDetections.get(rel);
if (cached) return cached;
const plugin = getPluginForFile(rel);
if (!plugin) {
cachedDetections.set(rel, []);
return [];
}
const content = readSafe(repoPath, rel);
if (!content) {
cachedDetections.set(rel, []);
return [];
}
try {
parser.setLanguage(plugin.language);
const tree = parser.parse(content);
const detections = plugin.scan(tree);
cachedDetections.set(rel, detections);
return detections;
} catch {
cachedDetections.set(rel, []);
return [];
}
};
const graphC = dbExecutor != null ? await this.extractConsumersGraph(dbExecutor, repoPath) : [];
const consumers = graphC.length > 0 ? graphC : await this.extractConsumersSourceScan(repoPath);
// Glob the source-scan file list at most once per extract() —
// both provider and consumer fallback paths share the same list.
let scannedFiles: string[] | null = null;
const getScannedFiles = async (): Promise<string[]> => {
if (scannedFiles) return scannedFiles;
scannedFiles = await this.scanFiles(repoPath);
return scannedFiles;
};
const graphProviders =
dbExecutor != null ? await this.extractProvidersGraph(dbExecutor, getDetections) : [];
const providers =
graphProviders.length > 0
? graphProviders
: this.extractProvidersSourceScan(await getScannedFiles(), getDetections);
const graphConsumers =
dbExecutor != null ? await this.extractConsumersGraph(dbExecutor, getDetections) : [];
const consumers =
graphConsumers.length > 0
? graphConsumers
: this.extractConsumersSourceScan(await getScannedFiles(), getDetections);
return [...providers, ...consumers];
}
private async scanFiles(repoPath: string): Promise<string[]> {
return glob(HTTP_SCAN_GLOB, {
cwd: repoPath,
ignore: ['**/node_modules/**', '**/.git/**', '**/dist/**', '**/build/**', '**/vendor/**'],
nodir: true,
});
}
// ─── Graph-assisted providers ──────────────────────────────────────
private async extractProvidersGraph(
db: CypherExecutor,
repoPath: string,
getDetections: (rel: string) => HttpDetection[],
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
let rows: Record<string, unknown>[];
@@ -152,22 +234,55 @@ export class HttpRouteExtractor implements ContractExtractor {
const routePath = String(row.routePath ?? '');
const routeSource = String(row.routeSource ?? row.routeReason ?? '');
let method = methodFromRouteReason(routeSource);
const content = readSafe(repoPath, filePath);
if (!method && content) {
method = this.inferMethodFromFileScan(content, routePath, 'provider');
// Look up handler name (and backfill method if missing) from the
// plugin's scan of the handler file. This replaces the old
// regex-based `inferMethodFromFileScan` and `pickJavaHandlerName`
// helpers — tree-sitter gives both pieces of information
// structurally. Always run the lookup: even when method is set by
// `methodFromRouteReason`, we still need the handler name.
const detections = filePath ? getDetections(filePath) : [];
const providerDetections = detections.filter((d) => d.role === 'provider');
let handlerName: string | null = null;
const normalizedRoute = normalizeHttpPath(routePath);
// Candidates share the same normalized path. When multiple
// detections at the same path exist (e.g. GET + POST /api/orders
// in one router), a blind `.find()` silently returned the first
// verb — attaching the wrong handler and, when method was not
// already pinned by the route reason, the wrong method too.
// Disambiguate by method when we know it; refuse to guess when
// we don't.
const candidates = providerDetections.filter(
(d) => normalizeHttpPath(d.path) === normalizedRoute,
);
let match: (typeof candidates)[number] | undefined;
const ambiguousCandidates = !method && candidates.length > 1;
if (method) {
match = candidates.find((d) => d.method === method);
} else if (candidates.length === 1) {
match = candidates[0];
}
// else: multiple candidates + unknown method → leave match
// undefined so handlerName stays null and skip symbol
// enrichment below, keeping the file-basename fallback instead
// of letting pickSymbolUid silently pick the first Function /
// Method in the file (which reintroduces the mis-attribution
// we were trying to avoid). Method stays at the conservative
// 'GET' default set below.
if (match) {
if (!method) method = match.method;
handlerName = match.name;
}
if (!method) method = 'GET';
const pathNorm = normalizeHttpPath(routePath);
const cid = contractIdFor(method, pathNorm);
const handlerName =
content && routePath ? pickJavaHandlerName(content, routePath, method) : null;
let symbolUid = '';
let symbolName = path.basename(filePath) || 'handler';
let symPath = filePath;
const fileId = row.fileId ?? row[0];
if (fileId) {
if (fileId && !ambiguousCandidates) {
try {
const syms = await db(CONTAINS_QUERY, { fileId });
if (syms.length > 0) {
@@ -201,145 +316,44 @@ export class HttpRouteExtractor implements ContractExtractor {
return out;
}
private inferMethodFromFileScan(
content: string,
routePath: string,
_role: string,
): string | null {
const tail = routePath.split('/').filter(Boolean).pop() || '';
for (const m of ['GET', 'POST', 'PUT', 'DELETE', 'PATCH'] as const) {
const mapNames: Record<string, string> = {
GET: 'GetMapping',
POST: 'PostMapping',
PUT: 'PutMapping',
DELETE: 'DeleteMapping',
PATCH: 'PatchMapping',
};
if (
content.includes(`@${mapNames[m]}`) &&
(content.includes(tail) || routePath.includes(tail))
) {
return m;
}
}
return null;
}
// ─── Source-scan providers ─────────────────────────────────────────
private async extractProvidersSourceScan(repoPath: string): Promise<ExtractedContract[]> {
const files = await glob('**/*.{ts,tsx,js,jsx,java,vue,svelte,php,py}', {
cwd: repoPath,
ignore: ['**/node_modules/**', '**/.git/**', '**/dist/**', '**/build/**'],
nodir: true,
});
private extractProvidersSourceScan(
files: string[],
getDetections: (rel: string) => HttpDetection[],
): ExtractedContract[] {
const out: ExtractedContract[] = [];
for (const rel of files) {
const content = readSafe(repoPath, rel);
if (!content) continue;
out.push(...this.scanSpringProviders(content, rel));
out.push(...this.scanExpressProviders(content, rel));
out.push(...this.scanLaravelProviders(content, rel));
out.push(...this.scanFastApiProviders(content, rel));
const detections = getDetections(rel);
for (const d of detections) {
if (d.role !== 'provider') continue;
const pathNorm = normalizeHttpPath(d.path);
out.push({
contractId: contractIdFor(d.method, pathNorm),
type: 'http',
role: 'provider',
symbolUid: '',
symbolRef: { filePath: rel, name: d.name ?? 'handler' },
symbolName: d.name ?? 'handler',
confidence: d.confidence,
meta: {
method: d.method,
path: pathNorm,
pathSegments: pathNorm.split('/').filter(Boolean),
extractionStrategy: 'source_scan',
framework: d.framework,
},
});
}
}
return this.dedupeContracts(out);
}
private dedupeContracts(items: ExtractedContract[]): ExtractedContract[] {
const seen = new Set<string>();
const out: ExtractedContract[] = [];
for (const c of items) {
const k = `${c.contractId}|${c.symbolRef.filePath}|${c.symbolRef.name}`;
if (seen.has(k)) continue;
seen.add(k);
out.push(c);
}
return out;
}
private scanSpringProviders(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
let classPrefix = '';
const classRm = content.match(/@RequestMapping\s*\(\s*"([^"]+)"/);
if (classRm) classPrefix = classRm[1].replace(/\/+$/, '');
const re = /@(Get|Post|Put|Delete|Patch)Mapping\s*\(\s*"([^"]+)"/gi;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const method = m[1].toUpperCase();
let p = m[2];
if (classPrefix) p = `${classPrefix}/${p.replace(/^\//, '')}`;
const pathNorm = normalizeHttpPath(p);
const sub = content.slice(m.index);
const nameM = sub.match(/(?:public|protected|private)\s+[\w<>,\s\[\]]+\s+(\w+)\s*\(/);
const name = nameM ? nameM[1] : m[0];
out.push(this.makeProvider(filePath, method, pathNorm, name, 0.8));
}
return out;
}
private scanExpressProviders(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
const re = /(?:router|app)\.(get|post|put|delete|patch)\s*\(\s*['"]([^'"]+)['"]/gi;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const method = m[1].toUpperCase();
const pathNorm = normalizeHttpPath(m[2]);
out.push(this.makeProvider(filePath, method, pathNorm, 'handler', 0.8));
}
return out;
}
private scanLaravelProviders(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
const re = /Route::(get|post|put|delete|patch)\s*\(\s*['"]([^'"]+)['"]/gi;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const method = m[1].toUpperCase();
const pathNorm = normalizeHttpPath(m[2]);
out.push(this.makeProvider(filePath, method, pathNorm, 'route', 0.8));
}
return out;
}
private scanFastApiProviders(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
const re = /@app\.(get|post|put|delete|patch)\s*\(\s*['"]([^'"]+)['"]/gi;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const method = m[1].toUpperCase();
const pathNorm = normalizeHttpPath(m[2]);
out.push(this.makeProvider(filePath, method, pathNorm, 'handler', 0.8));
}
return out;
}
private makeProvider(
filePath: string,
method: string,
pathNorm: string,
name: string,
confidence: number,
): ExtractedContract {
const cid = contractIdFor(method, pathNorm);
return {
contractId: cid,
type: 'http',
role: 'provider',
symbolUid: '',
symbolRef: { filePath, name },
symbolName: name,
confidence,
meta: {
method,
path: pathNorm,
pathSegments: pathNorm.split('/').filter(Boolean),
extractionStrategy: 'source_scan',
},
};
}
// ─── Graph-assisted consumers ──────────────────────────────────────
private async extractConsumersGraph(
db: CypherExecutor,
repoPath: string,
getDetections: (rel: string) => HttpDetection[],
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
let rows: Record<string, unknown>[];
@@ -353,11 +367,23 @@ export class HttpRouteExtractor implements ContractExtractor {
const routePath = String(row.routePath ?? '');
const pathNorm = normalizeHttpPath(routePath);
let method = 'GET';
const content = readSafe(repoPath, filePath);
if (content) {
const inferred = this.inferFetchMethod(content, pathNorm);
if (inferred) method = inferred;
// Prefer the plugin's detected method if we can find a matching
// fetch/axios call in the same file.
const detections = filePath ? getDetections(filePath) : [];
// Symmetric to the provider path: if multiple consumer calls in
// the same file share the same normalized path (e.g. a GET
// fetch AND a POST fetch to `/api/orders`), `.find()` silently
// picked the first verb and keyed the contract id on the wrong
// method. With no upstream method signal here, refuse to guess
// when candidates are ambiguous — leave `method` at its
// conservative 'GET' default.
const consumerCandidates = detections.filter(
(d) => d.role === 'consumer' && normalizeConsumerPath(d.path) === pathNorm,
);
if (consumerCandidates.length === 1) {
method = consumerCandidates[0].method;
}
const cid = contractIdFor(method, pathNorm);
let symbolUid = '';
let symbolName = 'fetch';
@@ -395,81 +421,47 @@ export class HttpRouteExtractor implements ContractExtractor {
return out;
}
private inferFetchMethod(content: string, pathNorm: string): string | null {
const esc = pathNorm.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
const fetchRe = new RegExp(
`fetch\\s*\\(\\s*['"\`]([^'"\`]*${esc}[^'"\`]*)['"\`]\\s*,\\s*\\{[^}]*method:\\s*['"](\\w+)['"]`,
'i',
);
const m = content.match(fetchRe);
if (m) return m[2].toUpperCase();
return null;
}
// ─── Source-scan consumers ─────────────────────────────────────────
private async extractConsumersSourceScan(repoPath: string): Promise<ExtractedContract[]> {
const files = await glob('**/*.{ts,tsx,js,jsx,vue,svelte}', {
cwd: repoPath,
ignore: ['**/node_modules/**', '**/.git/**'],
nodir: true,
});
private extractConsumersSourceScan(
files: string[],
getDetections: (rel: string) => HttpDetection[],
): ExtractedContract[] {
const out: ExtractedContract[] = [];
for (const rel of files) {
const content = readSafe(repoPath, rel);
if (!content) continue;
out.push(...this.scanFetchConsumers(content, rel));
out.push(...this.scanAxiosConsumers(content, rel));
const detections = getDetections(rel);
for (const d of detections) {
if (d.role !== 'consumer') continue;
const pathNorm = normalizeConsumerPath(d.path);
out.push({
contractId: contractIdFor(d.method, pathNorm),
type: 'http',
role: 'consumer',
symbolUid: '',
symbolRef: { filePath: rel, name: 'fetch' },
symbolName: 'fetch',
confidence: d.confidence,
meta: {
method: d.method,
path: pathNorm,
extractionStrategy: 'source_scan',
framework: d.framework,
},
});
}
}
return this.dedupeContracts(out);
}
private scanFetchConsumers(content: string, filePath: string): ExtractedContract[] {
private dedupeContracts(items: ExtractedContract[]): ExtractedContract[] {
const seen = new Set<string>();
const out: ExtractedContract[] = [];
const re =
/fetch\s*\(\s*['"`]([^'"`]+)['"`](?:\s*,\s*\{[^}]*method:\s*['"](\w+)['"][^}]*\})?\s*\)/gi;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const pathNorm = normalizeHttpPath(this.templateToPattern(m[1]));
const method = (m[2] || 'GET').toUpperCase();
out.push(this.makeConsumer(filePath, method, pathNorm, 0.7));
for (const c of items) {
const k = `${c.contractId}|${c.symbolRef.filePath}|${c.symbolRef.name}`;
if (seen.has(k)) continue;
seen.add(k);
out.push(c);
}
return out;
}
private templateToPattern(url: string): string {
return url.replace(/\$\{[^}]+\}/g, '{param}');
}
private scanAxiosConsumers(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
const re = /axios\.(get|post|put|delete|patch)\s*\(\s*[`'"]([^`'"]+)[`'"]/gi;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const method = m[1].toUpperCase();
const pathNorm = normalizeHttpPath(this.templateToPattern(m[2]));
out.push(this.makeConsumer(filePath, method, pathNorm, 0.7));
}
return out;
}
private makeConsumer(
filePath: string,
method: string,
pathNorm: string,
confidence: number,
): ExtractedContract {
return {
contractId: contractIdFor(method, pathNorm),
type: 'http',
role: 'consumer',
symbolUid: '',
symbolRef: { filePath, name: 'fetch' },
symbolName: 'fetch',
confidence,
meta: {
method,
path: pathNorm,
extractionStrategy: 'source_scan',
},
};
}
}
@@ -0,0 +1,311 @@
import type { ContractType, CrossLink, GroupManifestLink, StoredContract } from '../types.js';
import type { CypherExecutor } from '../contract-extractor.js';
export interface ManifestExtractResult {
contracts: StoredContract[];
crossLinks: CrossLink[];
}
/**
* Canonicalize an HTTP path for matching against Route.name in the graph.
* Mirrors core/ingestion/pipeline.ts ensureSlash semantics:
* - Ensures a leading slash.
* - Strips trailing slashes (except the root "/").
* - Normalizes consecutive slashes.
* - Does NOT lowercase (route matching is case-sensitive).
*/
function normalizeRoutePath(raw: string): string {
const trimmed = raw.trim();
if (!trimmed) return '/';
const withLeading = trimmed.startsWith('/') ? trimmed : `/${trimmed}`;
const collapsed = withLeading.replace(/\/+/g, '/');
if (collapsed === '/') return '/';
return collapsed.replace(/\/+$/, '');
}
/**
* Split a manifest HTTP contract into its optional `METHOD::` prefix and
* its path portion.
*
* `buildContractId` recommends the explicit-method form `GET::/api/orders`
* in group.yaml; if we hand that raw string to `normalizeRoutePath` we get
* `/GET::/api/orders`, which can never match `Route.name = "/api/orders"`
* in the graph. This helper extracts the path portion so the Cypher
* lookup uses the canonical route name.
*
* The method prefix regex mirrors `buildContractId` (line ~251) for
* symmetry: case-insensitive `[A-Za-z]+` followed by `::`. The captured
* method is upper-cased for downstream use; method-constrained matching
* against `HANDLES_ROUTE` is a future enhancement (not yet wired).
*
* Edge cases:
* - `"::/api/orders"` — empty method portion, no alpha prefix match, so
* the whole string is treated as a bare path (matches buildContractId
* which also requires `[A-Za-z]+`).
* - `"GET::"` — method with empty path, returns `{ method: 'GET', path: '' }`;
* `normalizeRoutePath('')` resolves to `/` for caller.
*/
function parseHttpContract(raw: string): { method: string | null; path: string } {
const match = raw.match(/^([A-Za-z]+)::/);
if (!match) return { method: null, path: raw };
return { method: match[1].toUpperCase(), path: raw.slice(match[0].length) };
}
/**
* Stable synthetic symbolUid for a manifest-declared contract whose target
* symbol could not be resolved against the per-repo graph (resolveSymbol
* returned null). Two reasons we don't leave the uid empty:
*
* 1. The bridge stores Contract nodes keyed in part by symbolUid; an empty
* uid means downstream Cypher queries that anchor on `provider.symbolUid`
* can't tell two different unresolved manifest contracts apart.
* 2. The cross-impact bridge query in cross-impact.ts joins local impact
* results to bridge contracts via `WHERE provider.symbolUid IN $localUids`.
* If the local impact engine produces a deterministic identifier for the
* unresolved target, it must agree with the value the bridge stored. A
* synthetic uid keyed off (repo, contractId) is the only thing both sides
* can derive without knowing about each other.
*
* Format: `manifest::<repo>::<contractId>`. Stable across syncs, scoped to a
* single repo within a group, and never collides with real indexer uids
* (which never start with `manifest::`).
*/
export function manifestSymbolUid(repo: string, contractId: string): string {
return `manifest::${repo}::${contractId}`;
}
export class ManifestExtractor {
async extractFromManifest(
links: GroupManifestLink[],
dbExecutors?: Map<string, CypherExecutor>,
): Promise<ManifestExtractResult> {
const contracts: StoredContract[] = [];
const crossLinks: CrossLink[] = [];
for (const link of links) {
const contractId = this.buildContractId(link.type, link.contract);
const providerRepo = link.role === 'provider' ? link.from : link.to;
const consumerRepo = link.role === 'provider' ? link.to : link.from;
const providerSymbol = await this.resolveSymbol(providerRepo, link, dbExecutors);
const consumerSymbol = await this.resolveSymbol(consumerRepo, link, dbExecutors);
const providerRef = providerSymbol || { filePath: '', name: link.contract };
const consumerRef = consumerSymbol || { filePath: '', name: link.contract };
// When the resolver finds a real graph symbol we keep its uid, otherwise
// fall back to the deterministic synthetic uid (see manifestSymbolUid).
const providerUid = providerSymbol?.uid || manifestSymbolUid(providerRepo, contractId);
const consumerUid = consumerSymbol?.uid || manifestSymbolUid(consumerRepo, contractId);
contracts.push({
contractId,
type: link.type,
role: 'provider',
symbolUid: providerUid,
symbolRef: providerRef,
symbolName: link.contract,
confidence: 1.0,
meta: { source: 'manifest' },
repo: providerRepo,
});
contracts.push({
contractId,
type: link.type,
role: 'consumer',
symbolUid: consumerUid,
symbolRef: consumerRef,
symbolName: link.contract,
confidence: 1.0,
meta: { source: 'manifest' },
repo: consumerRepo,
});
crossLinks.push({
from: { repo: consumerRepo, symbolUid: consumerUid, symbolRef: consumerRef },
to: { repo: providerRepo, symbolUid: providerUid, symbolRef: providerRef },
type: link.type,
contractId,
matchType: 'manifest',
confidence: 1.0,
});
}
return { contracts, crossLinks };
}
private async resolveSymbol(
repoPathKey: string,
link: GroupManifestLink,
dbExecutors?: Map<string, CypherExecutor>,
): Promise<{ filePath: string; name: string; uid: string } | null> {
const executor = dbExecutors?.get(repoPathKey);
if (!executor) return null;
// NOTE: All lookups use EXACT equality on the relevant name field and
// deterministic ORDER BY before LIMIT 1. Previous versions used CONTAINS
// for fuzzy matching (plus an unconditional ".proto" fallback for gRPC)
// which produced silent false positives: e.g. manifest "/orders" would
// match "/suborders", and a gRPC manifest entry in a repo with any
// .proto file would attach to a random proto symbol.
//
// If resolveSymbol returns null, the extractor falls back to a
// deterministic synthetic uid via `manifestSymbolUid(repo, contractId)`
// (see the function's docstring for why synthetic rather than empty).
// Cross-impact still works: the bridge query joins on the synthetic
// uid, and the local impact engine derives the same uid for the
// unresolved symbol — name-based hints are the additional safety net.
try {
let rows: Record<string, unknown>[];
if (link.type === 'http') {
// Route.name is the canonicalized URL path (see
// core/ingestion/pipeline.ts ensureSlash + generateId('Route', ...)).
// Normalize the manifest contract the same way so a user-written
// "/api/orders" matches "api/orders" in the graph.
//
// The contract may also use the explicit-method form "GET::/api/orders"
// recommended by buildContractId. Strip the METHOD:: prefix before
// normalizing — otherwise `normalizeRoutePath('GET::/api/orders')`
// returns `/GET::/api/orders` and never matches Route.name. The
// captured method is not yet used to constrain the Cypher query
// (method-aware HANDLES_ROUTE matching is a future enhancement).
const parsed = parseHttpContract(link.contract);
const normalized = normalizeRoutePath(parsed.path);
rows = await executor(
`MATCH (handler)-[r:CodeRelation {type: 'HANDLES_ROUTE'}]->(route:Route)
WHERE route.name = $normalized
RETURN handler.id AS uid, handler.name AS name, handler.filePath AS filePath
ORDER BY handler.filePath ASC
LIMIT 1`,
{ normalized },
);
} else if (link.type === 'topic') {
// Topic names aren't a first-class NodeLabel in the graph —
// topics are referenced by function/method symbols (Kafka
// listeners, publishers). Restrict to symbol-like labels to
// avoid cross-matching Files/Variables/Imports that happen to
// share the topic name.
rows = await executor(
`MATCH (n:Function|Method|Class|Interface) WHERE n.name = $contract
RETURN n.id AS uid, n.name AS name, n.filePath AS filePath
ORDER BY n.filePath ASC
LIMIT 1`,
{ contract: link.contract },
);
} else if (link.type === 'grpc') {
// Contract is "Service/Method" or just "Service" (or package.Service
// variants). Prefer matching by method name when present, otherwise
// by service name. NO .proto path fallback — that's guaranteed to
// return a wrong symbol in any repo with more than one proto file.
// Label filters scope lookups: methods → Function|Method, services
// → Class|Interface (no label match = no silent wrong hits on
// File/Variable nodes that happen to share the name).
const parts = link.contract.split('/');
const serviceName = parts[0]?.trim() ?? '';
const methodName = parts[1]?.trim() ?? '';
if (methodName) {
rows = await executor(
`MATCH (n:Function|Method) WHERE n.name = $methodName
RETURN n.id AS uid, n.name AS name, n.filePath AS filePath
ORDER BY n.filePath ASC
LIMIT 1`,
{ methodName },
);
} else if (serviceName) {
rows = await executor(
`MATCH (n:Class|Interface) WHERE n.name = $serviceName
RETURN n.id AS uid, n.name AS name, n.filePath AS filePath
ORDER BY n.filePath ASC
LIMIT 1`,
{ serviceName },
);
} else {
rows = [];
}
} else if (link.type === 'lib') {
// Only exact match on the symbol's name. Previous fallback to
// CONTAINS on n.filePath would promote "react" to "react-native"
// or "@types/react" — silent wrong attribution. Restrict to
// package-level labels so we don't return arbitrary symbols
// named after a library.
rows = await executor(
`MATCH (n:Package|Module) WHERE n.name = $contract
RETURN n.id AS uid, n.name AS name, n.filePath AS filePath
ORDER BY n.filePath ASC
LIMIT 1`,
{ contract: link.contract },
);
} else {
return null;
}
if (rows.length > 0) {
return {
filePath: rows[0].filePath as string,
name: rows[0].name as string,
uid: String(rows[0].uid ?? ''),
};
}
} catch (err) {
// Log but don't throw: a broken graph query in one repo shouldn't
// fail the whole manifest extraction. Unresolved contracts still
// get a synthetic symbolUid below, so cross-impact can proceed.
const message = err instanceof Error ? err.message : String(err);
console.warn(
`[manifest-extractor] resolveSymbol failed for ${link.type}:${link.contract} ` +
`in ${repoPathKey}: ${message}`,
);
}
return null;
}
/**
* Build a canonical contract id for a manifest link.
*
* HTTP is the only type with two valid forms:
* - Explicit method: `"GET::/api/orders"` → `"http::GET::/api/orders"`
* (matches exactly against `HttpRouteExtractor` provider/consumer
* contracts, which are also keyed by `http::<METHOD>::<path>`).
* - Method-agnostic: `"/api/orders"` → `"http::*::/api/orders"`
* — the `*` is a wildcard and is intended to match any concrete
* HTTP method on that path. Wildcard-aware matching is the
* responsibility of the sync / cross-impact layer (see #793);
* downstream code should treat `http::*::<path>` as matching
* every `http::<METHOD>::<path>` for the same path.
*
* Recommend the explicit-method form in group.yaml whenever the
* manifest author knows the method — it round-trips through exact
* equality matching without requiring wildcard logic downstream.
*
* NOTE on exhaustiveness: the switch covers every current
* `ContractType` variant and falls through to a `never` assertion so
* TypeScript fails the build if a new variant is added without a
* corresponding case.
*/
private buildContractId(type: ContractType, contract: string): string {
switch (type) {
case 'http': {
// Canonicalize method casing and path separators so logically
// equivalent inputs (`get::/api/orders` vs `GET::/api/orders`,
// or trailing-slash variants) produce the same contractId and
// matching `manifestSymbolUid` fallback. Without this, raw
// user casing leaks into cross-impact join keys and fragments
// matches across repos.
const { method, path: rawPath } = parseHttpContract(contract);
const normalizedPath = normalizeRoutePath(rawPath);
return method ? `http::${method}::${normalizedPath}` : `http::*::${normalizedPath}`;
}
case 'grpc':
return `grpc::${contract}`;
case 'topic':
return `topic::${contract}`;
case 'lib':
return `lib::${contract}`;
case 'custom':
return `custom::${contract}`;
default: {
const _exhaustive: never = type;
throw new Error(`Unhandled ContractType: ${String(_exhaustive)}`);
}
}
}
}
@@ -1,214 +1,49 @@
import * as fs from 'node:fs';
import * as path from 'node:path';
import { glob } from 'glob';
import Parser from 'tree-sitter';
import type { ContractExtractor, CypherExecutor } from '../contract-extractor.js';
import type { ExtractedContract, RepoHandle } from '../types.js';
import { readSafe } from './fs-utils.js';
import { scanFile, unquoteLiteral } from './tree-sitter-scanner.js';
import {
TOPIC_SCAN_GLOB,
getProviderForFile,
type Broker,
type TopicMeta,
} from './topic-patterns/index.js';
type Broker = 'kafka' | 'rabbitmq' | 'nats';
/**
* Language-agnostic orchestrator for topic (message broker) contract
* extraction. All grammar-specific knowledge lives in `topic-patterns/*`
* — this file must not import any tree-sitter grammar directly.
*
* Flow per file:
* 1. `getProviderForFile(rel)` → compiled plugin (or `undefined` if the
* file's extension isn't registered, in which case we skip it).
* 2. `scanFile(parser, provider, content)` → list of `{meta, valueText}`
* pairs, one per matched literal.
* 3. `unquoteLiteral(valueText)` → the raw topic string.
* 4. `makeContract(topic, meta, relPath)` → `ExtractedContract`.
*
* Adding a new language is a one-file edit in `topic-patterns/index.ts`.
*/
function readSafe(repoPath: string, rel: string): string | null {
const abs = path.resolve(repoPath, rel);
const base = path.resolve(repoPath);
const relToBase = path.relative(base, abs);
if (relToBase.startsWith('..') || path.isAbsolute(relToBase)) return null;
try {
return fs.readFileSync(abs, 'utf-8');
} catch {
return null;
}
}
function makeContract(
topicName: string,
role: 'provider' | 'consumer',
filePath: string,
symbolName: string,
confidence: number,
broker: Broker,
): ExtractedContract {
function makeContract(topicName: string, meta: TopicMeta, filePath: string): ExtractedContract {
return {
contractId: `topic::${topicName}`,
type: 'topic',
role,
role: meta.role,
symbolUid: '',
symbolRef: { filePath: filePath.replace(/\\/g, '/'), name: symbolName },
symbolName,
confidence,
symbolRef: { filePath: filePath.replace(/\\/g, '/'), name: meta.symbolName },
symbolName: meta.symbolName,
confidence: meta.confidence,
meta: {
broker,
broker: meta.broker satisfies Broker,
topicName,
extractionStrategy: 'source_scan',
extractionStrategy: 'tree_sitter',
},
};
}
interface PatternDef {
regex: RegExp;
role: 'provider' | 'consumer';
broker: Broker;
confidence: number;
topicGroup: number;
symbolName: string;
}
// --- Kafka patterns ---
const KAFKA_PATTERNS: PatternDef[] = [
// Java: @KafkaListener(topics = "xxx")
{
regex: /@KafkaListener\s*\(\s*topics\s*=\s*"([^"]+)"/g,
role: 'consumer',
broker: 'kafka',
confidence: 0.8,
topicGroup: 1,
symbolName: 'kafkaListener',
},
// Java: kafkaTemplate.send("xxx"
{
regex: /kafkaTemplate\.send\s*\(\s*"([^"]+)"/gi,
role: 'provider',
broker: 'kafka',
confidence: 0.8,
topicGroup: 1,
symbolName: 'kafkaTemplate.send',
},
// Node: producer.send({ topic: 'xxx'
{
regex: /producer\.send\s*\(\s*\{\s*topic:\s*['"]([^'"]+)['"]/g,
role: 'provider',
broker: 'kafka',
confidence: 0.8,
topicGroup: 1,
symbolName: 'producer.send',
},
// Node: consumer.subscribe({ topic: 'xxx'
{
regex: /consumer\.subscribe\s*\(\s*\{\s*topic:\s*['"]([^'"]+)['"]/g,
role: 'consumer',
broker: 'kafka',
confidence: 0.8,
topicGroup: 1,
symbolName: 'consumer.subscribe',
},
// Go: consumer.ConsumePartition("xxx"
{
regex: /\.ConsumePartition\s*\(\s*"([^"]+)"/g,
role: 'consumer',
broker: 'kafka',
confidence: 0.7,
topicGroup: 1,
symbolName: 'ConsumePartition',
},
// Python: KafkaConsumer('xxx'
{
regex: /KafkaConsumer\s*\(\s*['"]([^'"]+)['"]/g,
role: 'consumer',
broker: 'kafka',
confidence: 0.7,
topicGroup: 1,
symbolName: 'KafkaConsumer',
},
// Python: producer.send('xxx' or producer.produce('xxx'
{
regex: /producer\.(?:send|produce)\s*\(\s*['"]([^'"]+)['"]/g,
role: 'provider',
broker: 'kafka',
confidence: 0.7,
topicGroup: 1,
symbolName: 'producer.send',
},
];
// --- RabbitMQ patterns ---
const RABBITMQ_PATTERNS: PatternDef[] = [
// Java: @RabbitListener(queues = "xxx")
{
regex: /@RabbitListener\s*\(\s*queues\s*=\s*"([^"]+)"/g,
role: 'consumer',
broker: 'rabbitmq',
confidence: 0.8,
topicGroup: 1,
symbolName: 'rabbitListener',
},
// Java: rabbitTemplate.convertAndSend("xxx"
{
regex: /rabbitTemplate\.convertAndSend\s*\(\s*"([^"]+)"/gi,
role: 'provider',
broker: 'rabbitmq',
confidence: 0.8,
topicGroup: 1,
symbolName: 'rabbitTemplate.convertAndSend',
},
// Node: channel.consume("xxx"
{
regex: /channel\.consume\s*\(\s*"([^"]+)"/g,
role: 'consumer',
broker: 'rabbitmq',
confidence: 0.8,
topicGroup: 1,
symbolName: 'channel.consume',
},
// Node: channel.publish("xxx"
{
regex: /channel\.publish\s*\(\s*"([^"]+)"/g,
role: 'provider',
broker: 'rabbitmq',
confidence: 0.8,
topicGroup: 1,
symbolName: 'channel.publish',
},
// Node: channel.sendToQueue("xxx"
{
regex: /channel\.sendToQueue\s*\(\s*"([^"]+)"/g,
role: 'provider',
broker: 'rabbitmq',
confidence: 0.8,
topicGroup: 1,
symbolName: 'channel.sendToQueue',
},
// Python: channel.basic_consume(queue='xxx'
{
regex: /channel\.basic_consume\s*\(\s*queue\s*=\s*['"]([^'"]+)['"]/g,
role: 'consumer',
broker: 'rabbitmq',
confidence: 0.7,
topicGroup: 1,
symbolName: 'basic_consume',
},
// Python: channel.basic_publish(exchange='xxx'
{
regex: /channel\.basic_publish\s*\([^)]*exchange\s*=\s*['"]([^'"]+)['"]/g,
role: 'provider',
broker: 'rabbitmq',
confidence: 0.7,
topicGroup: 1,
symbolName: 'basic_publish',
},
];
// --- NATS patterns ---
const NATS_PATTERNS: PatternDef[] = [
// Go/Node: nc.Subscribe("xxx" or nc.subscribe("xxx"
{
regex: /nc\.(?:S|s)ubscribe\s*\(\s*"([^"]+)"/g,
role: 'consumer',
broker: 'nats',
confidence: 0.8,
topicGroup: 1,
symbolName: 'nc.Subscribe',
},
// Go/Node: nc.Publish("xxx" or nc.publish("xxx"
{
regex: /nc\.(?:P|p)ublish\s*\(\s*"([^"]+)"/g,
role: 'provider',
broker: 'nats',
confidence: 0.8,
topicGroup: 1,
symbolName: 'nc.Publish',
},
];
const ALL_PATTERNS: PatternDef[] = [...KAFKA_PATTERNS, ...RABBITMQ_PATTERNS, ...NATS_PATTERNS];
export class TopicExtractor implements ContractExtractor {
type = 'topic' as const;
@@ -221,46 +56,48 @@ export class TopicExtractor implements ContractExtractor {
repoPath: string,
_repo: RepoHandle,
): Promise<ExtractedContract[]> {
const files = await glob('**/*.{ts,tsx,js,jsx,java,go,py}', {
const files = await glob(TOPIC_SCAN_GLOB, {
cwd: repoPath,
ignore: ['**/node_modules/**', '**/.git/**', '**/vendor/**', '**/dist/**', '**/build/**'],
ignore: [
'**/node_modules/**',
'**/.git/**',
'**/vendor/**',
'**/dist/**',
'**/build/**',
// Language-level test file conventions. Go test files
// `*_test.go` live next to source; other languages either use
// separate test directories (Python's `tests/`, Java's
// `src/test/`) or are already covered by the dist/build ignores.
// Pushed to the glob level so the orchestrator stays
// language-agnostic.
'**/*_test.go',
],
nodir: true,
});
// One parser reused across files; the scanner calls `setLanguage` per
// file based on which plugin the registry returns.
const parser = new Parser();
const out: ExtractedContract[] = [];
for (const rel of files) {
const provider = getProviderForFile(rel);
if (!provider) continue;
const content = readSafe(repoPath, rel);
if (!content) continue;
out.push(...this.scanFile(content, rel));
}
return this.dedupe(out);
}
private scanFile(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
for (const pattern of ALL_PATTERNS) {
// Reset regex state for each file
const re = new RegExp(pattern.regex.source, pattern.regex.flags);
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const topicName = m[pattern.topicGroup];
const matches = scanFile(parser, provider, content);
for (const match of matches) {
const valueNode = match.captures.value;
if (!valueNode) continue;
const topicName = unquoteLiteral(valueNode.text);
if (!topicName) continue;
out.push(
makeContract(
topicName,
pattern.role,
filePath,
pattern.symbolName,
pattern.confidence,
pattern.broker,
),
);
out.push(makeContract(topicName, match.meta, rel));
}
}
return out;
return this.dedupe(out);
}
private dedupe(items: ExtractedContract[]): ExtractedContract[] {
@@ -0,0 +1,123 @@
import Go from 'tree-sitter-go';
import { compilePatterns, type LanguagePatterns } from '../tree-sitter-scanner.js';
import type { TopicMeta } from './types.js';
/**
* Go topic extraction patterns.
*
* Detects Sarama, segmentio/kafka-go and nats.go producer/consumer APIs:
* - `X.ConsumePartition("topic", ...)`
* - `sarama.ProducerMessage{Topic: "xxx"}`
* - `kafka.Writer{Topic: "xxx"}` / `kafka.WriterConfig{Topic: ...}`
* - `kafka.Reader{Topic: "xxx"}` / `kafka.ReaderConfig{Topic: ...}`
* - `nc.Subscribe("topic", ...)` / `js.Subscribe("topic", ...)`
* - `nc.Publish("topic", ...)` / `js.Publish("topic", ...)`
*
* Every query MUST bind `@value` to the topic literal node.
*/
const GO_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
name: 'go-topic',
language: Go,
patterns: [
{
meta: {
role: 'consumer',
broker: 'kafka',
confidence: 0.7,
symbolName: 'ConsumePartition',
},
query: `
(call_expression
function: (selector_expression
field: (field_identifier) @method (#eq? @method "ConsumePartition"))
arguments: (argument_list . (interpreted_string_literal) @value))
`,
},
{
meta: {
role: 'provider',
broker: 'kafka',
confidence: 0.75,
symbolName: 'sarama.ProducerMessage',
},
query: `
(composite_literal
type: (qualified_type
package: (package_identifier) @pkg (#eq? @pkg "sarama")
name: (type_identifier) @ty (#eq? @ty "ProducerMessage"))
body: (literal_value
(keyed_element
(literal_element (identifier) @field (#eq? @field "Topic"))
(literal_element (interpreted_string_literal) @value))))
`,
},
{
meta: {
role: 'provider',
broker: 'kafka',
confidence: 0.75,
symbolName: 'kafka.Writer',
},
query: `
(composite_literal
type: (qualified_type
package: (package_identifier) @pkg (#eq? @pkg "kafka")
name: (type_identifier) @ty (#match? @ty "^(Writer|WriterConfig)$"))
body: (literal_value
(keyed_element
(literal_element (identifier) @field (#eq? @field "Topic"))
(literal_element (interpreted_string_literal) @value))))
`,
},
{
meta: {
role: 'consumer',
broker: 'kafka',
confidence: 0.75,
symbolName: 'kafka.Reader',
},
query: `
(composite_literal
type: (qualified_type
package: (package_identifier) @pkg (#eq? @pkg "kafka")
name: (type_identifier) @ty (#match? @ty "^(Reader|ReaderConfig)$"))
body: (literal_value
(keyed_element
(literal_element (identifier) @field (#eq? @field "Topic"))
(literal_element (interpreted_string_literal) @value))))
`,
},
{
meta: {
role: 'consumer',
broker: 'nats',
confidence: 0.8,
symbolName: 'nc.Subscribe',
},
query: `
(call_expression
function: (selector_expression
operand: (identifier) @obj (#match? @obj "^(nc|js)$")
field: (field_identifier) @method (#match? @method "^[Ss]ubscribe$"))
arguments: (argument_list . (interpreted_string_literal) @value))
`,
},
{
meta: {
role: 'provider',
broker: 'nats',
confidence: 0.8,
symbolName: 'nc.Publish',
},
query: `
(call_expression
function: (selector_expression
operand: (identifier) @obj (#match? @obj "^(nc|js)$")
field: (field_identifier) @method (#match? @method "^[Pp]ublish$"))
arguments: (argument_list . (interpreted_string_literal) @value))
`,
},
],
};
export const GO_TOPIC_PROVIDER = compilePatterns(GO_TOPIC_SPEC);
@@ -0,0 +1,49 @@
import * as path from 'node:path';
import type { CompiledPatterns } from '../tree-sitter-scanner.js';
import type { TopicMeta } from './types.js';
import { JAVA_TOPIC_PROVIDER } from './java.js';
import { GO_TOPIC_PROVIDER } from './go.js';
import { PYTHON_TOPIC_PROVIDER } from './python.js';
import {
JAVASCRIPT_TOPIC_PROVIDER,
TYPESCRIPT_TOPIC_PROVIDER,
TSX_TOPIC_PROVIDER,
} from './node.js';
export type { TopicMeta, Broker } from './types.js';
/**
* File-extension → compiled-plugin registry for topic extraction. The
* top-level orchestrator (`topic-extractor.ts`) looks up the plugin for
* each file it visits and delegates the scanning to `tree-sitter-scanner`.
*
* Keys are lowercase extensions including the leading dot. To add a new
* language, drop a `topic-patterns/<lang>.ts` that exports a compiled
* provider, import it here and register the extension(s). No edits to
* `topic-extractor.ts` are required.
*/
const REGISTRY: Record<string, CompiledPatterns<TopicMeta>> = {
'.java': JAVA_TOPIC_PROVIDER,
'.go': GO_TOPIC_PROVIDER,
'.py': PYTHON_TOPIC_PROVIDER,
'.js': JAVASCRIPT_TOPIC_PROVIDER,
'.jsx': JAVASCRIPT_TOPIC_PROVIDER,
'.ts': TYPESCRIPT_TOPIC_PROVIDER,
'.tsx': TSX_TOPIC_PROVIDER,
};
/**
* Glob pattern for files worth scanning. Kept here so adding a new
* language to the registry also widens the glob automatically via a
* single edit.
*/
export const TOPIC_SCAN_GLOB = '**/*.{ts,tsx,js,jsx,java,go,py}';
/**
* Return the compiled provider registered for the given file's
* extension, or `undefined` if the extension is not registered.
*/
export function getProviderForFile(rel: string): CompiledPatterns<TopicMeta> | undefined {
const ext = path.extname(rel).toLowerCase();
return REGISTRY[ext];
}
@@ -0,0 +1,83 @@
import Java from 'tree-sitter-java';
import { compilePatterns, type LanguagePatterns } from '../tree-sitter-scanner.js';
import type { TopicMeta } from './types.js';
/**
* Java topic extraction patterns.
*
* Detects Kafka and RabbitMQ (Spring conventions) producer/consumer APIs:
* - `@KafkaListener(topics = "xxx")`
* - `@RabbitListener(queues = "xxx")`
* - `kafkaTemplate.send("xxx", ...)`
* - `rabbitTemplate.convertAndSend("xxx", ...)`
*
* Every query MUST bind `@value` to the topic literal node.
*/
const JAVA_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
name: 'java-topic',
language: Java,
patterns: [
{
meta: {
role: 'consumer',
broker: 'kafka',
confidence: 0.8,
symbolName: 'kafkaListener',
},
query: `
(annotation
name: (identifier) @name (#eq? @name "KafkaListener")
arguments: (annotation_argument_list
(element_value_pair
key: (identifier) @key (#eq? @key "topics")
value: (string_literal) @value)))
`,
},
{
meta: {
role: 'consumer',
broker: 'rabbitmq',
confidence: 0.8,
symbolName: 'rabbitListener',
},
query: `
(annotation
name: (identifier) @name (#eq? @name "RabbitListener")
arguments: (annotation_argument_list
(element_value_pair
key: (identifier) @key (#eq? @key "queues")
value: (string_literal) @value)))
`,
},
{
meta: {
role: 'provider',
broker: 'kafka',
confidence: 0.8,
symbolName: 'kafkaTemplate.send',
},
query: `
(method_invocation
object: (identifier) @obj (#eq? @obj "kafkaTemplate")
name: (identifier) @method (#eq? @method "send")
arguments: (argument_list . (string_literal) @value))
`,
},
{
meta: {
role: 'provider',
broker: 'rabbitmq',
confidence: 0.8,
symbolName: 'rabbitTemplate.convertAndSend',
},
query: `
(method_invocation
object: (identifier) @obj (#eq? @obj "rabbitTemplate")
name: (identifier) @method (#eq? @method "convertAndSend")
arguments: (argument_list . (string_literal) @value))
`,
},
],
};
export const JAVA_TOPIC_PROVIDER = compilePatterns(JAVA_TOPIC_SPEC);
@@ -0,0 +1,165 @@
import JavaScript from 'tree-sitter-javascript';
import TypeScript from 'tree-sitter-typescript';
import {
compilePatterns,
type LanguagePatterns,
type PatternSpec,
} from '../tree-sitter-scanner.js';
import type { TopicMeta } from './types.js';
/**
* Node.js / TypeScript topic extraction patterns.
*
* Detects kafkajs, amqplib (RabbitMQ), and nats.js producer/consumer APIs:
* - `producer.send({ topic: 'xxx', ... })` (kafkajs)
* - `consumer.subscribe({ topic: 'xxx', ... })` (kafkajs)
* - `channel.consume("queue", ...)` / `channel.publish(...)` / `channel.sendToQueue(...)`
* - `nc.subscribe("topic")` / `js.subscribe("topic")`
* - `nc.publish("topic", ...)` / `js.publish("topic", ...)`
*
* The JavaScript and TypeScript tree-sitter grammars share node type
* names for every construct we query here, so the pattern sources are
* defined once and compiled against each grammar variant. We export three
* providers because Parser.Query objects are NOT portable across grammar
* instances — `.js` files use the JavaScript grammar, `.ts` uses
* TypeScript.typescript, and `.tsx` uses TypeScript.tsx.
*
* Every query MUST bind `@value` to the topic literal node.
*/
const NODE_TOPIC_PATTERNS: PatternSpec<TopicMeta>[] = [
{
meta: {
role: 'provider',
broker: 'kafka',
confidence: 0.8,
symbolName: 'producer.send',
},
query: `
(call_expression
function: (member_expression
object: (identifier) @obj (#eq? @obj "producer")
property: (property_identifier) @prop (#eq? @prop "send"))
arguments: (arguments
(object
(pair
key: (property_identifier) @key (#eq? @key "topic")
value: [(string) (template_string)] @value))))
`,
},
{
meta: {
role: 'consumer',
broker: 'kafka',
confidence: 0.8,
symbolName: 'consumer.subscribe',
},
query: `
(call_expression
function: (member_expression
object: (identifier) @obj (#eq? @obj "consumer")
property: (property_identifier) @prop (#eq? @prop "subscribe"))
arguments: (arguments
(object
(pair
key: (property_identifier) @key (#eq? @key "topic")
value: [(string) (template_string)] @value))))
`,
},
{
meta: {
role: 'consumer',
broker: 'rabbitmq',
confidence: 0.8,
symbolName: 'channel.consume',
},
query: `
(call_expression
function: (member_expression
object: (identifier) @obj (#eq? @obj "channel")
property: (property_identifier) @prop (#eq? @prop "consume"))
arguments: (arguments . [(string) (template_string)] @value))
`,
},
{
meta: {
role: 'provider',
broker: 'rabbitmq',
confidence: 0.8,
symbolName: 'channel.publish',
},
query: `
(call_expression
function: (member_expression
object: (identifier) @obj (#eq? @obj "channel")
property: (property_identifier) @prop (#eq? @prop "publish"))
arguments: (arguments . [(string) (template_string)] @value))
`,
},
{
meta: {
role: 'provider',
broker: 'rabbitmq',
confidence: 0.8,
symbolName: 'channel.sendToQueue',
},
query: `
(call_expression
function: (member_expression
object: (identifier) @obj (#eq? @obj "channel")
property: (property_identifier) @prop (#eq? @prop "sendToQueue"))
arguments: (arguments . [(string) (template_string)] @value))
`,
},
{
meta: {
role: 'consumer',
broker: 'nats',
confidence: 0.8,
symbolName: 'nc.subscribe',
},
query: `
(call_expression
function: (member_expression
object: (identifier) @obj (#match? @obj "^(nc|js)$")
property: (property_identifier) @prop (#match? @prop "^[Ss]ubscribe$"))
arguments: (arguments . [(string) (template_string)] @value))
`,
},
{
meta: {
role: 'provider',
broker: 'nats',
confidence: 0.8,
symbolName: 'nc.publish',
},
query: `
(call_expression
function: (member_expression
object: (identifier) @obj (#match? @obj "^(nc|js)$")
property: (property_identifier) @prop (#match? @prop "^[Pp]ublish$"))
arguments: (arguments . [(string) (template_string)] @value))
`,
},
];
const JAVASCRIPT_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
name: 'javascript-topic',
language: JavaScript,
patterns: NODE_TOPIC_PATTERNS,
};
const TYPESCRIPT_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
name: 'typescript-topic',
language: TypeScript.typescript,
patterns: NODE_TOPIC_PATTERNS,
};
const TSX_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
name: 'tsx-topic',
language: TypeScript.tsx,
patterns: NODE_TOPIC_PATTERNS,
};
export const JAVASCRIPT_TOPIC_PROVIDER = compilePatterns(JAVASCRIPT_TOPIC_SPEC);
export const TYPESCRIPT_TOPIC_PROVIDER = compilePatterns(TYPESCRIPT_TOPIC_SPEC);
export const TSX_TOPIC_PROVIDER = compilePatterns(TSX_TOPIC_SPEC);
@@ -0,0 +1,119 @@
import Python from 'tree-sitter-python';
import { compilePatterns, type LanguagePatterns } from '../tree-sitter-scanner.js';
import type { TopicMeta } from './types.js';
/**
* Python topic extraction patterns.
*
* Detects kafka-python, pika (RabbitMQ), and nats-py producer/consumer APIs:
* - `KafkaConsumer('topic', ...)`
* - `producer.send('topic', ...)` / `producer.produce('topic', ...)`
* - `channel.basic_consume(queue='xxx', ...)`
* - `channel.basic_publish(exchange='xxx', ...)`
* - `await nc.subscribe('topic')`
* - `await nc.publish('topic', ...)`
*
* Every query MUST bind `@value` to the topic literal node.
*/
const PYTHON_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
name: 'python-topic',
language: Python,
patterns: [
{
meta: {
role: 'consumer',
broker: 'kafka',
confidence: 0.7,
symbolName: 'KafkaConsumer',
},
query: `
(call
function: (identifier) @func (#eq? @func "KafkaConsumer")
arguments: (argument_list . (string) @value))
`,
},
{
meta: {
role: 'provider',
broker: 'kafka',
confidence: 0.7,
symbolName: 'producer.send',
},
query: `
(call
function: (attribute
object: (identifier) @obj (#eq? @obj "producer")
attribute: (identifier) @method (#match? @method "^(send|produce)$"))
arguments: (argument_list . (string) @value))
`,
},
{
meta: {
role: 'consumer',
broker: 'rabbitmq',
confidence: 0.7,
symbolName: 'basic_consume',
},
query: `
(call
function: (attribute
object: (identifier) @obj (#eq? @obj "channel")
attribute: (identifier) @method (#eq? @method "basic_consume"))
arguments: (argument_list
(keyword_argument
name: (identifier) @kw (#eq? @kw "queue")
value: (string) @value)))
`,
},
{
meta: {
role: 'provider',
broker: 'rabbitmq',
confidence: 0.7,
symbolName: 'basic_publish',
},
query: `
(call
function: (attribute
object: (identifier) @obj (#eq? @obj "channel")
attribute: (identifier) @method (#eq? @method "basic_publish"))
arguments: (argument_list
(keyword_argument
name: (identifier) @kw (#eq? @kw "exchange")
value: (string) @value)))
`,
},
{
meta: {
role: 'consumer',
broker: 'nats',
confidence: 0.75,
symbolName: 'nc.subscribe',
},
query: `
(call
function: (attribute
object: (identifier) @obj (#eq? @obj "nc")
attribute: (identifier) @method (#eq? @method "subscribe"))
arguments: (argument_list . (string) @value))
`,
},
{
meta: {
role: 'provider',
broker: 'nats',
confidence: 0.75,
symbolName: 'nc.publish',
},
query: `
(call
function: (attribute
object: (identifier) @obj (#eq? @obj "nc")
attribute: (identifier) @method (#eq? @method "publish"))
arguments: (argument_list . (string) @value))
`,
},
],
};
export const PYTHON_TOPIC_PROVIDER = compilePatterns(PYTHON_TOPIC_SPEC);
@@ -0,0 +1,27 @@
/**
* Shared types for the topic-extractor language plugins.
*
* Each plugin lives in its own file (java.ts, go.ts, ...) and owns the
* tree-sitter grammar import + query sources. The top-level
* `topic-extractor.ts` orchestrator only knows about this type module and
* the plugin registry (`./index.ts`). It MUST NOT import any grammar or
* query text directly — that's the whole point of the split.
*/
export type Broker = 'kafka' | 'rabbitmq' | 'nats';
/**
* Per-pattern payload every topic plugin attaches to its query. Whatever
* the pattern matches, the orchestrator receives this object verbatim
* and uses it to build an `ExtractedContract`.
*
* Plugins produce one `TopicMeta` per pattern (not per match) because a
* single query uniquely identifies its broker/role/confidence triple.
*/
export interface TopicMeta {
role: 'provider' | 'consumer';
broker: Broker;
confidence: number;
/** Short human-readable label of the API being detected. */
symbolName: string;
}
@@ -0,0 +1,193 @@
import Parser from 'tree-sitter';
/**
* Shared, language-agnostic tree-sitter scanning utilities used by group
* extractors (topic, http, grpc, ...).
*
* Design goals:
* - The top-level extractors must not import any tree-sitter grammar.
* - Per-language plugins own their grammar import, their query sources,
* and the mapping from capture → meta.
* - This module provides the plumbing: compile queries once per plugin,
* parse a file with a given grammar, run all patterns, and return the
* captured `string_literal`-style nodes together with the plugin's meta.
*/
/**
* One pattern owned by a language plugin. Each pattern owns a tree-sitter
* S-expression query. Plugins can freely choose which capture names to
* use — the scanner exposes every capture in the returned `captures`
* map and does not privilege any particular name.
*
* `TMeta` is the plugin-specific payload the orchestrator receives back
* when this pattern matches — e.g. for topic extraction it carries the
* broker name, role, confidence, symbol name.
*/
export interface PatternSpec<TMeta> {
/** Tree-sitter S-expression. */
query: string;
/** Plugin-specific payload returned on every match. */
meta: TMeta;
}
/**
* A set of patterns owned by one language plugin, bound to a specific
* tree-sitter grammar.
*
* `language` is typed as `unknown` because tree-sitter's TypeScript
* declarations use `any` for the grammar object, and the grammar modules
* export different shapes (plain grammar vs. namespace with `typescript`
* / `tsx` members). Callers pass the concrete grammar object; this
* module forwards it to `parser.setLanguage` / `new Parser.Query`.
*/
export interface LanguagePatterns<TMeta> {
/** Human-readable plugin name for diagnostics. */
name: string;
/** tree-sitter grammar object. */
language: unknown;
/** Patterns authored against `language`. */
patterns: PatternSpec<TMeta>[];
}
/**
* Compiled form of a `LanguagePatterns` bundle. Queries are compiled
* eagerly at module load time so a broken grammar/query pair fails
* loudly the first time the plugin is imported, instead of silently
* at scan time when no contract is produced.
*/
export interface CompiledPatterns<TMeta> {
name: string;
language: unknown;
patterns: CompiledPattern<TMeta>[];
}
export interface CompiledPattern<TMeta> {
query: Parser.Query;
meta: TMeta;
}
/**
* Map from capture name → syntax node. Every named capture the query
* binds is exposed as an entry. If a query captures the same name more
* than once (unusual), the first occurrence wins — plugins that need
* all occurrences should use distinct capture names or fall back to
* `match.captures` array directly by iterating `query.matches()`
* themselves.
*/
export type CaptureMap = Record<string, Parser.SyntaxNode>;
/**
* One match returned by `scanFile` / `runCompiledPatterns`. The caller
* receives the full capture map plus the plugin meta, and is
* responsible for turning it into a domain object.
*/
export interface ScanMatch<TMeta> {
meta: TMeta;
captures: CaptureMap;
}
/**
* Compile a LanguagePatterns bundle. Call this once per plugin, at
* module load time, and export the result. Throws if any pattern
* fails to compile against the grammar — that's a bug in the plugin
* author's query, not a runtime condition.
*/
export function compilePatterns<TMeta>(bundle: LanguagePatterns<TMeta>): CompiledPatterns<TMeta> {
const compiled: CompiledPattern<TMeta>[] = [];
for (const spec of bundle.patterns) {
try {
const query = new Parser.Query(bundle.language, spec.query);
compiled.push({ query, meta: spec.meta });
} catch (err) {
const message = err instanceof Error ? err.message : String(err);
throw new Error(
`[tree-sitter-scanner] Failed to compile pattern in ${bundle.name}: ${message}\n` +
`Query source:\n${spec.query}`,
);
}
}
return { name: bundle.name, language: bundle.language, patterns: compiled };
}
/**
* Run every compiled pattern in `plugin` against an already-parsed
* tree. Use this when a plugin needs multiple query bundles against
* the same file (e.g. one query for class-level prefixes and another
* for method-level annotations) and wants to avoid re-parsing.
*/
export function runCompiledPatterns<TMeta>(
plugin: CompiledPatterns<TMeta>,
tree: Parser.Tree,
): ScanMatch<TMeta>[] {
const out: ScanMatch<TMeta>[] = [];
for (const compiled of plugin.patterns) {
let matches: Parser.QueryMatch[];
try {
matches = compiled.query.matches(tree.rootNode);
} catch {
continue;
}
for (const match of matches) {
const captures: CaptureMap = {};
for (const cap of match.captures) {
if (!(cap.name in captures)) captures[cap.name] = cap.node;
}
out.push({ meta: compiled.meta, captures });
}
}
return out;
}
/**
* Parse `content` with the plugin's grammar and run every compiled
* pattern against the AST. Returns one `ScanMatch` per matched query
* occurrence, carrying the plugin's meta payload.
*
* Errors are swallowed at the file level (malformed file must not abort
* the whole extract). Individual pattern failures are swallowed too so
* a single unusable query doesn't block the rest of the plugin.
*/
export function scanFile<TMeta>(
parser: Parser,
plugin: CompiledPatterns<TMeta>,
content: string,
): ScanMatch<TMeta>[] {
let tree: Parser.Tree;
try {
parser.setLanguage(plugin.language);
tree = parser.parse(content);
} catch {
return [];
}
return runCompiledPatterns(plugin, tree);
}
/**
* Strip enclosing quotes from a tree-sitter string literal node's text.
* Handles single / double / template quotes, Python triple-quoted strings,
* and Go raw string literals (backticks).
*
* Returns null for empty/nullish input so callers can uniformly skip
* captures whose value is missing.
*/
export function unquoteLiteral(raw: string): string | null {
if (!raw) return null;
// Python triple-quoted
if (
(raw.startsWith('"""') && raw.endsWith('"""')) ||
(raw.startsWith("'''") && raw.endsWith("'''"))
) {
return raw.slice(3, -3);
}
const first = raw[0];
const last = raw[raw.length - 1];
if ((first === '"' || first === "'" || first === '`') && last === first && raw.length >= 2) {
return raw.slice(1, -1);
}
// Some grammars expose the string content without quotes already (e.g.
// Python `string_content` child). Return as-is.
return raw;
}
+123 -13
View File
@@ -5,6 +5,15 @@ export interface MatchResult {
unmatched: StoredContract[];
}
export interface WildcardMatchResult {
matched: CrossLink[];
remaining: StoredContract[];
}
function isGrpcWildcard(cid: string): boolean {
return cid.startsWith('grpc::') && cid.endsWith('/*');
}
export function normalizeContractId(id: string): string {
const colonIdx = id.indexOf('::');
if (colonIdx === -1) return id;
@@ -24,6 +33,22 @@ export function normalizeContractId(id: string): string {
return id;
}
case 'grpc': {
// Canonical form: `grpc::<lowercased-package-or-service>[/<method>]`.
//
// The package/service segment is lowercased because gRPC package
// names are effectively case-insensitive across language bindings
// (`auth.AuthService`, `auth.authservice`, `AUTH.AUTHSERVICE` all
// describe the same wire protocol service). The RPC method segment
// is preserved as-is because the HTTP/2 path used on the wire is
// case-sensitive per the gRPC spec (`/Service/MethodName`), and
// method names in generated clients match the proto source exactly.
//
// A package-only id (no slash) and a package/method id are treated
// as DISTINCT canonical forms: `grpc::userservice` does not match
// `grpc::userservice/Login`. That's by design — callers that want
// service-level manifest matching against method-level providers
// should use the gRPC wildcard form `grpc::UserService/*` which is
// handled by runWildcardMatch below.
const slashIdx = rest.indexOf('/');
if (slashIdx > 0) {
const pkg = rest.substring(0, slashIdx).toLowerCase();
@@ -31,12 +56,12 @@ export function normalizeContractId(id: string): string {
return `grpc::${pkg}${method}`;
}
if (slashIdx === 0) {
// Malformed "package/method" with leading slash — do not lowercase the whole string
// (method segment is case-sensitive per spec).
// Malformed "/method" with leading slash — keep as-is so two
// equally malformed ids can still match each other.
return `grpc::${rest}`;
}
// No slash: spec is ambiguous (package-only vs full service.method). MVP: lowercase
// the whole token; differs from pkg/method split above where RPC method keeps case.
// No slash: package/service only. Lowercase to match the package
// segment produced by the pkg/method branch above.
return `grpc::${rest.toLowerCase()}`;
}
case 'topic':
@@ -66,27 +91,36 @@ function findMatchingKeys(contractId: string, index: Map<string, StoredContract[
return [];
}
export function runExactMatch(contracts: StoredContract[]): MatchResult {
export function buildProviderIndex(contracts: StoredContract[]): Map<string, StoredContract[]> {
const providers = contracts.filter((c) => c.role === 'provider');
const consumers = contracts.filter((c) => c.role === 'consumer');
const providerIndex = new Map<string, StoredContract[]>();
const index = new Map<string, StoredContract[]>();
for (const p of providers) {
const key = normalizeContractId(p.contractId);
const list = providerIndex.get(key) || [];
const list = index.get(key) || [];
list.push(p);
providerIndex.set(key, list);
index.set(key, list);
}
return index;
}
export function runExactMatch(
contracts: StoredContract[],
providerIndex?: Map<string, StoredContract[]>,
): MatchResult {
const index = providerIndex ?? buildProviderIndex(contracts);
// Skip gRPC wildcard consumers — they go to wildcard pass only
const consumers = contracts.filter((c) => c.role === 'consumer' && !isGrpcWildcard(c.contractId));
const matched: CrossLink[] = [];
const matchedConsumerIds = new Set<string>();
const matchedProviderIds = new Set<string>();
for (const consumer of consumers) {
const matchingKeys = findMatchingKeys(consumer.contractId, providerIndex);
const matchingKeys = findMatchingKeys(consumer.contractId, index);
if (matchingKeys.length === 0) continue;
const allMatchingProviders = matchingKeys.flatMap((k) => providerIndex.get(k) || []);
const allMatchingProviders = matchingKeys.flatMap((k) => index.get(k) || []);
for (const provider of allMatchingProviders) {
if (provider.repo === consumer.repo) {
if (!provider.service || !consumer.service || provider.service === consumer.service) {
@@ -118,10 +152,86 @@ export function runExactMatch(contracts: StoredContract[]): MatchResult {
}
}
const unmatched = contracts.filter((c) => {
// normalUnmatched: contracts that weren't matched in exact pass
const normalUnmatched = contracts.filter((c) => {
if (isGrpcWildcard(c.contractId)) return false; // excluded from exact, handled separately
const id = `${c.repo}::${c.contractId}`;
return c.role === 'provider' ? !matchedProviderIds.has(id) : !matchedConsumerIds.has(id);
});
// Re-add gRPC wildcard contracts — they were never in exact matching
const grpcWildcards = contracts.filter((c) => isGrpcWildcard(c.contractId));
const unmatched = [...normalUnmatched, ...grpcWildcards];
return { matched, unmatched };
}
export function runWildcardMatch(
unmatched: StoredContract[],
providerIndex: Map<string, StoredContract[]>,
): WildcardMatchResult {
const wildcardConsumers = unmatched.filter(
(c) => c.role === 'consumer' && isGrpcWildcard(c.contractId),
);
const matched: CrossLink[] = [];
const matchedConsumerIds = new Set<string>();
for (const consumer of wildcardConsumers) {
const normalized = normalizeContractId(consumer.contractId);
// "grpc::com.example.userservice/*" → "com.example.userservice"
// "grpc::userservice/*" → "userservice"
const fqService = normalized.slice(normalized.indexOf('::') + 2, -2); // strip "grpc::" and "/*"
for (const [key, providers] of providerIndex) {
// Only match against non-wildcard gRPC providers (method-level IDs)
if (!key.startsWith('grpc::') || key.endsWith('/*')) continue;
const afterPrefix = key.slice(6); // strip "grpc::"
const slashIdx = afterPrefix.indexOf('/');
if (slashIdx < 0) continue;
const providerFqService = afterPrefix.slice(0, slashIdx);
// Match: exact FQ service, or bare-name match when consumer has no package
const isMatch =
providerFqService === fqService ||
(!fqService.includes('.') && providerFqService.endsWith('.' + fqService));
if (!isMatch) continue;
for (const provider of providers) {
// Skip same-repo same-service (same logic as runExactMatch)
if (provider.repo === consumer.repo) {
if (!provider.service || !consumer.service || provider.service === consumer.service) {
continue;
}
}
matched.push({
from: {
repo: consumer.repo,
service: consumer.service,
symbolUid: consumer.symbolUid,
symbolRef: consumer.symbolRef,
},
to: {
repo: provider.repo,
service: provider.service,
symbolUid: provider.symbolUid,
symbolRef: provider.symbolRef,
},
type: consumer.type,
contractId: consumer.contractId, // consumer's wildcard ID
matchType: 'wildcard',
confidence: Math.min(provider.confidence, consumer.confidence),
});
matchedConsumerIds.add(`${consumer.repo}::${consumer.contractId}`);
}
}
}
const remaining = unmatched.filter((c) => {
if (c.role !== 'consumer' || !isGrpcWildcard(c.contractId)) return true;
return !matchedConsumerIds.has(`${c.repo}::${c.contractId}`);
});
return { matched, remaining };
}
+124
View File
@@ -0,0 +1,124 @@
import type { CrossLink, CrossLinkEndpoint, StoredContract } from './types.js';
function contractKey(contract: StoredContract): string {
return [contract.repo, contract.contractId, contract.role, contract.symbolRef.filePath].join(
'\0',
);
}
function endpointKey(endpoint: CrossLinkEndpoint): string {
return [
endpoint.repo,
endpoint.service ?? '',
endpoint.symbolRef.filePath,
endpoint.symbolRef.name,
].join('\0');
}
/**
* Score a contract by how much information it carries, so `dedupeContracts`
* can prefer the "richer" record when two contracts collide on the same
* `(repo, contractId, role, filePath)` key.
*
* Weights express a priority ordering, not calibrated probabilities:
* +3 — `symbolUid` resolved (tier 1 of the downstream lookup — highest
* signal because it's the strongest anchor for cross-impact traversal
* and the only one that's robust to renames)
* +2 — any of `filePath`, `symbolRef.name`, or `symbolName` that's more
* specific than the contractId itself (tier 2 signal — resolves
* uniquely in most cases and survives across syncs)
* +1 — `service` tag (monorepo attribution — useful but not sufficient
* on its own) or non-manifest origin (auto-extracted contracts are
* preferred over manifest-declared synthetic ones because the former
* are grounded in real source code)
*
* The absolute numbers don't matter, only their relative ordering.
*/
function contractRichness(contract: StoredContract): number {
let score = 0;
if (contract.symbolUid) score += 3;
if (contract.symbolRef.filePath) score += 2;
if (contract.symbolRef.name && contract.symbolRef.name !== contract.contractId) score += 2;
if (contract.symbolName && contract.symbolName !== contract.contractId) score += 2;
if (contract.service) score += 1;
if (contract.meta.source !== 'manifest') score += 1;
return score;
}
function mergeContracts(existing: StoredContract, incoming: StoredContract): StoredContract {
const [primary, secondary] =
contractRichness(incoming) > contractRichness(existing)
? [incoming, existing]
: [existing, incoming];
const symbolRefName = primary.symbolRef.name || secondary.symbolRef.name;
return {
...secondary,
...primary,
symbolUid: primary.symbolUid || secondary.symbolUid,
symbolRef: {
filePath: primary.symbolRef.filePath || secondary.symbolRef.filePath,
name: symbolRefName,
},
symbolName: primary.symbolName || secondary.symbolName || symbolRefName,
confidence: Math.max(existing.confidence, incoming.confidence),
service: primary.service ?? secondary.service,
meta: { ...secondary.meta, ...primary.meta },
};
}
function mergeEndpoints(
existing: CrossLinkEndpoint,
incoming: CrossLinkEndpoint,
): CrossLinkEndpoint {
return {
repo: existing.repo,
service: existing.service ?? incoming.service,
symbolUid: existing.symbolUid || incoming.symbolUid,
symbolRef: {
filePath: existing.symbolRef.filePath || incoming.symbolRef.filePath,
name: existing.symbolRef.name || incoming.symbolRef.name,
},
};
}
function crossLinkKey(link: CrossLink): string {
return [
link.type,
link.contractId,
link.matchType,
endpointKey(link.from),
endpointKey(link.to),
].join('\0');
}
export function dedupeContracts(items: StoredContract[]): StoredContract[] {
const deduped = new Map<string, StoredContract>();
for (const contract of items) {
const key = contractKey(contract);
const existing = deduped.get(key);
deduped.set(key, existing ? mergeContracts(existing, contract) : contract);
}
return [...deduped.values()];
}
export function dedupeCrossLinks(items: CrossLink[]): CrossLink[] {
const deduped = new Map<string, CrossLink>();
for (const link of items) {
const key = crossLinkKey(link);
const existing = deduped.get(key);
if (!existing) {
deduped.set(key, link);
continue;
}
const keepIncoming = link.confidence > existing.confidence;
const primary = keepIncoming ? link : existing;
const secondary = keepIncoming ? existing : link;
deduped.set(key, {
...primary,
confidence: Math.max(existing.confidence, link.confidence),
from: mergeEndpoints(primary.from, secondary.from),
to: mergeEndpoints(primary.to, secondary.to),
});
}
return [...deduped.values()];
}
+15 -1
View File
@@ -1,5 +1,5 @@
export type ContractType = 'http' | 'grpc' | 'topic' | 'lib' | 'custom';
export type MatchType = 'exact' | 'manifest' | 'bm25' | 'embedding';
export type MatchType = 'exact' | 'manifest' | 'wildcard' | 'bm25' | 'embedding';
export type ContractRole = 'provider' | 'consumer';
export interface GroupConfig {
@@ -131,3 +131,17 @@ export interface OutOfScopeLink {
contractId: string;
confidence: number;
}
/** Opaque handle to an open bridge LadybugDB. */
export interface BridgeHandle {
/** Internal — do not access directly. */
readonly _db: unknown;
readonly _conn: unknown;
readonly groupDir: string;
}
export interface BridgeMeta {
version: number;
generatedAt: string;
missingRepos: string[];
}
@@ -0,0 +1,391 @@
/**
* BindingAccumulator — read-append-only accumulator that collects TypeEnv
* bindings across files in the GitNexus analyzer pipeline.
*
* **Current behavior (both execution paths):** The accumulator carries only
* file-scope (`scope = ''`) entries. Function-scope bindings are stripped
* at both write sites:
*
* - **Worker path**: `parse-worker.ts` serializes only
* `typeEnv.fileScope()` entries across the IPC boundary.
* - **Sequential path**: `type-env.ts::flush()` iterates only the FILE_SCOPE
* entry of the env map and writes `BindingEntry` records with
* `scope: ''` hardcoded.
*
* The narrowing exists because function-scope bindings have zero downstream
* consumers today and were previously costing ~4.9 MB of heap + IPC on
* every pipeline run. See `type-env.ts::flush()` and the `FileScopeBindings`
* JSDoc in `parse-worker.ts` for the paired Phase 9 reversion checklist.
*
* **Historical quality asymmetry (Phase 9 consideration):** Even though
* both paths now carry only file-scope data, the two paths were built
* under different resolution capabilities, and a future Phase 9 reverter
* that widens them back to all scopes will inherit that asymmetry:
*
* - **Sequential path** had (and would regain) access to the full
* `SymbolTable` and `importedBindings`, so its bindings benefit from
* Tier 2 cross-file propagation.
* - **Worker path** runs without `SymbolTable` / `importedBindings` and
* can only produce Tier 0 (annotation-declared) and local Tier 1
* (same-file constructor inference) bindings.
*
* Phase 9 consumers that trust every entry equally will silently produce
* worse results for large repos (worker-dominant) than small ones
* (sequential-dominant). If Phase 9 needs homogeneous quality, either
* (a) tag entries with their tier at insert time so consumers can filter,
* or (b) post-process worker-path entries through a follow-up resolution
* pass after the main-thread `SymbolTable` is complete.
*
* **Lifecycle contract**: single-use — `append* → finalize → consume → dispose`.
* After `dispose()` the accumulator is permanently dead: any mutating call
* (`appendFile`) throws, and read methods return empty/undefined as if the
* accumulator had never been appended to. The instance is not recyclable;
* construct a new one for a new pipeline run. Finalization and disposal are
* orthogonal state dimensions and may be invoked in either order.
*/
export interface BindingEntry {
readonly scope: string; // '' for file-level, 'funcName@startIndex' for function-local
readonly varName: string;
readonly typeName: string;
}
/**
* Minimal graph-node shape required by `enrichExportedTypeMap()`. Intentionally
* narrower than the full `GraphNode` type in `graph/types.ts` so tests can
* construct a minimal mock without depending on the full graph module, and
* so the enrichment logic is a pure function over this contract.
*
* Matches the shape of the real `KnowledgeGraph` node's `properties.isExported`
* access path — tests that use a different shape silently pass while
* production fails.
*/
export interface EnrichmentGraphNode {
readonly id: string;
readonly properties?: { readonly isExported?: boolean } | undefined;
}
/**
* Minimal graph lookup interface used by `enrichExportedTypeMap()`.
* Consumes only the method the enrichment loop actually calls.
*/
export interface EnrichmentGraphLookup {
getNode(id: string): EnrichmentGraphNode | undefined;
}
/**
* Merge file-scope bindings from a (finalized) `BindingAccumulator` into an
* `exportedTypeMap` for symbols whose graph nodes are marked as exported.
*
* This is the single source of truth for the worker-path ExportedTypeMap
* enrichment loop. Previously the logic lived inline in `pipeline.ts` and
* the test suite reimplemented it as a `runEnrichmentLoop` helper — a
* drift-prone pattern that meant tests could pass while production regressed.
* Extracting it here makes the production code call the same function the
* tests call.
*
* **Node ID candidate order**: `Function:{filePath}:{name}` →
* `Variable:{filePath}:{name}` → `Const:{filePath}:{name}`. First match wins.
*
* **Tier 0 priority**: if `exportedTypeMap` already has an entry for a
* `(filePath, name)` pair, the accumulator entry does NOT overwrite it —
* the SymbolTable tier-0 pass is authoritative. Without this guard, a
* worker-path binding could clobber a higher-quality type from SymbolTable.
*
* **Finalize precondition**: the accumulator should be finalized before
* calling this function. The lifecycle contract is
* `append → finalize → enrich → dispose`. Finalization is not asserted
* here (the test suite and pipeline both honor it separately), but any
* append happening concurrently with this enrichment would be a lifecycle
* bug at the caller level.
*
* @returns The number of new entries written into `exportedTypeMap`
* (0 on empty accumulator or when every candidate was filtered
* out by the export check or the Tier 0 guard).
*/
export function enrichExportedTypeMap(
bindingAccumulator: BindingAccumulator,
graph: EnrichmentGraphLookup,
exportedTypeMap: Map<string, Map<string, string>>,
): number {
if (bindingAccumulator.fileCount === 0) return 0;
let enriched = 0;
for (const filePath of bindingAccumulator.files()) {
for (const [name, type] of bindingAccumulator.fileScopeEntries(filePath)) {
// Three-candidate-ID lookup mirrors the sequential-path export check
// in `collectExportedBindings()` (call-processor.ts).
const functionNodeId = `Function:${filePath}:${name}`;
const variableNodeId = `Variable:${filePath}:${name}`;
const constNodeId = `Const:${filePath}:${name}`;
const node =
graph.getNode(functionNodeId) ??
graph.getNode(variableNodeId) ??
graph.getNode(constNodeId);
if (!node?.properties?.isExported) continue;
let fileExports = exportedTypeMap.get(filePath);
if (!fileExports) {
fileExports = new Map();
exportedTypeMap.set(filePath, fileExports);
}
// Tier 0 priority: SymbolTable-populated entries are authoritative.
if (!fileExports.has(name)) {
fileExports.set(name, type);
enriched++;
}
}
}
return enriched;
}
const ENTRY_OVERHEAD = 64; // bytes per entry (object overhead + property refs)
const MAP_ENTRY_OVERHEAD = 80; // bytes per file entry in the map
export class BindingAccumulator {
// Storage is split into two parallel maps so file-scope reads are fast.
// - _allByFile holds every BindingEntry (used by getFile, memory estimate).
// - _fileScopeByFile is a nested Map<filePath, Map<varName, typeName>> for
// O(1) point-lookup via fileScopeGet(). For iteration-based consumers
// (enrichExportedTypeMap), fileScopeEntries() iterates the inner Map.
// Both maps carry the same key set modulo the `scope === ''` precondition:
// _allByFile has a key as soon as any entry is appended; _fileScopeByFile
// only has a key once a file-scope entry arrives. Code that iterates via
// files() uses _allByFile so files with only function-scope entries
// remain visible.
//
// Note: Map.set semantics mean a duplicate varName for the same file
// overwrites the previous value (last-write-wins). This is the correct
// behavior — duplicate top-level bindings in the same file shouldn't
// happen in well-formed source, and if they do the last declaration
// is typically the one the compiler sees.
private readonly _allByFile = new Map<string, BindingEntry[]>();
private readonly _fileScopeByFile = new Map<string, Map<string, string>>();
private _totalBindings = 0;
private _finalized = false;
private _disposed = false;
/**
* Append bindings for a file. Safe to call multiple times for the same file.
* Throws if the accumulator has been finalized. Skips if entries is empty.
*
* The `entries` parameter is `readonly` — this method never mutates the
* caller's array. Internally, the first `appendFile` call per filePath
* makes a defensive copy (`slice()`), and subsequent calls push into the
* accumulator's own storage.
*/
appendFile(filePath: string, entries: readonly BindingEntry[]): void {
if (this._finalized) {
throw new Error(
'[BindingAccumulator] appendFile after finalize — no further appends allowed',
);
}
// Single-use lifecycle: once disposed, the accumulator is dead. A
// post-dispose append almost always indicates a missed wiring step
// (the consumer is reading state that was supposed to be released),
// so convert the silent use-after-dispose into a loud failure.
if (this._disposed) {
throw new Error('BindingAccumulator: use after dispose');
}
if (entries.length === 0) {
return;
}
// Note on the file-scope-only invariant:
// The accumulator does NOT reject function-scope entries at this
// boundary. The narrowing contract is enforced by the two production
// write sites — `parse-worker.ts` (which uses `typeEnv.fileScope()`
// and hardcodes `scope: ''` in the pipeline adapter) and
// `type-env.ts::flush()` (which iterates only `env.get(FILE_SCOPE)`).
// The class JSDoc documents the invariant and the Phase 9 reversion
// path. Making `appendFile` runtime-reject non-file-scope entries
// would break the accumulator's own storage-split tests which
// legitimately exercise mixed-scope entries. If a future write path
// violates the invariant, tests should fail via missing exports in
// the enrichment loop, not via an assertion here.
// All-scope store.
const existingAll = this._allByFile.get(filePath);
if (existingAll !== undefined) {
for (const e of entries) {
existingAll.push(e);
}
} else {
this._allByFile.set(filePath, entries.slice());
}
// File-scope fast-path store (nested Map for O(1) point-lookup via fileScopeGet).
// Populated lazily on first file-scope entry per file.
let fileScopeMap = this._fileScopeByFile.get(filePath);
for (const e of entries) {
if (e.scope === '') {
if (fileScopeMap === undefined) {
fileScopeMap = new Map();
this._fileScopeByFile.set(filePath, fileScopeMap);
}
fileScopeMap.set(e.varName, e.typeName);
}
}
this._totalBindings += entries.length;
}
/** Lock the accumulator — no further appends. Idempotent. */
finalize(): void {
// Dev-mode invariant: verify the parallel storage split is consistent.
// `_fileScopeByFile` must be a proper projection of `_allByFile`
// where the outer key is a subset and the inner entries are exactly
// the `scope === ''` subset of `_allByFile[key]`. A drift would
// indicate a bug in `appendFile()` where one map was updated but
// not the other.
if (process.env.NODE_ENV !== 'production' && !this._finalized) {
for (const [filePath, fileScopeMap] of this._fileScopeByFile) {
const allEntries = this._allByFile.get(filePath);
if (allEntries === undefined) {
throw new Error(
`[BindingAccumulator] storage split drift: file ${filePath} has file-scope entries ` +
`but no _allByFile entry`,
);
}
// Count unique file-scope varNames in _allByFile (to match Map dedup
// semantics in _fileScopeByFile where Map.set deduplicates same-name).
const projectedNames = new Set(
allEntries.filter((e) => e.scope === '').map((e) => e.varName),
);
if (projectedNames.size !== fileScopeMap.size) {
throw new Error(
`[BindingAccumulator] storage split drift: file ${filePath} has ` +
`${fileScopeMap.size} file-scope names in Map but ${projectedNames.size} unique ` +
`file-scope varNames in _allByFile`,
);
}
}
}
this._finalized = true;
}
/**
* Release the accumulator's heap footprint. Clears both internal storage
* maps and resets `_totalBindings` to zero. Idempotent — calling twice
* is a no-op. Orthogonal to `finalize()` — calling `dispose()` does not
* change the finalized state.
*
* **Single-use lifecycle.** This is a one-way terminal transition: the
* accumulator is not recyclable. Any subsequent `appendFile` call throws
* (`'BindingAccumulator: use after dispose'`), regardless of whether
* `finalize()` was called first. Post-dispose reads do not throw —
* they return empty/undefined state matching a never-appended-to
* accumulator:
* - `fileCount === 0`
* - `totalBindings === 0`
* - `files()` yields an empty iterator
* - `getFile(x)` returns `undefined` for all `x`
* - `fileScopeEntries(x)` returns `[]` for all `x`
* - `fileScopeGet(x, y)` returns `undefined` for all `x, y`
* - `estimateMemoryBytes()` returns `0`
*
* Lifecycle note: the pipeline disposes the accumulator inside the
* `finally` of the `crossFile` phase, which is scheduled after every
* other accumulator consumer (Phase 9 call/assignment processing and
* the ExportedTypeMap enrichment loop). The dispose call therefore
* runs once, on both the happy path and the throw path of the
* crossFile phase.
*/
dispose(): void {
this._allByFile.clear();
this._fileScopeByFile.clear();
this._totalBindings = 0;
this._disposed = true;
}
/** Get all bindings for a file, or undefined if the file is unknown. */
getFile(filePath: string): readonly BindingEntry[] | undefined {
return this._allByFile.get(filePath);
}
/**
* Get only scope='' (file-level) entries as [varName, typeName] tuples.
* For iteration-based consumers (e.g., `enrichExportedTypeMap`).
* Returns an empty array for an unknown file.
*
* O(1) map lookup + O(n_file_scope) tuple reconstruction from the inner
* Map. Does NOT walk function-scope entries.
*
* For point-lookup consumers (e.g., Phase 9 fallback), prefer
* `fileScopeGet(filePath, name)` — O(1) with no allocation.
*/
fileScopeEntries(filePath: string): readonly (readonly [string, string])[] {
const map = this._fileScopeByFile.get(filePath);
return map ? [...map.entries()] : [];
}
/**
* O(1) point-lookup for a single file-scope binding by (filePath, name).
* Returns the typeName if found, `undefined` otherwise.
*
* This is the preferred lookup path for Phase 9 consumers that resolve
* a single callee's return type — avoids the O(n_file_scope) iteration
* and defensive-copy allocation of `fileScopeEntries()`.
*/
fileScopeGet(filePath: string, name: string): string | undefined {
return this._fileScopeByFile.get(filePath)?.get(name);
}
/** Iterate over all file paths in insertion order. */
files(): IterableIterator<string> {
return this._allByFile.keys();
}
/** Number of distinct files with at least one binding. */
get fileCount(): number {
return this._allByFile.size;
}
/** Total number of binding entries across all files. */
get totalBindings(): number {
return this._totalBindings;
}
/** Whether the accumulator has been finalized. */
get finalized(): boolean {
return this._finalized;
}
/**
* Whether the accumulator has been disposed. Exposed for symmetry with
* `finalized` so debug tooling and future Phase 9 consumers can detect a
* disposed accumulator without inspecting empty state heuristically.
*
* Disposal and finalization are orthogonal: a disposed accumulator may or
* may not be finalized, and vice versa. See `dispose()` for the full
* lifecycle contract.
*/
get disposed(): boolean {
return this._disposed;
}
/**
* Rough memory estimate in bytes (intentionally pessimistic).
* Formula: sum of (ENTRY_OVERHEAD + char bytes of scope+varName+typeName) per entry
* + MAP_ENTRY_OVERHEAD + char bytes of filePath per file.
*
* Note: V8 stores all-ASCII strings as Latin-1 (1 byte/char) and only upgrades
* to UCS-2 (2 bytes/char) for non-Latin-1 code points. Source paths and type names
* are typically all-ASCII, so actual heap cost is roughly half what this returns.
* The pessimistic factor is intentional — better to over-budget than under-budget.
*
* **⚠ Cost profile**: O(totalBindings) — iterates every entry in
* `_allByFile` and reads three string `.length` properties per entry.
* At a typical repo scale (10k files × ~20 file-scope bindings) this is
* ~200k property reads per call. Call at most once per pipeline run,
* NOT per file, per chunk, or per progress tick. The current single
* call site is the dev-mode telemetry log at the pipeline finalize
* seam. Adding a per-file-progress caller would silently make it
* quadratic in repo size.
*/
estimateMemoryBytes(): number {
let total = 0;
for (const [filePath, entries] of this._allByFile) {
total += MAP_ENTRY_OVERHEAD + filePath.length * 2;
for (const e of entries) {
total += ENTRY_OVERHEAD + (e.scope.length + e.varName.length + e.typeName.length) * 2;
}
}
return total;
}
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,177 @@
import type { SyntaxNode } from '../utils/ast-helpers.js';
import type { NodeLabel } from 'gitnexus-shared';
import type {
ClassExtractionConfig,
ClassExtractor,
ClassLikeNodeLabel,
ExtractedClassSymbol,
} from '../class-types.js';
const DEFAULT_SCOPE_NAME_NODE_TYPES = new Set([
'nested_namespace_specifier',
'scoped_identifier',
'scoped_type_identifier',
'qualified_name',
'namespace_name',
'namespace_identifier',
'package_identifier',
'type_identifier',
'identifier',
'name',
'constant',
]);
const DEFAULT_TYPE_NAME_NODE_TYPES = new Set([
'type_identifier',
'identifier',
'simple_identifier',
'namespace_identifier',
'constant',
'name',
]);
const DEFAULT_LABEL_BY_NODE_TYPE: Record<string, ClassLikeNodeLabel> = {
class_declaration: 'Class',
abstract_class_declaration: 'Class',
interface_declaration: 'Interface',
struct_declaration: 'Struct',
record_declaration: 'Record',
enum_declaration: 'Enum',
class_definition: 'Class',
struct_specifier: 'Struct',
class_specifier: 'Class',
enum_specifier: 'Enum',
struct_item: 'Struct',
enum_item: 'Enum',
class: 'Class',
object_declaration: 'Class',
companion_object: 'Class',
protocol_declaration: 'Interface',
extension_declaration: 'Class',
};
const CLASS_LIKE_LABELS = new Set<ClassLikeNodeLabel>([
'Class',
'Struct',
'Interface',
'Enum',
'Record',
]);
const normalizeQualifiedName = (value: string): string =>
value
.replace(/\s+/g, '')
.replace(/^::/, '')
.replace(/::/g, '.')
.replace(/\\/g, '.')
.replace(/\.+/g, '.')
.replace(/^\.+|\.+$/g, '');
const splitQualifiedName = (value: string): string[] => {
const normalized = normalizeQualifiedName(value);
return normalized ? normalized.split('.').filter(Boolean) : [];
};
const extractScopeSegmentsFromNode = (
scopeNode: SyntaxNode,
scopeNameNodeTypes: ReadonlySet<string>,
): string[] => {
const nameNode =
scopeNode.childForFieldName?.('name') ??
scopeNode.namedChildren?.find((child) => scopeNameNodeTypes.has(child.type));
return nameNode ? splitQualifiedName(nameNode.text) : [];
};
const extractTypeNameFromNode = (node: SyntaxNode): string | undefined => {
const nameField = node.childForFieldName?.('name');
if (nameField) return nameField.text;
const nameChild = node.namedChildren?.find((child) =>
DEFAULT_TYPE_NAME_NODE_TYPES.has(child.type),
);
return nameChild?.text;
};
const isClassLikeLabel = (label: NodeLabel | null | undefined): label is ClassLikeNodeLabel =>
label !== undefined && label !== null && CLASS_LIKE_LABELS.has(label as ClassLikeNodeLabel);
export function createClassExtractor(config: ClassExtractionConfig): ClassExtractor {
const typeDeclarationSet = new Set(config.typeDeclarationNodes);
const fileScopeSet = new Set(config.fileScopeNodeTypes ?? []);
const ancestorScopeSet = new Set(config.ancestorScopeNodeTypes ?? []);
const scopeNameNodeTypes = new Set([
...DEFAULT_SCOPE_NAME_NODE_TYPES,
...(config.scopeNameNodeTypes ?? []),
]);
const buildQualifiedName = (node: SyntaxNode, simpleName: string): string => {
let root = node;
while (root.parent) root = root.parent;
const readScopeSegments = (scopeNode: SyntaxNode): string[] =>
config.extractScopeSegments?.(scopeNode) ??
extractScopeSegmentsFromNode(scopeNode, scopeNameNodeTypes);
const fileScopeSegments: string[] = [];
for (const child of root.namedChildren ?? []) {
if (fileScopeSet.has(child.type)) {
fileScopeSegments.push(...readScopeSegments(child));
}
}
const ancestorScopes: string[][] = [];
let current = node.parent;
while (current) {
if (ancestorScopeSet.has(current.type)) {
const segments = readScopeSegments(current);
if (segments.length > 0) ancestorScopes.push(segments);
}
current = current.parent;
}
return [
...fileScopeSegments,
...ancestorScopes.reverse().flat(),
...splitQualifiedName(simpleName),
]
.filter(Boolean)
.join('.');
};
const extract = (
node: SyntaxNode,
fallback?: {
name?: string;
type?: NodeLabel | null;
},
): ExtractedClassSymbol | null => {
if (!typeDeclarationSet.has(node.type)) return null;
const name = config.extractName?.(node) ?? extractTypeNameFromNode(node) ?? fallback?.name;
const type =
config.extractType?.(node) ??
DEFAULT_LABEL_BY_NODE_TYPE[node.type] ??
(isClassLikeLabel(fallback?.type) ? fallback.type : undefined);
if (!name || !type) return null;
return {
name,
type,
qualifiedName: buildQualifiedName(node, name) || name,
};
};
return {
language: config.language,
isTypeDeclaration(node: SyntaxNode): boolean {
return typeDeclarationSet.has(node.type);
},
extract,
extractQualifiedName(node: SyntaxNode, simpleName: string): string | null {
return extract(node, { name: simpleName })?.qualifiedName ?? null;
},
};
}
@@ -0,0 +1,44 @@
import type { NodeLabel, SupportedLanguages } from 'gitnexus-shared';
import type { SyntaxNode } from './utils/ast-helpers.js';
export type ClassLikeNodeLabel = Extract<
NodeLabel,
'Class' | 'Struct' | 'Interface' | 'Enum' | 'Record'
>;
export interface ExtractedClassSymbol {
name: string;
type: ClassLikeNodeLabel;
qualifiedName: string;
}
/**
* Cross-language qualified type names are normalized to dot-separated scope
* segments:
* - file/package scope contributes leading segments when the language has one
* - lexical namespace/module/type scope contributes enclosing segments
* - the simple type name is always the trailing segment
*/
export interface ClassExtractor {
language: SupportedLanguages;
isTypeDeclaration(node: SyntaxNode): boolean;
extract(
node: SyntaxNode,
fallback?: {
name?: string;
type?: NodeLabel | null;
},
): ExtractedClassSymbol | null;
extractQualifiedName(node: SyntaxNode, simpleName: string): string | null;
}
export interface ClassExtractionConfig {
language: SupportedLanguages;
typeDeclarationNodes: string[];
fileScopeNodeTypes?: string[];
ancestorScopeNodeTypes?: string[];
scopeNameNodeTypes?: string[];
extractName?: (node: SyntaxNode) => string | undefined;
extractType?: (node: SyntaxNode) => ClassLikeNodeLabel | undefined;
extractScopeSegments?: (node: SyntaxNode) => string[] | null | undefined;
}
@@ -92,7 +92,7 @@ function isCopybook(filePath: string): boolean {
export const processCobol = (
graph: KnowledgeGraph,
files: CobolFile[],
allPathSet: Set<string>,
allPathSet: ReadonlySet<string>,
): CobolProcessResult => {
const result: CobolProcessResult = {
programs: 0,
+2 -2
View File
@@ -1,7 +1,7 @@
// gitnexus/src/core/ingestion/field-types.ts
import type { TypeEnvironment } from './type-env.js';
import type { SymbolTable } from './symbol-table.js';
import type { SymbolTableReader } from './model/symbol-table.js';
import { SupportedLanguages } from 'gitnexus-shared';
/**
@@ -57,7 +57,7 @@ export interface FieldExtractorContext {
/** Type environment for resolution */
typeEnv: TypeEnvironment;
/** Symbol table for FQN lookups */
symbolTable: SymbolTable;
symbolTable: SymbolTableReader;
/** Current file path */
filePath: string;
/** Language ID */
@@ -1,3 +1,4 @@
import { isVerboseIngestionEnabled } from './utils/verbose.js';
import fs from 'fs/promises';
import path from 'path';
import { glob } from 'glob';
@@ -43,6 +44,7 @@ export const walkRepositoryPaths = async (
const entries: ScannedFile[] = [];
let processed = 0;
let skippedLarge = 0;
const skippedLargePaths: string[] = [];
for (let start = 0; start < filtered.length; start += READ_CONCURRENCY) {
const batch = filtered.slice(start, start + READ_CONCURRENCY);
@@ -52,6 +54,7 @@ export const walkRepositoryPaths = async (
const stat = await fs.stat(fullPath);
if (stat.size > MAX_FILE_SIZE) {
skippedLarge++;
skippedLargePaths.push(relativePath.replace(/\\/g, '/'));
return null;
}
return { path: relativePath.replace(/\\/g, '/'), size: stat.size };
@@ -73,6 +76,11 @@ export const walkRepositoryPaths = async (
console.warn(
` Skipped ${skippedLarge} large files (>${MAX_FILE_SIZE / 1024}KB, likely generated/vendored)`,
);
if (isVerboseIngestionEnabled()) {
for (const p of skippedLargePaths) {
console.warn(` - ${p}`);
}
}
}
return entries;
@@ -19,47 +19,34 @@ import { ASTCache } from './ast-cache.js';
import Parser from 'tree-sitter';
import { isLanguageAvailable, loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { generateId } from '../../lib/utils.js';
import { getLanguageFromFilename } from 'gitnexus-shared';
import { getLanguageFromFilename, type SupportedLanguages } from 'gitnexus-shared';
import { isVerboseIngestionEnabled } from './utils/verbose.js';
import { yieldToEventLoop } from './utils/event-loop.js';
import { SupportedLanguages } from 'gitnexus-shared';
import { getProvider } from './languages/index.js';
import { getTreeSitterBufferSize } from './constants.js';
import type { ExtractedHeritage } from './workers/parse-worker.js';
import type { ResolutionContext } from './resolution-context.js';
import { TIER_CONFIDENCE } from './resolution-context.js';
import type {
ExtractedHeritage,
HeritageResolutionStrategy,
HeritageStrategyLookup,
} from './model/heritage-map.js';
import { resolveExtendsType } from './model/heritage-map.js';
import type { ResolutionContext } from './model/resolution-context.js';
import { TIER_CONFIDENCE } from './model/resolution-context.js';
/**
* Determine whether a heritage.extends capture is actually an IMPLEMENTS relationship.
* Uses the symbol table first (authoritative — Tier 1); falls back to provider-defined
* heuristics for external symbols not present in the graph:
* - interfaceNamePattern: matched against parent name (e.g., /^I[A-Z]/ for C#/Java)
* - heritageDefaultEdge: 'IMPLEMENTS' causes all unresolved parents to map to IMPLEMENTS
* - All others: default EXTENDS
* Derive the heritage-resolution strategy for a language from its
* `LanguageProvider`. This is the production wiring that `buildHeritageMap`
* and the standalone `resolveExtendsType` call site use — the model layer
* itself stays unaware of the provider registry.
*/
/** Exported for implementor-map construction (C#/Java: `extends` rows in base_list may be interfaces). */
export const resolveExtendsType = (
parentName: string,
currentFilePath: string,
ctx: ResolutionContext,
language: SupportedLanguages,
): { type: 'EXTENDS' | 'IMPLEMENTS'; idPrefix: string } => {
const resolved = ctx.resolve(parentName, currentFilePath);
if (resolved && resolved.candidates.length > 0) {
const isInterface = resolved.candidates[0].type === 'Interface';
return isInterface
? { type: 'IMPLEMENTS', idPrefix: 'Interface' }
: { type: 'EXTENDS', idPrefix: 'Class' };
}
// Unresolved symbol — fall back to provider-defined heuristics
const provider = getProvider(language);
if (provider.interfaceNamePattern?.test(parentName)) {
return { type: 'IMPLEMENTS', idPrefix: 'Interface' };
}
if (provider.heritageDefaultEdge === 'IMPLEMENTS') {
return { type: 'IMPLEMENTS', idPrefix: 'Interface' };
}
return { type: 'EXTENDS', idPrefix: 'Class' };
export const getHeritageStrategyForLanguage: HeritageStrategyLookup = (
lang: SupportedLanguages,
): HeritageResolutionStrategy => {
const provider = getProvider(lang);
return {
interfaceNamePattern: provider.interfaceNamePattern,
defaultEdge: provider.heritageDefaultEdge ?? 'EXTENDS',
};
};
/**
@@ -180,7 +167,7 @@ export const processHeritage = async (
parentClassName,
file.path,
ctx,
language,
getHeritageStrategyForLanguage(language),
);
const child = resolveHeritageId(
@@ -296,7 +283,7 @@ export const processHeritageFromExtracted = async (
h.parentName,
h.filePath,
ctx,
fileLanguage,
getHeritageStrategyForLanguage(fileLanguage),
);
const child = resolveHeritageId(
@@ -372,7 +359,7 @@ export const processHeritageFromExtracted = async (
/**
* Walk source files with the same heritage captures as parse-worker, producing
* {@link ExtractedHeritage} rows without mutating the graph. Used on the
* sequential pipeline path so `buildImplementorMap(..., ctx)` can run before
* sequential pipeline path so `buildHeritageMap(..., ctx)` can run before
* `processCalls` (worker path defers calls until heritage from all chunks exists).
*/
export async function extractExtractedHeritageFromFiles(
@@ -12,7 +12,11 @@ import type { ExtractedImport } from './workers/parse-worker.js';
import { getTreeSitterBufferSize } from './constants.js';
import { loadImportConfigs } from './language-config.js';
import { buildSuffixIndex } from './import-resolvers/utils.js';
import type { ResolutionContext, ModuleAliasMap } from './resolution-context.js';
import type {
ResolutionContext,
ModuleAliasMap,
NamedImportMap,
} from './model/resolution-context.js';
import type {
ImportResult,
ResolveCtx,
@@ -20,8 +24,7 @@ import type {
} from './import-resolvers/types.js';
import type { NamedBinding } from './named-bindings/types.js';
import type { SyntaxNode } from './utils/ast-helpers.js';
const isDev = process.env.NODE_ENV === 'development';
import { isDev } from './utils/env.js';
// Type: Map<FilePath, Set<ResolvedFilePath>>
// Stores all files that a given file imports from
@@ -61,30 +64,6 @@ function wireImplicitImports(
// Avoids expanding every Go package import into N individual ImportMap edges.
export type PackageMap = Map<string, Set<string>>;
// Type: Map<ImportingFilePath, Map<LocalName, {sourcePath, exportedName}>>
// Tracks which specific names a file imports from which sources (TS/Python only).
// Used to tighten Tier 2a resolution: `import { User } from './models'`
// means only `User` (not `Repo`) is visible from models.ts via this import.
// Stores both the resolved source path and the original exported name so that
// aliased imports (`import { User as U }`) can resolve U → User in the source file.
export interface NamedImportBinding {
sourcePath: string;
exportedName: string;
}
export type NamedImportMap = Map<string, Map<string, NamedImportBinding>>;
/**
* Check if a file path is directly inside a package directory identified by its suffix.
* Used by the symbol resolver for Go and C# directory-level import matching.
*/
export function isFileInPackageDir(filePath: string, dirSuffix: string): boolean {
// Prepend '/' so paths like "internal/auth/service.go" match suffix "/internal/auth/"
const normalized = '/' + filePath.replace(/\\/g, '/');
if (!normalized.includes(dirSuffix)) return false;
const afterDir = normalized.substring(normalized.indexOf(dirSuffix) + dirSuffix.length);
return !afterDir.includes('/');
}
// ImportResolutionContext is defined in ./import-resolvers/types.ts — re-exported here for consumers.
export function buildImportResolutionContext(allPaths: string[]): ImportResolutionContext {
@@ -2,7 +2,7 @@ import fs from 'fs/promises';
import path from 'path';
import type { ImportConfigs } from './import-resolvers/types.js';
const isDev = process.env.NODE_ENV === 'development';
import { isDev } from './utils/env.js';
// ============================================================================
// LANGUAGE-SPECIFIC CONFIG TYPES
@@ -9,9 +9,10 @@
* so adding a language to the enum without creating a provider is a compiler error.
*/
import type { SupportedLanguages } from 'gitnexus-shared';
import type { SupportedLanguages, MroStrategy } from 'gitnexus-shared';
import type { LanguageTypeConfig } from './type-extractors/types.js';
import type { CallRouter } from './call-routing.js';
import type { ClassExtractor } from './class-types.js';
import type { ExportChecker } from './export-detection.js';
import type { FieldExtractor } from './field-extractor.js';
import type { MethodExtractor } from './method-types.js';
@@ -25,15 +26,34 @@ import type { NodeLabel } from 'gitnexus-shared';
export type CaptureMap = Record<string, SyntaxNode | undefined>;
// ── Strategy tag types ─────────────────────────────────────────────────────
/** MRO strategy for multiple inheritance resolution. */
export type MroStrategy =
| 'first-wins'
| 'c3'
| 'leftmost-base'
| 'implements-split'
| 'qualified-syntax';
/** How a language handles imports — determines wildcard synthesis behavior. */
export type ImportSemantics = 'named' | 'wildcard' | 'namespace';
// NOTE: `MroStrategy` is defined in `gitnexus-shared` and re-exported above
// so `core/ingestion/model/resolve.ts` can consume it without importing from
// this file (which would pull in the full language-registry dependency graph).
/**
* How a language handles imports — determines wildcard synthesis behavior.
*
* Import resolution is a graph-traversal policy with multiple distinct strategies,
* analogous to MRO for method resolution. Each tag picks a strategy:
*
* | Tag | Mechanism | Traversal | Languages |
* |-----------------------|------------------------------------------------|---------------------|--------------------------------------------|
* | `named` | Per-symbol imports | None (use-site) | JS/TS, Java, C#, Rust, PHP, Kotlin, Vue |
* | `wildcard-transitive` | Textual paste, symbols chain through files | BFS closure | C, C++ (future: Obj-C, Fortran, Nim) |
* | `wildcard-leaf` | Whole public API, single hop | None (direct only) | Go, Ruby, Swift, Dart |
* | `namespace` | Qualified handle; symbols resolved at call site| None at import | Python |
* | `explicit-reexport` | Opt-in per-symbol re-export (SCAFFOLD) | Topological DAG | (future: TS `export *`, Rust `pub use`) |
*
* The `explicit-reexport` tag is a compile-time scaffold; no provider claims it yet.
* It falls through to `wildcard-leaf` behavior in synthesis so today's TS/Rust
* handling is unchanged. A future PR will implement the DAG walk for `export *`.
*/
export type ImportSemantics =
| 'named'
| 'wildcard-transitive'
| 'wildcard-leaf'
| 'namespace'
| 'explicit-reexport';
/**
* Everything a language needs to provide.
@@ -70,10 +90,12 @@ interface LanguageProviderConfig {
/** Named binding extraction from import statements.
* Default: undefined (language uses wildcard/whole-module imports). */
readonly namedBindingExtractor?: NamedBindingExtractorFn;
/** How this language handles imports.
/** How this language handles imports. See `ImportSemantics` for the full taxonomy.
* - 'named': per-symbol imports (JS/TS, Java, C#, Rust, PHP, Kotlin)
* - 'wildcard': whole-module imports, needs synthesis (Go, Ruby, C/C++, Swift)
* - 'namespace': namespace imports, needs moduleAliasMap (Python)
* - 'wildcard-transitive': textual-include closure; imports chain through files (C, C++)
* - 'wildcard-leaf': whole-module single-hop imports; no transitive chaining (Go, Ruby, Swift, Dart)
* - 'namespace': qualified namespace imports, needs moduleAliasMap (Python)
* - 'explicit-reexport': opt-in per-symbol re-export (scaffold; no provider uses yet)
* Default: 'named'. */
readonly importSemantics?: ImportSemantics;
/** Language-specific transformation of raw import path text before resolution.
@@ -91,6 +113,16 @@ interface LanguageProviderConfig {
projectConfig: unknown,
) => void;
// ── Enclosing owner resolution ─────────────────────────────────
/** Resolve a container node during enclosing-owner tree walks.
* Called when a CLASS_CONTAINER_TYPES node is found while walking up.
* - Return a different SyntaxNode to remap the container (e.g., Ruby
* singleton_class → enclosing class/module).
* - Return null to skip this container and keep walking up.
* - Omit (undefined) to use the container node as-is (default).
* Default: undefined (no remapping). */
readonly resolveEnclosingOwner?: (node: SyntaxNode) => SyntaxNode | null;
// ── Enclosing function resolution ───────────────────────────────
/** Resolve the enclosing function name + label from an AST ancestor node
* that is NOT a standard FUNCTION_NODE_TYPE. For languages where the
@@ -131,6 +163,10 @@ interface LanguageProviderConfig {
* declarations. Produces MethodInfo[] with name, parameters, visibility, isAbstract,
* isFinal, annotations metadata. Default: undefined (no method extraction). */
readonly methodExtractor?: MethodExtractor;
/** Class/type extractor for deriving canonical qualified names for class-like symbols.
* Uses the same provider-driven strategy pattern as method/field extraction so
* namespace/package/module rules stay language-specific. */
readonly classExtractor?: ClassExtractor;
/** Extract a semantic description for a definition node (e.g., PHP Eloquent
* property arrays, relation method descriptions).
* Default: undefined (no description extraction). */
+185 -5
View File
@@ -9,13 +9,26 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as cCppConfig } from '../type-extractors/c-cpp.js';
import { cCppExportChecker } from '../export-detection.js';
import { resolveCImport, resolveCppImport } from '../import-resolvers/standard.js';
import { C_QUERIES, CPP_QUERIES } from '../tree-sitter-queries.js';
import { isCppInsideClassOrStruct } from '../utils/ast-helpers.js';
/**
* Node types for standard function declarations that need C/C++ declarator handling.
* Used by cCppExtractFunctionName to determine how to extract the function name.
*/
const FUNCTION_DECLARATION_TYPES = new Set([
'function_declaration',
'function_definition',
'async_function_declaration',
'generator_function_declaration',
'function_item',
]);
import type { SyntaxNode } from '../utils/ast-helpers.js';
import type { NodeLabel } from 'gitnexus-shared';
import type { LanguageProvider } from '../language-provider.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import {
@@ -132,6 +145,165 @@ const C_BUILT_INS: ReadonlySet<string> = new Set([
'put',
]);
const cClassExtractor = createClassExtractor({
language: SupportedLanguages.C,
typeDeclarationNodes: ['struct_specifier', 'enum_specifier'],
});
const cppClassExtractor = createClassExtractor({
language: SupportedLanguages.CPlusPlus,
typeDeclarationNodes: ['class_specifier', 'struct_specifier', 'enum_specifier'],
ancestorScopeNodeTypes: ['namespace_definition', 'class_specifier', 'struct_specifier'],
});
/**
* C/C++ function name extraction — unwraps pointer_declarator / reference_declarator /
* function_declarator / qualified_identifier chains to find the actual function name.
* Handles field_identifier (method inside class body) and parenthesized_declarator.
*/
const cCppExtractFunctionName = (
node: SyntaxNode,
): { funcName: string | null; label: NodeLabel } | null => {
if (!FUNCTION_DECLARATION_TYPES.has(node.type)) return null;
let funcName: string | null = null;
let label: NodeLabel = 'Function';
// C/C++: function_definition -> [pointer_declarator ->] function_declarator -> qualified_identifier/identifier
// Unwrap pointer_declarator / reference_declarator wrappers to reach function_declarator
let declarator = node.childForFieldName?.('declarator');
if (!declarator) {
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (c?.type === 'function_declarator') {
declarator = c;
break;
}
}
}
while (
declarator &&
(declarator.type === 'pointer_declarator' || declarator.type === 'reference_declarator')
) {
let nextDeclarator = declarator.childForFieldName?.('declarator');
if (!nextDeclarator) {
for (let i = 0; i < declarator.childCount; i++) {
const c = declarator.child(i);
if (
c?.type === 'function_declarator' ||
c?.type === 'pointer_declarator' ||
c?.type === 'reference_declarator'
) {
nextDeclarator = c;
break;
}
}
}
declarator = nextDeclarator;
}
if (declarator) {
let innerDeclarator = declarator.childForFieldName?.('declarator');
if (!innerDeclarator) {
for (let i = 0; i < declarator.childCount; i++) {
const c = declarator.child(i);
if (
c?.type === 'qualified_identifier' ||
c?.type === 'identifier' ||
c?.type === 'field_identifier' ||
c?.type === 'parenthesized_declarator'
) {
innerDeclarator = c;
break;
}
}
}
if (innerDeclarator?.type === 'qualified_identifier') {
let nameNode = innerDeclarator.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < innerDeclarator.childCount; i++) {
const c = innerDeclarator.child(i);
if (c?.type === 'identifier') {
nameNode = c;
break;
}
}
}
if (nameNode?.text) {
funcName = nameNode.text;
label = 'Method';
}
} else if (
innerDeclarator?.type === 'identifier' ||
innerDeclarator?.type === 'field_identifier'
) {
// field_identifier is used for method names inside C++ class bodies
funcName = innerDeclarator.text;
if (innerDeclarator.type === 'field_identifier') label = 'Method';
} else if (innerDeclarator?.type === 'parenthesized_declarator') {
let nestedId: SyntaxNode | null = null;
for (let i = 0; i < innerDeclarator.childCount; i++) {
const c = innerDeclarator.child(i);
if (c?.type === 'qualified_identifier' || c?.type === 'identifier') {
nestedId = c;
break;
}
}
if (nestedId?.type === 'qualified_identifier') {
let nameNode = nestedId.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < nestedId.childCount; i++) {
const c = nestedId.child(i);
if (c?.type === 'identifier') {
nameNode = c;
break;
}
}
}
if (nameNode?.text) {
funcName = nameNode.text;
label = 'Method';
}
} else if (nestedId?.type === 'identifier') {
funcName = nestedId.text;
}
}
}
// Fallback for other node types in FUNCTION_DECLARATION_TYPES (e.g. function_item for Rust in C++ tree)
if (!funcName) {
let nameNode = node.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (
c?.type === 'identifier' ||
c?.type === 'property_identifier' ||
c?.type === 'simple_identifier'
) {
nameNode = c;
break;
}
}
}
funcName = nameNode?.text ?? null;
}
return { funcName, label };
};
/** Check if a C/C++ function_definition is inside a class or struct body.
* Used by cppLabelOverride to skip duplicate function captures
* that are already covered by definition.method queries. */
function isCppInsideClassOrStruct(functionNode: SyntaxNode): boolean {
let ancestor: SyntaxNode | null = functionNode?.parent ?? null;
while (ancestor) {
if (ancestor.type === 'class_specifier' || ancestor.type === 'struct_specifier') return true;
ancestor = ancestor.parent;
}
return false;
}
/** Label override shared by C and C++: skip function_definition captures inside class/struct
* bodies (they're duplicates of definition.method captures). */
const cppLabelOverride: NonNullable<LanguageProvider['labelOverride']> = (
@@ -149,9 +321,13 @@ export const cProvider = defineLanguage({
typeConfig: cCppConfig,
exportChecker: cCppExportChecker,
importResolver: resolveCImport,
importSemantics: 'wildcard',
importSemantics: 'wildcard-transitive',
fieldExtractor: createFieldExtractor(cFieldConfig),
methodExtractor: createMethodExtractor(cMethodConfig),
methodExtractor: createMethodExtractor({
...cMethodConfig,
extractFunctionName: cCppExtractFunctionName,
}),
classExtractor: cClassExtractor,
labelOverride: cppLabelOverride,
builtInNames: C_BUILT_INS,
});
@@ -163,10 +339,14 @@ export const cppProvider = defineLanguage({
typeConfig: cCppConfig,
exportChecker: cCppExportChecker,
importResolver: resolveCppImport,
importSemantics: 'wildcard',
importSemantics: 'wildcard-transitive',
mroStrategy: 'leftmost-base',
fieldExtractor: createFieldExtractor(cppFieldConfig),
methodExtractor: createMethodExtractor(cppMethodConfig),
methodExtractor: createMethodExtractor({
...cppMethodConfig,
extractFunctionName: cCppExtractFunctionName,
}),
classExtractor: cppClassExtractor,
labelOverride: cppLabelOverride,
builtInNames: C_BUILT_INS,
});
@@ -7,6 +7,7 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as csharpConfig } from '../type-extractors/csharp.js';
import { csharpExportChecker } from '../export-detection.js';
@@ -125,5 +126,24 @@ export const csharpProvider = defineLanguage({
mroStrategy: 'implements-split',
fieldExtractor: createFieldExtractor(csharpFieldConfig),
methodExtractor: createMethodExtractor(csharpMethodConfig),
classExtractor: createClassExtractor({
language: SupportedLanguages.CSharp,
typeDeclarationNodes: [
'class_declaration',
'interface_declaration',
'struct_declaration',
'enum_declaration',
'record_declaration',
],
fileScopeNodeTypes: ['file_scoped_namespace_declaration'],
ancestorScopeNodeTypes: [
'namespace_declaration',
'class_declaration',
'interface_declaration',
'struct_declaration',
'enum_declaration',
'record_declaration',
],
}),
builtInNames: BUILT_INS,
});
+29 -6
View File
@@ -2,7 +2,7 @@
* Dart Language Provider
*
* Dart traits:
* - importSemantics: 'wildcard' (Dart imports bring everything public into scope)
* - importSemantics: 'wildcard-leaf' (Dart imports bring everything public into scope)
* - exportChecker: public if no leading underscore
* - Dart SDK imports (dart:*) and external packages are skipped
* - enclosingFunctionFinder: Dart's tree-sitter grammar places function_body
@@ -12,8 +12,9 @@
import type { SyntaxNode } from '../utils/ast-helpers.js';
import type { NodeLabel } from 'gitnexus-shared';
import { FUNCTION_NODE_TYPES, extractFunctionName } from '../utils/ast-helpers.js';
import { FUNCTION_NODE_TYPES } from '../utils/ast-helpers.js';
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as dartConfig } from '../type-extractors/dart.js';
import { dartExportChecker } from '../export-detection.js';
@@ -21,6 +22,8 @@ import { resolveDartImport } from '../import-resolvers/dart.js';
import { DART_QUERIES } from '../tree-sitter-queries.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { dartConfig as dartFieldConfig } from '../field-extractors/configs/dart.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import { dartMethodConfig } from '../method-extractors/configs/dart.js';
/**
* Resolve the enclosing function from a `function_body` node by looking at its
@@ -28,8 +31,8 @@ import { dartConfig as dartFieldConfig } from '../field-extractors/configs/dart.
* function_body are siblings under program or class_body, unlike most languages
* where the function declaration wraps both.
*
* Delegates name extraction to the shared `extractFunctionName` which already
* handles Dart's function_signature and method_signature node types.
* Extracts the function name inline — Dart uses function_signature and
* method_signature (which wraps function_signature) as its FUNCTION_NODE_TYPES.
*/
const dartEnclosingFunctionFinder = (
node: SyntaxNode,
@@ -37,7 +40,21 @@ const dartEnclosingFunctionFinder = (
if (node.type !== 'function_body') return null;
const prev = node.previousSibling;
if (!prev || !FUNCTION_NODE_TYPES.has(prev.type)) return null;
const { funcName, label } = extractFunctionName(prev);
// method_signature wraps function_signature — unwrap to reach the name
let target = prev;
let label: NodeLabel = 'Function';
if (prev.type === 'method_signature') {
label = 'Method';
for (let i = 0; i < prev.childCount; i++) {
const c = prev.child(i);
if (c?.type === 'function_signature') {
target = c;
break;
}
}
}
const funcName = target.childForFieldName?.('name')?.text ?? null;
return funcName ? { funcName, label } : null;
};
@@ -73,8 +90,14 @@ export const dartProvider = defineLanguage({
typeConfig: dartConfig,
exportChecker: dartExportChecker,
importResolver: resolveDartImport,
importSemantics: 'wildcard',
importSemantics: 'wildcard-leaf',
fieldExtractor: createFieldExtractor(dartFieldConfig),
methodExtractor: createMethodExtractor(dartMethodConfig),
classExtractor: createClassExtractor({
language: SupportedLanguages.Dart,
typeDeclarationNodes: ['class_definition', 'extension_declaration', 'enum_declaration'],
ancestorScopeNodeTypes: ['class_definition', 'extension_declaration', 'enum_declaration'],
}),
enclosingFunctionFinder: dartEnclosingFunctionFinder,
builtInNames: BUILT_INS,
});
+22 -2
View File
@@ -5,11 +5,12 @@
* LanguageProvider, following the Strategy pattern used by the pipeline.
*
* Key Go traits:
* - importSemantics: 'wildcard' (Go imports entire packages)
* - importSemantics: 'wildcard-leaf' (Go imports entire packages)
* - callRouter: present (Go method calls may need routing)
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as goConfig } from '../type-extractors/go.js';
import { goExportChecker } from '../export-detection.js';
@@ -17,6 +18,8 @@ import { resolveGoImport } from '../import-resolvers/go.js';
import { GO_QUERIES } from '../tree-sitter-queries.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { goConfig as goFieldConfig } from '../field-extractors/configs/go.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import { goMethodConfig } from '../method-extractors/configs/go.js';
export const goProvider = defineLanguage({
id: SupportedLanguages.Go,
@@ -25,6 +28,23 @@ export const goProvider = defineLanguage({
typeConfig: goConfig,
exportChecker: goExportChecker,
importResolver: resolveGoImport,
importSemantics: 'wildcard',
importSemantics: 'wildcard-leaf',
fieldExtractor: createFieldExtractor(goFieldConfig),
methodExtractor: createMethodExtractor(goMethodConfig),
classExtractor: createClassExtractor({
language: SupportedLanguages.Go,
typeDeclarationNodes: ['type_declaration'],
fileScopeNodeTypes: ['package_clause'],
extractName(node) {
const typeSpec = node.namedChildren.find((child) => child.type === 'type_spec');
return typeSpec?.childForFieldName('name')?.text;
},
extractType(node) {
const typeSpec = node.namedChildren.find((child) => child.type === 'type_spec');
const typeNode = typeSpec?.childForFieldName('type');
if (typeNode?.type === 'struct_type') return 'Struct';
if (typeNode?.type === 'interface_type') return 'Interface';
return undefined;
},
}),
});
@@ -8,6 +8,7 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { javaTypeConfig } from '../type-extractors/jvm.js';
import { javaExportChecker } from '../export-detection.js';
@@ -31,4 +32,20 @@ export const javaProvider = defineLanguage({
mroStrategy: 'implements-split',
fieldExtractor: createFieldExtractor(javaConfig),
methodExtractor: createMethodExtractor(javaMethodConfig),
classExtractor: createClassExtractor({
language: SupportedLanguages.Java,
typeDeclarationNodes: [
'class_declaration',
'interface_declaration',
'enum_declaration',
'record_declaration',
],
fileScopeNodeTypes: ['package_declaration'],
ancestorScopeNodeTypes: [
'class_declaration',
'interface_declaration',
'enum_declaration',
'record_declaration',
],
}),
});
@@ -8,6 +8,7 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { kotlinTypeConfig } from '../type-extractors/jvm.js';
import { kotlinExportChecker } from '../export-detection.js';
@@ -15,12 +16,26 @@ import { resolveKotlinImport } from '../import-resolvers/jvm.js';
import { extractKotlinNamedBindings } from '../named-bindings/kotlin.js';
import { appendKotlinWildcard } from '../import-resolvers/jvm.js';
import { KOTLIN_QUERIES } from '../tree-sitter-queries.js';
import { isKotlinClassMethod } from '../utils/ast-helpers.js';
import type { SyntaxNode } from '../utils/ast-helpers.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { kotlinConfig } from '../field-extractors/configs/jvm.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import { kotlinMethodConfig } from '../method-extractors/configs/jvm.js';
/** Check if a Kotlin function_declaration capture is inside a class_body (i.e., a method).
* Kotlin grammar uses function_declaration for both top-level functions and class methods.
* Returns true when the captured definition node has a class_body ancestor. */
function isKotlinClassMethod(
captureNode: { parent?: SyntaxNode | null } | null | undefined,
): boolean {
let ancestor = captureNode?.parent;
while (ancestor) {
if (ancestor.type === 'class_body') return true;
ancestor = ancestor.parent;
}
return false;
}
const BUILT_INS: ReadonlySet<string> = new Set([
'println',
'print',
@@ -92,6 +107,16 @@ export const kotlinProvider = defineLanguage({
mroStrategy: 'implements-split',
fieldExtractor: createFieldExtractor(kotlinConfig),
methodExtractor: createMethodExtractor(kotlinMethodConfig),
classExtractor: createClassExtractor({
language: SupportedLanguages.Kotlin,
typeDeclarationNodes: ['class_declaration', 'object_declaration', 'companion_object'],
fileScopeNodeTypes: ['package_header'],
ancestorScopeNodeTypes: ['class_declaration', 'object_declaration', 'companion_object'],
extractType(node) {
if (node.type !== 'class_declaration') return undefined;
return node.children.some((child) => child?.text === 'interface') ? 'Interface' : 'Class';
},
}),
builtInNames: BUILT_INS,
labelOverride: (functionNode, defaultLabel) => {
if (defaultLabel !== 'Function') return defaultLabel;
+24 -11
View File
@@ -7,6 +7,7 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as phpConfig } from '../type-extractors/php.js';
import { phpExportChecker } from '../export-detection.js';
@@ -17,6 +18,8 @@ import { findDescendant, extractStringContent, type SyntaxNode } from '../utils/
import type { NodeLabel } from 'gitnexus-shared';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { phpConfig as phpFieldConfig } from '../field-extractors/configs/php.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import { phpMethodConfig } from '../method-extractors/configs/php.js';
const BUILT_INS: ReadonlySet<string> = new Set([
'echo',
@@ -157,18 +160,22 @@ function extractPhpPropertyDescription(propName: string, propDeclNode: SyntaxNod
* Returns description like "hasMany(Post)" or null.
*/
function extractEloquentRelationDescription(methodNode: SyntaxNode): string | null {
function findRelationCall(node: SyntaxNode): SyntaxNode | null {
if (node.type === 'member_call_expression') {
function findRelationCall(root: SyntaxNode): SyntaxNode | null {
const stack: SyntaxNode[] = [root];
while (stack.length > 0) {
const node = stack.pop()!;
if (node.type === 'member_call_expression') {
const children = node.children ?? [];
const objectNode = children.find(
(c: SyntaxNode) => c.type === 'variable_name' && c.text === '$this',
);
const nameNode = children.find((c: SyntaxNode) => c.type === 'name');
if (objectNode && nameNode && ELOQUENT_RELATIONS.has(nameNode.text)) return node;
}
const children = node.children ?? [];
const objectNode = children.find(
(c: SyntaxNode) => c.type === 'variable_name' && c.text === '$this',
);
const nameNode = children.find((c: SyntaxNode) => c.type === 'name');
if (objectNode && nameNode && ELOQUENT_RELATIONS.has(nameNode.text)) return node;
}
for (const child of node.children ?? []) {
const found = findRelationCall(child);
if (found) return found;
for (let i = children.length - 1; i >= 0; i--) {
stack.push(children[i]);
}
}
return null;
}
@@ -231,6 +238,12 @@ export const phpProvider = defineLanguage({
importResolver: resolvePhpImport,
namedBindingExtractor: extractPhpNamedBindings,
fieldExtractor: createFieldExtractor(phpFieldConfig),
methodExtractor: createMethodExtractor(phpMethodConfig),
classExtractor: createClassExtractor({
language: SupportedLanguages.PHP,
typeDeclarationNodes: ['class_declaration', 'interface_declaration', 'enum_declaration'],
ancestorScopeNodeTypes: ['namespace_definition'],
}),
descriptionExtractor: phpDescriptionExtractor,
isRouteFile: isPhpRouteFile,
builtInNames: BUILT_INS,
@@ -11,6 +11,7 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as pythonConfig } from '../type-extractors/python.js';
import { pythonExportChecker } from '../export-detection.js';
@@ -19,6 +20,8 @@ import { extractPythonNamedBindings } from '../named-bindings/python.js';
import { PYTHON_QUERIES } from '../tree-sitter-queries.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { pythonConfig as pythonFieldConfig } from '../field-extractors/configs/python.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import { pythonMethodConfig } from '../method-extractors/configs/python.js';
const BUILT_INS: ReadonlySet<string> = new Set([
'print',
@@ -61,5 +64,11 @@ export const pythonProvider = defineLanguage({
importSemantics: 'namespace',
mroStrategy: 'c3',
fieldExtractor: createFieldExtractor(pythonFieldConfig),
methodExtractor: createMethodExtractor(pythonMethodConfig),
classExtractor: createClassExtractor({
language: SupportedLanguages.Python,
typeDeclarationNodes: ['class_definition'],
ancestorScopeNodeTypes: ['class_definition'],
}),
builtInNames: BUILT_INS,
});
+49 -1
View File
@@ -8,7 +8,10 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import type { NodeLabel } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import type { SyntaxNode } from '../utils/ast-helpers.js';
import { typeConfig as rubyConfig } from '../type-extractors/ruby.js';
import { routeRubyCall } from '../call-routing.js';
import { rubyExportChecker } from '../export-detection.js';
@@ -16,6 +19,27 @@ import { resolveRubyImport } from '../import-resolvers/ruby.js';
import { RUBY_QUERIES } from '../tree-sitter-queries.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { rubyConfig as rubyFieldConfig } from '../field-extractors/configs/ruby.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import { rubyMethodConfig } from '../method-extractors/configs/ruby.js';
/** Ruby method/singleton_method: extract name from 'name' field, label as Method. */
const rubyExtractFunctionName = (
node: SyntaxNode,
): { funcName: string | null; label: NodeLabel } | null => {
if (node.type !== 'method' && node.type !== 'singleton_method') return null;
let nameNode = node.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (c?.type === 'identifier') {
nameNode = c;
break;
}
}
}
return { funcName: nameNode?.text ?? null, label: 'Method' };
};
const BUILT_INS: ReadonlySet<string> = new Set([
'puts',
@@ -83,7 +107,31 @@ export const rubyProvider = defineLanguage({
exportChecker: rubyExportChecker,
importResolver: resolveRubyImport,
callRouter: routeRubyCall,
importSemantics: 'wildcard',
importSemantics: 'wildcard-leaf',
resolveEnclosingOwner(node) {
// Ruby singleton_class (class << self) should resolve to the enclosing
// class or module for owner/container resolution (HAS_METHOD edges, class IDs).
if (node.type === 'singleton_class') {
let ancestor = node.parent;
while (ancestor) {
if (ancestor.type === 'class' || ancestor.type === 'module') {
return ancestor;
}
ancestor = ancestor.parent;
}
return null; // no enclosing class/module — skip
}
return node; // use as-is for all other container types
},
fieldExtractor: createFieldExtractor(rubyFieldConfig),
methodExtractor: createMethodExtractor({
...rubyMethodConfig,
extractFunctionName: rubyExtractFunctionName,
}),
classExtractor: createClassExtractor({
language: SupportedLanguages.Ruby,
typeDeclarationNodes: ['class'],
ancestorScopeNodeTypes: ['module', 'class'],
}),
builtInNames: BUILT_INS,
});
@@ -11,7 +11,10 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import type { NodeLabel } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import type { SyntaxNode } from '../utils/ast-helpers.js';
import { typeConfig as rustConfig } from '../type-extractors/rust.js';
import { rustExportChecker } from '../export-detection.js';
import { resolveRustImport } from '../import-resolvers/rust.js';
@@ -19,6 +22,37 @@ import { extractRustNamedBindings } from '../named-bindings/rust.js';
import { RUST_QUERIES } from '../tree-sitter-queries.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { rustConfig as rustFieldConfig } from '../field-extractors/configs/rust.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import { rustMethodConfig } from '../method-extractors/configs/rust.js';
/** Rust impl_item: find the function_item child and extract its name as a Method. */
const rustExtractFunctionName = (
node: SyntaxNode,
): { funcName: string | null; label: NodeLabel } | null => {
if (node.type !== 'impl_item') return null;
let funcItem: SyntaxNode | null = null;
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (c?.type === 'function_item') {
funcItem = c;
break;
}
}
if (!funcItem) return null;
let nameNode = funcItem.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < funcItem.childCount; i++) {
const c = funcItem.child(i);
if (c?.type === 'identifier') {
nameNode = c;
break;
}
}
}
return { funcName: nameNode?.text ?? null, label: 'Method' };
};
const BUILT_INS: ReadonlySet<string> = new Set([
'unwrap',
@@ -87,5 +121,14 @@ export const rustProvider = defineLanguage({
namedBindingExtractor: extractRustNamedBindings,
mroStrategy: 'qualified-syntax',
fieldExtractor: createFieldExtractor(rustFieldConfig),
methodExtractor: createMethodExtractor({
...rustMethodConfig,
extractFunctionName: rustExtractFunctionName,
}),
classExtractor: createClassExtractor({
language: SupportedLanguages.Rust,
typeDeclarationNodes: ['struct_item', 'enum_item'],
ancestorScopeNodeTypes: ['mod_item', 'struct_item', 'enum_item'],
}),
builtInNames: BUILT_INS,
});
+32 -2
View File
@@ -5,20 +5,25 @@
* LanguageProvider, following the Strategy pattern used by the pipeline.
*
* Key Swift traits:
* - importSemantics: 'wildcard' (Swift imports entire modules)
* - importSemantics: 'wildcard-leaf' (Swift imports entire modules)
* - heritageDefaultEdge: 'IMPLEMENTS' (protocols are more common than class inheritance)
* - implicitImportWirer: all files in the same SPM target see each other
*/
import { SupportedLanguages } from 'gitnexus-shared';
import type { NodeLabel } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as swiftConfig } from '../type-extractors/swift.js';
import { swiftExportChecker } from '../export-detection.js';
import { resolveSwiftImport } from '../import-resolvers/swift.js';
import { SWIFT_QUERIES } from '../tree-sitter-queries.js';
import type { SwiftPackageConfig } from '../language-config.js';
import type { SyntaxNode } from '../utils/ast-helpers.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { swiftConfig as swiftFieldConfig } from '../field-extractors/configs/swift.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import { swiftMethodConfig } from '../method-extractors/configs/swift.js';
/**
* Group Swift files by SPM target for implicit module visibility.
@@ -107,6 +112,15 @@ function wireSwiftImplicitImports(
}
}
/** Swift init/deinit declarations have special names and Constructor label. */
const swiftExtractFunctionName = (
node: SyntaxNode,
): { funcName: string | null; label: NodeLabel } | null => {
if (node.type === 'init_declaration') return { funcName: 'init', label: 'Constructor' };
if (node.type === 'deinit_declaration') return { funcName: 'deinit', label: 'Constructor' };
return null; // fall through to generic
};
const BUILT_INS: ReadonlySet<string> = new Set([
'print',
'debugPrint',
@@ -224,9 +238,25 @@ export const swiftProvider = defineLanguage({
typeConfig: swiftConfig,
exportChecker: swiftExportChecker,
importResolver: resolveSwiftImport,
importSemantics: 'wildcard',
importSemantics: 'wildcard-leaf',
heritageDefaultEdge: 'IMPLEMENTS',
fieldExtractor: createFieldExtractor(swiftFieldConfig),
methodExtractor: createMethodExtractor({
...swiftMethodConfig,
extractFunctionName: swiftExtractFunctionName,
}),
classExtractor: createClassExtractor({
language: SupportedLanguages.Swift,
typeDeclarationNodes: ['class_declaration', 'protocol_declaration'],
ancestorScopeNodeTypes: ['class_declaration', 'protocol_declaration'],
extractType(node) {
if (node.type === 'protocol_declaration') return 'Interface';
if (node.type !== 'class_declaration') return undefined;
if (node.children.some((child) => child?.text === 'struct')) return 'Struct';
if (node.children.some((child) => child?.text === 'enum')) return 'Enum';
return 'Class';
},
}),
implicitImportWirer: wireSwiftImplicitImports,
builtInNames: BUILT_INS,
});
@@ -8,7 +8,11 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import type { NodeLabel } from 'gitnexus-shared';
import { defineLanguage } from '../language-provider.js';
import { createClassExtractor } from '../class-extractors/generic.js';
import type { ClassExtractionConfig } from '../class-types.js';
import type { SyntaxNode } from '../utils/ast-helpers.js';
import { typeConfig as typescriptConfig } from '../type-extractors/typescript.js';
import { tsExportChecker } from '../export-detection.js';
import { resolveTypescriptImport, resolveJavascriptImport } from '../import-resolvers/standard.js';
@@ -23,6 +27,31 @@ import {
javascriptMethodConfig,
} from '../method-extractors/configs/typescript-javascript.js';
/**
* TypeScript/JavaScript: arrow_function and function_expression get their name
* from the parent variable_declarator (e.g. `const foo = () => {}`).
*/
const tsExtractFunctionName = (
node: SyntaxNode,
): { funcName: string | null; label: NodeLabel } | null => {
if (node.type !== 'arrow_function' && node.type !== 'function_expression') return null;
const parent = node.parent;
if (parent?.type !== 'variable_declarator') return null;
let nameNode = parent.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < parent.childCount; i++) {
const c = parent.child(i);
if (c?.type === 'identifier') {
nameNode = c;
break;
}
}
}
return { funcName: nameNode?.text ?? null, label: 'Function' };
};
export const BUILT_INS: ReadonlySet<string> = new Set([
'console',
'log',
@@ -120,6 +149,22 @@ export const BUILT_INS: ReadonlySet<string> = new Set([
'valueOf',
]);
const tsJsClassConfig: ClassExtractionConfig = {
language: SupportedLanguages.TypeScript,
typeDeclarationNodes: [
'class_declaration',
'abstract_class_declaration',
'interface_declaration',
'enum_declaration',
],
ancestorScopeNodeTypes: [
'class_declaration',
'abstract_class_declaration',
'interface_declaration',
'enum_declaration',
],
};
export const typescriptProvider = defineLanguage({
id: SupportedLanguages.TypeScript,
extensions: ['.ts', '.tsx'],
@@ -129,7 +174,11 @@ export const typescriptProvider = defineLanguage({
importResolver: resolveTypescriptImport,
namedBindingExtractor: extractTsNamedBindings,
fieldExtractor: typescriptFieldExtractor,
methodExtractor: createMethodExtractor(typescriptMethodConfig),
methodExtractor: createMethodExtractor({
...typescriptMethodConfig,
extractFunctionName: tsExtractFunctionName,
}),
classExtractor: createClassExtractor(tsJsClassConfig),
builtInNames: BUILT_INS,
});
@@ -142,6 +191,13 @@ export const javascriptProvider = defineLanguage({
importResolver: resolveJavascriptImport,
namedBindingExtractor: extractTsNamedBindings,
fieldExtractor: createFieldExtractor(javascriptConfig),
methodExtractor: createMethodExtractor(javascriptMethodConfig),
methodExtractor: createMethodExtractor({
...javascriptMethodConfig,
extractFunctionName: tsExtractFunctionName,
}),
classExtractor: createClassExtractor({
...tsJsClassConfig,
language: SupportedLanguages.JavaScript,
}),
builtInNames: BUILT_INS,
});
@@ -12,6 +12,7 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as typescriptConfig } from '../type-extractors/typescript.js';
import { tsExportChecker } from '../export-detection.js';
@@ -55,6 +56,22 @@ const VUE_SPECIFIC_BUILT_INS = [
const VUE_BUILT_INS: ReadonlySet<string> = new Set([...TS_BUILT_INS, ...VUE_SPECIFIC_BUILT_INS]);
const vueClassExtractor = createClassExtractor({
language: SupportedLanguages.Vue,
typeDeclarationNodes: [
'class_declaration',
'abstract_class_declaration',
'interface_declaration',
'enum_declaration',
],
ancestorScopeNodeTypes: [
'class_declaration',
'abstract_class_declaration',
'interface_declaration',
'enum_declaration',
],
});
export const vueProvider = defineLanguage({
id: SupportedLanguages.Vue,
extensions: ['.vue'],
@@ -64,5 +81,6 @@ export const vueProvider = defineLanguage({
importResolver: resolveVueImport,
namedBindingExtractor: extractTsNamedBindings,
fieldExtractor: typescriptFieldExtractor,
classExtractor: vueClassExtractor,
builtInNames: VUE_BUILT_INS,
});
@@ -23,7 +23,7 @@ interface MdFile {
export const processMarkdown = (
graph: KnowledgeGraph,
files: MdFile[],
allPathSet: Set<string>,
allPathSet: ReadonlySet<string>,
): { sections: number; links: number } => {
let totalSections = 0;
let totalLinks = 0;
@@ -87,7 +87,7 @@ function extractCppMethodName(node: SyntaxNode): string | undefined {
function extractCppReturnType(node: SyntaxNode): string | undefined {
const typeNode = node.childForFieldName('type');
if (typeNode) {
const typeText = extractSimpleTypeName(typeNode) ?? typeNode.text?.trim();
const typeText = typeNode.text?.trim();
// C++11 trailing return type: `auto foo() -> ReturnType`
// When the declared type is `auto`, check for a trailing_return_type on the
// function_declarator which holds the actual return type.
@@ -99,7 +99,7 @@ function extractCppReturnType(node: SyntaxNode): string | undefined {
if (child?.type === 'trailing_return_type') {
// trailing_return_type contains a type_descriptor with the real type
const typeDesc = child.firstNamedChild;
if (typeDesc) return extractSimpleTypeName(typeDesc) ?? typeDesc.text?.trim();
if (typeDesc) return typeDesc.text?.trim();
}
}
}
@@ -115,7 +115,7 @@ function extractCppReturnType(node: SyntaxNode): string | undefined {
first.type === 'sized_type_specifier' ||
first.type === 'template_type')
) {
return extractSimpleTypeName(first) ?? first.text?.trim();
return first.text?.trim();
}
return undefined;
}
@@ -148,6 +148,7 @@ function extractCppParameters(node: SyntaxNode): ParameterInfo[] {
type: typeNode
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: false,
isVariadic: false,
});
@@ -162,6 +163,7 @@ function extractCppParameters(node: SyntaxNode): ParameterInfo[] {
type: typeNode
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: true,
isVariadic: false,
});
@@ -177,6 +179,7 @@ function extractCppParameters(node: SyntaxNode): ParameterInfo[] {
type: typeNode
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: false,
isVariadic: true,
});
@@ -187,6 +190,7 @@ function extractCppParameters(node: SyntaxNode): ParameterInfo[] {
params.push({
name: '...',
type: null,
rawType: null,
isOptional: false,
isVariadic: true,
});
@@ -194,6 +198,25 @@ function extractCppParameters(node: SyntaxNode): ParameterInfo[] {
}
}
}
// C/C++: bare `...` token in parameter list is an unnamed child (not a named node).
// Check all children for the unnamed `...` token when no variadic was detected above.
if (!params.some((p) => p.isVariadic)) {
for (let i = 0; i < paramList.childCount; i++) {
const child = paramList.child(i);
if (child && !child.isNamed && child.text === '...') {
params.push({
name: '...',
type: null,
rawType: null,
isOptional: false,
isVariadic: true,
});
break;
}
}
}
return params;
}
@@ -292,8 +315,8 @@ function hasVirtualSpecifier(node: SyntaxNode, keyword: string): boolean {
// This includes namespace-wrapped and nested classes.
// - Friend declarations are not extracted.
// - Template method declarations with explicit specialization.
// - const-qualified method overloads (e.g. begin() vs begin() const) collapse
// to the same name — the schema has no isConst field to distinguish them.
// - const-qualified method overloads (e.g. begin() vs begin() const) are
// disambiguated via isConst flag and $const ID suffix.
export const cppMethodConfig: MethodExtractionConfig = {
language: SupportedLanguages.CPlusPlus,
typeDeclarationNodes: ['class_specifier', 'struct_specifier', 'union_specifier'],
@@ -334,6 +357,20 @@ export const cppMethodConfig: MethodExtractionConfig = {
isOverride(node) {
return hasVirtualSpecifier(node, 'override');
},
isConst(node) {
// const qualifier appears as a type_qualifier child of function_declarator,
// after the parameter_list: e.g. `int size() const` → funcDecl has
// type_qualifier child with text "const". Not to be confused with return-type
// const (e.g. `const int& begin()`) which is at a different AST level.
const funcDecl = findFunctionDeclarator(node);
if (!funcDecl) return false;
for (let i = 0; i < funcDecl.namedChildCount; i++) {
const child = funcDecl.namedChild(i);
if (child?.type === 'type_qualifier' && child.text === 'const') return true;
}
return false;
},
};
// ---------------------------------------------------------------------------
@@ -1,4 +1,5 @@
// gitnexus/src/core/ingestion/method-extractors/configs/csharp.ts
// Verified against tree-sitter-c-sharp 0.23.1
import { SupportedLanguages } from 'gitnexus-shared';
import type {
@@ -85,6 +86,7 @@ function extractParametersFromList(paramList: SyntaxNode): ParameterInfo[] {
type: typeNode
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: false,
isVariadic: true,
});
@@ -126,6 +128,7 @@ function extractParametersFromList(paramList: SyntaxNode): ParameterInfo[] {
params.push({
name: nameNode.text,
type: typeName,
rawType: typeNode?.text?.trim() ?? null,
isOptional,
isVariadic: false,
});
@@ -186,6 +189,7 @@ export const csharpMethodConfig: MethodExtractionConfig = {
'destructor_declaration',
'operator_declaration',
'conversion_operator_declaration',
'local_function_statement',
],
bodyNodeTypes: ['declaration_list'],
@@ -221,7 +225,7 @@ export const csharpMethodConfig: MethodExtractionConfig = {
// Constructors and destructors have no return type
// operator_declaration and conversion_operator_declaration use 'type' field, not 'returns'
const returnsNode = node.childForFieldName('returns');
if (returnsNode) return extractSimpleTypeName(returnsNode) ?? returnsNode.text?.trim();
if (returnsNode) return returnsNode.text?.trim();
// Fallback for operator/conversion declarations that use 'type' as return type field
if (node.type === 'operator_declaration' || node.type === 'conversion_operator_declaration') {
const typeNode = node.childForFieldName('type');
@@ -0,0 +1,409 @@
// gitnexus/src/core/ingestion/method-extractors/configs/dart.ts
// Verified against tree-sitter-dart 1.0.0 (80e23c07)
import { SupportedLanguages } from 'gitnexus-shared';
import type {
MethodExtractionConfig,
ParameterInfo,
MethodVisibility,
} from '../../method-types.js';
import { extractSimpleTypeName } from '../../type-extractors/shared.js';
import type { SyntaxNode } from '../../utils/ast-helpers.js';
// ---------------------------------------------------------------------------
// Dart helpers
// ---------------------------------------------------------------------------
/** Type node types that represent a return type in function/getter/setter signatures. */
const TYPE_NODE_TYPES = new Set([
'type_identifier',
'generic_type',
'function_type',
'nullable_type',
'void_type',
'record_type',
]);
/**
* Dart method_signature is a WRAPPER node containing one inner signature:
* function_signature, constructor_signature, getter_signature, setter_signature,
* operator_signature, or factory_constructor_signature.
*
* Name, parameters, and return type live on the INNER signature, not on
* method_signature itself.
*/
function getInnerSignature(node: SyntaxNode): SyntaxNode | null {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (
child &&
(child.type === 'function_signature' ||
child.type === 'constructor_signature' ||
child.type === 'getter_signature' ||
child.type === 'setter_signature' ||
child.type === 'operator_signature' ||
child.type === 'factory_constructor_signature')
) {
return child;
}
}
// `declaration` nodes (abstract methods) also wrap function_signature as a
// named child — handled by the loop above.
return null;
}
/**
* Extract the method name from a method_signature node.
*
* Descends into the inner signature to find the name field/identifier.
*/
function extractDartName(node: SyntaxNode): string | undefined {
const inner = getInnerSignature(node);
if (!inner) return undefined;
// constructor_signature name field may include "ClassName.namedCtor" via multiple children.
// getter_signature, setter_signature, function_signature all have a 'name' field.
if (inner.type === 'operator_signature') {
// operator_signature has no 'name' field; name is 'operator' + the operator symbol
for (let i = 0; i < inner.namedChildCount; i++) {
const child = inner.namedChild(i);
if (child?.type === 'binary_operator') {
return `operator ${child.text.trim()}`;
}
}
// Check for unnamed operator tokens like []= or ~
for (let i = 0; i < inner.childCount; i++) {
const child = inner.child(i);
if (child && !child.isNamed && child.text.trim() !== 'operator') {
const text = child.text.trim();
if (text && !TYPE_NODE_TYPES.has(child.type)) {
return `operator ${text}`;
}
}
}
return undefined;
}
if (inner.type === 'getter_signature') {
const nameNode = inner.childForFieldName('name');
return nameNode?.text;
}
if (inner.type === 'setter_signature') {
const nameNode = inner.childForFieldName('name');
return nameNode ? `set ${nameNode.text}` : undefined;
}
if (inner.type === 'factory_constructor_signature') {
// Collect all identifier children to form "ClassName" or "ClassName.named"
const parts: string[] = [];
for (let i = 0; i < inner.childCount; i++) {
const child = inner.child(i);
if (child?.isNamed && child.type === 'identifier') {
parts.push(child.text);
}
}
return parts.length > 0 ? parts.join('.') : undefined;
}
// function_signature and constructor_signature both have a 'name' field
const nameNode = inner.childForFieldName('name');
if (nameNode) {
// constructor_signature: name field may be multiple identifiers joined by '.'
return nameNode.text;
}
return undefined;
}
/**
* Extract the return type from the inner signature.
*
* function_signature children include type nodes before the name.
* getter_signature children include type nodes before 'get' keyword.
* constructor/setter signatures have no return type.
*/
function extractDartReturnType(node: SyntaxNode): string | undefined {
const inner = getInnerSignature(node);
if (!inner) return undefined;
// Constructors and setters have no return type
if (
inner.type === 'constructor_signature' ||
inner.type === 'setter_signature' ||
inner.type === 'factory_constructor_signature'
) {
return undefined;
}
// For function_signature, getter_signature, operator_signature:
// The type node is a named child before the name/operator
for (let i = 0; i < inner.namedChildCount; i++) {
const child = inner.namedChild(i);
if (child && TYPE_NODE_TYPES.has(child.type)) {
return child.text?.trim();
}
}
return undefined;
}
/**
* Extract parameters from the inner signature's formal_parameter_list.
*
* Dart parameters can be:
* - Positional required: `int x`
* - Optional positional: `[int? x]` — wrapped in optional_formal_parameters with '['
* - Optional named: `{int? x}` or `{required int x}` — wrapped in optional_formal_parameters with '{'
*/
function extractDartParameters(node: SyntaxNode): ParameterInfo[] {
const inner = getInnerSignature(node);
if (!inner) return [];
// getter_signature has no parameters
if (inner.type === 'getter_signature') return [];
// Find formal_parameter_list — it's a child, not a field in function_signature
let paramList: SyntaxNode | null = null;
if (inner.type === 'constructor_signature' || inner.type === 'factory_constructor_signature') {
paramList = inner.childForFieldName('parameters');
}
if (!paramList) {
for (let i = 0; i < inner.namedChildCount; i++) {
const child = inner.namedChild(i);
if (child?.type === 'formal_parameter_list') {
paramList = child;
break;
}
}
}
if (!paramList) return [];
return extractParamsFromList(paramList, false);
}
/**
* Extract ParameterInfo entries from a formal_parameter_list or optional_formal_parameters node.
*/
function extractParamsFromList(listNode: SyntaxNode, isOptionalBlock: boolean): ParameterInfo[] {
const params: ParameterInfo[] = [];
for (let i = 0; i < listNode.namedChildCount; i++) {
const child = listNode.namedChild(i);
if (!child) continue;
if (child.type === 'formal_parameter') {
params.push(extractSingleParam(child, isOptionalBlock));
} else if (child.type === 'optional_formal_parameters') {
// Determine if these are named ({}) or positional ([]) optional params
// by checking the surrounding delimiters
params.push(...extractParamsFromList(child, true));
}
}
return params;
}
/**
* Extract a single ParameterInfo from a formal_parameter node.
*/
function extractSingleParam(param: SyntaxNode, isOptionalBlock: boolean): ParameterInfo {
const nameNode = param.childForFieldName('name');
const name = nameNode?.text ?? '<unknown>';
// Find the type node
let typeName: string | null = null;
let rawTypeName: string | null = null;
for (let i = 0; i < param.namedChildCount; i++) {
const child = param.namedChild(i);
if (child && TYPE_NODE_TYPES.has(child.type)) {
rawTypeName = child.text?.trim() ?? null;
typeName = extractSimpleTypeName(child) ?? rawTypeName;
break;
}
// Also check type_identifier
if (child?.type === 'type_identifier') {
rawTypeName = child.text?.trim() ?? null;
typeName = rawTypeName;
break;
}
}
// Check for 'required' keyword:
// 1. Among children of the param node itself
let hasRequired = false;
for (let i = 0; i < param.childCount; i++) {
const child = param.child(i);
if (child && child.text.trim() === 'required') {
hasRequired = true;
break;
}
}
// 2. In tree-sitter-dart, `required` may be an anonymous sibling token
// immediately preceding the formal_parameter inside optional_formal_parameters.
if (!hasRequired) {
let prev = param.previousSibling;
// Skip comma separators
while (prev && !prev.isNamed && prev.text.trim() === ',') {
prev = prev.previousSibling;
}
if (prev && !prev.isNamed && prev.text.trim() === 'required') {
hasRequired = true;
}
}
// A parameter is optional if it's inside an optional_formal_parameters block
// and does NOT have the 'required' keyword
const isOptional = isOptionalBlock && !hasRequired;
return {
name,
type: typeName,
rawType: rawTypeName,
isOptional,
isVariadic: false, // Dart has no variadic params
};
}
/**
* Dart visibility: underscore prefix = private, else public.
*
* We resolve the name by descending into the inner signature.
*/
function extractDartVisibility(node: SyntaxNode): MethodVisibility {
const name = extractDartName(node);
if (!name) return 'public';
// Strip 'set ' or 'operator ' prefix to get the raw name
const rawName = name.startsWith('set ')
? name.slice(4)
: name.startsWith('operator ')
? name.slice(9)
: name;
return rawName.startsWith('_') ? 'private' : 'public';
}
/**
* In tree-sitter-dart, `static` is an anonymous child token of
* `method_signature` (or `declaration`), not a previous sibling.
*
* We check children first, then fall back to previous siblings for
* grammar variants.
*/
function isDartStatic(node: SyntaxNode): boolean {
// In tree-sitter-dart, `static` is an anonymous child token of method_signature
// (or declaration), not a previous sibling.
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (child && !child.isNamed && child.text.trim() === 'static') return true;
// Stop once we hit the inner signature — static always precedes it
if (child?.isNamed) break;
}
// Also check previous siblings (fallback for grammar variants)
let sibling = node.previousSibling;
while (sibling) {
if (sibling.isNamed && sibling.type !== 'annotation') break;
if (!sibling.isNamed && sibling.text.trim() === 'static') return true;
sibling = sibling.previousSibling;
}
return false;
}
/**
* A Dart method is abstract if it has no function_body sibling following it.
* In the tree-sitter grammar, function_body is a sibling of method_signature
* in class_body.
*/
function isDartAbstract(node: SyntaxNode, _ownerNode: SyntaxNode): boolean {
// `declaration` nodes in class_body represent abstract methods (no body, followed by ';').
// Note: extension bodies cannot have abstract members in Dart, but `declaration` nodes
// do not appear in extension_body in practice since extensions must provide implementations.
if (node.type === 'declaration') return true;
// For method_signature nodes, check if the next named sibling is a function_body
const next = node.nextNamedSibling;
return !next || next.type !== 'function_body';
}
/**
* Check for `async`, `async*`, or `sync*` keyword in the function_body sibling.
* The keyword appears as an unnamed child of function_body, or
* as a sibling keyword before function_body.
*
* Dart has three async-like forms: `async` (Future), `async*` (Stream), `sync*` (Iterable).
* All three are treated as async for graph purposes.
*/
function isDartAsync(node: SyntaxNode): boolean {
let sibling: SyntaxNode | null = node.nextSibling;
let limit = 3;
while (sibling && limit > 0) {
if (!sibling.isNamed) {
const text = sibling.text.trim();
if (text === 'async' || text === 'async*' || text === 'sync*') return true;
}
if (sibling.isNamed && sibling.type === 'function_body') {
// Check first child of function_body for async/async*/sync*
for (let i = 0; i < sibling.childCount; i++) {
const child = sibling.child(i);
if (child) {
const text = child.text.trim();
if (text === 'async' || text === 'async*' || text === 'sync*') return true;
}
// Stop at first substantial child
if (child?.isNamed) break;
}
break;
}
sibling = sibling.nextSibling;
limit--;
}
return false;
}
/**
* Extract annotations that appear as sibling nodes before the method_signature
* in class_body. Each annotation node is prefixed with '@'.
*/
function extractDartAnnotations(node: SyntaxNode): string[] {
const annotations: string[] = [];
let sibling = node.previousNamedSibling;
while (sibling && sibling.type === 'annotation') {
// annotation node text already includes '@', e.g. "@override"
const text = sibling.text?.trim();
if (text) {
// Normalize: strip arguments from annotation if present, keep just the name
// e.g. "@deprecated" -> "@deprecated", "@JsonKey(name: 'id')" -> "@JsonKey"
const match = text.match(/^@(\w+)/);
if (match) {
annotations.unshift('@' + match[1]);
} else {
annotations.unshift(text);
}
}
sibling = sibling.previousNamedSibling;
}
return annotations;
}
// ---------------------------------------------------------------------------
// Dart config
// ---------------------------------------------------------------------------
export const dartMethodConfig: MethodExtractionConfig = {
language: SupportedLanguages.Dart,
typeDeclarationNodes: ['class_definition', 'mixin_declaration', 'extension_declaration'],
methodNodeTypes: ['method_signature', 'declaration'],
bodyNodeTypes: ['class_body', 'extension_body'],
extractName: extractDartName,
extractReturnType: extractDartReturnType,
extractParameters: extractDartParameters,
extractVisibility: extractDartVisibility,
isStatic: isDartStatic,
isAbstract: isDartAbstract,
isFinal: () => false, // Dart methods cannot be 'final'
isAsync: isDartAsync,
extractAnnotations: extractDartAnnotations,
};
@@ -0,0 +1,195 @@
// gitnexus/src/core/ingestion/method-extractors/configs/go.ts
// Verified against tree-sitter-go 0.23.4
import { SupportedLanguages } from 'gitnexus-shared';
import type {
MethodExtractionConfig,
ParameterInfo,
MethodVisibility,
} from '../../method-types.js';
import { extractSimpleTypeName } from '../../type-extractors/shared.js';
import type { SyntaxNode } from '../../utils/ast-helpers.js';
// ---------------------------------------------------------------------------
// Go helpers
// ---------------------------------------------------------------------------
/**
* Extract the method/function name.
* - method_declaration: name is a `field_identifier`
* - function_declaration: name is an `identifier`
*/
function extractGoName(node: SyntaxNode): string | undefined {
const nameNode = node.childForFieldName('name');
return nameNode?.text;
}
/**
* Extract return type from the `result` field.
*
* Go supports single return (`int`) and multi-return (`(User, error)`).
* Multi-return appears as a `parameter_list` — extract the first type.
*/
function extractGoReturnType(node: SyntaxNode): string | undefined {
const result = node.childForFieldName('result');
if (!result) return undefined;
// Single return type (type_identifier, pointer_type, etc.)
if (result.type !== 'parameter_list') {
return result.text?.trim();
}
// Multi-return: (Type, error) — extract first parameter's type
for (let i = 0; i < result.namedChildCount; i++) {
const param = result.namedChild(i);
if (param?.type === 'parameter_declaration') {
const typeNode = param.childForFieldName('type');
if (typeNode) return typeNode.text?.trim();
}
}
return undefined;
}
/**
* Extract parameters from the `parameters` field.
*
* Go parameter_list contains parameter_declaration nodes with optional
* `name` and required `type` fields. Go allows multiple names for one type:
* `func(a, b int)` — each name shares the type.
*
* Handles variadic_parameter_declaration (`...string`).
*/
function extractGoParameters(node: SyntaxNode): ParameterInfo[] {
const paramList = node.childForFieldName('parameters');
if (!paramList) return [];
const params: ParameterInfo[] = [];
for (let i = 0; i < paramList.namedChildCount; i++) {
const param = paramList.namedChild(i);
if (!param) continue;
if (param.type === 'parameter_declaration') {
const typeNode = param.childForFieldName('type');
const typeName = typeNode
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null;
// Go allows multiple names for one type: func(a, b int)
const names: string[] = [];
for (let j = 0; j < param.namedChildCount; j++) {
const child = param.namedChild(j);
if (child?.type === 'identifier') {
names.push(child.text);
}
}
const rawType = typeNode?.text?.trim() ?? null;
if (names.length === 0) {
// Unnamed parameter: func(int, string)
params.push({
name: `_${i}`,
type: typeName,
rawType,
isOptional: false,
isVariadic: false,
});
} else {
for (const name of names) {
params.push({ name, type: typeName, rawType, isOptional: false, isVariadic: false });
}
}
} else if (param.type === 'variadic_parameter_declaration') {
const nameNode = param.childForFieldName('name');
const typeNode = param.childForFieldName('type');
const typeName = typeNode
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null;
params.push({
name: nameNode?.text ?? `_${i}`,
type: typeName,
rawType: typeNode?.text?.trim() ?? null,
isOptional: false,
isVariadic: true,
});
}
}
return params;
}
/**
* Go visibility: uppercase first character = exported (public), lowercase = unexported (private).
*/
function extractGoVisibility(node: SyntaxNode): MethodVisibility {
const name = extractGoName(node);
if (!name || name.length === 0) return 'private';
const first = name[0];
return first === first.toUpperCase() && first !== first.toLowerCase() ? 'public' : 'private';
}
/**
* Extract receiver type from the `receiver` field.
*
* The receiver is a parameter_list with one parameter_declaration:
* (r *Repo) → pointer_type → type_identifier "Repo"
* (r Repo) → type_identifier "Repo"
*/
function extractGoReceiverType(node: SyntaxNode): string | undefined {
const receiver = node.childForFieldName('receiver');
if (!receiver) return undefined;
for (let i = 0; i < receiver.namedChildCount; i++) {
const param = receiver.namedChild(i);
if (param?.type === 'parameter_declaration') {
const typeNode = param.childForFieldName('type');
if (!typeNode) continue;
// Unwrap pointer_type: *User → User
const inner = typeNode.type === 'pointer_type' ? typeNode.firstNamedChild : typeNode;
return inner?.text;
}
}
return undefined;
}
/**
* Resolve owner name from the receiver type.
* For function_declaration (no receiver), returns undefined.
*/
function extractGoOwnerName(node: SyntaxNode): string | undefined {
return extractGoReceiverType(node);
}
// ---------------------------------------------------------------------------
// Config
// ---------------------------------------------------------------------------
export const goMethodConfig: MethodExtractionConfig = {
language: SupportedLanguages.Go,
// Each method_declaration/function_declaration is treated as its own "container"
// for extractFromNode() — not used with extract() in the traditional sense.
// method_elem covers interface method signatures (abstract methods).
typeDeclarationNodes: ['method_declaration', 'function_declaration', 'method_elem'],
methodNodeTypes: ['method_declaration', 'function_declaration', 'method_elem'],
bodyNodeTypes: [],
extractName: extractGoName,
extractReturnType: extractGoReturnType,
extractParameters: extractGoParameters,
extractVisibility: extractGoVisibility,
extractReceiverType: extractGoReceiverType,
extractOwnerName: extractGoOwnerName,
isStatic(node) {
// Go functions (no receiver) are effectively static
return node.type === 'function_declaration';
},
isAbstract(node, _ownerNode) {
// Go interface method signatures (method_elem) are abstract — no body
return node.type === 'method_elem';
},
isFinal(_node) {
return false; // Go has no final methods
},
};
@@ -19,7 +19,9 @@ const INTERFACE_OWNER_TYPES = new Set(['interface_declaration', 'annotation_type
function extractReturnTypeFromField(node: SyntaxNode): string | undefined {
const typeNode = node.childForFieldName('type');
if (!typeNode) return undefined;
return extractSimpleTypeName(typeNode) ?? typeNode.text?.trim();
// Use .text to preserve full generic types (e.g. List<User>, Stream<T>)
// needed by the call resolver for return-type inference.
return typeNode.text?.trim();
}
function extractAnnotations(node: SyntaxNode, modifierType: string): string[] {
@@ -68,6 +70,7 @@ function extractJavaParameters(node: SyntaxNode): ParameterInfo[] {
params.push({
name: nameNode.text,
type: typeNode ? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim()) : null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: false,
isVariadic: false,
});
@@ -76,6 +79,7 @@ function extractJavaParameters(node: SyntaxNode): ParameterInfo[] {
// Varargs: type_identifier + "..." + variable_declarator
let paramName: string | undefined;
let paramType: string | null = null;
let paramRawType: string | null = null;
for (let j = 0; j < param.namedChildCount; j++) {
const c = param.namedChild(j);
if (!c) continue;
@@ -90,13 +94,15 @@ function extractJavaParameters(node: SyntaxNode): ParameterInfo[] {
c.type === 'floating_point_type' ||
c.type === 'boolean_type'
) {
paramType = extractSimpleTypeName(c) ?? c.text?.trim();
paramRawType = c.text?.trim() ?? null;
paramType = extractSimpleTypeName(c) ?? paramRawType;
}
}
if (paramName) {
params.push({
name: paramName,
type: paramType,
rawType: paramRawType,
isOptional: false,
isVariadic: true,
});
@@ -192,6 +198,7 @@ function extractKotlinParameters(node: SyntaxNode): ParameterInfo[] {
let paramName: string | undefined;
let paramType: string | null = null;
let paramRawType: string | null = null;
let hasDefault = false;
const isVariadic = nextIsVariadic;
nextIsVariadic = false;
@@ -206,7 +213,8 @@ function extractKotlinParameters(node: SyntaxNode): ParameterInfo[] {
part.type === 'nullable_type' ||
part.type === 'function_type'
) {
paramType = extractSimpleTypeName(part) ?? part.text?.trim();
paramRawType = part.text?.trim() ?? null;
paramType = extractSimpleTypeName(part) ?? paramRawType;
}
}
@@ -223,6 +231,7 @@ function extractKotlinParameters(node: SyntaxNode): ParameterInfo[] {
params.push({
name: paramName,
type: paramType,
rawType: paramRawType,
isOptional: hasDefault,
isVariadic: isVariadic,
});
@@ -252,7 +261,7 @@ function extractKotlinReturnType(node: SyntaxNode): string | undefined {
child.type === 'nullable_type' ||
child.type === 'function_type')
) {
return extractSimpleTypeName(child) ?? child.text?.trim();
return child.text?.trim();
}
if (child.type === 'function_body') break;
}
@@ -264,6 +273,7 @@ export const kotlinMethodConfig: MethodExtractionConfig = {
typeDeclarationNodes: ['class_declaration', 'object_declaration', 'companion_object'],
methodNodeTypes: ['function_declaration'],
bodyNodeTypes: ['class_body'],
staticOwnerTypes: new Set(['companion_object', 'object_declaration']),
extractName(node) {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
@@ -0,0 +1,326 @@
// gitnexus/src/core/ingestion/method-extractors/configs/php.ts
// Verified against tree-sitter-php 0.23.12
import { SupportedLanguages } from 'gitnexus-shared';
import type {
MethodExtractionConfig,
ParameterInfo,
MethodVisibility,
} from '../../method-types.js';
import { extractSimpleTypeName } from '../../type-extractors/shared.js';
import type { SyntaxNode } from '../../utils/ast-helpers.js';
// ---------------------------------------------------------------------------
// PHP helpers
// ---------------------------------------------------------------------------
/** Regex to extract PHPDoc @return annotations: `@return User` */
const PHPDOC_RETURN_RE = /@return\s+(\S+)/;
/** Node types to skip when walking backwards through siblings for PHPDoc. */
const PHPDOC_SKIP_NODE_TYPES: ReadonlySet<string> = new Set(['attribute_list', 'attribute']);
/**
* Normalize a PHPDoc return type for the MethodExtractor.
* Strips nullable prefix, null/false/void unions, namespace prefixes, and
* rejects uninformative types (mixed, void, self, static, object, array).
*/
function normalizePhpReturnType(raw: string): string | undefined {
let type = raw.startsWith('?') ? raw.slice(1) : raw;
const parts = type
.split('|')
.filter((p) => p !== 'null' && p !== 'false' && p !== 'void' && p !== 'mixed');
if (parts.length !== 1) return undefined;
type = parts[0];
const segments = type.split('\\');
type = segments[segments.length - 1];
if (
type === 'mixed' ||
type === 'void' ||
type === 'self' ||
type === 'static' ||
type === 'object' ||
type === 'array'
)
return undefined;
if (/^\w+(\[\])?$/.test(type) || /^\w+\s*</.test(type)) return type;
return undefined;
}
/**
* Walk backwards through preceding siblings of `node` to find a PHPDoc
* `@return Type` annotation. Skips `attribute_list` nodes (PHP 8 attributes).
*/
function extractPhpDocReturnType(node: SyntaxNode): string | undefined {
let sibling = node.previousSibling;
while (sibling) {
if (sibling.type === 'comment') {
const match = PHPDOC_RETURN_RE.exec(sibling.text);
if (match) return normalizePhpReturnType(match[1]);
} else if (sibling.isNamed && !PHPDOC_SKIP_NODE_TYPES.has(sibling.type)) {
break;
}
sibling = sibling.previousSibling;
}
return undefined;
}
const PHP_VIS = new Set<MethodVisibility>(['public', 'private', 'protected']);
/**
* Find the visibility keyword from a visibility_modifier named child.
* PHP tree-sitter emits `visibility_modifier` as a named node with text
* "public", "private", or "protected".
*/
function findPhpVisibility(node: SyntaxNode): MethodVisibility {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === 'visibility_modifier') {
const text = child.text.trim() as MethodVisibility;
if (PHP_VIS.has(text)) return text;
}
}
return 'public'; // PHP methods are public by default
}
/**
* Check for a named modifier child of a specific type.
* PHP tree-sitter uses distinct node types: abstract_modifier, final_modifier,
* static_modifier — rather than a wrapper `modifiers` node with keyword children.
*/
function hasModifierNode(node: SyntaxNode, modifierType: string): boolean {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === modifierType) return true;
}
return false;
}
/**
* Extract the return type from a PHP method_declaration node.
*
* In tree-sitter-php, the return type is not exposed via a named field.
* It appears as a type node (primitive_type, named_type, union_type,
* optional_type, nullable_type, intersection_type) after the formal_parameters
* and a `:` token separator.
*
* When the AST return type is missing or uninformative (`array` / `iterable`),
* falls back to parsing PHPDoc `@return Type` from preceding doc comments.
*/
function extractPhpReturnType(node: SyntaxNode): string | undefined {
const TYPE_NODE_TYPES = new Set([
'primitive_type',
'named_type',
'union_type',
'optional_type',
'nullable_type',
'intersection_type',
]);
let astType: string | undefined;
let seenParams = false;
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (!child) continue;
if (child.type === 'formal_parameters') {
seenParams = true;
continue;
}
// After the parameters node, look for the colon and then the type
if (seenParams && child.isNamed && TYPE_NODE_TYPES.has(child.type)) {
astType = child.text?.trim();
break;
}
// Stop at body or semicolon
if (child.type === 'compound_statement' || (!child.isNamed && child.text === ';')) {
break;
}
}
// If AST type is missing or uninformative, try PHPDoc @return fallback
if (!astType || astType === 'array' || astType === 'iterable') {
const docType = extractPhpDocReturnType(node);
if (docType) return docType;
}
return astType;
}
/**
* Extract parameters from a PHP method_declaration node.
*
* PHP parameter types in tree-sitter-php:
* - `simple_parameter`: regular parameter with optional type and default
* - `variadic_parameter`: `...$param` with optional type
* - `property_promotion_parameter`: constructor promotion `private string $name`
* (may also be variadic via an ERROR node containing `...`)
*/
function extractPhpParameters(node: SyntaxNode): ParameterInfo[] {
const paramList = node.childForFieldName('parameters');
if (!paramList) return [];
const params: ParameterInfo[] = [];
for (let i = 0; i < paramList.namedChildCount; i++) {
const param = paramList.namedChild(i);
if (!param) continue;
if (param.type === 'simple_parameter') {
const nameNode = param.childForFieldName('name');
if (!nameNode) continue;
const typeNode = param.childForFieldName('type');
const typeName = typeNode ? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim()) : null;
// Detect optional: '=' token among children indicates a default value
let isOptional = false;
for (let j = 0; j < param.childCount; j++) {
const c = param.child(j);
if (c && !c.isNamed && c.text === '=') {
isOptional = true;
break;
}
}
params.push({
name: stripDollar(nameNode.text),
type: typeName ?? null,
rawType: typeNode?.text?.trim() ?? null,
isOptional,
isVariadic: false,
});
} else if (param.type === 'variadic_parameter') {
const nameNode = param.childForFieldName('name');
if (!nameNode) continue;
const typeNode = param.childForFieldName('type');
const typeName = typeNode ? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim()) : null;
params.push({
name: stripDollar(nameNode.text),
type: typeName ?? null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: false,
isVariadic: true,
});
} else if (param.type === 'property_promotion_parameter') {
const nameNode = param.childForFieldName('name');
if (!nameNode) continue;
const typeNode = param.childForFieldName('type');
const typeName = typeNode ? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim()) : null;
// Detect variadic: an ERROR child containing "..." indicates variadic promotion
let isVariadic = false;
for (let j = 0; j < param.childCount; j++) {
const c = param.child(j);
if (c && (c.text === '...' || (c.type === 'ERROR' && c.text === '...'))) {
isVariadic = true;
break;
}
}
params.push({
name: stripDollar(nameNode.text),
type: typeName ?? null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: false,
isVariadic,
});
}
}
return params;
}
/** Strip leading $ from PHP variable names. */
function stripDollar(name: string): string {
return name.startsWith('$') ? name.slice(1) : name;
}
/**
* Extract PHP 8 attributes (#[...]) from a method_declaration node.
*
* AST structure: attribute_list → attribute_group → attribute → name child.
* Names are prefixed with '#' to distinguish from Java/Kotlin @ annotations.
*/
function extractPhpAnnotations(node: SyntaxNode): string[] {
const annotations: string[] = [];
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (!child || child.type !== 'attribute_list') continue;
for (let j = 0; j < child.namedChildCount; j++) {
const group = child.namedChild(j);
if (!group || group.type !== 'attribute_group') continue;
for (let k = 0; k < group.namedChildCount; k++) {
const attr = group.namedChild(k);
if (!attr || attr.type !== 'attribute') continue;
const nameNode = attr.firstNamedChild;
if (nameNode && nameNode.type === 'name') {
annotations.push('#' + nameNode.text);
}
}
}
}
return annotations;
}
// ---------------------------------------------------------------------------
// PHP config
// ---------------------------------------------------------------------------
export const phpMethodConfig: MethodExtractionConfig = {
language: SupportedLanguages.PHP,
typeDeclarationNodes: [
'class_declaration',
'interface_declaration',
'trait_declaration',
'enum_declaration',
],
methodNodeTypes: ['method_declaration', 'function_definition'],
bodyNodeTypes: ['declaration_list'],
extractName(node) {
return node.childForFieldName('name')?.text;
},
extractReturnType: extractPhpReturnType,
extractParameters: extractPhpParameters,
extractVisibility: findPhpVisibility,
isStatic(node) {
return hasModifierNode(node, 'static_modifier');
},
isAbstract(node, ownerNode) {
if (hasModifierNode(node, 'abstract_modifier')) return true;
// Interface methods are implicitly abstract when they have no body.
// Check ownerNode first, then fall back to walking the parent chain
// (needed when called from extractFromNode where ownerNode === node).
let isInterface = ownerNode.type === 'interface_declaration';
if (!isInterface) {
let p = node.parent;
while (p) {
if (p.type === 'interface_declaration') {
isInterface = true;
break;
}
p = p.parent;
}
}
if (isInterface) {
const body = node.childForFieldName('body');
if (body) return false;
for (let i = 0; i < node.namedChildCount; i++) {
if (node.namedChild(i)?.type === 'compound_statement') return false;
}
return true;
}
return false;
},
isFinal(node) {
return hasModifierNode(node, 'final_modifier');
},
extractAnnotations: extractPhpAnnotations,
};

Some files were not shown because too many files have changed in this diff Show More