Compare commits

...
Author SHA1 Message Date
Gergo Magyar e898e19714 refactor: rename FileAllScopeBindings to FileScopeBindings and update related references
- Updated the interface name from FileAllScopeBindings to FileScopeBindings to better reflect its purpose.
- Adjusted all occurrences of the renamed interface across multiple files including parsing-processor.ts, pipeline.ts, parse-worker.ts, and type-env.ts.
- Enhanced comments and documentation to clarify the narrowing of scope bindings and the rationale behind the changes.
- Improved error handling and validation in the pipeline for file-scope bindings.
- Added integration tests to ensure the correct behavior of the BindingAccumulator and its interaction with the TypeEnv flush process.
2026-04-09 19:43:59 +01:00
Gergo Magyar 448a4b22a6 fix(SM-14): address PR #743 post-fix review findings
Four items from the deep review on commit d3c25d20 — three Low
findings plus one informational note.

Low #1 — Expose `get disposed(): boolean` for API symmetry
- BindingAccumulator's `_disposed` field was set but never read or
  exposed. `_finalized` had a public getter (`get finalized()`) but
  `_disposed` did not. Added the matching `get disposed()` getter so
  debug tooling and future Phase 9 consumers can detect a disposed
  accumulator without inspecting empty state heuristically.
- JSDoc notes that disposal and finalization are orthogonal lifecycle
  dimensions — a disposed accumulator may or may not be finalized.

Low #2 — Test for Tier 0 "don't overwrite" protection
- Production enrichment loop at pipeline.ts:1104-1108 has a priority
  guard:
    if (!fileExports.has(name)) { fileExports.set(name, type); }
  preventing a worker-path binding from clobbering a higher-quality
  Tier 0 SymbolTable entry. The existing `runEnrichmentLoop` test
  helper in binding-accumulator.test.ts was missing this guard, and
  no test exercised the priority branch.
- Fixed the helper to mirror the production guard.
- Added a new test: "does not overwrite existing SymbolTable entry
  (Tier 0 priority)" — pre-populates exportedTypeMap with an
  "SymbolTableAuthoritativeType" entry, runs the enrichment loop
  against an accumulator with "WorkerInferredType" for the same name,
  asserts the authoritative type survives.

Low #3 — Move finalize() to before the enrichment loop
- Previously, finalize() was called at pipeline.ts:1715 (inside
  runPipelineFromRepo, AFTER runChunkedParseAndResolve had already
  returned). The enrichment loop at pipeline.ts:1087 (inside
  runChunkedParseAndResolve) consumed the still-mutable accumulator.
  The `finalized` state was therefore not a reliable "all reads are
  done" signal — it was a "no more writes" signal that arrived later
  than the actual last read.
- Moved finalize() to immediately before the enrichment loop at line
  1087. By that point all worker-path appends (line 934) and all
  sequential-path flushes (via processCalls at line 1051, also inside
  runChunkedParseAndResolve) have completed. Grep confirmed no further
  `bindingAccumulator.appendFile` calls exist outside runChunkedParseAndResolve.
- Lifecycle contract is now explicit:
    append phase → finalize → consume → dispose
- Replaced the old finalize() call at line 1715 with an explanatory
  comment pointing to the new seam.

Informational — parsing-processor.ts TypeEnv clarification
- parsing-processor.ts builds a FieldExtractor-only TypeEnv that is
  intentionally NOT flushed into the accumulator — the accumulator
  feed happens later in call-processor.ts via its own flush() call.
  A future reader might see `buildTypeEnv()` here and try to add a
  flush call, double-counting entries and tripping the single-use
  invariant.
- Added a multi-line comment explaining the ownership rule and
  cross-referencing PR #743 and plan 2026-04-09-005.

Verification
- `tsc --noEmit` clean
- 3110 unit tests pass (+1 new Tier 0 priority test)
- 1766 resolver integration tests pass — critically, the finalize()
  relocation did not regress any real pipeline path, proving all
  writes complete before the new finalize point
- Zero regressions

Plan: docs/plans/2026-04-09-005-fix-sm14-sequential-path-memory-regression-plan.md
Review: https://github.com/abhigyanpatwari/GitNexus/pull/743#issuecomment-4216262583
2026-04-09 18:52:51 +01:00
Gergo Magyar d3c25d2093 fix(SM-14): close sequential-path memory regression (Codex adversarial review)
Addresses the medium-severity finding from Codex's adversarial review of
commit 803631fe: the sequential path's `typeEnv.flush()` was still
writing every scope (file + function) into the BindingAccumulator, and
the accumulator stayed alive through Phase 14 and runGraphAnalysisPhases
with no reader. On fallback runs (workers disabled or unavailable),
large repos accumulated heap for nothing.

Applies BOTH Codex remediations — narrowing AND disposal:

R1 — Narrow typeEnv.flush() to file-scope only
- The sequential path now mirrors the worker-path narrowing from
  commit 803631fe. `flush()` iterates only `env.get(FILE_SCOPE)`
  instead of the nested `for (scope, scopeMap) of env` loop, writing
  entries with `scope: ''` hardcoded. Function-scope bindings never
  reach the accumulator from either execution path until a Phase 9
  consumer lands.
- The earlier rationale for keeping sequential-path full-scope data
  ("preserve a Phase 9 prototyping sample") did not survive the
  Codex challenge — Phase 9 authors use synthetic fixtures, and a
  live-repo sample from the sequential-only path isn't representative
  of production worker-dominant runs.
- Phase 9 reversion path documented inline at the flush() seam.

R2 — Add BindingAccumulator.dispose()
- New public method clears `_allByFile`, `_fileScopeByFile`, and
  explicitly resets `_totalBindings = 0` (feasibility review caught
  that clearing the maps alone would leave the `totalBindings` getter
  reporting stale counts). Idempotent and orthogonal to finalize() —
  calling dispose() doesn't change the finalized state.
- Post-dispose contract: all read methods return empty/undefined
  state matching a never-appended accumulator. Documented in class
  JSDoc with the lifecycle sequence.
- Used before finalize(): accumulator behaves like a fresh one,
  appends still succeed.
- Used after finalize(): reads return empty but appends still throw
  the existing "finalized" error.

R3 — Wire dispose() into the pipeline
- Inserted immediately after the dev telemetry log at pipeline.ts
  line ~1723, before `runCrossFileBindingPropagation` (Phase 14) and
  `runGraphAnalysisPhases`. Sequence is:
    enrichment loop → finalize → telemetry → dispose → Phase 14
  The telemetry log captures peak state before disposal, then the
  heap footprint is released for the long tail of graph analysis.
- Verified both runCrossFileBindingPropagation and
  runGraphAnalysisPhases signatures do NOT take a bindingAccumulator
  parameter — grep confirmed the last usage is at line 1720.

Tests (+6 scenarios)
- test/unit/type-env.test.ts: existing "flushes function-scoped
  bindings into accumulator" test was rewritten as a negative
  assertion ("does NOT flush function-scoped bindings, narrowed per
  PR #743 Codex review"). Plus a new "narrows mixed file-scope and
  function-scope env to file-scope only" test that builds a
  realistic TypeScript file with both scopes and asserts only the
  file-scope entry lands in the accumulator. This is the R1 red/
  green signal — both tests were written test-first and failed
  against the pre-narrowing flush() body.
- test/unit/binding-accumulator.test.ts: new `describe('dispose', ...)`
  block with 5 scenarios — empty all read methods, idempotency, pre-
  finalize behavior, post-finalize behavior, and
  `estimateMemoryBytes() === 0` guard.

Verification
- `tsc --noEmit` clean
- 3109 unit tests pass (+6 net: +3 narrowing tests — 2 new + 1
  rewritten — plus +5 dispose scenarios − 1 pre-existing test
  replaced = net +6)
- 1766 resolver integration tests pass
- Zero regressions

Plan: docs/plans/2026-04-09-005-fix-sm14-sequential-path-memory-regression-plan.md
Codex review: branch diff against main, verdict needs-attention
Previous commit: 803631fe (worker-path narrowing)
2026-04-09 18:39:37 +01:00
Gergo Magyar 803631fef7 fix(SM-14): address PR #743 BindingAccumulator review findings
Addresses the 5 findings from PR #743 review comment 4211636245.

Critical (R1) — Strip function-scope bindings from worker IPC
- parse-worker.ts previously serialized typeEnv.allScopes() over the
  worker IPC boundary on every batch, pushing ~4.9 MB of function-scope
  bindings (e.g. `handleRequest@15 → db: Database`) into the accumulator
  with zero downstream consumers. The only reader is the ExportedTypeMap
  enrichment loop in pipeline.ts, which calls fileScopeEntries() —
  the `scope = ''` subset only.
- Narrowed parse-worker.ts to use typeEnv.fileScope() and emit
  [varName, typeName] pairs. FileAllScopeBindings.bindings type narrowed
  from [string, string, string][] to [string, string][].
- pipeline.ts adapter updated to unpack the new two-element tuples and
  construct BindingEntry with scope: '' hardcoded.
- Sequential path (call-processor.ts → typeEnv.flush()) is UNCHANGED
  and still writes all scopes — preserves a working sample of higher-
  quality bindings for Phase 9 prototyping without IPC cost.
- Phase 9 reversion path documented inline on both FileAllScopeBindings
  and the pipeline adapter: change fileScope() → allScopes(), widen the
  tuple back to 3 elements, done. Field name `allScopeBindings` kept to
  keep that revert mechanically trivial.

Medium #1 (R2) — Worker-vs-sequential quality asymmetry
- BindingAccumulator class JSDoc now documents that entries are NOT
  homogeneous in resolution quality: sequential path has SymbolTable
  + importedBindings access (Tier 2 cross-file propagation); worker
  path has Tier 0 + local Tier 1 only. Phase 9 consumers that trust
  every entry equally will silently produce worse results for large
  repos (worker path dominant) than for small ones.

Medium #2 (R3) — Integration test for ExportedTypeMap enrichment
- Added a 4-scenario test suite mocking KnowledgeGraph nodes and
  running the exact enrichment loop from pipeline.ts:1082-1110 inline:
    (a) Exported Function node → enriched
    (b) Non-exported Variable → filtered out (isExported gate)
    (c) Exported Const node → enriched
    (d) No matching graph node → silently skipped via continue
- Locks in the `{Label}:{filePath}:{name}` node-ID format contract.
  If the ID format drifts for any language, this test fires.

Low #1 (R4) — Storage split for O(n_file_scope) reads
- BindingAccumulator now stores two parallel maps:
    _allByFile:       Map<string, BindingEntry[]>     — full entry list
    _fileScopeByFile: Map<string, [string, string][]> — scope='' fast path
- appendFile iterates input once, populates both maps synchronously.
  fileScopeEntries becomes O(1) map lookup + O(n_file_scope) return —
  no longer walks function-scope entries to filter.
- 4 new tests: mixed scopes correctness, only-function-scope file still
  visible via files()/fileCount, multiple appends accumulate
  consistently, 1001-entry performance guard.

Low #2 (R5) — Duplicate iteration logic
- Resolved as a side effect of R1: after the worker narrowing, the
  parse-worker loop iterates `fileScope()` (flat map) and
  typeEnv.flush() iterates `allScopes()` (nested map). Different data
  shapes — no common helper to extract.

Swift CI gap (R7)
- Acknowledged as out of scope. Not SM-14 specific — 95 Swift tests
  skipped across the broader test file.

Verification
- `tsc --noEmit` clean
- 3103 unit tests pass (+10 new scenarios in binding-accumulator.test.ts)
- 1766 resolver integration tests pass
- Zero regressions

Plan: docs/plans/2026-04-09-004-fix-sm14-binding-accumulator-review-findings-plan.md
Review: https://github.com/abhigyanpatwari/GitNexus/pull/743#issuecomment-4211636245
2026-04-09 18:05:20 +01:00
Gergo Magyar 47edb754d1 Merge remote-tracking branch 'origin/main' into sm14-binding-accumulator 2026-04-09 17:42:15 +01:00
d09078925e Extract resolveFreeCall from resolveCallTarget (SM-13) (#756)
* Initial plan

* feat(SM-13): extract resolveFreeCall from resolveCallTarget

Extract the free-function call resolution path into a dedicated
`resolveFreeCall(calledName, filePath, ctx)` function that uses
`lookupExact` + import-scoped resolution via `ctx.resolve()`.

- Free function calls (foo()) now route through `resolveFreeCall`
- Swift/Kotlin implicit constructors (User()) delegate to
  `resolveStaticCall` within `resolveFreeCall`
- `resolveCallTarget` dispatches `callForm === 'free'` early,
  removing the inline freeFormHasClassTarget logic
- S0 block simplified to only handle `callForm === 'constructor'`
- Global (Tier 3) fallthrough preserved via ctx.resolve() until Phase 5
- 9 new unit tests for resolveFreeCall
- All 163 unit tests pass, all 1199 integration resolver tests pass

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c5f2e73a-259a-438c-b5c8-286b82e3c215

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore: revert unrelated package-lock.json change

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c5f2e73a-259a-438c-b5c8-286b82e3c215

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(SM-13): address PR #756 review findings on resolveFreeCall

Addresses all 7 findings from the PR #756 review comment.

Code (R1, finding #1)
- Replace the literal `'Class' | 'Struct' | 'Record'` check in
  `hasClassTarget` with `INSTANTIABLE_CLASS_TYPES.has(c.type)`. Converts
  an invariant that was previously comment-enforced ("keep this list
  aligned with INSTANTIABLE_CLASS_TYPES") into one enforced structurally.
  Any future extension of the set propagates here automatically. The
  narrower Swift extension dedup block below still uses literal
  `'Class' | 'Struct'` by design — Swift extensions only produce Class
  duplicates in practice, Record is deliberately excluded there, and
  the inline comment now documents that asymmetry.

Tests (+12 regression scenarios)

Finding #2 — language coverage
- Go free function (doStuff())
- Python free function (def helper(): ... helper())
- Rust free function outside any impl block
- Java statically-imported function
- JavaScript module-level function
Each exercises `_resolveCallTargetForTesting` with `callForm='free'`
and the language-specific file extension. `resolveFreeCall` has no
file-extension branching, so these guard the dispatch chain per
language without assuming extractor-specific symbol shapes.

Finding #3 — argCount threading
- 2-arg overload selected when argCount=2
- 0-arg overload selected when argCount=0

Finding #5 — Tier 3 (global) resolution
- Function globally visible but not imported. Asserts exact
  `TIER_CONFIDENCE.global === 0.5` and `reason === 'global'` to catch
  silent drift if the tier table is ever refactored.

Finding #6 — preComputedArgTypes worker path
- String overload matched via preComputedArgTypes=['String']
- Int overload matched via preComputedArgTypes=['int'] (lowercase,
  mirroring the parse-worker's inferred-literal shape; stored 'Int' is
  normalized via normalizeJvmTypeName at comparison time)

Finding #7 — Enum null-route documentation
- Enum-only free call asserts `toBeNull()` with an explanatory comment
  linking to the INSTANTIABLE_CLASS_TYPES rationale. NOT marked skipped
  — current behavior is intentional, not broken.

Finding #4 — Swift extension dedup guard
- Two same-name Class entries at different path lengths; exercises the
  full dispatch chain:
    1. filterCallableCandidates with 'free' strips Class → length 0
    2. hasClassTarget triggers resolveStaticCall
    3. Homonym ambiguity null-routes per SM-12 round-1 contract
    4. Constructor-form retry repopulates with both Classes
    5. Dedup block sorts by filePath.length → shortest path wins

Verification
- `tsc --noEmit` clean
- 3064 unit tests pass (+12)
- 1766 integration tests pass
- Zero regressions

Plan: docs/plans/2026-04-09-003-fix-sm13-resolve-free-call-review-findings-plan.md
Review: https://github.com/abhigyanpatwari/GitNexus/pull/756#issuecomment-4213879002

* refactor(SM-13): extract dedupSwiftExtensionCandidates shared helper

Follow-up to the PR #756 review fix. SM-13 duplicated the Swift
extension same-name collision dedup block between `resolveCallTarget`
and `resolveFreeCall` — two copies of identical 15-line logic with the
same heuristic (`filePath.length` sort, Class/Struct-only, `length > 1`
guard). Extract a single shared helper so the two sites cannot drift.

Changes
- New `dedupSwiftExtensionCandidates(candidates, tier)` helper defined
  alongside `tryOverloadDisambiguation`, with JSDoc documenting:
  - The Swift extension scenario it addresses
  - Why it is intentionally narrower than INSTANTIABLE_CLASS_TYPES
    (Class/Struct only, not Record — C#/Kotlin records don't exhibit
    the multi-file definition pattern, widening risks accidental
    dedup of legitimately distinct record types)
  - The return-null-on-no-match contract so callers can fall through
- `resolveCallTarget` tail dedup (was lines 1593-1610): replaced with
  a single `dedupSwiftExtensionCandidates` call
- `resolveFreeCall` tail dedup (was lines 1994-2012): same replacement
- Net line count: -32 insertions, -9 deletions in the consumer sites,
  +36 for the shared helper + JSDoc

Verification
- `tsc --noEmit` clean
- 3064 unit tests pass (including the R7 Swift dedup guard test added
  in the previous commit that exercises the full free-form retry
  chain through this helper)
- 1766 integration tests pass
- Zero regressions

Follows-up on: https://github.com/abhigyanpatwari/GitNexus/pull/756

* docs(SM-13): address PR #756 final review — comment cleanup only

Three documentation-only findings from the approval review. No
behavior change, no new tests, no code path modifications.

Finding #1 — stale line-number comment
- The comment inside `resolveFreeCall` at the `hasClassTarget` site
  referenced "lines ~1994-2008" for the Swift extension dedup block.
  Those lines were the inlined pre-SM-13 version; the block has since
  been extracted to `dedupSwiftExtensionCandidates`. Replaced the line
  reference with the helper name so future readers don't chase dead
  line numbers.

Finding #2 — fuzzy-widening asymmetry undocumented
- `resolveFreeCall` intentionally has no `widenCache` parameter and no
  D2 fuzzy-widening pass (unlike `resolveCallTarget`'s member-call
  path). Added an explicit "Asymmetry vs `resolveCallTarget`" paragraph
  to the JSDoc so a caller comparing the two signatures knows the
  skipped pass is deliberate and tied to Phase 5.

Finding #3 — constructor-form retry reasons undocumented
- `resolveStaticCall` can return null for three distinct reasons
  (empty instantiable pool, homonym ambiguity, ownerless Constructor
  nodes). The retry below it unconditionally re-filters with
  `'constructor'` form, which is correct for all three but not
  obvious. Added a structured three-case comment enumerating each
  reason and linking (a) to the SM-12 null-route contract, (b) to
  the R7 dedup test, and (c) to the currently-uncovered ownerless-
  Constructor path (noted as a future test candidate).

Verification
- `tsc --noEmit` clean
- 175 `resolveFreeCall` + `resolveStaticCall` + sibling tests pass
  (sanity check — no behavior change expected)
- No regressions

Follows-up on: https://github.com/abhigyanpatwari/GitNexus/pull/756#issuecomment-4215739052

* test(SM-13): cover ownerless-Constructor retry + PHP free function

Two low-severity test gaps from PR #756 review comment 4215739052 —
previously addressed doc-only, now have concrete test coverage.

Finding #3 low — ownerless-Constructor retry path (previously comment-only)
- The retry after resolveStaticCall returns null handles three distinct
  null-return reasons. Cases (a) and (b) were already tested (Interface/
  Trait null-route from SM-12, Swift shadowing dedup from R7). Case (c) —
  resolveStaticCall step-4 bailout when the tiered pool contains
  ownerless Constructor nodes — was only covered by a comment.
- New test: Class + ownerless Constructor in tiered pool, callForm='free'.
  Exercises the full chain:
    1. resolveStaticCall step 3 walks classCandidates via
       lookupMethodByOwner — ownerless Constructor not in methodByOwner,
       nothing found.
    2. Step 4 detects Constructor in tiered pool, bails with null.
    3. resolveFreeCall retry re-runs filterCallableCandidates with
       'constructor' form, which prefers Constructor over Class per
       CONSTRUCTOR_TARGET_TYPES ordering.
    4. Single survivor returned.
- Asserts the Constructor node (not the Class) is the resolved target.

Low — PHP free function coverage gap
- The language coverage table in the same review flagged PHP free
  functions (top-level `function helper()` outside any class) as
  uncovered. Added a test mirroring the existing Go/Python/Rust/Java/
  JS language tests — exercises the `.php` dispatch path for free
  calls. Ruby and C/C++ remain uncovered; deferred to a future round
  since those languages also have other gaps in the broader test file.

Verification
- `tsc --noEmit` clean
- 3066 unit tests pass (+2 new regression tests)
- 1766 integration tests pass
- Zero regressions

Follows-up on: https://github.com/abhigyanpatwari/GitNexus/pull/756#issuecomment-4215739052

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-09 17:41:28 +01:00
JaysonAlbertandgfwangjie 338cb01ee0 [codex] fix large repository graph loading (#732)
* fix(web): stream large graph responses

* fix(server): harden graph streaming

* fix(ci): stabilize graph loading coverage

---------

Co-authored-by: gfwangjie <gfwangjie@gf.com.cn>
2026-04-09 17:40:24 +01:00
4450a14b98 feat(SM-12): Extract resolveStaticCall from resolveCallTarget (#754)
* Initial plan

* feat(SM-12): extract resolveStaticCall from resolveCallTarget

- Add resolveStaticCall(className, methodName, currentFile, ctx, argCount?) using
  lookupClassByName + lookupMethodByOwner for O(1) constructor/static resolution
- Add S0 fast path in resolveCallTarget for constructor/free-form class calls
- Export resolveStaticCall from call-processor.ts
- Add 11 unit tests covering constructor resolution, confidence tiers,
  arity disambiguation, and resolveCallTarget delegation

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c9471ca9-57ff-4dae-956e-e7ffdc326bc4

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore: revert unrelated package-lock.json change

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c9471ca9-57ff-4dae-956e-e7ffdc326bc4

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor: shorten verbose test name per code review feedback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c9471ca9-57ff-4dae-956e-e7ffdc326bc4

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(SM-12): address PR #754 review findings

Addresses Claude's review comments on PR #754:

Performance
- Pass pre-computed `tiered` result into `resolveStaticCall` as optional
  `tieredOverride` parameter, eliminating the duplicate `ctx.resolve(className,
  currentFile)` on every constructor call path.
- Cache `freeFormHasClassTarget` in `resolveCallTarget` so the S0 fast path
  and the free-form constructor retry share a single `.some()` scan.

Architecture
- Reconcile `CLASS_LIKE_TYPES` (call-processor) with `CLASS_TYPES`
  (symbol-table): `CLASS_LIKE_TYPES = [...CLASS_TYPES, 'Impl']`. This makes
  the relationship explicit — the call resolver's set is a strict superset
  of the heritage-index set, guaranteeing anything reachable via
  `lookupClassByName` also passes the resolver filter. Trait is now included
  (harmless: traits have no Constructor nodes, so step-3 returns undefined
  and step-5 still returns the class-like node when unique). Documented
  the Interface inclusion rationale (static methods + MRO walker).
- Collapse `resolveStaticCall`'s `methodName` parameter into `className` —
  all call sites passed identical values. Named constructors (Dart
  `User.fromJson()`) arrive as member calls and go through
  `resolveMemberCall`. Documented the reserved path for when a language
  surfaces a static-method-shaped call with a distinct member name.
- Document the known gap: `callForm === 'member'` constructor patterns
  (e.g. Python `models.User()`) are handled by the tail fallback, not S0.

Tests
- Add tiered-override test asserting `ctx.resolve` is not re-invoked when
  a pre-computed result is passed in.
- Add language-specific `_resolveCallTargetForTesting` integration tests
  for Java (`new User()`), Python (`User()`), and Kotlin (`User()`).

Verification: 3031 unit + 1766 integration tests pass, zero regressions.

* fix(SM-12): restrict resolveStaticCall fallback to instantiable kinds

Addresses the high-severity finding from the Codex adversarial review of
PR #754: `resolveStaticCall`'s step-5 "return the class itself when no
Constructor node is found" fallback reused `CLASS_LIKE_TYPES`, which —
after SM-11 and PR #754's reconciliation — now includes `Interface`,
`Trait`, and `Impl`. That is the method-dispatch set, not the
instantiable set, so constructor-shaped calls could resolve to
non-instantiable nodes and emit false `CALLS` edges.

Concrete failure: Rust same-file `impl User { ... }` alongside
`struct User { ... }` — both land at same-file tier, the Impl is not
filtered out, and the step-5 fallback produces a `CALLS` edge to the
`Impl` block instead of the `Struct`. The same widening exposed
Interface / Trait targets in Java / C# / PHP / Scala.

Fix
- Introduce `INSTANTIABLE_CLASS_TYPES = {'Class', 'Struct', 'Record'}`
  as a sibling to `CLASS_LIKE_TYPES`, documenting the contract
  explicitly and cross-referencing `CONSTRUCTOR_TARGET_TYPES`.
- Update `CLASS_LIKE_TYPES` JSDoc to clarify it is the method-dispatch
  set and add an anti-pattern warning against reusing it for
  constructor-fallback filtering.
- Tighten `resolveStaticCall` step 5: filter `classCandidates` through
  `INSTANTIABLE_CLASS_TYPES` before the `length === 1` check. This
  strips `Impl` from the Rust shadowing scenario (leaving `Struct` as
  the sole instantiable target) and null-routes Interface / Trait /
  `Impl`-alone scenarios, matching the SM-10 R3 null-route precedent.
- Step 3 (explicit Constructor lookup via `lookupMethodByOwner`) is
  intentionally unchanged — its `def.type === 'Constructor'` check is
  the correct contract, and legitimate Constructor nodes attached to
  `Impl` owners still resolve correctly.

Tests (+10 regression scenarios)
- Positive guards: Struct, Record fallback paths.
- Null-route: Interface (Java/C#/TS), PHP Trait, Rust Trait.
- Rust same-file shadowing: Struct wins over Impl.
- Rust Impl-alone: null-routes (no Struct present).
- Step-3 preservation: Constructor owned by Impl still resolves to the
  Constructor node, proving step-5 tightening doesn't leak into step 3.
- Full cascade via `_resolveCallTargetForTesting` for Interface and
  Trait — confirms no downstream path silently re-introduces the edge.

Verification
- `tsc --noEmit` clean
- 3041 unit tests pass (+10)
- 1766 integration tests pass
- Zero regressions

Plan: docs/plans/2026-04-09-002-fix-sm12-constructor-fallback-instantiable-only-plan.md
Codex review job: review-mnrao7fr-nv9y0e

* fix(SM-12): address PR #754 second review round

Addresses the 9 findings from the follow-up review on PR #754.

Performance
- Align `freeFormHasClassTarget` with `INSTANTIABLE_CLASS_TYPES`: drop
  `Enum` (S0 would always return null for it — wasted lookup work) and
  add `Record` (C# records and Kotlin data classes were bypassing S0
  entirely). The trigger set and the fallback filter set now agree by
  construction, documented inline.

Documentation
- Remove stale single-line JSDoc on `CLASS_LIKE_TYPES` (line 57) that
  duplicated the full multi-line block immediately below it — tooling
  picks up the first block so the old one-liner was shadowing the
  current explanation.
- Rewrite the `resolveStaticCall` JSDoc step list to match the actual
  step boundaries in the implementation (steps 3, 4, 5 were blurred in
  the old description).
- Add inline comment on step 3 documenting the same-name lookup
  assumption (`${candidate.nodeId}\0${className}`) and the symmetric
  miss case for Python `__init__`-style constructors.
- Add inline comment on step 4 documenting that it also catches the
  ambiguous-step-3 case, and warning against removing the check
  without handling that path explicitly.
- Add inline comment on step 5 enumerating the three length outcomes
  (0 / 1 / >1) so future readers see the dominant null-route case.
- Document Ruby `User.new` as a known gap alongside Python
  `models.User()` in the S0 header comment.

Tests (+2 scenarios)
- Record free-form constructor call via `_resolveCallTargetForTesting`
  exercises the aligned `freeFormHasClassTarget` trigger end-to-end,
  closing the gap where the direct `resolveStaticCall` test passed
  but the integration path was silently bypassing S0.
- Arity threading via `_resolveCallTargetForTesting` asserts that
  `call.argCount` flows through resolveCallTarget → S0 →
  resolveStaticCall → lookupMethodByOwner, catching any future
  regression where the argCount is dropped at the S0 call site.

Verification
- `tsc --noEmit` clean
- 3043 unit tests pass (+2)
- 1766 integration tests pass
- Zero regressions

Plan: docs/plans/2026-04-09-002-fix-sm12-constructor-fallback-instantiable-only-plan.md
Review: https://github.com/abhigyanpatwari/GitNexus/pull/754#issuecomment-4213536094

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-09 12:09:07 +01:00
Kunal Hemnani 3f0b8c1a5b fix(ingestion): replace lookupExact with lookupExactAll in named-binding-processor (#755) 2026-04-09 11:57:36 +01:00
bb68cc1eb0 Extract resolveMemberCall from resolveCallTarget (SM-11) (#744)
* Initial plan

* feat(SM-11): extract resolveMemberCall from resolveCallTarget

- Create resolveMemberCall(ownerType, methodName, currentFile, ctx, heritageMap?)
  that uses owner-scoped + MRO resolution only (no fuzzy lookup)
- resolveCallTarget delegates member calls (D0 path) to resolveMemberCall
- walkMixedChain uses resolveMemberCall for owner-scoped member-call resolution
- Add 7 unit tests for resolveMemberCall covering direct, inherited, MRO,
  null cases, and confidence tier assertions
- Export resolveMemberCall for external use

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/3b7889a9-5f2f-4572-8904-45084210f10d

* fix(SM-11): address PR #744 review

Blocking fixes:

- B1: Revert unrelated package-lock.json gitnexus-shared addition

- B2: Document confidence-tier semantic change on resolveMemberCall

Performance / coupling fixes:

- S1: walkMixedChain now calls resolveMethodByOwner directly (hot path) to avoid throwaway ResolveResult allocation per chain step

- S2: Thread tier from resolveMethodByOwner via { def, tier } tuple; eliminates double ctx.resolve

Alignment with semantic-model plan (Phase 3 target):

- resolveMethodByOwner now iterates ALL class-like candidates from ctx.resolve, deduplicating matches by nodeId. Absorbs D4's ownerId-filtering into the owner-scoped path.

- Handles homonym classes (two Users in different files) without falling through to D1-D4 fuzzy widening

- Shared-ancestor MRO walks automatically dedup (both homonyms walk to same base method)

- Unified direct-vs-MRO lookup under a single canWalkMRO check

Tests added:

- T1: Three D0 skip-condition tests via new _resolveCallTargetForTesting internal export (overloadHints, preComputedArgTypes, hasActiveModuleAlias)

- T2: Rust qualified-syntax null test (trait-inherited method) + direct impl control

- T3: C++ leftmost-base diamond inheritance test

- B2 lock-in: cross-file class tier assertion

- Homonym disambiguation: only-one-owns-method, both-own-method ambiguity, shared-ancestor MRO convergence

Verification:

- tsc --noEmit: clean

- vitest run test/unit/: 3014 passed

- vitest run test/integration/resolvers/: 1746 passed

* test(SM-11): address second PR #744 review round + per-language integration tests

Review fixes (https://github.com/abhigyanpatwari/GitNexus/pull/744#issuecomment-4211877593):

P1 (Performance): Replace Map allocation in resolveMethodByOwner with a firstDef+ambiguous flag pattern. Zero allocation for the common single-candidate case on the hot path — the previous Map approach allocated on every member call regardless of whether deduplication was needed.

P2 (Test gap): Strengthen the module-alias D0 skip test with a homonym fixture (two Users in different files). Previously the test passed whether or not D0 was actually bypassed; the new version proves D0 must be skipped by showing that resolveMemberCall directly returns null (ambiguous) but D1-D4 with alias narrowing picks the right one. Also fixes the underlying D2-vs-alias widening interaction: when filteredCandidates was narrowed by module-alias disambiguation, D2 no longer widens back to the full fuzzy pool (introduces aliasNarrowed boolean flag).

L1 (Language coverage): Add C# and Kotlin implements-split tests at the resolveMemberCall layer.

L2 (Maintainability): Export OverloadHints as @internal so the test can use a direct cast instead of fragile Parameters<...> type inference.

Per-language integration tests:

- rust-child-extends-parent: Direct impl method resolution via D0 (with honest documentation of the trait-method-as-Function gap that is Phase 5 / SM-16 scope)

- java-interface-default-method: User implements Validator with default method resolved via implements-split MRO

- csharp-interface-default-method: Same pattern for C# 8.0+ default interface methods

- kotlin-interface-default-method: Same pattern for Kotlin interfaces with default implementations

- python-multi-level-mro: 3-level C3 linearization (Grandparent ← Parent ← Child)

- cpp-diamond-inheritance: Classic diamond (Base ← A, B ← Derived) via leftmost-base MRO

Verification:

- tsc --noEmit: clean

- vitest run test/unit/: 3015 passed

- vitest run test/integration/resolvers/: 1763 passed (+17 new per-language tests)

* fix(SM-11): Codex adversarial review corrections + deeper D0 fixes

Addresses the three high-severity findings from the Codex adversarial review of PR #744 (https://github.com/abhigyanpatwari/GitNexus/pull/744#issuecomment-4212075120), plus four deeper fixes discovered during regression triage. All discovered issues are now addressed end-to-end rather than papered over with tail-return fallbacks.

Codex review findings:

R1 (C++ diamond): The cpp-diamond-inheritance fixture used non-virtual inheritance, which is genuinely ambiguous in real C++ (two Base subobjects). Changed A and B to use 'virtual public Base' so there's a single shared Base subobject and d.method() is an unambiguous call that the leftmost-base MRO walk correctly resolves.

R2 (C# default-interface): The csharp-interface-default-method fixture called user.Validate() via a User-typed variable, but C# does not inherit default interface methods as callable class members — the call is only valid through an interface-typed variable. Changed App.cs to 'IValidator user = new User(...)' which is the idiomatic dispatch pattern.

R3 (resolveCallTarget tail-return): When D1-D4 receiver filtering produced zero file-matched and zero owner-matched candidates for a member call, the function fell through to the permissive single-candidate tail return — silently emitting CALLS edges for methods that don't belong to the receiver. Added an explicit null-route inside the D1-D4 block that fires only when both filters yielded 0.

R4 (Rust negative assertion): Added the c.trait_only() negative integration test in rust.test.ts demonstrating that direct member calls on Rust structs do not walk trait ancestry. The test now passes because of R3 (previously fell through to the tail return).

Regression triage discoveries:

1. D0 was dead code on the sequential pipeline. The sequential path sets overloadHints for every call regardless of whether the method is overloaded, and the original D0 skip condition '!overloadHints && !preComputedArgTypes' was therefore always false. The Java/C#/C++ SM-9/SM-10 inheritance tests were passing ONLY via the tail-return fallback. Fix: narrow the skip to 'overloadHints && filteredCandidates.length > 1' — skip D0 only when there are actually multiple candidates that need overload disambiguation.

2. lookupMethodByOwner couldn't disambiguate arity-differing overloads (e.g. C++ greet() vs greet(string)). With D0 now firing on the sequential path, same-name/different-arity overloads would collapse to an arbitrary first pick. Fix: added an optional argCount parameter to lookupMethodByOwner + lookupMethodByOwnerWithMRO that filters the overload set by parameterCount/requiredParameterCount before the returnType dedup.

3. Python and Rust class methods are captured as Function nodes (not Method) with ownerId set to the class. The methodByOwner index only accepted 'Method' and 'Constructor' types, so Python class methods and Rust trait methods were invisible to D0. Fix: extended the methodByOwner indexing condition to include 'Function' when ownerId is set. This also unlocks the Rust trait-method negative assertion by ensuring the qualified-syntax MRO strategy has something to return null for.

4. D0 was being skipped when a local variable shadowed an imported module name (Python 'from models.c import C; c = C()' creates both a module alias 'c → models/c.py' AND a typed local 'c'). Fix: the D0 skip now gates on 'aliasNarrowed' (a new boolean tracking whether the alias block actually narrowed filteredCandidates) instead of 'hasActiveModuleAlias'. If the method isn't in the aliased module, the receiver is a typed local variable and D0 should run.

5. PHP trait walk missed the HasTimestamps trait because lookupClassByName did not include 'Trait' type. buildHeritageMap uses lookupClassByName to resolve parent names, so 'BaseModel use HasTimestamps' was failing to register an ancestor edge for BaseModel → HasTimestamps. Fix: added 'Trait' to CLASS_TYPES. The trait is now a valid class-like type for heritage resolution (PHP use, Rust impl Trait for Struct, Scala traits).

Test updates:

- Updated the 'no heritageMap' unit test in call-processor.test.ts to assert the correct null-route behavior instead of the old tail-return fallback.

- Added a new unit test asserting Trait inclusion in the class set.

- Updated the 'does NOT include other type-like labels' test to remove Trait from its rejection set.

Verification:

- tsc --noEmit: clean

- vitest run test/unit/: 3016 passed (+1 new Trait inclusion test)

- vitest run test/integration/resolvers/: 1764 passed (+1 new Rust negative assertion)

- Zero regressions

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergo Magyar <magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-09 09:52:12 +01:00
Roshan Warrierandtxhno d6debf3324 fix(symbol-table): index constructors in methodByOwner (#753)
Co-authored-by: txhno <198242577+txhno@users.noreply.github.com>
2026-04-09 08:26:15 +01:00
Pratyush Sharma 9ab92a97d0 fix(deps): pin tree-sitter-c override to resolve peer dep conflict (#720) (#723) 2026-04-09 06:40:17 +01:00
Murat Çelik 4fde5f241b feat: print skipped large file paths in verbose analyze output (#745) 2026-04-09 06:18:32 +01:00
evolution 9f9bbcd744 feat: support GITNEXUS_HOME env var to customize global directory (#746) 2026-04-09 06:18:17 +01:00
Cocoon-Break fd67cfd5a7 docs: fix web UI install link spacing in README (#731) 2026-04-09 06:16:14 +01:00
Pratyush Sharma 3f28f7ead5 fix(web): correct dev-mode serve command in OnboardingGuide (#725) 2026-04-09 06:15:14 +01:00
d9ba9aa998 SM-10: Add MRO fast path before D2 fuzzy widening in resolveCallTarget (#741)
* Initial plan

* Add MRO fast path before D2 fuzzy widening in resolveCallTarget

When receiverTypeName is known, try resolveMethodByOwner (owner-scoped
+ MRO lookup) before falling back to the expensive lookupFuzzy in D2.
This short-circuits cross-file member call resolution for the common
non-overloaded case.

The fast path is skipped when overload disambiguation hints are
available (overloadHints or preComputedArgTypes) to avoid picking the
wrong overload for same-return-type overloaded methods.

Passes heritageMap to resolveCallTarget from all 4 call sites:
- Language seed path (processCalls)
- Sequential path (processCalls)
- walkMixedChain fallback
- Worker path (processCallsFromExtracted)

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9e49521f-2472-47bc-96e9-be4a46b073f0

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(SM-10): address PR #741 review

Correctness:
- Module-alias guard for D0. When call.receiverName matches an active
  entry in ctx.moduleAliasMap for the current file, D0 is now skipped
  and resolution falls through to D1-D4 which respects the
  alias-narrowed candidate pool. Prevents a homonymous class in a
  different file from being picked by ctx.resolve(receiverTypeName)
  inside resolveMethodByOwner. New unit test pins the contract.

Unit tests (call-processor.test.ts — 3 new):
- D0 hit: child.parentMethod() resolves via MRO walk when
  heritageMap is provided.
- D0 skipped: same scenario still resolves via D1-D4 when heritageMap
  is undefined (backward-compat guard).
- Module-alias guard: two files both define class User with a save()
  method; 'import auth_mod as auth' in app.py must resolve
  auth.user.save() to auth_mod.py, not user_mod.py.

Integration language coverage (+3 fixtures/tests):
- swift-child-extends-parent — first-wins, gated on swiftAvailable.
- ruby-child-extends-parent   — first-wins.
- php-child-extends-parent    — first-wins (uses ParentClass since
  'Parent' is a PHP reserved word).

* test(SM-10): address second PR #741 review round

Unit tests (call-processor.test.ts, +2 new):
- overloadHints guard: Java source with two same-return-type overloads
  method(int) and method(String), int added first so lookupMethodByOwner
  would return it. processCalls auto-generates overloadHints for Java,
  forcing D0 to be skipped. o.method("hello") must resolve to
  method(String) via literal-inferred disambiguation.
- preComputedArgTypes guard: worker-path equivalent via
  processCallsFromExtracted with ExtractedCall.argTypes=['String'].
  Same two overloads, same correctness guarantee.

Integration tests (+2 fixtures + test blocks):
- go-child-extends-parent    — struct embedding, first-wins
  (Go structs are labeled 'Struct' not 'Class' in GitNexus).
- dart-child-extends-parent  — extends, first-wins, gated on
  dartAvailable like other Dart tests.

Documentation:
- Expanded the fallthrough comment in resolveMethodByOwner to clarify
  that unknown-extension paths land on plain lookupMethodByOwner
  without an ancestor walk, and that D1-D4 still runs on D0 miss.

* test(SM-10): D0 miss with heritageMap present falls through to D1-D4

Closes the last remaining gap from PR #741 review round 3. The existing
'D0 skipped' test only covered the heritageMap=undefined case, leaving
the miss-with-heritageMap path implicitly covered by integration tests
only. This adds a focused unit test where:

- Class Obj has a method doWork findable via tiered resolution
  (import-scoped) but intentionally NOT registered in methodByOwner
  (no ownerId), so lookupMethodByOwner misses.
- heritageMap is provided but built from an empty heritage array, so
  getAncestors(class:Obj) returns []. The MRO walk yields no parents.
- lookupMethodByOwnerWithMRO therefore returns undefined → D0 miss.
- D1 resolves the receiver type; D2 widens via lookupFuzzy;
  D3 file-filter picks the single matching candidate.
- A CALLS edge must still be emitted — D0 miss must not swallow
  the call.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-08 23:24:07 +01:00
abhigyanpatwariandClaude Opus 4.6 89feea744d refactor(sm-14): address PR review feedback
- Remove redundant typeEnvBindings worker payload — allScopeBindings is a strict
  superset (file-scope entries with scope=''). Removes duplicate IPC data and
  the dead fallback branch in pipeline.ts.
- Add single-use guard to TypeEnvironment.flush() — throws on second call to
  prevent silent duplication. Update JSDoc to clarify "copy" semantics
  (env is not actually drained — TypeEnv is per-file and discarded immediately).
- Remove redundant inner guard in parse-worker.ts allScopeBindings serialization.
- Rename allAllScopeBindings -> allScopeBindingsByFile for clarity.
- Document estimateMemoryBytes pessimistic ASCII assumption (V8 uses Latin-1
  for all-ASCII strings, so actual heap cost is ~half).
- Add test for single-use flush() guard.

Addresses PR #743 review feedback.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 03:20:18 +05:30
abhigyanpatwariandClaude Opus 4.6 ec85c1e8bc style: fix prettier formatting
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 02:48:54 +05:30
abhigyanpatwariandClaude Opus 4.6 0f83912636 test(sm-14): add pipeline integration simulation for BindingAccumulator
Simulates the worker deserialization -> accumulator -> fileScopeEntries flow
to verify end-to-end correctness.

Part of #679.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 02:31:23 +05:30
abhigyanpatwariandClaude Opus 4.6 750f4336f6 feat(ingestion): wire flush() into sequential processCalls path
Pass bindingAccumulator to processCalls on the sequential code path so
TypeEnv scopes are flushed into the accumulator for Phase 9+ cross-file
type propagation, matching the worker path wired in Task 4.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 02:31:23 +05:30
abhigyanpatwariandClaude Opus 4.6 a7fdb36113 refactor(pipeline): wire BindingAccumulator into chunked parse pipeline
Replace ad-hoc workerTypeEnvBindings array with BindingAccumulator class.
Add allScopeBindings to WorkerExtractedData for multi-scope binding
collection with backward-compat fallback to old typeEnvBindings format.
Return BindingAccumulator from runChunkedParseAndResolve and finalize
before Phase 14 cross-file binding propagation.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 02:30:14 +05:30
abhigyanpatwariandClaude Sonnet 4.6 2f72fcc8d1 feat(parse-worker): extend worker serialization to include all scopes
Add FileAllScopeBindings interface and allScopeBindings field to
ParseWorkerResult, serializing all TypeEnv scopes (including
function-local) for BindingAccumulator cross-file type propagation.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-09 02:30:14 +05:30
abhigyanpatwariandClaude Sonnet 4.6 56e2e7f520 feat(type-env): add flush() method to TypeEnvironment
Adds flush(filePath, accumulator) to the TypeEnvironment interface and
buildTypeEnv return object, draining all scoped bindings into a
BindingAccumulator. Adds 4 unit tests covering file-scope, function-scope,
empty env, and multi-file accumulation.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-09 02:30:13 +05:30
abhigyanpatwariandClaude Opus 4.6 3111f4dcd3 feat(sm-14): add BindingAccumulator class with unit tests
Read-append-only accumulator that collects (filePath, scope, varName) -> typeName
bindings from TypeEnv outputs across all files. Supports finalization, file-scope
filtering, iteration, and memory estimation.

Part of #679.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 02:30:13 +05:30
c19e76a4a3 feat(SM-9): Add lookupMethodByOwnerWithMRO using HeritageMap (#740)
* Initial plan

* feat(SM-9): add lookupMethodByOwnerWithMRO with HeritageMap parent chain walking

- Export c3Linearize from mro-processor.ts for reuse
- Add lookupMethodByOwnerWithMRO in call-processor.ts with MRO strategy support
- Update resolveMethodByOwner to fall back to MRO walk when HeritageMap available
- Thread heritageMap through walkMixedChain for chain resolution
- Add 10 unit tests covering all acceptance criteria

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/cc58249b-42f1-45a9-89fb-e3917e4d0171

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* feat(SM-9): add Java integration test with class Child extends Parent fixture

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/cc58249b-42f1-45a9-89fb-e3917e4d0171

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docs: address code review comments on MRO strategy documentation

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/cc58249b-42f1-45a9-89fb-e3917e4d0171

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* perf(SM-9): address PR #740 review comments

- Eliminate double direct lookup in resolveMethodByOwner: delegate
  straight to lookupMethodByOwnerWithMRO when a HeritageMap is
  available (the MRO helper already does the direct lookup before
  walking ancestors). Fallback path handles the no-HeritageMap case.
- Memoize C3 linearization per HeritageMap via a WeakMap keyed cache.
  HeritageMap is immutable after build, so C3 results are stable for
  its lifetime; WeakMap lets the cache auto-drain when the HeritageMap
  is GC'd. Null sentinel caches linearization failures so cyclic
  hierarchies are not reprocessed. Eliminates per-call buildParentMap +
  c3Linearize on Python codebases.
- ancestors variable typed as readonly to accept the cached result
  without copying.
- Add four missing MRO unit tests: Kotlin implements-split, C#
  implements-split, JavaScript first-wins (separate provider from TS),
  and C++ leftmost-base diamond (first diamond test for C++).

* fix(SM-9): CI prettier + address PR #740 follow-up review

- Fix CI prettier failure in test/integration/resolvers/java.test.ts
  (auto-formatted — was introduced in 37563a31 before my first fix
  commit but had not been caught locally).
- Pin caller on the SM-9 Java integration test (parentMethodCall.source
  === 'run') so a regression that misattributes the CALLS edge fails.
- Add two implements-split unit tests:
  * Ambiguous default from two interfaces → BFS first-wins. Pins the
    contract that lookupMethodByOwnerWithMRO returns a defined result
    (full ambiguity detection is deferred to computeMRO graph pass).
  * Class method precedence over interface default: Child extends Base
    implements IFoo where both define handle() — documents that BFS
    visits the extends edge first, matching Java's class-wins rule.
- Add @internal JSDoc on lookupMethodByOwnerWithMRO clarifying it is
  exported only for testing; resolveMethodByOwner is the proper entry
  point for callers.

* test(SM-9): per-language integration fixtures and tests for inherited method resolution

Extends the SM-9 integration coverage beyond Java with six new
child-extends-parent fixtures, one per MRO strategy:

- python-child-extends-parent       → C3 strategy
- typescript-child-extends-parent   → first-wins
- javascript-child-extends-parent   → first-wins (separate provider)
- kotlin-child-extends-parent       → implements-split
- csharp-child-extends-parent       → implements-split
- cpp-child-extends-parent          → leftmost-base

Each fixture follows the java-child-extends-parent pattern:
- Parent class with a single method
- Child class extending Parent, no override
- App class/function that instantiates Child and calls
  the parent method — exercises the full ingestion pipeline,
  HeritageMap construction, and lookupMethodByOwnerWithMRO walk.

For every fixture the matching integration test asserts:
- Parent and Child classes are detected
- Child → Parent EXTENDS edge is emitted
- The parent-method call resolves to the correct target file
- The caller is pinned (source === 'run' / 'Run') to catch
  edge misattribution regressions

Rust is intentionally omitted — its qualified-syntax strategy
returns undefined from lookupMethodByOwnerWithMRO by design, so
there is no inherited-method resolution to assert against.

All 1739 integration resolver tests pass (+18 new SM-9 tests
across 6 languages).

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-08 20:46:45 +01:00
b75e76d44a feat(SM-8): Build HeritageMap from accumulated ExtractedHeritage[] (#739)
* Initial plan

* feat(SM-8): add HeritageMap with MRO-aware parent/ancestor lookup

- New heritage-map.ts: HeritageMap interface with getParents() and getAncestors()
- buildHeritageMap() consumes ExtractedHeritage[], resolves names via lookupClassByName
- Cycle protection and bounded depth (MAX_ANCESTOR_DEPTH=32) in getAncestors
- Worker path: HeritageMap built from deferredWorkerHeritage, threaded into processCallsFromExtracted
- Sequential path: Heritage accumulated across chunks, HeritageMap built after all chunks, passed to processCalls
- 18 unit tests covering parent lookup, multi-level, diamond, cycles, missing parent, bounded depth

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c413e0a3-5d63-4ddb-8ece-02fe6ed99efd

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: rename cycle test for clarity per code review

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c413e0a3-5d63-4ddb-8ece-02fe6ed99efd

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor(SM-8): merge implementor map into heritage map

- Add `getImplementorFiles(interfaceName)` to HeritageMap interface
- Build implementor index (interface name → file paths) alongside parent
  lookup in `buildHeritageMap`, using same `resolveExtendsType` logic
- Remove `ImplementorMap` type, `buildImplementorMap`, `mergeImplementorMaps`
  from call-processor.ts
- Update `findInterfaceDispatchTargets`, `processCalls`, and
  `processCallsFromExtracted` to use HeritageMap for both parent
  lookup and implementor dispatch
- Pipeline: single `buildHeritageMap` call replaces separate
  buildImplementorMap + buildHeritageMap for both worker and
  sequential paths
- Migrate implementor tests from call-processor.test.ts to
  heritage-map.test.ts (4 new getImplementorFiles tests)
- Update interface dispatch test to use buildHeritageMap instead
  of hand-constructed ImplementorMap

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/085dffb4-b31e-4aa5-9aa3-4314bc0010e7

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: rename implementor test for clarity per code review

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/085dffb4-b31e-4aa5-9aa3-4314bc0010e7

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(SM-8): address PR #739 review comments

- pipeline.ts: cache chunk file contents from Pass 1 to eliminate
  double-read of sequential chunks in Pass 2. Peak memory drains
  incrementally as Pass 2 processes each chunk.
- heritage-map.ts: document Rust trait-impl omission from implementor
  index and the interface-name collision limitation.
- heritage-map.test.ts: add six tests covering the extends->IMPLEMENTS
  path across C# (interfaceNamePattern), Swift (heritageDefaultEdge),
  Java (symbol-table Interface lookup), Kotlin, PHP, and the Rust
  trait-impl omission.
- pipeline.ts: comment why the heritage accumulation uses a manual
  push loop instead of spread (ref #650).

* test(SM-8): address second PR #739 review pass

- Add TypeScript implements test to getImplementorFiles (closes
  the .ts coverage gap flagged by the bot reviewer).
- Tighten deep-chain boundary assertion from toBeLessThanOrEqual(32)
  to toBe(32) so a future regression returning fewer ancestors
  fails loudly. Added an ancestors[31] === 'class:Level32' check
  to pin the upper boundary.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-08 19:00:08 +01:00
MyShiningand许恩宁 83b5bec293 [cli] Replace owner-filtered method lookups in type-env (#736)
* refactor(type-env): use owner method lookup

* test(type-env): cover owner lookup edge cases

* test(type-env): cover inherited overload ambiguity

---------

Co-authored-by: 许恩宁 <xuenning@qiyi.com>
2026-04-08 17:41:09 +01:00
MyShiningand许恩宁 3388ae16d7 [cli] Replace Phase P class checks with class lookup index (#734)
* refactor(call-processor): use class lookup index in phase p

* test(call-processor): cover class lookup fallback

---------

Co-authored-by: 许恩宁 <xuenning@qiyi.com>
2026-04-08 14:48:54 +01:00
MyShiningand许恩宁 d784f591b2 [cli] Replace class-type fuzzy lookups in type-env.ts (#733)
* refactor(type-env): use class lookup index for type resolution

* test(type-env): add lookupClassByName regression coverage

* test(type-env): expand class lookup regression coverage

---------

Co-authored-by: 许恩宁 <xuenning@qiyi.com>
2026-04-08 14:12:01 +01:00
Kunal Hemnani 0f43190543 feat(symbol-table): add fuzzy lookup counters (#708) 2026-04-08 07:53:30 +01:00
Roshan Warrier fe87ff8f74 fix(symbol-table): index constructors in methodByOwner (#694) 2026-04-08 06:32:02 +01:00
Deepak Chauhan be2401061e [cli] Add qualified class lookups to SymbolTable (#716) 2026-04-07 22:57:18 +01:00
Deepak Chauhan b73233d232 feat(symbol-table): add class name lookup index (#707) 2026-04-07 13:29:50 +01:00
Tushar Dhawas (Kyo) 1c8ae5eb46 refactor: extract CLASS_LIKE_TYPES constant (#693)
* refactor: extract CLASS_LIKE_TYPES constant

* chore: apply prettier formatting
2026-04-07 11:56:03 +01:00
Zander Raycraft b73928f732 scarf (#688) 2026-04-06 18:39:43 -05:00
Dmytro Semchuk 19faf3b326 fix(docs): fix codex duplicate typo in main readme file (#687) 2026-04-06 22:20:35 +01:00
Gergő Magyar cb772b9e29 feat: lookupMethodByOwner index for O(1) cross-class chain resolution (#665)
Add eagerly-populated methodByOwner index to SymbolTable, keyed by
ownerNodeId\0methodName. Used by walkMixedChain as a fast path for
resolving intermediate method calls in cross-class chains like
user.getAddress().getCity().getZipCode(), avoiding expensive fuzzy
lookups when the owner type is already known.

Handles overloaded methods: returns the first match when all overloads
share the same returnType, undefined when return types differ (ambiguous).

- Add lookupMethodByOwner to SymbolTable interface + implementation
- Add resolveMethodByOwner helper in call-processor.ts
- Add fast path in walkMixedChain before resolveCallTarget fallback
- Add Java cross-class chain fixture + 6 integration tests
- Add 148 unit tests for methodByOwner index behavior
2026-04-06 10:20:04 +01:00
ivkondandClaude Opus 4.6 10f8815639 fix(ignore): respect negation patterns in .gitnexusignore (#654)
* fix(ignore): respect negation patterns in .gitnexusignore childrenIgnored

childrenIgnored checked `ig.ignores(rel) || ig.ignores(rel + '/')` which
short-circuited on the bare path — directory-only negation patterns like
`!iOS/` were missed because `ig.ignores('iOS')` treats the path as a file.
Now only checks with trailing slash since childrenIgnored is only called
for directories. Bare-name patterns (e.g. `local`) still match per gitignore spec.

Fixes #596

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test(ignore): add edge-case for bare `!dir` negation pattern

Verifies that `!iOS` (without trailing slash) also un-ignores the iOS/
directory — confirms the `ignore` package normalizes both `!dir` and
`!dir/` forms consistently when tested with a trailing-slash path.

Addresses non-blocking review suggestion on #654.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs(ignore): link ignore package docs for bare-name normalization

Adds references to the `ignore` package documentation in both the
childrenIgnored comment and the bare-negation test, explaining why
`!iOS` (without trailing slash) also re-includes the iOS/ directory.

Addresses non-blocking review suggestion on #654.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-06 08:32:47 +01:00
Abhigyan Patwari 6ead5e5986 fix(setup): prefer global gitnexus binary over npx for MCP config (#653) 2026-04-06 07:46:10 +01:00
Abhigyan Patwari 14791ded4a fix(server): return clean CORS rejection instead of 500 error (#646) 2026-04-06 07:45:19 +01:00
tantkandtian kian tan 9eeb20bb04 fix: replace Array.push(...spread) with loop to prevent stack overflow (#650)
* fix: replace Array.push(...spread) with loop to prevent stack overflow

On large codebases (78K+ C files), deferred arrays in
runChunkedParseAndResolve accumulate 100K+ entries. The spread
operator in push(...array) puts every element on the call stack
as a function argument, exceeding the maximum call stack size.

Replace all 11 occurrences of `arr.push(...other)` with
`for (const _item of other) arr.push(_item)` which uses
constant stack space regardless of array size.

Fixes #649

* fix: also replace push(...spread) in parsing-processor.ts

* style: format pipeline.ts to match prettier config

---------

Co-authored-by: tian kian tan <tan@example.com>
2026-04-05 21:54:09 +01:00
Gergő Magyar 5a7c0fdbb1 feat: same-arity overload disambiguation via type-hash suffix (#651) (#658)
* feat: same-arity overload disambiguation via type-hash suffix (#651)

Add ~type1,type2 suffix to Method/Constructor node IDs when same-arity
overloads with different parameter types exist in the same class. Also add
$const suffix for C++ const-qualified method overloads via new isConst field.

Key changes:
- typeTagForId() detects same-arity collisions and appends ~typeTag
- constTagForId() detects const/non-const collisions and appends $const
- TS/JS excluded from type-hashing (overload signatures collapse to impl body)
- Sequential findEnclosingFunction fixed: falls through on ambiguous same-class
  candidates instead of picking first; fallback path includes typeTag + constTag
- Per-call-site integration tests across Java, C#, Kotlin, C++, TypeScript
- Cross-file + chain resolution tests for all 5 languages
- C++ isConst extraction via tree-sitter type_qualifier in function_declarator

1710 integration + 18 unit tests pass.

* fix: preserve generic/template args in type-hash, perf + type safety fixes

- Add rawType field to ParameterInfo preserving full type text (vector<int>)
  while type stays simplified (vector). typeTagForId uses rawType for tags.
- Populate rawType in all 11 language method extractors
- Add buildCollisionGroups() to pre-group methods by name#arity (O(N) once
  per class instead of O(N) per method call)
- Cache method extraction in call-processor findEnclosingFunction fallback
- Fix null guards on getLanguageFromFilename in all findEnclosing paths
- Tighten SKIP_TYPE_HASH_LANGUAGES to ReadonlySet<SupportedLanguages>
- Document ID stability invariant on first overload introduction
- C++ integration tests: template overloads (vector<int> vs vector<string>),
  cross-file template + chain resolution, out-of-class method definitions

1718 integration + 20 unit tests pass.

* fix: add rawType to method-extraction unit test assertions

All 26 parameter .toEqual() assertions in method-extraction.test.ts
needed the new rawType field added to match ParameterInfo schema change.

* perf: cache tempMap/groups per class, consolidate extractFromNode

- Cache derived method map + collision groups per classNode.id in
  parsing-processor (avoids rebuild per method in same class)
- Replace per-call extractFromNode with cached class extraction +
  funcName:line lookup in call-processor fallback (avoids AST walk
  per call site)
- Remove dead clearEnclosingFunctionCache export, fix JSDoc

* test: add sequential-path integration test for same-arity overloads

Add skipWorkers option to PipelineOptions to force sequential parsing.
New test suite verifies type-hash disambiguation produces identical
results through the sequential path (parsing-processor + call-processor
findEnclosingFunction) as the worker path.
2026-04-05 21:51:55 +01:00
Gergő Magyar 0561d24efd feat: METHOD_IMPLEMENTS edges, overload disambiguation, MethodExtractor unification (#574) (#642) 2026-04-04 18:41:47 +01:00
Abhigyan Patwari 153262304c fix(mcp): unify stdout silencing to prevent embedder/pool-adapter conflicts (#645) 2026-04-04 11:56:49 +01:00
Abhigyan Patwari 16cf4c503e fix(web): replace aggressive heartbeat disconnect with graceful reconnection (#643) 2026-04-04 11:56:35 +01:00
Abhigyan Patwari 57951a197b fix(web): scope all backend calls to the active repo, not always the first (#644) 2026-04-04 11:55:34 +01:00
Gergő MagyarandClaude Opus 4.6 63fc4c795f feat: MethodExtractor configs for Python, PHP, Swift, Dart, Rust, Ruby (#624)
* feat: MethodExtractor configs for Python, PHP, Swift, Dart, Rust, Ruby with exhaustive integration tests

Add per-language MethodExtractionConfig for all remaining tree-sitter languages
(RFC #568 PR 2). Each config follows the established createMethodExtractor()
factory pattern — no new types, no parse-worker changes.

Configs:
- Python: @abstractmethod, @staticmethod/@classmethod, *args/**kwargs, type hints, _/__ visibility
- PHP: abstract/final/static keywords, PHP 8 #[] attributes, __construct/__destruct
- Swift: 5-level visibility, protocol-as-abstract, static/class methods, @ attributes
- Dart: _ convention visibility, abstract (no body), method_signature unwrapping
- Rust: pub visibility, &self receiver, trait_item + impl_item, #[] attributes
- Ruby: positional visibility via sibling-walk, singleton_method as static

Integration fixtures (18 directories) covering 3 resolution patterns:
- Method enrichment: parameterTypes, isAbstract, isFinal, annotations on graph nodes
- Overload dispatch: arity-based CALLS resolution via parameterTypes
- Abstract dispatch: abstract/concrete method distinction (Python, PHP, Rust, Swift)

Go deferred — requires factory changes for receiver-based method extraction.

Closes #571

* fix: address code review findings across 6 MethodExtractor configs

Fix all actionable items from the PR #624 deep-dive review:

Dart (critical — fixes 6 CI failures):
- isDartStatic: check children first, siblings as fallback
- isDartAbstract: handle declaration nodes for abstract methods
- extractSingleParam: detect required keyword as sibling token
- Add declaration to methodNodeTypes, mixin_declaration to typeDeclarationNodes
- Add member call query for variable assignments in tree-sitter-queries

Python:
- hasDecorator now matches dotted paths (e.g. @abc.abstractmethod)
- Fix version comment from ^0.23.6 to 0.23.4

PHP:
- Add enum_declaration to typeDeclarationNodes (PHP 8.1+)
- Add version comment for 0.23.12

Swift:
- Add isOverride using hasKeyword/hasModifier pattern

Rust:
- Fix version comment from ^0.23.2 to 0.23.1

Also: identifier fallback in generic.ts for mixin owner names,
Dart integration test label fix (Method vs Function), version
comment for tree-sitter-dart 1.0.0.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: Dart extension_declaration and Ruby module_function support

Dart:
- Add extension_declaration to typeDeclarationNodes and extension_body
  to bodyNodeTypes — extension methods are now extracted into the graph
- Add extension_declaration and mixin_declaration to CLASS_CONTAINER_TYPES
  for HAS_METHOD edge resolution

Ruby:
- module_function now maps to visibility 'private' in extractRubyVisibility
- module_function methods marked isStatic via backward-walk in isStatic
- Override semantics: private/public after module_function resets isStatic

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(go): Go MethodExtractor config with receiver-based extraction

Add Go as the 13th language with a per-language MethodExtractor config.
Go methods are top-level (not nested in struct bodies), so this adds
extractFromNode() to the MethodExtractor interface for direct method
node extraction without an enclosing class.

Config extracts:
- Name from field_identifier (methods) / identifier (functions)
- Return type including multi-return (first type from parameter_list)
- Parameters with variadic support
- Visibility via uppercase/lowercase convention
- Receiver type with pointer unwrapping (*User → User)
- isStatic for functions (no receiver)

Infrastructure:
- extractOwnerName optional hook on MethodExtractionConfig
- extractFromNode on MethodExtractor (factory auto-implements)
- Parse-worker uses extractFromNode when no enclosing class found
- method_declaration added to CLASS_CONTAINER_TYPES

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test: method enrichment integration tests for 7 languages + TS abstract class fix

Add method-enrichment integration test fixtures and test blocks for
Go, C++, Java, Kotlin, TypeScript, JavaScript, and C#. Each fixture
tests: class detection, HAS_METHOD edges, EXTENDS edges, isAbstract,
isStatic, annotations, parameterTypes, and CALLS edge resolution.

Fixes found during testing:
- Remove method_declaration from CLASS_CONTAINER_TYPES (added for Go
  but broke Java/C# HAS_METHOD edge resolution — method_declaration
  is also Java's method node type)
- Add abstract_class_declaration query to TypeScript tree-sitter
  queries (was missing, so abstract classes were invisible to pipeline)

1699 integration tests pass across 20 test files, 0 regressions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: format typeDeclarationNodes array for better readability in PHP config

* fix: Go interface methods + Rust impl-for-Struct owner resolution

Go:
- Add method_elem to methodNodeTypes so interface method signatures
  are extractable as abstract methods
- Integration test: Animal interface detected, Speak isAbstract,
  CALLS edges from app.go

Rust:
- Add extractOwnerName to resolve impl Trait for Struct to the
  concrete Struct (not the Trait) — fixes method misattribution
- Fix findEnclosingClassId to generate Struct: label (not Impl:)
  for impl blocks so HAS_METHOD edges resolve to struct nodes
- Tighten abstract-dispatch test: assert SqlRepo owns find/save

generic.ts:
- Fix extractOwnerName fallback: when hook returns a value, skip
  both name-field and type_identifier scan (was overwriting result)

1703 integration tests pass, 0 regressions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: code review response — Rust impl label, Swift params, Dart async, sequential methodExtractor

Address code review findings from PR #624:

- ast-helpers: Rust `impl Trait for Struct` uses Struct label (matches existing
  graph node), plain `impl Struct` uses Impl label (matches definition.impl)
- swift: fix parameter type extraction (user_type not type_annotation), detect
  default values as function_declaration siblings, add version comment
- dart: isDartAsync now detects async*/sync* generators, add clarifying comment
  for declaration nodes in extension bodies
- python: correct isFinal comment (PEP 591 @typing.final exists, just not modeled)
- parsing-processor: port methodExtractor enrichment to sequential path so
  isAbstract/isStatic/visibility/annotations/isFinal populate on <15-file repos
- tests: remove silent `if (prop !== undefined)` guards, assert properties
  directly, fix label queries (Dart Method vs Function, Swift Method for protocol
  methods), add Rust HAS_METHOD sourceLabel tests, Swift parameterTypes tests,
  and Dart async/sync* integration tests with fixture

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: Rust grammar gap + qualified method IDs to resolve same-file collisions

Phase 1 — Rust grammar:
- Add function_signature_item query to RUST_QUERIES so abstract trait methods
  (fn speak(&self) -> String;) become graph nodes with isAbstract=true

Phase 2 — Qualified method IDs:
- findEnclosingClassInfo returns {classId, className} for AST-based class lookup
- Both parsing paths (sequential + worker) qualify method/property IDs with
  enclosing class: Method:file:ClassName.method instead of Method:file:method
- extractFuncNameFromSourceId handles ClassName.method format
- Fixes silent data loss when same-name methods in different classes shared a
  file (e.g., Animal.speak and Dog.speak both now exist as distinct graph nodes)

Test updates:
- Rust: abstract+concrete trait methods both verified, function count adjusted
- Python: static method disambiguation now emits 2 CALLS edges (correct — no
  more ID collision masking the second call)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: owner-aware resolution for qualified method IDs

Address Codex adversarial review findings after qualified ID change:

- findEnclosingFunction: disambiguate candidates by ownerId when multiple
  same-name methods exist in file; qualify fallback-generated IDs
- findEnclosingFunctionId (worker): qualify sourceIds with enclosing class
  name so CALLS source attribution matches definition-phase node IDs
- buildExportedTypeMapFromGraph: use lookupExactAll + nodeId match instead
  of lookupExactFull which returns first definition for bare name

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: methodExtractor variadic arity, return type preservation, PHP abstract dispatch

Three bugs in the methodExtractor enrichment path broke 17 integration tests:

1. Variadic parameterCount: buildMethodProps and parse-worker set
   parameterCount = info.parameters.length even for variadic functions,
   causing arity filtering to reject valid calls. Now checks isVariadic
   and sets parameterCount = undefined (matching extractMethodSignature).

2. C++ bare `...` token: extractCppParameters only iterated named
   children, missing the unnamed `...` token in C-style variadics like
   log_entry(const char* fmt, ...). Added fallback scan of all children.

3. Return type stripping: All 11 language extractReturnType functions
   used extractSimpleTypeName() which strips generic parameters
   (List<User> → "List", Task<User> → "Task"). Changed to .text?.trim()
   to preserve full generic types needed for for-loop iterable resolution,
   async-await binding, and return-type inference.

Also fixes PHP abstract dispatch test that matched SqlRepository instead
of the interface due to ambiguous filePath.includes('Repository') filter,
and adds parent-walk fallback in PHP isAbstract for extractFromNode path.

* chore: remove plan and review artifacts from PR

* fix: address Round 4 review findings + infrastructure improvements

- Ruby: add singleton_class support for class << self methods (4 new tests)
- PHP: add enum_declaration to CLASS_CONTAINER_TYPES
- Dart: add mixin/extension labels to CONTAINER_TYPE_TO_LABEL
- Swift: add TODO for unverifiable struct/enum node types on Node 22
- C#: add grammar version comment (0.23.1)
- Ruby: fix version comment range to pin (0.23.1)
- Rust/ast-helpers: add cross-reference comments for impl_item duplication
- ast-helpers: document CLASS_CONTAINER_TYPES ↔ typeDeclarationNodes invariant
- generic.ts: replace Array.includes with Set for O(1) dedup in addNestedBodies
- Go/Python/Ruby: align isAbstract signature with 2-param interface contract
- CLAUDE.md: fix malformed backtick around gitnexus:start HTML comment
- parsing-processor: add per-class method extraction cache (eliminates O(N*M))
- ast-helpers: add scoped_type_identifier to impl_item resolution
- call-processor: add dev-mode warnings at silent candidates[0] fallbacks
- MCP context(): surface methodMetadata for Method/Function/Constructor nodes
- resources.ts: update schema to list all stored Method properties

* fix: singleton_class HAS_METHOD edge regression in findEnclosingClassInfo

singleton_class (class << self) was added to CLASS_CONTAINER_TYPES but
has no name field — its receiver `self` has node type 'self', not
'identifier'. findEnclosingClassInfo now walks up to the enclosing
class/module to inherit its name, matching ruby.ts:extractOwnerName.

Also fixes findEnclosingClassNode in parse-worker.ts to skip
singleton_class and return the actual class/module node.

Adds integration test assertions for from_habitat (class << self method):
HAS_METHOD edge from Animal, isStatic=true, parameterCount=1.

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 16:11:31 +01:00
Abhigyan Patwari 5c4fca21c3 Merge pull request #626 from ivkond/feat/intra-repo-service-tracking-clean
[group] Intra-repo service communication tracking
2026-04-03 17:04:24 +05:30
Nguyen Hai SonandClaude Opus 4.6 dd0f5eed7d feat(vue): Vue SFC support + destructured call result tracking (#604)
* feat(vue): add Vue SFC (.vue) support for indexing

Vue Single File Components are now fully supported in the indexing pipeline.
The implementation extracts <script> / <script setup> blocks from .vue files
and parses them using the existing TypeScript tree-sitter grammar — no new
npm dependencies required.

Key changes:
- SFC script extractor: regex-based extraction of <script setup lang="ts">
  blocks with correct line offset mapping back to the .vue file
- Vue language provider: reuses TypeScript queries, type config, field
  extractors, and named binding extraction
- Import resolution: .vue added to EXTENSIONS so `import Foo from './Foo'`
  resolves to Foo.vue; Vue import resolver delegates to TS resolver for
  tsconfig path alias support
- Export detection: <script setup> top-level bindings are implicitly exported
- Template component detection: PascalCase tags in <template> emit CALLS edges
- Line offsets applied to all emitted positions (startLine, endLine, route
  lineNumbers, decorator positions) in both worker and sequential paths

Validated on a 3,553-file Vue project:
  Before: 24,693 nodes | 73,614 edges | 0 symbols from .vue
  After:  30,495 nodes | 112,324 edges | 5,213 symbols from .vue
          18,682 imports from .vue | 5,826 vue-to-vue imports

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(typescript): track destructured call results in TypeEnv

Extend `extractPendingAssignment` to handle object destructuring from
function calls and await expressions:

  const { isMaker } = useUserRole()
  const { data } = await fetchData()
  const { name } = repo.getProfile()

Previously, only `const { x } = someVariable` (identifier RHS) produced
TypeEnv bindings. Call-expression RHS was silently skipped, leaving
destructured properties untracked.

The fix emits a synthetic `callResult` item plus N `fieldAccess` items
per destructured property, which the existing fixpoint resolver processes
in 2 iterations. No changes needed to type-env.ts, PendingAssignment
types, or call-processor — the existing infrastructure handles it.

Also extracts a `collectDestructuredFields` helper to share the
object_pattern property iteration logic between the identifier and
call-expression branches.

Note: Full property-type resolution requires the callee to have a
declared returnType in the SymbolTable. Arrow-function composables
without type annotations (common in Vue/React) won't resolve property
types until return-type inference is added in a future change.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(vue): address PR review issues for Vue SFC support

- Extract duplicated isVueSetupTopLevel to vue-sfc-extractor.ts shared
  utility, removing identical copies from parse-worker.ts and
  parsing-processor.ts
- Fix VUE_BUILT_INS to be a superset of TS BUILT_INS by importing and
  spreading the TypeScript set, preventing spurious unresolved calls for
  standard built-ins (Symbol, BigInt, WeakMap, array methods, etc.)
- Add Vue template component CALLS edge resolution in both sequential
  and worker paths (call-processor.ts), matching PascalCase template
  tags against imported .vue file basenames via the import map
- Add integration test for template PascalCase CALLS edges
  (App.vue → Button.vue)
- Add integration test for isExported: false on non-setup <script>
  blocks (OldStyle.vue options API)
- Add comment explaining TEMPLATE_RE greedy regex behavior for nested
  template tags
- Fix stale language count comment (14 → 15) and remove dead code
  branch in test

Made-with: Cursor

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 14:18:55 +05:30
Chirag Nighut e3d73a7aed Java method reference (#622) 2026-04-02 15:16:35 +01:00
ivkondandClaude Opus 4.6 255e3e79eb fix(group): address 4 HIGH-priority issues from PR #626 review
1. Path traversal via group name — add validateGroupName() with regex
   [a-zA-Z0-9][a-zA-Z0-9_-]*, called in getGroupDir (defense in depth)

2. gRPC proto regex can't handle nested braces — replace serviceRe with
   extractServiceBlocks() brace-depth counter (init depth=1, skip
   malformed protos)

3. Service boundary detector directory exclusions — add EXCLUDED_DIRS
   set (vendor, target, build, dist, __pycache__, .venv, venv, .tox,
   .mypy_cache, .gradle, .mvn, out, bin) replacing inline node_modules

4. Double-close of LadybugDB pools — remove blanket closeLbug() from
   cli/group.ts; sync.ts per-id cleanup is sufficient

Tests: 22 new tests across 5 files. Full suite: 4706 passed, 0 failed.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 12:55:33 +03:00
ivkondandClaude Opus 4.6 be4458b650 docs: add design spec and implementation plan for PR #626 HIGH fixes
Spec covers 4 HIGH-priority issues from review: path traversal via
group name, gRPC proto regex nested braces, service boundary detector
directory exclusions, double-close of LadybugDB pools.

Plan: 6 tasks with TDD, ordered by complexity.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 12:55:18 +03:00
ivkondandClaude Opus 4.6 4fed097abb feat(group): add sync pipeline, CLI, MCP tools, and monorepo fixture
Wire extractors into the sync pipeline with service boundary detection.
GroupService provides high-level API for all group operations.

- Sync pipeline: orchestrates extraction (HTTP, gRPC, topics) with
  service boundary assignment and exact matching
- GroupService: groupList, groupSync, groupContracts, groupQuery,
  groupStatus (groupImpact deferred to cross-repo follow-up PR)
- CLI: group create/add/remove/list/sync/contracts/query/status
- MCP tools: group_list, group_sync, group_contracts, group_query,
  group_status
- Monorepo fixture: 3 services (auth/orders/gateway) connected via
  gRPC + Kafka + HTTP — all intra-repo cross-links discovered
- Documentation: CLI commands and MCP tools added to both READMEs

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 00:40:31 +03:00
ivkondandClaude Opus 4.6 4fa395f4b6 feat(group): add service boundary detection and contract extractors
Service communication detection for microservice monorepos:

- ServiceBoundaryDetector: auto-detects service boundaries via markers
  (package.json, go.mod, Dockerfile, pom.xml, Cargo.toml, build.gradle,
  pyproject.toml, etc.)
- HttpRouteExtractor: graph-assisted (Strategy A) with source-scan
  fallback (Strategy B) for Spring, Express, Laravel, FastAPI providers
  and fetch/axios consumers
- GrpcExtractor: parses .proto files, detects Go/Java/Python/TS gRPC
  servers (RegisterXxxServer, @GrpcService, add_XxxServicer_to_server,
  @GrpcMethod) and clients (NewXxxClient, newBlockingStub, XxxStub)
- TopicExtractor: Kafka (@KafkaListener, producer.send), RabbitMQ
  (@RabbitListener, channel.publish/consume), NATS (nc.Subscribe/Publish)
  across Java, Node, Go, and Python

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 00:40:12 +03:00
ivkondandClaude Opus 4.6 52277247fe feat(group): add group infrastructure and contract matching
Core foundation for repository group analysis:
- Type system: ContractType, ExtractedContract, StoredContract, CrossLink
  with optional `service` field for intra-repo matching
- Config parser for group.yaml (repos, detection flags, matching thresholds)
- Contract registry storage with atomic writes
- Exact matching engine with per-type normalization (HTTP, gRPC, topic)
  and intra-repo support (different services within same repo can match)
- Extract LadybugDB pool-adapter from MCP backend for reuse by sync pipeline
- Git staleness checker for group status reporting

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 00:39:43 +03:00
Gergő Magyar ba5de0bde4 feat(cpp): C/C++ MethodExtractor config with pure virtual detection (#617)
* feat(cpp): C/C++ MethodExtractor config with pure virtual detection (#572)

- Pure virtual (= 0) detected as isAbstract via token scanning
- virtual/final/override via hasKeyword and virtual_specifier children
- Access specifier visibility via backward sibling walk (public:/private:/protected:)
- Pointer/reference parameter types extracted correctly
- Constructor and destructor support via declaration node type
- Static detection via storage_class_specifier
- 16 new tests covering all acceptance criteria

* fix(cpp): isVirtual infers from override/final + out-of-class resolution

- isVirtual returns true for override/final methods (C++ mandates these
  are virtual)
- Add findClassNodeByQualifiedName to parse-worker: resolves Foo::bar()
  back to the Foo class declaration for method extractor enrichment
- Handles pointer/ref return types, constructors, destructors
- Integration test for virtual/static/constructor inline methods
- 233 unit+integration tests pass, 97 C++ resolver tests pass

* fix(cpp): address review — deep pointers, templates, unions, trailing returns

- Fix extractParamName: recursive unwrap for int** ptr → "ptr" (not "**ptr")
- Fix findFunctionDeclarator: recursive unwrap for multi-level pointer chains
- Template methods: generic extractor unwraps template_declaration to inner node
- union_specifier: added to typeDeclarationNodes, visibility defaults to public
- Trailing return type: auto foo() -> T now extracts T instead of "auto"
- Fix version comment: ^0.22.4 → ^0.23.4 to match package.json
- 4 new tests: double pointer params, template methods, union methods, trailing returns

* fix(cpp): template method visibility + union isTypeDeclaration test

extractCppVisibility now walks from the template_declaration parent
when the node is wrapped by a template, restoring correct access-
specifier resolution for templated class methods.

Also adds missing isTypeDeclaration assertion for union_specifier and
expands the template method test with explicit visibility checks.

* fix(cpp): address deep gap analysis review findings

- findClassNodeByQualifiedName: recursive pointer/reference
  declarator unwrap, fixing out-of-class linking for deep pointer
  return types (e.g. int** Foo::bar())
- findClassNodeByQualifiedName: recurse into namespace_definition
  blocks so namespace-wrapped classes resolve correctly
- Suppress = delete / = default special members from extraction
  via delete_method_clause / default_method_clause node detection
- Update known-gaps: namespace-wrapped classes, const-overload collapse
- Add tree-sitter-c version comment for consistency
- toBeFalsy() → toBe(undefined) for precise isVirtual assertion
- Tests: = delete, = default, = 0 non-regression, operator overloads,
  deep pointer return types, default visibility (class vs struct),
  multiple access specifier sections
2026-04-01 18:07:11 +01:00
Abhigyan PatwariandClaude Opus 4.6 dc86ea96dc chore: release v1.5.3 — update CHANGELOG and package-lock
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 21:35:24 +05:30
Abhigyan PatwariandClaude Opus 4.6 d1adc8331a chore: release v1.6.0 — update CHANGELOG and package-lock
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 21:32:11 +05:30
80d363f145 fix(wiki): Azure OpenAI compat and HTML viewer script injection (#618)
* fix(wiki): Azure OpenAI compat and HTML viewer script injection

- Use max_completion_tokens instead of deprecated max_tokens for all models
- Skip sending temperature for Azure provider (some models reject non-default values)
- Simplify Azure interactive setup: endpoint + deployment + key (3 prompts instead of 7)
- Escape </script> in embedded JSON to prevent premature script tag closure

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(test): align wiki-llm-client test with max_completion_tokens change

The test expected max_tokens for non-reasoning models, but the source
now uses max_completion_tokens for all models since max_tokens is
deprecated by newer OpenAI models.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Abhigyan Patwari <abhigyan@Abhigyans-MacBook-Air.local>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 21:30:43 +05:30
Gergő Magyar 12be2025f1 feat(ts,js): TypeScript/JavaScript MethodExtractor config (#588)
* feat(ts,js): MethodExtractor config for TypeScript and JavaScript (#570)

Add per-language method extraction config following the established
JVM and C# patterns. Shared config base mirrors the field extractor's
typescript-javascript.ts pattern — TS-only node types are harmless
no-ops for JS.

Key features:
- isAbstract for abstract class methods and interface methods
- Parameter extraction with isOptional (?:, defaults) and isVariadic (...)
- Decorator extraction from preceding body-level siblings
- isAsync and isOverride detection
- Visibility via accessibility_modifier two-pass pattern
- Return type extraction unwrapping type_annotation

* test(ts,js): add override, getter/setter, destructured param tests

Address code review findings:
- Add override method detection test
- Add getter/setter extraction test
- Add destructured parameter with type annotation test
- Tighten constructor and private method assertions

* refactor(ts,js): address code review findings

- Replace O(M*N) decorator index scan with previousNamedSibling walk
- Remove dead findVisibility 'modifiers' fallback (TS uses
  accessibility_modifier, not a modifiers wrapper)
- Document call_signature/construct_signature as known gaps
- Document that TS constructors are method_definition nodes
- Remove unused findVisibility import

* fix(ts,js): type guard before cast, add generator/computed/overload tests

- Use type guard pattern (Set.has check before as-cast) in visibility
  extraction to ensure string is validated before narrowing
- Add generator method test (*items()) — confirms extraction works
- Add computed property name test ([Symbol.iterator]) — documents
  bracket-in-name behavior as intentional
- Add class-level method overload test — verifies overload signatures
  + implementation are all extracted

* fix(ts,js): detect #private methods as visibility 'private'

ES2022 private class methods (#name) use private_property_identifier
as their name node type. Detect this and return 'private' visibility
instead of the default 'public'.

* fix(ts,js): address review findings + close ingestion gaps

- hasKeyword/findVisibility: skip name field child to prevent false
  positives on soft-keyword method names (e.g. `abstract()`, `static()`)
- extractTsJsParameters: filter TS `this` parameter (compile-time only)
- extractMethodSignature: mirror `this`-param skip in fallback path
- tree-sitter queries: capture abstract_method_signature,
  method_signature, and private_property_identifier for TS; add
  private_property_identifier for JS
- Remove dead childForFieldName('name') fallbacks and typeFromAnnotation
  fallback
- Add 10+ unit tests, 4 integration tests through query pipeline

* test(ts): update HAS_METHOD count for interface method_signature capture

The new method_signature query now captures ILogger.log() as a Method
node with a HAS_METHOD edge, increasing the expected count from 4 to 5.

* fix(ts,js): address second review — async generator test, declare module gap

- Add async generator method test (async *values() → isAsync: true)
- Document declare module/global augmentation as known gap
2026-04-01 14:09:59 +01:00
385 changed files with 33299 additions and 2934 deletions
+1 -1
View File
@@ -63,7 +63,7 @@ Generic “core standards” playbooks are often long and stack-specific. For th
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus**. Use the GitNexus MCP tools to understand code, assess impact, and navigate safely. For current symbol stats, run `npx gitnexus analyze` and inspect `.gitnexus/meta.json`.
This project is indexed by GitNexus as **GitNexus** (3298 symbols, 7954 relationships, 185 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
+57
View File
@@ -58,8 +58,65 @@ This repository is a **monorepo** with two main products: the **CLI / MCP packag
| Web UI behavior | `gitnexus-web/src/` (components, workers, graph client). |
| CI | `.github/workflows/*.yml`, `.github/actions/setup-gitnexus/`. |
## Known limitations
### Overloaded method resolution
Method and Constructor node IDs include an arity suffix (`#<paramCount>`) to
disambiguate overloaded methods. Two overloads with different parameter counts
produce distinct graph nodes: `Method:file:Class.method#1` vs
`Method:file:Class.method#2`.
**Same-arity overload disambiguation:** When two overloads share the same
parameter count but differ in types (e.g. `save(int)` vs `save(String)`), a
type-hash suffix `~type1,type2` is appended to produce distinct node IDs:
`Method:file:Class.save#1~int` vs `Method:file:Class.save#1~String`. The suffix
is only added when a same-arity collision is detected within a class and all
parameters have non-null type annotations. Languages without type info (Python,
Ruby, JS) fall back to arity-only IDs. TypeScript/JavaScript overload signatures
are intentionally excluded from type-hashing because they are declaration-only
contracts that should collapse to the implementation body's node ID. See issue
\#651.
**C++ const-qualified overload disambiguation:** Methods overloaded by const
qualification (e.g. `begin()` vs `begin() const`) are disambiguated via an
`isConst` property and a `$const` ID suffix appended to the const-qualified
variant when a non-const collision exists. The `$const` suffix appears after the
type-hash suffix: e.g. `Method:file:Container.begin#0$const`.
**Generic/template type preservation in type-hash:** The type-hash suffix uses
`rawType` (full AST text including generic/template args) rather than the
simplified `type` from `extractSimpleTypeName`. This means C++ template overloads
like `process(vector<int>)` vs `process(vector<string>)` produce distinct IDs:
`~vector<int>` vs `~vector<std::string>`. Java generic overloads like
`process(List<String>)` vs `process(List<Integer>)` are a compile error due to
type erasure, so this gap is theoretical for Java.
**ID stability on first overload:** Type and const tags are collision-only. When
a class has `save(int)` as its only `save` method, the ID is `save#1` (no tag).
Adding `save(String)` changes the original to `save#1~int`. This is correct for
fresh analysis but means IDs are not stable across overload additions. Future
incremental re-analysis should account for this.
**Variadic method matching:** When one side is variadic (`parameterCount`
undefined) and the other has a fixed count, `METHOD_IMPLEMENTS` edges are
emitted with confidence 0.7 instead of 1.0. Variadic methods like
`foo(String... args)` may superficially match `foo(String s)` by type but
are not guaranteed to be interchangeable across all languages (Java/Kotlin
accept this via varargs sugar; TypeScript, C#, Rust do not).
**Confidence tiering** for `METHOD_IMPLEMENTS` edges:
| Match quality | Confidence | When |
|---|---|---|
| Exact parameter types match | 1.0 | Both sides have `parameterTypes` arrays and they match |
| Arity (count) matches | 1.0 | Both sides have `parameterCount`, types unavailable |
| Variadic vs fixed | 0.7 | One side is variadic, other has fixed count |
| Lenient (insufficient info) | 0.7 | One or both sides lack type and count data |
## Related docs
- [MIGRATION.md](MIGRATION.md) — breaking changes and migration guidance.
- [RUNBOOK.md](RUNBOOK.md) — operational commands and recovery.
- [GUARDRAILS.md](GUARDRAILS.md) — safety boundaries for humans and agents.
- [TESTING.md](TESTING.md) — how to run tests.
+13
View File
@@ -10,6 +10,19 @@ All notable changes to GitNexus will be documented in this file.
- Added automatic cleanup of stale KuzuDB index files
- LadybugDB v0.15 requires explicit VECTOR extension loading for semantic search
## [1.5.3] - 2026-04-01
### Added
- **TypeScript/JavaScript MethodExtractor config** — shared extraction config covering abstract methods, visibility modifiers, async/override keywords, decorators, rest/optional/destructured parameters, and return types (#588) — @compound-ai
### Fixed
- **Azure OpenAI compatibility** — use `max_completion_tokens` instead of deprecated `max_tokens` (newer models reject `max_tokens`); skip `temperature` for Azure provider (some models reject non-default values) (#618)
- **Simplified Azure interactive setup** — 3 prompts (endpoint, deployment, key) instead of 7 (#618)
- **Wiki HTML viewer script injection** — escape `</script>` in embedded JSON so LLM-generated markdown no longer breaks the viewer (#618)
- Ensure import rewrites survive npm publish lifecycle
## [1.4.0] - 2026-03-13
### Added
+103 -1
View File
@@ -49,4 +49,106 @@ If always-on instructions grow, load deep conventions via conditional reads (e.g
## GitNexus rules
GitNexus MCP rules are in the `<!-- gitnexus:start -->` … `<!-- gitnexus:end -->` block in **[AGENTS.md](AGENTS.md)** — load that section when working with MCP tools or the graph index.
GitNexus MCP rules are in the `<!-- gitnexus:start -->` … `<!-- gitnexus:end -->` block in **[AGENTS.md](AGENTS.md)** — load that section when working with MCP tools or the graph index.
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (3298 symbols, 7954 relationships, 185 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## When Debugging
1. `gitnexus_query({query: "<error or symptom>"})` — find execution flows related to the issue
2. `gitnexus_context({name: "<suspect function>"})` — see all callers, callees, and process participation
3. `READ gitnexus://repo/GitNexus/process/{processName}` — trace the full execution flow step by step
4. For regressions: `gitnexus_detect_changes({scope: "compare", base_ref: "main"})` — see what your branch changed
## When Refactoring
- **Renaming**: MUST use `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` first. Review the preview — graph edits are safe, text_search edits need manual review. Then run with `dry_run: false`.
- **Extracting/Splitting**: MUST run `gitnexus_context({name: "target"})` to see all incoming/outgoing refs, then `gitnexus_impact({target: "target", direction: "upstream"})` to find all external callers before moving code.
- After any refactor: run `gitnexus_detect_changes({scope: "all"})` to verify only expected files changed.
## Never Do
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Tools Quick Reference
| Tool | When to use | Command |
|------|-------------|---------|
| `query` | Find code by concept | `gitnexus_query({query: "auth validation"})` |
| `context` | 360-degree view of one symbol | `gitnexus_context({name: "validateUser"})` |
| `impact` | Blast radius before editing | `gitnexus_impact({target: "X", direction: "upstream"})` |
| `detect_changes` | Pre-commit scope check | `gitnexus_detect_changes({scope: "staged"})` |
| `rename` | Safe multi-file rename | `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` |
| `cypher` | Custom graph queries | `gitnexus_cypher({query: "MATCH ..."})` |
## Impact Risk Levels
| Depth | Meaning | Action |
|-------|---------|--------|
| d=1 | WILL BREAK — direct callers/importers | MUST update these |
| d=2 | LIKELY AFFECTED — indirect deps | Should test |
| d=3 | MAY NEED TESTING — transitive | Test if critical path |
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
## Self-Check Before Finishing
Before completing any code modification task, verify:
1. `gitnexus_impact` was run for all modified symbols
2. No HIGH/CRITICAL risk warnings were ignored
3. `gitnexus_detect_changes()` confirms changes match expected scope
4. All d=1 (WILL BREAK) dependents were updated
## Keeping the Index Fresh
After committing code changes, the GitNexus index becomes stale. Re-run analyze to update it:
```bash
npx gitnexus analyze
```
If the index previously included embeddings, preserve them by adding `--embeddings`:
```bash
npx gitnexus analyze --embeddings
```
To check whether embeddings exist, inspect `.gitnexus/meta.json` — the `stats.embeddings` field shows the count (0 means no embeddings). **Running analyze without `--embeddings` will delete any previously generated embeddings.**
> Claude Code users: A PostToolUse hook handles this automatically after `git commit` and `git merge`.
## CLI
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
<!-- gitnexus:end -->
+27
View File
@@ -0,0 +1,27 @@
# Migration Guide
## OVERRIDES → METHOD_OVERRIDES (PR #642)
The `OVERRIDES` relationship type has been renamed to `METHOD_OVERRIDES` for
consistency with the new `METHOD_IMPLEMENTS` edge type.
### Do I need to migrate?
**No.** Backward compatibility is handled automatically at runtime:
- `local-backend.ts` dual-reads both `OVERRIDES` and `METHOD_OVERRIDES` in all
impact-analysis and context queries. Existing stored graphs with `OVERRIDES`
edges continue to return correct results without any manual intervention.
- The `REL_TYPES` array in `schema-constants.ts` includes both names so Cypher
queries that reference either will work.
### What happens on re-index?
Running `npx gitnexus analyze` on a repository produces `METHOD_OVERRIDES`
edges going forward. The old `OVERRIDES` edges are replaced as part of the
normal full re-index.
### When will the legacy alias be removed?
The `OVERRIDES` compat alias will remain until a future major version. Removal
will be announced in this file and in the changelog before it happens.
+17 -3
View File
@@ -52,7 +52,7 @@ https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
| **What** | Index repos locally, connect AI agents via MCP | Visual graph explorer + AI chat in browser |
| **For** | Daily development with Cursor, Claude Code, Codex, Windsurf, OpenCode | Quick exploration, demos, one-off analysis |
| **Scale** | Full repos, any size | Limited by browser memory (~5k files), or unlimited via backend mode |
| **Install** | `npm install -g gitnexus` | No install —[gitnexus.vercel.app](https://gitnexus.vercel.app) |
| **Install** | `npm install -g gitnexus` | No install — [gitnexus.vercel.app](https://gitnexus.vercel.app) |
| **Storage** | LadybugDB native (fast, persistent) | LadybugDB WASM (in-memory, per session) |
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
| **Privacy** | Everything local, no network | Everything in-browser, no server |
@@ -119,7 +119,6 @@ To configure MCP for your editor, run `npx gitnexus setup` once — or set it up
| **Codex** | Yes | Yes | — | MCP + Skills |
| **Windsurf** | Yes | — | — | MCP |
| **OpenCode** | Yes | Yes | — | MCP + Skills |
| **Codex** | Yes | — | — | MCP |
> **Claude Code** gets the deepest integration: MCP tools + agent skills + PreToolUse hooks that enrich searches with graph context + PostToolUse hooks that auto-reindex after commits.
@@ -206,11 +205,21 @@ gitnexus clean --all --force # Delete all indexes
gitnexus wiki [path] # Generate repository wiki from knowledge graph
gitnexus wiki --model <model> # Wiki with custom LLM model (default: gpt-4o-mini)
gitnexus wiki --base-url <url> # Wiki with custom LLM API base URL
# Repository groups (multi-repo / monorepo service tracking)
gitnexus group create <name> # Create a repository group
gitnexus group add <name> <repo> # Add a repo to a group
gitnexus group remove <name> <repo> # Remove a repo from a group
gitnexus group list [name] # List groups, or show one group's config
gitnexus group sync <name> # Extract contracts and match across repos/services
gitnexus group contracts <name> # Inspect extracted contracts and cross-links
gitnexus group query <name> <q> # Search execution flows across all repos in a group
gitnexus group status <name> # Check staleness of repos in a group
```
### What Your AI Agent Gets
**7 tools** exposed via MCP:
**16 tools** exposed via MCP (11 per-repo + 5 group):
| Tool | What It Does | `repo` Param |
| ------------------ | ----------------------------------------------------------------- | -------------- |
@@ -221,6 +230,11 @@ gitnexus wiki --base-url <url> # Wiki with custom LLM API base URL
| `detect_changes` | Git-diff impact — maps changed lines to affected processes | Optional |
| `rename` | Multi-file coordinated rename with graph + text search | Optional |
| `cypher` | Raw Cypher graph queries | Optional |
| `group_list` | List configured repository groups | — |
| `group_sync` | Extract contracts and match across repos/services | — |
| `group_contracts`| Inspect extracted contracts and cross-links | — |
| `group_query` | Search execution flows across all repos in a group | — |
| `group_status` | Check staleness of repos in a group | — |
> When only one repo is indexed, the `repo` parameter is optional. With multiple repos, specify which one: `query({query: "auth", repo: "my-app"})`.
@@ -0,0 +1,725 @@
# PR #626 HIGH-Priority Fixes Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Fix 4 HIGH-priority issues from PR #626 code review before merge.
**Architecture:** Minimal targeted fixes — each task is independent. TDD: tests first, then implementation. No refactoring beyond what's needed.
**Tech Stack:** TypeScript, Vitest, Node.js fs/path APIs
**Spec:** `docs/superpowers/specs/2026-04-02-pr626-high-fixes-design.md`
**Paths:** All file paths are relative to the monorepo root (`GitNexus/`). Git commands run from the root. The `gitnexus/` prefix is a package subdirectory, not a separate repo.
---
### Task 1: Path Traversal — Validate Group Name
**Files:**
- Modify: `gitnexus/src/core/group/storage.ts:17-19` (getGroupDir) and `:63-68` (createGroupDir)
- Test: `gitnexus/test/unit/group/storage.test.ts`
- [ ] **Step 1: Write failing tests for validateGroupName**
In `gitnexus/test/unit/group/storage.test.ts`, add `createGroupDir` and `validateGroupName` to the existing import from `'../../../src/core/group/storage.js'` (line 6-11). Then add these describe blocks at the end of the outer `describe('Group storage', ...)`:
```typescript
describe('validateGroupName', () => {
it('test_validateGroupName_traversal_path_throws', () => {
expect(() => validateGroupName('../../evil')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_slash_in_name_throws', () => {
expect(() => validateGroupName('foo/bar')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_empty_string_throws', () => {
expect(() => validateGroupName('')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_starts_with_dash_throws', () => {
expect(() => validateGroupName('-leading-dash')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_starts_with_underscore_throws', () => {
expect(() => validateGroupName('_leading')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_dots_throws', () => {
expect(() => validateGroupName('com.example')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_valid_alphanumeric_passes', () => {
expect(() => validateGroupName('my-group_01')).not.toThrow();
});
it('test_validateGroupName_single_char_passes', () => {
expect(() => validateGroupName('A')).not.toThrow();
});
it('test_validateGroupName_all_digits_passes', () => {
expect(() => validateGroupName('123')).not.toThrow();
});
});
describe('getGroupDir rejects invalid names', () => {
it('test_getGroupDir_traversal_throws', () => {
expect(() => getGroupDir(tmpDir, '../../etc')).toThrow(/Invalid group name/);
});
it('test_getGroupDir_valid_name_returns_path', () => {
const dir = getGroupDir(tmpDir, 'company');
expect(dir).toBe(path.join(tmpDir, 'groups', 'company'));
});
});
describe('createGroupDir rejects invalid names', () => {
it('test_createGroupDir_traversal_throws', async () => {
await expect(createGroupDir(tmpDir, '../evil')).rejects.toThrow(/Invalid group name/);
});
});
```
- [ ] **Step 2: Run tests to verify they fail**
Run: `cd gitnexus && npx vitest run test/unit/group/storage.test.ts`
Expected: FAIL — `validateGroupName` is not exported, `getGroupDir` does not throw.
- [ ] **Step 3: Implement validateGroupName and wire into getGroupDir and createGroupDir**
In `gitnexus/src/core/group/storage.ts`, add the validation function before `getGroupDir` and call it:
```typescript
const GROUP_NAME_RE = /^[a-zA-Z0-9][a-zA-Z0-9_-]*$/;
export function validateGroupName(name: string): void {
if (!GROUP_NAME_RE.test(name)) {
throw new Error(
`Invalid group name "${name}". Names must start with a letter or digit and contain only [a-zA-Z0-9_-].`,
);
}
}
export function getGroupDir(gitnexusDir: string, groupName: string): string {
validateGroupName(groupName);
return path.join(gitnexusDir, 'groups', groupName);
}
```
`createGroupDir` already calls `getGroupDir` at line 68, so it inherits validation automatically. No change needed in `createGroupDir`.
- [ ] **Step 4: Run tests to verify they pass**
Run: `cd gitnexus && npx vitest run test/unit/group/storage.test.ts`
Expected: ALL PASS
- [ ] **Step 5: Commit**
```bash
cd gitnexus && git add src/core/group/storage.ts test/unit/group/storage.test.ts
git commit -m "fix(group): validate group name to prevent path traversal
Add validateGroupName() with regex [a-zA-Z0-9][a-zA-Z0-9_-]*.
Called in getGroupDir (defense in depth) which covers all CLI entry
points: create, add, remove, status, sync.
Addresses PR #626 review item 1 (HIGH).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
---
### Task 2: Directory Exclusions in Service Boundary Detector
**Files:**
- Modify: `gitnexus/src/core/group/service-boundary-detector.ts:24-51` (add constant), `:78` (walkForBoundaries), `:130` (hasSourceFilesInSubdirs)
- Test: `gitnexus/test/unit/group/service-boundary-detector.test.ts`
- [ ] **Step 1: Write failing tests for excluded directories**
Add this describe block inside the existing `detectServiceBoundaries` describe in `gitnexus/test/unit/group/service-boundary-detector.test.ts`:
```typescript
it('test_detect_skips_vendor_directory', async () => {
writeFile('services/auth/package.json', '{}');
writeFile('services/auth/src/index.ts', '');
// vendor should be skipped — its contents should not create a boundary
writeFile('vendor/some-dep/package.json', '{}');
writeFile('vendor/some-dep/src/lib.go', '');
const boundaries = await detectServiceBoundaries(tmpDir);
const paths = boundaries.map((b) => b.servicePath);
expect(paths).toContain('services/auth');
expect(paths).not.toContain('vendor/some-dep');
});
it('test_detect_skips_target_directory', async () => {
writeFile('services/api/go.mod', 'module api');
writeFile('services/api/main.go', '');
writeFile('target/classes/Main.java', '');
writeFile('target/pom.xml', '<project/>');
const boundaries = await detectServiceBoundaries(tmpDir);
const paths = boundaries.map((b) => b.servicePath);
expect(paths).toContain('services/api');
expect(paths).not.toContain('target');
});
it('test_detect_skips_pycache_directory', async () => {
writeFile('services/ml/pyproject.toml', '[project]');
writeFile('services/ml/model.py', '');
// __pycache__ with a marker + source files — would be detected as
// a boundary if not excluded, since it has package.json + .py file
writeFile('__pycache__/package.json', '{}');
writeFile('__pycache__/cached.py', '');
const boundaries = await detectServiceBoundaries(tmpDir);
const paths = boundaries.map((b) => b.servicePath);
expect(paths).toContain('services/ml');
expect(paths.every((p) => !p.includes('__pycache__'))).toBe(true);
});
it('test_detect_skips_dotfile_directories_regression', async () => {
writeFile('services/api/package.json', '{}');
writeFile('services/api/src/index.ts', '');
writeFile('.hidden/package.json', '{}');
writeFile('.hidden/src/index.ts', '');
const boundaries = await detectServiceBoundaries(tmpDir);
const paths = boundaries.map((b) => b.servicePath);
expect(paths).toContain('services/api');
expect(paths).not.toContain('.hidden');
});
it('test_detect_does_not_skip_regular_source_directories', async () => {
writeFile('services/api/package.json', '{}');
writeFile('services/api/src/index.ts', '');
const boundaries = await detectServiceBoundaries(tmpDir);
expect(boundaries).toHaveLength(1);
expect(boundaries[0].serviceName).toBe('api');
});
```
- [ ] **Step 2: Run tests to verify `vendor` and `target` tests fail**
Run: `cd gitnexus && npx vitest run test/unit/group/service-boundary-detector.test.ts`
Expected: `test_detect_skips_vendor_directory` and `test_detect_skips_target_directory` FAIL (vendor/target not excluded). Other new tests may pass since dotfile exclusion already exists.
- [ ] **Step 3: Add EXCLUDED_DIRS constant and update both walking functions**
In `gitnexus/src/core/group/service-boundary-detector.ts`:
After `SOURCE_EXTENSIONS` (after line 51), add:
```typescript
const EXCLUDED_DIRS = new Set([
'node_modules',
'vendor',
'target',
'build',
'dist',
'__pycache__',
'.venv',
'venv',
'.tox',
'.mypy_cache',
'.gradle',
'.mvn',
'out',
'bin',
]);
```
In `walkForBoundaries`, replace line 78:
```typescript
if (entry.name.startsWith('.') || entry.name === 'node_modules') continue;
```
with:
```typescript
if (entry.name.startsWith('.') || EXCLUDED_DIRS.has(entry.name)) continue;
```
In `hasSourceFilesInSubdirs`, replace line 130:
```typescript
if (entry.isDirectory() && !entry.name.startsWith('.') && entry.name !== 'node_modules') {
```
with:
```typescript
if (entry.isDirectory() && !entry.name.startsWith('.') && !EXCLUDED_DIRS.has(entry.name)) {
```
- [ ] **Step 4: Run tests to verify they pass**
Run: `cd gitnexus && npx vitest run test/unit/group/service-boundary-detector.test.ts`
Expected: ALL PASS
- [ ] **Step 5: Commit**
```bash
cd gitnexus && git add src/core/group/service-boundary-detector.ts test/unit/group/service-boundary-detector.test.ts
git commit -m "fix(group): add directory exclusions to service boundary detector
Add EXCLUDED_DIRS set: vendor, target, build, dist, __pycache__,
.venv, venv, .tox, .mypy_cache, .gradle, .mvn, out, bin.
Applied in walkForBoundaries and hasSourceFilesInSubdirs.
Replaces inline node_modules check.
Addresses PR #626 review item 3 (HIGH).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
---
### Task 3: Remove Double-Close of LadybugDB Pools
**Files:**
- Modify: `gitnexus/src/cli/group.ts:160` (remove import), `:187-189` (remove finally block body)
- Test: `gitnexus/test/unit/group/sync.test.ts` (add pool cleanup test)
- Test: `gitnexus/test/integration/group/group-cli.test.ts` (verify no blanket close in source)
- [ ] **Step 1: Write unit tests for per-id pool cleanup in sync.ts**
Add to `gitnexus/test/unit/group/sync.test.ts`, inside the existing `describe('syncGroup', ...)`:
```typescript
it('test_syncGroup_closes_only_opened_pools', async () => {
const config = makeConfig({
'app/backend': 'backend-repo',
'app/frontend': 'frontend-repo',
});
const closedIds: string[] = [];
// Mock initLbug/closeLbug via per-repo override that tracks pool lifecycle
const { vi } = await import('vitest');
const poolAdapter = await import('../../../src/core/lbug/pool-adapter.js');
const initSpy = vi.spyOn(poolAdapter, 'initLbug').mockResolvedValue(undefined);
const closeSpy = vi.spyOn(poolAdapter, 'closeLbug').mockImplementation(async (id?: string) => {
if (id) closedIds.push(id);
});
try {
await syncGroup(config, {
resolveRepoHandle: async (_name, groupPath) => ({
id: groupPath.replace(/\//g, '-'),
path: groupPath,
repoPath: '/tmp/' + groupPath,
storagePath: '/tmp/' + groupPath + '/.gitnexus',
}),
skipWrite: true,
}).catch(() => {});
// Regardless of extraction errors, closeLbug should be called per id
// closeLbug should only receive specific pool ids, never undefined/empty
for (const id of closedIds) {
expect(id).toBeTruthy();
expect(typeof id).toBe('string');
}
// No blanket close (no-arg call)
const blanketCalls = closeSpy.mock.calls.filter((args) => args.length === 0 || !args[0]);
expect(blanketCalls).toHaveLength(0);
} finally {
initSpy.mockRestore();
closeSpy.mockRestore();
}
});
```
- [ ] **Step 2: Run sync unit test to verify it passes (sync.ts already does per-id cleanup)**
Run: `cd gitnexus && npx vitest run test/unit/group/sync.test.ts`
Expected: PASS — sync.ts already cleans up correctly. This test locks the behavior.
- [ ] **Step 3: Write test verifying CLI source has no blanket closeLbug()**
Add to `gitnexus/test/integration/group/group-cli.test.ts`:
```typescript
it('test_sync_command_source_does_not_call_blanket_closeLbug', () => {
const cliGroupPath = path.join(repoRoot, 'src', 'cli', 'group.ts');
const source = fs.readFileSync(cliGroupPath, 'utf-8');
// closeLbug() without arguments (blanket close) must not appear.
// closeLbug(id) with argument is fine (that's in sync.ts, not here).
// Match closeLbug() but not closeLbug(someArg)
const blanketClosePattern = /closeLbug\s*\(\s*\)/;
expect(source).not.toMatch(blanketClosePattern);
});
```
- [ ] **Step 4: Run test to verify it fails**
Run: `cd gitnexus && npx vitest run test/integration/group/group-cli.test.ts`
Expected: FAIL — `closeLbug()` (no args) exists at line 188.
- [ ] **Step 5: Remove blanket closeLbug() from cli/group.ts**
In `gitnexus/src/cli/group.ts`:
Remove the `closeLbug` import at line 160:
```typescript
const { closeLbug } = await import('../core/lbug/pool-adapter.js');
```
Replace the try/finally wrapper (lines 162-189):
```typescript
try {
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const config = await loadGroupConfig(groupDir);
console.log(`Syncing group "${name}" (${Object.keys(config.repos).length} repos)...\n`);
const result = await syncGroup(config, {
groupDir,
allowStale: Boolean(opts.allowStale),
verbose: Boolean(opts.verbose),
skipEmbeddings: Boolean(opts.skipEmbeddings),
exactOnly: Boolean(opts.exactOnly),
});
if (opts.json) {
console.log(JSON.stringify(result, null, 2));
} else {
console.log(`\nMatching cascade:`);
const exactLinks = result.crossLinks.filter((l) => l.matchType === 'exact');
console.log(` exact: ${exactLinks.length} cross-links (confidence 1.0)`);
console.log(` unmatched: ${result.unmatched.length} contracts`);
console.log(
`\nWrote contracts.json (${result.contracts.length} contracts, ${result.crossLinks.length} cross-links)`,
);
}
} finally {
await closeLbug().catch(() => {});
}
```
Becomes (remove try/finally entirely, since sync.ts handles its own cleanup):
```typescript
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const config = await loadGroupConfig(groupDir);
console.log(`Syncing group "${name}" (${Object.keys(config.repos).length} repos)...\n`);
const result = await syncGroup(config, {
groupDir,
allowStale: Boolean(opts.allowStale),
verbose: Boolean(opts.verbose),
skipEmbeddings: Boolean(opts.skipEmbeddings),
exactOnly: Boolean(opts.exactOnly),
});
if (opts.json) {
console.log(JSON.stringify(result, null, 2));
} else {
console.log(`\nMatching cascade:`);
const exactLinks = result.crossLinks.filter((l) => l.matchType === 'exact');
console.log(` exact: ${exactLinks.length} cross-links (confidence 1.0)`);
console.log(` unmatched: ${result.unmatched.length} contracts`);
console.log(
`\nWrote contracts.json (${result.contracts.length} contracts, ${result.crossLinks.length} cross-links)`,
);
}
```
- [ ] **Step 6: Run tests to verify they pass**
Run: `cd gitnexus && npx vitest run test/integration/group/group-cli.test.ts test/unit/group/sync.test.ts`
Expected: ALL PASS
- [ ] **Step 7: Commit**
```bash
cd gitnexus && git add src/cli/group.ts test/integration/group/group-cli.test.ts test/unit/group/sync.test.ts
git commit -m "fix(group): remove blanket closeLbug() from CLI sync command
sync.ts already closes pools per-id in its finally block.
The blanket closeLbug() in cli/group.ts tears down ALL active pools
including unrelated ones in MCP server context.
Addresses PR #626 review item 4 (HIGH).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
---
### Task 4: gRPC Proto Regex — Brace-Depth Counter
**Files:**
- Modify: `gitnexus/src/core/group/extractors/grpc-extractor.ts:101-130` (parseProtoFile)
- Test: `gitnexus/test/unit/group/grpc-extractor.test.ts`
- [ ] **Step 1: Write failing tests for nested braces in proto services**
Add this describe block inside the existing `proto file parsing` describe in `gitnexus/test/unit/group/grpc-extractor.test.ts`:
```typescript
it('test_extract_proto_with_google_api_http_nested_braces', async () => {
writeFile(
'api/gateway.proto',
`syntax = "proto3";
package gateway.v1;
import "google/api/annotations.proto";
service GatewayService {
rpc GetUser (GetUserRequest) returns (UserResponse) {
option (google.api.http) = {
get: "/v1/users/{user_id}"
};
}
rpc CreateUser (CreateUserRequest) returns (UserResponse) {
option (google.api.http) = {
post: "/v1/users"
body: "*"
};
}
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const providers = contracts.filter(
(c) => c.role === 'provider' && c.symbolRef.filePath === 'api/gateway.proto',
);
expect(providers).toHaveLength(2);
const ids = providers.map((c) => c.contractId).sort();
expect(ids).toEqual([
'grpc::gateway.v1.GatewayService/CreateUser',
'grpc::gateway.v1.GatewayService/GetUser',
]);
});
it('test_extract_proto_with_multiple_services', async () => {
writeFile(
'api/multi.proto',
`syntax = "proto3";
package multi;
service ServiceA {
rpc MethodA (Req) returns (Res);
}
service ServiceB {
rpc MethodB1 (Req) returns (Res);
rpc MethodB2 (Req) returns (Res);
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const providers = contracts.filter(
(c) => c.role === 'provider' && c.symbolRef.filePath === 'api/multi.proto',
);
expect(providers).toHaveLength(3);
const ids = providers.map((c) => c.contractId).sort();
expect(ids).toEqual([
'grpc::multi.ServiceA/MethodA',
'grpc::multi.ServiceB/MethodB1',
'grpc::multi.ServiceB/MethodB2',
]);
});
it('test_extract_proto_with_nested_option_blocks_in_rpc', async () => {
writeFile(
'api/nested.proto',
`syntax = "proto3";
package nested;
service DeepService {
rpc DeepMethod (Req) returns (Res) {
option (google.api.http) = {
post: "/v1/deep"
body: "*"
additional_bindings {
get: "/v1/deep/{id}"
}
};
}
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const providers = contracts.filter(
(c) => c.role === 'provider' && c.symbolRef.filePath === 'api/nested.proto',
);
expect(providers).toHaveLength(1);
expect(providers[0].contractId).toBe('grpc::nested.DeepService/DeepMethod');
});
it('test_extract_proto_malformed_unclosed_brace_skips_service', async () => {
writeFile(
'api/broken.proto',
`syntax = "proto3";
package broken;
service IncompleteService {
rpc SomeMethod (Req) returns (Res);
// Missing closing brace — EOF before depth returns to 0
`,
);
// Should not throw; incomplete service is silently skipped
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const providers = contracts.filter(
(c) => c.role === 'provider' && c.symbolRef.filePath === 'api/broken.proto',
);
// The old regex would find partial match; the new parser should skip it
expect(providers).toHaveLength(0);
});
```
- [ ] **Step 2: Run tests to verify the nested brace test fails**
Run: `cd gitnexus && npx vitest run test/unit/group/grpc-extractor.test.ts`
Expected: `test_extract_proto_with_google_api_http_nested_braces` FAIL — regex stops at first `}` inside the `option` block.
- [ ] **Step 3: Replace serviceRe regex with extractServiceBlocks function**
In `gitnexus/src/core/group/extractors/grpc-extractor.ts`, replace the `parseProtoFile` method (lines 101-130):
```typescript
private parseProtoFile(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
const pkgMatch = content.match(/^package\s+([\w.]+)\s*;/m);
const pkg = pkgMatch ? pkgMatch[1] : '';
for (const { name: serviceName, body } of extractServiceBlocks(content)) {
const rpcRe = /rpc\s+(\w+)\s*\(/g;
let rpcMatch: RegExpExecArray | null;
while ((rpcMatch = rpcRe.exec(body)) !== null) {
const methodName = rpcMatch[1];
const cid = contractId(pkg, serviceName, methodName);
out.push(
makeContract(cid, 'provider', filePath, `${serviceName}.${methodName}`, 0.85, {
package: pkg,
service: serviceName,
method: methodName,
source: 'proto',
}),
);
}
}
return out;
}
```
Add this function before the class (e.g. after `serviceOnlyContractId`, around line 26):
```typescript
function extractServiceBlocks(content: string): Array<{ name: string; body: string }> {
const results: Array<{ name: string; body: string }> = [];
const headerRe = /service\s+(\w+)\s*\{/g;
let headerMatch: RegExpExecArray | null;
while ((headerMatch = headerRe.exec(content)) !== null) {
const serviceName = headerMatch[1];
const bodyStart = headerMatch.index + headerMatch[0].length;
let depth = 1;
let pos = bodyStart;
while (pos < content.length && depth > 0) {
const ch = content[pos];
if (ch === '{') depth++;
else if (ch === '}') depth--;
pos++;
}
// If EOF before depth returns to 0, skip incomplete service
if (depth !== 0) continue;
// body is between opening { (consumed by regex) and closing } (pos is one past it)
const body = content.slice(bodyStart, pos - 1);
results.push({ name: serviceName, body });
}
return results;
}
```
- [ ] **Step 4: Run tests to verify they pass**
Run: `cd gitnexus && npx vitest run test/unit/group/grpc-extractor.test.ts`
Expected: ALL PASS (including existing regression tests)
- [ ] **Step 5: Commit**
```bash
cd gitnexus && git add src/core/group/extractors/grpc-extractor.ts test/unit/group/grpc-extractor.test.ts
git commit -m "fix(group): replace gRPC proto regex with brace-depth counter
The serviceRe regex used [^}]* which stopped at the first '}'.
Proto services with google.api.http annotations contain nested {}
blocks, causing methods to be missed.
New extractServiceBlocks() uses a brace-depth counter (init depth=1
after opening {, scan char-by-char). Malformed protos with unclosed
braces are silently skipped.
Addresses PR #626 review item 2 (HIGH).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
---
### Task 5: Run Full Test Suite
- [ ] **Step 1: Run all group-related tests**
Run: `cd gitnexus && npx vitest run test/unit/group/ test/integration/group/`
Expected: ALL PASS
- [ ] **Step 2: Run full test suite to catch regressions**
Run: `cd gitnexus && npx vitest run`
Expected: ALL PASS, 0 failures
- [ ] **Step 3: Run typecheck**
Run: `cd gitnexus && npx tsc --noEmit`
Expected: No errors
---
### Task 6: CLI Integration Smoke Test
- [ ] **Step 1: Add CLI smoke test for path traversal**
Add to `gitnexus/test/integration/group/group-cli.test.ts` inside the existing `group CLI` describe:
```typescript
it('test_create_with_invalid_name_fails', () => {
const result = runGroup(['create', '../../evil']);
expect(result.status).not.toBe(0);
expect(result.stderr).toContain('Invalid group name');
});
```
- [ ] **Step 2: Run test**
Run: `cd gitnexus && npx vitest run test/integration/group/group-cli.test.ts`
Expected: ALL PASS
- [ ] **Step 3: Commit**
```bash
cd gitnexus && git add test/integration/group/group-cli.test.ts
git commit -m "test(group): add CLI smoke test for path traversal rejection
Verifies that 'group create ../../evil' fails with Invalid group name.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
@@ -0,0 +1,175 @@
# PR #626 HIGH-Priority Fixes Design
**Date:** 2026-04-02
**PR:** abhigyanpatwari/GitNexus#626 — Intra-repo service communication tracking
**Scope:** 4 HIGH-priority issues identified by abhigyanpatwari and xkonjin
**Approach:** Minimal targeted fixes (option A) — no refactoring, no scope creep
---
## Fix 1: Path Traversal via Group Name
**File:** `gitnexus/src/core/group/storage.ts`
**Risk:** A group name like `../../etc` creates directories outside the intended path.
### Solution
Add `validateGroupName(name: string): void` that enforces `/^[a-zA-Z0-9][a-zA-Z0-9_-]*$/`.
- Call in `createGroupDir` (primary entry point)
- Call in `getGroupDir` (defense in depth)
- Throw descriptive error on invalid names
**Legacy:** Groups already on disk with names outside this pattern are not auto-renamed; only new `create` / resolved paths are validated.
### Why regex over path.resolve + startsWith
- abhigyanpatwari explicitly requested `[a-zA-Z0-9_-]`
- Stricter: disallows spaces, dots, Unicode edge cases
- Simpler to reason about
### Tests
- `../../evil` throws
- `foo/bar` throws
- Empty string throws
- `my-group_01` passes
- `A` (single char) passes
- CLI smoke: one integration test that hits `getGroupDir` / `createGroupDir` (e.g. `group create` or `group add`) with an invalid name proves wiring for every subcommand that resolves a group through storage
### CLI/API entry points accepting groupName
All paths flow through `getGroupDir` (which validates), so coverage is implicit. For reference:
| Command | Entry | Calls |
|---------|-------|-------|
| `group create` | `cli/group.ts` action | `createGroupDir` -> `getGroupDir` |
| `group add` | `cli/group.ts` action | `getGroupDir` |
| `group remove` | `cli/group.ts` action | `getGroupDir` |
| `group list` | `cli/group.ts` action | reads `groups/` dir directly — no traversal risk (reads, not writes) |
| `group status` | `cli/group.ts` action | `getGroupDir` |
| `group sync` | `cli/group.ts` action | `getGroupDir` |
**`listGroups`:** Reads directory names from disk without validation. Not a write path, so no traversal risk. May surface manually-created directories with non-conforming names — accepted as-is, not in scope.
---
## Fix 2: gRPC Proto Regex -> Brace-Depth Counter
**File:** `gitnexus/src/core/group/extractors/grpc-extractor.ts`
**Risk:** `serviceRe = /service\s+(\w+)\s*\{([^}]*)}/gs` stops at first `}`. Proto services with `google.api.http` annotations inside RPCs contain nested `{ }` blocks.
### Solution
Replace `serviceRe` regex with `extractServiceBlocks(content: string): Array<{ name: string; body: string }>`:
1. Use regex only to find `service <Name> {` start positions (regex consumes the opening `{`)
2. Initialise depth to 1 immediately after the opening `{`
3. Scan forward char by char: `{` -> depth++, `}` -> depth--; collect into body
4. Stop when depth reaches 0 (the matching closing `}`)
5. Return name + body pairs
Inner `rpcRe` regex remains unchanged — it operates on the already-extracted body.
**Malformed input:** If EOF is reached before `depth` returns to 0, skip the incomplete service (do not add to results). Lock this in the test.
**Scope limitation (v1):** Brace-depth only — no lexer for string literals or comments containing `{`/`}`. Sufficient for `google.api.http` annotations. Known false positive: braces inside `//` comments or quoted strings within proto options. Accepted for v1; a proper proto lexer is out of scope.
### Tests
- Proto with single service, no nesting (regression)
- Proto with `google.api.http` nested braces inside RPC options
- Proto with multiple services
- Proto with nested `option` blocks inside RPC (e.g. `google.api.http`)
- Malformed proto with unclosed brace (graceful handling)
---
## Fix 3: Directory Exclusions in Service Boundary Detector
**File:** `gitnexus/src/core/group/service-boundary-detector.ts`
**Risk:** Walks entire repo tree, only skipping dotfiles and `node_modules`. Extremely slow on repos with `vendor/`, `target/`, `__pycache__/`, `.venv/`.
### Solution
Create `EXCLUDED_DIRS` as a `Set<string>` (alongside existing `SERVICE_MARKERS`, `SOURCE_EXTENSIONS`), for example:
```text
node_modules, vendor, target, build, dist,
__pycache__, .venv, venv, .tox, .mypy_cache,
.gradle, .mvn, out, bin
```
(Implement as `new Set([...])` — the list above is the membership, not a string literal.)
Apply in both:
- `walkForBoundaries` (line 77-78) — replace current inline `=== 'node_modules'` check with `EXCLUDED_DIRS.has(entry.name)`
- `hasSourceFilesInSubdirs` (line 130) — replace `entry.name !== 'node_modules'` with `!EXCLUDED_DIRS.has(entry.name)`
Note: remove the old `=== 'node_modules'` literal from both locations — it is covered by `EXCLUDED_DIRS`.
Dotfile exclusion (`.` prefix) remains as a separate check since it's a pattern, not a name.
Exclusions apply only to `isDirectory()` entries — file names are never checked against `EXCLUDED_DIRS`.
**Tradeoff:** Rare layouts that keep source under names like `out/` or `bin/` will be skipped; accepted for performance on typical monorepos.
**Case sensitivity:** `Set.has` is case-sensitive (matches current `=== 'node_modules'` behavior). Windows case-insensitive FS not handled — accepted as-is, consistent with existing code.
### Tests
- Directory named `vendor/` is skipped
- Directory named `target/` is skipped
- Directory named `__pycache__/` is skipped
- Regular source directories are NOT skipped
- Dotfile directories still skipped (regression)
---
## Fix 4: Double-Close of LadybugDB Pools
**Files:**
- `gitnexus/src/core/group/sync.ts` (lines 155-157) — per-id cleanup (KEEP)
- `gitnexus/src/cli/group.ts` (line 188) — blanket `closeLbug()` (REMOVE)
**Risk:** In MCP server context, `closeLbug()` without arguments tears down ALL active pools, including ones from unrelated operations.
### Solution
Remove the `closeLbug()` call (no arguments) from `cli/group.ts` finally block. The per-id cleanup in `sync.ts` is sufficient:
```typescript
// sync.ts — KEEP: cleans up only pools opened by this sync
finally {
for (const id of [...new Set(openPoolIds)]) {
await closeLbug(id).catch(() => {});
}
}
```
```typescript
// cli/group.ts — REMOVE: blanket close that kills all pools
finally {
await closeLbug().catch(() => {}); // DELETE THIS
}
```
Remove the `closeLbug` import from `cli/group.ts` — after removing the `finally` call it has no remaining usages.
### Tests (unit level — mock pool adapter)
- `syncGroup` closes only the pools it opened (mock `closeLbug`, assert called with specific ids)
- Two-pool scenario: sync opens pools A and B, both closed in finally; pool C (opened elsewhere) not touched
- CLI `sync` command does not call blanket `closeLbug()` (verify no zero-arg call in source — static check or grep-based test)
---
## Out of Scope
- JSON -> LadybugDB migration (tracked in #606)
- MEDIUM/LOW issues (items 5-10 from review summary)
- Test gap coverage beyond what's needed for these 4 fixes
- Any refactoring or architectural changes
## Execution Order
Fixes are independent — can be implemented in parallel or any order.
Recommended order for review clarity: 1 -> 3 -> 4 -> 2 (simplest to most complex).
+2 -1
View File
@@ -97,7 +97,8 @@ export type RelationshipType =
| 'CONTAINS'
| 'CALLS'
| 'INHERITS'
| 'OVERRIDES'
| 'METHOD_OVERRIDES'
| 'METHOD_IMPLEMENTS'
| 'IMPORTS'
| 'USES'
| 'DEFINES'
@@ -41,6 +41,7 @@ const EXTENSION_MAP: Record<SupportedLanguages, readonly string[]> = {
[SupportedLanguages.Kotlin]: ['.kt', '.kts'],
[SupportedLanguages.Swift]: ['.swift'],
[SupportedLanguages.Dart]: ['.dart'],
[SupportedLanguages.Vue]: ['.vue'],
[SupportedLanguages.Cobol]: ['.cbl', '.cob', '.cpy', '.cobol'],
} satisfies Record<SupportedLanguages, readonly string[]>; // Ensure exhaustiveness
@@ -98,6 +99,7 @@ const SYNTAX_MAP: Record<SupportedLanguages, string> = {
[SupportedLanguages.Kotlin]: 'kotlin',
[SupportedLanguages.Swift]: 'swift',
[SupportedLanguages.Dart]: 'dart',
[SupportedLanguages.Vue]: 'typescript',
[SupportedLanguages.Cobol]: 'cobol',
} satisfies Record<SupportedLanguages, string>; // Ensure exhaustiveness
+1
View File
@@ -19,6 +19,7 @@ export enum SupportedLanguages {
Kotlin = 'kotlin',
Swift = 'swift',
Dart = 'dart',
Vue = 'vue',
/** Standalone regex processor — no tree-sitter, no LanguageProvider. */
Cobol = 'cobol',
}
+3 -1
View File
@@ -55,7 +55,9 @@ export const REL_TYPES = [
'HAS_METHOD',
'HAS_PROPERTY',
'ACCESSES',
'OVERRIDES',
'METHOD_OVERRIDES',
'OVERRIDES', // Legacy compat alias — kept until all stored indexes are migrated
'METHOD_IMPLEMENTS',
'MEMBER_OF',
'STEP_IN_PROCESS',
'HANDLES_ROUTE',
@@ -0,0 +1,109 @@
import { test, expect } from '@playwright/test';
/**
* E2E tests for heartbeat disconnect/reconnect behavior.
*
* Verifies the key regression: when the heartbeat fails, the UI shows a
* "reconnecting" banner instead of resetting to the onboarding screen.
*
* Strategy: block /api/heartbeat via route interception BEFORE loading the
* graph. The heartbeat EventSource can never connect, so onReconnecting
* fires on the first retry attempt. This reliably tests the banner behavior
* without depending on setOffline timing (which varies across CI environments).
*/
const BACKEND_URL = process.env.BACKEND_URL ?? 'http://localhost:4747';
const FRONTEND_URL = process.env.FRONTEND_URL ?? 'http://localhost:5173';
test.beforeAll(async () => {
if (process.env.E2E) return;
try {
const [backendRes, frontendRes] = await Promise.allSettled([
fetch(`${BACKEND_URL}/api/repos`),
fetch(FRONTEND_URL),
]);
if (
backendRes.status === 'rejected' ||
(backendRes.status === 'fulfilled' && !backendRes.value.ok)
) {
test.skip(true, 'gitnexus serve not available');
return;
}
if (
frontendRes.status === 'rejected' ||
(frontendRes.status === 'fulfilled' && !frontendRes.value.ok)
) {
test.skip(true, 'Vite dev server not available');
return;
}
if (backendRes.status === 'fulfilled') {
const repos = await backendRes.value.json();
if (!repos.length) {
test.skip(true, 'No indexed repos');
return;
}
}
} catch {
test.skip(true, 'servers not available');
}
});
test.describe('Heartbeat Reconnect', () => {
test('shows reconnecting banner instead of onboarding reset when heartbeat is unavailable', async ({
page,
}) => {
// Block the heartbeat BEFORE navigating — the EventSource will fail
// immediately on every connection attempt, triggering onReconnecting.
await page.route('**/api/heartbeat', (route) => route.abort('connectionrefused'));
// Load the app and connect to a repo (all other endpoints work normally)
await page.goto('/');
const landingCard = page.locator('[data-testid="landing-repo-card"]').first();
try {
await landingCard.waitFor({ state: 'visible', timeout: 15_000 });
await landingCard.click();
} catch {
// auto-connect may skip the landing screen
}
// Wait for graph to load (heartbeat is blocked, but graph loads fine)
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// The reconnecting banner should appear (heartbeat is failing)
const banner = page.getByText('Server connection lost');
await expect(banner).toBeVisible({ timeout: 15_000 });
// The graph canvas should STILL be visible — NOT reset to onboarding
await expect(page.locator('canvas').first()).toBeVisible();
});
test('banner clears when heartbeat becomes available', async ({ page }) => {
// Start with heartbeat blocked
await page.route('**/api/heartbeat', (route) => route.abort('connectionrefused'));
await page.goto('/');
const landingCard = page.locator('[data-testid="landing-repo-card"]').first();
try {
await landingCard.waitFor({ state: 'visible', timeout: 15_000 });
await landingCard.click();
} catch {
// auto-connect may skip the landing screen
}
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// Verify banner appears
const banner = page.getByText('Server connection lost');
await expect(banner).toBeVisible({ timeout: 15_000 });
// Unblock heartbeat — the real server is running, so reconnect will succeed
await page.unroute('**/api/heartbeat');
// Banner should disappear as heartbeat reconnects
await expect(banner).not.toBeVisible({ timeout: 30_000 });
// Graph should still be there
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible();
});
});
+112
View File
@@ -0,0 +1,112 @@
import { test, expect } from '@playwright/test';
/**
* E2E tests for multi-repo scoping and URL persistence.
*
* Verifies that:
* - Connecting via ?server= loads data and sets ?project= in the URL
* - The repo name appears in the UI after connecting
* - F5 with ?server=&project= reconnects to the correct repo
*
* Runs against the single indexed repo in CI — validates the plumbing
* works end-to-end even with one repo.
*/
const BACKEND_URL = process.env.BACKEND_URL ?? 'http://localhost:4747';
const FRONTEND_URL = process.env.FRONTEND_URL ?? 'http://localhost:5173';
let firstRepoName: string;
test.beforeAll(async () => {
if (process.env.E2E) {
// Still need to fetch the repo name for assertions
try {
const res = await fetch(`${BACKEND_URL}/api/repos`);
const repos = await res.json();
firstRepoName = repos[0]?.name ?? '';
} catch {
firstRepoName = '';
}
return;
}
try {
const [backendRes, frontendRes] = await Promise.allSettled([
fetch(`${BACKEND_URL}/api/repos`),
fetch(FRONTEND_URL),
]);
if (
backendRes.status === 'rejected' ||
(backendRes.status === 'fulfilled' && !backendRes.value.ok)
) {
test.skip(true, 'gitnexus serve not available');
return;
}
if (
frontendRes.status === 'rejected' ||
(frontendRes.status === 'fulfilled' && !frontendRes.value.ok)
) {
test.skip(true, 'Vite dev server not available');
return;
}
if (backendRes.status === 'fulfilled') {
const repos = await backendRes.value.json();
if (!repos.length) {
test.skip(true, 'No indexed repos');
return;
}
firstRepoName = repos[0].name;
}
} catch {
test.skip(true, 'servers not available');
}
});
test.describe('Multi-Repo Scoping', () => {
test('auto-connect via ?server= sets ?project= in URL', async ({ page }) => {
// Navigate with ?server= param (the bookmarkable shortcut)
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
// Wait for graph to load
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// URL should now contain ?project= with the repo name
const url = new URL(page.url());
const project = url.searchParams.get('project');
expect(project).toBeTruthy();
expect(project).toBe(firstRepoName);
});
test('?server= is preserved in URL for F5 recovery', async ({ page }) => {
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// URL should still have ?server=
const url = new URL(page.url());
expect(url.searchParams.get('server')).toBeTruthy();
// F5 should reconnect (not show onboarding)
await page.reload();
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
});
test('node count in status bar matches backend data', async ({ page }) => {
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// Fetch expected node count from backend
const res = await fetch(`${BACKEND_URL}/api/repo?repo=${encodeURIComponent(firstRepoName)}`);
const repoInfo = await res.json();
const expectedNodes = repoInfo.stats?.nodes;
if (expectedNodes) {
// Status bar shows node count — use the status-ready area to avoid
// matching multiple elements (file tree, header may also show counts)
const statusBar = page.locator('footer');
const nodeText = statusBar.getByText(/\d+ nodes/).first();
await expect(nodeText).toBeVisible({ timeout: 10_000 });
const text = await nodeText.textContent();
const displayedNodes = parseInt(text?.match(/(\d+)\s*nodes/)?.[1] ?? '0', 10);
expect(displayedNodes).toBeGreaterThan(0);
}
});
});
+24 -26
View File
@@ -1,4 +1,4 @@
import { test, expect, type TestInfo } from '@playwright/test';
import { test, expect } from '@playwright/test';
/**
* E2E tests for the GitNexus web UI — exploring view features.
@@ -58,36 +58,41 @@ test.beforeAll(async () => {
* For these tests we require at least one indexed repo, so pick the first
* landing card when present and then wait for the exploring view.
*/
async function waitForGraphLoaded(page: import('@playwright/test').Page, testInfo: TestInfo) {
async function waitForGraphLoaded(page: import('@playwright/test').Page) {
await page.goto('/');
const landingCard = page.locator('[data-testid="landing-repo-card"]').first();
const landingCards = page.locator('[data-testid="landing-repo-card"]');
const preferredLandingCard = landingCards
.filter({ hasText: /GitNexus|local-integration/ })
.first();
try {
await landingCard.waitFor({ state: 'visible', timeout: 15_000 });
await landingCards.first().waitFor({ state: 'visible', timeout: 15_000 });
const landingCard =
(await preferredLandingCard.count()) > 0 ? preferredLandingCard : landingCards.first();
await landingCard.click();
} catch {
// Landing screen may not appear (e.g. ?server auto-connect)
}
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
await expect(page.getByText(/\d+ nodes/).first()).toBeVisible();
await page.screenshot({ path: testInfo.outputPath('graph-loaded.png') });
const statusBar = page.getByRole('contentinfo');
await expect(statusBar.getByText('Ready', { exact: true })).toBeVisible({ timeout: 45_000 });
await expect(statusBar).toContainText(/nodes/, {
timeout: 20_000,
});
}
test.describe('Server Connection & Graph Loading', () => {
test('selects a repo from landing and loads graph', async ({ page }, testInfo) => {
await waitForGraphLoaded(page, testInfo);
await page.screenshot({ path: testInfo.outputPath('graph-loaded-full.png'), fullPage: true });
test('selects a repo from landing and loads graph', async ({ page }) => {
await waitForGraphLoaded(page);
});
});
test.describe('Nexus AI', () => {
test('panel opens and agent initializes without error', async ({ page }, testInfo) => {
await waitForGraphLoaded(page, testInfo);
test('panel opens and agent initializes without error', async ({ page }) => {
await waitForGraphLoaded(page);
await page.getByRole('button', { name: 'Nexus AI' }).click();
await expect(page.getByText('Ask me anything')).toBeVisible({ timeout: 15_000 });
await page.screenshot({ path: testInfo.outputPath('nexus-ai-panel.png'), fullPage: true });
const errorBanner = page.getByText('Database not ready');
expect(await errorBanner.isVisible().catch(() => false)).toBe(false);
@@ -95,8 +100,8 @@ test.describe('Nexus AI', () => {
});
test.describe('Processes Panel', () => {
test('shows process list and View button works', async ({ page }, testInfo) => {
await waitForGraphLoaded(page, testInfo);
test('shows process list and View button works', async ({ page }) => {
await waitForGraphLoaded(page);
await page.getByRole('button', { name: 'Nexus AI' }).click();
await page.getByText('Processes').click();
@@ -104,7 +109,6 @@ test.describe('Processes Panel', () => {
await expect(page.locator('[data-testid="process-list-loaded"]')).toBeVisible({
timeout: 15_000,
});
await page.screenshot({ path: testInfo.outputPath('processes-panel.png'), fullPage: true });
const processRow = page.locator('[data-testid="process-row"]').first();
await expect(processRow).toBeVisible({ timeout: 10_000 });
@@ -114,14 +118,10 @@ test.describe('Processes Panel', () => {
await viewBtn.waitFor({ state: 'visible', timeout: 5_000 });
await viewBtn.click();
await expect(page.locator('[data-testid="process-modal"]')).toBeVisible({ timeout: 5_000 });
await page.screenshot({
path: testInfo.outputPath('process-view-clicked.png'),
fullPage: true,
});
});
test('lightbulb highlights nodes in graph', async ({ page }, testInfo) => {
await waitForGraphLoaded(page, testInfo);
test('lightbulb highlights nodes in graph', async ({ page }) => {
await waitForGraphLoaded(page);
await page.getByRole('button', { name: 'Nexus AI' }).click();
await page.getByText('Processes').click();
@@ -137,13 +137,12 @@ test.describe('Processes Panel', () => {
await lightbulb.waitFor({ state: 'visible', timeout: 5_000 });
await lightbulb.click();
await expect(processRow).toHaveClass(/bg-amber-950/, { timeout: 5_000 });
await page.screenshot({ path: testInfo.outputPath('after-highlight.png'), fullPage: true });
});
});
test.describe('Turn Off All Highlights', () => {
test('selecting a node dims others, button clears it', async ({ page }, testInfo) => {
await waitForGraphLoaded(page, testInfo);
test('selecting a node dims others, button clears it', async ({ page }) => {
await waitForGraphLoaded(page);
await expect(page.locator('canvas').first()).toBeVisible({ timeout: 10_000 });
@@ -160,6 +159,5 @@ test.describe('Turn Off All Highlights', () => {
await expect(highlightToggle).toHaveAttribute('title', 'Turn on AI highlights', {
timeout: 5_000,
});
await page.screenshot({ path: testInfo.outputPath('highlights-cleared.png'), fullPage: true });
});
});
+78 -47
View File
@@ -1,4 +1,4 @@
import { useCallback, useEffect, useRef } from 'react';
import { useCallback, useEffect, useRef, useState } from 'react';
import { AppStateProvider, useAppState } from './hooks/useAppState';
import { DropZone } from './components/DropZone';
import { LoadingOverlay } from './components/LoadingOverlay';
@@ -36,7 +36,6 @@ const AppContent = () => {
refreshLLMSettings,
initializeAgent,
startEmbeddingsWithFallback,
embeddingStatus,
codeReferences,
selectedNode,
isCodePanelOpen,
@@ -45,17 +44,27 @@ const AppContent = () => {
availableRepos,
setAvailableRepos,
switchRepo,
setCurrentRepo,
} = useAppState();
const graphCanvasRef = useRef<GraphCanvasHandle>(null);
const [serverDisconnected, setServerDisconnected] = useState(false);
const handleServerConnect = useCallback(
async (result: ConnectResult): Promise<void> => {
// Extract project name from repoPath
// Use the canonical repo name from the server response so all subsequent
// backend calls (queries, search, grep, readFile) scope to this repo.
const repoName = result.repoInfo.name;
const repoPath = result.repoInfo.repoPath ?? result.repoInfo.path;
const parts = (repoPath || '').split('/').filter((p) => p && !p.startsWith('.'));
const projectName = parts[parts.length - 1] || parts[0] || 'server-project';
const projectName =
repoName || repoPath?.split('/').filter(Boolean).pop() || 'server-project';
setProjectName(projectName);
setCurrentRepo(projectName);
// Update URL so F5 / bookmarks preserve which repo is open
const url = new URL(window.location.href);
url.searchParams.set('project', projectName);
window.history.replaceState(null, '', url.toString());
// Build KnowledgeGraph from server data for visualization
const graph = createKnowledgeGraph();
@@ -80,10 +89,18 @@ const AppContent = () => {
console.warn('Failed to initialize agent:', err);
}
},
[setViewMode, setGraph, setProjectName, initializeAgent, startEmbeddingsWithFallback],
[
setViewMode,
setGraph,
setProjectName,
setCurrentRepo,
initializeAgent,
startEmbeddingsWithFallback,
],
);
// Auto-connect when ?server query param is present (bookmarkable shortcut)
// Auto-connect when ?server query param is present (bookmarkable shortcut).
// Also reads ?project= to connect to a specific repo.
const autoConnectRan = useRef(false);
useEffect(() => {
if (autoConnectRan.current) return;
@@ -91,9 +108,12 @@ const AppContent = () => {
if (!params.has('server')) return;
autoConnectRan.current = true;
// Clean the URL so a refresh won't re-trigger
const cleanUrl = window.location.pathname + window.location.hash;
window.history.replaceState(null, '', cleanUrl);
const serverUrl = params.get('server') || window.location.origin;
const projectParam = params.get('project') || undefined;
// Keep ?server= in the URL so F5 reconnects to the same server.
// autoConnectRan.current prevents re-trigger within the same session.
// handleServerConnect() will add/update ?project= after connecting.
setProgress({
phase: 'extracting',
@@ -103,36 +123,39 @@ const AppContent = () => {
});
setViewMode('loading');
const serverUrl = params.get('server') || window.location.origin;
const baseUrl = normalizeServerUrl(serverUrl);
connectToServer(serverUrl, (phase, downloaded, total) => {
if (phase === 'validating') {
setProgress({
phase: 'extracting',
percent: 5,
message: 'Connecting to server...',
detail: 'Validating server',
});
} else if (phase === 'downloading') {
const pct = total ? Math.round((downloaded / total) * 90) + 5 : 50;
const mb = (downloaded / (1024 * 1024)).toFixed(1);
setProgress({
phase: 'extracting',
percent: pct,
message: 'Downloading graph...',
detail: `${mb} MB downloaded`,
});
} else if (phase === 'extracting') {
setProgress({
phase: 'extracting',
percent: 97,
message: 'Processing...',
detail: 'Extracting file contents',
});
}
})
connectToServer(
serverUrl,
(phase, downloaded, total) => {
if (phase === 'validating') {
setProgress({
phase: 'extracting',
percent: 5,
message: 'Connecting to server...',
detail: 'Validating server',
});
} else if (phase === 'downloading') {
const pct = total ? Math.round((downloaded / total) * 90) + 5 : 50;
const mb = (downloaded / (1024 * 1024)).toFixed(1);
setProgress({
phase: 'extracting',
percent: pct,
message: 'Downloading graph...',
detail: `${mb} MB downloaded`,
});
} else if (phase === 'extracting') {
setProgress({
phase: 'extracting',
percent: 97,
message: 'Processing...',
detail: 'Extracting file contents',
});
}
},
undefined,
projectParam,
)
.then(async (result) => {
await handleServerConnect(result);
setProgress(null);
@@ -169,21 +192,18 @@ const AppContent = () => {
// ── Server heartbeat: detect when server goes down while exploring ────────
// Uses SSE (EventSource) for instant detection — no polling delay.
// On disconnect: show a reconnecting banner instead of resetting to onboarding.
// The heartbeat retries indefinitely with capped backoff and recovers automatically.
useEffect(() => {
if (viewMode !== 'exploring') return;
const cleanup = connectHeartbeat(
() => {}, // onConnect — already connected, no action needed
() => {
// Server went down — return to onboarding
setViewMode('onboarding');
setGraph(null);
setProgress(null);
},
() => setServerDisconnected(false),
() => setServerDisconnected(true),
);
return cleanup;
}, [viewMode, setViewMode, setGraph, setProgress]);
}, [viewMode]);
// Render based on view mode
if (viewMode === 'onboarding') {
@@ -196,7 +216,12 @@ const AppContent = () => {
await handleServerConnect(result);
setProgress(null);
if (serverUrl) {
setServerBaseUrl(normalizeServerUrl(serverUrl));
const base = normalizeServerUrl(serverUrl);
setServerBaseUrl(base);
// Add ?server= so F5 reconnects to this server
const url = new URL(window.location.href);
url.searchParams.set('server', base);
window.history.replaceState(null, '', url.toString());
}
}}
/>
@@ -268,6 +293,12 @@ const AppContent = () => {
<StatusBar />
{serverDisconnected && (
<div className="fixed bottom-12 left-1/2 z-50 -translate-x-1/2 rounded-lg border border-yellow-500/30 bg-yellow-900/80 px-4 py-2 text-sm text-yellow-200 shadow-lg backdrop-blur">
Server connection lost — reconnecting&hellip;
</div>
)}
{/* Settings Panel (modal) */}
<SettingsPanel
isOpen={isSettingsPanelOpen}
@@ -54,6 +54,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
clearCodeReferences,
setSelectedNode,
codeReferenceFocus,
projectName,
} = useAppState();
const nodeById = useMemo(() => {
@@ -174,7 +175,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
return () => {
rafIds.forEach((id) => cancelAnimationFrame(id));
};
}, [codeReferenceFocus?.ts, aiReferences]);
}, [codeReferenceFocus, aiReferences]);
const refsWithSnippets = useMemo(() => {
return aiReferences.map((ref) => {
@@ -223,10 +224,11 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
const isWholeFile = selectedIsFile || startLine === undefined;
const options = isWholeFile
? undefined
? { repo: projectName }
: {
startLine: Math.max(0, startLine - CONTEXT_LINES),
endLine: (endLine ?? startLine) + CONTEXT_LINES,
repo: projectName,
};
readFile(selectedFilePath, options)
@@ -251,6 +253,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
selectedNode?.properties?.startLine,
selectedNode?.properties?.endLine,
selectedIsFile,
projectName,
]);
// Scroll to the selected node's startLine after content loads
@@ -201,7 +201,7 @@ interface OnboardingGuideProps {
}
export const OnboardingGuide = ({ isPolling }: OnboardingGuideProps) => {
const primary = isDev ? 'cd gitnexus && npm run serve' : 'npx gitnexus@latest serve';
const primary = isDev ? 'npm run --prefix gitnexus serve' : 'npx gitnexus@latest serve';
const termLabel = isDev ? 'Start backend' : 'Terminal';
// Step states: step 1 = copy command, step 2 = run/wait, step 3 = auto-connect
@@ -277,7 +277,9 @@ export const OnboardingGuide = ({ isPolling }: OnboardingGuideProps) => {
state={step2State}
number={2}
title={isPolling ? 'Waiting for server to start' : 'Paste and run in your terminal'}
description={isPolling ? undefined : 'Open a new terminal window, paste, and hit Enter.'}
description={
isPolling ? undefined : 'Open a terminal at the project root, paste, and hit Enter.'
}
>
{isPolling && <PollingBar />}
</StepRow>
+1 -1
View File
@@ -64,7 +64,7 @@ export const StatusBar = () => {
</a>
{/* Right - Stats */}
<div className="flex items-center gap-3">
<div className="flex items-center gap-3" data-testid="graph-stats">
{graph && (
<>
<span>{nodeCount} nodes</span>
+11
View File
@@ -145,6 +145,7 @@ interface AppState {
availableRepos: BackendRepo[];
setAvailableRepos: (repos: BackendRepo[]) => void;
switchRepo: (repoName: string) => Promise<void>;
setCurrentRepo: (repoName: string) => void;
// Worker API (shared across app)
runQuery: (cypher: string) => Promise<any[]>;
@@ -456,6 +457,10 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
// Backend client — direct HTTP calls (no Worker/Comlink)
const repoRef = useRef<string | undefined>(undefined);
const setCurrentRepo = useCallback((repoName: string) => {
repoRef.current = repoName;
}, []);
const runQuery = useCallback(async (cypher: string): Promise<any[]> => {
return backendRunQuery(cypher, repoRef.current);
}, []);
@@ -1077,6 +1082,11 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
setProjectName(pName);
repoRef.current = pName;
// Update URL so F5 / bookmarks open the correct repo
const url = new URL(window.location.href);
url.searchParams.set('project', pName);
window.history.replaceState(null, '', url.toString());
const newGraph = createKnowledgeGraph();
for (const node of result.nodes) newGraph.addNode(node);
for (const rel of result.relationships) newGraph.addRelationship(rel);
@@ -1219,6 +1229,7 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
availableRepos,
setAvailableRepos,
switchRepo,
setCurrentRepo,
runQuery,
isDatabaseReady,
// Embedding state and methods
+90 -12
View File
@@ -302,16 +302,26 @@ export const fetchServerInfo = async (): Promise<ServerInfo> => {
};
/**
* Connect an SSE heartbeat to the backend. Fires `onDisconnect` when the
* server goes down (after one retry to avoid false positives from transient
* network hiccups). Returns a cleanup function.
* Connect an SSE heartbeat to the backend. Retries indefinitely with capped
* exponential backoff so transient hiccups don't reset the UI.
*
* - `onConnect` fires on every successful (re)connection.
* - `onReconnecting` fires on the first retry after a drop — use it to show
* a "reconnecting" banner while keeping the current view intact.
*
* Returns a cleanup function that tears down the EventSource and timers.
*/
export const connectHeartbeat = (onConnect: () => void, onDisconnect: () => void): (() => void) => {
export const connectHeartbeat = (
onConnect: () => void,
onReconnecting: () => void,
): (() => void) => {
let closed = false;
let retryTimer: ReturnType<typeof setTimeout> | null = null;
let es: EventSource | null = null;
let attempt = 0;
const MAX_RETRIES = 3;
/** Whether we've already fired onReconnecting for the current drop. */
let notifiedReconnecting = false;
const MAX_BACKOFF_MS = 15_000;
const connect = () => {
if (closed) return;
@@ -319,6 +329,7 @@ export const connectHeartbeat = (onConnect: () => void, onDisconnect: () => void
es.onopen = () => {
if (!closed) {
attempt = 0;
notifiedReconnecting = false;
onConnect();
}
};
@@ -326,13 +337,15 @@ export const connectHeartbeat = (onConnect: () => void, onDisconnect: () => void
es?.close();
es = null;
if (closed) return;
if (attempt < MAX_RETRIES) {
const delay = 1_000 * Math.pow(2, attempt);
attempt++;
retryTimer = setTimeout(connect, delay);
} else {
onDisconnect();
if (!notifiedReconnecting) {
notifiedReconnecting = true;
onReconnecting();
}
const delay = Math.min(1_000 * Math.pow(2, attempt), MAX_BACKOFF_MS);
attempt++;
retryTimer = setTimeout(connect, delay);
};
};
@@ -391,13 +404,18 @@ export const fetchGraph = async (
onProgress?: (downloaded: number, total: number | null) => void;
},
): Promise<{ nodes: GraphNode[]; relationships: GraphRelationship[] }> => {
const params = [repoParam(repo), opts?.includeContent ? 'includeContent=true' : '']
const params = [repoParam(repo), opts?.includeContent ? 'includeContent=true' : '', 'stream=true']
.filter(Boolean)
.join('&');
const url = `${_backendUrl}/api/graph${params ? `?${params}` : ''}`;
const response = await fetchWithTimeout(url, { signal: opts?.signal }, 60_000);
await assertOk(response);
const contentType = response.headers.get('Content-Type') || '';
if (contentType.includes('application/x-ndjson')) {
return parseNdjsonGraphResponse(response, opts?.onProgress);
}
if (!opts?.onProgress || !response.body) {
return response.json() as Promise<{ nodes: GraphNode[]; relationships: GraphRelationship[] }>;
}
@@ -426,6 +444,66 @@ export const fetchGraph = async (
return JSON.parse(new TextDecoder().decode(combined));
};
const parseNdjsonGraphResponse = async (
response: Response,
onProgress?: (downloaded: number, total: number | null) => void,
): Promise<{ nodes: GraphNode[]; relationships: GraphRelationship[] }> => {
if (!response.body) {
throw new BackendError('No response body', response.status, 'server');
}
const contentLength = response.headers.get('Content-Length');
const total = contentLength ? parseInt(contentLength, 10) : null;
const reader = response.body.getReader();
const decoder = new TextDecoder();
const nodes: GraphNode[] = [];
const relationships: GraphRelationship[] = [];
let buffer = '';
let downloaded = 0;
const parseLine = (line: string) => {
const trimmed = line.trim();
if (!trimmed) return;
const record = JSON.parse(trimmed) as
| { type: 'node'; data: GraphNode }
| { type: 'relationship'; data: GraphRelationship }
| { type: 'error'; error: string };
if (record.type === 'node') {
nodes.push(record.data);
return;
}
if (record.type === 'relationship') {
relationships.push(record.data);
return;
}
if (record.type === 'error') {
throw new BackendError(record.error, response.status || 500, 'server');
}
};
while (true) {
const { done, value } = await reader.read();
if (done) break;
downloaded += value.length;
onProgress?.(downloaded, total);
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split('\n');
buffer = lines.pop() || '';
for (const line of lines) {
parseLine(line);
}
}
buffer += decoder.decode();
parseLine(buffer);
return { nodes, relationships };
};
/** Execute a Cypher query. Returns rows. */
export const runQuery = async (
cypher: string,
+147
View File
@@ -0,0 +1,147 @@
import { describe, it, expect, vi, beforeEach, afterEach } from 'vitest';
import { connectHeartbeat } from '../../src/services/backend-client';
// Mock EventSource to simulate SSE behavior
class MockEventSource {
onopen: (() => void) | null = null;
onerror: (() => void) | null = null;
closed = false;
close() {
this.closed = true;
}
}
let lastEventSource: MockEventSource | null = null;
beforeEach(() => {
lastEventSource = null;
vi.stubGlobal(
'EventSource',
vi.fn().mockImplementation(() => {
lastEventSource = new MockEventSource();
return lastEventSource;
}),
);
vi.useFakeTimers();
});
afterEach(() => {
vi.useRealTimers();
vi.unstubAllGlobals();
});
describe('connectHeartbeat', () => {
it('calls onConnect when EventSource opens', () => {
const onConnect = vi.fn();
const onReconnecting = vi.fn();
connectHeartbeat(onConnect, onReconnecting);
lastEventSource!.onopen!();
expect(onConnect).toHaveBeenCalledOnce();
expect(onReconnecting).not.toHaveBeenCalled();
});
it('calls onReconnecting on first error, then retries', () => {
const onConnect = vi.fn();
const onReconnecting = vi.fn();
connectHeartbeat(onConnect, onReconnecting);
// Simulate connection drop
lastEventSource!.onerror!();
expect(onReconnecting).toHaveBeenCalledOnce();
expect(lastEventSource!.closed).toBe(true);
// Advance past first retry delay (1s)
vi.advanceTimersByTime(1_000);
// A new EventSource should have been created
expect(EventSource).toHaveBeenCalledTimes(2);
});
it('fires onReconnecting only once per disconnect', () => {
const onConnect = vi.fn();
const onReconnecting = vi.fn();
connectHeartbeat(onConnect, onReconnecting);
// First error
lastEventSource!.onerror!();
expect(onReconnecting).toHaveBeenCalledOnce();
// Second retry fires error again
vi.advanceTimersByTime(1_000);
lastEventSource!.onerror!();
expect(onReconnecting).toHaveBeenCalledOnce(); // still 1
// Third retry fires error
vi.advanceTimersByTime(2_000);
lastEventSource!.onerror!();
expect(onReconnecting).toHaveBeenCalledOnce(); // still 1
});
it('retries indefinitely instead of giving up after 3 attempts', () => {
const onConnect = vi.fn();
const onReconnecting = vi.fn();
connectHeartbeat(onConnect, onReconnecting);
// Simulate 10 consecutive failures — should never stop retrying
for (let i = 0; i < 10; i++) {
lastEventSource!.onerror!();
// Advance past the max backoff (15s) to ensure the next retry fires
vi.advanceTimersByTime(16_000);
}
// Should have created 11 EventSources (1 initial + 10 retries)
expect(EventSource).toHaveBeenCalledTimes(11);
});
it('resets reconnecting state when connection recovers', () => {
const onConnect = vi.fn();
const onReconnecting = vi.fn();
connectHeartbeat(onConnect, onReconnecting);
// Drop
lastEventSource!.onerror!();
expect(onReconnecting).toHaveBeenCalledOnce();
// Retry succeeds
vi.advanceTimersByTime(1_000);
lastEventSource!.onopen!();
expect(onConnect).toHaveBeenCalledOnce();
// Drop again — should fire onReconnecting again (reset after recovery)
lastEventSource!.onerror!();
expect(onReconnecting).toHaveBeenCalledTimes(2);
});
it('caps backoff at 15 seconds', () => {
const onConnect = vi.fn();
const onReconnecting = vi.fn();
connectHeartbeat(onConnect, onReconnecting);
// Fail many times to push backoff past the cap
for (let i = 0; i < 6; i++) {
lastEventSource!.onerror!();
// The delay for attempt i is min(1000 * 2^i, 15000)
// i=0: 1s, i=1: 2s, i=2: 4s, i=3: 8s, i=4: 15s (capped), i=5: 15s (capped)
vi.advanceTimersByTime(16_000);
}
// All retries should have fired — 7 EventSources total
expect(EventSource).toHaveBeenCalledTimes(7);
});
it('stops retrying when cleanup is called', () => {
const onConnect = vi.fn();
const onReconnecting = vi.fn();
const cleanup = connectHeartbeat(onConnect, onReconnecting);
lastEventSource!.onerror!();
cleanup();
// Advance time — no new EventSource should be created
vi.advanceTimersByTime(30_000);
expect(EventSource).toHaveBeenCalledTimes(1);
});
});
@@ -1,5 +1,5 @@
import { describe, expect, it } from 'vitest';
import { normalizeServerUrl } from '../../src/services/backend-client';
import { afterEach, describe, expect, it, vi } from 'vitest';
import { fetchGraph, normalizeServerUrl, setBackendUrl } from '../../src/services/backend-client';
describe('normalizeServerUrl', () => {
it('adds http:// to localhost', () => {
@@ -31,3 +31,137 @@ describe('normalizeServerUrl', () => {
expect(normalizeServerUrl('https://gitnexus.example.com')).toBe('https://gitnexus.example.com');
});
});
afterEach(() => {
vi.restoreAllMocks();
});
describe('fetchGraph', () => {
it('requests streamed graph responses from the backend', async () => {
setBackendUrl('http://localhost:4747');
const fetchMock = vi.fn().mockResolvedValue(
new Response('{"nodes":[],"relationships":[]}', {
status: 200,
headers: {
'Content-Type': 'application/json',
},
}),
);
vi.stubGlobal('fetch', fetchMock);
await fetchGraph('big-repo');
expect(fetchMock).toHaveBeenCalledWith(
expect.stringContaining('/api/graph?repo=big-repo&stream=true'),
expect.any(Object),
);
});
it('parses NDJSON graph streams incrementally', async () => {
setBackendUrl('http://localhost:4747');
const encoder = new TextEncoder();
const stream = new ReadableStream<Uint8Array>({
start(controller) {
controller.enqueue(
encoder.encode(
[
'{"type":"node","data":{"id":"File:src/app.ts","label":"File","properties":{"name":"app.ts","filePath":"src/app.ts"}}}\n',
'{"type":"relationship","data":{"id":"File:src/app.ts_CONTAINS_Function:src/app.ts:main","type":"CONTAINS","sourceId":"File:src/app.ts","targetId":"Function:src/app.ts:main"}}\n',
].join(''),
),
);
controller.close();
},
});
vi.stubGlobal(
'fetch',
vi.fn().mockResolvedValue(
new Response(stream, {
status: 200,
headers: {
'Content-Type': 'application/x-ndjson',
},
}),
),
);
const progress = vi.fn();
const result = await fetchGraph('big-repo', { onProgress: progress });
expect(result.nodes).toHaveLength(1);
expect(result.relationships).toHaveLength(1);
expect(result.nodes[0].id).toBe('File:src/app.ts');
expect(result.relationships[0].type).toBe('CONTAINS');
expect(progress).toHaveBeenCalled();
});
it('parses NDJSON graph lines split across chunks', async () => {
setBackendUrl('http://localhost:4747');
const encoder = new TextEncoder();
const stream = new ReadableStream<Uint8Array>({
start(controller) {
controller.enqueue(
encoder.encode(
'{"type":"node","data":{"id":"File:src/app.ts","label":"File","properties":{"name":"app.ts"',
),
);
controller.enqueue(
encoder.encode(
',"filePath":"src/app.ts"}}}\n{"type":"relationship","data":{"id":"File:src/app.ts_CONTAINS_Function:src/app.ts:main","type":"CONTAINS","sourceId":"File:src/app.ts","targetId":"Function:src/app.ts:main"}}\n',
),
);
controller.close();
},
});
vi.stubGlobal(
'fetch',
vi.fn().mockResolvedValue(
new Response(stream, {
status: 200,
headers: {
'Content-Type': 'application/x-ndjson',
},
}),
),
);
const result = await fetchGraph('big-repo');
expect(result.nodes).toHaveLength(1);
expect(result.relationships).toHaveLength(1);
expect(result.nodes[0].properties.filePath).toBe('src/app.ts');
});
it('throws backend errors emitted in the NDJSON stream', async () => {
setBackendUrl('http://localhost:4747');
const encoder = new TextEncoder();
const stream = new ReadableStream<Uint8Array>({
start(controller) {
controller.enqueue(encoder.encode('{"type":"error","error":"stream failed"}\n'));
controller.close();
},
});
vi.stubGlobal(
'fetch',
vi.fn().mockResolvedValue(
new Response(stream, {
status: 200,
headers: {
'Content-Type': 'application/x-ndjson',
},
}),
),
);
await expect(fetchGraph('big-repo')).rejects.toMatchObject({
message: 'stream failed',
});
});
});
+10
View File
@@ -164,6 +164,16 @@ gitnexus clean # Delete index for current repo
gitnexus clean --all --force # Delete all indexes
gitnexus wiki [path] # Generate LLM-powered docs from knowledge graph
gitnexus wiki --model <model> # Wiki with custom LLM model (default: gpt-4o-mini)
# Repository groups (multi-repo / monorepo service tracking)
gitnexus group create <name> # Create a repository group
gitnexus group add <name> <repo> # Add a repo to a group
gitnexus group remove <name> <repo> # Remove a repo from a group
gitnexus group list [name] # List groups, or show one group's config
gitnexus group sync <name> # Extract contracts and match across repos/services
gitnexus group contracts <name> # Inspect extracted contracts and cross-links
gitnexus group query <name> <q> # Search execution flows across all repos in a group
gitnexus group status <name> # Check staleness of repos in a group
```
## Remote Embeddings
+37 -2
View File
@@ -1,17 +1,18 @@
{
"name": "gitnexus",
"version": "1.5.2",
"version": "1.5.3",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "gitnexus",
"version": "1.5.2",
"version": "1.5.3",
"license": "PolyForm-Noncommercial-1.0.0",
"dependencies": {
"@huggingface/transformers": "^3.0.0",
"@ladybugdb/core": "^0.15.2",
"@modelcontextprotocol/sdk": "^1.0.0",
"@scarf/scarf": "^1.4.0",
"cli-progress": "^3.12.0",
"commander": "^12.0.0",
"cors": "^2.8.5",
@@ -21,6 +22,7 @@
"graphology-indices": "^0.17.0",
"graphology-utils": "^2.3.0",
"ignore": "^7.0.5",
"js-yaml": "^4.1.1",
"lru-cache": "^11.0.0",
"mnemonist": "^0.39.0",
"onnxruntime-node": "^1.24.0",
@@ -46,6 +48,7 @@
"@types/cli-progress": "^3.11.6",
"@types/cors": "^2.8.17",
"@types/express": "^4.17.21",
"@types/js-yaml": "^4.0.9",
"@types/node": "^20.0.0",
"@types/uuid": "^10.0.0",
"@vitest/coverage-v8": "^4.0.18",
@@ -1876,6 +1879,13 @@
"dev": true,
"license": "MIT"
},
"node_modules/@scarf/scarf": {
"version": "1.4.0",
"resolved": "https://registry.npmjs.org/@scarf/scarf/-/scarf-1.4.0.tgz",
"integrity": "sha512-xxeapPiUXdZAE3che6f3xogoJPeZgig6omHEy1rIY5WVsB3H2BHNnZH+gHG6x91SCWyQCzWGsuL2Hh3ClO5/qQ==",
"hasInstallScript": true,
"license": "Apache-2.0"
},
"node_modules/@standard-schema/spec": {
"version": "1.1.0",
"resolved": "https://registry.npmjs.org/@standard-schema/spec/-/spec-1.1.0.tgz",
@@ -1993,6 +2003,13 @@
"dev": true,
"license": "MIT"
},
"node_modules/@types/js-yaml": {
"version": "4.0.9",
"resolved": "https://registry.npmjs.org/@types/js-yaml/-/js-yaml-4.0.9.tgz",
"integrity": "sha512-k4MGaQl5TGo/iipqb2UDG2UwjXziSWkh0uysQelTlJpX1qGlpUZYm8PnO4DxG1qBomtJUdYJ6qR6xdIah10JLg==",
"dev": true,
"license": "MIT"
},
"node_modules/@types/mime": {
"version": "1.3.5",
"resolved": "https://registry.npmjs.org/@types/mime/-/mime-1.3.5.tgz",
@@ -2286,6 +2303,12 @@
"url": "https://github.com/chalk/ansi-styles?sponsor=1"
}
},
"node_modules/argparse": {
"version": "2.0.1",
"resolved": "https://registry.npmjs.org/argparse/-/argparse-2.0.1.tgz",
"integrity": "sha512-8+9WqebbFzpX9OR+Wa6O29asIogeRMzcGtAINdpMHHyAg10f05aSFVBbcEqGf/PXw1EjAZ+q2/bEBg3DvurK3Q==",
"license": "Python-2.0"
},
"node_modules/array-flatten": {
"version": "1.1.1",
"resolved": "https://registry.npmjs.org/array-flatten/-/array-flatten-1.1.1.tgz",
@@ -3567,6 +3590,18 @@
"dev": true,
"license": "MIT"
},
"node_modules/js-yaml": {
"version": "4.1.1",
"resolved": "https://registry.npmjs.org/js-yaml/-/js-yaml-4.1.1.tgz",
"integrity": "sha512-qQKT4zQxXl8lLwBtHMWwaTcGfFOZviOJet3Oy/xmGk2gZH677CJM9EvtfdSkgWcATZhj/55JZ0rmy3myCT5lsA==",
"license": "MIT",
"dependencies": {
"argparse": "^2.0.1"
},
"bin": {
"js-yaml": "bin/js-yaml.js"
}
},
"node_modules/json-schema-traverse": {
"version": "1.0.0",
"resolved": "https://registry.npmjs.org/json-schema-traverse/-/json-schema-traverse-1.0.0.tgz",
+7 -2
View File
@@ -1,6 +1,6 @@
{
"name": "gitnexus",
"version": "1.5.2",
"version": "1.5.3",
"description": "Graph-powered code intelligence for AI agents. Index any codebase, query via MCP or CLI.",
"author": "Abhigyan Patwari",
"license": "PolyForm-Noncommercial-1.0.0",
@@ -53,15 +53,18 @@
"@huggingface/transformers": "^3.0.0",
"@ladybugdb/core": "^0.15.2",
"@modelcontextprotocol/sdk": "^1.0.0",
"@scarf/scarf": "^1.4.0",
"cli-progress": "^3.12.0",
"commander": "^12.0.0",
"cors": "^2.8.5",
"express": "^4.19.2",
"gitnexus-shared": "file:../gitnexus-shared",
"glob": "^11.0.0",
"graphology": "^0.25.4",
"graphology-indices": "^0.17.0",
"graphology-utils": "^2.3.0",
"ignore": "^7.0.5",
"js-yaml": "^4.1.1",
"lru-cache": "^11.0.0",
"mnemonist": "^0.39.0",
"onnxruntime-node": "^1.24.0",
@@ -90,6 +93,7 @@
"@types/cli-progress": "^3.11.6",
"@types/cors": "^2.8.17",
"@types/express": "^4.17.21",
"@types/js-yaml": "^4.0.9",
"@types/node": "^20.0.0",
"@types/uuid": "^10.0.0",
"@vitest/coverage-v8": "^4.0.18",
@@ -100,7 +104,8 @@
"overrides": {
"@huggingface/transformers": {
"onnxruntime-node": "$onnxruntime-node"
}
},
"tree-sitter-c": "0.23.2"
},
"engines": {
"node": ">=20.0.0"
+29 -2
View File
@@ -42,10 +42,28 @@ const GITNEXUS_END_MARKER = '<!-- gitnexus:end -->';
* - Exact tool commands with parameters — vague directives get ignored
* - Self-review checklist — forces model to verify its own work
*/
async function findGroupsContainingRegistryName(registryName: string): Promise<string[]> {
const { listGroups, getDefaultGitnexusDir, getGroupDir } =
await import('../core/group/storage.js');
const { loadGroupConfig } = await import('../core/group/config-parser.js');
const names = await listGroups();
const hits: string[] = [];
for (const g of names) {
try {
const config = await loadGroupConfig(getGroupDir(getDefaultGitnexusDir(), g));
if (Object.values(config.repos).some((r) => r === registryName)) hits.push(config.name);
} catch {
// skip invalid or unreadable groups
}
}
return hits;
}
function generateGitNexusContent(
projectName: string,
stats: RepoStats,
generatedSkills?: GeneratedSkillInfo[],
groupNames?: string[],
): string {
const generatedRows =
generatedSkills && generatedSkills.length > 0
@@ -155,7 +173,15 @@ To check whether embeddings exist, inspect \`.gitnexus/meta.json\` — the \`sta
> Claude Code users: A PostToolUse hook handles this automatically after \`git commit\` and \`git merge\`.
## CLI
${
groupNames && groupNames.length > 0
? `## Cross-Repo Groups
This repository is listed under GitNexus **group(s): ${groupNames.join(', ')}** (see \`~/.gitnexus/groups/\`). For blast radius across repository boundaries, use MCP tools \`group_impact\`, \`group_sync\`, \`group_query\`, \`group_contracts\`, \`group_status\`, and \`group_list\`. From the terminal: \`npx gitnexus group list\`, \`npx gitnexus group sync <name>\`, \`npx gitnexus group impact <name> --target <symbol> --repo <group-path>\`.
`
: ''
}## CLI
${skillsTable}
@@ -305,7 +331,8 @@ export async function generateAIContextFiles(
generatedSkills?: GeneratedSkillInfo[],
options?: AIContextOptions,
): Promise<{ files: string[] }> {
const content = generateGitNexusContent(projectName, stats, generatedSkills);
const groupNames = await findGroupsContainingRegistryName(projectName);
const content = generateGitNexusContent(projectName, stats, generatedSkills, groupNames);
const createdFiles: string[] = [];
if (!options?.skipAgentsMd) {
+298
View File
@@ -0,0 +1,298 @@
// gitnexus/src/cli/group.ts
import { createRequire } from 'node:module';
import type { Command } from 'commander';
const _require = createRequire(import.meta.url);
const yaml = _require('js-yaml') as typeof import('js-yaml');
export function registerGroupCommands(program: Command): void {
const group = program
.command('group')
.description('Manage repository groups for cross-index impact analysis');
group
.command('create <name>')
.description('Create a new group with template group.yaml')
.option('--force', 'Overwrite existing group')
.action(async (name: string, opts: { force?: boolean }) => {
const { createGroupDir, getDefaultGitnexusDir } = await import('../core/group/storage.js');
const dir = await createGroupDir(getDefaultGitnexusDir(), name, opts.force);
console.log(`Created group "${name}" at ${dir}`);
console.log('Edit group.yaml to add repos, then run: gitnexus group sync ' + name);
});
group
.command('add <group> <groupPath> <registryName>')
.description(
'Add a repo to a group. <groupPath> = hierarchy path (e.g. hr/hiring/backend), <registryName> = name from registry',
)
.action(async (groupName: string, groupPath: string, registryName: string) => {
const { getGroupDir, getDefaultGitnexusDir } = await import('../core/group/storage.js');
const { loadGroupConfig } = await import('../core/group/config-parser.js');
const path = await import('node:path');
const fs = await import('node:fs/promises');
const groupDir = getGroupDir(getDefaultGitnexusDir(), groupName);
const config = await loadGroupConfig(groupDir);
config.repos[groupPath] = registryName;
await fs.writeFile(path.join(groupDir, 'group.yaml'), yaml.dump(config), 'utf-8');
console.log(`Added ${registryName} as "${groupPath}" to group "${groupName}"`);
console.log(`Run: gitnexus group sync ${groupName}`);
});
group
.command('remove <group> <path>')
.description('Remove a repo from a group')
.action(async (groupName: string, repoPath: string) => {
const { getGroupDir, getDefaultGitnexusDir } = await import('../core/group/storage.js');
const { loadGroupConfig } = await import('../core/group/config-parser.js');
const path = await import('node:path');
const fs = await import('node:fs/promises');
const groupDir = getGroupDir(getDefaultGitnexusDir(), groupName);
const config = await loadGroupConfig(groupDir);
if (!(repoPath in config.repos)) {
console.error(`Repo path "${repoPath}" not found in group "${groupName}"`);
process.exitCode = 1;
return;
}
delete config.repos[repoPath];
await fs.writeFile(path.join(groupDir, 'group.yaml'), yaml.dump(config), 'utf-8');
console.log(`Removed "${repoPath}" from group "${groupName}"`);
});
group
.command('list [name]')
.description('List all groups or details of one')
.action(async (name?: string) => {
const { listGroups, getDefaultGitnexusDir, getGroupDir } =
await import('../core/group/storage.js');
if (!name) {
const groups = await listGroups();
if (groups.length === 0) {
console.log('No groups configured. Create one with: gitnexus group create <name>');
return;
}
console.log('Groups:');
groups.forEach((g) => console.log(` ${g}`));
return;
}
const { loadGroupConfig } = await import('../core/group/config-parser.js');
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const config = await loadGroupConfig(groupDir);
console.log(`Group: ${config.name}`);
if (config.description) console.log(`Description: ${config.description}`);
console.log(`\nRepos (${Object.keys(config.repos).length}):`);
for (const [p, id] of Object.entries(config.repos)) {
console.log(` ${p} -> ${id}`);
}
if (config.links.length > 0) {
console.log(`\nManifest links (${config.links.length}):`);
for (const link of config.links) {
console.log(` ${link.from} -> ${link.to} [${link.type}: ${link.contract}]`);
}
}
});
group
.command('status <name>')
.description('Check staleness of group and repos')
.action(async (name: string) => {
const { readContractRegistry, getGroupDir, getDefaultGitnexusDir } =
await import('../core/group/storage.js');
const { LocalBackend } = await import('../mcp/local/local-backend.js');
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const registry = await readContractRegistry(groupDir);
console.log(
`Group: ${name}${registry ? ` (last sync: ${registry.generatedAt})` : ' (never synced)'}\n`,
);
const backend = new LocalBackend();
try {
await backend.init();
const raw = await backend.getGroupService().groupStatus({ name });
const st = raw as {
repos?: Record<
string,
{
indexStale: boolean;
contractsStale: boolean;
missing: boolean;
commitsBehind?: number;
}
>;
missingRepos?: string[];
};
console.log(' Repo index / contracts staleness:');
for (const [repoPath, row] of Object.entries(st.repos || {})) {
if (row.missing) {
console.log(` ${repoPath.padEnd(25)} MISSING (not in registry or unreadable)`);
continue;
}
const idx = row.indexStale
? `STALE (${row.commitsBehind ?? '?'} commits behind)`
: 'OK ';
const ctr = row.contractsStale ? ' CONTRACTS_STALE' : '';
console.log(` ${repoPath.padEnd(25)} ${idx}${ctr}`);
}
if ((st.missingRepos || []).length > 0) {
console.log(`\n Last sync missing repos: ${st.missingRepos!.join(', ')}`);
}
} finally {
await backend.dispose().catch(() => {});
}
});
group
.command('sync <name>')
.description('Sync Contract Registry — extract contracts and build cross-links')
.option('--skip-embeddings', 'Exact + BM25 only (no embedding fallback)')
.option('--exact-only', 'Exact match only')
.option('--allow-stale', 'Skip stale index warnings')
.option('--verbose', 'Show each cross-link detail')
.option('--json', 'JSON output')
.action(async (name: string, opts: Record<string, boolean | undefined>) => {
const { getGroupDir, getDefaultGitnexusDir } = await import('../core/group/storage.js');
const { loadGroupConfig } = await import('../core/group/config-parser.js');
const { syncGroup } = await import('../core/group/sync.js');
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const config = await loadGroupConfig(groupDir);
console.log(`Syncing group "${name}" (${Object.keys(config.repos).length} repos)...\n`);
const result = await syncGroup(config, {
groupDir,
allowStale: Boolean(opts.allowStale),
verbose: Boolean(opts.verbose),
skipEmbeddings: Boolean(opts.skipEmbeddings),
exactOnly: Boolean(opts.exactOnly),
});
if (opts.json) {
console.log(JSON.stringify(result, null, 2));
} else {
console.log(`\nMatching cascade:`);
const exactLinks = result.crossLinks.filter((l) => l.matchType === 'exact');
console.log(` exact: ${exactLinks.length} cross-links (confidence 1.0)`);
console.log(` unmatched: ${result.unmatched.length} contracts`);
console.log(
`\nWrote contracts.json (${result.contracts.length} contracts, ${result.crossLinks.length} cross-links)`,
);
}
});
group
.command('query <name> <query>')
.description('Search execution flows across all repos in a group')
.option('--subgroup <path>', 'Limit search scope')
.option('--limit <n>', 'Max merged results', '5')
.option('--json', 'JSON output')
.action(
async (
name: string,
queryText: string,
opts: Record<string, string | boolean | undefined>,
) => {
const { LocalBackend } = await import('../mcp/local/local-backend.js');
const limit = parseInt(String(opts.limit ?? '5'), 10) || 5;
const subgroup = opts.subgroup as string | undefined;
const backend = new LocalBackend();
try {
await backend.init();
console.log(`Searching "${queryText}" across group "${name}"...\n`);
const raw = await backend.getGroupService().groupQuery({
name,
query: queryText,
limit,
subgroup,
});
const merged = raw as {
results: Array<Record<string, unknown>>;
per_repo: Array<{ repo: string; count: number }>;
};
if (opts.json) {
console.log(JSON.stringify(raw, null, 2));
} else {
console.log(`Results (top ${merged.results.length}):\n`);
for (const p of merged.results) {
const label = (p.summary || p.heuristicLabel || p.name || 'unnamed') as string;
console.log(` [${p._repo}] ${label} (rrf: ${(p._rrf_score as number).toFixed(4)})`);
}
if (merged.results.length === 0) {
console.log(' No matching execution flows found.');
}
}
} finally {
await backend.dispose().catch(() => {});
}
},
);
group
.command('contracts <name>')
.description('Inspect Contract Registry')
.option('--type <type>', 'Filter by contract type')
.option('--repo <repo>', 'Filter by repo')
.option('--unmatched', 'Show only unmatched contracts')
.option('--json', 'JSON output')
.action(async (name: string, opts: Record<string, string | boolean | undefined>) => {
const { LocalBackend } = await import('../mcp/local/local-backend.js');
const backend = new LocalBackend();
try {
await backend.init();
const raw = await backend.getGroupService().groupContracts({
name,
type: opts.type as string | undefined,
repo: opts.repo as string | undefined,
unmatchedOnly: Boolean(opts.unmatched),
});
if (raw && typeof raw === 'object' && 'error' in raw) {
console.error(String((raw as { error: string }).error));
process.exitCode = 1;
return;
}
const { contracts, crossLinks } = raw as {
contracts: Array<{
role: string;
contractId: string;
repo: string;
symbolRef: { name: string };
}>;
crossLinks: Array<{
from: { repo: string };
to: { repo: string };
matchType: string;
confidence: number;
contractId: string;
}>;
};
if (opts.json) {
console.log(JSON.stringify({ contracts, crossLinks }, null, 2));
} else {
console.log(`Contracts (${contracts.length}):`);
for (const c of contracts) {
console.log(` [${c.role}] ${c.contractId} (${c.repo}) ${c.symbolRef.name}`);
}
console.log(`\nCross-links (${crossLinks.length}):`);
for (const l of crossLinks) {
console.log(
` ${l.from.repo} -> ${l.to.repo} [${l.matchType}, conf=${l.confidence}] ${l.contractId}`,
);
}
}
} finally {
await backend.dispose().catch(() => {});
}
});
}
+3
View File
@@ -6,6 +6,7 @@
import { Command } from 'commander';
import { createRequire } from 'node:module';
import { createLazyAction } from './lazy-action.js';
import { registerGroupCommands } from './group.js';
const _require = createRequire(import.meta.url);
const pkg = _require('../../package.json');
@@ -148,4 +149,6 @@ program
.option('--idle-timeout <seconds>', 'Auto-shutdown after N seconds idle (0 = disabled)', '0')
.action(createLazyAction(() => import('./eval-server.js'), 'evalServerCommand'));
registerGroupCommands(program);
program.parse(process.argv);
+4 -1
View File
@@ -14,7 +14,10 @@ process.on('unhandledRejection', (reason: any) => {
export const serveCommand = async (options?: { port?: string; host?: string }) => {
const port = Number(options?.port ?? 4747);
const host = options?.host ?? '127.0.0.1';
// Default to 'localhost' so the OS decides whether to bind to 127.0.0.1 or
// ::1 based on system configuration, avoiding spurious CORS errors when the
// hosted frontend at gitnexus.vercel.app connects to localhost.
const host = options?.host ?? 'localhost';
try {
await createServer(port, host);
+35 -2
View File
@@ -9,7 +9,7 @@
import fs from 'fs/promises';
import path from 'path';
import os from 'os';
import { execFile } from 'child_process';
import { execFile, execFileSync } from 'child_process';
import { promisify } from 'util';
import { fileURLToPath } from 'url';
import { glob } from 'glob';
@@ -25,11 +25,44 @@ interface SetupResult {
errors: string[];
}
/**
* Resolve the absolute path to the `gitnexus` binary if it's installed
* globally (or via npm -g / yarn global). Returns null when not found.
*/
function resolveGitnexusBin(): string | null {
try {
const cmd = process.platform === 'win32' ? 'where' : 'which';
const resolved = execFileSync(cmd, ['gitnexus'], {
encoding: 'utf-8',
timeout: 5000,
stdio: ['ignore', 'pipe', 'ignore'],
})
.split('\n')[0]
.trim();
return resolved || null;
} catch {
return null;
}
}
/**
* The MCP server entry for all editors.
* On Windows, npx must be invoked via cmd /c since it's a .cmd script.
*
* Prefers the globally-installed `gitnexus` binary (starts in ~1 s) over
* `npx -y gitnexus@latest` (cold-cache install of native deps can take
* >60 s, exceeding Claude Code's 30 s MCP connection timeout).
*
* Falls back to npx when the binary isn't on PATH — e.g. first-time
* users who ran `npx gitnexus analyze` but haven't done `npm i -g`.
*/
function getMcpEntry() {
const bin = resolveGitnexusBin();
if (bin) {
return { command: bin, args: ['mcp'] };
}
// Fallback: npx (works without a global install, but slow cold-start)
if (process.platform === 'win32') {
return {
command: 'cmd',
+18 -58
View File
@@ -232,65 +232,28 @@ export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptio
llmConfig = { ...llmConfig, provider: 'cursor', model, apiKey: '', baseUrl: '' };
} else if (choice === '3') {
// Azure OpenAI guided setup
console.log('\n Azure OpenAI setup.');
console.log(
' You need: your resource name, deployment name, and API key from the Azure portal.\n',
);
// Azure OpenAI guided setup — minimal prompts
console.log('\n Azure OpenAI setup.\n');
const resourceName = (
await prompt(' Azure resource name (e.g. my-openai-resource): ')
).trim();
if (!resourceName) {
console.log('\n No resource name provided. Aborting.\n');
const endpoint = (
await prompt(' Endpoint URL (e.g. https://my-resource.openai.azure.com): ')
)
.trim()
.replace(/\/+$/, '');
if (!endpoint) {
console.log('\n No endpoint provided. Aborting.\n');
process.exitCode = 1;
return;
}
const deploymentName = (
await prompt(' Deployment name (the name you gave your model deployment): ')
).trim();
const deploymentName = (await prompt(' Deployment name: ')).trim();
if (!deploymentName) {
console.log('\n No deployment name provided. Aborting.\n');
process.exitCode = 1;
return;
}
// Offer v1 or legacy URL
console.log('\n API format:');
console.log(' [1] v1 API — recommended (no api-version needed)');
console.log(' [2] Legacy — uses api-version query param\n');
const apiFormat = await prompt(' Select format (1/2, default: 1): ');
let azureApiVersion: string | undefined;
let azureBaseUrl: string;
if (apiFormat === '2') {
const versionInput = await prompt(' api-version (default: 2024-10-21): ');
azureApiVersion = versionInput || '2024-10-21';
azureBaseUrl = `https://${resourceName}.openai.azure.com/openai/deployments/${deploymentName}`;
} else {
azureBaseUrl = `https://${resourceName}.openai.azure.com/openai/v1`;
azureApiVersion = undefined;
}
defaultModel = deploymentName;
// Ask if this is a reasoning model deployment
const reasoningAnswer = await prompt(
' Is this a reasoning model (o1, o3, o4-mini)? (y/N): ',
);
const isReasoningModelDeployment = ['y', 'yes'].includes(reasoningAnswer.toLowerCase());
if (isReasoningModelDeployment) {
console.log(
' Note: temperature and max_tokens will be omitted for this deployment (Azure reasoning model requirement).\n',
);
}
const modelInput = await prompt(` Model / deployment name (default: ${defaultModel}): `);
const model = modelInput || defaultModel;
// API key
// API key — use env var if available
const envKey = process.env.GITNEXUS_API_KEY || process.env.OPENAI_API_KEY || '';
let azureKey: string;
if (envKey) {
@@ -311,26 +274,23 @@ export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptio
return;
}
// Save Azure config including optional apiVersion and isReasoningModel
const azureConfig: Parameters<typeof saveCLIConfig>[0] = {
// Always use v1 API format — no need for api-version
const azureBaseUrl = `${endpoint}/openai/v1`;
await saveCLIConfig({
apiKey: azureKey,
baseUrl: azureBaseUrl,
model,
model: deploymentName,
provider: 'azure',
isReasoningModel: isReasoningModelDeployment,
};
if (azureApiVersion) azureConfig.apiVersion = azureApiVersion;
await saveCLIConfig(azureConfig);
});
console.log(' Config saved to ~/.gitnexus/config.json\n');
llmConfig = {
...llmConfig,
apiKey: azureKey,
baseUrl: azureBaseUrl,
model,
model: deploymentName,
provider: 'azure',
apiVersion: azureApiVersion,
isReasoningModel: isReasoningModelDeployment,
};
} else {
// OpenAI-compatible provider (OpenAI, OpenRouter, Custom)
+8 -3
View File
@@ -398,11 +398,16 @@ export const createIgnoreFilter = async (repoPath: string, options?: IgnoreOptio
// defense-in-depth — do not remove `dot: false` assuming this covers it.
if (DEFAULT_IGNORE_LIST.has(p.name)) return true;
// Check against .gitignore / .gitnexusignore patterns.
// Test both bare path and path with trailing slash to handle
// bare-name patterns (e.g. `local`) and dir-only patterns (e.g. `local/`).
// Since childrenIgnored is only called for directories, always test with
// a trailing slash. This ensures directory-only negation patterns (e.g.
// `!iOS/`) are applied correctly — without the slash, `ig.ignores('iOS')`
// treats the path as a file and misses the negation.
// Bare-name patterns (e.g. `local`) still match `local/` per gitignore spec:
// the `ignore` package normalizes `dir` and `dir/` to match directories.
// See: https://github.com/kaelzhang/node-ignore#2-filenames-and-dirnames
if (ig) {
const rel = p.relative();
if (rel && (ig.ignores(rel) || ig.ignores(rel + '/'))) return true;
if (rel && ig.ignores(rel + '/')) return true;
}
return false;
},
+1 -1
View File
@@ -93,7 +93,7 @@ export async function augment(pattern: string, cwd?: string): Promise<string> {
if (!repo) return '';
// Lazy-load lbug adapter (skip unnecessary init)
const { initLbug, executeQuery, isLbugReady } = await import('../../mcp/core/lbug-adapter.js');
const { initLbug, executeQuery, isLbugReady } = await import('../lbug/pool-adapter.js');
const { searchFTSFromLbug } = await import('../search/bm25-index.js');
const repoId = repo.name.toLowerCase();
+39
View File
@@ -0,0 +1,39 @@
/**
* Git working tree vs index commit staleness (used by MCP resources, group status, etc.).
* Lives in core/ so application code does not depend on the MCP package layer.
*/
import { execFileSync } from 'node:child_process';
export interface StalenessInfo {
isStale: boolean;
commitsBehind: number;
hint?: string;
}
/**
* Check how many commits the index is behind HEAD (synchronous; uses git CLI).
*/
export function checkStaleness(repoPath: string, lastCommit: string): StalenessInfo {
try {
const result = execFileSync('git', ['rev-list', '--count', `${lastCommit}..HEAD`], {
cwd: repoPath,
encoding: 'utf-8',
stdio: ['pipe', 'pipe', 'pipe'],
}).trim();
const commitsBehind = parseInt(result, 10) || 0;
if (commitsBehind > 0) {
return {
isStale: true,
commitsBehind,
hint: `⚠️ Index is ${commitsBehind} commit${commitsBehind > 1 ? 's' : ''} behind HEAD. Run analyze tool to update.`,
};
}
return { isStale: false, commitsBehind: 0 };
} catch {
return { isStale: false, commitsBehind: 0 };
}
}
+98
View File
@@ -0,0 +1,98 @@
import { createRequire } from 'node:module';
import type { GroupConfig, GroupManifestLink, ContractType, ContractRole } from './types.js';
const _require = createRequire(import.meta.url);
const yaml = _require('js-yaml') as typeof import('js-yaml');
const VALID_CONTRACT_TYPES: ContractType[] = ['http', 'grpc', 'topic', 'lib', 'custom'];
const VALID_ROLES: ContractRole[] = ['provider', 'consumer'];
const DEFAULT_DETECT = {
http: true,
grpc: true,
topics: true,
shared_libs: true,
embedding_fallback: true,
};
const DEFAULT_MATCHING = {
bm25_threshold: 0.7,
embedding_threshold: 0.65,
max_candidates_per_step: 3,
};
export function parseGroupConfig(yamlContent: string): GroupConfig {
const raw = yaml.load(yamlContent, { schema: yaml.JSON_SCHEMA }) as Record<string, unknown>;
if (!raw || typeof raw !== 'object' || Array.isArray(raw)) {
throw new Error('Invalid YAML: expected an object');
}
if (raw.version === undefined) throw new Error('version is required in group.yaml');
if (raw.version !== 1) {
throw new Error(`Unsupported group.yaml version: ${raw.version}. Expected 1.`);
}
if (!raw.name || typeof raw.name !== 'string') throw new Error('name is required in group.yaml');
if (!raw.repos || typeof raw.repos !== 'object' || Array.isArray(raw.repos)) {
throw new Error('repos is required in group.yaml (must be a mapping)');
}
const repos = raw.repos as Record<string, string>;
const repoPaths = new Set(Object.keys(repos));
const rawLinks = (raw.links as unknown[]) || [];
const links: GroupManifestLink[] = rawLinks.map((l: unknown, i: number) => {
const link = l as Record<string, unknown>;
if (!link.from || !repoPaths.has(link.from as string)) {
throw new Error(`links[${i}].from "${link.from}" does not match any repo path in group`);
}
if (!link.to || !repoPaths.has(link.to as string)) {
throw new Error(`links[${i}].to "${link.to}" does not match any repo path in group`);
}
if (!VALID_CONTRACT_TYPES.includes(link.type as ContractType)) {
throw new Error(
`links[${i}].type "${link.type}" is invalid. Expected: ${VALID_CONTRACT_TYPES.join(', ')}`,
);
}
if (!VALID_ROLES.includes(link.role as ContractRole)) {
throw new Error(`links[${i}].role "${link.role}" is invalid. Expected: provider | consumer`);
}
if (
link.contract === undefined ||
link.contract === null ||
String(link.contract).trim() === ''
) {
throw new Error(`links[${i}].contract is required`);
}
return {
from: link.from as string,
to: link.to as string,
type: link.type as ContractType,
contract: String(link.contract),
role: link.role as ContractRole,
};
});
const detect = { ...DEFAULT_DETECT, ...((raw.detect as object) || {}) };
const matching = { ...DEFAULT_MATCHING, ...((raw.matching as object) || {}) };
const packages = (raw.packages as Record<string, Record<string, string>>) || {};
return {
version: 1,
name: raw.name as string,
description: (raw.description as string) || '',
repos,
links,
packages,
detect,
matching,
};
}
export async function loadGroupConfig(groupDir: string): Promise<GroupConfig> {
const fsp = await import('node:fs/promises');
const path = await import('node:path');
const yamlPath = path.join(groupDir, 'group.yaml');
const content = await fsp.readFile(yamlPath, 'utf-8');
return parseGroupConfig(content);
}
@@ -0,0 +1,16 @@
import type { ContractType, ExtractedContract, RepoHandle } from './types.js';
export interface ContractExtractor {
type: ContractType;
canExtract(repo: RepoHandle): Promise<boolean>;
extract(
dbExecutor: CypherExecutor | null,
repoPath: string,
repo: RepoHandle,
): Promise<ExtractedContract[]>;
}
export type CypherExecutor = (
query: string,
params?: Record<string, unknown>,
) => Promise<Record<string, unknown>[]>;
@@ -0,0 +1,357 @@
import * as fs from 'node:fs';
import * as path from 'node:path';
import { glob } from 'glob';
import type { ContractExtractor, CypherExecutor } from '../contract-extractor.js';
import type { ExtractedContract, RepoHandle } from '../types.js';
function readSafe(repoPath: string, rel: string): string | null {
const abs = path.resolve(repoPath, rel);
const base = path.resolve(repoPath);
const relToBase = path.relative(base, abs);
if (relToBase.startsWith('..') || path.isAbsolute(relToBase)) return null;
try {
return fs.readFileSync(abs, 'utf-8');
} catch {
return null;
}
}
function contractId(pkg: string, service: string, method: string): string {
const prefix = pkg ? `${pkg}.${service}` : service;
return `grpc::${prefix}/${method}`;
}
function serviceOnlyContractId(serviceName: string): string {
return `grpc::${serviceName}/*`;
}
function extractServiceBlocks(content: string): Array<{ name: string; body: string }> {
const results: Array<{ name: string; body: string }> = [];
// v1: brace-depth only — braces inside comments or string literals are not filtered (see spec Fix 2)
const headerRe = /service\s+(\w+)\s*\{/g;
let headerMatch: RegExpExecArray | null;
while ((headerMatch = headerRe.exec(content)) !== null) {
const serviceName = headerMatch[1];
const bodyStart = headerMatch.index + headerMatch[0].length;
let depth = 1;
let pos = bodyStart;
while (pos < content.length && depth > 0) {
const ch = content[pos];
if (ch === '{') depth++;
else if (ch === '}') depth--;
pos++;
}
// If EOF before depth returns to 0, skip incomplete service
if (depth !== 0) continue;
// body is between opening { (consumed by regex) and closing } (pos is one past it)
const body = content.slice(bodyStart, pos - 1);
results.push({ name: serviceName, body });
}
return results;
}
function makeContract(
cid: string,
role: 'provider' | 'consumer',
filePath: string,
symbolName: string,
confidence: number,
meta: Record<string, unknown>,
): ExtractedContract {
return {
contractId: cid,
type: 'grpc',
role,
symbolUid: '',
symbolRef: { filePath: filePath.replace(/\\/g, '/'), name: symbolName },
symbolName,
confidence,
meta: { ...meta, extractionStrategy: 'source_scan' },
};
}
export class GrpcExtractor implements ContractExtractor {
type = 'grpc' as const;
async canExtract(_repo: RepoHandle): Promise<boolean> {
return true;
}
async extract(
_dbExecutor: CypherExecutor | null,
repoPath: string,
_repo: RepoHandle,
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
// Proto files — definitive provider source
const protoFiles = await glob('**/*.proto', {
cwd: repoPath,
ignore: ['**/node_modules/**', '**/.git/**', '**/vendor/**'],
nodir: true,
});
for (const rel of protoFiles) {
const content = readSafe(repoPath, rel);
if (content) out.push(...this.parseProtoFile(content, rel));
}
// Source files — server/client detection
const sourceFiles = await glob('**/*.{go,java,py,ts,tsx,js,jsx}', {
cwd: repoPath,
ignore: ['**/node_modules/**', '**/.git/**', '**/vendor/**', '**/dist/**', '**/build/**'],
nodir: true,
});
for (const rel of sourceFiles) {
const content = readSafe(repoPath, rel);
if (!content) continue;
const ext = path.extname(rel).toLowerCase();
if (ext === '.go') {
out.push(...this.scanGoProviders(content, rel));
out.push(...this.scanGoConsumers(content, rel));
} else if (ext === '.java') {
out.push(...this.scanJavaProviders(content, rel));
out.push(...this.scanJavaConsumers(content, rel));
} else if (ext === '.py') {
out.push(...this.scanPythonProviders(content, rel));
out.push(...this.scanPythonConsumers(content, rel));
} else if (['.ts', '.tsx', '.js', '.jsx'].includes(ext)) {
out.push(...this.scanTsProviders(content, rel));
}
}
return this.dedupe(out);
}
private parseProtoFile(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
const pkgMatch = content.match(/^package\s+([\w.]+)\s*;/m);
const pkg = pkgMatch ? pkgMatch[1] : '';
for (const { name: serviceName, body } of extractServiceBlocks(content)) {
const rpcRe = /rpc\s+(\w+)\s*\(/g;
let rpcMatch: RegExpExecArray | null;
while ((rpcMatch = rpcRe.exec(body)) !== null) {
const methodName = rpcMatch[1];
const cid = contractId(pkg, serviceName, methodName);
out.push(
makeContract(cid, 'provider', filePath, `${serviceName}.${methodName}`, 0.85, {
package: pkg,
service: serviceName,
method: methodName,
source: 'proto',
}),
);
}
}
return out;
}
private scanGoProviders(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
// pb.RegisterXxxServer(
const registerRe = /\w+\.Register(\w+)Server\s*\(/g;
let m: RegExpExecArray | null;
while ((m = registerRe.exec(content)) !== null) {
const serviceName = m[1];
out.push(
makeContract(
serviceOnlyContractId(serviceName),
'provider',
filePath,
`Register${serviceName}Server`,
0.8,
{ service: serviceName, source: 'go_register' },
),
);
}
// pb.UnimplementedXxxServer
const unimplRe = /\w+\.Unimplemented(\w+)Server\b/g;
while ((m = unimplRe.exec(content)) !== null) {
const serviceName = m[1];
out.push(
makeContract(
serviceOnlyContractId(serviceName),
'provider',
filePath,
`Unimplemented${serviceName}Server`,
0.8,
{ service: serviceName, source: 'go_unimplemented' },
),
);
}
return out;
}
private scanGoConsumers(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
const re = /\w+\.New(\w+)Client\s*\(/g;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const serviceName = m[1];
out.push(
makeContract(
serviceOnlyContractId(serviceName),
'consumer',
filePath,
`New${serviceName}Client`,
0.7,
{ service: serviceName, source: 'go_client' },
),
);
}
return out;
}
private scanJavaProviders(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
// @GrpcService
if (content.includes('@GrpcService')) {
const implBaseRe = /extends\s+(\w+)Grpc\.(\w+)ImplBase/;
const m = content.match(implBaseRe);
if (m) {
out.push(
makeContract(serviceOnlyContractId(m[1]), 'provider', filePath, m[2], 0.8, {
service: m[1],
source: 'java_grpc_service',
}),
);
} else {
// Try extracting service name from class name
const classRe =
/class\s+(\w*?)(?:Grpc)?(?:Service)?\s+extends\s+(\w+)(?:Grpc\.(\w+))?ImplBase/;
const cm = content.match(classRe);
if (cm) {
const svcName = cm[2].replace(/Grpc$/, '');
out.push(
makeContract(serviceOnlyContractId(svcName), 'provider', filePath, cm[1], 0.8, {
service: svcName,
source: 'java_grpc_service',
}),
);
}
}
}
// extends XxxImplBase (without @GrpcService)
if (!content.includes('@GrpcService')) {
const implRe = /extends\s+(\w+?)(?:Grpc\.(\w+))?ImplBase/;
const m = content.match(implRe);
if (m) {
const svcName = m[2] || m[1].replace(/Grpc$/, '');
out.push(
makeContract(serviceOnlyContractId(svcName), 'provider', filePath, svcName, 0.8, {
service: svcName,
source: 'java_impl_base',
}),
);
}
}
return out;
}
private scanJavaConsumers(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
// XxxGrpc.newBlockingStub( or XxxGrpc.newStub(
const re = /(\w+)Grpc\.new(?:Blocking)?Stub\s*\(/g;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const serviceName = m[1];
out.push(
makeContract(
serviceOnlyContractId(serviceName),
'consumer',
filePath,
`${serviceName}Stub`,
0.7,
{ service: serviceName, source: 'java_stub' },
),
);
}
return out;
}
private scanPythonProviders(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
// add_XxxServicer_to_server(
const re = /add_(\w+?)Servicer_to_server\s*\(/g;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const serviceName = m[1];
out.push(
makeContract(
serviceOnlyContractId(serviceName),
'provider',
filePath,
`add_${serviceName}Servicer_to_server`,
0.8,
{ service: serviceName, source: 'python_servicer' },
),
);
}
return out;
}
private scanPythonConsumers(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
// XxxStub(
const re = /(\w+)Stub\s*\(/g;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const name = m[1];
// Filter out common false positives
if (['Mock', 'Test', 'Fake', 'Stub'].includes(name)) continue;
out.push(
makeContract(serviceOnlyContractId(name), 'consumer', filePath, `${name}Stub`, 0.7, {
service: name,
source: 'python_stub',
}),
);
}
return out;
}
private scanTsProviders(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
// @GrpcMethod('ServiceName', 'MethodName')
const re = /@GrpcMethod\s*\(\s*['"](\w+)['"]\s*,\s*['"](\w+)['"]\s*\)/g;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const serviceName = m[1];
const methodName = m[2];
const cid = contractId('', serviceName, methodName);
out.push(
makeContract(cid, 'provider', filePath, `${serviceName}.${methodName}`, 0.8, {
service: serviceName,
method: methodName,
source: 'ts_grpc_method',
}),
);
}
return out;
}
private dedupe(items: ExtractedContract[]): ExtractedContract[] {
const seen = new Set<string>();
const out: ExtractedContract[] = [];
for (const c of items) {
const k = `${c.contractId}|${c.role}|${c.symbolRef.filePath}`;
if (seen.has(k)) continue;
seen.add(k);
out.push(c);
}
return out;
}
}
@@ -0,0 +1,475 @@
import * as fs from 'node:fs';
import * as path from 'node:path';
import { glob } from 'glob';
import type { ContractExtractor, CypherExecutor } from '../contract-extractor.js';
import type { ExtractedContract, RepoHandle } from '../types.js';
const HANDLES_ROUTE_QUERY = `
MATCH (handlerFile:File)-[r:CodeRelation {type: 'HANDLES_ROUTE'}]->(route:Route)
RETURN handlerFile.id AS fileId, handlerFile.filePath AS filePath,
route.name AS routePath, route.id AS routeId,
route.responseKeys AS responseKeys,
r.reason AS routeSource`;
const FETCHES_QUERY = `
MATCH (callerFile:File)-[r:CodeRelation {type: 'FETCHES'}]->(route:Route)
RETURN callerFile.id AS fileId, callerFile.filePath AS filePath,
route.name AS routePath, route.id AS routeId,
r.reason AS fetchReason`;
const CONTAINS_QUERY = `
MATCH (file:File {id: $fileId})<-[:CodeRelation {type: 'CONTAINS'}]-(sym)
WHERE sym.startLine IS NOT NULL
RETURN sym.id AS uid, sym.name AS name, sym.filePath AS filePath, labels(sym) AS labels
ORDER BY sym.startLine`;
export function normalizeHttpPath(p: string): string {
let s = p.trim().split('?')[0].toLowerCase().replace(/\/+$/, '');
s = s.replace(/:\w+/g, '{param}');
s = s.replace(/\{[^}]+\}/g, '{param}');
s = s.replace(/\[[^\]]+\]/g, '{param}');
return s;
}
function methodFromRouteReason(reason: string): string | null {
const r = reason || '';
if (/GetMapping|decorator-Get/i.test(r)) return 'GET';
if (/PostMapping|decorator-Post/i.test(r)) return 'POST';
if (/PutMapping|decorator-Put/i.test(r)) return 'PUT';
if (/DeleteMapping|decorator-Delete/i.test(r)) return 'DELETE';
if (/PatchMapping|decorator-Patch/i.test(r)) return 'PATCH';
return null;
}
function contractIdFor(method: string, pathNorm: string): string {
return `http::${method.toUpperCase()}::${pathNorm}`;
}
function readSafe(repoPath: string, rel: string): string | null {
const abs = path.resolve(repoPath, rel);
const base = path.resolve(repoPath);
const relToBase = path.relative(base, abs);
if (relToBase.startsWith('..') || path.isAbsolute(relToBase)) return null;
try {
return fs.readFileSync(abs, 'utf-8');
} catch {
return null;
}
}
function pickJavaHandlerName(
content: string,
routePath: string,
httpMethod: string,
): string | null {
const tail = routePath.split('/').filter(Boolean).pop() || '';
const mapNames: Record<string, string> = {
GET: 'GetMapping',
POST: 'PostMapping',
PUT: 'PutMapping',
DELETE: 'DeleteMapping',
PATCH: 'PatchMapping',
};
const ann = mapNames[httpMethod] || 'GetMapping';
const lines = content.split(/\r?\n/);
for (let i = 0; i < lines.length; i++) {
const line = lines[i];
if (!line.includes(`@${ann}`)) continue;
if (!line.includes(`"${tail}"`) && !line.includes(`'${tail}'`) && tail && !line.includes(tail))
continue;
for (let j = i + 1; j < Math.min(i + 8, lines.length); j++) {
const m = lines[j].match(/(?:public|protected|private)\s+[\w<>,\s\[\]]+\s+(\w+)\s*\(/);
if (m) return m[1];
}
}
return null;
}
function pickSymbolUid(
rows: Record<string, unknown>[],
preferredName: string | null,
): { uid: string; name: string; filePath: string } {
const norm = (x: unknown) => String(x ?? '');
const labeled = rows.filter((r) => {
const labels = r.labels ?? r[3];
const s = JSON.stringify(labels);
return s.includes('Method') || s.includes('Function');
});
const pool = labeled.length > 0 ? labeled : rows;
if (preferredName) {
const hit = pool.find((r) => norm(r.name ?? r[1]) === preferredName);
if (hit) {
return {
uid: norm(hit.uid ?? hit[0]),
name: norm(hit.name ?? hit[1]),
filePath: norm(hit.filePath ?? hit[2]),
};
}
}
const first = pool[0] || rows[0];
return {
uid: norm(first?.uid ?? first?.[0]),
name: norm(first?.name ?? first?.[1]),
filePath: norm(first?.filePath ?? first?.[2]),
};
}
export class HttpRouteExtractor implements ContractExtractor {
type = 'http' as const;
async canExtract(_repo: RepoHandle): Promise<boolean> {
return true;
}
async extract(
dbExecutor: CypherExecutor | null,
repoPath: string,
repo: RepoHandle,
): Promise<ExtractedContract[]> {
const graphP = dbExecutor != null ? await this.extractProvidersGraph(dbExecutor, repoPath) : [];
const providers = graphP.length > 0 ? graphP : await this.extractProvidersSourceScan(repoPath);
const graphC = dbExecutor != null ? await this.extractConsumersGraph(dbExecutor, repoPath) : [];
const consumers = graphC.length > 0 ? graphC : await this.extractConsumersSourceScan(repoPath);
return [...providers, ...consumers];
}
private async extractProvidersGraph(
db: CypherExecutor,
repoPath: string,
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
let rows: Record<string, unknown>[];
try {
rows = await db(HANDLES_ROUTE_QUERY);
} catch {
return [];
}
for (const row of rows) {
const filePath = String(row.filePath ?? '');
const routePath = String(row.routePath ?? '');
const routeSource = String(row.routeSource ?? row.routeReason ?? '');
let method = methodFromRouteReason(routeSource);
const content = readSafe(repoPath, filePath);
if (!method && content) {
method = this.inferMethodFromFileScan(content, routePath, 'provider');
}
if (!method) method = 'GET';
const pathNorm = normalizeHttpPath(routePath);
const cid = contractIdFor(method, pathNorm);
const handlerName =
content && routePath ? pickJavaHandlerName(content, routePath, method) : null;
let symbolUid = '';
let symbolName = path.basename(filePath) || 'handler';
let symPath = filePath;
const fileId = row.fileId ?? row[0];
if (fileId) {
try {
const syms = await db(CONTAINS_QUERY, { fileId });
if (syms.length > 0) {
const picked = pickSymbolUid(syms, handlerName);
symbolUid = picked.uid;
symbolName = picked.name;
symPath = picked.filePath || filePath;
}
} catch {
/* ignore */
}
}
out.push({
contractId: cid,
type: 'http',
role: 'provider',
symbolUid,
symbolRef: { filePath: symPath, name: symbolName },
symbolName,
confidence: 0.9,
meta: {
method,
path: pathNorm,
pathSegments: pathNorm.split('/').filter(Boolean),
extractionStrategy: 'graph_assisted',
routeSource,
},
});
}
return out;
}
private inferMethodFromFileScan(
content: string,
routePath: string,
_role: string,
): string | null {
const tail = routePath.split('/').filter(Boolean).pop() || '';
for (const m of ['GET', 'POST', 'PUT', 'DELETE', 'PATCH'] as const) {
const mapNames: Record<string, string> = {
GET: 'GetMapping',
POST: 'PostMapping',
PUT: 'PutMapping',
DELETE: 'DeleteMapping',
PATCH: 'PatchMapping',
};
if (
content.includes(`@${mapNames[m]}`) &&
(content.includes(tail) || routePath.includes(tail))
) {
return m;
}
}
return null;
}
private async extractProvidersSourceScan(repoPath: string): Promise<ExtractedContract[]> {
const files = await glob('**/*.{ts,tsx,js,jsx,java,vue,svelte,php,py}', {
cwd: repoPath,
ignore: ['**/node_modules/**', '**/.git/**', '**/dist/**', '**/build/**'],
nodir: true,
});
const out: ExtractedContract[] = [];
for (const rel of files) {
const content = readSafe(repoPath, rel);
if (!content) continue;
out.push(...this.scanSpringProviders(content, rel));
out.push(...this.scanExpressProviders(content, rel));
out.push(...this.scanLaravelProviders(content, rel));
out.push(...this.scanFastApiProviders(content, rel));
}
return this.dedupeContracts(out);
}
private dedupeContracts(items: ExtractedContract[]): ExtractedContract[] {
const seen = new Set<string>();
const out: ExtractedContract[] = [];
for (const c of items) {
const k = `${c.contractId}|${c.symbolRef.filePath}|${c.symbolRef.name}`;
if (seen.has(k)) continue;
seen.add(k);
out.push(c);
}
return out;
}
private scanSpringProviders(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
let classPrefix = '';
const classRm = content.match(/@RequestMapping\s*\(\s*"([^"]+)"/);
if (classRm) classPrefix = classRm[1].replace(/\/+$/, '');
const re = /@(Get|Post|Put|Delete|Patch)Mapping\s*\(\s*"([^"]+)"/gi;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const method = m[1].toUpperCase();
let p = m[2];
if (classPrefix) p = `${classPrefix}/${p.replace(/^\//, '')}`;
const pathNorm = normalizeHttpPath(p);
const sub = content.slice(m.index);
const nameM = sub.match(/(?:public|protected|private)\s+[\w<>,\s\[\]]+\s+(\w+)\s*\(/);
const name = nameM ? nameM[1] : m[0];
out.push(this.makeProvider(filePath, method, pathNorm, name, 0.8));
}
return out;
}
private scanExpressProviders(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
const re = /(?:router|app)\.(get|post|put|delete|patch)\s*\(\s*['"]([^'"]+)['"]/gi;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const method = m[1].toUpperCase();
const pathNorm = normalizeHttpPath(m[2]);
out.push(this.makeProvider(filePath, method, pathNorm, 'handler', 0.8));
}
return out;
}
private scanLaravelProviders(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
const re = /Route::(get|post|put|delete|patch)\s*\(\s*['"]([^'"]+)['"]/gi;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const method = m[1].toUpperCase();
const pathNorm = normalizeHttpPath(m[2]);
out.push(this.makeProvider(filePath, method, pathNorm, 'route', 0.8));
}
return out;
}
private scanFastApiProviders(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
const re = /@app\.(get|post|put|delete|patch)\s*\(\s*['"]([^'"]+)['"]/gi;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const method = m[1].toUpperCase();
const pathNorm = normalizeHttpPath(m[2]);
out.push(this.makeProvider(filePath, method, pathNorm, 'handler', 0.8));
}
return out;
}
private makeProvider(
filePath: string,
method: string,
pathNorm: string,
name: string,
confidence: number,
): ExtractedContract {
const cid = contractIdFor(method, pathNorm);
return {
contractId: cid,
type: 'http',
role: 'provider',
symbolUid: '',
symbolRef: { filePath, name },
symbolName: name,
confidence,
meta: {
method,
path: pathNorm,
pathSegments: pathNorm.split('/').filter(Boolean),
extractionStrategy: 'source_scan',
},
};
}
private async extractConsumersGraph(
db: CypherExecutor,
repoPath: string,
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
let rows: Record<string, unknown>[];
try {
rows = await db(FETCHES_QUERY);
} catch {
return [];
}
for (const row of rows) {
const filePath = String(row.filePath ?? '');
const routePath = String(row.routePath ?? '');
const pathNorm = normalizeHttpPath(routePath);
let method = 'GET';
const content = readSafe(repoPath, filePath);
if (content) {
const inferred = this.inferFetchMethod(content, pathNorm);
if (inferred) method = inferred;
}
const cid = contractIdFor(method, pathNorm);
let symbolUid = '';
let symbolName = 'fetch';
let symPath = filePath;
const fileId = row.fileId ?? row[0];
if (fileId) {
try {
const syms = await db(CONTAINS_QUERY, { fileId });
if (syms.length > 0) {
const picked = pickSymbolUid(syms, null);
symbolUid = picked.uid;
symbolName = picked.name;
symPath = picked.filePath || filePath;
}
} catch {
/* ignore */
}
}
out.push({
contractId: cid,
type: 'http',
role: 'consumer',
symbolUid,
symbolRef: { filePath: symPath, name: symbolName },
symbolName,
confidence: 0.9,
meta: {
method,
path: pathNorm,
extractionStrategy: 'graph_assisted',
fetchReason: String(row.fetchReason ?? ''),
},
});
}
return out;
}
private inferFetchMethod(content: string, pathNorm: string): string | null {
const esc = pathNorm.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
const fetchRe = new RegExp(
`fetch\\s*\\(\\s*['"\`]([^'"\`]*${esc}[^'"\`]*)['"\`]\\s*,\\s*\\{[^}]*method:\\s*['"](\\w+)['"]`,
'i',
);
const m = content.match(fetchRe);
if (m) return m[2].toUpperCase();
return null;
}
private async extractConsumersSourceScan(repoPath: string): Promise<ExtractedContract[]> {
const files = await glob('**/*.{ts,tsx,js,jsx,vue,svelte}', {
cwd: repoPath,
ignore: ['**/node_modules/**', '**/.git/**'],
nodir: true,
});
const out: ExtractedContract[] = [];
for (const rel of files) {
const content = readSafe(repoPath, rel);
if (!content) continue;
out.push(...this.scanFetchConsumers(content, rel));
out.push(...this.scanAxiosConsumers(content, rel));
}
return this.dedupeContracts(out);
}
private scanFetchConsumers(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
const re =
/fetch\s*\(\s*['"`]([^'"`]+)['"`](?:\s*,\s*\{[^}]*method:\s*['"](\w+)['"][^}]*\})?\s*\)/gi;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const pathNorm = normalizeHttpPath(this.templateToPattern(m[1]));
const method = (m[2] || 'GET').toUpperCase();
out.push(this.makeConsumer(filePath, method, pathNorm, 0.7));
}
return out;
}
private templateToPattern(url: string): string {
return url.replace(/\$\{[^}]+\}/g, '{param}');
}
private scanAxiosConsumers(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
const re = /axios\.(get|post|put|delete|patch)\s*\(\s*[`'"]([^`'"]+)[`'"]/gi;
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const method = m[1].toUpperCase();
const pathNorm = normalizeHttpPath(this.templateToPattern(m[2]));
out.push(this.makeConsumer(filePath, method, pathNorm, 0.7));
}
return out;
}
private makeConsumer(
filePath: string,
method: string,
pathNorm: string,
confidence: number,
): ExtractedContract {
return {
contractId: contractIdFor(method, pathNorm),
type: 'http',
role: 'consumer',
symbolUid: '',
symbolRef: { filePath, name: 'fetch' },
symbolName: 'fetch',
confidence,
meta: {
method,
path: pathNorm,
extractionStrategy: 'source_scan',
},
};
}
}
@@ -0,0 +1,277 @@
import * as fs from 'node:fs';
import * as path from 'node:path';
import { glob } from 'glob';
import type { ContractExtractor, CypherExecutor } from '../contract-extractor.js';
import type { ExtractedContract, RepoHandle } from '../types.js';
type Broker = 'kafka' | 'rabbitmq' | 'nats';
function readSafe(repoPath: string, rel: string): string | null {
const abs = path.resolve(repoPath, rel);
const base = path.resolve(repoPath);
const relToBase = path.relative(base, abs);
if (relToBase.startsWith('..') || path.isAbsolute(relToBase)) return null;
try {
return fs.readFileSync(abs, 'utf-8');
} catch {
return null;
}
}
function makeContract(
topicName: string,
role: 'provider' | 'consumer',
filePath: string,
symbolName: string,
confidence: number,
broker: Broker,
): ExtractedContract {
return {
contractId: `topic::${topicName}`,
type: 'topic',
role,
symbolUid: '',
symbolRef: { filePath: filePath.replace(/\\/g, '/'), name: symbolName },
symbolName,
confidence,
meta: {
broker,
topicName,
extractionStrategy: 'source_scan',
},
};
}
interface PatternDef {
regex: RegExp;
role: 'provider' | 'consumer';
broker: Broker;
confidence: number;
topicGroup: number;
symbolName: string;
}
// --- Kafka patterns ---
const KAFKA_PATTERNS: PatternDef[] = [
// Java: @KafkaListener(topics = "xxx")
{
regex: /@KafkaListener\s*\(\s*topics\s*=\s*"([^"]+)"/g,
role: 'consumer',
broker: 'kafka',
confidence: 0.8,
topicGroup: 1,
symbolName: 'kafkaListener',
},
// Java: kafkaTemplate.send("xxx"
{
regex: /kafkaTemplate\.send\s*\(\s*"([^"]+)"/gi,
role: 'provider',
broker: 'kafka',
confidence: 0.8,
topicGroup: 1,
symbolName: 'kafkaTemplate.send',
},
// Node: producer.send({ topic: 'xxx'
{
regex: /producer\.send\s*\(\s*\{\s*topic:\s*['"]([^'"]+)['"]/g,
role: 'provider',
broker: 'kafka',
confidence: 0.8,
topicGroup: 1,
symbolName: 'producer.send',
},
// Node: consumer.subscribe({ topic: 'xxx'
{
regex: /consumer\.subscribe\s*\(\s*\{\s*topic:\s*['"]([^'"]+)['"]/g,
role: 'consumer',
broker: 'kafka',
confidence: 0.8,
topicGroup: 1,
symbolName: 'consumer.subscribe',
},
// Go: consumer.ConsumePartition("xxx"
{
regex: /\.ConsumePartition\s*\(\s*"([^"]+)"/g,
role: 'consumer',
broker: 'kafka',
confidence: 0.7,
topicGroup: 1,
symbolName: 'ConsumePartition',
},
// Python: KafkaConsumer('xxx'
{
regex: /KafkaConsumer\s*\(\s*['"]([^'"]+)['"]/g,
role: 'consumer',
broker: 'kafka',
confidence: 0.7,
topicGroup: 1,
symbolName: 'KafkaConsumer',
},
// Python: producer.send('xxx' or producer.produce('xxx'
{
regex: /producer\.(?:send|produce)\s*\(\s*['"]([^'"]+)['"]/g,
role: 'provider',
broker: 'kafka',
confidence: 0.7,
topicGroup: 1,
symbolName: 'producer.send',
},
];
// --- RabbitMQ patterns ---
const RABBITMQ_PATTERNS: PatternDef[] = [
// Java: @RabbitListener(queues = "xxx")
{
regex: /@RabbitListener\s*\(\s*queues\s*=\s*"([^"]+)"/g,
role: 'consumer',
broker: 'rabbitmq',
confidence: 0.8,
topicGroup: 1,
symbolName: 'rabbitListener',
},
// Java: rabbitTemplate.convertAndSend("xxx"
{
regex: /rabbitTemplate\.convertAndSend\s*\(\s*"([^"]+)"/gi,
role: 'provider',
broker: 'rabbitmq',
confidence: 0.8,
topicGroup: 1,
symbolName: 'rabbitTemplate.convertAndSend',
},
// Node: channel.consume("xxx"
{
regex: /channel\.consume\s*\(\s*"([^"]+)"/g,
role: 'consumer',
broker: 'rabbitmq',
confidence: 0.8,
topicGroup: 1,
symbolName: 'channel.consume',
},
// Node: channel.publish("xxx"
{
regex: /channel\.publish\s*\(\s*"([^"]+)"/g,
role: 'provider',
broker: 'rabbitmq',
confidence: 0.8,
topicGroup: 1,
symbolName: 'channel.publish',
},
// Node: channel.sendToQueue("xxx"
{
regex: /channel\.sendToQueue\s*\(\s*"([^"]+)"/g,
role: 'provider',
broker: 'rabbitmq',
confidence: 0.8,
topicGroup: 1,
symbolName: 'channel.sendToQueue',
},
// Python: channel.basic_consume(queue='xxx'
{
regex: /channel\.basic_consume\s*\(\s*queue\s*=\s*['"]([^'"]+)['"]/g,
role: 'consumer',
broker: 'rabbitmq',
confidence: 0.7,
topicGroup: 1,
symbolName: 'basic_consume',
},
// Python: channel.basic_publish(exchange='xxx'
{
regex: /channel\.basic_publish\s*\([^)]*exchange\s*=\s*['"]([^'"]+)['"]/g,
role: 'provider',
broker: 'rabbitmq',
confidence: 0.7,
topicGroup: 1,
symbolName: 'basic_publish',
},
];
// --- NATS patterns ---
const NATS_PATTERNS: PatternDef[] = [
// Go/Node: nc.Subscribe("xxx" or nc.subscribe("xxx"
{
regex: /nc\.(?:S|s)ubscribe\s*\(\s*"([^"]+)"/g,
role: 'consumer',
broker: 'nats',
confidence: 0.8,
topicGroup: 1,
symbolName: 'nc.Subscribe',
},
// Go/Node: nc.Publish("xxx" or nc.publish("xxx"
{
regex: /nc\.(?:P|p)ublish\s*\(\s*"([^"]+)"/g,
role: 'provider',
broker: 'nats',
confidence: 0.8,
topicGroup: 1,
symbolName: 'nc.Publish',
},
];
const ALL_PATTERNS: PatternDef[] = [...KAFKA_PATTERNS, ...RABBITMQ_PATTERNS, ...NATS_PATTERNS];
export class TopicExtractor implements ContractExtractor {
type = 'topic' as const;
async canExtract(_repo: RepoHandle): Promise<boolean> {
return true;
}
async extract(
_dbExecutor: CypherExecutor | null,
repoPath: string,
_repo: RepoHandle,
): Promise<ExtractedContract[]> {
const files = await glob('**/*.{ts,tsx,js,jsx,java,go,py}', {
cwd: repoPath,
ignore: ['**/node_modules/**', '**/.git/**', '**/vendor/**', '**/dist/**', '**/build/**'],
nodir: true,
});
const out: ExtractedContract[] = [];
for (const rel of files) {
const content = readSafe(repoPath, rel);
if (!content) continue;
out.push(...this.scanFile(content, rel));
}
return this.dedupe(out);
}
private scanFile(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
for (const pattern of ALL_PATTERNS) {
// Reset regex state for each file
const re = new RegExp(pattern.regex.source, pattern.regex.flags);
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const topicName = m[pattern.topicGroup];
if (!topicName) continue;
out.push(
makeContract(
topicName,
pattern.role,
filePath,
pattern.symbolName,
pattern.confidence,
pattern.broker,
),
);
}
}
return out;
}
private dedupe(items: ExtractedContract[]): ExtractedContract[] {
const seen = new Set<string>();
const out: ExtractedContract[] = [];
for (const c of items) {
const k = `${c.contractId}|${c.role}|${c.symbolRef.filePath}`;
if (seen.has(k)) continue;
seen.add(k);
out.push(c);
}
return out;
}
}
+127
View File
@@ -0,0 +1,127 @@
import type { StoredContract, CrossLink } from './types.js';
export interface MatchResult {
matched: CrossLink[];
unmatched: StoredContract[];
}
export function normalizeContractId(id: string): string {
const colonIdx = id.indexOf('::');
if (colonIdx === -1) return id;
const type = id.substring(0, colonIdx);
const rest = id.substring(colonIdx + 2);
switch (type) {
case 'http': {
const parts = rest.split('::');
if (parts.length >= 2) {
const method = parts[0].toUpperCase();
let pathPart = parts.slice(1).join('::');
pathPart = pathPart.replace(/\/+$/, '');
return `http::${method}::${pathPart}`;
}
return id;
}
case 'grpc': {
const slashIdx = rest.indexOf('/');
if (slashIdx > 0) {
const pkg = rest.substring(0, slashIdx).toLowerCase();
const method = rest.substring(slashIdx);
return `grpc::${pkg}${method}`;
}
if (slashIdx === 0) {
// Malformed "package/method" with leading slash — do not lowercase the whole string
// (method segment is case-sensitive per spec).
return `grpc::${rest}`;
}
// No slash: spec is ambiguous (package-only vs full service.method). MVP: lowercase
// the whole token; differs from pkg/method split above where RPC method keeps case.
return `grpc::${rest.toLowerCase()}`;
}
case 'topic':
return `topic::${rest.trim().toLowerCase()}`;
case 'lib':
return `lib::${rest.toLowerCase()}`;
default:
return id;
}
}
function findMatchingKeys(contractId: string, index: Map<string, StoredContract[]>): string[] {
const normalized = normalizeContractId(contractId);
if (index.has(normalized)) return [normalized];
if (normalized.startsWith('http::*::')) {
const pathPart = normalized.substring('http::*::'.length);
const matches: string[] = [];
for (const key of index.keys()) {
if (key.startsWith('http::') && key.endsWith(`::${pathPart}`)) {
matches.push(key);
}
}
return matches;
}
return [];
}
export function runExactMatch(contracts: StoredContract[]): MatchResult {
const providers = contracts.filter((c) => c.role === 'provider');
const consumers = contracts.filter((c) => c.role === 'consumer');
const providerIndex = new Map<string, StoredContract[]>();
for (const p of providers) {
const key = normalizeContractId(p.contractId);
const list = providerIndex.get(key) || [];
list.push(p);
providerIndex.set(key, list);
}
const matched: CrossLink[] = [];
const matchedConsumerIds = new Set<string>();
const matchedProviderIds = new Set<string>();
for (const consumer of consumers) {
const matchingKeys = findMatchingKeys(consumer.contractId, providerIndex);
if (matchingKeys.length === 0) continue;
const allMatchingProviders = matchingKeys.flatMap((k) => providerIndex.get(k) || []);
for (const provider of allMatchingProviders) {
if (provider.repo === consumer.repo) {
if (!provider.service || !consumer.service || provider.service === consumer.service) {
continue;
}
}
matched.push({
from: {
repo: consumer.repo,
service: consumer.service,
symbolUid: consumer.symbolUid,
symbolRef: consumer.symbolRef,
},
to: {
repo: provider.repo,
service: provider.service,
symbolUid: provider.symbolUid,
symbolRef: provider.symbolRef,
},
type: consumer.type,
contractId: consumer.contractId,
matchType: 'exact',
confidence: 1.0,
});
matchedConsumerIds.add(`${consumer.repo}::${consumer.contractId}`);
matchedProviderIds.add(`${provider.repo}::${provider.contractId}`);
}
}
const unmatched = contracts.filter((c) => {
const id = `${c.repo}::${c.contractId}`;
return c.role === 'provider' ? !matchedProviderIds.has(id) : !matchedConsumerIds.has(id);
});
return { matched, unmatched };
}
@@ -0,0 +1,177 @@
import fs from 'node:fs/promises';
import path from 'node:path';
export interface ServiceBoundary {
servicePath: string;
serviceName: string;
markers: string[];
confidence: number;
}
const SERVICE_MARKERS = [
'package.json',
'go.mod',
'Dockerfile',
'pom.xml',
'build.gradle',
'build.gradle.kts',
'Cargo.toml',
'pyproject.toml',
'requirements.txt',
'mix.exs',
] as const;
const SOURCE_EXTENSIONS = new Set([
'.ts',
'.tsx',
'.js',
'.jsx',
'.mjs',
'.cjs',
'.go',
'.java',
'.kt',
'.kts',
'.py',
'.pyi',
'.rs',
'.c',
'.cpp',
'.h',
'.hpp',
'.cs',
'.rb',
'.php',
'.swift',
'.dart',
'.ex',
'.exs',
'.erl',
'.proto',
]);
const EXCLUDED_DIRS = new Set([
'node_modules',
'vendor',
'target',
'build',
'dist',
'__pycache__',
'.venv',
'venv',
'.tox',
'.mypy_cache',
'.gradle',
'.mvn',
'out',
'bin',
]);
export async function detectServiceBoundaries(repoPath: string): Promise<ServiceBoundary[]> {
const boundaries: ServiceBoundary[] = [];
await walkForBoundaries(repoPath, repoPath, boundaries);
return boundaries;
}
async function walkForBoundaries(
dir: string,
repoRoot: string,
results: ServiceBoundary[],
): Promise<void> {
let entries: import('node:fs').Dirent[];
try {
entries = await fs.readdir(dir, { withFileTypes: true });
} catch {
return;
}
const isRoot = path.resolve(dir) === path.resolve(repoRoot);
const foundMarkers: string[] = [];
let hasSourceFiles = false;
const subdirs: string[] = [];
for (const entry of entries) {
if (entry.name.startsWith('.') || EXCLUDED_DIRS.has(entry.name)) continue;
if (entry.isDirectory()) {
subdirs.push(path.join(dir, entry.name));
} else if (entry.isFile()) {
if (SERVICE_MARKERS.includes(entry.name as (typeof SERVICE_MARKERS)[number])) {
foundMarkers.push(entry.name);
}
const ext = path.extname(entry.name).toLowerCase();
if (SOURCE_EXTENSIONS.has(ext)) {
hasSourceFiles = true;
}
}
}
// Check subdirectories for source files if not found at this level
if (!hasSourceFiles && foundMarkers.length > 0) {
hasSourceFiles = await hasSourceFilesInSubdirs(subdirs);
}
if (!isRoot && foundMarkers.length >= 1 && hasSourceFiles) {
const relativePath = path.relative(repoRoot, dir).replace(/\\/g, '/');
const serviceName = path.basename(dir);
const confidence = computeConfidence(foundMarkers.length);
results.push({
servicePath: relativePath,
serviceName,
markers: foundMarkers,
confidence,
});
}
// Recurse into subdirectories
for (const subdir of subdirs) {
await walkForBoundaries(subdir, repoRoot, results);
}
}
async function hasSourceFilesInSubdirs(subdirs: string[]): Promise<boolean> {
for (const subdir of subdirs) {
let entries: import('node:fs').Dirent[];
try {
entries = await fs.readdir(subdir, { withFileTypes: true });
} catch {
continue;
}
for (const entry of entries) {
if (entry.isFile()) {
const ext = path.extname(entry.name).toLowerCase();
if (SOURCE_EXTENSIONS.has(ext)) return true;
}
if (entry.isDirectory() && !entry.name.startsWith('.') && !EXCLUDED_DIRS.has(entry.name)) {
const deeper = await hasSourceFilesInSubdirs([path.join(subdir, entry.name)]);
if (deeper) return true;
}
}
}
return false;
}
function computeConfidence(markerCount: number): number {
if (markerCount >= 3) return 1.0;
if (markerCount === 2) return 0.9;
return 0.75;
}
export function assignService(filePath: string, boundaries: ServiceBoundary[]): string | undefined {
const normalized = filePath.replace(/\\/g, '/');
let bestMatch: ServiceBoundary | undefined;
let bestLength = 0;
for (const boundary of boundaries) {
const prefix = boundary.servicePath + '/';
if (normalized.startsWith(prefix) && boundary.servicePath.length > bestLength) {
bestMatch = boundary;
bestLength = boundary.servicePath.length;
}
}
return bestMatch?.servicePath;
}
+223
View File
@@ -0,0 +1,223 @@
/**
* Group orchestration shared by MCP (LocalBackend) and CLI.
* DB access is injected via GroupToolPort so this module stays free of LocalBackend private API.
*/
import { checkStaleness } from '../git-staleness.js';
import { loadGroupConfig } from './config-parser.js';
import { getDefaultGitnexusDir, getGroupDir, listGroups, readContractRegistry } from './storage.js';
import { syncGroup } from './sync.js';
export interface GroupRepoHandle {
id: string;
name: string;
repoPath: string;
storagePath: string;
indexedAt?: string;
lastCommit?: string;
}
export interface GroupToolPort {
resolveRepo(repoParam?: string): Promise<GroupRepoHandle>;
impact(
repo: GroupRepoHandle,
params: {
target: string;
direction: 'upstream' | 'downstream';
maxDepth?: number;
relationTypes?: string[];
includeTests?: boolean;
minConfidence?: number;
},
): Promise<unknown>;
query(
repo: GroupRepoHandle,
params: {
query: string;
task_context?: string;
goal?: string;
limit?: number;
max_symbols?: number;
include_content?: boolean;
},
): Promise<unknown>;
impactByUid(
repoId: string,
uid: string,
direction: string,
opts: {
maxDepth: number;
relationTypes: string[];
minConfidence: number;
includeTests: boolean;
},
): Promise<unknown | null>;
}
function repoInSubgroup(repoPath: string, subgroup?: string): boolean {
if (!subgroup?.trim()) return true;
const s = subgroup.replace(/\/+$/, '');
return repoPath === s || repoPath.startsWith(`${s}/`);
}
export class GroupService {
constructor(private readonly port: GroupToolPort) {}
async groupList(params: Record<string, unknown>): Promise<unknown> {
const name = typeof params.name === 'string' ? params.name.trim() : '';
if (!name) {
const groups = await listGroups();
return { groups };
}
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const config = await loadGroupConfig(groupDir);
return {
name: config.name,
description: config.description,
repos: config.repos,
links: config.links,
};
}
async groupSync(params: Record<string, unknown>): Promise<unknown> {
const name = String(params.name ?? '').trim();
if (!name) return { error: 'name is required' };
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const config = await loadGroupConfig(groupDir);
const result = await syncGroup(config, {
groupDir,
exactOnly: Boolean(params.exactOnly),
skipEmbeddings: Boolean(params.skipEmbeddings),
allowStale: Boolean(params.allowStale),
verbose: Boolean(params.verbose),
});
return {
contracts: result.contracts.length,
crossLinks: result.crossLinks.length,
unmatched: result.unmatched.length,
missingRepos: result.missingRepos,
};
}
async groupContracts(params: Record<string, unknown>): Promise<unknown> {
const name = String(params.name ?? '').trim();
if (!name) return { error: 'name is required' };
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const registry = await readContractRegistry(groupDir);
if (!registry) {
return { error: `No contracts.json for group "${name}". Run group_sync first.` };
}
let contracts = registry.contracts;
if (params.type) contracts = contracts.filter((c) => c.type === params.type);
if (params.repo) contracts = contracts.filter((c) => c.repo === params.repo);
if (params.unmatchedOnly) {
const matchedIds = new Set(
registry.crossLinks.flatMap((l) => [
`${l.from.repo}::${l.contractId}`,
`${l.to.repo}::${l.contractId}`,
]),
);
contracts = contracts.filter((c) => !matchedIds.has(`${c.repo}::${c.contractId}`));
}
return { contracts, crossLinks: registry.crossLinks };
}
async groupQuery(params: Record<string, unknown>): Promise<unknown> {
const name = String(params.name ?? '').trim();
const queryText = String(params.query ?? '').trim();
if (!name || !queryText) return { error: 'name and query are required' };
const limit = typeof params.limit === 'number' && params.limit > 0 ? params.limit : 5;
const subgroup = typeof params.subgroup === 'string' ? params.subgroup : undefined;
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const config = await loadGroupConfig(groupDir);
const perRepo: Array<{ repo: string; score: number; processes: unknown[] }> = [];
for (const [repoPath, registryName] of Object.entries(config.repos)) {
if (!repoInSubgroup(repoPath, subgroup)) continue;
try {
const repoObj = await this.port.resolveRepo(registryName);
const queryResult = (await this.port.query(repoObj, {
query: queryText,
limit,
max_symbols: 10,
include_content: false,
})) as { processes?: Array<Record<string, unknown>> };
const processes = queryResult.processes || [];
const scored = processes.map((p, idx) => ({
...p,
_rrf_score: 1 / (idx + 1 + 60),
_repo: repoPath,
}));
perRepo.push({ repo: repoPath, score: 0, processes: scored });
} catch {
perRepo.push({ repo: repoPath, score: 0, processes: [] });
}
}
const allProcesses = perRepo.flatMap((r) => r.processes as Array<Record<string, unknown>>);
allProcesses.sort((a, b) => (b._rrf_score as number) - (a._rrf_score as number));
const topN = allProcesses.slice(0, limit);
return {
group: name,
query: queryText,
results: topN,
per_repo: perRepo.map((r) => ({ repo: r.repo, count: r.processes.length })),
};
}
async groupStatus(params: Record<string, unknown>): Promise<unknown> {
const name = String(params.name ?? '').trim();
if (!name) return { error: 'name is required' };
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const config = await loadGroupConfig(groupDir);
const registry = await readContractRegistry(groupDir);
const repoStatuses: Record<
string,
{
indexStale: boolean;
contractsStale: boolean;
missing: boolean;
commitsBehind?: number;
}
> = {};
const fsp = await import('node:fs/promises');
const pathMod = await import('node:path');
for (const [repoPath, registryName] of Object.entries(config.repos)) {
try {
const repoObj = await this.port.resolveRepo(registryName);
const metaPath = pathMod.join(repoObj.storagePath, 'meta.json');
const metaRaw = await fsp.readFile(metaPath, 'utf-8').catch(() => '{}');
const meta = JSON.parse(metaRaw) as { lastCommit?: string; indexedAt?: string };
const staleness = meta.lastCommit
? checkStaleness(repoObj.repoPath, meta.lastCommit)
: { isStale: true, commitsBehind: -1 };
const snapshot = registry?.repoSnapshots[repoPath];
const contractsStale =
snapshot && meta.indexedAt ? snapshot.indexedAt !== meta.indexedAt : !snapshot;
repoStatuses[repoPath] = {
indexStale: staleness.isStale,
contractsStale: Boolean(contractsStale),
missing: false,
commitsBehind: staleness.commitsBehind,
};
} catch {
repoStatuses[repoPath] = { indexStale: false, contractsStale: false, missing: true };
}
}
return {
group: name,
lastSync: registry?.generatedAt || null,
missingRepos: registry?.missingRepos || [],
repos: repoStatuses,
};
}
}
+109
View File
@@ -0,0 +1,109 @@
import * as fs from 'node:fs';
import * as fsp from 'node:fs/promises';
import * as path from 'node:path';
import * as os from 'node:os';
import type { ContractRegistry } from './types.js';
const CONTRACTS_FILE = 'contracts.json';
export function getDefaultGitnexusDir(): string {
return process.env.GITNEXUS_HOME || path.join(os.homedir(), '.gitnexus');
}
export function getGroupsBaseDir(gitnexusDir?: string): string {
return path.join(gitnexusDir || getDefaultGitnexusDir(), 'groups');
}
const GROUP_NAME_RE = /^[a-zA-Z0-9][a-zA-Z0-9_-]*$/;
export function validateGroupName(name: string): void {
if (!GROUP_NAME_RE.test(name)) {
throw new Error(
`Invalid group name "${name}". Names must start with a letter or digit and contain only [a-zA-Z0-9_-].`,
);
}
}
export function getGroupDir(gitnexusDir: string, groupName: string): string {
validateGroupName(groupName);
return path.join(gitnexusDir, 'groups', groupName);
}
export async function writeContractRegistry(
groupDir: string,
registry: ContractRegistry,
): Promise<void> {
const targetPath = path.join(groupDir, CONTRACTS_FILE);
const tmpPath = `${targetPath}.tmp.${Date.now()}`;
await fsp.writeFile(tmpPath, JSON.stringify(registry, null, 2), 'utf-8');
await fsp.rename(tmpPath, targetPath);
}
export async function readContractRegistry(groupDir: string): Promise<ContractRegistry | null> {
const filePath = path.join(groupDir, CONTRACTS_FILE);
try {
const content = await fsp.readFile(filePath, 'utf-8');
return JSON.parse(content) as ContractRegistry;
} catch (err: unknown) {
if ((err as NodeJS.ErrnoException).code === 'ENOENT') return null;
throw err;
}
}
export async function listGroups(gitnexusDir?: string): Promise<string[]> {
const groupsDir = getGroupsBaseDir(gitnexusDir);
try {
const entries = await fsp.readdir(groupsDir, { withFileTypes: true });
const names: string[] = [];
for (const entry of entries) {
if (entry.isDirectory()) {
const yamlPath = path.join(groupsDir, entry.name, 'group.yaml');
if (fs.existsSync(yamlPath)) {
names.push(entry.name);
}
}
}
return names;
} catch (err: unknown) {
if ((err as NodeJS.ErrnoException).code === 'ENOENT') return [];
throw err;
}
}
export async function createGroupDir(
gitnexusDir: string,
groupName: string,
force: boolean = false,
): Promise<string> {
const groupDir = getGroupDir(gitnexusDir, groupName);
if (fs.existsSync(path.join(groupDir, 'group.yaml')) && !force) {
throw new Error(`Group "${groupName}" already exists. Use --force to overwrite.`);
}
await fsp.mkdir(groupDir, { recursive: true });
const template = `version: 1
name: ${groupName}
description: ""
repos: {}
links: []
packages: {}
detect:
http: true
grpc: true
topics: true
shared_libs: true
embedding_fallback: true
matching:
bm25_threshold: 0.7
embedding_threshold: 0.65
max_candidates_per_step: 3
`;
await fsp.writeFile(path.join(groupDir, 'group.yaml'), template, 'utf-8');
return groupDir;
}
+185
View File
@@ -0,0 +1,185 @@
import fs from 'node:fs/promises';
import path from 'node:path';
import { Buffer } from 'node:buffer';
import { initLbug, closeLbug, executeParameterized } from '../lbug/pool-adapter.js';
import { readRegistry, type RegistryEntry } from '../../storage/repo-manager.js';
import type { GroupConfig, RepoHandle, RepoSnapshot, StoredContract, CrossLink } from './types.js';
import { HttpRouteExtractor } from './extractors/http-route-extractor.js';
import { GrpcExtractor } from './extractors/grpc-extractor.js';
import { TopicExtractor } from './extractors/topic-extractor.js';
import { runExactMatch } from './matching.js';
import { detectServiceBoundaries, assignService } from './service-boundary-detector.js';
import type { CypherExecutor } from './contract-extractor.js';
import { writeContractRegistry } from './storage.js';
import type { ContractRegistry } from './types.js';
export interface SyncOptions {
extractorOverride?:
| ((repo: RepoHandle) => Promise<StoredContract[]>)
| (() => Promise<StoredContract[]>);
resolveRepoHandle?: (registryName: string, groupPath: string) => Promise<RepoHandle | null>;
skipWrite?: boolean;
groupDir?: string;
allowStale?: boolean;
verbose?: boolean;
exactOnly?: boolean;
skipEmbeddings?: boolean;
}
export interface SyncResult {
contracts: StoredContract[];
crossLinks: CrossLink[];
unmatched: StoredContract[];
missingRepos: string[];
repoSnapshots: Record<string, RepoSnapshot>;
}
export function stableRepoPoolId(entry: RegistryEntry, allEntries: RegistryEntry[]): string {
const base = entry.name.toLowerCase();
const resolved = path.resolve(entry.path);
for (const other of allEntries) {
if (other.name.toLowerCase() === base && path.resolve(other.path) !== resolved) {
const hash = Buffer.from(entry.path).toString('base64url').slice(0, 6);
return `${base}-${hash}`;
}
}
return base;
}
function defaultResolveHandle(allEntries: RegistryEntry[]) {
return async (registryName: string, groupPath: string): Promise<RepoHandle | null> => {
const e = allEntries.find((en) => en.name === registryName);
if (!e) return null;
const poolId = stableRepoPoolId(e, allEntries);
return {
id: poolId,
path: groupPath,
repoPath: e.path,
storagePath: e.storagePath,
};
};
}
export async function syncGroup(config: GroupConfig, opts?: SyncOptions): Promise<SyncResult> {
const missingRepos: string[] = [];
const repoSnapshots: Record<string, RepoSnapshot> = {};
let autoContracts: StoredContract[] = [];
let dbExecutors: Map<string, CypherExecutor> | undefined;
const eo = opts?.extractorOverride;
if (eo && eo.length === 0) {
autoContracts = await (eo as () => Promise<StoredContract[]>)();
} else {
const entries = await readRegistry();
const resolve = opts?.resolveRepoHandle ?? defaultResolveHandle(entries);
const httpEx = new HttpRouteExtractor();
const grpcEx = new GrpcExtractor();
const topicEx = new TopicExtractor();
dbExecutors = new Map<string, CypherExecutor>();
const openPoolIds: string[] = [];
try {
for (const [groupPath, regName] of Object.entries(config.repos)) {
const handle = await resolve(regName, groupPath);
if (!handle) {
missingRepos.push(groupPath);
continue;
}
const poolId = handle.id;
const lbugPath = path.join(handle.storagePath, 'lbug');
try {
await initLbug(poolId, lbugPath);
openPoolIds.push(poolId);
const executor: CypherExecutor = (query, params) =>
executeParameterized(poolId, query, params ?? {});
dbExecutors.set(groupPath, executor);
const boundaries = await detectServiceBoundaries(handle.repoPath);
if (config.detect.http) {
const extracted = await httpEx.extract(executor, handle.repoPath, handle);
for (const c of extracted) {
autoContracts.push({
...c,
repo: groupPath,
service: assignService(c.symbolRef.filePath, boundaries),
});
}
}
if (config.detect.grpc) {
const extracted = await grpcEx.extract(executor, handle.repoPath, handle);
for (const c of extracted) {
autoContracts.push({
...c,
repo: groupPath,
service: assignService(c.symbolRef.filePath, boundaries),
});
}
}
if (config.detect.topics) {
const extracted = await topicEx.extract(executor, handle.repoPath, handle);
for (const c of extracted) {
autoContracts.push({
...c,
repo: groupPath,
service: assignService(c.symbolRef.filePath, boundaries),
});
}
}
const metaPath = path.join(handle.storagePath, 'meta.json');
try {
const raw = await fs.readFile(metaPath, 'utf-8');
const m = JSON.parse(raw) as { indexedAt?: string; lastCommit?: string };
repoSnapshots[groupPath] = {
indexedAt: m.indexedAt || '',
lastCommit: m.lastCommit || '',
};
} catch {
const e = entries.find((en) => en.name === regName);
repoSnapshots[groupPath] = {
indexedAt: e?.indexedAt || '',
lastCommit: e?.lastCommit || '',
};
}
} catch {
missingRepos.push(groupPath);
}
}
} finally {
for (const id of [...new Set(openPoolIds)]) {
await closeLbug(id).catch(() => {});
}
}
}
const { matched, unmatched } = runExactMatch(autoContracts);
const crossLinks: CrossLink[] = matched;
const allContracts: StoredContract[] = autoContracts;
const registry: ContractRegistry = {
version: 1,
generatedAt: new Date().toISOString(),
repoSnapshots,
missingRepos,
contracts: allContracts,
crossLinks,
};
if (opts?.groupDir && !opts.skipWrite) {
await writeContractRegistry(opts.groupDir, registry);
}
return {
contracts: allContracts,
crossLinks,
unmatched,
missingRepos,
repoSnapshots,
};
}
+133
View File
@@ -0,0 +1,133 @@
export type ContractType = 'http' | 'grpc' | 'topic' | 'lib' | 'custom';
export type MatchType = 'exact' | 'manifest' | 'bm25' | 'embedding';
export type ContractRole = 'provider' | 'consumer';
export interface GroupConfig {
version: number;
name: string;
description: string;
repos: Record<string, string>;
links: GroupManifestLink[];
packages: Record<string, Record<string, string>>;
detect: DetectConfig;
matching: MatchingConfig;
}
export interface GroupManifestLink {
from: string;
to: string;
type: ContractType;
contract: string;
role: ContractRole;
}
export interface DetectConfig {
http: boolean;
grpc: boolean;
topics: boolean;
shared_libs: boolean;
embedding_fallback: boolean;
}
export interface MatchingConfig {
bm25_threshold: number;
embedding_threshold: number;
max_candidates_per_step: number;
}
export interface SymbolRef {
filePath: string;
name: string;
}
export interface ExtractedContract {
contractId: string;
type: ContractType;
role: ContractRole;
symbolUid: string;
symbolRef: SymbolRef;
symbolName: string;
confidence: number;
meta: Record<string, unknown>;
/** Service boundary within a monorepo (relative path from repo root, e.g. "services/auth"). */
service?: string;
}
export interface CrossLinkEndpoint {
repo: string;
/** Service boundary within a monorepo (relative path from repo root). */
service?: string;
symbolUid: string;
symbolRef: SymbolRef;
}
export interface CrossLink {
from: CrossLinkEndpoint;
to: CrossLinkEndpoint;
type: ContractType;
contractId: string;
matchType: MatchType;
confidence: number;
}
export interface RepoSnapshot {
indexedAt: string;
lastCommit: string;
}
export interface ContractRegistry {
version: number;
generatedAt: string;
repoSnapshots: Record<string, RepoSnapshot>;
missingRepos: string[];
contracts: StoredContract[];
crossLinks: CrossLink[];
}
export interface StoredContract extends ExtractedContract {
repo: string;
}
/** Repo within a group (group path + paths; name collision with MCP RepoHandle — import from group/types only). */
export interface RepoHandle {
id: string;
path: string;
repoPath: string;
storagePath: string;
}
export interface GroupImpactResult {
local: unknown;
group: string;
cross: CrossRepoImpact[];
outOfScope: OutOfScopeLink[];
truncated: boolean;
truncatedRepos: string[];
summary: {
direct: number;
processes_affected: number;
modules_affected: number;
cross_repo_hits: number;
};
risk: string;
}
export interface CrossRepoImpact {
repo: string;
repo_path: string;
contract: {
id: string;
type: ContractType;
match_type: MatchType;
confidence: number;
};
by_depth: Record<string, unknown[]>;
affected_processes: string[];
}
export interface OutOfScopeLink {
from: string;
to: string;
contractId: string;
confidence: number;
}
@@ -0,0 +1,368 @@
/**
* BindingAccumulator — read-append-only accumulator that collects TypeEnv
* bindings across files in the GitNexus analyzer pipeline.
*
* **Current behavior (both execution paths):** The accumulator carries only
* file-scope (`scope = ''`) entries. Function-scope bindings are stripped
* at both write sites:
*
* - **Worker path**: `parse-worker.ts` serializes only
* `typeEnv.fileScope()` entries across the IPC boundary.
* - **Sequential path**: `type-env.ts::flush()` iterates only the FILE_SCOPE
* entry of the env map and writes `BindingEntry` records with
* `scope: ''` hardcoded.
*
* The narrowing exists because function-scope bindings have zero downstream
* consumers today and were previously costing ~4.9 MB of heap + IPC on
* every pipeline run. See `type-env.ts::flush()` and the `FileScopeBindings`
* JSDoc in `parse-worker.ts` for the paired Phase 9 reversion checklist.
*
* **Historical quality asymmetry (Phase 9 consideration):** Even though
* both paths now carry only file-scope data, the two paths were built
* under different resolution capabilities, and a future Phase 9 reverter
* that widens them back to all scopes will inherit that asymmetry:
*
* - **Sequential path** had (and would regain) access to the full
* `SymbolTable` and `importedBindings`, so its bindings benefit from
* Tier 2 cross-file propagation.
* - **Worker path** runs without `SymbolTable` / `importedBindings` and
* can only produce Tier 0 (annotation-declared) and local Tier 1
* (same-file constructor inference) bindings.
*
* Phase 9 consumers that trust every entry equally will silently produce
* worse results for large repos (worker-dominant) than small ones
* (sequential-dominant). If Phase 9 needs homogeneous quality, either
* (a) tag entries with their tier at insert time so consumers can filter,
* or (b) post-process worker-path entries through a follow-up resolution
* pass after the main-thread `SymbolTable` is complete.
*
* **Lifecycle contract**: `append → finalize → consume → dispose`. See
* `finalize()` and `dispose()` for the state machine. Disposal is
* orthogonal to finalization: either order is legal.
*/
export interface BindingEntry {
readonly scope: string; // '' for file-level, 'funcName@startIndex' for function-local
readonly varName: string;
readonly typeName: string;
}
/**
* Minimal graph-node shape required by `enrichExportedTypeMap()`. Intentionally
* narrower than the full `GraphNode` type in `graph/types.ts` so tests can
* construct a minimal mock without depending on the full graph module, and
* so the enrichment logic is a pure function over this contract.
*
* Matches the shape of the real `KnowledgeGraph` node's `properties.isExported`
* access path — tests that use a different shape silently pass while
* production fails.
*/
export interface EnrichmentGraphNode {
readonly id: string;
readonly properties?: { readonly isExported?: boolean } | undefined;
}
/**
* Minimal graph lookup interface used by `enrichExportedTypeMap()`.
* Consumes only the method the enrichment loop actually calls.
*/
export interface EnrichmentGraphLookup {
getNode(id: string): EnrichmentGraphNode | undefined;
}
/**
* Merge file-scope bindings from a (finalized) `BindingAccumulator` into an
* `exportedTypeMap` for symbols whose graph nodes are marked as exported.
*
* This is the single source of truth for the worker-path ExportedTypeMap
* enrichment loop. Previously the logic lived inline in `pipeline.ts` and
* the test suite reimplemented it as a `runEnrichmentLoop` helper — a
* drift-prone pattern that meant tests could pass while production regressed.
* Extracting it here makes the production code call the same function the
* tests call.
*
* **Node ID candidate order**: `Function:{filePath}:{name}` →
* `Variable:{filePath}:{name}` → `Const:{filePath}:{name}`. First match wins.
*
* **Tier 0 priority**: if `exportedTypeMap` already has an entry for a
* `(filePath, name)` pair, the accumulator entry does NOT overwrite it —
* the SymbolTable tier-0 pass is authoritative. Without this guard, a
* worker-path binding could clobber a higher-quality type from SymbolTable.
*
* **Finalize precondition**: the accumulator should be finalized before
* calling this function. The lifecycle contract is
* `append → finalize → enrich → dispose`. Finalization is not asserted
* here (the test suite and pipeline both honor it separately), but any
* append happening concurrently with this enrichment would be a lifecycle
* bug at the caller level.
*
* @returns The number of new entries written into `exportedTypeMap`
* (0 on empty accumulator or when every candidate was filtered
* out by the export check or the Tier 0 guard).
*/
export function enrichExportedTypeMap(
bindingAccumulator: BindingAccumulator,
graph: EnrichmentGraphLookup,
exportedTypeMap: Map<string, Map<string, string>>,
): number {
if (bindingAccumulator.fileCount === 0) return 0;
let enriched = 0;
for (const filePath of bindingAccumulator.files()) {
for (const [name, type] of bindingAccumulator.fileScopeEntries(filePath)) {
// Three-candidate-ID lookup mirrors the sequential-path export check
// in `collectExportedBindings()` (call-processor.ts).
const functionNodeId = `Function:${filePath}:${name}`;
const variableNodeId = `Variable:${filePath}:${name}`;
const constNodeId = `Const:${filePath}:${name}`;
const node =
graph.getNode(functionNodeId) ??
graph.getNode(variableNodeId) ??
graph.getNode(constNodeId);
if (!node?.properties?.isExported) continue;
let fileExports = exportedTypeMap.get(filePath);
if (!fileExports) {
fileExports = new Map();
exportedTypeMap.set(filePath, fileExports);
}
// Tier 0 priority: SymbolTable-populated entries are authoritative.
if (!fileExports.has(name)) {
fileExports.set(name, type);
enriched++;
}
}
}
return enriched;
}
const ENTRY_OVERHEAD = 64; // bytes per entry (object overhead + property refs)
const MAP_ENTRY_OVERHEAD = 80; // bytes per file entry in the map
export class BindingAccumulator {
// Storage is split into two parallel maps so fileScopeEntries() is
// O(n_file_scope) instead of O(n_total).
// - _allByFile holds every BindingEntry (used by getFile, memory estimate).
// - _fileScopeByFile caches the flat [varName, typeName] view of the
// `scope === ''` subset, populated at insert time so reads are O(1) map
// lookup + O(n_file_scope) array return. Both maps carry the same key
// set modulo the `scope === ''` precondition: _allByFile has a key as
// soon as any entry is appended; _fileScopeByFile only has a key once a
// file-scope entry arrives. Code that iterates via files() uses
// _allByFile so files with only function-scope entries remain visible.
private readonly _allByFile = new Map<string, BindingEntry[]>();
private readonly _fileScopeByFile = new Map<string, [string, string][]>();
private _totalBindings = 0;
private _finalized = false;
private _disposed = false;
/**
* Append bindings for a file. Safe to call multiple times for the same file.
* Throws if the accumulator has been finalized. Skips if entries is empty.
*
* The `entries` parameter is `readonly` — this method never mutates the
* caller's array. Internally, the first `appendFile` call per filePath
* makes a defensive copy (`slice()`), and subsequent calls push into the
* accumulator's own storage.
*/
appendFile(filePath: string, entries: readonly BindingEntry[]): void {
if (this._finalized) {
throw new Error(
'[BindingAccumulator] appendFile after finalize — no further appends allowed',
);
}
if (entries.length === 0) {
return;
}
// Contract consistency: if this accumulator was previously disposed
// without being finalized, `dispose()` is documented to leave it
// "behaving like a fresh one" for subsequent appends. Clear the
// `_disposed` flag here so the `disposed` getter tracks the actual
// live state, not a stale signal from the prior lifecycle cycle.
if (this._disposed) {
this._disposed = false;
}
// Note on the file-scope-only invariant:
// The accumulator does NOT reject function-scope entries at this
// boundary. The narrowing contract is enforced by the two production
// write sites — `parse-worker.ts` (which uses `typeEnv.fileScope()`
// and hardcodes `scope: ''` in the pipeline adapter) and
// `type-env.ts::flush()` (which iterates only `env.get(FILE_SCOPE)`).
// The class JSDoc documents the invariant and the Phase 9 reversion
// path. Making `appendFile` runtime-reject non-file-scope entries
// would break the accumulator's own storage-split tests which
// legitimately exercise mixed-scope entries. If a future write path
// violates the invariant, tests should fail via missing exports in
// the enrichment loop, not via an assertion here.
// All-scope store.
const existingAll = this._allByFile.get(filePath);
if (existingAll !== undefined) {
for (const e of entries) {
existingAll.push(e);
}
} else {
this._allByFile.set(filePath, entries.slice());
}
// File-scope fast-path store. Populated lazily on first file-scope entry.
let existingFileScope = this._fileScopeByFile.get(filePath);
for (const e of entries) {
if (e.scope === '') {
if (existingFileScope === undefined) {
existingFileScope = [];
this._fileScopeByFile.set(filePath, existingFileScope);
}
existingFileScope.push([e.varName, e.typeName]);
}
}
this._totalBindings += entries.length;
}
/** Lock the accumulator — no further appends. Idempotent. */
finalize(): void {
// Dev-mode invariant: verify the parallel storage split is consistent.
// `_fileScopeByFile` must be a proper projection of `_allByFile`
// where the outer key is a subset and the inner entries are exactly
// the `scope === ''` subset of `_allByFile[key]`. A drift would
// indicate a bug in `appendFile()` where one map was updated but
// not the other.
if (process.env.NODE_ENV !== 'production' && !this._finalized) {
for (const [filePath, fileScopeTuples] of this._fileScopeByFile) {
const allEntries = this._allByFile.get(filePath);
if (allEntries === undefined) {
throw new Error(
`[BindingAccumulator] storage split drift: file ${filePath} has file-scope entries ` +
`but no _allByFile entry`,
);
}
const projectedCount = allEntries.filter((e) => e.scope === '').length;
if (projectedCount !== fileScopeTuples.length) {
throw new Error(
`[BindingAccumulator] storage split drift: file ${filePath} has ` +
`${fileScopeTuples.length} file-scope tuples but ${projectedCount} file-scope ` +
`entries in _allByFile`,
);
}
}
}
this._finalized = true;
}
/**
* Release the accumulator's heap footprint. Clears both internal storage
* maps and resets `_totalBindings` to zero. Idempotent and orthogonal to
* `finalize()` — calling `dispose()` does not change the finalized state.
*
* Post-dispose contract: all read methods return empty/undefined state
* matching a never-appended-to accumulator. Specifically:
* - `fileCount === 0`
* - `totalBindings === 0`
* - `files()` yields an empty iterator
* - `getFile(x)` returns `undefined` for all `x`
* - `fileScopeEntries(x)` returns `[]` for all `x`
* - `estimateMemoryBytes()` returns `0`
*
* If `dispose()` is called **before** `finalize()`, subsequent `appendFile`
* calls succeed — the accumulator behaves like a fresh one. If called
* **after** `finalize()`, subsequent `appendFile` calls throw the existing
* "finalized" error.
*
* Lifecycle note: the pipeline disposes the accumulator after the
* ExportedTypeMap enrichment loop consumes its file-scope entries, so
* the heap is released before Phase 14 (`runCrossFileBindingPropagation`)
* and `runGraphAnalysisPhases` begin their long-running work. When Phase 9
* wires a consumer into that stage, the dispose call should move later in
* the pipeline or be removed entirely.
*/
dispose(): void {
this._allByFile.clear();
this._fileScopeByFile.clear();
this._totalBindings = 0;
this._disposed = true;
}
/** Get all bindings for a file, or undefined if the file is unknown. */
getFile(filePath: string): readonly BindingEntry[] | undefined {
return this._allByFile.get(filePath);
}
/**
* Get only scope='' (file-level) entries as [varName, typeName] tuples.
* Backward-compatible with the old workerTypeEnvBindings pattern.
* Returns an empty array for an unknown file.
*
* O(1) map lookup + O(n_file_scope) defensive-copy construction — does
* NOT walk function-scope entries. See the `_fileScopeByFile` field
* comment for the storage split rationale.
*
* The return value is a shallow copy; mutating it does not affect
* subsequent reads or internal state. This encapsulation guard prevents
* a Phase 9 consumer from accidentally corrupting the accumulator via
* `acc.fileScopeEntries(p).push(...)` or similar.
*/
fileScopeEntries(filePath: string): readonly (readonly [string, string])[] {
const cached = this._fileScopeByFile.get(filePath);
return cached ? cached.slice() : [];
}
/** Iterate over all file paths in insertion order. */
files(): IterableIterator<string> {
return this._allByFile.keys();
}
/** Number of distinct files with at least one binding. */
get fileCount(): number {
return this._allByFile.size;
}
/** Total number of binding entries across all files. */
get totalBindings(): number {
return this._totalBindings;
}
/** Whether the accumulator has been finalized. */
get finalized(): boolean {
return this._finalized;
}
/**
* Whether the accumulator has been disposed. Exposed for symmetry with
* `finalized` so debug tooling and future Phase 9 consumers can detect a
* disposed accumulator without inspecting empty state heuristically.
*
* Disposal and finalization are orthogonal: a disposed accumulator may or
* may not be finalized, and vice versa. See `dispose()` for the full
* lifecycle contract.
*/
get disposed(): boolean {
return this._disposed;
}
/**
* Rough memory estimate in bytes (intentionally pessimistic).
* Formula: sum of (ENTRY_OVERHEAD + char bytes of scope+varName+typeName) per entry
* + MAP_ENTRY_OVERHEAD + char bytes of filePath per file.
*
* Note: V8 stores all-ASCII strings as Latin-1 (1 byte/char) and only upgrades
* to UCS-2 (2 bytes/char) for non-Latin-1 code points. Source paths and type names
* are typically all-ASCII, so actual heap cost is roughly half what this returns.
* The pessimistic factor is intentional — better to over-budget than under-budget.
*
* **⚠ Cost profile**: O(totalBindings) — iterates every entry in
* `_allByFile` and reads three string `.length` properties per entry.
* At a typical repo scale (10k files × ~20 file-scope bindings) this is
* ~200k property reads per call. Call at most once per pipeline run,
* NOT per file, per chunk, or per progress tick. The current single
* call site is the dev-mode telemetry log at the pipeline finalize
* seam. Adding a per-file-progress caller would silently make it
* quadratic in repo size.
*/
estimateMemoryBytes(): number {
let total = 0;
for (const [filePath, entries] of this._allByFile) {
total += MAP_ENTRY_OVERHEAD + filePath.length * 2;
for (const e of entries) {
total += ENTRY_OVERHEAD + (e.scope.length + e.varName.length + e.typeName.length) * 2;
}
}
return total;
}
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,177 @@
import type { SyntaxNode } from '../utils/ast-helpers.js';
import type { NodeLabel } from 'gitnexus-shared';
import type {
ClassExtractionConfig,
ClassExtractor,
ClassLikeNodeLabel,
ExtractedClassSymbol,
} from '../class-types.js';
const DEFAULT_SCOPE_NAME_NODE_TYPES = new Set([
'nested_namespace_specifier',
'scoped_identifier',
'scoped_type_identifier',
'qualified_name',
'namespace_name',
'namespace_identifier',
'package_identifier',
'type_identifier',
'identifier',
'name',
'constant',
]);
const DEFAULT_TYPE_NAME_NODE_TYPES = new Set([
'type_identifier',
'identifier',
'simple_identifier',
'namespace_identifier',
'constant',
'name',
]);
const DEFAULT_LABEL_BY_NODE_TYPE: Record<string, ClassLikeNodeLabel> = {
class_declaration: 'Class',
abstract_class_declaration: 'Class',
interface_declaration: 'Interface',
struct_declaration: 'Struct',
record_declaration: 'Record',
enum_declaration: 'Enum',
class_definition: 'Class',
struct_specifier: 'Struct',
class_specifier: 'Class',
enum_specifier: 'Enum',
struct_item: 'Struct',
enum_item: 'Enum',
class: 'Class',
object_declaration: 'Class',
companion_object: 'Class',
protocol_declaration: 'Interface',
extension_declaration: 'Class',
};
const CLASS_LIKE_LABELS = new Set<ClassLikeNodeLabel>([
'Class',
'Struct',
'Interface',
'Enum',
'Record',
]);
const normalizeQualifiedName = (value: string): string =>
value
.replace(/\s+/g, '')
.replace(/^::/, '')
.replace(/::/g, '.')
.replace(/\\/g, '.')
.replace(/\.+/g, '.')
.replace(/^\.+|\.+$/g, '');
const splitQualifiedName = (value: string): string[] => {
const normalized = normalizeQualifiedName(value);
return normalized ? normalized.split('.').filter(Boolean) : [];
};
const extractScopeSegmentsFromNode = (
scopeNode: SyntaxNode,
scopeNameNodeTypes: ReadonlySet<string>,
): string[] => {
const nameNode =
scopeNode.childForFieldName?.('name') ??
scopeNode.namedChildren?.find((child) => scopeNameNodeTypes.has(child.type));
return nameNode ? splitQualifiedName(nameNode.text) : [];
};
const extractTypeNameFromNode = (node: SyntaxNode): string | undefined => {
const nameField = node.childForFieldName?.('name');
if (nameField) return nameField.text;
const nameChild = node.namedChildren?.find((child) =>
DEFAULT_TYPE_NAME_NODE_TYPES.has(child.type),
);
return nameChild?.text;
};
const isClassLikeLabel = (label: NodeLabel | null | undefined): label is ClassLikeNodeLabel =>
label !== undefined && label !== null && CLASS_LIKE_LABELS.has(label as ClassLikeNodeLabel);
export function createClassExtractor(config: ClassExtractionConfig): ClassExtractor {
const typeDeclarationSet = new Set(config.typeDeclarationNodes);
const fileScopeSet = new Set(config.fileScopeNodeTypes ?? []);
const ancestorScopeSet = new Set(config.ancestorScopeNodeTypes ?? []);
const scopeNameNodeTypes = new Set([
...DEFAULT_SCOPE_NAME_NODE_TYPES,
...(config.scopeNameNodeTypes ?? []),
]);
const buildQualifiedName = (node: SyntaxNode, simpleName: string): string => {
let root = node;
while (root.parent) root = root.parent;
const readScopeSegments = (scopeNode: SyntaxNode): string[] =>
config.extractScopeSegments?.(scopeNode) ??
extractScopeSegmentsFromNode(scopeNode, scopeNameNodeTypes);
const fileScopeSegments: string[] = [];
for (const child of root.namedChildren ?? []) {
if (fileScopeSet.has(child.type)) {
fileScopeSegments.push(...readScopeSegments(child));
}
}
const ancestorScopes: string[][] = [];
let current = node.parent;
while (current) {
if (ancestorScopeSet.has(current.type)) {
const segments = readScopeSegments(current);
if (segments.length > 0) ancestorScopes.push(segments);
}
current = current.parent;
}
return [
...fileScopeSegments,
...ancestorScopes.reverse().flat(),
...splitQualifiedName(simpleName),
]
.filter(Boolean)
.join('.');
};
const extract = (
node: SyntaxNode,
fallback?: {
name?: string;
type?: NodeLabel | null;
},
): ExtractedClassSymbol | null => {
if (!typeDeclarationSet.has(node.type)) return null;
const name = config.extractName?.(node) ?? extractTypeNameFromNode(node) ?? fallback?.name;
const type =
config.extractType?.(node) ??
DEFAULT_LABEL_BY_NODE_TYPE[node.type] ??
(isClassLikeLabel(fallback?.type) ? fallback.type : undefined);
if (!name || !type) return null;
return {
name,
type,
qualifiedName: buildQualifiedName(node, name) || name,
};
};
return {
language: config.language,
isTypeDeclaration(node: SyntaxNode): boolean {
return typeDeclarationSet.has(node.type);
},
extract,
extractQualifiedName(node: SyntaxNode, simpleName: string): string | null {
return extract(node, { name: simpleName })?.qualifiedName ?? null;
},
};
}
@@ -0,0 +1,44 @@
import type { NodeLabel, SupportedLanguages } from 'gitnexus-shared';
import type { SyntaxNode } from './utils/ast-helpers.js';
export type ClassLikeNodeLabel = Extract<
NodeLabel,
'Class' | 'Struct' | 'Interface' | 'Enum' | 'Record'
>;
export interface ExtractedClassSymbol {
name: string;
type: ClassLikeNodeLabel;
qualifiedName: string;
}
/**
* Cross-language qualified type names are normalized to dot-separated scope
* segments:
* - file/package scope contributes leading segments when the language has one
* - lexical namespace/module/type scope contributes enclosing segments
* - the simple type name is always the trailing segment
*/
export interface ClassExtractor {
language: SupportedLanguages;
isTypeDeclaration(node: SyntaxNode): boolean;
extract(
node: SyntaxNode,
fallback?: {
name?: string;
type?: NodeLabel | null;
},
): ExtractedClassSymbol | null;
extractQualifiedName(node: SyntaxNode, simpleName: string): string | null;
}
export interface ClassExtractionConfig {
language: SupportedLanguages;
typeDeclarationNodes: string[];
fileScopeNodeTypes?: string[];
ancestorScopeNodeTypes?: string[];
scopeNameNodeTypes?: string[];
extractName?: (node: SyntaxNode) => string | undefined;
extractType?: (node: SyntaxNode) => ClassLikeNodeLabel | undefined;
extractScopeSegments?: (node: SyntaxNode) => string[] | null | undefined;
}
@@ -226,6 +226,7 @@ export const ENTRY_POINT_PATTERNS = {
/^onEvent$/, // BLoC event handler
/^mapEventToState$/, // Legacy BLoC pattern
],
[SupportedLanguages.Vue]: [], // Vue uses TypeScript queries — entry points handled via TS patterns
[SupportedLanguages.Cobol]: [], // Standalone regex processor — no tree-sitter entry points
} satisfies Record<SupportedLanguages, RegExp[]>;
@@ -15,12 +15,17 @@ import type { FieldVisibility } from '../../field-types.js';
/**
* Check whether any child of `node` (named or unnamed) has .text matching
* one of the given `keywords`.
* the given `keyword`.
*
* Skips the `name` field child to avoid false positives when a method is
* named after a contextual keyword (e.g. `abstract()` in TypeScript).
*/
export function hasKeyword(node: SyntaxNode, keyword: string): boolean {
const nameNode = node.childForFieldName('name');
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (child && child.text.trim() === keyword) return true;
if (!child || child === nameNode) continue;
if (child.text.trim() === keyword) return true;
}
return false;
}
@@ -46,6 +51,7 @@ export function hasModifier(node: SyntaxNode, modifierType: string, keyword: str
/**
* Return the first matching visibility keyword found either as a direct keyword
* child or inside a modifier wrapper node.
* Skips the `name` field child (same rationale as hasKeyword).
*/
export function findVisibility(
node: SyntaxNode,
@@ -53,10 +59,12 @@ export function findVisibility(
defaultVis: FieldVisibility,
modifierNodeType?: string,
): FieldVisibility {
const nameNode = node.childForFieldName('name');
// Direct keyword children
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
const text = child?.text.trim() as FieldVisibility | undefined;
if (!child || child === nameNode) continue;
const text = child.text.trim() as FieldVisibility | undefined;
if (text && (keywords as ReadonlySet<string>).has(text)) return text;
}
// Modifier wrapper
@@ -1,3 +1,4 @@
import { isVerboseIngestionEnabled } from './utils/verbose.js';
import fs from 'fs/promises';
import path from 'path';
import { glob } from 'glob';
@@ -43,6 +44,7 @@ export const walkRepositoryPaths = async (
const entries: ScannedFile[] = [];
let processed = 0;
let skippedLarge = 0;
const skippedLargePaths: string[] = [];
for (let start = 0; start < filtered.length; start += READ_CONCURRENCY) {
const batch = filtered.slice(start, start + READ_CONCURRENCY);
@@ -52,6 +54,7 @@ export const walkRepositoryPaths = async (
const stat = await fs.stat(fullPath);
if (stat.size > MAX_FILE_SIZE) {
skippedLarge++;
skippedLargePaths.push(relativePath.replace(/\\/g, '/'));
return null;
}
return { path: relativePath.replace(/\\/g, '/'), size: stat.size };
@@ -73,6 +76,11 @@ export const walkRepositoryPaths = async (
console.warn(
` Skipped ${skippedLarge} large files (>${MAX_FILE_SIZE / 1024}KB, likely generated/vendored)`,
);
if (isVerboseIngestionEnabled()) {
for (const p of skippedLargePaths) {
console.warn(` - ${p}`);
}
}
}
return entries;
@@ -891,6 +891,7 @@ export const AST_FRAMEWORK_PATTERNS_BY_LANGUAGE = {
patterns: FRAMEWORK_AST_PATTERNS.riverpod,
},
],
[SupportedLanguages.Vue]: [], // Vue uses TypeScript AST framework detection
[SupportedLanguages.Cobol]: [], // Standalone regex processor — no AST framework patterns
} satisfies Record<SupportedLanguages, AstFrameworkPatternConfig[]>;
+167
View File
@@ -0,0 +1,167 @@
/**
* Heritage Map
*
* Unified inheritance data structure built from accumulated
* {@link ExtractedHeritage} records **after all chunks complete** (between
* chunk processing and call resolution). Consumes `ExtractedHeritage[]` and
* resolves type names to nodeIds via `lookupClassByName`, NOT graph-edge
* queries.
*
* Combines two previously separate concerns:
* 1. **Parent/ancestor lookup** (MRO-aware method resolution)
* 2. **Implementor lookup** (interface dispatch — which files contain
* classes implementing a given interface)
*/
import type { ExtractedHeritage } from './workers/parse-worker.js';
import type { ResolutionContext } from './resolution-context.js';
import { getLanguageFromFilename } from 'gitnexus-shared';
import { resolveExtendsType } from './heritage-processor.js';
// ---------------------------------------------------------------------------
// Public types
// ---------------------------------------------------------------------------
/** Maximum ancestor chain depth to prevent runaway traversal. */
const MAX_ANCESTOR_DEPTH = 32;
export interface HeritageMap {
/** Direct parents of `childNodeId` (extends + implements + trait-impl). */
getParents(childNodeId: string): string[];
/** Full ancestor chain (BFS, bounded depth, cycle-safe). */
getAncestors(childNodeId: string): string[];
/**
* File paths of classes that directly implement or extend-as-interface the
* given interface/abstract-class **name**. Replaces the standalone
* `ImplementorMap` — used by interface-dispatch in call resolution.
*/
getImplementorFiles(interfaceName: string): ReadonlySet<string>;
}
/** Shared empty set returned when no implementors are found. */
const EMPTY_SET: ReadonlySet<string> = new Set();
// ---------------------------------------------------------------------------
// Builder
// ---------------------------------------------------------------------------
/**
* Build a HeritageMap from accumulated ExtractedHeritage records.
*
* Resolves class/interface/struct/trait names to nodeIds via
* `ctx.symbols.lookupClassByName`. When a name resolves to multiple
* candidates, all are recorded (partial-class / cross-file scenario).
* Unresolvable names are silently skipped — a missing parent is better
* than a wrong edge.
*
* Also builds the implementor index (interface name → implementing file
* paths) that was previously maintained by `buildImplementorMap` in
* call-processor.ts.
*/
export const buildHeritageMap = (
heritage: readonly ExtractedHeritage[],
ctx: ResolutionContext,
): HeritageMap => {
// childNodeId → Set<parentNodeId> (Set to deduplicate cross-chunk duplicates)
const directParents = new Map<string, Set<string>>();
// interfaceName → Set<filePath> (implementor lookup for interface dispatch)
const implementorFiles = new Map<string, Set<string>>();
for (const h of heritage) {
// ── Parent lookup (nodeId-based) ────────────────────────────────
const childDefs = ctx.symbols.lookupClassByName(h.className);
const parentDefs = ctx.symbols.lookupClassByName(h.parentName);
if (childDefs.length > 0 && parentDefs.length > 0) {
for (const child of childDefs) {
for (const parent of parentDefs) {
// Skip self-references
if (child.nodeId === parent.nodeId) continue;
let parents = directParents.get(child.nodeId);
if (!parents) {
parents = new Set();
directParents.set(child.nodeId, parents);
}
parents.add(parent.nodeId);
}
}
}
// ── Implementor index (name-based) ──────────────────────────────
//
// Known limitation: Rust `kind: 'trait-impl'` entries are intentionally NOT
// added to the implementor index. Interface dispatch resolution currently
// does not traverse Rust trait objects, so recording them here would
// inflate the index without a consumer. Revisit if/when trait-object
// dispatch is added.
//
// Known limitation: `getImplementorFiles` is keyed by interface **name**
// (string), so two interfaces with the same unqualified name in different
// packages (e.g. `pkgA.IRepository` vs `pkgB.IRepository`) collide. This
// matches the behavior of the prior standalone `ImplementorMap` and is
// not a regression introduced by this consolidation.
let isImpl = false;
if (h.kind === 'implements') {
isImpl = true;
} else if (h.kind === 'extends') {
const lang = getLanguageFromFilename(h.filePath);
if (lang) {
const { type } = resolveExtendsType(h.parentName, h.filePath, ctx, lang);
isImpl = type === 'IMPLEMENTS';
}
}
if (isImpl) {
let files = implementorFiles.get(h.parentName);
if (!files) {
files = new Set();
implementorFiles.set(h.parentName, files);
}
files.add(h.filePath);
}
}
// --- Public API ---------------------------------------------------
const getParents = (childNodeId: string): string[] => {
const parents = directParents.get(childNodeId);
return parents ? [...parents] : [];
};
const getAncestors = (childNodeId: string): string[] => {
const result: string[] = [];
const visited = new Set<string>();
visited.add(childNodeId); // prevent cycles through the start node
// BFS with bounded depth
let frontier = getParents(childNodeId);
let depth = 0;
while (frontier.length > 0 && depth < MAX_ANCESTOR_DEPTH) {
const nextFrontier: string[] = [];
for (const parentId of frontier) {
if (visited.has(parentId)) continue;
visited.add(parentId);
result.push(parentId);
// Expand parent's own parents for next level
const grandparents = directParents.get(parentId);
if (grandparents) {
for (const gp of grandparents) {
if (!visited.has(gp)) nextFrontier.push(gp);
}
}
}
frontier = nextFrontier;
depth++;
}
return result;
};
const getImplementorFiles = (interfaceName: string): ReadonlySet<string> => {
return implementorFiles.get(interfaceName) ?? EMPTY_SET;
};
return { getParents, getAncestors, getImplementorFiles };
};
@@ -372,7 +372,7 @@ export const processHeritageFromExtracted = async (
/**
* Walk source files with the same heritage captures as parse-worker, producing
* {@link ExtractedHeritage} rows without mutating the graph. Used on the
* sequential pipeline path so `buildImplementorMap(..., ctx)` can run before
* sequential pipeline path so `buildHeritageMap(..., ctx)` can run before
* `processCalls` (worker path defers calls until heritage from all chunks exists).
*/
export async function extractExtractedHeritageFromFiles(
@@ -11,6 +11,7 @@ export const EXTENSIONS = [
'.ts',
'.jsx',
'.js',
'.vue',
'/index.tsx',
'/index.ts',
'/index.jsx',
@@ -0,0 +1,13 @@
/**
* Vue import resolver — delegates to TypeScript's standard resolver.
*
* Vue <script> blocks use the same import syntax as TypeScript (including
* tsconfig path aliases like `@/`), so no custom resolution logic is needed.
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { resolveStandard } from './standard.js';
import type { ImportResolverFn } from './types.js';
export const resolveVueImport: ImportResolverFn = (raw, fp, ctx) =>
resolveStandard(raw, fp, ctx, SupportedLanguages.TypeScript);
@@ -12,6 +12,7 @@
import type { SupportedLanguages } from 'gitnexus-shared';
import type { LanguageTypeConfig } from './type-extractors/types.js';
import type { CallRouter } from './call-routing.js';
import type { ClassExtractor } from './class-types.js';
import type { ExportChecker } from './export-detection.js';
import type { FieldExtractor } from './field-extractor.js';
import type { MethodExtractor } from './method-types.js';
@@ -131,6 +132,10 @@ interface LanguageProviderConfig {
* declarations. Produces MethodInfo[] with name, parameters, visibility, isAbstract,
* isFinal, annotations metadata. Default: undefined (no method extraction). */
readonly methodExtractor?: MethodExtractor;
/** Class/type extractor for deriving canonical qualified names for class-like symbols.
* Uses the same provider-driven strategy pattern as method/field extraction so
* namespace/package/module rules stay language-specific. */
readonly classExtractor?: ClassExtractor;
/** Extract a semantic description for a definition node (e.g., PHP Eloquent
* property arrays, relation method descriptions).
* Default: undefined (no description extraction). */
+185 -1
View File
@@ -9,19 +9,34 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as cCppConfig } from '../type-extractors/c-cpp.js';
import { cCppExportChecker } from '../export-detection.js';
import { resolveCImport, resolveCppImport } from '../import-resolvers/standard.js';
import { C_QUERIES, CPP_QUERIES } from '../tree-sitter-queries.js';
import { isCppInsideClassOrStruct } from '../utils/ast-helpers.js';
/**
* Node types for standard function declarations that need C/C++ declarator handling.
* Used by cCppExtractFunctionName to determine how to extract the function name.
*/
const FUNCTION_DECLARATION_TYPES = new Set([
'function_declaration',
'function_definition',
'async_function_declaration',
'generator_function_declaration',
'function_item',
]);
import type { SyntaxNode } from '../utils/ast-helpers.js';
import type { NodeLabel } from 'gitnexus-shared';
import type { LanguageProvider } from '../language-provider.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import {
cConfig as cFieldConfig,
cppConfig as cppFieldConfig,
} from '../field-extractors/configs/c-cpp.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import { cMethodConfig, cppMethodConfig } from '../method-extractors/configs/c-cpp.js';
const C_BUILT_INS: ReadonlySet<string> = new Set([
'printf',
@@ -130,6 +145,165 @@ const C_BUILT_INS: ReadonlySet<string> = new Set([
'put',
]);
const cClassExtractor = createClassExtractor({
language: SupportedLanguages.C,
typeDeclarationNodes: ['struct_specifier', 'enum_specifier'],
});
const cppClassExtractor = createClassExtractor({
language: SupportedLanguages.CPlusPlus,
typeDeclarationNodes: ['class_specifier', 'struct_specifier', 'enum_specifier'],
ancestorScopeNodeTypes: ['namespace_definition', 'class_specifier', 'struct_specifier'],
});
/**
* C/C++ function name extraction — unwraps pointer_declarator / reference_declarator /
* function_declarator / qualified_identifier chains to find the actual function name.
* Handles field_identifier (method inside class body) and parenthesized_declarator.
*/
const cCppExtractFunctionName = (
node: SyntaxNode,
): { funcName: string | null; label: NodeLabel } | null => {
if (!FUNCTION_DECLARATION_TYPES.has(node.type)) return null;
let funcName: string | null = null;
let label: NodeLabel = 'Function';
// C/C++: function_definition -> [pointer_declarator ->] function_declarator -> qualified_identifier/identifier
// Unwrap pointer_declarator / reference_declarator wrappers to reach function_declarator
let declarator = node.childForFieldName?.('declarator');
if (!declarator) {
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (c?.type === 'function_declarator') {
declarator = c;
break;
}
}
}
while (
declarator &&
(declarator.type === 'pointer_declarator' || declarator.type === 'reference_declarator')
) {
let nextDeclarator = declarator.childForFieldName?.('declarator');
if (!nextDeclarator) {
for (let i = 0; i < declarator.childCount; i++) {
const c = declarator.child(i);
if (
c?.type === 'function_declarator' ||
c?.type === 'pointer_declarator' ||
c?.type === 'reference_declarator'
) {
nextDeclarator = c;
break;
}
}
}
declarator = nextDeclarator;
}
if (declarator) {
let innerDeclarator = declarator.childForFieldName?.('declarator');
if (!innerDeclarator) {
for (let i = 0; i < declarator.childCount; i++) {
const c = declarator.child(i);
if (
c?.type === 'qualified_identifier' ||
c?.type === 'identifier' ||
c?.type === 'field_identifier' ||
c?.type === 'parenthesized_declarator'
) {
innerDeclarator = c;
break;
}
}
}
if (innerDeclarator?.type === 'qualified_identifier') {
let nameNode = innerDeclarator.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < innerDeclarator.childCount; i++) {
const c = innerDeclarator.child(i);
if (c?.type === 'identifier') {
nameNode = c;
break;
}
}
}
if (nameNode?.text) {
funcName = nameNode.text;
label = 'Method';
}
} else if (
innerDeclarator?.type === 'identifier' ||
innerDeclarator?.type === 'field_identifier'
) {
// field_identifier is used for method names inside C++ class bodies
funcName = innerDeclarator.text;
if (innerDeclarator.type === 'field_identifier') label = 'Method';
} else if (innerDeclarator?.type === 'parenthesized_declarator') {
let nestedId: SyntaxNode | null = null;
for (let i = 0; i < innerDeclarator.childCount; i++) {
const c = innerDeclarator.child(i);
if (c?.type === 'qualified_identifier' || c?.type === 'identifier') {
nestedId = c;
break;
}
}
if (nestedId?.type === 'qualified_identifier') {
let nameNode = nestedId.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < nestedId.childCount; i++) {
const c = nestedId.child(i);
if (c?.type === 'identifier') {
nameNode = c;
break;
}
}
}
if (nameNode?.text) {
funcName = nameNode.text;
label = 'Method';
}
} else if (nestedId?.type === 'identifier') {
funcName = nestedId.text;
}
}
}
// Fallback for other node types in FUNCTION_DECLARATION_TYPES (e.g. function_item for Rust in C++ tree)
if (!funcName) {
let nameNode = node.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (
c?.type === 'identifier' ||
c?.type === 'property_identifier' ||
c?.type === 'simple_identifier'
) {
nameNode = c;
break;
}
}
}
funcName = nameNode?.text ?? null;
}
return { funcName, label };
};
/** Check if a C/C++ function_definition is inside a class or struct body.
* Used by cppLabelOverride to skip duplicate function captures
* that are already covered by definition.method queries. */
function isCppInsideClassOrStruct(functionNode: SyntaxNode): boolean {
let ancestor: SyntaxNode | null = functionNode?.parent ?? null;
while (ancestor) {
if (ancestor.type === 'class_specifier' || ancestor.type === 'struct_specifier') return true;
ancestor = ancestor.parent;
}
return false;
}
/** Label override shared by C and C++: skip function_definition captures inside class/struct
* bodies (they're duplicates of definition.method captures). */
const cppLabelOverride: NonNullable<LanguageProvider['labelOverride']> = (
@@ -149,6 +323,11 @@ export const cProvider = defineLanguage({
importResolver: resolveCImport,
importSemantics: 'wildcard',
fieldExtractor: createFieldExtractor(cFieldConfig),
methodExtractor: createMethodExtractor({
...cMethodConfig,
extractFunctionName: cCppExtractFunctionName,
}),
classExtractor: cClassExtractor,
labelOverride: cppLabelOverride,
builtInNames: C_BUILT_INS,
});
@@ -163,6 +342,11 @@ export const cppProvider = defineLanguage({
importSemantics: 'wildcard',
mroStrategy: 'leftmost-base',
fieldExtractor: createFieldExtractor(cppFieldConfig),
methodExtractor: createMethodExtractor({
...cppMethodConfig,
extractFunctionName: cCppExtractFunctionName,
}),
classExtractor: cppClassExtractor,
labelOverride: cppLabelOverride,
builtInNames: C_BUILT_INS,
});
@@ -7,6 +7,7 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as csharpConfig } from '../type-extractors/csharp.js';
import { csharpExportChecker } from '../export-detection.js';
@@ -125,5 +126,24 @@ export const csharpProvider = defineLanguage({
mroStrategy: 'implements-split',
fieldExtractor: createFieldExtractor(csharpFieldConfig),
methodExtractor: createMethodExtractor(csharpMethodConfig),
classExtractor: createClassExtractor({
language: SupportedLanguages.CSharp,
typeDeclarationNodes: [
'class_declaration',
'interface_declaration',
'struct_declaration',
'enum_declaration',
'record_declaration',
],
fileScopeNodeTypes: ['file_scoped_namespace_declaration'],
ancestorScopeNodeTypes: [
'namespace_declaration',
'class_declaration',
'interface_declaration',
'struct_declaration',
'enum_declaration',
'record_declaration',
],
}),
builtInNames: BUILT_INS,
});
+27 -4
View File
@@ -12,8 +12,9 @@
import type { SyntaxNode } from '../utils/ast-helpers.js';
import type { NodeLabel } from 'gitnexus-shared';
import { FUNCTION_NODE_TYPES, extractFunctionName } from '../utils/ast-helpers.js';
import { FUNCTION_NODE_TYPES } from '../utils/ast-helpers.js';
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as dartConfig } from '../type-extractors/dart.js';
import { dartExportChecker } from '../export-detection.js';
@@ -21,6 +22,8 @@ import { resolveDartImport } from '../import-resolvers/dart.js';
import { DART_QUERIES } from '../tree-sitter-queries.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { dartConfig as dartFieldConfig } from '../field-extractors/configs/dart.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import { dartMethodConfig } from '../method-extractors/configs/dart.js';
/**
* Resolve the enclosing function from a `function_body` node by looking at its
@@ -28,8 +31,8 @@ import { dartConfig as dartFieldConfig } from '../field-extractors/configs/dart.
* function_body are siblings under program or class_body, unlike most languages
* where the function declaration wraps both.
*
* Delegates name extraction to the shared `extractFunctionName` which already
* handles Dart's function_signature and method_signature node types.
* Extracts the function name inline — Dart uses function_signature and
* method_signature (which wraps function_signature) as its FUNCTION_NODE_TYPES.
*/
const dartEnclosingFunctionFinder = (
node: SyntaxNode,
@@ -37,7 +40,21 @@ const dartEnclosingFunctionFinder = (
if (node.type !== 'function_body') return null;
const prev = node.previousSibling;
if (!prev || !FUNCTION_NODE_TYPES.has(prev.type)) return null;
const { funcName, label } = extractFunctionName(prev);
// method_signature wraps function_signature — unwrap to reach the name
let target = prev;
let label: NodeLabel = 'Function';
if (prev.type === 'method_signature') {
label = 'Method';
for (let i = 0; i < prev.childCount; i++) {
const c = prev.child(i);
if (c?.type === 'function_signature') {
target = c;
break;
}
}
}
const funcName = target.childForFieldName?.('name')?.text ?? null;
return funcName ? { funcName, label } : null;
};
@@ -75,6 +92,12 @@ export const dartProvider = defineLanguage({
importResolver: resolveDartImport,
importSemantics: 'wildcard',
fieldExtractor: createFieldExtractor(dartFieldConfig),
methodExtractor: createMethodExtractor(dartMethodConfig),
classExtractor: createClassExtractor({
language: SupportedLanguages.Dart,
typeDeclarationNodes: ['class_definition', 'extension_declaration', 'enum_declaration'],
ancestorScopeNodeTypes: ['class_definition', 'extension_declaration', 'enum_declaration'],
}),
enclosingFunctionFinder: dartEnclosingFunctionFinder,
builtInNames: BUILT_INS,
});
@@ -10,6 +10,7 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as goConfig } from '../type-extractors/go.js';
import { goExportChecker } from '../export-detection.js';
@@ -17,6 +18,8 @@ import { resolveGoImport } from '../import-resolvers/go.js';
import { GO_QUERIES } from '../tree-sitter-queries.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { goConfig as goFieldConfig } from '../field-extractors/configs/go.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import { goMethodConfig } from '../method-extractors/configs/go.js';
export const goProvider = defineLanguage({
id: SupportedLanguages.Go,
@@ -27,4 +30,21 @@ export const goProvider = defineLanguage({
importResolver: resolveGoImport,
importSemantics: 'wildcard',
fieldExtractor: createFieldExtractor(goFieldConfig),
methodExtractor: createMethodExtractor(goMethodConfig),
classExtractor: createClassExtractor({
language: SupportedLanguages.Go,
typeDeclarationNodes: ['type_declaration'],
fileScopeNodeTypes: ['package_clause'],
extractName(node) {
const typeSpec = node.namedChildren.find((child) => child.type === 'type_spec');
return typeSpec?.childForFieldName('name')?.text;
},
extractType(node) {
const typeSpec = node.namedChildren.find((child) => child.type === 'type_spec');
const typeNode = typeSpec?.childForFieldName('type');
if (typeNode?.type === 'struct_type') return 'Struct';
if (typeNode?.type === 'interface_type') return 'Interface';
return undefined;
},
}),
});
@@ -23,6 +23,7 @@ import { phpProvider } from './php.js';
import { rubyProvider } from './ruby.js';
import { swiftProvider } from './swift.js';
import { dartProvider } from './dart.js';
import { vueProvider } from './vue.js';
import { cobolProvider } from './cobol.js';
export const providers = {
@@ -40,6 +41,7 @@ export const providers = {
[SupportedLanguages.Ruby]: rubyProvider,
[SupportedLanguages.Swift]: swiftProvider,
[SupportedLanguages.Dart]: dartProvider,
[SupportedLanguages.Vue]: vueProvider,
[SupportedLanguages.Cobol]: cobolProvider,
} satisfies Record<SupportedLanguages, LanguageProvider>;
@@ -8,6 +8,7 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { javaTypeConfig } from '../type-extractors/jvm.js';
import { javaExportChecker } from '../export-detection.js';
@@ -31,4 +32,20 @@ export const javaProvider = defineLanguage({
mroStrategy: 'implements-split',
fieldExtractor: createFieldExtractor(javaConfig),
methodExtractor: createMethodExtractor(javaMethodConfig),
classExtractor: createClassExtractor({
language: SupportedLanguages.Java,
typeDeclarationNodes: [
'class_declaration',
'interface_declaration',
'enum_declaration',
'record_declaration',
],
fileScopeNodeTypes: ['package_declaration'],
ancestorScopeNodeTypes: [
'class_declaration',
'interface_declaration',
'enum_declaration',
'record_declaration',
],
}),
});
@@ -8,6 +8,7 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { kotlinTypeConfig } from '../type-extractors/jvm.js';
import { kotlinExportChecker } from '../export-detection.js';
@@ -15,12 +16,26 @@ import { resolveKotlinImport } from '../import-resolvers/jvm.js';
import { extractKotlinNamedBindings } from '../named-bindings/kotlin.js';
import { appendKotlinWildcard } from '../import-resolvers/jvm.js';
import { KOTLIN_QUERIES } from '../tree-sitter-queries.js';
import { isKotlinClassMethod } from '../utils/ast-helpers.js';
import type { SyntaxNode } from '../utils/ast-helpers.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { kotlinConfig } from '../field-extractors/configs/jvm.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import { kotlinMethodConfig } from '../method-extractors/configs/jvm.js';
/** Check if a Kotlin function_declaration capture is inside a class_body (i.e., a method).
* Kotlin grammar uses function_declaration for both top-level functions and class methods.
* Returns true when the captured definition node has a class_body ancestor. */
function isKotlinClassMethod(
captureNode: { parent?: SyntaxNode | null } | null | undefined,
): boolean {
let ancestor = captureNode?.parent;
while (ancestor) {
if (ancestor.type === 'class_body') return true;
ancestor = ancestor.parent;
}
return false;
}
const BUILT_INS: ReadonlySet<string> = new Set([
'println',
'print',
@@ -92,6 +107,16 @@ export const kotlinProvider = defineLanguage({
mroStrategy: 'implements-split',
fieldExtractor: createFieldExtractor(kotlinConfig),
methodExtractor: createMethodExtractor(kotlinMethodConfig),
classExtractor: createClassExtractor({
language: SupportedLanguages.Kotlin,
typeDeclarationNodes: ['class_declaration', 'object_declaration', 'companion_object'],
fileScopeNodeTypes: ['package_header'],
ancestorScopeNodeTypes: ['class_declaration', 'object_declaration', 'companion_object'],
extractType(node) {
if (node.type !== 'class_declaration') return undefined;
return node.children.some((child) => child?.text === 'interface') ? 'Interface' : 'Class';
},
}),
builtInNames: BUILT_INS,
labelOverride: (functionNode, defaultLabel) => {
if (defaultLabel !== 'Function') return defaultLabel;
@@ -7,6 +7,7 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as phpConfig } from '../type-extractors/php.js';
import { phpExportChecker } from '../export-detection.js';
@@ -17,6 +18,8 @@ import { findDescendant, extractStringContent, type SyntaxNode } from '../utils/
import type { NodeLabel } from 'gitnexus-shared';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { phpConfig as phpFieldConfig } from '../field-extractors/configs/php.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import { phpMethodConfig } from '../method-extractors/configs/php.js';
const BUILT_INS: ReadonlySet<string> = new Set([
'echo',
@@ -231,6 +234,12 @@ export const phpProvider = defineLanguage({
importResolver: resolvePhpImport,
namedBindingExtractor: extractPhpNamedBindings,
fieldExtractor: createFieldExtractor(phpFieldConfig),
methodExtractor: createMethodExtractor(phpMethodConfig),
classExtractor: createClassExtractor({
language: SupportedLanguages.PHP,
typeDeclarationNodes: ['class_declaration', 'interface_declaration', 'enum_declaration'],
ancestorScopeNodeTypes: ['namespace_definition'],
}),
descriptionExtractor: phpDescriptionExtractor,
isRouteFile: isPhpRouteFile,
builtInNames: BUILT_INS,
@@ -11,6 +11,7 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as pythonConfig } from '../type-extractors/python.js';
import { pythonExportChecker } from '../export-detection.js';
@@ -19,6 +20,8 @@ import { extractPythonNamedBindings } from '../named-bindings/python.js';
import { PYTHON_QUERIES } from '../tree-sitter-queries.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { pythonConfig as pythonFieldConfig } from '../field-extractors/configs/python.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import { pythonMethodConfig } from '../method-extractors/configs/python.js';
const BUILT_INS: ReadonlySet<string> = new Set([
'print',
@@ -61,5 +64,11 @@ export const pythonProvider = defineLanguage({
importSemantics: 'namespace',
mroStrategy: 'c3',
fieldExtractor: createFieldExtractor(pythonFieldConfig),
methodExtractor: createMethodExtractor(pythonMethodConfig),
classExtractor: createClassExtractor({
language: SupportedLanguages.Python,
typeDeclarationNodes: ['class_definition'],
ancestorScopeNodeTypes: ['class_definition'],
}),
builtInNames: BUILT_INS,
});
@@ -8,7 +8,10 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import type { NodeLabel } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import type { SyntaxNode } from '../utils/ast-helpers.js';
import { typeConfig as rubyConfig } from '../type-extractors/ruby.js';
import { routeRubyCall } from '../call-routing.js';
import { rubyExportChecker } from '../export-detection.js';
@@ -16,6 +19,27 @@ import { resolveRubyImport } from '../import-resolvers/ruby.js';
import { RUBY_QUERIES } from '../tree-sitter-queries.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { rubyConfig as rubyFieldConfig } from '../field-extractors/configs/ruby.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import { rubyMethodConfig } from '../method-extractors/configs/ruby.js';
/** Ruby method/singleton_method: extract name from 'name' field, label as Method. */
const rubyExtractFunctionName = (
node: SyntaxNode,
): { funcName: string | null; label: NodeLabel } | null => {
if (node.type !== 'method' && node.type !== 'singleton_method') return null;
let nameNode = node.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (c?.type === 'identifier') {
nameNode = c;
break;
}
}
}
return { funcName: nameNode?.text ?? null, label: 'Method' };
};
const BUILT_INS: ReadonlySet<string> = new Set([
'puts',
@@ -85,5 +109,14 @@ export const rubyProvider = defineLanguage({
callRouter: routeRubyCall,
importSemantics: 'wildcard',
fieldExtractor: createFieldExtractor(rubyFieldConfig),
methodExtractor: createMethodExtractor({
...rubyMethodConfig,
extractFunctionName: rubyExtractFunctionName,
}),
classExtractor: createClassExtractor({
language: SupportedLanguages.Ruby,
typeDeclarationNodes: ['class'],
ancestorScopeNodeTypes: ['module', 'class'],
}),
builtInNames: BUILT_INS,
});
@@ -11,7 +11,10 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import type { NodeLabel } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import type { SyntaxNode } from '../utils/ast-helpers.js';
import { typeConfig as rustConfig } from '../type-extractors/rust.js';
import { rustExportChecker } from '../export-detection.js';
import { resolveRustImport } from '../import-resolvers/rust.js';
@@ -19,6 +22,37 @@ import { extractRustNamedBindings } from '../named-bindings/rust.js';
import { RUST_QUERIES } from '../tree-sitter-queries.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { rustConfig as rustFieldConfig } from '../field-extractors/configs/rust.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import { rustMethodConfig } from '../method-extractors/configs/rust.js';
/** Rust impl_item: find the function_item child and extract its name as a Method. */
const rustExtractFunctionName = (
node: SyntaxNode,
): { funcName: string | null; label: NodeLabel } | null => {
if (node.type !== 'impl_item') return null;
let funcItem: SyntaxNode | null = null;
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (c?.type === 'function_item') {
funcItem = c;
break;
}
}
if (!funcItem) return null;
let nameNode = funcItem.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < funcItem.childCount; i++) {
const c = funcItem.child(i);
if (c?.type === 'identifier') {
nameNode = c;
break;
}
}
}
return { funcName: nameNode?.text ?? null, label: 'Method' };
};
const BUILT_INS: ReadonlySet<string> = new Set([
'unwrap',
@@ -87,5 +121,14 @@ export const rustProvider = defineLanguage({
namedBindingExtractor: extractRustNamedBindings,
mroStrategy: 'qualified-syntax',
fieldExtractor: createFieldExtractor(rustFieldConfig),
methodExtractor: createMethodExtractor({
...rustMethodConfig,
extractFunctionName: rustExtractFunctionName,
}),
classExtractor: createClassExtractor({
language: SupportedLanguages.Rust,
typeDeclarationNodes: ['struct_item', 'enum_item'],
ancestorScopeNodeTypes: ['mod_item', 'struct_item', 'enum_item'],
}),
builtInNames: BUILT_INS,
});
@@ -11,14 +11,19 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import type { NodeLabel } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as swiftConfig } from '../type-extractors/swift.js';
import { swiftExportChecker } from '../export-detection.js';
import { resolveSwiftImport } from '../import-resolvers/swift.js';
import { SWIFT_QUERIES } from '../tree-sitter-queries.js';
import type { SwiftPackageConfig } from '../language-config.js';
import type { SyntaxNode } from '../utils/ast-helpers.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { swiftConfig as swiftFieldConfig } from '../field-extractors/configs/swift.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import { swiftMethodConfig } from '../method-extractors/configs/swift.js';
/**
* Group Swift files by SPM target for implicit module visibility.
@@ -107,6 +112,15 @@ function wireSwiftImplicitImports(
}
}
/** Swift init/deinit declarations have special names and Constructor label. */
const swiftExtractFunctionName = (
node: SyntaxNode,
): { funcName: string | null; label: NodeLabel } | null => {
if (node.type === 'init_declaration') return { funcName: 'init', label: 'Constructor' };
if (node.type === 'deinit_declaration') return { funcName: 'deinit', label: 'Constructor' };
return null; // fall through to generic
};
const BUILT_INS: ReadonlySet<string> = new Set([
'print',
'debugPrint',
@@ -227,6 +241,22 @@ export const swiftProvider = defineLanguage({
importSemantics: 'wildcard',
heritageDefaultEdge: 'IMPLEMENTS',
fieldExtractor: createFieldExtractor(swiftFieldConfig),
methodExtractor: createMethodExtractor({
...swiftMethodConfig,
extractFunctionName: swiftExtractFunctionName,
}),
classExtractor: createClassExtractor({
language: SupportedLanguages.Swift,
typeDeclarationNodes: ['class_declaration', 'protocol_declaration'],
ancestorScopeNodeTypes: ['class_declaration', 'protocol_declaration'],
extractType(node) {
if (node.type === 'protocol_declaration') return 'Interface';
if (node.type !== 'class_declaration') return undefined;
if (node.children.some((child) => child?.text === 'struct')) return 'Struct';
if (node.children.some((child) => child?.text === 'enum')) return 'Enum';
return 'Class';
},
}),
implicitImportWirer: wireSwiftImplicitImports,
builtInNames: BUILT_INS,
});
@@ -8,7 +8,11 @@
*/
import { SupportedLanguages } from 'gitnexus-shared';
import type { NodeLabel } from 'gitnexus-shared';
import { defineLanguage } from '../language-provider.js';
import { createClassExtractor } from '../class-extractors/generic.js';
import type { ClassExtractionConfig } from '../class-types.js';
import type { SyntaxNode } from '../utils/ast-helpers.js';
import { typeConfig as typescriptConfig } from '../type-extractors/typescript.js';
import { tsExportChecker } from '../export-detection.js';
import { resolveTypescriptImport, resolveJavascriptImport } from '../import-resolvers/standard.js';
@@ -17,8 +21,38 @@ import { TYPESCRIPT_QUERIES, JAVASCRIPT_QUERIES } from '../tree-sitter-queries.j
import { typescriptFieldExtractor } from '../field-extractors/typescript.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { javascriptConfig } from '../field-extractors/configs/typescript-javascript.js';
import { createMethodExtractor } from '../method-extractors/generic.js';
import {
typescriptMethodConfig,
javascriptMethodConfig,
} from '../method-extractors/configs/typescript-javascript.js';
const BUILT_INS: ReadonlySet<string> = new Set([
/**
* TypeScript/JavaScript: arrow_function and function_expression get their name
* from the parent variable_declarator (e.g. `const foo = () => {}`).
*/
const tsExtractFunctionName = (
node: SyntaxNode,
): { funcName: string | null; label: NodeLabel } | null => {
if (node.type !== 'arrow_function' && node.type !== 'function_expression') return null;
const parent = node.parent;
if (parent?.type !== 'variable_declarator') return null;
let nameNode = parent.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < parent.childCount; i++) {
const c = parent.child(i);
if (c?.type === 'identifier') {
nameNode = c;
break;
}
}
}
return { funcName: nameNode?.text ?? null, label: 'Function' };
};
export const BUILT_INS: ReadonlySet<string> = new Set([
'console',
'log',
'warn',
@@ -115,6 +149,22 @@ const BUILT_INS: ReadonlySet<string> = new Set([
'valueOf',
]);
const tsJsClassConfig: ClassExtractionConfig = {
language: SupportedLanguages.TypeScript,
typeDeclarationNodes: [
'class_declaration',
'abstract_class_declaration',
'interface_declaration',
'enum_declaration',
],
ancestorScopeNodeTypes: [
'class_declaration',
'abstract_class_declaration',
'interface_declaration',
'enum_declaration',
],
};
export const typescriptProvider = defineLanguage({
id: SupportedLanguages.TypeScript,
extensions: ['.ts', '.tsx'],
@@ -124,6 +174,11 @@ export const typescriptProvider = defineLanguage({
importResolver: resolveTypescriptImport,
namedBindingExtractor: extractTsNamedBindings,
fieldExtractor: typescriptFieldExtractor,
methodExtractor: createMethodExtractor({
...typescriptMethodConfig,
extractFunctionName: tsExtractFunctionName,
}),
classExtractor: createClassExtractor(tsJsClassConfig),
builtInNames: BUILT_INS,
});
@@ -136,5 +191,13 @@ export const javascriptProvider = defineLanguage({
importResolver: resolveJavascriptImport,
namedBindingExtractor: extractTsNamedBindings,
fieldExtractor: createFieldExtractor(javascriptConfig),
methodExtractor: createMethodExtractor({
...javascriptMethodConfig,
extractFunctionName: tsExtractFunctionName,
}),
classExtractor: createClassExtractor({
...tsJsClassConfig,
language: SupportedLanguages.JavaScript,
}),
builtInNames: BUILT_INS,
});
@@ -0,0 +1,86 @@
/**
* Vue language provider.
*
* Vue SFCs are preprocessed by extracting the <script> / <script setup>
* block content, which is then parsed as TypeScript. This provider reuses
* nearly all TypeScript infrastructure — queries, type config, field
* extraction, and named binding extraction.
*
* Export detection for <script setup> is handled directly in the parse
* worker (all top-level bindings are implicitly exported). The export
* checker here is used as fallback for non-setup <script> blocks.
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { createClassExtractor } from '../class-extractors/generic.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as typescriptConfig } from '../type-extractors/typescript.js';
import { tsExportChecker } from '../export-detection.js';
import { resolveVueImport } from '../import-resolvers/vue.js';
import { extractTsNamedBindings } from '../named-bindings/typescript.js';
import { TYPESCRIPT_QUERIES } from '../tree-sitter-queries.js';
import { typescriptFieldExtractor } from '../field-extractors/typescript.js';
import { BUILT_INS as TS_BUILT_INS } from './typescript.js';
const VUE_SPECIFIC_BUILT_INS = [
'ref',
'reactive',
'computed',
'watch',
'watchEffect',
'onMounted',
'onUnmounted',
'onBeforeMount',
'onBeforeUnmount',
'onUpdated',
'onBeforeUpdate',
'nextTick',
'defineProps',
'defineEmits',
'defineExpose',
'defineOptions',
'defineSlots',
'defineModel',
'withDefaults',
'toRef',
'toRefs',
'unref',
'isRef',
'shallowRef',
'triggerRef',
'provide',
'inject',
'useSlots',
'useAttrs',
] as const;
const VUE_BUILT_INS: ReadonlySet<string> = new Set([...TS_BUILT_INS, ...VUE_SPECIFIC_BUILT_INS]);
const vueClassExtractor = createClassExtractor({
language: SupportedLanguages.Vue,
typeDeclarationNodes: [
'class_declaration',
'abstract_class_declaration',
'interface_declaration',
'enum_declaration',
],
ancestorScopeNodeTypes: [
'class_declaration',
'abstract_class_declaration',
'interface_declaration',
'enum_declaration',
],
});
export const vueProvider = defineLanguage({
id: SupportedLanguages.Vue,
extensions: ['.vue'],
treeSitterQueries: TYPESCRIPT_QUERIES,
typeConfig: typescriptConfig,
exportChecker: tsExportChecker,
importResolver: resolveVueImport,
namedBindingExtractor: extractTsNamedBindings,
fieldExtractor: typescriptFieldExtractor,
classExtractor: vueClassExtractor,
builtInNames: VUE_BUILT_INS,
});
@@ -0,0 +1,409 @@
// gitnexus/src/core/ingestion/method-extractors/configs/c-cpp.ts
// Verified against tree-sitter-cpp ^0.23.4
import { SupportedLanguages } from 'gitnexus-shared';
import type {
MethodExtractionConfig,
ParameterInfo,
MethodVisibility,
} from '../../method-types.js';
import { hasKeyword } from '../../field-extractors/configs/helpers.js';
import { extractSimpleTypeName } from '../../type-extractors/shared.js';
import type { SyntaxNode } from '../../utils/ast-helpers.js';
// ---------------------------------------------------------------------------
// C/C++ helpers
// ---------------------------------------------------------------------------
/**
* Find the function_declarator inside a method node, handling pointer/reference
* return types where the function_declarator is nested inside a pointer_declarator
* or reference_declarator.
*/
function findFunctionDeclarator(node: SyntaxNode): SyntaxNode | null {
const declarator = node.childForFieldName('declarator');
if (!declarator) return null;
if (declarator.type === 'function_declarator') return declarator;
// Recursively unwrap pointer_declarator / reference_declarator chains
// (e.g. int** (*pfn)() has pointer_declarator → pointer_declarator → function_declarator)
let current: SyntaxNode | null = declarator;
while (current) {
for (let i = 0; i < current.namedChildCount; i++) {
const child = current.namedChild(i);
if (child?.type === 'function_declarator') return child;
}
// Go deeper into nested pointer/reference declarators
const next = current.namedChildren.find(
(c) => c.type === 'pointer_declarator' || c.type === 'reference_declarator',
);
current = next ?? null;
}
return null;
}
/**
* Detect `= delete` and `= default` special member function declarations.
* These are not callable methods and should be suppressed from extraction.
* tree-sitter-cpp ^0.23.4 emits `delete_method_clause` / `default_method_clause`
* as named children of the function_definition node.
*/
function isDeletedOrDefaulted(node: SyntaxNode): boolean {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === 'delete_method_clause' || child?.type === 'default_method_clause') {
return true;
}
}
return false;
}
/**
* Extract method name from a function_declarator.
* The name is the `declarator` field of the function_declarator — typically a
* field_identifier, but can be a destructor_name (~ClassName) or operator name.
*/
function extractCppMethodName(node: SyntaxNode): string | undefined {
const funcDecl = findFunctionDeclarator(node);
if (!funcDecl) return undefined;
// Suppress `= delete` and `= default` special members — these are not callable
// methods and should not appear in HAS_METHOD edges.
if (isDeletedOrDefaulted(node)) return undefined;
const nameNode = funcDecl.childForFieldName('declarator');
if (!nameNode) return undefined;
// destructor_name: ~ClassName
if (nameNode.type === 'destructor_name') return nameNode.text;
// operator_name: operator==, operator+, etc.
if (nameNode.type === 'operator_name') return nameNode.text;
return nameNode.text;
}
/**
* Extract return type from the `type` field of the method node.
* tree-sitter-cpp puts the return type as the `type` field on field_declaration
* and function_definition nodes.
*/
function extractCppReturnType(node: SyntaxNode): string | undefined {
const typeNode = node.childForFieldName('type');
if (typeNode) {
const typeText = typeNode.text?.trim();
// C++11 trailing return type: `auto foo() -> ReturnType`
// When the declared type is `auto`, check for a trailing_return_type on the
// function_declarator which holds the actual return type.
if (typeText === 'auto') {
const funcDecl = findFunctionDeclarator(node);
if (funcDecl) {
for (let i = 0; i < funcDecl.namedChildCount; i++) {
const child = funcDecl.namedChild(i);
if (child?.type === 'trailing_return_type') {
// trailing_return_type contains a type_descriptor with the real type
const typeDesc = child.firstNamedChild;
if (typeDesc) return typeDesc.text?.trim();
}
}
}
}
return typeText;
}
// Fallback: first type-like named child (for declarations without type field)
const first = node.firstNamedChild;
if (
first &&
(first.type === 'primitive_type' ||
first.type === 'type_identifier' ||
first.type === 'sized_type_specifier' ||
first.type === 'template_type')
) {
return first.text?.trim();
}
return undefined;
}
/**
* Extract parameters from the parameter_list inside the function_declarator.
*
* C/C++ uses parameter_declaration (required) and optional_parameter_declaration
* (with default value). Variadic `...` appears as a variadic_parameter_declaration.
*/
function extractCppParameters(node: SyntaxNode): ParameterInfo[] {
const funcDecl = findFunctionDeclarator(node);
if (!funcDecl) return [];
const paramList = funcDecl.childForFieldName('parameters');
if (!paramList) return [];
const params: ParameterInfo[] = [];
for (let i = 0; i < paramList.namedChildCount; i++) {
const param = paramList.namedChild(i);
if (!param) continue;
switch (param.type) {
case 'parameter_declaration': {
const typeNode = param.childForFieldName('type');
const declNode = param.childForFieldName('declarator');
// Extract name — may be wrapped in pointer_declarator or reference_declarator
const name = extractParamName(declNode);
params.push({
name: name ?? typeNode?.text?.trim() ?? '?',
type: typeNode
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: false,
isVariadic: false,
});
break;
}
case 'optional_parameter_declaration': {
const typeNode = param.childForFieldName('type');
const declNode = param.childForFieldName('declarator');
const name = extractParamName(declNode);
params.push({
name: name ?? typeNode?.text?.trim() ?? '?',
type: typeNode
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: true,
isVariadic: false,
});
break;
}
case 'variadic_parameter_declaration': {
// C-style `...` or typed variadic `T... args`
const typeNode = param.childForFieldName('type');
const declNode = param.childForFieldName('declarator');
const name = extractParamName(declNode);
params.push({
name: name ?? '...',
type: typeNode
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: false,
isVariadic: true,
});
break;
}
case 'variadic_parameter': {
// Bare `...` (C-style)
params.push({
name: '...',
type: null,
rawType: null,
isOptional: false,
isVariadic: true,
});
break;
}
}
}
// C/C++: bare `...` token in parameter list is an unnamed child (not a named node).
// Check all children for the unnamed `...` token when no variadic was detected above.
if (!params.some((p) => p.isVariadic)) {
for (let i = 0; i < paramList.childCount; i++) {
const child = paramList.child(i);
if (child && !child.isNamed && child.text === '...') {
params.push({
name: '...',
type: null,
rawType: null,
isOptional: false,
isVariadic: true,
});
break;
}
}
}
return params;
}
/** Extract parameter name, recursively unwrapping pointer/reference declarators. */
function extractParamName(declNode: SyntaxNode | null): string | undefined {
if (!declNode) return undefined;
if (declNode.type === 'identifier') return declNode.text;
// Recursively unwrap pointer_declarator / reference_declarator chains (e.g. int** ptr)
for (let i = 0; i < declNode.namedChildCount; i++) {
const child = declNode.namedChild(i);
if (!child) continue;
if (child.type === 'identifier') return child.text;
if (child.type === 'pointer_declarator' || child.type === 'reference_declarator') {
return extractParamName(child);
}
}
return undefined;
}
/**
* Detect C++ access specifier by walking backwards through siblings.
* Mirrors the field extractor pattern in c-cpp.ts.
*/
function extractCppVisibility(node: SyntaxNode): MethodVisibility {
// If this node was unwrapped from a template_declaration, the access_specifier
// is a sibling of the template_declaration in field_declaration_list, not of
// this node — climb up one level before walking backward.
const startNode = node.parent?.type === 'template_declaration' ? node.parent : node;
let sibling = startNode.previousNamedSibling;
while (sibling) {
if (sibling.type === 'access_specifier') {
const text = sibling.text.replace(':', '').trim();
if (text === 'public' || text === 'private' || text === 'protected') return text;
}
sibling = sibling.previousNamedSibling;
}
// Default: struct/union = public, class = private
const parent = startNode.parent?.parent;
return parent?.type === 'struct_specifier' || parent?.type === 'union_specifier'
? 'public'
: 'private';
}
/**
* Detect pure virtual methods (`= 0`).
* tree-sitter-cpp emits `=` (unnamed) followed by `number_literal` with text `0`.
*/
function isPureVirtual(node: SyntaxNode): boolean {
let foundEquals = false;
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (!child) continue;
if (child.text === '=') {
foundEquals = true;
} else if (foundEquals && child.type === 'number_literal' && child.text === '0') {
return true;
} else if (foundEquals) {
foundEquals = false; // Reset if something else follows `=`
}
}
return false;
}
/**
* Check for a virtual_specifier ('final' or 'override') inside the function_declarator.
* In tree-sitter-cpp, these are named children of the function_declarator, not the
* method node itself.
*/
function hasVirtualSpecifier(node: SyntaxNode, keyword: string): boolean {
const funcDecl = findFunctionDeclarator(node);
if (!funcDecl) return false;
for (let i = 0; i < funcDecl.namedChildCount; i++) {
const child = funcDecl.namedChild(i);
if (child?.type === 'virtual_specifier' && child.text === keyword) return true;
}
return false;
}
// ---------------------------------------------------------------------------
// C++ config
// ---------------------------------------------------------------------------
// C++ methods appear as field_declaration (declarations) or function_definition
// (inline definitions) inside field_declaration_list. The generic extractor
// iterates bodyNodeTypes children and matches against methodNodeTypes.
//
// Key difference from TS/JVM/C#: C++ has no dedicated method_declaration node.
// A field_declaration is a method if it contains a function_declarator.
// The generic extractor calls extractName() on every methodNodeType node — if
// extractName returns undefined (no function_declarator), the method is skipped.
//
// Known gaps:
// - Out-of-class method definitions (void Foo::bar() {}) are not linked as
// HAS_METHOD — they appear as top-level function_definition nodes.
// This includes namespace-wrapped and nested classes.
// - Friend declarations are not extracted.
// - Template method declarations with explicit specialization.
// - const-qualified method overloads (e.g. begin() vs begin() const) are
// disambiguated via isConst flag and $const ID suffix.
export const cppMethodConfig: MethodExtractionConfig = {
language: SupportedLanguages.CPlusPlus,
typeDeclarationNodes: ['class_specifier', 'struct_specifier', 'union_specifier'],
// declaration covers constructors/destructors; field_declaration covers method
// declarations; function_definition covers inline method definitions.
// Non-method declarations (variables, typedefs) are filtered by extractName
// returning undefined when no function_declarator is found.
methodNodeTypes: ['field_declaration', 'function_definition', 'declaration'],
bodyNodeTypes: ['field_declaration_list'],
extractName: extractCppMethodName,
extractReturnType: extractCppReturnType,
extractParameters: extractCppParameters,
extractVisibility: extractCppVisibility,
isStatic(node) {
return hasKeyword(node, 'static');
},
isAbstract(node) {
return isPureVirtual(node);
},
isFinal(node) {
return hasVirtualSpecifier(node, 'final');
},
isVirtual(node) {
// In C++, override and method-level final are only legal on virtual functions,
// so they imply virtual even without the explicit keyword.
return (
hasKeyword(node, 'virtual') ||
hasVirtualSpecifier(node, 'override') ||
hasVirtualSpecifier(node, 'final')
);
},
isOverride(node) {
return hasVirtualSpecifier(node, 'override');
},
isConst(node) {
// const qualifier appears as a type_qualifier child of function_declarator,
// after the parameter_list: e.g. `int size() const` → funcDecl has
// type_qualifier child with text "const". Not to be confused with return-type
// const (e.g. `const int& begin()`) which is at a different AST level.
const funcDecl = findFunctionDeclarator(node);
if (!funcDecl) return false;
for (let i = 0; i < funcDecl.namedChildCount; i++) {
const child = funcDecl.namedChild(i);
if (child?.type === 'type_qualifier' && child.text === 'const') return true;
}
return false;
},
};
// ---------------------------------------------------------------------------
// C config (minimal — C has no classes/methods, only struct function pointers)
// Verified against tree-sitter-c 0.23.2
// ---------------------------------------------------------------------------
// C does not have methods in the OOP sense. Structs with function pointer fields
// are handled by the field extractor. This config exists for completeness but
// will rarely match since C structs don't contain function_definition nodes.
export const cMethodConfig: MethodExtractionConfig = {
language: SupportedLanguages.C,
typeDeclarationNodes: ['struct_specifier'],
methodNodeTypes: ['function_definition'],
bodyNodeTypes: ['field_declaration_list'],
extractName: extractCppMethodName,
extractReturnType: extractCppReturnType,
extractParameters: extractCppParameters,
extractVisibility() {
return 'public'; // C has no access control
},
isStatic(node) {
return hasKeyword(node, 'static');
},
isAbstract() {
return false; // C has no virtual/abstract
},
isFinal() {
return false;
},
};
@@ -1,4 +1,5 @@
// gitnexus/src/core/ingestion/method-extractors/configs/csharp.ts
// Verified against tree-sitter-c-sharp 0.23.1
import { SupportedLanguages } from 'gitnexus-shared';
import type {
@@ -85,6 +86,7 @@ function extractParametersFromList(paramList: SyntaxNode): ParameterInfo[] {
type: typeNode
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: false,
isVariadic: true,
});
@@ -126,6 +128,7 @@ function extractParametersFromList(paramList: SyntaxNode): ParameterInfo[] {
params.push({
name: nameNode.text,
type: typeName,
rawType: typeNode?.text?.trim() ?? null,
isOptional,
isVariadic: false,
});
@@ -186,6 +189,7 @@ export const csharpMethodConfig: MethodExtractionConfig = {
'destructor_declaration',
'operator_declaration',
'conversion_operator_declaration',
'local_function_statement',
],
bodyNodeTypes: ['declaration_list'],
@@ -221,7 +225,7 @@ export const csharpMethodConfig: MethodExtractionConfig = {
// Constructors and destructors have no return type
// operator_declaration and conversion_operator_declaration use 'type' field, not 'returns'
const returnsNode = node.childForFieldName('returns');
if (returnsNode) return extractSimpleTypeName(returnsNode) ?? returnsNode.text?.trim();
if (returnsNode) return returnsNode.text?.trim();
// Fallback for operator/conversion declarations that use 'type' as return type field
if (node.type === 'operator_declaration' || node.type === 'conversion_operator_declaration') {
const typeNode = node.childForFieldName('type');
@@ -0,0 +1,409 @@
// gitnexus/src/core/ingestion/method-extractors/configs/dart.ts
// Verified against tree-sitter-dart 1.0.0 (80e23c07)
import { SupportedLanguages } from 'gitnexus-shared';
import type {
MethodExtractionConfig,
ParameterInfo,
MethodVisibility,
} from '../../method-types.js';
import { extractSimpleTypeName } from '../../type-extractors/shared.js';
import type { SyntaxNode } from '../../utils/ast-helpers.js';
// ---------------------------------------------------------------------------
// Dart helpers
// ---------------------------------------------------------------------------
/** Type node types that represent a return type in function/getter/setter signatures. */
const TYPE_NODE_TYPES = new Set([
'type_identifier',
'generic_type',
'function_type',
'nullable_type',
'void_type',
'record_type',
]);
/**
* Dart method_signature is a WRAPPER node containing one inner signature:
* function_signature, constructor_signature, getter_signature, setter_signature,
* operator_signature, or factory_constructor_signature.
*
* Name, parameters, and return type live on the INNER signature, not on
* method_signature itself.
*/
function getInnerSignature(node: SyntaxNode): SyntaxNode | null {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (
child &&
(child.type === 'function_signature' ||
child.type === 'constructor_signature' ||
child.type === 'getter_signature' ||
child.type === 'setter_signature' ||
child.type === 'operator_signature' ||
child.type === 'factory_constructor_signature')
) {
return child;
}
}
// `declaration` nodes (abstract methods) also wrap function_signature as a
// named child — handled by the loop above.
return null;
}
/**
* Extract the method name from a method_signature node.
*
* Descends into the inner signature to find the name field/identifier.
*/
function extractDartName(node: SyntaxNode): string | undefined {
const inner = getInnerSignature(node);
if (!inner) return undefined;
// constructor_signature name field may include "ClassName.namedCtor" via multiple children.
// getter_signature, setter_signature, function_signature all have a 'name' field.
if (inner.type === 'operator_signature') {
// operator_signature has no 'name' field; name is 'operator' + the operator symbol
for (let i = 0; i < inner.namedChildCount; i++) {
const child = inner.namedChild(i);
if (child?.type === 'binary_operator') {
return `operator ${child.text.trim()}`;
}
}
// Check for unnamed operator tokens like []= or ~
for (let i = 0; i < inner.childCount; i++) {
const child = inner.child(i);
if (child && !child.isNamed && child.text.trim() !== 'operator') {
const text = child.text.trim();
if (text && !TYPE_NODE_TYPES.has(child.type)) {
return `operator ${text}`;
}
}
}
return undefined;
}
if (inner.type === 'getter_signature') {
const nameNode = inner.childForFieldName('name');
return nameNode?.text;
}
if (inner.type === 'setter_signature') {
const nameNode = inner.childForFieldName('name');
return nameNode ? `set ${nameNode.text}` : undefined;
}
if (inner.type === 'factory_constructor_signature') {
// Collect all identifier children to form "ClassName" or "ClassName.named"
const parts: string[] = [];
for (let i = 0; i < inner.childCount; i++) {
const child = inner.child(i);
if (child?.isNamed && child.type === 'identifier') {
parts.push(child.text);
}
}
return parts.length > 0 ? parts.join('.') : undefined;
}
// function_signature and constructor_signature both have a 'name' field
const nameNode = inner.childForFieldName('name');
if (nameNode) {
// constructor_signature: name field may be multiple identifiers joined by '.'
return nameNode.text;
}
return undefined;
}
/**
* Extract the return type from the inner signature.
*
* function_signature children include type nodes before the name.
* getter_signature children include type nodes before 'get' keyword.
* constructor/setter signatures have no return type.
*/
function extractDartReturnType(node: SyntaxNode): string | undefined {
const inner = getInnerSignature(node);
if (!inner) return undefined;
// Constructors and setters have no return type
if (
inner.type === 'constructor_signature' ||
inner.type === 'setter_signature' ||
inner.type === 'factory_constructor_signature'
) {
return undefined;
}
// For function_signature, getter_signature, operator_signature:
// The type node is a named child before the name/operator
for (let i = 0; i < inner.namedChildCount; i++) {
const child = inner.namedChild(i);
if (child && TYPE_NODE_TYPES.has(child.type)) {
return child.text?.trim();
}
}
return undefined;
}
/**
* Extract parameters from the inner signature's formal_parameter_list.
*
* Dart parameters can be:
* - Positional required: `int x`
* - Optional positional: `[int? x]` — wrapped in optional_formal_parameters with '['
* - Optional named: `{int? x}` or `{required int x}` — wrapped in optional_formal_parameters with '{'
*/
function extractDartParameters(node: SyntaxNode): ParameterInfo[] {
const inner = getInnerSignature(node);
if (!inner) return [];
// getter_signature has no parameters
if (inner.type === 'getter_signature') return [];
// Find formal_parameter_list — it's a child, not a field in function_signature
let paramList: SyntaxNode | null = null;
if (inner.type === 'constructor_signature' || inner.type === 'factory_constructor_signature') {
paramList = inner.childForFieldName('parameters');
}
if (!paramList) {
for (let i = 0; i < inner.namedChildCount; i++) {
const child = inner.namedChild(i);
if (child?.type === 'formal_parameter_list') {
paramList = child;
break;
}
}
}
if (!paramList) return [];
return extractParamsFromList(paramList, false);
}
/**
* Extract ParameterInfo entries from a formal_parameter_list or optional_formal_parameters node.
*/
function extractParamsFromList(listNode: SyntaxNode, isOptionalBlock: boolean): ParameterInfo[] {
const params: ParameterInfo[] = [];
for (let i = 0; i < listNode.namedChildCount; i++) {
const child = listNode.namedChild(i);
if (!child) continue;
if (child.type === 'formal_parameter') {
params.push(extractSingleParam(child, isOptionalBlock));
} else if (child.type === 'optional_formal_parameters') {
// Determine if these are named ({}) or positional ([]) optional params
// by checking the surrounding delimiters
params.push(...extractParamsFromList(child, true));
}
}
return params;
}
/**
* Extract a single ParameterInfo from a formal_parameter node.
*/
function extractSingleParam(param: SyntaxNode, isOptionalBlock: boolean): ParameterInfo {
const nameNode = param.childForFieldName('name');
const name = nameNode?.text ?? '<unknown>';
// Find the type node
let typeName: string | null = null;
let rawTypeName: string | null = null;
for (let i = 0; i < param.namedChildCount; i++) {
const child = param.namedChild(i);
if (child && TYPE_NODE_TYPES.has(child.type)) {
rawTypeName = child.text?.trim() ?? null;
typeName = extractSimpleTypeName(child) ?? rawTypeName;
break;
}
// Also check type_identifier
if (child?.type === 'type_identifier') {
rawTypeName = child.text?.trim() ?? null;
typeName = rawTypeName;
break;
}
}
// Check for 'required' keyword:
// 1. Among children of the param node itself
let hasRequired = false;
for (let i = 0; i < param.childCount; i++) {
const child = param.child(i);
if (child && child.text.trim() === 'required') {
hasRequired = true;
break;
}
}
// 2. In tree-sitter-dart, `required` may be an anonymous sibling token
// immediately preceding the formal_parameter inside optional_formal_parameters.
if (!hasRequired) {
let prev = param.previousSibling;
// Skip comma separators
while (prev && !prev.isNamed && prev.text.trim() === ',') {
prev = prev.previousSibling;
}
if (prev && !prev.isNamed && prev.text.trim() === 'required') {
hasRequired = true;
}
}
// A parameter is optional if it's inside an optional_formal_parameters block
// and does NOT have the 'required' keyword
const isOptional = isOptionalBlock && !hasRequired;
return {
name,
type: typeName,
rawType: rawTypeName,
isOptional,
isVariadic: false, // Dart has no variadic params
};
}
/**
* Dart visibility: underscore prefix = private, else public.
*
* We resolve the name by descending into the inner signature.
*/
function extractDartVisibility(node: SyntaxNode): MethodVisibility {
const name = extractDartName(node);
if (!name) return 'public';
// Strip 'set ' or 'operator ' prefix to get the raw name
const rawName = name.startsWith('set ')
? name.slice(4)
: name.startsWith('operator ')
? name.slice(9)
: name;
return rawName.startsWith('_') ? 'private' : 'public';
}
/**
* In tree-sitter-dart, `static` is an anonymous child token of
* `method_signature` (or `declaration`), not a previous sibling.
*
* We check children first, then fall back to previous siblings for
* grammar variants.
*/
function isDartStatic(node: SyntaxNode): boolean {
// In tree-sitter-dart, `static` is an anonymous child token of method_signature
// (or declaration), not a previous sibling.
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (child && !child.isNamed && child.text.trim() === 'static') return true;
// Stop once we hit the inner signature — static always precedes it
if (child?.isNamed) break;
}
// Also check previous siblings (fallback for grammar variants)
let sibling = node.previousSibling;
while (sibling) {
if (sibling.isNamed && sibling.type !== 'annotation') break;
if (!sibling.isNamed && sibling.text.trim() === 'static') return true;
sibling = sibling.previousSibling;
}
return false;
}
/**
* A Dart method is abstract if it has no function_body sibling following it.
* In the tree-sitter grammar, function_body is a sibling of method_signature
* in class_body.
*/
function isDartAbstract(node: SyntaxNode, _ownerNode: SyntaxNode): boolean {
// `declaration` nodes in class_body represent abstract methods (no body, followed by ';').
// Note: extension bodies cannot have abstract members in Dart, but `declaration` nodes
// do not appear in extension_body in practice since extensions must provide implementations.
if (node.type === 'declaration') return true;
// For method_signature nodes, check if the next named sibling is a function_body
const next = node.nextNamedSibling;
return !next || next.type !== 'function_body';
}
/**
* Check for `async`, `async*`, or `sync*` keyword in the function_body sibling.
* The keyword appears as an unnamed child of function_body, or
* as a sibling keyword before function_body.
*
* Dart has three async-like forms: `async` (Future), `async*` (Stream), `sync*` (Iterable).
* All three are treated as async for graph purposes.
*/
function isDartAsync(node: SyntaxNode): boolean {
let sibling: SyntaxNode | null = node.nextSibling;
let limit = 3;
while (sibling && limit > 0) {
if (!sibling.isNamed) {
const text = sibling.text.trim();
if (text === 'async' || text === 'async*' || text === 'sync*') return true;
}
if (sibling.isNamed && sibling.type === 'function_body') {
// Check first child of function_body for async/async*/sync*
for (let i = 0; i < sibling.childCount; i++) {
const child = sibling.child(i);
if (child) {
const text = child.text.trim();
if (text === 'async' || text === 'async*' || text === 'sync*') return true;
}
// Stop at first substantial child
if (child?.isNamed) break;
}
break;
}
sibling = sibling.nextSibling;
limit--;
}
return false;
}
/**
* Extract annotations that appear as sibling nodes before the method_signature
* in class_body. Each annotation node is prefixed with '@'.
*/
function extractDartAnnotations(node: SyntaxNode): string[] {
const annotations: string[] = [];
let sibling = node.previousNamedSibling;
while (sibling && sibling.type === 'annotation') {
// annotation node text already includes '@', e.g. "@override"
const text = sibling.text?.trim();
if (text) {
// Normalize: strip arguments from annotation if present, keep just the name
// e.g. "@deprecated" -> "@deprecated", "@JsonKey(name: 'id')" -> "@JsonKey"
const match = text.match(/^@(\w+)/);
if (match) {
annotations.unshift('@' + match[1]);
} else {
annotations.unshift(text);
}
}
sibling = sibling.previousNamedSibling;
}
return annotations;
}
// ---------------------------------------------------------------------------
// Dart config
// ---------------------------------------------------------------------------
export const dartMethodConfig: MethodExtractionConfig = {
language: SupportedLanguages.Dart,
typeDeclarationNodes: ['class_definition', 'mixin_declaration', 'extension_declaration'],
methodNodeTypes: ['method_signature', 'declaration'],
bodyNodeTypes: ['class_body', 'extension_body'],
extractName: extractDartName,
extractReturnType: extractDartReturnType,
extractParameters: extractDartParameters,
extractVisibility: extractDartVisibility,
isStatic: isDartStatic,
isAbstract: isDartAbstract,
isFinal: () => false, // Dart methods cannot be 'final'
isAsync: isDartAsync,
extractAnnotations: extractDartAnnotations,
};
@@ -0,0 +1,195 @@
// gitnexus/src/core/ingestion/method-extractors/configs/go.ts
// Verified against tree-sitter-go 0.23.4
import { SupportedLanguages } from 'gitnexus-shared';
import type {
MethodExtractionConfig,
ParameterInfo,
MethodVisibility,
} from '../../method-types.js';
import { extractSimpleTypeName } from '../../type-extractors/shared.js';
import type { SyntaxNode } from '../../utils/ast-helpers.js';
// ---------------------------------------------------------------------------
// Go helpers
// ---------------------------------------------------------------------------
/**
* Extract the method/function name.
* - method_declaration: name is a `field_identifier`
* - function_declaration: name is an `identifier`
*/
function extractGoName(node: SyntaxNode): string | undefined {
const nameNode = node.childForFieldName('name');
return nameNode?.text;
}
/**
* Extract return type from the `result` field.
*
* Go supports single return (`int`) and multi-return (`(User, error)`).
* Multi-return appears as a `parameter_list` — extract the first type.
*/
function extractGoReturnType(node: SyntaxNode): string | undefined {
const result = node.childForFieldName('result');
if (!result) return undefined;
// Single return type (type_identifier, pointer_type, etc.)
if (result.type !== 'parameter_list') {
return result.text?.trim();
}
// Multi-return: (Type, error) — extract first parameter's type
for (let i = 0; i < result.namedChildCount; i++) {
const param = result.namedChild(i);
if (param?.type === 'parameter_declaration') {
const typeNode = param.childForFieldName('type');
if (typeNode) return typeNode.text?.trim();
}
}
return undefined;
}
/**
* Extract parameters from the `parameters` field.
*
* Go parameter_list contains parameter_declaration nodes with optional
* `name` and required `type` fields. Go allows multiple names for one type:
* `func(a, b int)` — each name shares the type.
*
* Handles variadic_parameter_declaration (`...string`).
*/
function extractGoParameters(node: SyntaxNode): ParameterInfo[] {
const paramList = node.childForFieldName('parameters');
if (!paramList) return [];
const params: ParameterInfo[] = [];
for (let i = 0; i < paramList.namedChildCount; i++) {
const param = paramList.namedChild(i);
if (!param) continue;
if (param.type === 'parameter_declaration') {
const typeNode = param.childForFieldName('type');
const typeName = typeNode
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null;
// Go allows multiple names for one type: func(a, b int)
const names: string[] = [];
for (let j = 0; j < param.namedChildCount; j++) {
const child = param.namedChild(j);
if (child?.type === 'identifier') {
names.push(child.text);
}
}
const rawType = typeNode?.text?.trim() ?? null;
if (names.length === 0) {
// Unnamed parameter: func(int, string)
params.push({
name: `_${i}`,
type: typeName,
rawType,
isOptional: false,
isVariadic: false,
});
} else {
for (const name of names) {
params.push({ name, type: typeName, rawType, isOptional: false, isVariadic: false });
}
}
} else if (param.type === 'variadic_parameter_declaration') {
const nameNode = param.childForFieldName('name');
const typeNode = param.childForFieldName('type');
const typeName = typeNode
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null;
params.push({
name: nameNode?.text ?? `_${i}`,
type: typeName,
rawType: typeNode?.text?.trim() ?? null,
isOptional: false,
isVariadic: true,
});
}
}
return params;
}
/**
* Go visibility: uppercase first character = exported (public), lowercase = unexported (private).
*/
function extractGoVisibility(node: SyntaxNode): MethodVisibility {
const name = extractGoName(node);
if (!name || name.length === 0) return 'private';
const first = name[0];
return first === first.toUpperCase() && first !== first.toLowerCase() ? 'public' : 'private';
}
/**
* Extract receiver type from the `receiver` field.
*
* The receiver is a parameter_list with one parameter_declaration:
* (r *Repo) → pointer_type → type_identifier "Repo"
* (r Repo) → type_identifier "Repo"
*/
function extractGoReceiverType(node: SyntaxNode): string | undefined {
const receiver = node.childForFieldName('receiver');
if (!receiver) return undefined;
for (let i = 0; i < receiver.namedChildCount; i++) {
const param = receiver.namedChild(i);
if (param?.type === 'parameter_declaration') {
const typeNode = param.childForFieldName('type');
if (!typeNode) continue;
// Unwrap pointer_type: *User → User
const inner = typeNode.type === 'pointer_type' ? typeNode.firstNamedChild : typeNode;
return inner?.text;
}
}
return undefined;
}
/**
* Resolve owner name from the receiver type.
* For function_declaration (no receiver), returns undefined.
*/
function extractGoOwnerName(node: SyntaxNode): string | undefined {
return extractGoReceiverType(node);
}
// ---------------------------------------------------------------------------
// Config
// ---------------------------------------------------------------------------
export const goMethodConfig: MethodExtractionConfig = {
language: SupportedLanguages.Go,
// Each method_declaration/function_declaration is treated as its own "container"
// for extractFromNode() — not used with extract() in the traditional sense.
// method_elem covers interface method signatures (abstract methods).
typeDeclarationNodes: ['method_declaration', 'function_declaration', 'method_elem'],
methodNodeTypes: ['method_declaration', 'function_declaration', 'method_elem'],
bodyNodeTypes: [],
extractName: extractGoName,
extractReturnType: extractGoReturnType,
extractParameters: extractGoParameters,
extractVisibility: extractGoVisibility,
extractReceiverType: extractGoReceiverType,
extractOwnerName: extractGoOwnerName,
isStatic(node) {
// Go functions (no receiver) are effectively static
return node.type === 'function_declaration';
},
isAbstract(node, _ownerNode) {
// Go interface method signatures (method_elem) are abstract — no body
return node.type === 'method_elem';
},
isFinal(_node) {
return false; // Go has no final methods
},
};
@@ -19,7 +19,9 @@ const INTERFACE_OWNER_TYPES = new Set(['interface_declaration', 'annotation_type
function extractReturnTypeFromField(node: SyntaxNode): string | undefined {
const typeNode = node.childForFieldName('type');
if (!typeNode) return undefined;
return extractSimpleTypeName(typeNode) ?? typeNode.text?.trim();
// Use .text to preserve full generic types (e.g. List<User>, Stream<T>)
// needed by the call resolver for return-type inference.
return typeNode.text?.trim();
}
function extractAnnotations(node: SyntaxNode, modifierType: string): string[] {
@@ -68,6 +70,7 @@ function extractJavaParameters(node: SyntaxNode): ParameterInfo[] {
params.push({
name: nameNode.text,
type: typeNode ? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim()) : null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: false,
isVariadic: false,
});
@@ -76,6 +79,7 @@ function extractJavaParameters(node: SyntaxNode): ParameterInfo[] {
// Varargs: type_identifier + "..." + variable_declarator
let paramName: string | undefined;
let paramType: string | null = null;
let paramRawType: string | null = null;
for (let j = 0; j < param.namedChildCount; j++) {
const c = param.namedChild(j);
if (!c) continue;
@@ -90,13 +94,15 @@ function extractJavaParameters(node: SyntaxNode): ParameterInfo[] {
c.type === 'floating_point_type' ||
c.type === 'boolean_type'
) {
paramType = extractSimpleTypeName(c) ?? c.text?.trim();
paramRawType = c.text?.trim() ?? null;
paramType = extractSimpleTypeName(c) ?? paramRawType;
}
}
if (paramName) {
params.push({
name: paramName,
type: paramType,
rawType: paramRawType,
isOptional: false,
isVariadic: true,
});
@@ -192,6 +198,7 @@ function extractKotlinParameters(node: SyntaxNode): ParameterInfo[] {
let paramName: string | undefined;
let paramType: string | null = null;
let paramRawType: string | null = null;
let hasDefault = false;
const isVariadic = nextIsVariadic;
nextIsVariadic = false;
@@ -206,7 +213,8 @@ function extractKotlinParameters(node: SyntaxNode): ParameterInfo[] {
part.type === 'nullable_type' ||
part.type === 'function_type'
) {
paramType = extractSimpleTypeName(part) ?? part.text?.trim();
paramRawType = part.text?.trim() ?? null;
paramType = extractSimpleTypeName(part) ?? paramRawType;
}
}
@@ -223,6 +231,7 @@ function extractKotlinParameters(node: SyntaxNode): ParameterInfo[] {
params.push({
name: paramName,
type: paramType,
rawType: paramRawType,
isOptional: hasDefault,
isVariadic: isVariadic,
});
@@ -252,7 +261,7 @@ function extractKotlinReturnType(node: SyntaxNode): string | undefined {
child.type === 'nullable_type' ||
child.type === 'function_type')
) {
return extractSimpleTypeName(child) ?? child.text?.trim();
return child.text?.trim();
}
if (child.type === 'function_body') break;
}
@@ -0,0 +1,326 @@
// gitnexus/src/core/ingestion/method-extractors/configs/php.ts
// Verified against tree-sitter-php 0.23.12
import { SupportedLanguages } from 'gitnexus-shared';
import type {
MethodExtractionConfig,
ParameterInfo,
MethodVisibility,
} from '../../method-types.js';
import { extractSimpleTypeName } from '../../type-extractors/shared.js';
import type { SyntaxNode } from '../../utils/ast-helpers.js';
// ---------------------------------------------------------------------------
// PHP helpers
// ---------------------------------------------------------------------------
/** Regex to extract PHPDoc @return annotations: `@return User` */
const PHPDOC_RETURN_RE = /@return\s+(\S+)/;
/** Node types to skip when walking backwards through siblings for PHPDoc. */
const PHPDOC_SKIP_NODE_TYPES: ReadonlySet<string> = new Set(['attribute_list', 'attribute']);
/**
* Normalize a PHPDoc return type for the MethodExtractor.
* Strips nullable prefix, null/false/void unions, namespace prefixes, and
* rejects uninformative types (mixed, void, self, static, object, array).
*/
function normalizePhpReturnType(raw: string): string | undefined {
let type = raw.startsWith('?') ? raw.slice(1) : raw;
const parts = type
.split('|')
.filter((p) => p !== 'null' && p !== 'false' && p !== 'void' && p !== 'mixed');
if (parts.length !== 1) return undefined;
type = parts[0];
const segments = type.split('\\');
type = segments[segments.length - 1];
if (
type === 'mixed' ||
type === 'void' ||
type === 'self' ||
type === 'static' ||
type === 'object' ||
type === 'array'
)
return undefined;
if (/^\w+(\[\])?$/.test(type) || /^\w+\s*</.test(type)) return type;
return undefined;
}
/**
* Walk backwards through preceding siblings of `node` to find a PHPDoc
* `@return Type` annotation. Skips `attribute_list` nodes (PHP 8 attributes).
*/
function extractPhpDocReturnType(node: SyntaxNode): string | undefined {
let sibling = node.previousSibling;
while (sibling) {
if (sibling.type === 'comment') {
const match = PHPDOC_RETURN_RE.exec(sibling.text);
if (match) return normalizePhpReturnType(match[1]);
} else if (sibling.isNamed && !PHPDOC_SKIP_NODE_TYPES.has(sibling.type)) {
break;
}
sibling = sibling.previousSibling;
}
return undefined;
}
const PHP_VIS = new Set<MethodVisibility>(['public', 'private', 'protected']);
/**
* Find the visibility keyword from a visibility_modifier named child.
* PHP tree-sitter emits `visibility_modifier` as a named node with text
* "public", "private", or "protected".
*/
function findPhpVisibility(node: SyntaxNode): MethodVisibility {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === 'visibility_modifier') {
const text = child.text.trim() as MethodVisibility;
if (PHP_VIS.has(text)) return text;
}
}
return 'public'; // PHP methods are public by default
}
/**
* Check for a named modifier child of a specific type.
* PHP tree-sitter uses distinct node types: abstract_modifier, final_modifier,
* static_modifier — rather than a wrapper `modifiers` node with keyword children.
*/
function hasModifierNode(node: SyntaxNode, modifierType: string): boolean {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === modifierType) return true;
}
return false;
}
/**
* Extract the return type from a PHP method_declaration node.
*
* In tree-sitter-php, the return type is not exposed via a named field.
* It appears as a type node (primitive_type, named_type, union_type,
* optional_type, nullable_type, intersection_type) after the formal_parameters
* and a `:` token separator.
*
* When the AST return type is missing or uninformative (`array` / `iterable`),
* falls back to parsing PHPDoc `@return Type` from preceding doc comments.
*/
function extractPhpReturnType(node: SyntaxNode): string | undefined {
const TYPE_NODE_TYPES = new Set([
'primitive_type',
'named_type',
'union_type',
'optional_type',
'nullable_type',
'intersection_type',
]);
let astType: string | undefined;
let seenParams = false;
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (!child) continue;
if (child.type === 'formal_parameters') {
seenParams = true;
continue;
}
// After the parameters node, look for the colon and then the type
if (seenParams && child.isNamed && TYPE_NODE_TYPES.has(child.type)) {
astType = child.text?.trim();
break;
}
// Stop at body or semicolon
if (child.type === 'compound_statement' || (!child.isNamed && child.text === ';')) {
break;
}
}
// If AST type is missing or uninformative, try PHPDoc @return fallback
if (!astType || astType === 'array' || astType === 'iterable') {
const docType = extractPhpDocReturnType(node);
if (docType) return docType;
}
return astType;
}
/**
* Extract parameters from a PHP method_declaration node.
*
* PHP parameter types in tree-sitter-php:
* - `simple_parameter`: regular parameter with optional type and default
* - `variadic_parameter`: `...$param` with optional type
* - `property_promotion_parameter`: constructor promotion `private string $name`
* (may also be variadic via an ERROR node containing `...`)
*/
function extractPhpParameters(node: SyntaxNode): ParameterInfo[] {
const paramList = node.childForFieldName('parameters');
if (!paramList) return [];
const params: ParameterInfo[] = [];
for (let i = 0; i < paramList.namedChildCount; i++) {
const param = paramList.namedChild(i);
if (!param) continue;
if (param.type === 'simple_parameter') {
const nameNode = param.childForFieldName('name');
if (!nameNode) continue;
const typeNode = param.childForFieldName('type');
const typeName = typeNode ? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim()) : null;
// Detect optional: '=' token among children indicates a default value
let isOptional = false;
for (let j = 0; j < param.childCount; j++) {
const c = param.child(j);
if (c && !c.isNamed && c.text === '=') {
isOptional = true;
break;
}
}
params.push({
name: stripDollar(nameNode.text),
type: typeName ?? null,
rawType: typeNode?.text?.trim() ?? null,
isOptional,
isVariadic: false,
});
} else if (param.type === 'variadic_parameter') {
const nameNode = param.childForFieldName('name');
if (!nameNode) continue;
const typeNode = param.childForFieldName('type');
const typeName = typeNode ? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim()) : null;
params.push({
name: stripDollar(nameNode.text),
type: typeName ?? null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: false,
isVariadic: true,
});
} else if (param.type === 'property_promotion_parameter') {
const nameNode = param.childForFieldName('name');
if (!nameNode) continue;
const typeNode = param.childForFieldName('type');
const typeName = typeNode ? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim()) : null;
// Detect variadic: an ERROR child containing "..." indicates variadic promotion
let isVariadic = false;
for (let j = 0; j < param.childCount; j++) {
const c = param.child(j);
if (c && (c.text === '...' || (c.type === 'ERROR' && c.text === '...'))) {
isVariadic = true;
break;
}
}
params.push({
name: stripDollar(nameNode.text),
type: typeName ?? null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: false,
isVariadic,
});
}
}
return params;
}
/** Strip leading $ from PHP variable names. */
function stripDollar(name: string): string {
return name.startsWith('$') ? name.slice(1) : name;
}
/**
* Extract PHP 8 attributes (#[...]) from a method_declaration node.
*
* AST structure: attribute_list → attribute_group → attribute → name child.
* Names are prefixed with '#' to distinguish from Java/Kotlin @ annotations.
*/
function extractPhpAnnotations(node: SyntaxNode): string[] {
const annotations: string[] = [];
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (!child || child.type !== 'attribute_list') continue;
for (let j = 0; j < child.namedChildCount; j++) {
const group = child.namedChild(j);
if (!group || group.type !== 'attribute_group') continue;
for (let k = 0; k < group.namedChildCount; k++) {
const attr = group.namedChild(k);
if (!attr || attr.type !== 'attribute') continue;
const nameNode = attr.firstNamedChild;
if (nameNode && nameNode.type === 'name') {
annotations.push('#' + nameNode.text);
}
}
}
}
return annotations;
}
// ---------------------------------------------------------------------------
// PHP config
// ---------------------------------------------------------------------------
export const phpMethodConfig: MethodExtractionConfig = {
language: SupportedLanguages.PHP,
typeDeclarationNodes: [
'class_declaration',
'interface_declaration',
'trait_declaration',
'enum_declaration',
],
methodNodeTypes: ['method_declaration', 'function_definition'],
bodyNodeTypes: ['declaration_list'],
extractName(node) {
return node.childForFieldName('name')?.text;
},
extractReturnType: extractPhpReturnType,
extractParameters: extractPhpParameters,
extractVisibility: findPhpVisibility,
isStatic(node) {
return hasModifierNode(node, 'static_modifier');
},
isAbstract(node, ownerNode) {
if (hasModifierNode(node, 'abstract_modifier')) return true;
// Interface methods are implicitly abstract when they have no body.
// Check ownerNode first, then fall back to walking the parent chain
// (needed when called from extractFromNode where ownerNode === node).
let isInterface = ownerNode.type === 'interface_declaration';
if (!isInterface) {
let p = node.parent;
while (p) {
if (p.type === 'interface_declaration') {
isInterface = true;
break;
}
p = p.parent;
}
}
if (isInterface) {
const body = node.childForFieldName('body');
if (body) return false;
for (let i = 0; i < node.namedChildCount; i++) {
if (node.namedChild(i)?.type === 'compound_statement') return false;
}
return true;
}
return false;
},
isFinal(node) {
return hasModifierNode(node, 'final_modifier');
},
extractAnnotations: extractPhpAnnotations,
};
@@ -0,0 +1,325 @@
// gitnexus/src/core/ingestion/method-extractors/configs/python.ts
// Verified against tree-sitter-python 0.23.4
import { SupportedLanguages } from 'gitnexus-shared';
import type {
MethodExtractionConfig,
ParameterInfo,
MethodVisibility,
} from '../../method-types.js';
import { hasKeyword } from '../../field-extractors/configs/helpers.js';
import { extractSimpleTypeName } from '../../type-extractors/shared.js';
import type { SyntaxNode } from '../../utils/ast-helpers.js';
// ---------------------------------------------------------------------------
// Python helpers
// ---------------------------------------------------------------------------
/** Names that represent the instance/class receiver — not real parameters. */
const SELF_NAMES = new Set(['self', 'cls']);
/**
* Unwrap a decorated_definition to its inner function_definition.
*
* tree-sitter-python wraps decorated functions/methods in a `decorated_definition`
* node that contains the decorators as children followed by the function_definition.
* This is different from TS/JS where decorators are siblings.
*/
function unwrapDecorated(node: SyntaxNode): SyntaxNode {
if (node.type === 'decorated_definition') {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child && child.type === 'function_definition') return child;
}
}
return node;
}
/**
* Collect decorator names from a decorated_definition wrapper.
*
* Returns decorator names prefixed with '@'. If the node is a plain
* function_definition (no decorators), check if its parent is a
* decorated_definition and collect from there.
*/
function collectDecorators(node: SyntaxNode): SyntaxNode[] {
let wrapper: SyntaxNode | null = null;
if (node.type === 'decorated_definition') {
wrapper = node;
} else if (node.parent?.type === 'decorated_definition') {
wrapper = node.parent;
}
if (!wrapper) return [];
const decorators: SyntaxNode[] = [];
for (let i = 0; i < wrapper.namedChildCount; i++) {
const child = wrapper.namedChild(i);
if (child && child.type === 'decorator') {
decorators.push(child);
}
}
return decorators;
}
function extractDecoratorName(decorator: SyntaxNode): string | undefined {
// decorator > identifier (simple)
// decorator > call > identifier (call-style, e.g. @lru_cache())
// decorator > attribute (dotted, e.g. @abc.abstractmethod)
const expr = decorator.firstNamedChild;
if (!expr) return undefined;
if (expr.type === 'identifier') return '@' + expr.text;
if (expr.type === 'attribute') return '@' + expr.text;
if (expr.type === 'call') {
const fn = expr.childForFieldName('function');
return fn ? '@' + fn.text : undefined;
}
return undefined;
}
function hasDecorator(node: SyntaxNode, name: string): boolean {
const decorators = collectDecorators(node);
for (const dec of decorators) {
const decName = extractDecoratorName(dec);
if (decName === '@' + name || decName?.endsWith('.' + name)) return true;
}
return false;
}
/**
* Extract parameters from a Python function_definition.
*
* Handles: identifier, default_parameter, typed_parameter, typed_default_parameter,
* list_splat_pattern (*args), dictionary_splat_pattern (**kwargs), and typed variants.
* Skips `self` and `cls` first parameters.
*/
function extractPythonParameters(node: SyntaxNode): ParameterInfo[] {
const funcNode = unwrapDecorated(node);
const paramList = funcNode.childForFieldName('parameters');
if (!paramList) return [];
const params: ParameterInfo[] = [];
let isFirst = true;
for (let i = 0; i < paramList.namedChildCount; i++) {
const param = paramList.namedChild(i);
if (!param) continue;
switch (param.type) {
case 'identifier': {
// Bare parameter: `self`, `cls`, or untyped `x`
if (isFirst && SELF_NAMES.has(param.text)) {
isFirst = false;
continue;
}
isFirst = false;
params.push({
name: param.text,
type: null,
rawType: null,
isOptional: false,
isVariadic: false,
});
break;
}
case 'default_parameter': {
// `x = value` — untyped with default
isFirst = false;
const nameNode = param.childForFieldName('name');
if (nameNode) {
params.push({
name: nameNode.text,
type: null,
rawType: null,
isOptional: true,
isVariadic: false,
});
}
break;
}
case 'typed_parameter': {
// `x: int` or `*args: str` or `**kwargs: int`
// The first named child can be identifier, list_splat_pattern, or dictionary_splat_pattern
const inner = param.firstNamedChild;
if (!inner) break;
if (isFirst && inner.type === 'identifier' && SELF_NAMES.has(inner.text)) {
isFirst = false;
continue;
}
isFirst = false;
const typeNode = param.childForFieldName('type');
const typeText = typeNode
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null;
const rawTypeText = typeNode?.text?.trim() ?? null;
if (inner.type === 'list_splat_pattern') {
const nameId = inner.firstNamedChild;
if (nameId) {
params.push({
name: nameId.text,
type: typeText,
rawType: rawTypeText,
isOptional: false,
isVariadic: true,
});
}
} else if (inner.type === 'dictionary_splat_pattern') {
const nameId = inner.firstNamedChild;
if (nameId) {
params.push({
name: nameId.text,
type: typeText,
rawType: rawTypeText,
isOptional: false,
isVariadic: true,
});
}
} else {
params.push({
name: inner.text,
type: typeText,
rawType: rawTypeText,
isOptional: false,
isVariadic: false,
});
}
break;
}
case 'typed_default_parameter': {
// `x: int = 5` — typed with default
isFirst = false;
const nameNode = param.childForFieldName('name');
const typeNode = param.childForFieldName('type');
if (nameNode) {
params.push({
name: nameNode.text,
type: typeNode
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: true,
isVariadic: false,
});
}
break;
}
case 'list_splat_pattern': {
// `*args` (untyped)
isFirst = false;
const nameId = param.firstNamedChild;
if (nameId) {
params.push({
name: nameId.text,
type: null,
rawType: null,
isOptional: false,
isVariadic: true,
});
}
break;
}
case 'dictionary_splat_pattern': {
// `**kwargs` (untyped)
isFirst = false;
const nameId = param.firstNamedChild;
if (nameId) {
params.push({
name: nameId.text,
type: null,
rawType: null,
isOptional: false,
isVariadic: true,
});
}
break;
}
default:
isFirst = false;
break;
}
}
return params;
}
/**
* Extract return type from the `return_type` field.
*
* tree-sitter-python uses a `type` field on function_definition for the return
* type annotation (e.g. `-> str`). The field contains a type node.
*/
function extractPythonReturnType(node: SyntaxNode): string | undefined {
const funcNode = unwrapDecorated(node);
const returnType = funcNode.childForFieldName('return_type');
if (!returnType) return undefined;
// Use .text to preserve full generic types (e.g. list[User], Dict[str, User])
// that the call resolver needs for for-loop iterable and return-type inference.
return returnType.text?.trim();
}
/**
* Extract visibility based on Python name-mangling convention.
* `__name` (not dunder) = private, `_name` = protected, else public.
*/
function extractPythonVisibility(node: SyntaxNode): MethodVisibility {
const funcNode = unwrapDecorated(node);
const nameNode = funcNode.childForFieldName('name');
const name = nameNode?.text;
if (!name) return 'public';
if (name.startsWith('__') && !name.endsWith('__')) return 'private';
if (name.startsWith('_') && !(name.startsWith('__') && name.endsWith('__'))) return 'protected';
return 'public';
}
// ---------------------------------------------------------------------------
// Config
// ---------------------------------------------------------------------------
export const pythonMethodConfig: MethodExtractionConfig = {
language: SupportedLanguages.Python,
typeDeclarationNodes: ['class_definition'],
// Both function_definition and decorated_definition must be listed:
// decorated methods appear as decorated_definition in the class body block,
// while undecorated methods appear as function_definition directly.
methodNodeTypes: ['function_definition', 'decorated_definition'],
bodyNodeTypes: ['block'],
extractName(node) {
const funcNode = unwrapDecorated(node);
const nameNode = funcNode.childForFieldName('name');
return nameNode?.text;
},
extractReturnType: extractPythonReturnType,
extractParameters: extractPythonParameters,
extractVisibility: extractPythonVisibility,
isStatic(node) {
return hasDecorator(node, 'staticmethod') || hasDecorator(node, 'classmethod');
},
isAbstract(node, _ownerNode) {
return hasDecorator(node, 'abstractmethod');
},
isFinal(_node) {
return false; // @typing.final (PEP 591) is captured in annotations; isFinal not modeled
},
extractAnnotations(node) {
const decorators = collectDecorators(node);
const annotations: string[] = [];
for (const dec of decorators) {
const name = extractDecoratorName(dec);
if (name) annotations.push(name);
}
return annotations;
},
isAsync(node) {
const funcNode = unwrapDecorated(node);
return hasKeyword(funcNode, 'async');
},
};
@@ -0,0 +1,303 @@
// gitnexus/src/core/ingestion/method-extractors/configs/ruby.ts
// Verified against tree-sitter-ruby 0.23.1
import { SupportedLanguages } from 'gitnexus-shared';
import type {
MethodExtractionConfig,
ParameterInfo,
MethodVisibility,
} from '../../method-types.js';
import type { SyntaxNode } from '../../utils/ast-helpers.js';
// ---------------------------------------------------------------------------
// Ruby helpers
// ---------------------------------------------------------------------------
const VISIBILITY_MODIFIERS = new Set(['private', 'protected', 'public']);
/** Regex to extract YARD `@return [Type]` annotations from comments. */
const YARD_RETURN_RE = /@return\s+\[([^\]]+)\]/;
/**
* Extract the simple type name from a YARD type string.
* Handles qualified types ("Models::User" -> "User"), generics ("Array<User>"
* -> "Array"), nullable ("String, nil" -> "String"), and rejects ambiguous
* unions ("String, Integer" -> undefined).
*/
function extractYardTypeName(yardType: string): string | undefined {
const trimmed = yardType.trim();
// Bracket-balanced split on commas to handle generics like Hash<Symbol, User>
const parts: string[] = [];
let depth = 0,
start = 0;
for (let i = 0; i < trimmed.length; i++) {
if (trimmed[i] === '<') depth++;
else if (trimmed[i] === '>') depth--;
else if (trimmed[i] === ',' && depth === 0) {
parts.push(trimmed.slice(start, i).trim());
start = i + 1;
}
}
parts.push(trimmed.slice(start).trim());
const filtered = parts.filter((p) => p !== '' && p !== 'nil');
if (filtered.length !== 1) return undefined; // ambiguous union
const typePart = filtered[0];
// Qualified: "Models::User" -> "User"
const segments = typePart.split('::');
const last = segments[segments.length - 1];
// Generic: "Array<User>" -> "Array"
const genericMatch = last.match(/^(\w+)\s*[<{(]/);
if (genericMatch) return genericMatch[1];
// Simple identifier
if (/^\w+$/.test(last)) return last;
return undefined;
}
/**
* Extract visibility for a Ruby method by walking backwards through the
* parent body_statement's named children from the method node's position.
*
* Ruby visibility modifiers (private, protected, public) appear as bare
* `identifier` nodes in the body_statement. The most recent modifier
* before the method determines its visibility. Default is public.
*
* Example AST for:
* class Foo
* private
* def secret; end
* end
*
* body_statement
* identifier "private" ← index 0
* method "def secret" ← index 1
*/
function extractRubyVisibility(node: SyntaxNode): MethodVisibility {
const parent = node.parent;
if (!parent) return 'public';
// Find the index of this method node in the parent's named children
let methodIndex = -1;
for (let i = 0; i < parent.namedChildCount; i++) {
if (parent.namedChild(i) === node) {
methodIndex = i;
break;
}
}
if (methodIndex < 0) return 'public';
// Walk backwards from the method node looking for a visibility modifier
for (let i = methodIndex - 1; i >= 0; i--) {
const sibling = parent.namedChild(i);
if (!sibling) continue;
if (sibling.type === 'identifier' && VISIBILITY_MODIFIERS.has(sibling.text)) {
return sibling.text as MethodVisibility;
}
// module_function makes instance methods private
if (sibling.type === 'identifier' && sibling.text === 'module_function') {
return 'private';
}
}
return 'public';
}
/**
* Extract parameters from a Ruby method's method_parameters node.
*
* Handles: identifier, optional_parameter (default), splat_parameter (*args),
* hash_splat_parameter (**kwargs), block_parameter (&block), keyword_parameter.
*/
function extractRubyParameters(node: SyntaxNode): ParameterInfo[] {
const paramList = node.childForFieldName('parameters');
if (!paramList) return [];
const params: ParameterInfo[] = [];
for (let i = 0; i < paramList.namedChildCount; i++) {
const param = paramList.namedChild(i);
if (!param) continue;
switch (param.type) {
case 'identifier': {
// Plain parameter: def foo(x)
params.push({
name: param.text,
type: null,
rawType: null,
isOptional: false,
isVariadic: false,
});
break;
}
case 'optional_parameter': {
// Default parameter: def foo(x = 10)
const nameNode = param.childForFieldName('name');
if (nameNode) {
params.push({
name: nameNode.text,
type: null,
rawType: null,
isOptional: true,
isVariadic: false,
});
}
break;
}
case 'splat_parameter': {
// Splat: def foo(*args)
const nameNode = param.childForFieldName('name');
if (nameNode) {
params.push({
name: nameNode.text,
type: null,
rawType: null,
isOptional: false,
isVariadic: true,
});
}
break;
}
case 'hash_splat_parameter': {
// Double splat: def foo(**kwargs)
const nameNode = param.childForFieldName('name');
if (nameNode) {
params.push({
name: nameNode.text,
type: null,
rawType: null,
isOptional: false,
isVariadic: true,
});
}
break;
}
case 'block_parameter': {
// Block: def foo(&block)
const nameNode = param.childForFieldName('name');
if (nameNode) {
params.push({
name: nameNode.text,
type: null,
rawType: null,
isOptional: false,
isVariadic: false,
});
}
break;
}
case 'keyword_parameter': {
// Keyword: def foo(name:) or def foo(name: "default")
const nameNode = param.childForFieldName('name');
const valueNode = param.childForFieldName('value');
if (nameNode) {
params.push({
name: nameNode.text,
type: null,
rawType: null,
isOptional: !!valueNode,
isVariadic: false,
});
}
break;
}
}
}
return params;
}
// ---------------------------------------------------------------------------
// Config
// ---------------------------------------------------------------------------
export const rubyMethodConfig: MethodExtractionConfig = {
language: SupportedLanguages.Ruby,
typeDeclarationNodes: ['class', 'module', 'singleton_class'],
methodNodeTypes: ['method', 'singleton_method'],
bodyNodeTypes: ['body_statement'],
extractOwnerName(node) {
// singleton_class (class << self) inherits the enclosing class/module name
if (node.type === 'singleton_class') {
let ancestor = node.parent;
while (ancestor) {
if (ancestor.type === 'class' || ancestor.type === 'module') {
const nameNode = ancestor.childForFieldName('name');
return nameNode?.text;
}
ancestor = ancestor.parent;
}
return undefined;
}
return undefined; // use default resolution for class/module
},
extractName(node) {
const nameNode = node.childForFieldName('name');
return nameNode?.text;
},
extractReturnType(node) {
// Walk backwards through preceding siblings looking for YARD @return [Type].
// Try direct siblings first, then fall back to parent (body_statement) siblings
// for class methods where the comment may be a sibling of the body_statement.
const search = (startNode: SyntaxNode): string | undefined => {
let sibling = startNode.previousSibling;
while (sibling) {
if (sibling.type === 'comment') {
const match = YARD_RETURN_RE.exec(sibling.text);
if (match) return extractYardTypeName(match[1]);
} else if (sibling.isNamed) {
break;
}
sibling = sibling.previousSibling;
}
return undefined;
};
const result = search(node);
if (result) return result;
if (node.parent?.type === 'body_statement') {
return search(node.parent);
}
return undefined;
},
extractParameters: extractRubyParameters,
extractVisibility: extractRubyVisibility,
isStatic(node) {
if (node.type === 'singleton_method') return true;
// module_function makes following methods callable at module level (static)
const parent = node.parent;
if (!parent) return false;
let methodIndex = -1;
for (let i = 0; i < parent.namedChildCount; i++) {
if (parent.namedChild(i) === node) {
methodIndex = i;
break;
}
}
for (let i = methodIndex - 1; i >= 0; i--) {
const sibling = parent.namedChild(i);
if (!sibling) continue;
if (sibling.type === 'identifier' && sibling.text === 'module_function') return true;
// Other visibility modifiers override module_function
if (sibling.type === 'identifier' && VISIBILITY_MODIFIERS.has(sibling.text)) return false;
}
return false;
},
isAbstract(_node, _ownerNode) {
return false; // Ruby has no abstract methods
},
isFinal(_node) {
return false; // Ruby has no final methods
},
};
@@ -0,0 +1,211 @@
// gitnexus/src/core/ingestion/method-extractors/configs/rust.ts
// Verified against tree-sitter-rust 0.23.1
import { SupportedLanguages } from 'gitnexus-shared';
import type {
MethodExtractionConfig,
ParameterInfo,
MethodVisibility,
} from '../../method-types.js';
import { extractSimpleTypeName } from '../../type-extractors/shared.js';
import type { SyntaxNode } from '../../utils/ast-helpers.js';
// ---------------------------------------------------------------------------
// Rust helpers
// ---------------------------------------------------------------------------
/**
* Extract method name from function_item or function_signature_item.
* Both use a `name` field containing an identifier.
*/
function extractRustMethodName(node: SyntaxNode): string | undefined {
const nameNode = node.childForFieldName('name');
return nameNode?.text;
}
/**
* Extract return type from the `return_type` field.
* tree-sitter-rust puts the return type (after `->`) as the `return_type` field.
*/
function extractRustReturnType(node: SyntaxNode): string | undefined {
const typeNode = node.childForFieldName('return_type');
if (!typeNode) return undefined;
return typeNode.text?.trim();
}
/**
* Extract parameters, skipping the self_parameter (handled by extractReceiverType).
*
* Rust parameters use `pattern` and `type` fields:
* parameter { pattern: identifier, type: primitive_type }
*/
function extractRustParameters(node: SyntaxNode): ParameterInfo[] {
const paramList = node.childForFieldName('parameters');
if (!paramList) return [];
const params: ParameterInfo[] = [];
for (let i = 0; i < paramList.namedChildCount; i++) {
const param = paramList.namedChild(i);
if (!param) continue;
// Skip self_parameter — it is the receiver, not a regular parameter
if (param.type === 'self_parameter') continue;
if (param.type === 'parameter') {
const patternNode = param.childForFieldName('pattern');
const typeNode = param.childForFieldName('type');
params.push({
name: patternNode?.text ?? '?',
type: typeNode ? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null) : null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: false,
isVariadic: false,
});
}
}
return params;
}
/**
* Detect visibility from visibility_modifier named child.
* `pub`, `pub(crate)`, `pub(super)`, `pub(in path)` → public.
* Absence → private (Rust default).
*/
function extractRustVisibility(node: SyntaxNode): MethodVisibility {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === 'visibility_modifier') return 'public';
}
return 'private';
}
/**
* Detect receiver type from the first parameter if it is a self_parameter.
*
* Variants:
* - `self` → "self"
* - `&self` → "&self"
* - `&mut self` → "&mut self"
* - `mut self` → "mut self"
* - `self: Box<Self>` → "Box<Self>" (explicit self type)
*/
function extractRustReceiverType(node: SyntaxNode): string | undefined {
const paramList = node.childForFieldName('parameters');
if (!paramList) return undefined;
const first = paramList.namedChild(0);
if (!first || first.type !== 'self_parameter') return undefined;
return first.text;
}
/**
* Check whether a function_item has the `async` keyword.
* tree-sitter-rust wraps it in a `function_modifiers` named child.
*/
function isRustAsync(node: SyntaxNode): boolean {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === 'function_modifiers' && child.text.includes('async')) return true;
}
return false;
}
/**
* Extract attributes from preceding sibling attribute_item nodes.
*
* In tree-sitter-rust, `#[inline]` is an `attribute_item` sibling that precedes
* the function_item in the declaration_list, not a child of the function_item.
*/
function extractRustAnnotations(node: SyntaxNode): string[] {
const annotations: string[] = [];
let sibling = node.previousNamedSibling;
while (sibling) {
if (sibling.type === 'attribute_item') {
annotations.unshift(sibling.text);
} else {
// Stop at the first non-attribute sibling — attributes are contiguous
break;
}
sibling = sibling.previousNamedSibling;
}
return annotations;
}
// ---------------------------------------------------------------------------
// Rust config
// ---------------------------------------------------------------------------
// Rust methods live inside `impl` blocks (concrete implementations) or
// `trait` blocks (trait definitions with required/default methods).
//
// `impl_item` body contains `function_item` nodes for concrete methods.
// `trait_item` body contains `function_item` (default methods) and
// `function_signature_item` (required/abstract methods without a body).
//
// ownerName resolution: `impl_item` has no `name` field — the generic
// extractor falls back to the first `type_identifier` child, which is the
// implementing type (e.g. `impl MyStruct { ... }` → "MyStruct").
// `trait_item` uses the standard `name` field.
//
// Known gaps:
// - Macro-generated methods (e.g. derive) are not visible in the AST.
// - Unsafe methods are not distinguished (no isUnsafe field in schema).
export const rustMethodConfig: MethodExtractionConfig = {
language: SupportedLanguages.Rust,
typeDeclarationNodes: ['impl_item', 'trait_item'],
methodNodeTypes: ['function_item', 'function_signature_item'],
bodyNodeTypes: ['declaration_list'],
// For `impl Trait for Struct`, resolve owner to the concrete Struct (after `for`).
// For plain `impl Struct`, resolve to Struct (first type_identifier).
// For `trait Foo`, let the default name-field resolution handle it.
extractOwnerName(node) {
if (node.type !== 'impl_item') return undefined;
const children = node.children ?? [];
const forIdx = children.findIndex((c: SyntaxNode) => c.text === 'for');
if (forIdx !== -1) {
// impl Trait for Struct — pick the type after `for`
const typeNode = children
.slice(forIdx + 1)
.find(
(c: SyntaxNode) => c.type === 'type_identifier' || c.type === 'scoped_type_identifier',
);
if (typeNode) return typeNode.text;
}
// Plain `impl Struct` — pick the first type_identifier
const first = children.find((c: SyntaxNode) => c.type === 'type_identifier');
return first?.text;
},
extractName: extractRustMethodName,
extractReturnType: extractRustReturnType,
extractParameters: extractRustParameters,
extractVisibility: extractRustVisibility,
isStatic(node) {
// A Rust method is an "associated function" (static) if it lacks a
// self_parameter as first parameter.
const paramList = node.childForFieldName('parameters');
if (!paramList) return true;
const first = paramList.namedChild(0);
return !first || first.type !== 'self_parameter';
},
isAbstract(node, ownerNode) {
// Only trait methods without a body (function_signature_item) are abstract.
// function_signature_item never has a body field.
if (ownerNode.type === 'trait_item' && node.type === 'function_signature_item') {
return true;
}
return false;
},
isFinal() {
// Rust has no `final` concept — all methods are effectively sealed
// (traits cannot be "overridden" the way Java methods can).
return false;
},
extractAnnotations: extractRustAnnotations,
extractReceiverType: extractRustReceiverType,
isAsync: isRustAsync,
};
@@ -0,0 +1,311 @@
// gitnexus/src/core/ingestion/method-extractors/configs/swift.ts
// Verified against tree-sitter-swift 0.6.0
import { SupportedLanguages } from 'gitnexus-shared';
import type {
MethodExtractionConfig,
ParameterInfo,
MethodVisibility,
} from '../../method-types.js';
import { findVisibility, hasKeyword, hasModifier } from '../../field-extractors/configs/helpers.js';
import { extractSimpleTypeName } from '../../type-extractors/shared.js';
import type { SyntaxNode } from '../../utils/ast-helpers.js';
// ---------------------------------------------------------------------------
// Swift helpers
// ---------------------------------------------------------------------------
const SWIFT_VIS = new Set<MethodVisibility>([
'public',
'private',
'fileprivate',
'internal',
'open',
]);
/**
* Extract the method name from a function_declaration or protocol_function_declaration.
*
* In tree-sitter-swift, the name is stored in a `simple_identifier` child
* (not a 'name' field) on both function_declaration and protocol_function_declaration.
*/
function extractSwiftName(node: SyntaxNode): string | undefined {
// Try field-based name first
const nameField = node.childForFieldName('name');
if (nameField) return nameField.text;
// Walk named children for simple_identifier (the function name)
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === 'simple_identifier') return child.text;
}
return undefined;
}
/**
* Extract the return type from a Swift function declaration.
*
* In tree-sitter-swift, the return type appears in a `type_annotation` child
* that follows the parameter list (after `->` in source). It may also appear
* as a direct type child (user_type, optional_type, tuple_type, array_type).
*/
function extractSwiftReturnType(node: SyntaxNode): string | undefined {
// Look for the return type — typically the last type_annotation or a type node
// that appears after the parameter list.
// tree-sitter-swift places the return type inside a type child after '->'
let seenParams = false;
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (!child) continue;
if (child.type === 'parameter') {
seenParams = true;
continue;
}
// The parameter list may be unnamed children; track when we pass ')'
if (seenParams || child.type === 'type_annotation') {
if (child.type === 'type_annotation') {
const inner = child.firstNamedChild;
if (inner) return inner.text?.trim();
}
if (
child.type === 'user_type' ||
child.type === 'optional_type' ||
child.type === 'tuple_type' ||
child.type === 'array_type' ||
child.type === 'dictionary_type' ||
child.type === 'function_type'
) {
return child.text?.trim();
}
}
}
// Fallback: scan all children (named + unnamed) for '->' then grab the next named child
let seenArrow = false;
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (!child) continue;
if (!child.isNamed && child.text.trim() === '->') {
seenArrow = true;
continue;
}
if (seenArrow && child.isNamed) {
return child.text?.trim();
}
}
return undefined;
}
/**
* Extract parameters from a Swift function declaration.
*
* In tree-sitter-swift, parameters are `parameter` named children directly on
* the function_declaration node. Each parameter has:
* - An external name (label) and/or internal name as simple_identifier children
* - A type_annotation child containing the type
* - An optional default value after '='
* - A possible `...` for variadic parameters
*/
function extractSwiftParameters(node: SyntaxNode): ParameterInfo[] {
const params: ParameterInfo[] = [];
// In tree-sitter-swift 0.6.0, parameters are direct children of function_declaration.
// Default value tokens ('=', literal) are siblings of the parameter node at the
// function_declaration level, not children of the parameter node.
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (!child?.isNamed || child.type !== 'parameter') continue;
// Extract parameter name — the last simple_identifier is the internal name
let paramName: string | undefined;
for (let j = 0; j < child.namedChildCount; j++) {
const part = child.namedChild(j);
if (part?.type === 'simple_identifier') {
paramName = part.text;
}
}
if (!paramName) continue;
// Extract type — tree-sitter-swift uses user_type (not type_annotation)
let typeName: string | null = null;
let rawTypeName: string | null = null;
for (let j = 0; j < child.namedChildCount; j++) {
const part = child.namedChild(j);
if (part?.type === 'user_type' || part?.type === 'type_annotation') {
rawTypeName = part.text?.trim() ?? null;
const inner = part.firstNamedChild;
if (inner) {
typeName = extractSimpleTypeName(inner) ?? inner.text?.trim() ?? null;
} else {
typeName = rawTypeName;
}
break;
}
// Handle built-in types (array_type, dictionary_type, optional_type, tuple_type)
if (part?.type.endsWith('_type') && part.type !== 'simple_identifier') {
rawTypeName = part.text?.trim() ?? null;
typeName = extractSimpleTypeName(part) ?? rawTypeName;
break;
}
}
// Check for default value: '=' token appears as a sibling after the parameter node
let isOptional = false;
const nextSibling = node.child(i + 1);
if (nextSibling && !nextSibling.isNamed && nextSibling.text.trim() === '=') {
isOptional = true;
}
// Check for variadic: '...' token among parameter children
let isVariadic = false;
for (let j = 0; j < child.childCount; j++) {
const c = child.child(j);
if (c && c.text.trim() === '...') {
isVariadic = true;
break;
}
}
params.push({
name: paramName,
type: typeName,
rawType: rawTypeName,
isOptional,
isVariadic,
});
}
return params;
}
/**
* Check if a method is inside a protocol.
*
* A protocol_function_declaration is always abstract. For function_declaration
* inside a protocol_body (if it appears there), it's also abstract when it has
* no body.
*/
function isSwiftAbstract(node: SyntaxNode, ownerNode: SyntaxNode): boolean {
// protocol_function_declaration nodes are inherently abstract
if (node.type === 'protocol_function_declaration') return true;
// function_declaration inside a protocol is abstract if it has no body
if (ownerNode.type === 'protocol_declaration') {
const body = node.childForFieldName('body');
if (!body) {
// Also check for function_body named child
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === 'function_body') return false;
}
return true;
}
return false;
}
return false;
}
/**
* Collect attribute nodes from a Swift function declaration.
*
* In tree-sitter-swift, attributes appear as `attribute` named children
* directly on the function_declaration node, or inside a `modifiers` wrapper.
* Each attribute node text starts with '@'.
*/
function extractSwiftAnnotations(node: SyntaxNode): string[] {
const annotations: string[] = [];
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (!child) continue;
if (child.type === 'attribute') {
const text = child.text?.trim();
if (text) {
// Normalize: strip arguments, keep just the name
// e.g. "@objc(myMethod)" -> "@objc", "@available(iOS 13, *)" -> "@available"
const match = text.match(/^@(\w+)/);
if (match) {
annotations.push('@' + match[1]);
} else {
annotations.push(text);
}
}
}
// Also check inside modifiers wrapper
if (child.type === 'modifiers') {
for (let j = 0; j < child.namedChildCount; j++) {
const mod = child.namedChild(j);
if (mod?.type === 'attribute') {
const text = mod.text?.trim();
if (text) {
const match = text.match(/^@(\w+)/);
if (match) {
annotations.push('@' + match[1]);
} else {
annotations.push(text);
}
}
}
}
}
}
return annotations;
}
// ---------------------------------------------------------------------------
// Swift config
// ---------------------------------------------------------------------------
export const swiftMethodConfig: MethodExtractionConfig = {
language: SupportedLanguages.Swift,
// tree-sitter-swift 0.6.0 may use class_declaration for classes, structs, enums, extensions,
// and actors — but this cannot be verified until the grammar installs on Node 22+.
// TODO: Verify struct_declaration, enum_declaration, extension_declaration, actor_declaration
// node types once tree-sitter-swift loads on Node 22, and add them here if they are distinct.
// protocol_declaration is a separate, confirmed node type.
typeDeclarationNodes: ['class_declaration', 'protocol_declaration'],
// function_declaration for class/struct methods, protocol_function_declaration for protocol methods
methodNodeTypes: ['function_declaration', 'protocol_function_declaration'],
bodyNodeTypes: ['class_body', 'protocol_body'],
extractName: extractSwiftName,
extractReturnType: extractSwiftReturnType,
extractParameters: extractSwiftParameters,
extractVisibility(node) {
return findVisibility(node, SWIFT_VIS, 'internal', 'modifiers');
},
isStatic(node) {
return (
hasKeyword(node, 'static') ||
hasKeyword(node, 'class') ||
hasModifier(node, 'modifiers', 'static') ||
hasModifier(node, 'modifiers', 'class')
);
},
isAbstract: isSwiftAbstract,
isFinal(node) {
return hasKeyword(node, 'final') || hasModifier(node, 'modifiers', 'final');
},
isAsync(node) {
return hasKeyword(node, 'async') || hasModifier(node, 'modifiers', 'async');
},
isOverride(node) {
return hasKeyword(node, 'override') || hasModifier(node, 'modifiers', 'override');
},
extractAnnotations: extractSwiftAnnotations,
};
@@ -0,0 +1,351 @@
// gitnexus/src/core/ingestion/method-extractors/configs/typescript-javascript.ts
// Verified against tree-sitter-typescript ^0.23.2, tree-sitter-javascript ^0.23.0
import { SupportedLanguages } from 'gitnexus-shared';
import type {
MethodExtractionConfig,
ParameterInfo,
MethodVisibility,
} from '../../method-types.js';
import { hasKeyword } from '../../field-extractors/configs/helpers.js';
import { extractSimpleTypeName } from '../../type-extractors/shared.js';
import type { SyntaxNode } from '../../utils/ast-helpers.js';
// ---------------------------------------------------------------------------
// TS/JS helpers
// ---------------------------------------------------------------------------
const VISIBILITY_KEYWORDS = new Set<MethodVisibility>(['public', 'private', 'protected']);
/**
* Extract parameters from formal_parameters.
*
* Handles both TS node types (required_parameter, optional_parameter, rest_parameter)
* and JS node types (identifier, assignment_pattern, rest_pattern), plus destructured
* parameters (object_pattern, array_pattern) in both grammars.
*/
function extractTsJsParameters(node: SyntaxNode): ParameterInfo[] {
const paramList = node.childForFieldName('parameters');
if (!paramList) return [];
const params: ParameterInfo[] = [];
for (let i = 0; i < paramList.namedChildCount; i++) {
const param = paramList.namedChild(i);
if (!param) continue;
switch (param.type) {
case 'required_parameter': {
const patternNode = param.childForFieldName('pattern');
if (!patternNode) break;
// Skip TS `this` parameter — it's a compile-time type constraint, not a real param
if (patternNode.type === 'this') break;
// Rest parameter: pattern is a rest_pattern (...args) — extract inner identifier
const isRest = patternNode.type === 'rest_pattern';
const nameNode = isRest ? patternNode.firstNamedChild : patternNode;
if (!nameNode) break;
// type field is a type_annotation — unwrap to get the inner type node
const typeAnnotation = param.childForFieldName('type');
const typeNode = typeAnnotation?.firstNamedChild;
// Default value: presence of a 'value' field means isOptional
const hasDefault = !!param.childForFieldName('value');
params.push({
name: nameNode.text,
type: typeNode
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: hasDefault,
isVariadic: isRest,
});
break;
}
case 'optional_parameter': {
const nameNode = param.childForFieldName('pattern');
if (!nameNode) break;
const typeAnnotation = param.childForFieldName('type');
const typeNode = typeAnnotation?.firstNamedChild;
params.push({
name: nameNode.text,
type: typeNode
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: true,
isVariadic: false,
});
break;
}
case 'rest_parameter': {
const nameNode = param.childForFieldName('pattern');
if (!nameNode) break;
const typeAnnotation = param.childForFieldName('type');
const typeNode = typeAnnotation?.firstNamedChild;
params.push({
name: nameNode.text,
type: typeNode
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null,
rawType: typeNode?.text?.trim() ?? null,
isOptional: false,
isVariadic: true,
});
break;
}
case 'identifier': {
// JS: bare parameter name, no type info
params.push({
name: param.text,
type: null,
rawType: null,
isOptional: false,
isVariadic: false,
});
break;
}
case 'assignment_pattern': {
// JS: param = defaultValue — the left side is the name, isOptional = true
const left = param.childForFieldName('left');
if (left) {
params.push({
name: left.text,
type: null,
rawType: null,
isOptional: true,
isVariadic: false,
});
}
break;
}
case 'rest_pattern': {
// JS: ...args
const inner = param.firstNamedChild;
if (inner) {
params.push({
name: inner.text,
type: null,
rawType: null,
isOptional: false,
isVariadic: true,
});
}
break;
}
case 'object_pattern':
case 'array_pattern': {
// Destructured parameter — use full text as name
params.push({
name: param.text,
type: null,
rawType: null,
isOptional: false,
isVariadic: false,
});
break;
}
}
}
return params;
}
/** Regex to extract @returns or @return from JSDoc comments: `@returns {Type}` */
const JSDOC_RETURN_RE = /@returns?\s*\{([^}]+)\}/;
/**
* Minimal sanitization for JSDoc return types — preserves generic wrappers
* (e.g. `Promise<User>`) so that extractReturnTypeName in call-processor
* can apply WRAPPER_GENERICS unwrapping. Only strips JSDoc-specific syntax markers.
*/
function sanitizeJsDocReturnType(raw: string): string | undefined {
let type = raw.trim();
// Strip JSDoc nullable/non-nullable prefixes: ?User → User, !User → User
if (type.startsWith('?') || type.startsWith('!')) type = type.slice(1);
// Strip module: prefix — module:models.User → models.User
if (type.startsWith('module:')) type = type.slice(7);
// Reject unions (ambiguous)
if (type.includes('|')) return undefined;
if (!type) return undefined;
return type;
}
/**
* Walk backwards through preceding siblings looking for a JSDoc comment containing
* `@returns {Type}` or `@return {Type}`. Stops at the first non-comment named node
* (excluding decorators, which precede methods in TS/JS).
*/
function extractJsDocReturnType(node: SyntaxNode): string | undefined {
let sibling = node.previousSibling;
while (sibling) {
if (sibling.type === 'comment') {
const match = JSDOC_RETURN_RE.exec(sibling.text);
if (match) return sanitizeJsDocReturnType(match[1]);
} else if (sibling.isNamed && sibling.type !== 'decorator') break;
sibling = sibling.previousSibling;
}
return undefined;
}
/**
* Extract return type from return_type field, unwrapping type_annotation.
* Falls back to JSDoc `@returns {Type}` when the AST has no return type annotation.
*
* tree-sitter-typescript uses `return_type` as the field name (not `type` like JVM).
* The return_type field points to a type_annotation node that must be unwrapped.
*/
function extractTsJsReturnType(node: SyntaxNode): string | undefined {
const returnType = node.childForFieldName('return_type');
if (returnType) {
if (returnType.type === 'type_annotation') {
const inner = returnType.firstNamedChild;
if (inner) return inner.text?.trim();
}
return returnType.text?.trim();
}
// AST has no return type annotation — try JSDoc fallback
return extractJsDocReturnType(node);
}
/**
* Extract visibility from accessibility_modifier or #private name.
*
* tree-sitter-typescript emits accessibility_modifier as a named child of method nodes
* (not as a modifiers wrapper like JVM). Pass 1 scans for that child; pass 2 checks for
* ES2022 private_property_identifier (#name). Default: public.
*/
function extractTsJsVisibility(node: SyntaxNode): MethodVisibility {
// Pass 1: check for accessibility_modifier named child (TS-specific)
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child && child.type === 'accessibility_modifier') {
const t = child.text.trim();
if (VISIBILITY_KEYWORDS.has(t as MethodVisibility)) return t as MethodVisibility;
}
}
// Pass 2: ES2022 private methods (#name) are inherently private
const nameNode = node.childForFieldName('name');
if (nameNode && nameNode.type === 'private_property_identifier') return 'private';
// No accessibility_modifier found — default to public.
// Note: tree-sitter-typescript does not wrap modifiers in a 'modifiers' node
// (unlike JVM), so there is no wrapper to scan.
return 'public';
}
/**
* Extract decorator names, prefixed with '@'.
*
* In tree-sitter-typescript, decorators are **siblings** of the method_definition in the
* class_body — they are NOT children of the method node. We find them by walking backwards
* from the method node through its preceding siblings in the parent body.
*/
function extractTsJsDecorators(node: SyntaxNode): string[] {
const decorators: string[] = [];
// Walk backwards via previousNamedSibling to collect consecutive decorator siblings.
// This avoids the O(N) index-finding scan through the parent's children.
let sibling = node.previousNamedSibling;
while (sibling && sibling.type === 'decorator') {
const name = extractDecoratorName(sibling);
if (name) decorators.unshift(name);
sibling = sibling.previousNamedSibling;
}
return decorators;
}
function extractDecoratorName(decorator: SyntaxNode): string | undefined {
const expr = decorator.firstNamedChild;
if (!expr) return undefined;
if (expr.type === 'call_expression') {
const fn = expr.childForFieldName('function');
return fn ? '@' + fn.text : undefined;
}
if (expr.type === 'identifier') return '@' + expr.text;
if (expr.type === 'member_expression') return '@' + expr.text;
return undefined;
}
// ---------------------------------------------------------------------------
// Config
// ---------------------------------------------------------------------------
// TS and JS share the same config base. TS-only node types (abstract_class_declaration,
// interface_declaration, abstract_method_signature, method_signature, interface_body) are
// included because the JS grammar never produces these nodes — they are harmless no-ops.
// This mirrors the field extractor's typescript-javascript.ts shared pattern.
//
// Note: TS and JS share a method config but NOT a field extractor because the TS field
// extractor needs a hand-written class for type_alias_declaration object literals and
// nested type discovery. Methods have no such requirement.
const shared: Omit<MethodExtractionConfig, 'language'> = {
typeDeclarationNodes: [
'class_declaration',
'abstract_class_declaration',
'interface_declaration',
],
// Note: TS constructors are method_definition nodes (name = 'constructor'), so no
// explicit constructor_declaration entry is needed (unlike JVM/C# configs).
// Known gaps:
// - call_signature and construct_signature (e.g., interface Fn { (x: string): void; })
// are not extracted — they have no name field and are uncommon in practice.
// - class_expression (const Foo = class { ... }) — methods inside class expressions
// are not discovered because class_expression is not in typeDeclarationNodes.
// - declare module / declare global augmentations — methods inside ambient_module_declaration
// wrappers are not surfaced because the top-level walker doesn't descend into them.
methodNodeTypes: [
'method_definition',
'method_signature',
'abstract_method_signature',
'function_declaration',
'generator_function_declaration',
'function_signature',
],
bodyNodeTypes: ['class_body', 'interface_body'],
extractName(node) {
const nameNode = node.childForFieldName('name');
return nameNode?.text;
},
extractReturnType: extractTsJsReturnType,
extractParameters: extractTsJsParameters,
extractVisibility: extractTsJsVisibility,
isStatic(node) {
return hasKeyword(node, 'static');
},
isAbstract(node, ownerNode) {
// Explicit abstract keyword on the method itself
if (hasKeyword(node, 'abstract')) return true;
// Interface methods are implicitly abstract — TS interfaces never have method bodies
// (unlike Java default methods), so no !body check needed
if (ownerNode.type === 'interface_declaration') return true;
return false;
},
isFinal(_node) {
return false; // TS/JS has no final/sealed methods
},
extractAnnotations: extractTsJsDecorators,
isAsync(node) {
return hasKeyword(node, 'async');
},
isOverride(node) {
return hasKeyword(node, 'override');
},
};
export const typescriptMethodConfig: MethodExtractionConfig = {
...shared,
language: SupportedLanguages.TypeScript,
};
export const javascriptMethodConfig: MethodExtractionConfig = {
...shared,
language: SupportedLanguages.JavaScript,
};
@@ -16,8 +16,8 @@ import type {
MethodInfo,
} from '../method-types.js';
/** Owner node types where member functions are effectively static (JVM semantics). */
const STATIC_OWNER_TYPES = new Set(['companion_object', 'object_declaration']);
/** Owner node types where member functions are effectively static (JVM/Ruby semantics). */
const STATIC_OWNER_TYPES = new Set(['companion_object', 'object_declaration', 'singleton_class']);
/**
* Create a MethodExtractor from a declarative config.
@@ -37,17 +37,27 @@ export function createMethodExtractor(config: MethodExtractionConfig): MethodExt
extract(node: SyntaxNode, context: MethodExtractorContext): ExtractedMethods | null {
if (!typeDeclarationSet.has(node.type)) return null;
// Resolve owner name: field-based → type_identifier → simple_identifier → "Companion"
// Resolve owner name: config hook → field-based → type_identifier → simple_identifier → "Companion"
let ownerName: string | undefined;
const nameField = node.childForFieldName('name');
if (nameField) {
ownerName = nameField.text;
} else {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child && (child.type === 'type_identifier' || child.type === 'simple_identifier')) {
ownerName = child.text;
break;
if (config.extractOwnerName) {
ownerName = config.extractOwnerName(node);
}
if (!ownerName) {
const nameField = node.childForFieldName('name');
if (nameField) {
ownerName = nameField.text;
} else {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (
child &&
(child.type === 'type_identifier' ||
child.type === 'simple_identifier' ||
child.type === 'identifier')
) {
ownerName = child.text;
break;
}
}
}
}
@@ -71,6 +81,13 @@ export function createMethodExtractor(config: MethodExtractionConfig): MethodExt
return { ownerName, methods };
},
extractFromNode(node: SyntaxNode, context: MethodExtractorContext): MethodInfo | null {
if (!methodNodeSet.has(node.type)) return null;
return buildMethod(node, node, context, config);
},
...(config.extractFunctionName ? { extractFunctionName: config.extractFunctionName } : {}),
};
}
@@ -102,10 +119,17 @@ function findBodies(node: SyntaxNode, bodyNodeSet: Set<string>): SyntaxNode[] {
return result;
}
function addNestedBodies(parent: SyntaxNode, bodyNodeSet: Set<string>, out: SyntaxNode[]): void {
function addNestedBodies(
parent: SyntaxNode,
bodyNodeSet: Set<string>,
out: SyntaxNode[],
seen?: Set<SyntaxNode>,
): void {
const visited = seen ?? new Set(out);
for (let i = 0; i < parent.namedChildCount; i++) {
const child = parent.namedChild(i);
if (child && bodyNodeSet.has(child.type) && !out.includes(child)) {
if (child && bodyNodeSet.has(child.type) && !visited.has(child)) {
visited.add(child);
out.push(child);
}
}
@@ -120,9 +144,15 @@ function extractMethodsFromBody(
out: MethodInfo[],
): void {
for (let i = 0; i < body.namedChildCount; i++) {
const child = body.namedChild(i);
let child = body.namedChild(i);
if (!child) continue;
// C++ template methods are wrapped in template_declaration — unwrap to the inner node
if (child.type === 'template_declaration') {
const inner = child.namedChildren.find((c) => methodNodeSet.has(c.type));
if (inner) child = inner;
}
if (methodNodeSet.has(child.type)) {
const method = buildMethod(child, ownerNode, context, config);
if (method) out.push(method);
@@ -170,6 +200,7 @@ function buildMethod(
...(config.isOverride?.(node) ? { isOverride: true } : {}),
...(config.isAsync?.(node) ? { isAsync: true } : {}),
...(config.isPartial?.(node) ? { isPartial: true } : {}),
...(config.isConst?.(node) ? { isConst: true } : {}),
annotations: config.extractAnnotations?.(node) ?? [],
sourceFile: context.filePath,
line: node.startPosition.row + 1,
@@ -10,6 +10,10 @@ export type MethodVisibility = FieldVisibility;
export interface ParameterInfo {
name: string;
type: string | null;
/** Full type text including generic/template args (e.g. 'vector<int>', 'List<String>').
* Used by typeTagForId for overload disambiguation where generic args matter.
* Falls back to `type` when not set. */
rawType?: string | null;
isOptional: boolean;
isVariadic: boolean;
}
@@ -27,6 +31,7 @@ export interface MethodInfo {
isOverride?: boolean;
isAsync?: boolean;
isPartial?: boolean;
isConst?: boolean;
annotations: string[];
sourceFile: string;
line: number;
@@ -46,6 +51,16 @@ export interface MethodExtractor {
language: SupportedLanguages;
extract(node: SyntaxNode, context: MethodExtractorContext): ExtractedMethods | null;
isTypeDeclaration(node: SyntaxNode): boolean;
/** Extract method info from a standalone method node (e.g. Go top-level method_declaration). */
extractFromNode?(node: SyntaxNode, context: MethodExtractorContext): MethodInfo | null;
/** Extract function name + label from an AST node during parent-walk.
* Languages with non-standard AST structures (e.g. C/C++ declarator
* unwrapping, Swift init/deinit, Rust impl_item) provide this hook
* to replace the generic name-field lookup.
* Return null to fall through to the generic extractor. */
extractFunctionName?(
node: SyntaxNode,
): { funcName: string | null; label: import('gitnexus-shared').NodeLabel } | null;
}
export interface MethodExtractionConfig {
@@ -66,9 +81,17 @@ export interface MethodExtractionConfig {
isOverride?: (node: SyntaxNode) => boolean;
isAsync?: (node: SyntaxNode) => boolean;
isPartial?: (node: SyntaxNode) => boolean;
isConst?: (node: SyntaxNode) => boolean;
/** Resolve the owner name from a standalone method node (e.g. Go receiver type). */
extractOwnerName?: (node: SyntaxNode) => string | undefined;
/** Extract a primary constructor from the owner node itself (e.g. C# 12 class Point(int x, int y)). */
extractPrimaryConstructor?: (
ownerNode: SyntaxNode,
context: MethodExtractorContext,
) => MethodInfo | null;
/** Extract function name + label from an AST node during parent-walk.
* Passed through to the MethodExtractor by createMethodExtractor. */
extractFunctionName?: (
node: SyntaxNode,
) => { funcName: string | null; label: import('gitnexus-shared').NodeLabel } | null;
}
+400 -14
View File
@@ -3,7 +3,7 @@
*
* Walks the inheritance DAG (EXTENDS/IMPLEMENTS edges), collects methods from
* each ancestor via HAS_METHOD edges, detects method-name collisions across
* parents, and applies language-specific resolution rules to emit OVERRIDES edges.
* parents, and applies language-specific resolution rules to emit METHOD_OVERRIDES edges.
*
* Language-specific rules:
* - C++: leftmost base class in declaration order wins
@@ -13,10 +13,10 @@
* - Rust: no auto-resolution — requires qualified syntax, resolvedTo = null
* - Default: single inheritance — first definition wins
*
* OVERRIDES edge direction: Class → Method (not Method → Method).
* METHOD_OVERRIDES edge direction: Class → Method (not Method → Method).
* The source is the child class that inherits conflicting methods,
* the target is the winning ancestor method node.
* Cypher: MATCH (c:Class)-[r:CodeRelation {type: 'OVERRIDES'}]->(m:Method)
* Cypher: MATCH (c:Class)-[r:CodeRelation {type: 'METHOD_OVERRIDES'}]->(m:Method)
*/
import { KnowledgeGraph } from '../graph/types.js';
@@ -47,6 +47,7 @@ export interface MROResult {
entries: MROEntry[];
overrideEdges: number;
ambiguityCount: number;
methodImplementsEdges: number;
}
// ---------------------------------------------------------------------------
@@ -126,7 +127,7 @@ function gatherAncestors(classId: string, parentMap: Map<string, string[]>): str
* Returns an array of ancestor IDs in C3 order (excluding the class itself),
* or null if linearization fails (inconsistent or cyclic hierarchy).
*/
function c3Linearize(
export function c3Linearize(
classId: string,
parentMap: Map<string, string[]>,
cache: Map<string, string[] | null>,
@@ -289,6 +290,10 @@ export function computeMRO(graph: KnowledgeGraph): MROResult {
let overrideEdges = 0;
let ambiguityCount = 0;
// Pre-computed maps to avoid redundant BFS in emitMethodImplementsEdges
const ancestorsMap = new Map<string, string[]>();
const edgeTypesMap = new Map<string, Map<string, 'EXTENDS' | 'IMPLEMENTS'>>();
// Process every class that has at least one parent
for (const [classId, directParents] of parentMap) {
if (directParents.length === 0) continue;
@@ -302,12 +307,16 @@ export function computeMRO(graph: KnowledgeGraph): MROResult {
// Compute linearized MRO depending on language strategy
const provider = getProvider(language);
const ancestors = gatherAncestors(classId, parentMap);
ancestorsMap.set(classId, ancestors);
edgeTypesMap.set(classId, buildTransitiveEdgeTypes(classId, parentMap, parentEdgeType));
let mroOrder: string[];
if (provider.mroStrategy === 'c3') {
const c3Result = c3Linearize(classId, parentMap, c3Cache);
mroOrder = c3Result ?? gatherAncestors(classId, parentMap);
mroOrder = c3Result ?? ancestors;
} else {
mroOrder = gatherAncestors(classId, parentMap);
mroOrder = ancestors;
}
// Get the parent names for the MRO entry
@@ -348,11 +357,9 @@ export function computeMRO(graph: KnowledgeGraph): MROResult {
// Detect collisions: methods defined in 2+ different ancestors
const ambiguities: MethodAmbiguity[] = [];
// Compute transitive edge types once per class (only needed for implements-split languages)
// Use pre-computed transitive edge types (only needed for implements-split languages)
const needsEdgeTypes = provider.mroStrategy === 'implements-split';
const classEdgeTypes = needsEdgeTypes
? buildTransitiveEdgeTypes(classId, parentMap, parentEdgeType)
: undefined;
const classEdgeTypes = needsEdgeTypes ? edgeTypesMap.get(classId) : undefined;
for (const [methodName, defs] of methodsByName) {
if (defs.length < 2) continue;
@@ -401,13 +408,13 @@ export function computeMRO(graph: KnowledgeGraph): MROResult {
ambiguityCount++;
}
// Emit OVERRIDES edge if resolution found
// Emit METHOD_OVERRIDES edge if resolution found
if (resolution.resolvedTo !== null) {
graph.addRelationship({
id: generateId('OVERRIDES', `${classId}->${resolution.resolvedTo}`),
id: generateId('METHOD_OVERRIDES', `${classId}->${resolution.resolvedTo}`),
sourceId: classId,
targetId: resolution.resolvedTo,
type: 'OVERRIDES',
type: 'METHOD_OVERRIDES',
confidence: resolution.confidence,
reason: resolution.reason,
});
@@ -424,7 +431,386 @@ export function computeMRO(graph: KnowledgeGraph): MROResult {
});
}
return { entries, overrideEdges, ambiguityCount };
const methodImplementsEdges = emitMethodImplementsEdges(
graph,
parentMap,
methodMap,
parentEdgeType,
ancestorsMap,
edgeTypesMap,
);
return { entries, overrideEdges, ambiguityCount, methodImplementsEdges };
}
// ---------------------------------------------------------------------------
// METHOD_IMPLEMENTS edge emission
// ---------------------------------------------------------------------------
/**
* Check if two parameter type arrays match.
* When either side has no type info, fall back to parameterCount comparison
* (arity-compatible matching). If both have parameterCount and they differ,
* return no match. If counts match, return confident match. If either count
* is undefined, return lenient (non-confident) match.
*
* Returns `{ match, confident }`:
* - Exact type match → `{ match: true, confident: true }`
* - Arity match (both have parameterCount, counts equal) → `{ match: true, confident: true }`
* - Lenient (either side lacks types AND lacks parameterCount) → `{ match: true, confident: false }`
* - No match → `{ match: false, confident: false }`
*/
function parameterTypesMatch(
a: string[],
b: string[],
aParamCount?: number,
bParamCount?: number,
): { match: boolean; confident: boolean } {
// If one side is variadic and the other isn't, types may match superficially
// but the methods aren't guaranteed to be interchangeable
if ((aParamCount === undefined) !== (bParamCount === undefined)) {
return { match: true, confident: false };
}
if (a.length === 0 || b.length === 0) {
// Fall back to arity check when type info is missing
if (aParamCount !== undefined && bParamCount !== undefined) {
return { match: aParamCount === bParamCount, confident: aParamCount === bParamCount };
}
return { match: true, confident: false }; // lenient when either count is unknown
}
if (a.length !== b.length) return { match: false, confident: false };
const exact = a.every((t, i) => t === b[i]);
return { match: exact, confident: exact };
}
/**
* For each concrete class that implements/extends an interface or trait,
* find methods in the class that implement methods defined in the interface
* and emit METHOD_IMPLEMENTS edges: ConcreteMethod → InterfaceMethod.
*
* Method node IDs include a `#<paramCount>` arity suffix, so overloaded
* methods with different parameter counts are distinct nodes in the graph.
* For same-arity overloads with different parameter types, a `~type1,type2`
* suffix is appended when type info is available (issue #651), producing
* distinct nodes that `parameterTypesMatch` can resolve to correct edges.
*/
function emitMethodImplementsEdges(
graph: KnowledgeGraph,
parentMap: Map<string, string[]>,
methodMap: Map<string, string[]>,
parentEdgeType: Map<string, Map<string, 'EXTENDS' | 'IMPLEMENTS'>>,
ancestorsMap: Map<string, string[]>,
edgeTypesMap: Map<string, Map<string, 'EXTENDS' | 'IMPLEMENTS'>>,
): number {
let edgeCount = 0;
for (const [classId, parentIds] of parentMap) {
const classNode = graph.getNode(classId);
if (!classNode) continue;
// Interfaces and traits declare contracts — they don't implement them
if (classNode.label === 'Interface' || classNode.label === 'Trait') continue;
// Get this class's own methods
const ownMethodIds = methodMap.get(classId) ?? [];
// Build a lookup: methodName → Array<{methodId, parameterTypes, parameterCount}> for own methods
const ownMethodsByName = new Map<
string,
Array<{ methodId: string; parameterTypes: string[]; parameterCount?: number }>
>();
for (const methodId of ownMethodIds) {
const methodNode = graph.getNode(methodId);
if (!methodNode || methodNode.label === 'Property') continue;
// Abstract methods don't satisfy interface contracts
if (methodNode.properties.isAbstract === true) continue;
const name = methodNode.properties.name as string;
const parameterTypes = (methodNode.properties.parameterTypes as string[] | undefined) ?? [];
const parameterCount = methodNode.properties.parameterCount as number | undefined;
let bucket = ownMethodsByName.get(name);
if (!bucket) {
bucket = [];
ownMethodsByName.set(name, bucket);
}
bucket.push({ methodId, parameterTypes, parameterCount });
}
// Use pre-computed ancestors and edge types; fall back to computing if missing (safety)
const allAncestors = ancestorsMap.get(classId) ?? gatherAncestors(classId, parentMap);
const ancestorEdgeTypes =
edgeTypesMap.get(classId) ?? buildTransitiveEdgeTypes(classId, parentMap, parentEdgeType);
// Dedup set: avoid duplicate edges from diamond paths
const emitted = new Set<string>();
// For each ancestor, check if it's an interface/trait or classified as IMPLEMENTS
for (const ancestorId of allAncestors) {
const ancestorNode = graph.getNode(ancestorId);
if (!ancestorNode) continue;
const isInterfaceLike = ancestorNode.label === 'Interface' || ancestorNode.label === 'Trait';
const classifiedEdgeType = ancestorEdgeTypes.get(ancestorId);
if (!isInterfaceLike && classifiedEdgeType !== 'IMPLEMENTS') continue;
// Get ancestor's methods
const ancestorMethodIds = methodMap.get(ancestorId) ?? [];
for (const ancestorMethodId of ancestorMethodIds) {
const ancestorMethodNode = graph.getNode(ancestorMethodId);
if (!ancestorMethodNode || ancestorMethodNode.label === 'Property') continue;
const ancestorName = ancestorMethodNode.properties.name as string;
const ancestorParamTypes =
(ancestorMethodNode.properties.parameterTypes as string[] | undefined) ?? [];
const ancestorParamCount = ancestorMethodNode.properties.parameterCount as
| number
| undefined;
// Find matching method in own class by name + parameterTypes/arity
const candidates = ownMethodsByName.get(ancestorName);
// Unit 3: If no own method matches, walk the EXTENDS chain to find inherited concrete method
if (!candidates || candidates.length === 0) {
const inherited = findInheritedMethod(
classId,
ancestorName,
ancestorParamTypes,
ancestorParamCount,
graph,
parentMap,
methodMap,
parentEdgeType,
ancestorMethodId,
);
if (inherited) {
const edgeKey = `${inherited.methodId}->${ancestorMethodId}`;
if (!emitted.has(edgeKey)) {
emitted.add(edgeKey);
graph.addRelationship({
id: generateId('METHOD_IMPLEMENTS', edgeKey),
sourceId: inherited.methodId,
targetId: ancestorMethodId,
type: 'METHOD_IMPLEMENTS',
confidence: inherited.confident ? 1.0 : 0.7,
reason: '',
});
edgeCount++;
}
}
continue;
}
// Unit 4: Filter candidates by type/arity match, then check for ambiguity
const matching: Array<{
methodId: string;
parameterTypes: string[];
parameterCount?: number;
confident: boolean;
}> = [];
for (const c of candidates) {
const result = parameterTypesMatch(
c.parameterTypes,
ancestorParamTypes,
c.parameterCount,
ancestorParamCount,
);
if (result.match) {
matching.push({ ...c, confident: result.confident });
}
}
if (matching.length === 0) continue;
// If multiple candidates match at name+arity level, emit no edge (ambiguous)
if (matching.length > 1) continue;
const winner = matching[0];
const edgeKey = `${winner.methodId}->${ancestorMethodId}`;
if (emitted.has(edgeKey)) continue;
emitted.add(edgeKey);
graph.addRelationship({
id: generateId('METHOD_IMPLEMENTS', edgeKey),
sourceId: winner.methodId,
targetId: ancestorMethodId,
type: 'METHOD_IMPLEMENTS',
confidence: winner.confident ? 1.0 : 0.7,
reason: '',
});
edgeCount++;
}
}
}
return edgeCount;
}
/**
* Walk the class's EXTENDS chain to find the nearest concrete method matching
* the given name and parameter signature. If the EXTENDS chain yields no match,
* fall back to IMPLEMENTS parents and check for non-abstract default methods
* (e.g. Java default interface methods, Kotlin interface defaults).
* Returns the first matching method found in BFS order, or null.
*/
function findInheritedMethod(
classId: string,
methodName: string,
targetParamTypes: string[],
targetParamCount: number | undefined,
graph: KnowledgeGraph,
parentMap: Map<string, string[]>,
methodMap: Map<string, string[]>,
parentEdgeType: Map<string, Map<string, 'EXTENDS' | 'IMPLEMENTS'>>,
/** Method ID to exclude from results (prevents self-edges when the ancestor
* method being matched lives on an IMPLEMENTS parent). */
excludeMethodId?: string,
): { methodId: string; parameterTypes: string[]; confident: boolean } | null {
const visited = new Set<string>();
const queue: string[] = [];
// Seed with direct EXTENDS parents only
const directParents = parentMap.get(classId) ?? [];
const directEdges = parentEdgeType.get(classId);
for (const pid of directParents) {
const et = directEdges?.get(pid);
if (et === 'EXTENDS') {
// Also check that the parent is not an Interface/Trait
const parentNode = graph.getNode(pid);
if (parentNode && parentNode.label !== 'Interface' && parentNode.label !== 'Trait') {
queue.push(pid);
}
}
}
// Level-order BFS: process all ancestors at the current depth before
// advancing. Once any match is found at depth D, finish that depth and stop.
// Diamond dedup: same methodId via two paths at the same depth = 1 match.
let currentLevel = [...queue];
while (currentLevel.length > 0) {
const matches = new Map<
string,
{ methodId: string; parameterTypes: string[]; confident: boolean }
>();
const nextLevel: string[] = [];
for (const ancestorId of currentLevel) {
if (visited.has(ancestorId)) continue;
visited.add(ancestorId);
// Check this ancestor's methods
const methods = methodMap.get(ancestorId) ?? [];
for (const mid of methods) {
const mNode = graph.getNode(mid);
if (!mNode || mNode.label === 'Property') continue;
// Abstract inherited methods don't count as concrete implementations
if (mNode.properties.isAbstract === true) continue;
if (mNode.properties.name !== methodName) continue;
const mParamTypes = (mNode.properties.parameterTypes as string[] | undefined) ?? [];
const mParamCount = mNode.properties.parameterCount as number | undefined;
const ptResult = parameterTypesMatch(
mParamTypes,
targetParamTypes,
mParamCount,
targetParamCount,
);
if (ptResult.match) {
matches.set(mid, {
methodId: mid,
parameterTypes: mParamTypes,
confident: ptResult.confident,
});
}
}
// Collect EXTENDS parents for the next depth level
const grandparents = parentMap.get(ancestorId) ?? [];
const ancestorEdges = parentEdgeType.get(ancestorId);
for (const gp of grandparents) {
if (visited.has(gp)) continue;
const gpEdge = ancestorEdges?.get(gp);
if (gpEdge === 'EXTENDS') {
const gpNode = graph.getNode(gp);
if (gpNode && gpNode.label !== 'Interface' && gpNode.label !== 'Trait') {
nextLevel.push(gp);
}
}
}
}
// If any matches found at this depth, decide and stop
if (matches.size === 1) return matches.values().next().value!;
if (matches.size > 1) return null; // ambiguous at same depth
currentLevel = nextLevel;
}
// ── Second pass: walk IMPLEMENTS parents AND their interface ancestry ──
// Only reached when the EXTENDS chain yielded no match.
// BFS through interface/trait hierarchy to find default (non-abstract) methods.
const implBfsQueue: string[] = [];
for (const pid of directParents) {
const et = directEdges?.get(pid);
if (et === 'IMPLEMENTS') {
implBfsQueue.push(pid);
}
}
// Collect all matches from the IMPLEMENTS BFS — return null if ambiguous (>1 match)
const implMatches: Array<{
methodId: string;
parameterTypes: string[];
confident: boolean;
}> = [];
const implVisited = new Set<string>();
while (implBfsQueue.length > 0) {
const ifaceId = implBfsQueue.shift()!;
if (implVisited.has(ifaceId)) continue;
implVisited.add(ifaceId);
// Only process Interface/Trait nodes — Dart `implements Class` does not
// inherit method bodies, so Class/Struct/Enum parents must be skipped.
const ifaceNode = graph.getNode(ifaceId);
if (!ifaceNode || (ifaceNode.label !== 'Interface' && ifaceNode.label !== 'Trait')) continue;
// Check this interface/trait's methods for a non-abstract default
const methods = methodMap.get(ifaceId) ?? [];
for (const mid of methods) {
if (mid === excludeMethodId) continue; // prevent self-edges
const mNode = graph.getNode(mid);
if (!mNode || mNode.label === 'Property') continue;
if (mNode.properties.isAbstract === true) continue;
if (mNode.properties.name !== methodName) continue;
const mParamTypes = (mNode.properties.parameterTypes as string[] | undefined) ?? [];
const mParamCount = mNode.properties.parameterCount as number | undefined;
const ptResult = parameterTypesMatch(
mParamTypes,
targetParamTypes,
mParamCount,
targetParamCount,
);
if (ptResult.match) {
implMatches.push({
methodId: mid,
parameterTypes: mParamTypes,
confident: ptResult.confident,
});
}
}
// Walk this interface's parents (interface-extends-interface chains)
const ifaceParents = parentMap.get(ifaceId) ?? [];
for (const gp of ifaceParents) {
if (!implVisited.has(gp)) implBfsQueue.push(gp);
}
}
// Ambiguous: multiple interfaces provide the same default method
if (implMatches.length === 1) return implMatches[0];
return null; // 0 matches or ambiguous (>1)
}
/**
@@ -41,7 +41,7 @@ export function walkBindingChain(
const targetName = binding.exportedName;
const resolvedDefs =
targetName !== lookupName || depth > 0
? symbolTable.lookupFuzzy(targetName).filter((def) => def.filePath === binding.sourcePath)
? symbolTable.lookupExactAll(binding.sourcePath, targetName)
: allDefs.filter((def) => def.filePath === binding.sourcePath);
if (resolvedDefs.length > 0) return resolvedDefs;
+235 -94
View File
@@ -4,21 +4,30 @@ import Parser from 'tree-sitter';
import { loadParser, loadLanguage, isLanguageAvailable } from '../tree-sitter/parser-loader.js';
import { getProvider } from './languages/index.js';
import { generateId } from '../../lib/utils.js';
import { SymbolTable } from './symbol-table.js';
import type { SymbolTable } from './symbol-table.js';
import { ASTCache } from './ast-cache.js';
import { getLanguageFromFilename } from 'gitnexus-shared';
import { getLanguageFromFilename, SupportedLanguages } from 'gitnexus-shared';
import { extractVueScript, isVueSetupTopLevel } from './vue-sfc-extractor.js';
import { yieldToEventLoop } from './utils/event-loop.js';
import {
getDefinitionNodeFromCaptures,
findEnclosingClassId,
extractMethodSignature,
findEnclosingClassInfo,
getLabelFromCaptures,
CLASS_CONTAINER_TYPES,
type SyntaxNode,
type EnclosingClassInfo,
} from './utils/ast-helpers.js';
import { detectFrameworkFromAST } from './framework-detection.js';
import { buildTypeEnv } from './type-env.js';
import type { FieldInfo, FieldExtractorContext } from './field-types.js';
import type { MethodInfo } from './method-types.js';
import {
buildMethodProps,
arityForIdFromInfo,
typeTagForId,
constTagForId,
buildCollisionGroups,
} from './utils/method-props.js';
import type { LanguageProvider } from './language-provider.js';
import { WorkerPool } from './workers/worker-pool.js';
import type {
@@ -33,7 +42,7 @@ import type {
ExtractedDecoratorRoute,
ExtractedToolDef,
FileConstructorBindings,
FileTypeEnvBindings,
FileScopeBindings,
ExtractedORMQuery,
} from './workers/parse-worker.js';
import { getTreeSitterBufferSize, TREE_SITTER_MAX_BUFFER } from './constants.js';
@@ -51,7 +60,7 @@ export interface WorkerExtractedData {
toolDefs: ExtractedToolDef[];
ormQueries: ExtractedORMQuery[];
constructorBindings: FileConstructorBindings[];
typeEnvBindings: FileTypeEnvBindings[];
fileScopeBindings: FileScopeBindings[];
}
// ============================================================================
@@ -85,7 +94,7 @@ const processParsingWithWorkers = async (
toolDefs: [],
ormQueries: [],
constructorBindings: [],
typeEnvBindings: [],
fileScopeBindings: [],
};
const total = files.length;
@@ -109,7 +118,7 @@ const processParsingWithWorkers = async (
const allToolDefs: ExtractedToolDef[] = [];
const allORMQueries: ExtractedORMQuery[] = [];
const allConstructorBindings: FileConstructorBindings[] = [];
const allTypeEnvBindings: FileTypeEnvBindings[] = [];
const fileScopeBindingsByFile: FileScopeBindings[] = [];
for (const result of chunkResults) {
for (const node of result.nodes) {
graph.addNode({
@@ -131,20 +140,22 @@ const processParsingWithWorkers = async (
returnType: sym.returnType,
declaredType: sym.declaredType,
ownerId: sym.ownerId,
qualifiedName: sym.qualifiedName,
});
}
allImports.push(...result.imports);
allCalls.push(...result.calls);
allAssignments.push(...result.assignments);
allHeritage.push(...result.heritage);
allRoutes.push(...result.routes);
allFetchCalls.push(...result.fetchCalls);
allDecoratorRoutes.push(...result.decoratorRoutes);
allToolDefs.push(...result.toolDefs);
if (result.ormQueries) allORMQueries.push(...result.ormQueries);
allConstructorBindings.push(...result.constructorBindings);
allTypeEnvBindings.push(...result.typeEnvBindings);
for (const _item of result.imports) allImports.push(_item);
for (const _item of result.calls) allCalls.push(_item);
for (const _item of result.assignments) allAssignments.push(_item);
for (const _item of result.heritage) allHeritage.push(_item);
for (const _item of result.routes) allRoutes.push(_item);
for (const _item of result.fetchCalls) allFetchCalls.push(_item);
for (const _item of result.decoratorRoutes) allDecoratorRoutes.push(_item);
for (const _item of result.toolDefs) allToolDefs.push(_item);
if (result.ormQueries) for (const _item of result.ormQueries) allORMQueries.push(_item);
for (const _item of result.constructorBindings) allConstructorBindings.push(_item);
if (result.fileScopeBindings)
for (const _item of result.fileScopeBindings) fileScopeBindingsByFile.push(_item);
}
// Merge and log skipped languages from workers
@@ -174,7 +185,7 @@ const processParsingWithWorkers = async (
toolDefs: allToolDefs,
ormQueries: allORMQueries,
constructorBindings: allConstructorBindings,
typeEnvBindings: allTypeEnvBindings,
fileScopeBindings: fileScopeBindingsByFile,
};
};
@@ -184,14 +195,17 @@ const processParsingWithWorkers = async (
// Inline caches to avoid repeated parent-walks per node (same pattern as parse-worker.ts).
// Keyed by tree-sitter node reference — cleared at the start of each file.
const classIdCache = new Map<SyntaxNode, string | null>();
const classInfoCache = new Map<SyntaxNode, EnclosingClassInfo | null>();
const exportCache = new Map<SyntaxNode, boolean>();
const cachedFindEnclosingClassId = (node: SyntaxNode, filePath: string): string | null => {
const cached = classIdCache.get(node);
const cachedFindEnclosingClassInfo = (
node: SyntaxNode,
filePath: string,
): EnclosingClassInfo | null => {
const cached = classInfoCache.get(node);
if (cached !== undefined) return cached;
const result = findEnclosingClassId(node, filePath);
classIdCache.set(node, result);
const result = findEnclosingClassInfo(node, filePath);
classInfoCache.set(node, result);
return result;
};
@@ -210,6 +224,18 @@ const cachedExportCheck = (
// FieldExtractor cache for sequential path — same pattern as parse-worker.ts
const seqFieldInfoCache = new Map<number, Map<string, FieldInfo>>();
// MethodExtractor cache for sequential path — avoids re-traversing the same class
// body once per method. Keyed on classNode.id (tree-sitter node identity number).
const seqMethodExtractCache = new Map<
number,
{ ownerName: string | undefined; methods: MethodInfo[] } | null
>();
// Derived method map + collision groups cache — avoids rebuilding per method.
const seqMethodMapCache = new Map<
number,
{ map: Map<string, MethodInfo>; groups: Map<string, MethodInfo[]> }
>();
function seqFindEnclosingClassNode(node: SyntaxNode): SyntaxNode | null {
let current = node.parent;
while (current) {
@@ -259,9 +285,11 @@ const processParsingSequential = async (
const file = files[i];
// Reset memoization before each new file (node refs are per-tree)
classIdCache.clear();
classInfoCache.clear();
exportCache.clear();
seqFieldInfoCache.clear();
seqMethodExtractCache.clear();
seqMethodMapCache.clear();
onFileProgress?.(i + 1, total, file.path);
@@ -280,6 +308,18 @@ const processParsingSequential = async (
// Skip files larger than the max tree-sitter buffer (32 MB)
if (file.content.length > TREE_SITTER_MAX_BUFFER) continue;
// Vue SFC preprocessing: extract <script> block content
let parseContent = file.content;
let lineOffset = 0;
let isVueSetup = false;
if (language === SupportedLanguages.Vue) {
const extracted = extractVueScript(file.content);
if (!extracted) continue; // skip .vue files with no script block
parseContent = extracted.scriptContent;
lineOffset = extracted.lineOffset;
isVueSetup = extracted.isSetup;
}
try {
await loadLanguage(language, file.path);
} catch {
@@ -288,8 +328,8 @@ const processParsingSequential = async (
let tree;
try {
tree = parser.parse(file.content, undefined, {
bufferSize: getTreeSitterBufferSize(file.content.length),
tree = parser.parse(parseContent, undefined, {
bufferSize: getTreeSitterBufferSize(parseContent.length),
});
} catch (parseError) {
console.warn(`Skipping unparseable file: ${file.path}`);
@@ -315,9 +355,19 @@ const processParsingSequential = async (
continue;
}
// Build per-file type environment for FieldExtractor context (lightweight — skipped if no fieldExtractor)
// Build per-file type environment for FieldExtractor context (lightweight — skipped if no fieldExtractor).
//
// Note: this TypeEnv is intentionally NOT flushed into the BindingAccumulator.
// The accumulator feed happens later in `call-processor.ts` via its own
// `typeEnv.flush(accumulator)` call. Flushing here would double-count
// file-scope bindings and break the single-use invariant of `flush()`.
// See the BindingAccumulator class JSDoc for the full accumulator
// lifecycle and flush-site ownership rules.
const typeEnv = provider.fieldExtractor
? buildTypeEnv(tree, language, { enclosingFunctionFinder: provider.enclosingFunctionFinder })
? buildTypeEnv(tree, language, {
enclosingFunctionFinder: provider.enclosingFunctionFinder,
extractFunctionName: provider.methodExtractor?.extractFunctionName,
})
: null;
matches.forEach((match) => {
@@ -327,94 +377,184 @@ const processParsingSequential = async (
captureMap[c.name] = c.node;
});
const nodeLabel = getLabelFromCaptures(captureMap, provider);
if (!nodeLabel) return;
const definitionNodeForRange = getDefinitionNodeFromCaptures(captureMap);
const definitionNode = getDefinitionNodeFromCaptures(captureMap);
const defaultNodeLabel = getLabelFromCaptures(captureMap, provider);
if (!defaultNodeLabel) return;
const nameNode = captureMap['name'];
const extractedClassSymbol =
definitionNode && provider.classExtractor?.isTypeDeclaration(definitionNode)
? provider.classExtractor.extract(definitionNode, {
name: nameNode?.text,
type: defaultNodeLabel,
})
: null;
const nodeLabel = extractedClassSymbol?.type ?? defaultNodeLabel;
// Synthesize name for constructors without explicit @name capture (e.g. Swift init)
if (!nameNode && nodeLabel !== 'Constructor') return;
const nodeName = nameNode ? nameNode.text : 'init';
if (!nameNode && nodeLabel !== 'Constructor' && !extractedClassSymbol) return;
const nodeName = extractedClassSymbol?.name ?? (nameNode ? nameNode.text : 'init');
const definitionNodeForRange = getDefinitionNodeFromCaptures(captureMap);
const startLine = definitionNodeForRange
? definitionNodeForRange.startPosition.row
? definitionNodeForRange.startPosition.row + lineOffset
: nameNode
? nameNode.startPosition.row
: 0;
const nodeId = generateId(nodeLabel, `${file.path}:${nodeName}`);
? nameNode.startPosition.row + lineOffset
: lineOffset;
const definitionNode = getDefinitionNodeFromCaptures(captureMap);
// Compute enclosing class BEFORE node ID — needed to qualify method IDs
const needsOwner =
nodeLabel === 'Method' ||
nodeLabel === 'Constructor' ||
nodeLabel === 'Property' ||
nodeLabel === 'Function';
const enclosingClassInfo = needsOwner
? cachedFindEnclosingClassInfo(nameNode || definitionNodeForRange, file.path)
: null;
const enclosingClassId = enclosingClassInfo?.classId ?? null;
// Qualify method/property IDs with enclosing class name to avoid collisions
// e.g. "Method:animal.dart:Animal.speak" vs "Method:animal.dart:Dog.speak"
const qualifiedName = enclosingClassInfo
? `${enclosingClassInfo.className}.${nodeName}`
: nodeName;
// Extract method metadata for Function/Method/Constructor nodes BEFORE generating
// the node ID — parameterCount is needed to disambiguate overloaded methods.
// Use the per-language MethodExtractor for method metadata (isAbstract, isStatic,
// visibility, annotations, parameterCount, parameterTypes, returnType, etc.).
const isMethodLike =
nodeLabel === 'Function' || nodeLabel === 'Method' || nodeLabel === 'Constructor';
let methodProps: Record<string, unknown> = {};
let arityForId: number | undefined; // raw param count for ID, even for variadic
let seqDefMethodInfo: MethodInfo | undefined;
let seqDefMethods: MethodInfo[] | undefined;
let seqClassNodeId: number | undefined;
if (isMethodLike && definitionNode) {
let enriched = false;
if (provider.methodExtractor) {
// Try class-based extraction (method inside a class/struct/trait body)
const classNode = seqFindEnclosingClassNode(definitionNode);
if (classNode) {
// Cache extract() results per class node to avoid re-traversing the
// same class body for every method it contains (O(N) -> O(1) per hit).
let result:
| { ownerName: string | undefined; methods: MethodInfo[] }
| null
| undefined = seqMethodExtractCache.get(classNode.id);
if (result === undefined) {
result =
provider.methodExtractor.extract(classNode, {
filePath: file.path,
language,
}) ?? null;
seqMethodExtractCache.set(classNode.id, result);
}
if (result?.methods?.length) {
const defLine = definitionNode.startPosition.row + 1;
const info = result.methods.find((m) => m.name === nodeName && m.line === defLine);
if (info) {
enriched = true;
arityForId = arityForIdFromInfo(info);
methodProps = buildMethodProps(info);
seqDefMethodInfo = info;
seqDefMethods = result.methods;
seqClassNodeId = classNode.id;
}
}
}
// For top-level methods (e.g. Go method_declaration), try extractFromNode
if (!enriched && provider.methodExtractor.extractFromNode) {
const info = provider.methodExtractor.extractFromNode(definitionNode, {
filePath: file.path,
language,
});
if (info) {
enriched = true;
arityForId = arityForIdFromInfo(info);
methodProps = buildMethodProps(info);
}
}
}
}
// Append #<paramCount> to Method/Constructor IDs to disambiguate overloads.
// Functions are not suffixed — they don't overload by name in the same scope.
// When same-arity collisions exist, append ~type1,type2 for further disambiguation.
const needsAritySuffix = nodeLabel === 'Method' || nodeLabel === 'Constructor';
let arityTag = needsAritySuffix && arityForId !== undefined ? `#${arityForId}` : '';
if (arityTag && seqDefMethods && seqDefMethodInfo && seqClassNodeId !== undefined) {
// Use cached method map + collision groups (built once per class, not per method)
let cached = seqMethodMapCache.get(seqClassNodeId);
if (!cached) {
const tempMap = new Map<string, MethodInfo>();
for (const m of seqDefMethods) tempMap.set(`${m.name}:${m.line}`, m);
cached = { map: tempMap, groups: buildCollisionGroups(tempMap) };
seqMethodMapCache.set(seqClassNodeId, cached);
}
arityTag += typeTagForId(
cached.map,
nodeName,
arityForId,
seqDefMethodInfo,
language,
cached.groups,
);
arityTag += constTagForId(
cached.map,
nodeName,
arityForId,
seqDefMethodInfo,
cached.groups,
);
}
const nodeId = generateId(nodeLabel, `${file.path}:${qualifiedName}${arityTag}`);
const classNodeForSymbol = definitionNodeForRange || definitionNode || nameNode;
const qualifiedTypeName =
extractedClassSymbol?.qualifiedName ??
(classNodeForSymbol && provider.classExtractor?.isTypeDeclaration(classNodeForSymbol)
? (provider.classExtractor.extractQualifiedName(classNodeForSymbol, nodeName) ?? nodeName)
: undefined);
const frameworkHint = definitionNode
? detectFrameworkFromAST(language, (definitionNode.text || '').slice(0, 300))
: null;
// Extract method signature for Method/Constructor nodes
const methodSig =
nodeLabel === 'Function' || nodeLabel === 'Method' || nodeLabel === 'Constructor'
? extractMethodSignature(definitionNode)
: undefined;
// Language-specific return type fallback (e.g. Ruby YARD @return [Type])
// Also upgrades uninformative AST types like PHP `array` with PHPDoc `@return User[]`
if (
methodSig &&
(!methodSig.returnType ||
methodSig.returnType === 'array' ||
methodSig.returnType === 'iterable') &&
definitionNode
) {
const tc = provider.typeConfig;
if (tc?.extractReturnType) {
const docReturn = tc.extractReturnType(definitionNode);
if (docReturn) methodSig.returnType = docReturn;
}
}
const node: GraphNode = {
id: nodeId,
label: nodeLabel as NodeLabel,
properties: {
name: nodeName,
filePath: file.path,
startLine: definitionNodeForRange ? definitionNodeForRange.startPosition.row : startLine,
endLine: definitionNodeForRange ? definitionNodeForRange.endPosition.row : startLine,
startLine: definitionNodeForRange
? definitionNodeForRange.startPosition.row + lineOffset
: startLine,
endLine: definitionNodeForRange
? definitionNodeForRange.endPosition.row + lineOffset
: startLine,
language: language,
isExported: cachedExportCheck(
provider.exportChecker,
nameNode || definitionNodeForRange,
nodeName,
),
isExported:
language === SupportedLanguages.Vue && isVueSetup
? isVueSetupTopLevel(nameNode || definitionNodeForRange)
: cachedExportCheck(
provider.exportChecker,
nameNode || definitionNodeForRange,
nodeName,
),
...(qualifiedTypeName !== undefined ? { qualifiedName: qualifiedTypeName } : {}),
...(frameworkHint
? {
astFrameworkMultiplier: frameworkHint.entryPointMultiplier,
astFrameworkReason: frameworkHint.reason,
}
: {}),
...(methodSig
? {
parameterCount: methodSig.parameterCount,
...(methodSig.requiredParameterCount !== undefined
? { requiredParameterCount: methodSig.requiredParameterCount }
: {}),
...(methodSig.parameterTypes ? { parameterTypes: methodSig.parameterTypes } : {}),
returnType: methodSig.returnType,
}
: {}),
...methodProps,
},
};
graph.addNode(node);
// Compute enclosing class for Method/Constructor/Property/Function — used for both ownerId and HAS_METHOD
// Function is included because Kotlin/Rust/Python capture class methods as Function nodes
const needsOwner =
nodeLabel === 'Method' ||
nodeLabel === 'Constructor' ||
nodeLabel === 'Property' ||
nodeLabel === 'Function';
const enclosingClassId = needsOwner
? cachedFindEnclosingClassId(nameNode || definitionNodeForRange, file.path)
: null;
// enclosingClassId already computed above (before nodeId generation)
// Extract declared type and field metadata for Property nodes
let declaredType: string | undefined;
@@ -441,7 +581,7 @@ const processParsingSequential = async (
}
}
}
// All 14 languages register a FieldExtractor — no fallback needed.
// All 15 tree-sitter languages register a FieldExtractor — no fallback needed.
}
// Apply field metadata to the graph node retroactively
@@ -451,12 +591,13 @@ const processParsingSequential = async (
if (declaredType !== undefined) node.properties.declaredType = declaredType;
symbolTable.add(file.path, nodeName, nodeId, nodeLabel, {
parameterCount: methodSig?.parameterCount,
requiredParameterCount: methodSig?.requiredParameterCount,
parameterTypes: methodSig?.parameterTypes,
returnType: methodSig?.returnType,
parameterCount: methodProps.parameterCount as number | undefined,
requiredParameterCount: methodProps.requiredParameterCount as number | undefined,
parameterTypes: methodProps.parameterTypes as string[] | undefined,
returnType: methodProps.returnType as string | undefined,
declaredType,
ownerId: enclosingClassId ?? undefined,
qualifiedName: qualifiedTypeName,
});
const fileId = generateId('File', file.path);
+177 -62
View File
@@ -1,4 +1,9 @@
import { createKnowledgeGraph } from '../graph/graph.js';
import {
BindingAccumulator,
enrichExportedTypeMap,
type BindingEntry,
} from './binding-accumulator.js';
import { processStructure } from './structure-processor.js';
import { processMarkdown } from './markdown-processor.js';
import { processCobol, isCobolFile, isJclFile } from './cobol-processor.js';
@@ -21,9 +26,8 @@ import {
buildImportedRawReturnTypes,
type ExportedTypeMap,
buildExportedTypeMapFromGraph,
buildImplementorMap,
mergeImplementorMaps,
} from './call-processor.js';
import { buildHeritageMap } from './heritage-map.js';
import { nextjsFileToRouteURL, normalizeFetchURL } from './route-extractors/nextjs.js';
import { expoFileToRouteURL } from './route-extractors/expo.js';
import { phpFileToRouteURL } from './route-extractors/php.js';
@@ -484,6 +488,8 @@ async function runCrossFileBindingPropagation(
export interface PipelineOptions {
/** Skip MRO, community detection, and process extraction for faster test runs. */
skipGraphPhases?: boolean;
/** Force sequential parsing (no worker pool). Useful for testing the sequential path. */
skipWorkers?: boolean;
}
// ── Extracted pipeline phases ──────────────────────────────────────────────
@@ -641,6 +647,7 @@ async function runChunkedParseAndResolve(
repoPath: string,
pipelineStart: number,
onProgress: ProgressFn,
options?: PipelineOptions,
): Promise<{
exportedTypeMap: ExportedTypeMap;
allFetchCalls: ExtractedFetchCall[];
@@ -648,6 +655,7 @@ async function runChunkedParseAndResolve(
allDecoratorRoutes: ExtractedDecoratorRoute[];
allToolDefs: ExtractedToolDef[];
allORMQueries: ExtractedORMQuery[];
bindingAccumulator: BindingAccumulator;
}> {
const symbolTable = ctx.symbols;
@@ -719,7 +727,10 @@ async function runChunkedParseAndResolve(
// Create worker pool once, reuse across chunks
let workerPool: WorkerPool | undefined;
if (totalParseable >= MIN_FILES_FOR_WORKERS || totalBytes >= MIN_BYTES_FOR_WORKERS) {
if (
!options?.skipWorkers &&
(totalParseable >= MIN_FILES_FOR_WORKERS || totalBytes >= MIN_BYTES_FOR_WORKERS)
) {
try {
let workerUrl = new URL('./workers/parse-worker.js', import.meta.url);
// When running under vitest, import.meta.url points to src/ where no .js exists.
@@ -777,7 +788,7 @@ async function runChunkedParseAndResolve(
// Phase 14: Collect exported type bindings for cross-file propagation
const exportedTypeMap: ExportedTypeMap = new Map();
// Accumulate file-scope TypeEnv bindings from workers (closes worker/sequential quality gap)
const workerTypeEnvBindings: { filePath: string; bindings: [string, string][] }[] = [];
const bindingAccumulator = new BindingAccumulator();
// Accumulate fetch() calls from workers for Next.js route matching
const allFetchCalls: ExtractedFetchCall[] = [];
// Accumulate framework-extracted routes (Laravel, etc.) for Route node creation
@@ -873,11 +884,12 @@ async function runChunkedParseAndResolve(
);
}
}
deferredWorkerCalls.push(...chunkWorkerData.calls);
deferredWorkerHeritage.push(...chunkWorkerData.heritage);
deferredConstructorBindings.push(...chunkWorkerData.constructorBindings);
for (const _item of chunkWorkerData.calls) deferredWorkerCalls.push(_item);
for (const _item of chunkWorkerData.heritage) deferredWorkerHeritage.push(_item);
for (const _item of chunkWorkerData.constructorBindings)
deferredConstructorBindings.push(_item);
if (chunkWorkerData.assignments?.length) {
deferredAssignments.push(...chunkWorkerData.assignments);
for (const _item of chunkWorkerData.assignments) deferredAssignments.push(_item);
}
// Heritage + Routes — calls deferred until all chunks have contributed heritage
@@ -910,25 +922,47 @@ async function runChunkedParseAndResolve(
});
}),
]);
// Collect TypeEnv file-scope bindings for exported type enrichment
if (chunkWorkerData.typeEnvBindings?.length) {
workerTypeEnvBindings.push(...chunkWorkerData.typeEnvBindings);
// Collect file-scope bindings into BindingAccumulator. The worker
// IPC payload carries only file-scope entries (`scope = ''`
// hardcoded here). See the FileScopeBindings JSDoc in
// parse-worker.ts for the rationale and Phase 9 reversion path.
//
// Defensive validation at the IPC boundary: silently skip entries
// with non-string varName/typeName. If a future worker regression
// (or a Phase 9 reversion mistake that emits 3-tuples into the
// 2-tuple consumer) produces malformed data, logging is better
// than silently writing `undefined` into the enrichment map.
if (chunkWorkerData.fileScopeBindings?.length) {
for (const { filePath, bindings } of chunkWorkerData.fileScopeBindings) {
if (typeof filePath !== 'string' || filePath.length === 0) continue;
if (!Array.isArray(bindings)) continue;
const entries: BindingEntry[] = [];
for (const tuple of bindings) {
if (!Array.isArray(tuple) || tuple.length !== 2) continue;
const [varName, typeName] = tuple;
if (typeof varName !== 'string' || typeof typeName !== 'string') continue;
entries.push({ scope: '', varName, typeName });
}
if (entries.length > 0) {
bindingAccumulator.appendFile(filePath, entries);
}
}
}
// Collect fetch() calls for Next.js route matching
if (chunkWorkerData.fetchCalls?.length) {
allFetchCalls.push(...chunkWorkerData.fetchCalls);
for (const _item of chunkWorkerData.fetchCalls) allFetchCalls.push(_item);
}
if (chunkWorkerData.routes?.length) {
allExtractedRoutes.push(...chunkWorkerData.routes);
for (const _item of chunkWorkerData.routes) allExtractedRoutes.push(_item);
}
if (chunkWorkerData.decoratorRoutes?.length) {
allDecoratorRoutes.push(...chunkWorkerData.decoratorRoutes);
for (const _item of chunkWorkerData.decoratorRoutes) allDecoratorRoutes.push(_item);
}
if (chunkWorkerData.toolDefs?.length) {
allToolDefs.push(...chunkWorkerData.toolDefs);
for (const _item of chunkWorkerData.toolDefs) allToolDefs.push(_item);
}
if (chunkWorkerData.ormQueries?.length) {
allORMQueries.push(...chunkWorkerData.ormQueries);
for (const _item of chunkWorkerData.ormQueries) allORMQueries.push(_item);
}
} else {
await processImports(graph, chunkFiles, astCache, ctx, undefined, repoPath, allPaths);
@@ -942,11 +976,9 @@ async function runChunkedParseAndResolve(
// chunkContents + chunkFiles + chunkWorkerData go out of scope → GC reclaims
}
// Complete implementor map from all worker heritage, then resolve CALLS once (interface dispatch).
const fullWorkerImplementorMap =
deferredWorkerHeritage.length > 0
? buildImplementorMap(deferredWorkerHeritage, ctx)
: new Map<string, Set<string>>();
// Build unified HeritageMap (parent lookup + implementor index) after all chunks.
const fullWorkerHeritageMap =
deferredWorkerHeritage.length > 0 ? buildHeritageMap(deferredWorkerHeritage, ctx) : undefined;
if (deferredWorkerCalls.length > 0) {
await processCallsFromExtracted(
@@ -967,7 +999,7 @@ async function runChunkedParseAndResolve(
});
},
deferredConstructorBindings.length > 0 ? deferredConstructorBindings : undefined,
fullWorkerImplementorMap,
fullWorkerHeritageMap,
);
}
@@ -987,17 +1019,38 @@ async function runChunkedParseAndResolve(
// Synthesize wildcard import bindings once after ALL imports are processed,
// before any call resolution — same rationale as the worker-path inline synthesis.
if (sequentialChunkPaths.length > 0) synthesizeWildcardImportBindings(graph, ctx);
// Merge implementor-map deltas per chunk (O(heritage per chunk)), not O(|edges|) graph scans
// per chunk — mirrors worker-path deferred heritage without re-iterating all relationships.
const sequentialImplementorMap = new Map<string, Set<string>>();
// Pass 1: Extract heritage from all sequential chunks.
// Heritage must be fully accumulated BEFORE call resolution so the HeritageMap
// has the complete ancestor chain and implementor index (same constraint as
// the worker path).
//
// File contents are read once here and cached for Pass 2 to avoid a 2× I/O
// cost on the sequential path (ASTs are intentionally NOT cached — rebuilding
// them in Pass 2 keeps peak memory bounded to one chunk at a time).
const allSequentialHeritage: ExtractedHeritage[] = [];
const cachedSequentialChunkFiles: Array<Array<{ path: string; content: string }>> = [];
for (const chunkPaths of sequentialChunkPaths) {
const chunkContents = await readFileContents(repoPath, chunkPaths);
const chunkFiles = chunkPaths
.filter((p) => chunkContents.has(p))
.map((p) => ({ path: p, content: chunkContents.get(p)! }));
cachedSequentialChunkFiles.push(chunkFiles);
astCache = createASTCache(chunkFiles.length);
const sequentialHeritage = await extractExtractedHeritageFromFiles(chunkFiles, astCache);
mergeImplementorMaps(sequentialImplementorMap, buildImplementorMap(sequentialHeritage, ctx));
// Manual loop (not spread) — `push(...arr)` blows the stack on very large
// arrays, see #650. Pay the explicit iteration cost for safety.
for (const h of sequentialHeritage) allSequentialHeritage.push(h);
astCache.clear();
}
// Build unified HeritageMap from all sequential heritage (parent lookup + implementor index).
const sequentialHeritageMap =
allSequentialHeritage.length > 0 ? buildHeritageMap(allSequentialHeritage, ctx) : undefined;
// Pass 2: Process calls, heritage edges, fetch calls, and ORM queries per chunk.
// Reuse the file contents cached in Pass 1 instead of re-reading from disk.
for (let chunkIdx = 0; chunkIdx < sequentialChunkPaths.length; chunkIdx++) {
const chunkFiles = cachedSequentialChunkFiles[chunkIdx];
astCache = createASTCache(chunkFiles.length);
const rubyHeritage = await processCalls(
graph,
chunkFiles,
@@ -1008,7 +1061,8 @@ async function runChunkedParseAndResolve(
undefined,
undefined,
undefined,
sequentialImplementorMap,
sequentialHeritageMap,
bindingAccumulator,
);
await processHeritage(graph, chunkFiles, astCache, ctx);
if (rubyHeritage.length > 0) {
@@ -1017,13 +1071,17 @@ async function runChunkedParseAndResolve(
// Extract fetch() calls for Next.js route matching (sequential path)
const chunkFetchCalls = await extractFetchCallsFromFiles(chunkFiles, astCache);
if (chunkFetchCalls.length > 0) {
allFetchCalls.push(...chunkFetchCalls);
for (const _item of chunkFetchCalls) allFetchCalls.push(_item);
}
// Extract ORM queries (sequential path)
for (const f of chunkFiles) {
extractORMQueriesInline(f.path, f.content, allORMQueries);
}
astCache.clear();
// Release cached chunk content as soon as Pass 2 finishes with it so the
// Pass-1 content map drains incrementally rather than being held for the
// full duration of Pass 2.
cachedSequentialChunkFiles[chunkIdx] = [];
}
// Log resolution cache stats
@@ -1034,41 +1092,36 @@ async function runChunkedParseAndResolve(
console.log(
`🔍 Resolution cache: ${rcStats.cacheHits} hits, ${rcStats.cacheMisses} misses (${hitRate}% hit rate)`,
);
console.log(
`🔍 Fuzzy Lookups: ${rcStats.fuzzyCallCount} total, ${rcStats.fuzzyCallableCallCount} callable`,
);
}
// ── Worker path quality enrichment: merge TypeEnv file-scope bindings into ExportedTypeMap ──
// Workers return file-scope bindings from their TypeEnv fixpoint (includes inferred types
// like `const config = getConfig()` → Config). Filter by graph isExported to match
// the sequential path's collectExportedBindings behavior.
if (workerTypeEnvBindings.length > 0) {
let enriched = 0;
for (const { filePath, bindings } of workerTypeEnvBindings) {
for (const [name, type] of bindings) {
// Verify the symbol is exported via graph node
const nodeId = `Function:${filePath}:${name}`;
const varNodeId = `Variable:${filePath}:${name}`;
const constNodeId = `Const:${filePath}:${name}`;
const node =
graph.getNode(nodeId) ?? graph.getNode(varNodeId) ?? graph.getNode(constNodeId);
if (!node?.properties?.isExported) continue;
// ── Finalize the accumulator before the read phase begins. All worker-path
// appends (line ~934) and sequential-path flushes (via `processCalls` →
// `typeEnv.flush()` earlier in this function) have completed by here,
// so the finalize-write-lock is correct at this seam. Making the
// lifecycle contract explicit — `append → finalize → consume → dispose`.
// Previously `finalize()` was called much later in `runPipelineFromRepo`
// after the enrichment loop had already read the mutable accumulator.
bindingAccumulator.finalize();
let fileExports = exportedTypeMap.get(filePath);
if (!fileExports) {
fileExports = new Map();
exportedTypeMap.set(filePath, fileExports);
}
// Don't overwrite existing entries (Tier 0 from SymbolTable is authoritative)
if (!fileExports.has(name)) {
fileExports.set(name, type);
enriched++;
}
}
}
if (isDev && enriched > 0) {
console.log(
`🔗 Worker TypeEnv enrichment: ${enriched} fixpoint-inferred exports added to ExportedTypeMap`,
);
}
// ── Worker path quality enrichment: merge file-scope bindings into ExportedTypeMap ──
// Counterpart to `collectExportedBindings()` in call-processor.ts which
// handles the sequential path (main thread, full SymbolTable access).
// This call handles the worker path via the accumulator. Both sites
// populate the same `exportedTypeMap` with subtly different export-check
// semantics — sequential uses SymbolTable + graph lookup, `enrichExportedTypeMap`
// uses a three-candidate-ID graph lookup. They must stay in sync until
// Phase 9 unifies them. If you edit one, check the other.
//
// The enrichment loop itself lives in `binding-accumulator.ts` so tests
// can exercise the real production code instead of reimplementing it.
const enriched = enrichExportedTypeMap(bindingAccumulator, graph, exportedTypeMap);
if (isDev && enriched > 0) {
console.log(
`🔗 Worker TypeEnv enrichment: ${enriched} fixpoint-inferred exports added to ExportedTypeMap`,
);
}
// ── Final synthesis pass for whole-module-import languages ──
@@ -1096,6 +1149,7 @@ async function runChunkedParseAndResolve(
allDecoratorRoutes,
allToolDefs,
allORMQueries,
bindingAccumulator,
};
}
@@ -1103,7 +1157,7 @@ async function runChunkedParseAndResolve(
* Post-parse graph analysis: MRO, community detection, process extraction.
*
* @reads graph (all nodes and relationships from parse + resolve phases)
* @writes graph (Community nodes, Process nodes, MEMBER_OF edges, STEP_IN_PROCESS edges, OVERRIDES edges)
* @writes graph (Community nodes, Process nodes, MEMBER_OF edges, STEP_IN_PROCESS edges, METHOD_OVERRIDES edges)
*/
async function runGraphAnalysisPhases(
graph: ReturnType<typeof createKnowledgeGraph>,
@@ -1126,7 +1180,7 @@ async function runGraphAnalysisPhases(
const mroResult = computeMRO(graph);
if (isDev && mroResult.entries.length > 0) {
console.log(
`🔀 MRO: ${mroResult.entries.length} classes analyzed, ${mroResult.ambiguityCount} ambiguities found, ${mroResult.overrideEdges} OVERRIDES edges`,
`🔀 MRO: ${mroResult.entries.length} classes analyzed, ${mroResult.ambiguityCount} ambiguities, ${mroResult.overrideEdges} METHOD_OVERRIDES, ${mroResult.methodImplementsEdges} METHOD_IMPLEMENTS`,
);
}
@@ -1327,6 +1381,15 @@ export const runPipelineFromRepo = async (
const ctx = createResolutionContext();
const pipelineStart = Date.now();
// Hoisted reference for error-path cleanup. The accumulator is normally
// disposed at the happy-path seam after the dev telemetry log, but if any
// step between the runChunkedParseAndResolve return and that seam throws
// (ORM processing, tool node creation, Phase 14, graph analysis), the
// catch handler disposes it here so the heap footprint does not leak
// through the rethrow. See binding-accumulator.ts dispose() JSDoc for the
// lifecycle contract.
let bindingAccumulatorForCleanup: BindingAccumulator | undefined;
try {
// Phase 1+2: Scan paths, build structure, process markdown
const { scannedFiles, allPaths, totalFiles } = await runScanAndStructure(
@@ -1343,6 +1406,7 @@ export const runPipelineFromRepo = async (
allDecoratorRoutes,
allToolDefs,
allORMQueries,
bindingAccumulator,
} = await runChunkedParseAndResolve(
graph,
ctx,
@@ -1352,7 +1416,13 @@ export const runPipelineFromRepo = async (
repoPath,
pipelineStart,
onProgress,
options,
);
// Track the accumulator for error-path cleanup — the happy-path dispose
// is still at the post-telemetry seam below, this reference is only
// consulted by the catch handler if any step between here and there
// throws.
bindingAccumulatorForCleanup = bindingAccumulator;
// ── Phase 3.5: Route Registry (Next.js + PHP + Laravel + decorators) ──
type RouteEntry = { filePath: string; source: string };
@@ -1664,6 +1734,45 @@ export const runPipelineFromRepo = async (
processORMQueries(graph, allORMQueries, isDev);
}
// `bindingAccumulator.finalize()` was moved inside `runChunkedParseAndResolve`
// to immediately precede the enrichment loop — see the comment there for
// the ordering rationale. By the time execution
// reaches this point, the accumulator has already been finalized, consumed
// by the enrichment loop, and is ready for dispose() below after the dev
// telemetry log captures peak state.
if (isDev) {
if (bindingAccumulator.totalBindings > 0) {
const memKB = Math.round(bindingAccumulator.estimateMemoryBytes() / 1024);
console.log(
`📦 BindingAccumulator: ${bindingAccumulator.totalBindings} bindings across ${bindingAccumulator.fileCount} files (~${memKB} KB)`,
);
} else if (totalFiles > 0) {
// Zero-binding signal: if the pipeline parsed files but the
// accumulator is empty, something upstream dropped all bindings.
// Flag it so operators can spot a regression (e.g. a worker path
// that accidentally emits empty fileScopeBindings arrays for every
// file, or a TypeEnv build failure). Dev-mode only.
console.log(
`📦 BindingAccumulator: EMPTY — 0 bindings across 0 files despite ${totalFiles} parsed files. If the codebase has typed bindings, this indicates an upstream regression.`,
);
}
}
// Release the accumulator's heap footprint now. The ExportedTypeMap
// enrichment loop above is the only current consumer, and the dev
// telemetry log just captured peak state. Phase 14 and
// runGraphAnalysisPhases do not read the accumulator today — keeping
// it alive through those long-running phases pins heap for no reason.
// When Phase 9 wires a consumer into runCrossFileBindingPropagation,
// move this dispose() call to after that consumer completes or delete
// it entirely if the consumer takes lifecycle ownership.
bindingAccumulator.dispose();
// Happy-path dispose completed — clear the cleanup ref so the catch
// handler doesn't attempt a second (harmless but noisy) dispose if a
// later phase throws.
bindingAccumulatorForCleanup = undefined;
// ── Phase 14: Cross-file binding propagation (topological level sort) ──
await runCrossFileBindingPropagation(
graph,
@@ -1708,6 +1817,12 @@ export const runPipelineFromRepo = async (
return { graph, repoPath, totalFileCount: totalFiles, communityResult, processResult };
} catch (error) {
// Error-path cleanup: dispose the accumulator if a step after the
// destructure from runChunkedParseAndResolve but before the happy-path
// dispose threw. The reference is cleared on the happy path, so this
// is a no-op when the pipeline completed successfully and then threw
// from an unrelated post-dispose step (e.g., future cleanup code).
bindingAccumulatorForCleanup?.dispose();
ctx.clear();
throw error;
}
@@ -70,6 +70,8 @@ export interface ResolutionContext {
getStats(): {
fileCount: number;
globalSymbolCount: number;
fuzzyCallCount: number;
fuzzyCallableCallCount: number;
cacheHits: number;
cacheMisses: number;
};
+183 -1
View File
@@ -1,9 +1,27 @@
import type { NodeLabel } from 'gitnexus-shared';
export const CLASS_TYPES = new Set([
'Class',
'Struct',
'Interface',
'Enum',
'Record',
// Traits are class-like for heritage resolution: PHP `use Trait;`, Rust
// `impl Trait for Struct`, and Scala traits all contribute methods to the
// hierarchy of their using/implementing type. Including Trait here lets
// buildHeritageMap resolve `h.parentName` to a Trait nodeId so the MRO
// walker can visit the trait and find its methods.
'Trait',
]);
export interface SymbolDefinition {
nodeId: string;
filePath: string;
type: NodeLabel;
/** Canonical dot-separated qualified type name for class-like symbols
* (e.g. `App.Models.User`). Falls back to the simple symbol name when no
* package/namespace/module scope exists or no explicit qualified metadata is provided. */
qualifiedName?: string;
parameterCount?: number;
/** Number of required (non-optional, non-default) parameters.
* Enables range-based arity filtering: argCount >= requiredParameterCount && argCount <= parameterCount. */
@@ -36,6 +54,7 @@ export interface SymbolTable {
returnType?: string;
declaredType?: string;
ownerId?: string;
qualifiedName?: string;
},
) => void;
@@ -79,10 +98,57 @@ export interface SymbolTable {
*/
lookupFieldByOwner: (ownerNodeId: string, fieldName: string) => SymbolDefinition | undefined;
/**
* Look up a method by its owning class nodeId and method name.
* O(1) via dedicated eagerly-populated index keyed by `ownerNodeId\0methodName`.
* For overloaded methods (same owner + name): returns the first match when all
* overloads share the same returnType, undefined when return types differ (ambiguous).
* Used by walkMixedChain for deterministic cross-class chain resolution.
*/
/**
* Lookup a method by owner class + name, optionally filtered by arity.
*
* When `argCount` is provided, overloads whose parameter count doesn't
* accommodate the call's argument count are filtered out before the
* returnType dedup runs. This lets D0 (`resolveMemberCall`) disambiguate
* arity-differing overloads (e.g. C++ `greet()` vs `greet(string)`) that
* would otherwise collide on the shared `ownerId + methodName` key.
*
* Same-arity, same-returnType overloads (e.g. `save(int)` vs `save(String)`,
* both returning `void`) still collapse to the first match — callers must
* gate D0 on overload concern before invoking this function for that case.
*/
lookupMethodByOwner: (
ownerNodeId: string,
methodName: string,
argCount?: number,
) => SymbolDefinition | undefined;
/**
* Look up class-like definitions (Class, Struct, Interface, Enum, Record) by name.
* O(1) via dedicated eagerly-populated index keyed by symbol name.
* Returns all matching definitions across files (e.g. partial classes).
* Used by Phase 1 semantic-model tasks to replace filtered lookupFuzzy calls.
*/
lookupClassByName: (name: string) => SymbolDefinition[];
/**
* Look up class-like definitions by canonical qualified name.
* Qualified names are normalized to dot-separated scope segments across languages,
* e.g. `App.Models.User`, `com.example.User`, or `Admin.User`.
* Top-level class-like symbols with no explicit scope are indexed under their simple name.
*/
lookupClassByQualifiedName: (qualifiedName: string) => SymbolDefinition[];
/**
* Debugging: See how many symbols are tracked
*/
getStats: () => { fileCount: number; globalSymbolCount: number };
getStats: () => {
fileCount: number;
globalSymbolCount: number;
fuzzyCallCount: number;
fuzzyCallableCallCount: number;
};
/**
* Cleanup memory
@@ -109,6 +175,19 @@ export const createSymbolTable = (): SymbolTable => {
// Only Property symbols with ownerId and declaredType are indexed.
const fieldByOwner = new Map<string, SymbolDefinition>();
// 5. Eagerly-populated Method Index — keyed by "ownerNodeId\0methodName".
// Method symbols with ownerId are indexed. Supports overloads (array values).
const methodByOwner = new Map<string, SymbolDefinition[]>();
// 6. Eagerly-populated Class-type Index — keyed by symbol name.
// Only Class, Struct, Interface, Enum, Record symbols are indexed.
const classByName = new Map<string, SymbolDefinition[]>();
const classByQualifiedName = new Map<string, SymbolDefinition[]>();
let fuzzyCallCount = 0;
let fuzzyCallableCallCount = 0;
const CALLABLE_TYPES = new Set(['Function', 'Method', 'Constructor']);
const add = (
@@ -123,12 +202,17 @@ export const createSymbolTable = (): SymbolTable => {
returnType?: string;
declaredType?: string;
ownerId?: string;
qualifiedName?: string;
},
) => {
const qualifiedName = CLASS_TYPES.has(type)
? (metadata?.qualifiedName ?? name)
: metadata?.qualifiedName;
const def: SymbolDefinition = {
nodeId,
filePath,
type,
...(qualifiedName !== undefined ? { qualifiedName } : {}),
...(metadata?.parameterCount !== undefined
? { parameterCount: metadata.parameterCount }
: {}),
@@ -170,6 +254,43 @@ export const createSymbolTable = (): SymbolTable => {
}
globalIndex.get(name)!.push(def);
// C2. Methods, constructors, and ownerId-bound Functions go to
// methodByOwner index (in addition to globalIndex).
//
// Some language extractors emit class methods as `Function` with an
// `ownerId` — notably Python (`def method(self):` inside a class body),
// Rust trait methods, and Kotlin object/companion methods. Treating
// `Function` with ownerId the same as `Method` here makes D0
// (`resolveMemberCall`) work uniformly across all supported languages
// instead of silently falling through to D1-D4 fuzzy widening.
if ((type === 'Method' || type === 'Constructor' || type === 'Function') && metadata?.ownerId) {
const key = `${metadata.ownerId}\0${name}`;
const existing = methodByOwner.get(key);
if (existing) {
existing.push(def);
} else {
methodByOwner.set(key, [def]);
}
}
// C3. Class-like types go to classByName index (in addition to globalIndex).
if (CLASS_TYPES.has(type)) {
const existing = classByName.get(name);
if (existing) {
existing.push(def);
} else {
classByName.set(name, [def]);
}
const qualifiedKey = qualifiedName ?? name;
const qualifiedMatches = classByQualifiedName.get(qualifiedKey);
if (qualifiedMatches) {
qualifiedMatches.push(def);
} else {
classByQualifiedName.set(qualifiedKey, [def]);
}
}
// D. Invalidate the lazy callable index only when adding callable types
if (CALLABLE_TYPES.has(type)) {
callableIndex = null;
@@ -191,10 +312,12 @@ export const createSymbolTable = (): SymbolTable => {
};
const lookupFuzzy = (name: string): SymbolDefinition[] => {
fuzzyCallCount++;
return globalIndex.get(name) || [];
};
const lookupFuzzyCallable = (name: string): SymbolDefinition[] => {
fuzzyCallableCallCount++;
if (!callableIndex) {
// Build the callable index lazily on first use
callableIndex = new Map();
@@ -213,9 +336,60 @@ export const createSymbolTable = (): SymbolTable => {
return fieldByOwner.get(`${ownerNodeId}\0${fieldName}`);
};
const lookupMethodByOwner = (
ownerNodeId: string,
methodName: string,
argCount?: number,
): SymbolDefinition | undefined => {
const defs = methodByOwner.get(`${ownerNodeId}\0${methodName}`);
if (!defs || defs.length === 0) return undefined;
// Arity narrowing: when an argCount is provided and there are multiple
// overloads, keep only those whose parameterCount can accommodate the
// call. This resolves arity-differing overloads (e.g. C++ `greet()` vs
// `greet(string)`) that share the same `ownerId + methodName` key.
//
// Candidates with `parameterCount === undefined` (extractor didn't
// populate the count — typically variadic or unknown) are retained
// conservatively so that legitimate variadic matches still resolve.
let pool = defs;
if (argCount !== undefined && defs.length > 1) {
const arityMatched = defs.filter((d) => {
if (d.parameterCount === undefined) return true;
const min = d.requiredParameterCount ?? d.parameterCount;
return argCount >= min && argCount <= d.parameterCount;
});
// Only adopt the arity-narrowed pool when it found matches; if arity
// rules out every candidate, fall back to the unfiltered set so the
// caller's fuzzy path still has something to work with.
if (arityMatched.length > 0) pool = arityMatched;
}
if (pool.length === 1) return pool[0];
// Multiple overloads after arity narrowing: return first if all share
// the same defined returnType (safe for chain resolution), undefined if
// return types differ (truly ambiguous — can't determine which overload).
const firstReturnType = pool[0].returnType;
if (firstReturnType === undefined) return undefined;
for (let i = 1; i < pool.length; i++) {
if (pool[i].returnType !== firstReturnType) return undefined;
}
return pool[0];
};
const lookupClassByName = (name: string): SymbolDefinition[] => {
return classByName.get(name) ?? [];
};
const lookupClassByQualifiedName = (qualifiedName: string): SymbolDefinition[] => {
return classByQualifiedName.get(qualifiedName) ?? [];
};
const getStats = () => ({
fileCount: fileIndex.size,
globalSymbolCount: globalIndex.size,
fuzzyCallableCallCount: fuzzyCallableCallCount,
fuzzyCallCount: fuzzyCallCount,
});
const clear = () => {
@@ -223,6 +397,11 @@ export const createSymbolTable = (): SymbolTable => {
globalIndex.clear();
callableIndex = null;
fieldByOwner.clear();
methodByOwner.clear();
classByName.clear();
classByQualifiedName.clear();
fuzzyCallCount = 0;
fuzzyCallableCallCount = 0;
};
return {
@@ -233,6 +412,9 @@ export const createSymbolTable = (): SymbolTable => {
lookupFuzzy,
lookupFuzzyCallable,
lookupFieldByOwner,
lookupMethodByOwner,
lookupClassByName,
lookupClassByQualifiedName,
getStats,
clear,
};
@@ -11,6 +11,9 @@ export const TYPESCRIPT_QUERIES = `
(class_declaration
name: (type_identifier) @name) @definition.class
(abstract_class_declaration
name: (type_identifier) @name) @definition.class
(interface_declaration
name: (type_identifier) @name) @definition.interface
@@ -24,6 +27,18 @@ export const TYPESCRIPT_QUERIES = `
(method_definition
name: (property_identifier) @name) @definition.method
; ES2022 #private methods (private_property_identifier not matched by property_identifier)
(method_definition
name: (private_property_identifier) @name) @definition.method
; Abstract method signatures in abstract classes
(abstract_method_signature
name: (property_identifier) @name) @definition.method
; Interface method signatures
(method_signature
name: (property_identifier) @name) @definition.method
(lexical_declaration
(variable_declarator
name: (identifier) @name
@@ -145,6 +160,10 @@ export const JAVASCRIPT_QUERIES = `
(method_definition
name: (property_identifier) @name) @definition.method
; ES2022 #private methods
(method_definition
name: (private_property_identifier) @name) @definition.method
(lexical_declaration
(variable_declarator
name: (identifier) @name
@@ -330,6 +349,7 @@ export const JAVA_QUERIES = `
; Calls
(method_invocation name: (identifier) @call.name) @call
(method_invocation object: (_) name: (identifier) @call.name) @call
(method_reference) @call
; Constructor calls: new Foo()
(object_creation_expression type: (type_identifier) @call.name) @call
@@ -615,6 +635,7 @@ export const CSHARP_QUERIES = `
export const RUST_QUERIES = `
; Functions & Items
(function_item name: (identifier) @name) @definition.function
(function_signature_item name: (identifier) @name) @definition.function
(struct_item name: (type_identifier) @name) @definition.struct
(enum_item name: (type_identifier) @name) @definition.enum
(trait_item name: (type_identifier) @name) @definition.trait
@@ -1105,6 +1126,13 @@ export const DART_QUERIES = `
value: (identifier) @call.name
(selector (argument_part))) @call
; ── Calls: member calls in variable assignments (var x = obj.method()) ──────
(initialized_variable_definition
(selector
(unconditional_assignable_selector
(identifier) @call.name))
(selector (argument_part))) @call
; ── Re-exports (export 'foo.dart') ───────────────────────────────────────────
(import_or_export
(library_export
@@ -1163,5 +1191,6 @@ export const LANGUAGE_QUERIES: Record<SupportedLanguages, string> = {
[SupportedLanguages.Ruby]: RUBY_QUERIES,
[SupportedLanguages.Swift]: SWIFT_QUERIES,
[SupportedLanguages.Dart]: DART_QUERIES,
[SupportedLanguages.Vue]: TYPESCRIPT_QUERIES, // Vue <script> blocks are parsed as TypeScript
[SupportedLanguages.Cobol]: '', // Standalone regex processor — no tree-sitter queries
};
+137 -73
View File
@@ -1,13 +1,14 @@
import {
type SyntaxNode,
FUNCTION_NODE_TYPES,
extractFunctionName,
CLASS_CONTAINER_TYPES,
genericFuncName,
} from './utils/ast-helpers.js';
import { CALL_EXPRESSION_TYPES } from './utils/call-analysis.js';
import { SupportedLanguages } from 'gitnexus-shared';
import { TYPED_PARAMETER_TYPES } from './type-extractors/shared.js';
import { getProvider } from './languages/index.js';
import type { BindingAccumulator, BindingEntry } from './binding-accumulator.js';
import type {
ClassNameLookup,
ReturnTypeLookup,
@@ -42,7 +43,21 @@ type TypeEnv = Map<string, Map<string, string>>;
const FILE_SCOPE = '';
/** Shared empty map for files with no file-scope bindings. */
const EMPTY_FILE_SCOPE: ReadonlyMap<string, string> = new Map();
/**
* Create a fresh empty Map for the "no file-scope bindings" fallback.
*
* **Why not a shared sentinel**: we previously used a module-level
* `const EMPTY_FILE_SCOPE = new Map()` typed as `ReadonlyMap` and shared
* across every TypeEnv instance. That was a latent singleton-poisoning
* footgun: any caller that did `(fileScope() as Map).set(...)` — or any
* future refactor that widened the return type — would silently corrupt
* every subsequent "empty" return for the process lifetime. A Proxy
* wrapper was considered but broke Map's internal-slot methods (`.size`,
* iteration protocol). Allocating a fresh empty Map per call is a few
* bytes per file — immediately GC'd, no measurable cost even at 10k files
* — and eliminates the shared-mutation hazard entirely.
*/
const emptyFileScope = (): ReadonlyMap<string, string> => new Map();
/** Fallback for languages where class names aren't in a 'name' field (e.g. Kotlin uses type_identifier). */
const findTypeIdentifierChild = (node: SyntaxNode): SyntaxNode | null => {
@@ -73,6 +88,10 @@ export interface TypeEnvironment {
* Populated when a variable has BOTH a declared base type AND a more specific
* constructor type (e.g., `Animal a = new Dog()` → key maps to 'Dog'). */
readonly constructorTypeMap: ReadonlyMap<string, string>;
/** Copy all scoped bindings into a BindingAccumulator.
* Must be called at most once per TypeEnv instance — throws on second call.
* The source `env` is not cleared (TypeEnv is per-file and discarded immediately after). */
flush(filePath: string, accumulator: BindingAccumulator): void;
}
/**
@@ -132,6 +151,7 @@ const lookupInEnv = (
callNode: SyntaxNode,
patternOverrides?: PatternOverrides,
enclosingFunctionFinder?: (n: SyntaxNode) => { funcName: string; label: NodeLabel } | null,
extractFunctionNameHook?: (n: SyntaxNode) => { funcName: string | null; label: NodeLabel } | null,
): string | undefined => {
// Self/this receiver: resolve to enclosing class name via AST walk
if (varName === 'self' || varName === 'this' || varName === '$this') {
@@ -145,7 +165,11 @@ const lookupInEnv = (
}
// Determine the enclosing function scope for the call
const scopeKey = findEnclosingScopeKey(callNode, enclosingFunctionFinder);
const scopeKey = findEnclosingScopeKey(
callNode,
enclosingFunctionFinder,
extractFunctionNameHook,
);
// Check position-indexed pattern overrides first (e.g., Kotlin when/is smart casts).
// These take priority over flat scopeEnv because they represent per-branch narrowing.
@@ -361,11 +385,12 @@ const extractParentClassFromNode = (classNode: SyntaxNode): string | undefined =
const findEnclosingScopeKey = (
node: SyntaxNode,
enclosingFunctionFinder?: (n: SyntaxNode) => { funcName: string; label: NodeLabel } | null,
extractFunctionNameHook?: (n: SyntaxNode) => { funcName: string | null; label: NodeLabel } | null,
): string | undefined => {
let current = node.parent;
while (current) {
if (FUNCTION_NODE_TYPES.has(current.type)) {
const { funcName } = extractFunctionName(current);
const funcName = extractFunctionNameHook?.(current)?.funcName ?? genericFuncName(current);
if (funcName) return `${funcName}@${current.startIndex}`;
}
// Language-specific hook (e.g., Dart function_body → sibling function_signature)
@@ -389,7 +414,7 @@ const findEnclosingScopeKey = (
* using cross-file type information when available.
*
* Only `.has()` is exposed — the SymbolTable doesn't support iteration.
* Results are memoized to avoid redundant lookupFuzzy scans across declarations.
* Results are memoized to avoid redundant class-index scans across declarations.
*/
const createClassNameLookup = (
localNames: Set<string>,
@@ -404,7 +429,7 @@ const createClassNameLookup = (
const cached = memo.get(name);
if (cached !== undefined) return cached;
const result = symbolTable
.lookupFuzzy(name)
.lookupClassByName(name)
.some((def) => def.type === 'Class' || def.type === 'Enum' || def.type === 'Struct');
memo.set(name, result);
return result;
@@ -453,18 +478,23 @@ const SKIP_SUBTREE_TYPES = new Set([
]);
const CLASS_LIKE_TYPES = new Set(['Class', 'Struct', 'Interface']);
type ClassDefRef = { nodeId: string; type: string; filePath: string };
const lookupClassDefsByName = (
symbolTable: SymbolTable,
name: string,
allowedTypes: ReadonlySet<string> = CLASS_LIKE_TYPES,
): ClassDefRef[] => symbolTable.lookupClassByName(name).filter((d) => allowedTypes.has(d.type));
/** Memoize class definition lookups during fixpoint iteration.
* SymbolTable is immutable during type resolution, so results never change.
* Eliminates redundant array allocations + filter scans across iterations. */
const createClassDefCache = (symbolTable?: SymbolTable) => {
const cache = new Map<string, Array<{ nodeId: string; type: string }>>();
const cache = new Map<string, ClassDefRef[]>();
return (typeName: string) => {
let result = cache.get(typeName);
if (result === undefined) {
result = symbolTable
? symbolTable.lookupFuzzy(typeName).filter((d) => CLASS_LIKE_TYPES.has(d.type))
: [];
result = symbolTable ? lookupClassDefsByName(symbolTable, typeName) : [];
cache.set(typeName, result);
}
return result;
@@ -550,7 +580,7 @@ export const isSubclassOf = (
const walkParentChain = <T>(
typeName: string,
parentMap: ReadonlyMap<string, readonly string[]> | undefined,
getClassDefs: (name: string) => Array<{ nodeId: string; type: string }>,
getClassDefs: (name: string) => ClassDefRef[],
lookupOnClass: (nodeId: string) => T | undefined,
): T | undefined => {
if (!parentMap) return undefined;
@@ -586,15 +616,13 @@ const resolveFieldType = (
field: string,
scopeEnv: ReadonlyMap<string, string>,
symbolTable?: SymbolTable,
getClassDefs?: (typeName: string) => Array<{ nodeId: string; type: string }>,
getClassDefs?: (typeName: string) => ClassDefRef[],
parentMap?: ReadonlyMap<string, readonly string[]>,
): string | undefined => {
if (!symbolTable) return undefined;
const receiverType = scopeEnv.get(receiver);
if (!receiverType) return undefined;
const lookup =
getClassDefs ??
((name: string) => symbolTable.lookupFuzzy(name).filter((d) => CLASS_LIKE_TYPES.has(d.type)));
const lookup = getClassDefs ?? ((name: string) => lookupClassDefsByName(symbolTable, name));
const classDefs = lookup(receiverType);
if (classDefs.length !== 1) return undefined;
// Direct lookup first
@@ -610,40 +638,52 @@ const resolveFieldType = (
/** Resolve a method's return type given a receiver variable and method name.
* Uses SymbolTable to find class nodeIds for the receiver's type, then
* looks up the method via lookupFuzzyCallable filtered by ownerId.
* looks up the method via owner-scoped lookupMethodByOwner.
* Falls back to MRO parent chain walking if direct lookup fails (Phase 11A). */
const resolveMethodReturnType = (
receiver: string,
method: string,
scopeEnv: ReadonlyMap<string, string>,
symbolTable?: SymbolTable,
getClassDefs?: (typeName: string) => Array<{ nodeId: string; type: string }>,
getClassDefs?: (typeName: string) => ClassDefRef[],
parentMap?: ReadonlyMap<string, readonly string[]>,
): string | undefined => {
if (!symbolTable) return undefined;
const receiverType = scopeEnv.get(receiver);
let receiverType = scopeEnv.get(receiver);
// When substituteThisReceiver replaced $this/self with the enclosing class name,
// the receiver IS the type — look it up directly as a class name.
if (!receiverType) {
const lookup = getClassDefs ?? ((name: string) => lookupClassDefsByName(symbolTable, name));
if (lookup(receiver).length > 0) receiverType = receiver;
}
if (!receiverType) return undefined;
const lookup =
getClassDefs ??
((name: string) => symbolTable.lookupFuzzy(name).filter((d) => CLASS_LIKE_TYPES.has(d.type)));
const lookup = getClassDefs ?? ((name: string) => lookupClassDefsByName(symbolTable, name));
const classDefs = lookup(receiverType);
if (classDefs.length === 0) return undefined;
// Direct lookup first
const classNodeIds = new Set(classDefs.map((d) => d.nodeId));
const methods = symbolTable
.lookupFuzzyCallable(method)
.filter((d) => d.ownerId && classNodeIds.has(d.ownerId));
const directMethodLookups = classDefs.map((d) => ({
classDef: d,
methodDef: symbolTable.lookupMethodByOwner(d.nodeId, method),
}));
const hasAmbiguousDirectLookup = directMethodLookups.some(({ classDef, methodDef }) => {
if (methodDef) return false;
return symbolTable
.lookupExactAll(classDef.filePath, method)
.some((d) => d.ownerId === classDef.nodeId);
});
if (hasAmbiguousDirectLookup) return undefined;
const methods = directMethodLookups
.map(({ methodDef }) => methodDef)
.filter((d): d is NonNullable<typeof d> => d !== undefined);
if (methods.length === 1 && methods[0].returnType) {
return extractReturnTypeName(methods[0].returnType);
}
// MRO parent chain walking on miss
if (methods.length === 0) {
const inherited = walkParentChain(receiverType, parentMap, lookup, (nodeId) => {
const parentMethods = symbolTable
.lookupFuzzyCallable(method)
.filter((d) => d.ownerId === nodeId);
if (parentMethods.length !== 1 || !parentMethods[0].returnType) return undefined;
return extractReturnTypeName(parentMethods[0].returnType);
const parentMethod = symbolTable.lookupMethodByOwner(nodeId, method);
if (!parentMethod?.returnType) return undefined;
return extractReturnTypeName(parentMethod.returnType);
});
return inherited;
}
@@ -765,6 +805,11 @@ export interface BuildTypeEnvOptions {
enclosingFunctionFinder?: (
ancestorNode: SyntaxNode,
) => { funcName: string; label: NodeLabel } | null;
/** Language-specific function name extraction from an AST node.
* Replaces the generic name-field lookup for languages with non-standard
* AST structures (C/C++ declarator unwrapping, Swift init/deinit, etc.).
* When null is returned or not provided, falls back to node.childForFieldName('name')?.text. */
extractFunctionName?: (node: SyntaxNode) => { funcName: string | null; label: NodeLabel } | null;
}
/** Seed cross-file type bindings into the file scope.
@@ -794,7 +839,9 @@ export const buildTypeEnv = (
const symbolTable = options?.symbolTable;
const parentMap = options?.parentMap;
const extractFuncNameHook = options?.extractFunctionName;
const env: TypeEnv = new Map();
let flushed = false;
const patternOverrides: PatternOverrides = new Map();
// Phase P: maps `scope\0varName` → constructor type when a declaration has BOTH
// a base type annotation AND a more specific constructor initializer.
@@ -955,47 +1002,19 @@ export const buildTypeEnv = (
// This decouples type node capture from scopeEnv success — container types
// (User[], []User, List[User]) that fail extractSimpleTypeName still get
// their AST type node recorded for Strategy 1 for-loop resolution.
// Try direct extraction first (works for Go var_spec, Python assignment, Rust let_declaration).
// Try direct type field first, then unwrap wrapper nodes (C# field_declaration,
// local_declaration_statement wrap their type inside a variable_declaration child).
let typeNode = node.childForFieldName('type');
//
// Prefer language-specific locator when provided (keeps buildTypeEnv generic),
// then fall back to a small set of safe, cross-grammar heuristics.
let typeNode =
config.getDeclarationTypeNode?.(node) ?? node.childForFieldName('type') ?? null;
// Fallback: some grammars wrap type annotations in a `type_annotation` child
// instead of exposing a named `type` field on the declaration node.
if (!typeNode) {
// C# field_declaration / local_declaration_statement wrap type inside variable_declaration.
// Use manual loop instead of namedChildren.find() to avoid array allocation on hot path.
let wrapped = node.childForFieldName('declaration');
if (!wrapped) {
for (let i = 0; i < node.namedChildCount; i++) {
const c = node.namedChild(i);
if (c?.type === 'variable_declaration') {
wrapped = c;
break;
}
}
}
if (wrapped) {
typeNode = wrapped.childForFieldName('type');
// Kotlin: variable_declaration stores the type as user_type / nullable_type
// child rather than a named 'type' field.
if (!typeNode) {
for (let i = 0; i < wrapped.namedChildCount; i++) {
const c = wrapped.namedChild(i);
if (c && (c.type === 'user_type' || c.type === 'nullable_type')) {
typeNode = c;
break;
}
}
}
}
// Swift: property_declaration has type_annotation as a direct child (not a 'type' field).
// Extract the inner type node (array_type, user_type, etc.) for declarationTypeNodes.
if (!typeNode) {
for (let i = 0; i < node.namedChildCount; i++) {
const c = node.namedChild(i);
if (c?.type === 'type_annotation') {
// Use the inner type (array_type, user_type) rather than the annotation wrapper
typeNode = c.firstNamedChild ?? c;
break;
}
for (let i = 0; i < node.namedChildCount; i++) {
const c = node.namedChild(i);
if (c?.type === 'type_annotation') {
typeNode = c.firstNamedChild ?? c;
break;
}
}
}
@@ -1085,7 +1104,7 @@ export const buildTypeEnv = (
// Detect scope boundaries (function/method definitions)
let scope = currentScope;
if (FUNCTION_NODE_TYPES.has(node.type)) {
const { funcName } = extractFunctionName(node);
const funcName = extractFuncNameHook?.(node)?.funcName ?? genericFuncName(node);
if (funcName) scope = `${funcName}@${node.startIndex}`;
}
@@ -1234,11 +1253,56 @@ export const buildTypeEnv = (
return {
lookup: (varName, callNode) =>
lookupInEnv(env, varName, callNode, patternOverrides, options?.enclosingFunctionFinder),
lookupInEnv(
env,
varName,
callNode,
patternOverrides,
options?.enclosingFunctionFinder,
extractFuncNameHook,
),
constructorBindings: bindings,
fileScope: () => env.get(FILE_SCOPE) ?? EMPTY_FILE_SCOPE,
fileScope: () => env.get(FILE_SCOPE) ?? emptyFileScope(),
allScopes: () => env as ReadonlyMap<string, ReadonlyMap<string, string>>,
constructorTypeMap,
flush(filePath: string, accumulator: BindingAccumulator): void {
if (flushed) {
throw new Error(
`[TypeEnvironment] flush called twice for ${filePath} — flush is single-use`,
);
}
// Narrow flush() to iterate only the FILE_SCOPE entry, mirroring the
// worker-path narrowing in parse-worker.ts (commit 803631fe). Before
// this change, both execution paths had the same asymmetry bug: the
// worker path was fixed but the sequential path (this code) still
// wrote function-scope entries into long-lived accumulator storage
// that no consumer reads until Phase 9 lands.
//
// Phase 9 reversion: when a downstream consumer of function-scope
// bindings exists, restore the nested iteration:
//
// for (const [scope, scopeMap] of env) {
// for (const [varName, typeName] of scopeMap) {
// entries.push({ scope, varName, typeName });
// }
// }
//
// See BindingAccumulator class JSDoc and FileScopeBindings JSDoc in
// parse-worker.ts for the full reversion checklist.
const fileScope = env.get(FILE_SCOPE) ?? emptyFileScope();
const entries: BindingEntry[] = [];
for (const [varName, typeName] of fileScope) {
entries.push({ scope: '', varName, typeName });
}
if (entries.length > 0) {
accumulator.appendFile(filePath, entries);
}
// Mark the env as flushed AFTER the successful append. If appendFile
// throws (e.g., accumulator is already finalized due to a lifecycle
// ordering bug), the caller can catch and retry — the single-use
// guard now tracks "data was written", not "flush was attempted".
flushed = true;
},
};
};
@@ -559,6 +559,24 @@ const inferLiteralType: LiteralTypeInferrer = (node) => {
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
getDeclarationTypeNode: (node) => {
// C# field_declaration / local_declaration_statement wrap type inside variable_declaration.
// Prefer the wrapper node's `type` field when present.
const direct = node.childForFieldName('type');
if (direct) return direct;
const wrapped =
node.childForFieldName('declaration') ??
(() => {
for (let i = 0; i < node.namedChildCount; i++) {
const c = node.namedChild(i);
if (c?.type === 'variable_declaration') return c;
}
return null;
})();
return wrapped?.childForFieldName('type') ?? null;
},
forLoopNodeTypes: FOR_LOOP_NODE_TYPES,
patternBindingNodeTypes: new Set([
'is_pattern_expression',
@@ -850,6 +850,20 @@ const extractKotlinPatternBinding: PatternBindingExtractor = (
export const kotlinTypeConfig: LanguageTypeConfig = {
allowPatternBindingOverwrite: true,
declarationNodeTypes: KOTLIN_DECLARATION_NODE_TYPES,
getDeclarationTypeNode: (node) => {
// Kotlin property_declaration wraps the actual declaration in variable_declaration.
// The type is commonly a user_type / nullable_type child (positional, not 'type' field).
const varDecl =
node.type === 'property_declaration' ? findChild(node, 'variable_declaration') : node;
if (varDecl) {
return (
varDecl.childForFieldName('type') ??
findChild(varDecl, 'user_type') ??
findChild(varDecl, 'nullable_type')
);
}
return node.childForFieldName('type') ?? findChild(node, 'user_type') ?? null;
},
forLoopNodeTypes: KOTLIN_FOR_LOOP_NODE_TYPES,
patternBindingNodeTypes: new Set(['type_test', 'equality_expression']),
extractDeclaration: extractKotlinDeclaration,
@@ -6,7 +6,6 @@ import type {
InitializerExtractor,
ClassNameLookup,
ConstructorBindingScanner,
ReturnTypeExtractor,
PendingAssignmentExtractor,
ForLoopExtractor,
} from './types.js';
@@ -337,60 +336,6 @@ const scanConstructorBinding: ConstructorBindingScanner = (node) => {
return undefined;
};
/** Regex to extract PHPDoc @return annotations: `@return User` */
const PHPDOC_RETURN_RE = /@return\s+(\S+)/;
/**
* Normalize a PHPDoc return type for storage in the SymbolTable.
* Unlike normalizePhpType (which strips User[] → User for scopeEnv), this preserves
* array notation so lookupRawReturnType can extract element types for for-loop resolution.
* \App\Models\User[] → User[]
* ?User → User
* Collection<User> → Collection<User> (preserved for extractElementTypeFromString)
*/
const normalizePhpReturnType = (raw: string): string | undefined => {
// Strip nullable prefix: ?User[] → User[]
let type = raw.startsWith('?') ? raw.slice(1) : raw;
// Strip union with null/false/void: User[]|null → User[]
const parts = type
.split('|')
.filter((p) => p !== 'null' && p !== 'false' && p !== 'void' && p !== 'mixed');
if (parts.length !== 1) return undefined;
type = parts[0];
// Strip namespace: \App\Models\User[] → User[]
const segments = type.split('\\');
type = segments[segments.length - 1];
// Skip uninformative types
if (
type === 'mixed' ||
type === 'void' ||
type === 'self' ||
type === 'static' ||
type === 'object' ||
type === 'array'
)
return undefined;
if (/^\w+(\[\])?$/.test(type) || /^\w+\s*</.test(type)) return type;
return undefined;
};
/**
* Extract return type from PHPDoc `@return Type` annotation preceding a method.
* Walks backwards through preceding siblings looking for comment nodes.
* Preserves array notation (e.g., User[]) for for-loop element type extraction.
*/
const extractReturnType: ReturnTypeExtractor = (node) => {
let sibling = node.previousSibling;
while (sibling) {
if (sibling.type === 'comment') {
const match = PHPDOC_RETURN_RE.exec(sibling.text);
if (match) return normalizePhpReturnType(match[1]);
} else if (sibling.isNamed && !SKIP_NODE_TYPES.has(sibling.type)) break;
sibling = sibling.previousSibling;
}
return undefined;
};
/** PHP: $alias = $user → assignment_expression with variable_name left/right.
* PHP TypeEnv stores variables WITH $ prefix ($user → User), so we keep $ in lhs/rhs. */
const extractPendingAssignment: PendingAssignmentExtractor = (node, scopeEnv) => {
@@ -605,7 +550,6 @@ export const typeConfig: LanguageTypeConfig = {
extractParameter,
extractInitializer,
scanConstructorBinding,
extractReturnType,
extractForLoopBinding,
extractPendingAssignment,
};
@@ -4,7 +4,6 @@ import type {
TypeBindingExtractor,
InitializerExtractor,
ConstructorBindingScanner,
ReturnTypeExtractor,
PendingAssignmentExtractor,
ForLoopExtractor,
} from './types.js';
@@ -43,9 +42,6 @@ const YARD_PARAM_RE = /@param\s+(\w+)\s+\[([^\]]+)\]/g;
/** Alternate YARD order: `@param [Type] name` */
const YARD_PARAM_ALT_RE = /@param\s+\[([^\]]+)\]\s+(\w+)/g;
/** Regex to extract @return annotations: `@return [Type]` */
const YARD_RETURN_RE = /@return\s+\[([^\]]+)\]/;
/**
* Extract the simple type name from a YARD type string.
* Handles:
@@ -229,35 +225,6 @@ const extractInitializer: InitializerExtractor = (node, env, classNames): void =
}
};
/**
* Extract return type from YARD `@return [Type]` annotation preceding a method.
* Reuses the same comment-walking strategy as collectYardParams: try direct
* siblings first, fall back to parent (body_statement) siblings for class methods.
*/
const extractReturnType: ReturnTypeExtractor = (node) => {
const search = (startNode: SyntaxNode): string | undefined => {
let sibling = startNode.previousSibling;
while (sibling) {
if (sibling.type === 'comment') {
const match = YARD_RETURN_RE.exec(sibling.text);
if (match) return extractYardTypeName(match[1]);
} else if (sibling.isNamed) {
break;
}
sibling = sibling.previousSibling;
}
return undefined;
};
const result = search(node);
if (result) return result;
if (node.parent?.type === 'body_statement') {
return search(node.parent);
}
return undefined;
};
/**
* Ruby constructor binding scanner: captures both `user = User.new` and
* plain call assignments like `user = get_user()`.
@@ -452,7 +419,6 @@ export const typeConfig: LanguageTypeConfig = {
extractParameter,
extractInitializer,
scanConstructorBinding,
extractReturnType,
extractForLoopBinding,
extractPendingAssignment,
};
@@ -496,6 +496,17 @@ function extractSwiftElementTypeFromTypeNode(typeNode: SyntaxNode): string | und
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
getDeclarationTypeNode: (node) => {
// Swift: many declarations store type as a type_annotation child (not a 'type' field).
// Prefer a direct 'type' field if present, else unwrap type_annotation to its inner type.
const direct = node.childForFieldName('type');
if (direct) return direct;
for (let i = 0; i < node.namedChildCount; i++) {
const c = node.namedChild(i);
if (c?.type === 'type_annotation') return c.firstNamedChild ?? c;
}
return null;
},
forLoopNodeTypes: FOR_LOOP_NODE_TYPES,
extractDeclaration,
extractParameter,
@@ -6,6 +6,11 @@ export type TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>)
/** Extracts type bindings from a parameter node into the env map */
export type ParameterExtractor = (node: SyntaxNode, env: Map<string, string>) => void;
/** Optionally locates the type-annotation AST node for a declaration node.
* Used by buildTypeEnv to populate declarationTypeNodes and constructorTypeMap.
* If absent, buildTypeEnv falls back to generic heuristics (childForFieldName('type'), etc). */
export type DeclarationTypeNodeLocator = (node: SyntaxNode) => SyntaxNode | null;
/** Minimal interface for checking whether a name is a known class/struct.
* Narrower than ReadonlySet — only `.has()` is used by extractors. */
export type ClassNameLookup = { has(name: string): boolean };
@@ -25,11 +30,6 @@ export type ConstructorBindingScanner = (
node: SyntaxNode,
) => { varName: string; calleeName: string; receiverClassName?: string } | undefined;
/** Extracts a return type string from a method/function definition node.
* Used for languages where return types are expressed in comments (e.g. YARD @return [Type])
* rather than in AST fields. Returns undefined if no return type can be determined. */
export type ReturnTypeExtractor = (node: SyntaxNode) => string | undefined;
/** Infer the type name of a literal AST node for overload disambiguation.
* Returns the canonical type name (e.g. 'int', 'String', 'boolean') or undefined
* for non-literal nodes. Only used when resolveCallTarget has multiple candidates
@@ -142,6 +142,9 @@ export interface LanguageTypeConfig {
readonly allowPatternBindingOverwrite?: boolean;
/** Node types that represent typed declarations for this language */
declarationNodeTypes: ReadonlySet<string>;
/** Optional: language-specific way to find a declaration's type-annotation node.
* Prefer providing this for grammars where the type is wrapped (e.g., C#, Kotlin, Swift). */
getDeclarationTypeNode?: DeclarationTypeNodeLocator;
/** AST node types for for-each/for-in statements with explicit element types. */
forLoopNodeTypes?: ReadonlySet<string>;
/** Optional allowlist of AST node types on which extractPatternBinding should run.
@@ -162,9 +165,6 @@ export interface LanguageTypeConfig {
* Called on every AST node during buildTypeEnv walk; returns undefined for non-matches.
* The callee binding is unverified — the caller must confirm against the SymbolTable. */
scanConstructorBinding?: ConstructorBindingScanner;
/** Extract return type from comment-based annotations (e.g. YARD @return [Type]).
* Called as fallback when extractMethodSignature finds no AST-based return type. */
extractReturnType?: ReturnTypeExtractor;
/** Extract loop variable → type binding from a for-each AST node. */
extractForLoopBinding?: ForLoopExtractor;
/** Extract pending assignment for Tier 2 propagation.

Some files were not shown because too many files have changed in this diff Show More