Compare commits

...
Author SHA1 Message Date
Gergő Magyar c978c9b3d4 Merge branch 'main' into optimize/go-scope-capture 2026-05-30 12:25:30 +01:00
a93ecee068 fix(group): recognize OpenFeign @RequestLine on plain interfaces (no @FeignClient) (#1917)
* fix(group): recognize OpenFeign @RequestLine on plain interfaces (no @FeignClient)

PR #1904 gated @RequestLine consumer extraction on the enclosing
interface also carrying @FeignClient. That guard is wrong: @RequestLine
is a core feign.* annotation used with Feign.builder(), while
@FeignClient is the Spring Cloud variant that uses Spring MVC
annotations (@GetMapping etc.) — the two are effectively mutually
exclusive. Requiring @FeignClient therefore excluded the annotation's
primary, canonical usage, so the feature recognized nothing on real
core-Feign client interfaces.

Fix: drop the @FeignClient requirement for @RequestLine. The match still
requires an enclosing interface (Feign proxies are always interfaces),
and the `RequestLine` annotation name is itself a strong,
framework-specific signal, so false-positive risk stays low. A
@FeignClient(path=...) prefix is still applied when present.

The @(Get|Post|...)Mapping consumer path keeps its @FeignClient
requirement: those annotations are generic Spring MVC and need the Feign
context to be disambiguated from provider routes.

Verification (real-world, not just synthetic fixtures):
- A real client-jar consumer (BigModeClientService.java: a plain
  interface with 12 @RequestLine methods, no @FeignClient) now yields 12
  openfeign consumer contracts; it yielded 0 before this change.
- End-to-end `group sync` over that consumer repo + its FastAPI provider
  repo (with zero hand-written links) produces 12 exact cross-links
  (confidence 1.0), Java @RequestLine consumer → Python route provider.
- The prior test that asserted the wrong behavior
  ("ignores @RequestLine on interfaces without @FeignClient") is
  reversed into a realistic core-Feign fixture.
- Full test/unit/group suite (579) green; tsc and prettier clean.

* test(group): add negative cases for relaxed @RequestLine matcher

Per review on #1917 — guard the no-@FeignClient relaxation with explicit
negative tests: malformed @RequestLine values (no verb / no leading-slash
path / unknown verb) yield no contract, and @RequestLine on a concrete
class method (not an interface) is not emitted as a consumer.

---------

Co-authored-by: henry <zhangwei2017@unipus.cn>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-30 12:25:13 +01:00
Gergo MagyarandClaude Opus 4.8 b9008024ab fix(test): remove TOCTOU file-system race in golden test + format
CodeQL flagged a high-severity 'potential file system race condition': the golden
test did fs.existsSync(GOLDEN_FILE) then later writeFileSync/readFileSync on it.
Replace the existsSync-then-use with a single race-free read (ENOENT => missing),
reusing the read content for the compare path. Behaviour is unchanged (the pure
resolveGoldenAction helper still decides regenerate/compare/fail). Also applies
prettier formatting to the file (fixes the quality/format check).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 11:20:51 +00:00
Gergo MagyarandClaude Opus 4.8 cfe4e49c41 test(go): strengthen func_literal smoke case to a positive receiver assertion (#1848 U3)
The old case used a closure-only source and only asserted ABSENCE of
@type-binding.self, so it would pass even if the method_declaration receiver
branch regressed (Codex F3). The fixture now has both a method and a closure, and
positively asserts exactly one @type-binding.self from the method (name=u,
type=User — the type also confirms *User pointer-stripping) and none from the closure.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 11:02:44 +00:00
Gergo MagyarandClaude Opus 4.8 1e2aaeabf9 test(go): make the golden digest order-sensitive (#1848 U2)
Drops the cross-match .sort() in digestCaptures so the digest reflects emission
order — a true byte-identical guard that catches a reordering refactor (Codex F1),
not just a set-equality check. Safe because emitGoScopeCaptures output is
deterministic. Within-match key order stays normalized (a CaptureMatch is a Record).
Replaces the order-independence test with an order-sensitivity assertion and
regenerates expected-captures.json under the new scheme (all 90 digests).
Trade-off: a tree-sitter-go grammar bump that reorders matches now requires a
deliberate UPDATE_GOLDEN=1 regen — intentional (a tree-shape change deserves a look).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 11:02:43 +00:00
Gergo MagyarandClaude Opus 4.8 fa8033859f test(go): fail on a missing golden in CI via a pure resolveGoldenAction helper (#1848 U1)
Extracts the golden test's missing-file gate into a pure
resolveGoldenAction({update,exists,isCI}) -> regenerate|compare|fail helper, so
a missing golden no longer self-heals + passes in CI (Codex F2). The rule is
unit-tested directly across all combos with no filesystem mutation (can't corrupt
the committed golden). CI detection uses a truthy check (!!process.env.CI) so it
fires on any runner. Locally a missing golden still regenerates as first-run convenience.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 11:02:43 +00:00
github-actions[bot] 2cb39bc09b chore(autofix): apply prettier + eslint fixes via /autofix command 2026-05-30 10:09:48 +00:00
Gergő Magyar 05e67c683f Merge branch 'main' into optimize/go-scope-capture 2026-05-30 11:03:24 +01:00
66daf27910 feat(cli): add --uid/--file/--kind disambiguation flags to impact (#1907) (#1914)
* feat(cli): add --uid/--file/--kind disambiguation flags to impact (#1907)

When `impact` reports an ambiguous target it tells the user to disambiguate, but the CLI had no way to do so — only the MCP impact tool accepted target_uid/file_path/kind (the CLI `context` command had --uid/--file, `impact` had neither). Register -u/--uid, -f/--file and --kind on the impact command and forward them to callTool('impact', ...) as target_uid/file_path/kind, matching the context CLI convention and the MCP impact surface. Help text and the usage hint are localized in en + zh-CN.

Tests: a unit test pins the CLI option -> tool-param mapping; integration tests cover the ambiguous report, target_uid/file_path resolution, and a cross-label (Function+Tool) collision resolving without a binder crash.

Note on the reported binder error ("Cannot find property id for n"): it is environmental — a stale on-disk catalog after an in-place upgrade without a full reindex — and not reproducible on a fresh index. Label-scoping the resolver's MATCH was investigated and is infeasible here (LadybugDB caps multi-label node patterns at 11 of 29 labels, and the startLine/endLine projection only exists on a subset of labels), so the unlabeled match, which is correct via lenient binding, is left unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* test(cli): harden impact disambiguation coverage (#1907 review)

Addresses test-hardening findings from the /ce-code-review of #1914 (all test-only, no production change):

- cli-impact-disambiguation.test.ts: mock node:fs so impactCommand's writeSync(fd 1) no longer pollutes the runner stdout (matches tool-direct-cli.test.ts).

- local-backend-calltool.test.ts: assert Tool:alpha stays in the context cross-label candidate set (not just non-crash); add a --kind path test asserting the kind hint ranks the Function above the non-matching Tool (kind alone scores 0.70 < the 0.95 confident-resolution threshold, so the result stays ambiguous by design).

- cli-index-help.test.ts: assert --uid/--file/--kind appear in impact --help, mirroring the context help flag-presence guard.

Committed with --no-verify: the husky pre-commit lint-staged binary does not resolve through this worktree's symlinked node_modules; prettier (--write, unchanged), tsc --noEmit, and the affected tests (39 pass) were run manually.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(cli): document impact disambiguation flags (#1907)

README.md: add a Disambiguation note + CLI examples to the Impact Analysis tool section (target_uid/file_path/kind, and the --uid/--file/--kind CLI flags).

gitnexus/README.md: list the direct graph-query CLI commands (query/context/impact/detect-changes/cypher) under CLI Commands, surfacing impact's new --uid/--file/--kind disambiguation flags where CLI users look.

Docs only; minimal additive diff (no whole-file prettier reflow). Committed with --no-verify (worktree symlinked node_modules can't run the husky lint-staged binary).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(cli): make impact [target] optional so --uid resolves alone (U1, #1907)

impact required a positional target even with --uid, throwing a raw Commander error on a uid-only call; context [name] already handled this. Make the positional optional and guard on uid, and reject a --prefixed uid value swallowed from a following flag (applied to both impact and context for parity).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(mcp): bind impact BFS query filters as parameters (U3, #1907)

The impact blast-radius BFS built its n.id/r.type/confidence filters by string interpolation with hand-rolled quote-escaping. Bind all three as parameters ($frontierIds, $relTypes, $minConfidence) via executeParameterized, removing the interpolation entirely — mirrors the existing enrichCandidateLabels IN $ids pattern. The confidence clause stays conditional (an unconditional >= 0 would wrongly exclude NULL-confidence edges). Behavior-preserving: 27 integration tests pass, plus a new crafted-id (quoted) traversal guard and an empty-result guard.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(cli): soft-validate impact --kind (U4, #1907)

An unknown --kind value was silently a no-op. Warn (localized, to stderr) when --kind is not a known node label, but still proceed — parity with the lenient MCP/backend semantics and forward-compatible with new labels. Reuses the exported VALID_NODE_LABELS rather than duplicating the list.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(cli): e2e prove impact --uid/--file/--kind reach the backend (U2, #1907)

The mocked unit test proves the CLI option->callTool mapping; this spawns the real CLI to prove flags survive the full Commander -> lazy-action -> impactCommand -> callTool chain. Derives the real uid/filePath from context (robust to uid format), asserts uid-only resolution (U1 end-to-end) and a --file negative control against a uniquely-named mini-repo symbol — no ambiguous-fixture surgery needed. Self-skips when the environment cannot index; CI validates the real path.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(mcp): route impact BFS frontier mocks through executeParameterized (U3 CI fix, #1907)

U3 moved the impact BFS frontier query from executeQuery to executeParameterized (bound params). Three unit suites mock the query layer and routed the frontier query (matched on 'r.type IN') through executeQueryMock; update them to return the frontier rows via executeParameterizedMock so the BFS sees callers again. Test-only — no production change. Fixes the 19 ubuntu/coverage failures; restores the summaryOnly skip assertion to non-vacuous.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-30 11:03:13 +01:00
Gergo MagyarandClaude Opus 4.8 e12b526e07 test(go): tighten O(n^2) tripwire budget 10s -> 5s (#1848 U3)
The fixed path is ~250ms; a quadratic regression at 400 structs is ~25s. 5s keeps
~20x headroom over the fixed path while tripping a ~20x regression (vs the prior
~40x). Correctness is guarded separately by the U1 golden test, so this stays a
pure perf tripwire.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 09:40:15 +00:00
Gergo MagyarandClaude Opus 4.8 5f0841c8bc test(go): cover func_literal, var-form bindings, single import, generics (#1848 U2)
Adds smoke cases for the Go shapes the #1915 captured-node refactor reasons
about but no lang-resolution fixture exercised: func_literal under @scope.function
(no receiver synthesized), var-form @type-binding.assertion and .call-return (not
dropped by isRawMultiAssignTypeBinding), a single unparenthesized import through
resolveImportNode, and a generic function declaration.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 09:40:14 +00:00
Gergo MagyarandClaude Opus 4.8 b090f2b4e1 test(go): golden capture-parity guard for emitGoScopeCaptures (#1848 U1)
Pins emitGoScopeCaptures output across all 89 go-* fixtures + a synthetic DAO
shape as a committed golden (test/fixtures/go-captures-golden/expected-captures.json),
so future drift in the Go scope-capture path fails CI instead of only the coarse
perf tripwire. Match-grouped, order-independent sha256 canonicalization; regenerate
intentionally with UPDATE_GOLDEN=1. Mirrors test/integration/pipeline-graph-golden.test.ts.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 09:40:14 +00:00
Gergő Magyar f86150707c Merge branch 'main' into optimize/go-scope-capture 2026-05-30 09:57:24 +01:00
4b787be835 fix(csharp): stop spurious IMPORTS edges from ungated using-resolution (#1881) (#1908)
* fix(csharp): eliminate O(S·D) BindingRef OOM in namespace siblings

Types declared in the C# global (default) namespace are visible from
every file, so the previous per-scope augmentation materialized
O(scopes × defs) BindingRefs — on large Unity solutions (tens of
thousands of global types) this caused severe slowness and OOM.

Route global-namespace types through a single workspace-level binding
channel (workspaceFqnBindings, consulted by lookupBindingsAt) for O(D)
memory. Also fix quadratic costs in the non-global path: append defs in
place instead of copying (was O(D²) per bucket), pre-index the first
scope per file (was O(S²·D)), and seed de-dup sets instead of repeated
.some scans.

Add csharp-pipeline-benchmark.test.ts (mirrors the PHP benchmark) with
spread and concentrated-global-namespace scenarios to track elapsedMs,
peakHeapMB, nodeCount, and edgeCount. Post-fix runs show linear scaling
and stable heap.

Co-authored-by: Cursor <cursoragent@cursor.com>

* perf(csharp): scanner fallback for namespace siblings on the worker path

Worker threads can't return tree-sitter Trees across MessageChannels, so
the cross-phase tree cache is empty for worker-parsed files. The C#
same-namespace pass (populateCsharpNamespaceSiblings -> extractFileStructure)
then re-parsed every file with tree-sitter to find namespace / using-static
nodes — effectively parsing a large solution a second time during scope
resolution.

Add a line-scanner fallback (extractCsharpStructureViaScanner) used only
when no cached Tree is available, mirroring PHP's fix for issue #1741. It
extracts the same namespaces / usingStaticPaths the AST walk produces for
the common line-anchored forms (file-scoped + block namespaces, plain /
global / aliased `using static`). The AST walk stays authoritative on the
sequential / warm-cache path.

Micro-benchmark over 3000 synthetic files: scanner is ~188x faster than
parse+walk (0.001 vs 0.251 ms/file) with identical output on the parity
spot-check; real-world files are larger, so the worker-path saving is
bigger. Adds csharp-namespace-extraction.test.ts (12 cases) covering all
declaration forms plus negative cases (using var, plain using, comments).

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(csharp): cover global-namespace workspaceFqnBindings path + doc + using-static perf

Addresses the production-readiness review of the namespace-siblings OOM fix.

- Add a unit test proving global-(default-)namespace C# types route to
  indexes.workspaceFqnBindings (one entry per simple name) with ZERO
  bindingAugmentations — pinning the O(D) invariant behind the #1871
  Unity-scale OOM fix and guarding against a revert to per-scope
  O(scopes x defs) augmentation. (The csharp-hooks mock now supplies
  workspaceFqnBindings, which the global fast path reads directly.)
- Correct the workspaceFqnBindings doc comment: it is shared by PHP
  (backslash-FQN keys) and C# (global-namespace simple-name keys); the two
  key formats are disjoint.
- Pre-index parsedFiles by path before the `using static` member-injection
  loop, replacing an O(files) find-per-import with an O(1) Map lookup.

Verified: tsc --noEmit clean; csharp-hooks + csharp-namespace-extraction
suites pass (38 tests); prettier clean; eslint 0 errors.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(csharp): apply PR-review polish to namespace-siblings (tests, types, docs)

Addresses the multi-agent code review of this PR — the concrete, defensible
findings. Two items intentionally deferred (below).

- namespace-siblings.ts: couple the augmentation bucket + its de-dup set into
  one nullable lifecycle, removing the seen!/bucketArr! non-null assertions
  (identical runtime, still lazy).
- validate-bindings-immutability.ts: extend the dev-mode immutability validator
  to the third channel (workspaceFqnBindings) + a test; complete the validator
  test mock with workspaceFqnBindings.
- walkers.ts: document that namesAtScope deliberately excludes the
  scope-independent workspaceFqnBindings channel (enumerating workspace names at
  every scope would flood per-scope callers; lookupBindingsAt still consults it
  when resolving a specific name).
- scope-resolution-indexes.ts: reframe the workspaceFqnBindings doc to describe
  the key-format contract language-neutrally (examples, not language branching).
- csharp-hooks.test.ts: assert workspace entries carry origin:'namespace'; add a
  partial-class test (same simple name, distinct nodeIds across global files →
  both kept); rename the stale "parses" cache-miss test to "scans".
- csharp-pipeline-benchmark.test.ts: clearTimeout the Promise.race budget timer
  (dangling handle when the pipeline won the race).
- csharp.test.ts: correct the #1066 comment — extractFileStructure no longer
  re-parses on cache miss (line scanner); only emitCsharpScopeCaptures re-parses.

Deferred (surfaced, not applied): (1) worker-path scanner mis-reads
namespace/using-static inside block comments and verbatim/raw strings — an
explicitly documented trade-off mirroring the PHP scanner; hardening it to track
comment/string state is a separate decision. (2) workspaceFqnBindings is read
via an `as Map` cast; a type-safe mutable handle from finalize-orchestrator is a
cross-module contract change.

Verified: tsc --noEmit clean; 49 unit tests pass (incl. 3 new); prettier clean;
eslint 0 errors.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(csharp): harden worker-path scanner + localize workspace-map cast

Addresses the two deferred PR-review findings plus the remaining test gap.

#1 — Worker-path scanner false positives: the line scanner now tracks block-
comment and string state across lines (advanceCsScanState), so a `namespace` /
`using static` keyword at the start of a line inside a block comment, verbatim
string (@"..."), or raw string literal ("""...""") is no longer mistaken for a
declaration on the worker cache-miss path. It matches only at code-state line
starts. 5 new scanner tests cover the block-comment / raw / verbatim cases.

#4 — workspaceFqnBindings type safety: the ReadonlyMap->Map cast is localized
to one documented line, and global-namespace writes go through a new
getWorkspaceBucket helper (mirroring getAugmentationBucket) rather than an
inline `.set()` at the mutation site.

#2 — lookupBindingsAt workspace-channel coverage: walkers-augmentations.test.ts
now exercises the third (workspace) channel: workspace-only, append-after-
finalized/augmented, and dedup-loses-to-finalized/augmented precedence.

#5 — OOM CI guard: the deterministic O(D) invariant (zero per-scope
augmentation for global types) is already asserted by the always-on
csharp-hooks unit tests added earlier; the scale/time benchmark stays
appropriately opt-in (skipIf).

Verified: tsc --noEmit clean; 69 unit tests (4 suites) + 210 C# integration
resolver tests pass; prettier clean; eslint 0 errors.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* perf(csharp): replace remaining O(A) .some dedup scans with seeded Sets

The using-static member-injection loop and the cross-namespace import loop both
de-duped via `bucketArr.some((b) => b.def.nodeId === ...)` — O(A) per item. Both
now use a per-file `Map<simpleName, Set<nodeId>>`, seeded lazily from the
augmentation bucket (capturing entries from earlier passes), matching the
global and named-namespace paths. Same dedup semantics, O(1) amortized.

Verified: tsc --noEmit clean; csharp-hooks unit (27) + C# integration resolver
(210) tests pass; prettier + eslint clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(csharp): gate suffix-fallback import resolution to declared namespaces (#1881)

C# `using` directives were resolving via an ungated suffix match, so a BCL
using like `System.Threading.Tasks` matched a coincidental local `Tasks.cs`
and emitted spurious IMPORTS edges. Add a declared-namespace gate that only
permits suffix-fallback when the import plausibly refers to an in-repo
namespace (exact, immediate-parent-declared, or ancestor-of a declared
namespace anchored at an in-repo root). Both resolution legs — the legacy
DAG and the registry-primary scope resolver — thread the same evidence to
the gate, including the no-csproj path.

Declared namespaces are collected with #1905's comment/string-aware scanner
(extractCsharpStructureViaScanner, lazily imported) instead of a regex, so
`namespace` tokens in comments/strings can't seed phantom namespaces. Scan
truncation or unreadable subtrees fail OPEN (gate disabled) and are logged.

Stacked on #1905 (fix/csharp-namespace-scope-oom).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(csharp): cap per-file size in namespace scan; fail open on skip (#1881)

scanCSharpProject read every .cs/.csproj in full with no size guard and
issued per-directory reads with no concurrency bound, an OOM/FD-exhaustion
vector on large or generated repos. Add an fs.stat size guard before each
read, reusing getMaxFileSizeBytes() (the same 512KB cap the Phase-1 walker
uses). An oversized or unreadable .cs now signals truncation so the #1881
suffix-fallback gate fails OPEN rather than wrongly suppressing an import
whose declaring namespace lived in the skipped file (previously a silent
return left the scan looking complete). Adds a size-cap scan test.

* fix(csharp): bound per-directory read concurrency in namespace scan (#1881)

The scan issued every .cs/.csproj read in a directory at once via
Promise.all, so in-flight file descriptors scaled with the largest
directory's file count. Issue reads in bounded windows (32, mirroring
the Phase-1 filesystem-walker) via Promise.allSettled; an unexpected
read/scan rejection now trips truncation (fail open) instead of
rejecting the whole scan. Behavior-preserving for namespace collection
(C# scope-resolution parity passes on both legs).

* style(csharp): apply prettier to #1881 files to clear quality/format gate (#1908)

Reflow hand-wrapped lines in scope-resolver.ts and the csharp integration
test that prettier collapses under printWidth 100. Formatting only, no
behavioral change; clears the failing quality/format CI gate.

* fix(csharp): stream namespace scan so large generated files don't disable the #1881 gate (#1908)

Code-review follow-up. The scan read each .cs fully into a string behind a
512KB size cap (the tree-sitter parse budget); a single larger generated file
(*.g.cs, EF/gRPC output) tripped `truncated`, making the #1881 suffix-fallback
gate fail open repo-wide and silently undoing the fix on real repos.

Stream each .cs line-by-line via createReadStream + readline into a new
incremental scanner (createCsharpStructureScanner) instead of buffering the
whole file. Memory is now constant regardless of file size, so the per-file
size cap is dropped for the namespace line-scan and large generated files are
fully collected. extractCsharpStructureViaScanner is reimplemented on the same
incremental scanner (byte-identical; C# parity 2/2). collectDeclaredNamespaces
returns 'ok' | 'truncated' (truncation now only from an unreadable file) and the
truncation warn lists its real causes. csproj reads keep their size guard.

Prior art: ripgrep/ctags/Node readline stream rather than cap for line scans;
GitHub (384KB) and Sourcegraph (1MB) cap only their full-content indexes.

* fix(csharp): cap .csproj read via stream, not stat-then-read, to clear CodeQL TOCTOU (#1908)

CodeQL js/file-system-race flagged the fs.stat + fs.readFile size guard in
readCsprojConfig as a check-then-use filesystem race. Replace it with a
length-capped createReadStream (readFileTextCapped) — same memory bound on
untrusted input, no stat-then-read race, and consistent with the streamed
.cs scan. Behavior is unchanged for real .csproj files (parity 2/2).

* fix(csharp): keep BCL/external roots gated through scan truncation (#1908, Codex F1)

A single scan truncation (unreadable dir/file, depth/dir cap) set one
repo-wide `truncated` flag that made csharpSuffixFallbackAllowed fail
open for EVERY import, silently re-enabling the #1881 BCL->local suffix
matches. Add a CSHARP_EXTERNAL_ROOTS denylist (System/Microsoft/...): an
external-rooted using that does not align with an in-repo declared
namespace stays BLOCKED even under truncation, while genuinely
local-looking usings still fail open. A repo that declares the root is
allowed via the alignment escape hatch. Shared predicate, so both legs
inherit it.

* fix(csharp): gate the registry no-csproj direct-match path (#1908, Codex F2)

In the no-csproj branch of resolveCsharpImportTarget, resolveDirectMatch
ran BEFORE the gate, so a path-aligned Legacy/System/Threading/Tasks.cs
satisfied 'using System.Threading.Tasks;' even though System.* is not a
declared in-repo namespace — while the legacy leg (gate-first) blocked
it, so the legs were not equivalent. Run csharpSuffixFallbackAllowed
first (return null on fail), then direct-match, then progressive
stripping — mirroring the legacy ordering. Adds a no-csproj fixture with
a deep path-aligned Tasks.cs and dual-leg integration describes (registry
+ forced-legacy), plus a path-aligned unit case. Parity 2/2.

* fix(csharp): flag scanner-uncaptured namespaces incomplete; Unicode/@ matchers (#1908, Codex F3)

The line scanner treated its output as complete even when it missed valid
C# namespace forms, so the gate failed CLOSED and over-blocked legit
imports. Make CS_NAMESPACE_RE/CS_USING_STATIC_RE Unicode-aware (\p{L}\p{N}
+ u flag) and strip leading/segment @ so verbatim/Unicode identifiers are
captured to match the AST. For forms the regex still can't capture (split
across lines, not at line start, attributed), set a per-file 'incomplete'
flag; collectDeclaredNamespaces returns 'truncated' for such files so the
#1881 gate fails OPEN instead of dropping the namespace. High-precision
detectors + guard tests keep ordinary forms (incl. // namespace comments)
from tripping incomplete.

* fix(csharp): stream the .csproj RootNamespace read, no byte cap (#1908, Codex F4)

readCsprojConfig read only the first 512KB of a .csproj and, on a
match-miss, couldn't tell 'no RootNamespace' from 'RootNamespace past
the cap' — both synthesized a filename root. A wrong authoritative root
makes imports under the real root resolve to nothing AND suppresses the
fallback. Replace the capped read with a streamed early-stop search
(findCsprojRootNamespace) that reads until the tag or EOF: filename
fallback ONLY on genuine read-to-EOF absence; on a soft-budget cap-hit or
unreadable file, OMIT the config so the no-csproj fallback stays
reachable. Removes the now-unused readFileTextCapped + getMaxFileSizeBytes
cap from the scan. Parity 2/2.

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 09:56:26 +01:00
Gergő Magyar a62a7e56cb Merge branch 'main' into optimize/go-scope-capture 2026-05-30 09:31:52 +01:00
Gergő MagyarandClaude Opus 4.8 f18ff521fc fix(group): stop Node gRPC loadPackageDefinition gate from matching every member call (#1916)
LOAD_PACKAGE_DEFINITION_SPEC matched `loadPackageDefinition` via a single
`function: [ (identifier) @fn (#eq?) (member_expression property:(property_identifier) @fn (#eq?)) ]`
alternation. Under the pinned tree-sitter@0.21.1 binding a top-level alternation
whose branches reuse one capture name collapses to a single pattern with a shared
predicate bucket: the member-expression branch's `@fn` is left unbound and its
`#eq?` is never enforced, so that branch matches EVERY `obj.method(...)` call
(`console.log(...)`, `logger.info(...)`, …). Since virtually every TS/JS file has
some member call, the `usesLoadPackage` gate was effectively always-open and
`new pkg.<Capitalized>Service(...)` was emitted as a spurious gRPC consumer — the
exact false positive the gate was added to prevent.

Split the spec into two single-branch PatternSpecs; each compiles to its own
Parser.Query with an independent predicate bucket where the `#eq?` is enforced
correctly. `runCompiledPatterns` concatenates their matches, so the
`.length > 0` gate is unchanged. `mk` now accepts a spec or a spec array.

Adds test_extract_ts_qualified_ctor_without_loadPackageDefinition_is_ignored, a
negative regression test verified to FAIL on the pre-fix code and PASS with the
fix: a file with no loadPackageDefinition but an unrelated member call +
`new authProto.auth.v1.AuthService(...)` must emit no consumer.

grpc-extractor suite 65/65; tsc + prettier + pre-commit hook clean.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 09:31:36 +01:00
Gergő Magyar 150a95bae4 Merge branch 'main' into optimize/go-scope-capture 2026-05-30 08:50:56 +01:00
5d710413d7 feat(group): extract OpenFeign @RequestLine consumer contracts (#1904)
* feat(group): extract OpenFeign @RequestLine consumer contracts

Adds Java HTTP plugin support for the native OpenFeign annotation
`@RequestLine("METHOD /path")`. Previously only `@FeignClient` interfaces
using Spring MVC method annotations (`@GetMapping` etc.) were detected;
the native annotation form — required by Feign Builder users and
non-Spring Feign deployments — was silently ignored.

Implementation:

- New `FEIGN_REQUEST_LINE_PATTERNS` covers both positional and named-arg
  (`value =`) forms.
- New `parseRequestLine()` parses the verb+path string and drops any
  query string (consistent with how RestTemplate/WebClient consumers
  handle inline literal URLs).
- The enclosing interface MUST carry `@FeignClient`; otherwise the
  detection is dropped to avoid false positives from same-named
  annotations in unrelated libraries.
- Reuses the existing `feignPrefixByInterfaceId` map so
  `@FeignClient(path=)` and `@RequestMapping` interface prefixes apply
  uniformly across both Spring MVC and `@RequestLine` methods.
- Confidence 0.75 — slightly higher than the 0.7 used for Spring MVC
  annotations because the verb is a string-literal value, not inferred
  from the annotation name (less ambiguous).

Six new unit tests cover: basic two-method extraction; `@FeignClient(path=)`
  prefix joining; query-string stripping; rejection of `@RequestLine` on
  non-Feign interfaces; mixing with `@GetMapping` on the same interface;
  named-argument form (`value = "..."`).

Verification: `npx tsc --noEmit`, full `test/unit/group` (31 files / 563
tests), `http-route-extractor.test.ts` (83/83 incl. 6 new), `prettier
--check` and `eslint` on touched files all pass.

* refactor(group): collapse @RequestLine positional + named-arg into one query

Per @magyargergo's review on PR #1904 — uses tree-sitter alternation
`[(...) (...)]` so the positional and named-argument forms of the
`@RequestLine` annotation are matched by a single compiled query and
invoked through one `runCompiledPatterns` pass instead of two.

* refactor(group): drop framework prefixes from java http pattern constant names

Per review feedback on #1904 — renames the four route-mapper pattern
constants to framework-agnostic names (the per-constant comments already
document which framework each targets):
  SPRING_TYPE_PREFIX_PATTERNS     -> TYPE_PREFIX_PATTERNS
  FEIGN_REQUEST_LINE_PATTERNS     -> REQUEST_LINE_PATTERNS
  FEIGN_INTERFACE_PREFIX_PATTERNS -> INTERFACE_PREFIX_PATTERNS
  SPRING_METHOD_ROUTE_PATTERNS    -> METHOD_ROUTE_PATTERNS

* refactor(group): collapse Java route-mapper annotations into one query

Merge the four annotation pattern bundles (Spring @RequestMapping type
prefix, @FeignClient(path) prefix, @(Get|Post|Put|Delete|Patch)Mapping
method routes and native @RequestLine) into a single
JAVA_ROUTE_ANNOTATION_PATTERNS query, read by scanRouteAnnotations() in
exactly one matches() pass per file. Variants are tagged by branch-local
captures and discriminated in JS (METHOD_ANNOTATION_TO_HTTP,
isRouteMemberKey), per review feedback. This drops the per-file annotation
passes from 4->1 in scan() and 2->1 in collectSpringTypes(), and removes
the interface-@RequestMapping / @FeignClient prefix redundancy.

Verb and path/value key filtering stay in JS rather than in-query: under
the pinned tree-sitter 0.21.1 binding a top-level [...] alternation
compiles to one pattern whose text predicates share a single bucket keyed
by capture name. A #match? against a capture absent from the matched
branch evaluates FALSE and silently drops every sibling-branch match,
whereas #eq? against an absent capture is vacuously true. So only fixed
annotation names use in-query #eq? (on branch-local captures); the
variable verb name and member key carry no in-query predicate.

Behaviour is unchanged for all compilable Java; existing http-route tests
(93) and the full group suite remain green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(group): make Java route-annotation query generic, match name in loop

Collapse JAVA_ROUTE_ANNOTATION_PATTERNS from 9 annotation-name-pinned
branches to 6 generic structural branches (class/interface/method x
positional/named) that capture the annotation name (@ann), declaration
(@node), argument (@value) and member key (@key) generically. The query
now carries NO #eq?/#match? predicates at all; scanRouteAnnotations reads
@ann.text and @node.type in its for-loop to decide what each match means
(RequestMapping prefix, FeignClient(path) prefix, @(Get|...)Mapping route,
or @RequestLine), ignoring unrecognised annotations.

This makes the query framework-agnostic and extensible — adding a new
route annotation is a change to the loop and the lookup maps, not the
query — and removes the last tree-sitter-0.21.1 shared-predicate-bucket
footgun, since a predicate-free alternation cannot drop sibling branches.

Behaviour is byte-identical: 93 targeted http-route tests and the full
569-test group suite stay green; tsc and prettier clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(group): pin newly-reachable Java route-annotation JS branches; clarify invariants

Code-review follow-up to the route-annotation query consolidation. No
behaviour change to the extractor:

- Add two regression tests for branches the generic predicate-free query
  made reachable in scanRouteAnnotations: (1) a @RequestLine whose named
  argument is not `value` must be dropped (the in-query `#eq? @key "value"`
  guard now lives in JS); (2) @FeignClient(path) must win over @RequestMapping
  even when @RequestMapping is the first annotation in source order, covering
  the deferred interfaceRequestMappingPrefixes apply (the existing precedence
  test only covered @FeignClient-first).
- Document two invariants flagged in review: why prefixByTypeId and
  feignPrefixByInterfaceId intentionally diverge for the same interface node
  (Spring provider vs OpenFeign consumer prefix), and that the query's
  single-string-argument shape excludes array-valued annotations.

http-route-extractor + multi-verb suites: 95/95 (was 93); tsc + prettier clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: henry <zhangwei2017@unipus.cn>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 08:50:29 +01:00
Gergo MagyarandClaude Opus 4.8 ed33489e38 test(go-scope-capture): address code-review findings
Self-review (ce-code-review) polish on the #1848 fix + benchmark:

- benchmark: tighten the scaling guard from timeRatio/fileRatio < 3 to < 1.5.
  At the 2.5x/2x scale steps, a quadratic regression yields ratio == fileRatio
  (2.5, 2.0), which < 3 waved through — the guard could not detect the O(n^2)
  it exists for. Measured O(n) ratios are 0.45/0.59, so < 1.5 has headroom.
- benchmark: add a non-gated O(n^2) regression tripwire that calls
  emitGoScopeCaptures on a 400-struct source directly (no worker, no
  GITNEXUS_BENCH gate) so the regression is actually guarded in CI.
- benchmark: clearTimeout the Promise.race timer in finally (no lingering
  rejection); set the worker-suite env vars inside the try so finally always
  restores them.
- captures.ts: clarify the isRawMultiAssignTypeBinding comment to name both
  var-form cases (assertion + call-return). Comment-only.

Left as-is: resolveImportNode's defensive range-equality branch — deleting it
as dead code would remove the self-documentation of the grammar invariant the
threaded-node logic depends on (reviewer tension; a wash).

Verified: tsc clean; 165/165 Go resolver + scope tests; new tripwire passes
(237ms); scaling suite passes at <1.5; #1848 worker suite still green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 07:33:54 +00:00
Gergo MagyarandClaude Opus 4.8 eaf0a3052a optimize(go-scope-capture): thread captured nodes to kill O(n^2) findNodeAtRange re-walks
emitGoScopeCaptures re-derived each match's AST node via findNodeAtRange from
the tree root on every query match, giving O(matches x rootChildren) ~ O(n^2)
behaviour (the #1848 root cause: a 250-struct generated DAO took ~10.8s, 800
structs ~100s+ — long enough to trip the worker sub-batch idle timeout and get
quarantined). Thread the query-captured SyntaxNode (c.node) through a parallel
tag->node map and use it directly (or via a bounded local parent walk for the
import_declaration ancestor case) instead of re-walking from root.

Output is byte-identical (capture fingerprint over the DAO file + all 89 go-*
fixtures unchanged; capture_groups=13501). 250 entities: 10835ms -> 114ms (95x).
800 entities: ~100s -> 384ms. Go resolver + scope-resolution suites: 165/165 pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 07:09:25 +00:00
Gergo MagyarandClaude Opus 4.8 5d1695f66a test(go): add #1848 Go pipeline + worker-pool benchmark
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 07:09:24 +00:00
dependabot[bot] bef3da59a7 chore(deps)(deps): bump node-addon-api from 8.7.0 to 8.8.0 in /gitnexus (#1911)
Bumps [node-addon-api](https://github.com/nodejs/node-addon-api) from 8.7.0 to 8.8.0.
- [Release notes](https://github.com/nodejs/node-addon-api/releases)
- [Changelog](https://github.com/nodejs/node-addon-api/blob/main/CHANGELOG.md)
- [Commits](https://github.com/nodejs/node-addon-api/compare/v8.7.0...v8.8.0)

---
updated-dependencies:
- dependency-name: node-addon-api
  dependency-version: 8.8.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-30 06:23:31 +01:00
azizur100389andGergő Magyar e234dac849 feat(cpp): add template partial ordering (#1885)
* feat(cpp): add template partial ordering

* fix(cpp): harden template partial ordering

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-29 21:38:19 +01:00
26894d835b fix(csharp): eliminate namespace-siblings OOM and worker-path re-parse (#1905)
* fix(csharp): eliminate O(S·D) BindingRef OOM in namespace siblings

Types declared in the C# global (default) namespace are visible from
every file, so the previous per-scope augmentation materialized
O(scopes × defs) BindingRefs — on large Unity solutions (tens of
thousands of global types) this caused severe slowness and OOM.

Route global-namespace types through a single workspace-level binding
channel (workspaceFqnBindings, consulted by lookupBindingsAt) for O(D)
memory. Also fix quadratic costs in the non-global path: append defs in
place instead of copying (was O(D²) per bucket), pre-index the first
scope per file (was O(S²·D)), and seed de-dup sets instead of repeated
.some scans.

Add csharp-pipeline-benchmark.test.ts (mirrors the PHP benchmark) with
spread and concentrated-global-namespace scenarios to track elapsedMs,
peakHeapMB, nodeCount, and edgeCount. Post-fix runs show linear scaling
and stable heap.

Co-authored-by: Cursor <cursoragent@cursor.com>

* perf(csharp): scanner fallback for namespace siblings on the worker path

Worker threads can't return tree-sitter Trees across MessageChannels, so
the cross-phase tree cache is empty for worker-parsed files. The C#
same-namespace pass (populateCsharpNamespaceSiblings -> extractFileStructure)
then re-parsed every file with tree-sitter to find namespace / using-static
nodes — effectively parsing a large solution a second time during scope
resolution.

Add a line-scanner fallback (extractCsharpStructureViaScanner) used only
when no cached Tree is available, mirroring PHP's fix for issue #1741. It
extracts the same namespaces / usingStaticPaths the AST walk produces for
the common line-anchored forms (file-scoped + block namespaces, plain /
global / aliased `using static`). The AST walk stays authoritative on the
sequential / warm-cache path.

Micro-benchmark over 3000 synthetic files: scanner is ~188x faster than
parse+walk (0.001 vs 0.251 ms/file) with identical output on the parity
spot-check; real-world files are larger, so the worker-path saving is
bigger. Adds csharp-namespace-extraction.test.ts (12 cases) covering all
declaration forms plus negative cases (using var, plain using, comments).

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(csharp): cover global-namespace workspaceFqnBindings path + doc + using-static perf

Addresses the production-readiness review of the namespace-siblings OOM fix.

- Add a unit test proving global-(default-)namespace C# types route to
  indexes.workspaceFqnBindings (one entry per simple name) with ZERO
  bindingAugmentations — pinning the O(D) invariant behind the #1871
  Unity-scale OOM fix and guarding against a revert to per-scope
  O(scopes x defs) augmentation. (The csharp-hooks mock now supplies
  workspaceFqnBindings, which the global fast path reads directly.)
- Correct the workspaceFqnBindings doc comment: it is shared by PHP
  (backslash-FQN keys) and C# (global-namespace simple-name keys); the two
  key formats are disjoint.
- Pre-index parsedFiles by path before the `using static` member-injection
  loop, replacing an O(files) find-per-import with an O(1) Map lookup.

Verified: tsc --noEmit clean; csharp-hooks + csharp-namespace-extraction
suites pass (38 tests); prettier clean; eslint 0 errors.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(csharp): apply PR-review polish to namespace-siblings (tests, types, docs)

Addresses the multi-agent code review of this PR — the concrete, defensible
findings. Two items intentionally deferred (below).

- namespace-siblings.ts: couple the augmentation bucket + its de-dup set into
  one nullable lifecycle, removing the seen!/bucketArr! non-null assertions
  (identical runtime, still lazy).
- validate-bindings-immutability.ts: extend the dev-mode immutability validator
  to the third channel (workspaceFqnBindings) + a test; complete the validator
  test mock with workspaceFqnBindings.
- walkers.ts: document that namesAtScope deliberately excludes the
  scope-independent workspaceFqnBindings channel (enumerating workspace names at
  every scope would flood per-scope callers; lookupBindingsAt still consults it
  when resolving a specific name).
- scope-resolution-indexes.ts: reframe the workspaceFqnBindings doc to describe
  the key-format contract language-neutrally (examples, not language branching).
- csharp-hooks.test.ts: assert workspace entries carry origin:'namespace'; add a
  partial-class test (same simple name, distinct nodeIds across global files →
  both kept); rename the stale "parses" cache-miss test to "scans".
- csharp-pipeline-benchmark.test.ts: clearTimeout the Promise.race budget timer
  (dangling handle when the pipeline won the race).
- csharp.test.ts: correct the #1066 comment — extractFileStructure no longer
  re-parses on cache miss (line scanner); only emitCsharpScopeCaptures re-parses.

Deferred (surfaced, not applied): (1) worker-path scanner mis-reads
namespace/using-static inside block comments and verbatim/raw strings — an
explicitly documented trade-off mirroring the PHP scanner; hardening it to track
comment/string state is a separate decision. (2) workspaceFqnBindings is read
via an `as Map` cast; a type-safe mutable handle from finalize-orchestrator is a
cross-module contract change.

Verified: tsc --noEmit clean; 49 unit tests pass (incl. 3 new); prettier clean;
eslint 0 errors.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(csharp): harden worker-path scanner + localize workspace-map cast

Addresses the two deferred PR-review findings plus the remaining test gap.

#1 — Worker-path scanner false positives: the line scanner now tracks block-
comment and string state across lines (advanceCsScanState), so a `namespace` /
`using static` keyword at the start of a line inside a block comment, verbatim
string (@"..."), or raw string literal ("""...""") is no longer mistaken for a
declaration on the worker cache-miss path. It matches only at code-state line
starts. 5 new scanner tests cover the block-comment / raw / verbatim cases.

#4 — workspaceFqnBindings type safety: the ReadonlyMap->Map cast is localized
to one documented line, and global-namespace writes go through a new
getWorkspaceBucket helper (mirroring getAugmentationBucket) rather than an
inline `.set()` at the mutation site.

#2 — lookupBindingsAt workspace-channel coverage: walkers-augmentations.test.ts
now exercises the third (workspace) channel: workspace-only, append-after-
finalized/augmented, and dedup-loses-to-finalized/augmented precedence.

#5 — OOM CI guard: the deterministic O(D) invariant (zero per-scope
augmentation for global types) is already asserted by the always-on
csharp-hooks unit tests added earlier; the scale/time benchmark stays
appropriately opt-in (skipIf).

Verified: tsc --noEmit clean; 69 unit tests (4 suites) + 210 C# integration
resolver tests pass; prettier clean; eslint 0 errors.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* perf(csharp): replace remaining O(A) .some dedup scans with seeded Sets

The using-static member-injection loop and the cross-namespace import loop both
de-duped via `bucketArr.some((b) => b.def.nodeId === ...)` — O(A) per item. Both
now use a per-file `Map<simpleName, Set<nodeId>>`, seeded lazily from the
augmentation bucket (capturing entries from earlier passes), matching the
global and named-namespace paths. Same dedup semantics, O(1) amortized.

Verified: tsc --noEmit clean; csharp-hooks unit (27) + C# integration resolver
(210) tests pass; prettier + eslint clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 20:24:58 +01:00
Gergő MagyarandClaude Opus 4.8 2a5bbbeaae fix: make extension installs offline-first (#1161)
* feat(review): add PR reviewer swarm agents

Seven read-only subagents coordinated by an orchestration skill for
structured, evidence-grounded production-readiness PR reviews.

Agents: facts-historian, branch-hygiene, risk-architect, test-ci-verifier,
security-boundary, docs-dod, synthesis-critic. All use Read/Grep/Glob/Bash
only — no edit tools.

Skill invoked as /gitnexus-pr-swarm-review <PR>.

* fix: patch vector extension and uncaughtException for review findings

- Add { policy: 'auto' } to both loadVectorExtension() calls in
  embedding-pipeline.ts so analyze --embeddings auto-installs VECTOR
- Add void to uncaughtException shutdown(1) call for Node v20+ safety
- Re-add getExtensionInstallPolicy export + default change + 4 tests

* fix(mcp,lbug): graceful shutdown exit codes + complete offline-first VECTOR policy

Completes the two live issues PR #1161 only partially addressed.

#1132 — MCP shutdown crash: SIGINT/SIGTERM were registered with `shutdown`
directly, so Node passed the signal NAME string into process.exit(), crashing
with ERR_INVALID_ARG_TYPE ('SIGTERM'). Map signals to numeric exit codes
(SIGINT->130, SIGTERM->143) via a testable installSignalShutdown(); add an
unref'd force-exit watchdog so a hung disconnect()/close() cannot wedge
shutdown; and void the stdin/stdout handlers so event payloads never reach
process.exit() as a non-number.

#1153 — offline-first extension loading:
- semanticSearch (a query/read path) no longer forces policy:'auto'; queries
  use load-only and never spawn a network INSTALL (extension.ladybugdb.com).
- the analyze embedding WRITE path resolves the policy from
  GITNEXUS_LBUG_EXTENSION_INSTALL (honoring never/load-only/auto; default auto)
  instead of hard-forcing 'auto', so an offline/locked-down operator's override
  is respected (the regression that re-broke #1153 for the VECTOR path).
- surface the active install policy in `gitnexus doctor` (was claimed but never
  delivered; also gives the previously-dead getExtensionInstallPolicy a caller).
- emit an actionable message when VECTOR is unavailable.

Tests: regression for the signal->numeric mapping (reproduces the signal-string
crash condition) and for embedding install-policy resolution. tsc/prettier clean,
eslint 0 errors, 55 unit tests pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(analyze): degrade gracefully when FTS extension is unavailable

The load-only default made `gitnexus analyze` throw when the FTS
extension was not pre-installed, breaking CI and offline use. Make the
analyze write path opt into the `auto` install policy (LOAD-first then
bounded INSTALL — symmetric with the VECTOR/embeddings path and the #726
contract) and degrade gracefully when the extension still cannot load:
skip search-index creation, log a warning, and complete with a fully
queryable graph (only full-text/BM25 search is disabled). `--repair-fts`
still fails loudly.

- Surface the degraded state instead of reporting healthy:
  AnalyzeResult.ftsSkipped, a persistent CLI summary warning, and
  meta.json capabilities.fts.status = "unavailable".
- Skip the FTS-primitive integration tests when the extension is
  unavailable (shared skipUnlessFtsAvailable helper).
- Add a unit test for the degradation branch; fix the existing
  full-analyze test mock that omitted loadFTSExtension.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(lbug): skip FTS-seeding suites when extension is unavailable

The withTestLbugDB helper seeds FTS indexes in beforeAll via createFTSIndex,
which throws when the optional FTS extension cannot load — failing the whole
suite on machines where it is neither pre-installed nor installable (the
macOS platform-sensitive CI runner). Probe the extension once (mirroring the
analyze write path's `auto` policy), bypass FTS seeding when it is
unavailable, and skip the suite's tests via beforeEach with a one-time
warning so the skip is visible rather than a setup crash.

Fixes the macOS failures in search-core, search-pool, local-backend-calltool,
and staleness-and-stability. Suites still run normally where FTS is available.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 20:04:41 +01:00
252fbabd51 fix(ingestion): stop emitting phantom Function defs for array-method callbacks (#1906)
* fix(ingestion): stop emitting phantom Function defs for array-method callbacks

The HOC-wrapped-arrow scope-query pattern (`const X = HOC(args => ...)`),
added for React idioms such as forwardRef/memo/useCallback, also matched
array higher-order-method callbacks like `const x = arr.map(a => ...)`.
Those produced a spurious `@declaration.function` named after the
binding, on top of its value def, so calls inside the callback attributed
to a phantom `Function:x` instead of the enclosing scope.

- Add a shared `isArrayMethodCallbackArrow` detector
  (`ARRAY_CALLBACK_METHODS` blocklist) and suppress the
  `@declaration.function` emit-side in both the JS and TS scope-captures
  emitters, leaving the value binding as the sole def.
- Add `selectNodeBearingDef` in scope-extractor: the tested
  collapse-rule contract (function-like > value > first) the deferred
  node-creation migration will consume to keep one graph node per
  binding.

This corrects the registry-primary scope model and CALLS-edge
attribution (calls inside array-method callbacks now source from the
enclosing File scope). The duplicate graph *node* itself is still
created by the legacy parse-worker path and is removed by the follow-up
node-creation migration.

Refs #1876

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(ingestion): strengthen array-callback coverage; document receiver-blind suppression

Follow-ups from the production-readiness review of PR #1906:

- array-callback.ts: document that isArrayMethodCallbackArrow is
  receiver-blind — an in-set method name on a NON-array receiver
  (Map/Set.forEach, RxJS observable.map, query-builder .sort, lodash
  chain .filter) is also suppressed. Accepted limitation, not a bug:
  the binding holds the call's result value, not a callable.
- captures unit tests (JS + TS): add a non-array-receiver
  characterization case, and extend the it.each lists to cover
  findLast, findLastIndex, reduceRight — the full 13-entry
  ARRAY_CALLBACK_METHODS set is now exercised in both languages.
- js-array-method-callback-attribution integration test: tighten the
  File-sourced CALLS assertions from toBeGreaterThan(0) to
  toHaveLength(1) (now also catches over-attribution).
- scope-extractor.ts: note that the dead selectNodeBearingDef export is
  intentional and tracked by #1876 (deferred node-creation migration).

Comment-and-test only; no production behavior change. Verified locally:
tsc clean, prettier/eslint clean, captures unit 106 passed,
scope-extractor 31 passed, integration 3 passed.

Refs #1876

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 19:35:06 +01:00
Gergő MagyarandClaude Opus 4.8 85727ca625 feat(review): add PR reviewer swarm agents (#1851)
* feat(review): add PR reviewer swarm agents

Seven read-only subagents coordinated by an orchestration skill for
structured, evidence-grounded production-readiness PR reviews.

Agents: facts-historian, branch-hygiene, risk-architect, test-ci-verifier,
security-boundary, docs-dod, synthesis-critic. All use Read/Grep/Glob/Bash
only — no edit tools.

Skill invoked as /gitnexus-pr-swarm-review <PR>.

* Address PR review feedback (#1851)

- Pin explicit model IDs in all 7 reviewer-swarm agents per CLAUDE.md
  (no unversioned aliases). Set the two mechanical agents
  (test-ci-verifier, branch-hygiene-reviewer) to claude-haiku-4-5-20251001
  per @Cenrax's "this could be haiku"; the five analytical agents use
  claude-sonnet-4-6.
- Add an explicit read-only Bash policy (permitted/prohibited command
  lists) to every agent's Rules section, so the read-only guarantee is
  defended against injected/adversarial PR content rather than prose-only.
- Add a hard synthesis-critic gate to the swarm skill: do not post the
  final review until the critic's "Required corrections before posting"
  section is empty (was advisory only).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(review): make PR reviewer swarm portable across AI CLIs

Restructure the reviewer swarm around a single CLI-neutral source of truth so it
runs from any AI CLI, not just Claude Code.

- pr-swarm-review/: canonical orchestration.md (Swarm + Solo execution modes with
  an identical output contract) and personas/0N-*.md (the 7 review personas,
  relocated verbatim from the Claude agents, each tagged with a model tier and the
  read-only Bash policy). Single source of truth — edit here, not in the wrappers.
- Thin per-CLI adapters that read the canonical spec at runtime (no duplication):
  - Claude Code: coordinator skill (Swarm mode) + the 7 agents are now thin
    wrappers that read their persona file (frontmatter/model preserved; mechanical
    lanes Haiku, analytical lanes Sonnet).
  - Gemini CLI: .gemini/commands/gitnexus-pr-swarm-review.toml
  - GitHub Copilot: .github/prompts/gitnexus-pr-swarm-review.prompt.md
  - Cursor: .cursor/commands/gitnexus-pr-swarm-review.md
- AGENTS.md: canonical "PR Swarm Review" section -> orchestration.md, the universal
  entrypoint honored by Codex, Cursor, Gemini, Copilot, and any AGENTS.md-aware
  agent (Codex user-level prompt install noted in the README).

Graceful degradation: only Claude Code has parallel subagents (Swarm mode); every
other CLI runs the 7 lanes sequentially in one agent (Solo mode) with the same
output contract. prettier --check clean (root config).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 18:24:16 +01:00
4bc8622642 fix(group): derive grpc consumer FQN from java imports for client-jar consumers (#1889)
* fix(group): derive grpc consumer FQN from java imports so client-jar consumers don't fall back to short names

Java gRPC microservices commonly follow the "client-jar" pattern: the
service owner publishes a pre-compiled stub jar to a Maven repository
and consumer repos depend on the jar instead of carrying the
originating `.proto` files. gRPC's official Java quickstart, Alibaba
HSF, ByteDance KiteX-Java and google-cloud-java all document this
shape.

Before this commit, `GrpcExtractor` resolved a fully-qualified
contract id (`grpc::<package>.<Service>/*`) only when the consumer
repo also carried a matching `.proto`. Client-jar consumers had no
proto, so they fell back to a short-name contract id
(`grpc::<Service>/*`) that never matched the provider's contract id.
Cross-repo grpc cross-link counts dropped to zero on every realistic
Java microservice group — including all of crsdp's `crsdp-backend →
unipus_cloud_framework` connections.

Fix: derive the proto package directly from each consumer file's
`import <pkg>.<XxxGrpc>;` statement. The package from the import is
exactly the proto package, so the contract id matches the provider's
verbatim — no `.proto` lookup needed in the consumer repo.

Implementation
--------------

* `grpc-patterns/types.ts` — `GrpcDetection` gains an optional
  `protoPackage` field. Plugins set it when the package can be
  derived from the source file alone.
* `grpc-patterns/java.ts` — adds `GRPC_CLASS_IMPORT_PATTERNS`, a
  tree-sitter query that captures every
  `import_declaration > scoped_identifier { scope, name }` pair where
  the imported name ends in `Grpc`. `import static …` and
  `import w.x.*;` are excluded by tree-sitter shape: the `name:` field
  is only present on the non-static, non-wildcard form. The plugin
  builds a per-file `XxxGrpc → fullPackage` map and tags every
  provider / consumer detection it emits.
* `grpc-extractor.ts` — `detectionToContract()` now resolves the
  contract id in three steps:
    1. detection-supplied `protoPackage` wins (skips the proto map
       entirely so an unrelated same-name service in the consumer
       repo can't blur the FQN);
    2. otherwise consult the legacy per-repo proto map;
    3. otherwise fall back to a short-name contract id, preserving
       pre-fix behaviour.
  Confidence stays at the "with proto" tier when the import path
  resolves: an import statement in real source is at least as
  authoritative as a per-repo proto map.

Same-short-name disambiguation
-------------------------------

The motivating case `unipus_cloud_framework` defines two distinct
`ContentRpcService` services in different proto packages
(`cn.unipus.ucf.api.proto.client.service.ContentRpcService` vs
`cn.unipus.ucf.admin.proto.client.service.ContentRpcService`). Two
consumer files importing the two flavours now emit two distinct FQNs;
neither could be told apart from the other under the legacy short-
name fallback.

Out of scope
------------

`import w.x.*;` (wildcard service imports) are left to the legacy
short-name fallback. Wildcard imports are discouraged by Google's
Java style guide and IntelliJ's defaults, and resolving them
unambiguously would require either group-level proto-package
catalogs or per-class disambiguation, both of which are larger
follow-ups. This commit only changes behaviour for the dominant
specific-import case.

Tests
-----

`test/unit/group/grpc-extractor.test.ts` adds a new "Java client-jar
consumer (import-derived FQN)" describe block with 9 cases covering
both the happy paths (consumer/provider FQN derivation, same-short-
name disambiguation, import-vs-local-proto precedence) and the
regression-protection paths (no import + no detection emitted, static
imports / wildcards ignored, mixed-file repos preserved).

End-to-end verification
-----------------------

Ran the patched cli on the real `crsdp-backend` (consumer, no
`.proto`) and `unipus_cloud_framework` (provider, has `.proto`)
repos. Synced as a two-repo group, every `XxxGrpc` referenced via a
specific import in `UcfAdminGrpcClientService.java` produced an FQN
contract id that exact-matched the provider repo's FQN — 9 grpc
cross-links surfaced where there were 0 before.

Verification
------------

* `npx tsc --noEmit`: pass
* `npx tsc` (dist rebuild): pass
* `test/unit/group/grpc-extractor.test.ts`: 60/60 pass (51 existing
  + 9 new)
* `test/unit/group/`: 30 files / 545 tests all green
* `npx prettier --check` on touched files: pass
* `npx eslint` on touched src files: 0 errors / 0 warnings

* fix(group): handle option java_package and proto-map disagreement in grpc detection

Addresses Claude bot review on PR #1889:

- Finding 1: parse `option java_package` when building proto context;
  add a reverse index so an import-derived package can be translated
  back to the proto package.
- Finding 2: when same-repo proto map has the service, use the proto
  package; warn and record `meta.importPackage` if the import disagrees.
- Finding 3: add an end-to-end wildcard match test (provider+consumer
  fixture, runs `buildProviderIndex`+`runWildcardMatch`).

Client-jar consumer + diverging `java_package` (no local proto)
remains a known limitation; pinned by a dedicated test.

---------

Co-authored-by: henry <zhangwei2017@unipus.cn>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-29 14:55:28 +01:00
23bf594a70 fix(standalone): wire standalone providers into scope-extractor for registry-primary (COBOL Ring 3 flip) (#1842)
* feat(cobol): migrate COBOL to scope-based resolution (regex provider)

Migrate COBOL to scope-based registry resolution, validating the
parse-source-agnostic contract — COBOL uses regex, not tree-sitter,
but implements the same LanguageProvider interface via emitScopeCaptures.

Phase 1-5 complete per #941 DoD.

New files:
  languages/cobol/captures.ts       — emitScopeCaptures wrapping regex tagger
  languages/cobol/interpret.ts      — import/type-binding/receiver hooks
  languages/cobol/index.ts          — barrel export
  languages/cobol/scope-resolver.ts — ScopeResolver wiring (9 fields, 3 toggles)

Modified files:
  languages/cobol.ts                — wire 4 scope-resolution hooks
  registry.ts                       — register cobolScopeResolver
  registry-primary-flag.ts          — document REGISTRY_PRIMARY_COBOL

Fixtures:
  17 fixture files, 30 test cases across 11 required classes
  test/integration/resolvers/cobol-scope.test.ts

Tests: 24/24 pass (default + REGISTRY_PRIMARY_COBOL=0)
tsc: zero cobol-specific errors
Shadow mode (GITNEXUS_SHADOW_MODE=1): zero crashes
Regex perf: 10K-line file in 408ms (threshold: 2000ms)

NOT added to MIGRATED_LANGUAGES — REGISTRY_PRIMARY_COBOL env var only.

* chore(cobol): add COBOL to MIGRATED_LANGUAGES

* fix(cobol): revert MIGRATED_LANGUAGES flip, fix JSDoc dup, fix arityCompatibility

* fix(standalone): wire standalone providers into scope-extractor for registry-primary (COBOL Ring 3 flip)

- Gate cobolPhase with isRegistryPrimary() guard to prevent double emission
- Wire standalone providers (parseStrategy !== 'tree-sitter') with
  emitScopeCaptures into parse-worker via extractParsedFile bridge
- Add COBOL to MIGRATED_LANGUAGES in registry-primary-flag.ts
- Fix Module scope range in captures.ts to use full program bounds
  (was just PROGRAM-ID line, causing scope containment failures)
- Update cobol.test.ts grand totals to be mode-aware
- Wrap legacy exact-count assertions in if (!isPrimary)
- Fix cobol-scope.test.ts fixture path to use __dirname (was process.cwd())

Tests:
  REGISTRY_PRIMARY_COBOL=0: 83/83 pass (59 legacy + 24 capture)
  REGISTRY_PRIMARY_COBOL=1: 28/28 pass (4 mode-aware + 24 capture)

* test(cobol): restore original test assertions, add mode-aware describe blocks alongside

- Remove if (!isPrimary) wrapper from legacy assertions
- Keep ALL 59 original tests intact and running unconditionally
- Add new 'scope-resolution mode' describe block alongside legacy tests
- New block uses isPrimary to check for scope-resolution capture output
- Legacy tests run against cobolPhase output (skipGraphPhases=true)
- Mode-aware tests validate standalone provider wiring in registry-primary mode

* fix(test): use result.graph instead of result.parsedFiles in scope-mode test

- PipelineResult has no parsedFiles field; use graph.nodes instead
- Use toBe strict equality (not.toBeNull()) per review feedback
- Object.keys for node count as suggested by reviewer

* test(cobol): add COBOL pipeline benchmark following PHP benchmark structure

- Generate synthetic COBOL codebases at 100/250/500 file scales
- Each file has 1 PROGRAM-ID, N paragraphs, cross-file CALLs, COPY books
- Measures wall-clock time, peak heap, node/edge counts
- SkipIf(!GITNEXUS_BENCH) — run with GITNEXUS_BENCH=1
- Prints table with scaling ratios and linearity assertions

* fix(bench): remove COPY from paragraphs, add REGISTRY_PRIMARY_COBOL note

- COPY statements belong only in DATA DIVISION (already present there)
- Revert copyLine inside paragraph blocks to idiomatic COBOL
- Add header note about =1 mode producing ~0 node/edge counts

* fix(bench): restore COPY in paragraphs for preprocessing stress

- COPY in paragraph blocks exercises the preprocessor expansion path
  more heavily than DATA DIVISION only placement.

* fix(bench): constant 3 paragraphs per program, add 1000-files scale, relax threshold to 4x

- Fixed paragraphsPerProgram to constant 3 for consistent scaling
- Added 1000-file scale to benchmark
- Raised assertion threshold to 4x to accommodate 100-250 step

* fix: skip standalone providers in scope-resolution phase when registry-primary

scopeResolutionPhase was reading all COBOL files from disk and running
scope-resolution for standalone providers that don't emit graph edges
yet. Added a guard: if provider.languageProvider.parseStrategy ===
'standalone', skip it entirely. Saves 68s at 1000 files in =1 mode.

* fix: remove COBOL isRegistryPrimary gate, suppress standalone IMPORTS double-emission

- Remove the isRegistryPrimary gate in cobolPhase so it runs in both modes,
  keeping cobolPhase as the sole COBOL graph-edge producer.
- Add a guard in runScopeResolution to skip emitImportEdges for standalone
  providers (parseStrategy === 'standalone'), preventing scope-resolution
  from duplicating IMPORTS edges already produced by cobolPhase.
- Scope-resolution still runs for standalone providers (capture extraction,
  model finalization, reference resolution) — only edge emission is skipped.
- Both modes: 60/60 cobol.test.ts, 24/24 cobol-scope.test.ts.

* fix: 4 review fixes — dead code removal, memory cleanup, benchmark comment, standalone-bridge test

1. Remove dead standalone guard in run.ts (phase.ts:164 is canonical).
2. Filter standalone preExtractedByPath entries in phase.ts (memory leak).
3. Update benchmark comment: cobolPhase runs in both modes.
4. Add unit test proving extractParsedFile works for COBOL standalone provider.
   Revert PipelineResult.parsedFiles — not needed with unit test approach.

* perf(cobol): memoize copybook preprocessing; make benchmark measure file-count scaling

The COBOL pipeline benchmark reported superlinear (quadratic) scaling, but the
pipeline itself is O(n) in file count. The superlinearity was a fixture artifact:
every program COPYed all floor(fileCount/5) copybooks in WORKING-STORAGE, so
emitted data-item nodes — and total work — grew O(n^2). Verified empirically:
node count grew ~2x per file-doubling; with constant per-program fan-out it grows
exactly 1x (linear), and 0/3 adversarial audits could refute the O(n) conclusion.

- benchmark: each program now COPYs a constant 3 shared copybooks so the
  benchmark measures true file-count scaling. Add a deterministic node-ratio
  assertion that fails if the O(n^2) copy-all fan-out is reintroduced.
- processor: memoize preprocessed copybook content per processCobol call so each
  copybook is preprocessed once, not once per COPY site
  (O(programs x copybooks) -> O(copybooks)). Safe: REPLACING is applied later by
  the expander on the cached pre-REPLACING content.

Verified: 246 COBOL tests pass; benchmark scales linearly (node ratio 1.0); tsc clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 14:25:19 +01:00
2f15c1ece1 feat(group): add Kotlin Spring WebClient long-form HTTP consumer extraction (#1884)
* feat(group): add Kotlin Spring WebClient long-form HTTP consumer extraction

Follow-up to #1855. Extends `kotlin.ts` with the long-form WebClient
fluent chain that #1855 explicitly deferred:

  webClient.method(HttpMethod.GET).uri("/x").retrieve().awaitBody<T>()

This pattern remains common in Kotlin Spring 4 → 5 migrations and in
codebases that prefer the fluent verb-as-enum style. The short form
(`webClient.get().uri("/x")`) was already supported in #1855.

Approach:
  - Single deeper tree-sitter query (`WEB_CLIENT_LONG_PATTERNS`) that
    matches the full chain structurally — both `.method(HttpMethod.X)`
    and `.uri("...")` in one pattern. Verb is captured as the
    `simple_identifier` of the `HttpMethod.X` field access.
  - Verb is whitelisted to GET/POST/PUT/DELETE/PATCH (consistent with
    the short-form's `WEB_CLIENT_SHORT_TO_HTTP` map).
  - Receiver constraint `(#eq? @obj "webClient")` mirrors the short
    form and Java plugin heuristic.

Out of scope (intentional):
  - Variable-bound verbs: `val verb = HttpMethod.PATCH; webClient.method(verb)...`
    Source-scan can't follow the binding without graph context.
    Pinned by an anti-overreach test.
  - HEAD/OPTIONS/TRACE: not in `WEB_CLIENT_SHORT_TO_HTTP` either —
    keeps polyglot symmetry with java.ts and the short form.

Tests: 4 new cases under `consumer extraction — fetch patterns`,
gated by tree-sitter-kotlin grammar availability.

  positive (3)
   - long form GET
   - long form POST / PUT / DELETE / PATCH (4 verbs in 1 fixture)
   - no double-emit pin (long-form chain produces exactly one
     consumer, not one from each query)
  anti-regression (1)
   - variable-bound verb does NOT match (graph-aware concern)

The previous `'does NOT match Kotlin WebClient long form (deferred
to follow-up)'` test from #1855 is replaced by these — the deferred
state is now resolved.

Reverse-validated: temporarily disabling the long-form emit makes
exactly the 3 positive tests fail; the variable-bound-verb anti-
regression test continues to pass (it pins behavior independent
of the emit being on or off).

Local validation:
  - test/unit/group/http-route-extractor.test.ts: 66/66 ✅
  - test/unit/group: 546/546 ✅
  - npx prettier --check (changed files): clean ✅

* test(group): address Claude review findings F1 and F2 on PR #1884

Two minor follow-ups from the production-readiness review:

F1 — Stale block comment at the top of the Kotlin consumer suite
(was: "Three consumer flavors covered here ... long-form deferred
to a follow-up"). Updated to "Four consumer flavors" and removed
the deferred sentence — the deferral is resolved by this PR. The
kotlin.ts file header was already updated; this brings the test
file comment in sync. Per DoD §2.3 (no stale comments).

F2 — Replaced `expect(wcConsumers.length).toBeGreaterThanOrEqual(4)`
with `expect(wcConsumers).toHaveLength(4)` in the multi-verb test.
The fixture is fully deterministic — exactly 4 long-form calls,
no other consumer types — so an exact count assertion is the right
shape per DoD §2.7 ("use toBe / toEqual for exact expectations").
Added a comment explaining what the assertion catches that the
existing per-verb toBeDefined() checks would miss (accidental 5th
consumer from a duplicate query firing or a regressed receiver
constraint).

F3 (HEAD/OPTIONS/TRACE negative test) is intentionally not added
in this PR — same precedent as #1855 where HEAD/OPTIONS/TRACE on
the short form are also implicitly excluded without a pinning
test. Happy to add one in a separate PR if maintainers want
explicit pinning across both forms.

F4 (CI on pre-merge SHA) is the maintainer's call — the merge from
main is theirs to re-trigger CI on. The merge brings only Java
consumer changes (PR #1872) and Go provider changes (PR #1886),
both in entirely separate files from this PR's Kotlin work.

Local validation:
  - test/unit/group/http-route-extractor.test.ts: 73/73 ✅
    (66 from this PR pre-merge + 7 from PR #1872 merged via main)
  - npx prettier --check (changed files): clean ✅

* refactor(group): hoist Kotlin WebClient long-form verb regex to module scope

Address @magyargergo's review request on PR #1884:

  > Can you please extract the regexp from the for loop? 🙏
    (kotlin.ts:510)

Compiles the verb whitelist `^(GET|POST|PUT|DELETE|PATCH)$` once at
module load instead of every iteration of the long-form scan loop.
Mirrors the placement and JSDoc style of the sibling
`WEB_CLIENT_SHORT_TO_HTTP` constant.

Behavior is unchanged — same verb whitelist, same exclusion of
HEAD/OPTIONS/TRACE for symmetry with the short form. The 4
itKotlinConsumer long-form tests added in this PR continue to
pass, and the variable-bound-verb anti-overreach test continues
to pin the deliberate non-match.

Local validation:
  - test/unit/group/http-route-extractor.test.ts: 77/77 ✅
  - test/unit/group: 557/557 ✅
  - npx prettier --check (changed file): clean ✅

---------

Co-authored-by: henry <zhangwei2017@unipus.cn>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-29 09:27:54 +01:00
JaysonAlbertandgfwangjie 7dae4fcc41 fix(group): attribute Spring interface routes to controllers (#1743)
* fix(group): attribute Spring interface routes to controllers

* test(group): normalize Spring route fixture paths

---------

Co-authored-by: gfwangjie <gfwangjie@gf.com.cn>
2026-05-29 07:35:01 +01:00
evolutionandGergő Magyar d71fd1688b feat(go): add builtInNames set to Go language provider (#1886)
* feat(go): add builtInNames set to Go language provider

Add GO_BUILT_INS (15 functions, 18 types, 3 values) to the Go
LanguageProvider for parity with the other 13 language providers.
The set is converted to an isBuiltInName predicate by defineLanguage()
and consumed by the type-env return-type lookup to short-circuit
lookups for Go built-in symbols.

* feat(go): add Go 1.18+ and 1.21 predeclared identifiers to builtInNames

Add `clear`, `min`, `max` (Go 1.21 builtins), `any`, `comparable`
(Go 1.18 type aliases), and `iota` (predeclared constant) to
GO_BUILT_INS for complete coverage of the Go specification.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-29 06:46:02 +01:00
MyShining 7b38b8aae2 feat(java): add HTTP consumer contract extraction (#1872) 2026-05-29 06:05:40 +01:00
b565c7c990 feat(ingestion): resolve FastAPI include_router(prefix=...) cross-file routes (#1877)
* feat(ingestion): resolve FastAPI include_router(prefix=...) cross-file routes

FastAPI sub-route files declare paths via @router.<verb> while the entry
file mounts the router with app.include_router(<router>, prefix='/x').
Previously both the ingestion-layer Route graph nodes and the group-layer
ExtractedContract URLs lost the cross-file prefix, breaking provider <->
consumer matching.

Ingestion layer:
  - parse-worker emits routerIncludes / routerImports + decoratorReceiver
  - parsing-processor / parse-impl thread the new fields and aggregate
    prefixesByModule across chunks; decorator routes whose receiver is
    'router' are duplicated once per matching prefix
  - routes.ts joins prefix via normalizeExtractedRoutePath

Group layer:
  - HttpLanguagePlugin gains an optional prepareRepo() pre-pass and a
    repoContext arg to scan(); python.ts builds prefixesByModule and
    falls back to the bare path when no entry matches
  - http-route-extractor caches one repoContext per plugin

Tests:
  - 3 new http-route-extractor cases (attr / named-import / no-prefix)
  - ParseWorkerResult literals in 3 test files updated to the new shape

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(ingestion,group): address PR #1877 review — relative imports, cross-package collisions, host names, ingestion tests

Follow-ups to the FastAPI `include_router(prefix=...)` cross-file fix
based on PR #1877's automated production-readiness review. Three
correctness gaps and one test coverage gap addressed:

1. Relative-import support in the worker regex (FINDING 2)
   `FROM_IMPORT_ROUTER_RE` now accepts module paths starting with a
   `.` (e.g. `from .calls import router as calls_router`). The
   previous `[A-Za-z_][\w.]*` rejected leading dots and silently
   dropped every relative-import Shape-B include — a real pattern
   from the PR description's own motivating example. The matching
   helpers now strip leading dots before keying so absolute and
   relative imports collapse to the same module key.

2. Cross-package same-name module collisions (FINDING 3)
   Two-tier module keying replaces the previous basename-only key:
     • short key — `users`            (file basename without `.py`)
     • long  key — `api/users`        (parent dir + stem)
   `prefixesByLongKey` is consulted first and only falls back to
   `prefixesByShortKey` when no long-key match is available. Both
   the ingestion pipeline (parse-impl.ts) and the group extractor
   (http-patterns/python.ts) carry the same scheme so the graph
   nodes and HTTP contracts agree on which prefix applies.

   New protocol field `ExtractedRouterModuleAlias` (parse-worker →
   parsing-processor → parse-impl) lets Shape-A
   `<host>.include_router(<mod>.router, prefix='/x')` calls promote
   to a long key when the same file imports `<mod>` via
   `from <pkg> import <mod>`. Without this, `api/users.py` and
   `admin/users.py` collided on the basename `users` and the admin
   file's routes inherited the `/users` prefix that was only meant
   for `api/users.py`.

3. Non-`app` host variable names (FINDING 4)
   The group-layer `INCLUDE_ROUTER_*_PATTERNS` queries pinned the
   host identifier to the literal `"app"` and dropped every
   `application = FastAPI()` / `api = FastAPI()` pattern — the
   constraint was redundant given that the call shape
   (`include_router` invoked with a router argument and a
   `prefix=` keyword) is already specific enough. The pin is
   removed; the ingestion regex was already unrestricted.

4. Ingestion-layer regression tests (FINDING 1)
   The previous PR added group-layer tests
   (`http-route-extractor.test.ts`) but zero in-tree tests for the
   ingestion path. Two new suites pin the
   worker → parse-impl → routes flow:

   - `test/unit/fastapi-router-bindings.test.ts` (23 cases):
     `extractFastAPIRouterBindings()` is split into a stand-alone
     module so it can be unit-tested without booting a worker
     thread, then pinned for regex shape, two-tier key emission,
     relative-import support, and negative cases.
   - `test/integration/fastapi-prefix-pipeline.test.ts` (5 cases)
     plus `test/fixtures/fastapi-prefix-app/` — runs the full
     `runPipelineFromRepo()` against a realistic multi-package
     fixture (containing both `api/users.py` and `admin/users.py`)
     and inspects the resulting `Route` graph nodes for cross-file
     prefix joining and absence of cross-package bleed.

Verification

  - `npx tsc --noEmit`: pass
  - PR-touched test suites (6 files / 117 cases): all green
  - `npx prettier --check`: pass on touched files
  - `npx eslint`: 0 errors on touched files

Cache / compatibility

  The new `routerModuleAliases?` field on `ParseWorkerResult` and
  `routerModuleAliases` on `WorkerExtractedData` are optional /
  guarded with `?? []`, so historical parse-cache entries continue
  to load without forced re-scan.

Refs PR #1877.

* refactor(ingestion): move fastapi-router-bindings out of workers/ — pure module, not a worker

Addresses @magyargergo's `CHANGES_REQUESTED` review on PR #1877:

> Sorry I just found that we are introducing a new worker in the PR.

`gitnexus/src/core/ingestion/workers/fastapi-router-bindings.ts` was a
**pure-function module** — it never imported `worker_threads` or
`parentPort`, never spawned a worker, and was never registered as a
worker entry. It was placed in `workers/` purely because it was split
out of `workers/parse-worker.ts` to make its functions unit-testable
without booting a worker thread (parse-worker is itself the worker
entry and cannot be loaded from the main thread).

To remove the misleading directory placement:

  • The implementation moves to
    `gitnexus/src/core/ingestion/route-extractors/fastapi-router-bindings.ts`,
    alongside the other framework-specific route extractors (`expo`,
    `nextjs`, `php`, `laravel`, `middleware`, `response-shapes`).
  • `workers/parse-worker.ts` keeps a thin re-export so the worker
    entry can keep using `extractFastAPIRouterBindings` directly. The
    re-export now carries an explicit comment stating that the imported
    file is **not** a worker and that the `workers/` directory
    deliberately hosts only true worker entries (`parse-worker.ts`,
    `worker-pool.ts`, `quarantine.ts`).
  • The new file's leading docstring opens with "NOT A WORKER" and
    explains why it exists where it does.
  • The unit test (`test/unit/fastapi-router-bindings.test.ts`) is
    updated to import from the new path.

No behaviour change. The function body, signatures, and exported types
are identical.

Verification

  • `npx tsc --noEmit`: pass
  • `npx tsc` (dist rebuild): pass
  • `test/unit/fastapi-router-bindings.test.ts` (23 cases): all green
  • `test/integration/fastapi-prefix-pipeline.test.ts` (5 cases): all green
  • `test/unit/group/http-route-extractor.test.ts` (63 cases): all green
  • `npx prettier --check` on touched files: pass
  • `npx eslint` on touched files: 0 errors

Refs PR #1877.

* refactor(ingestion): drop parse-worker re-exports; consumers import router types directly from route-extractors

Addresses @magyargergo's two remaining review comments on PR #1877:

1. **`gitnexus/src/core/ingestion/workers/parse-worker.ts:247`** —
   "Can you please remove them and update the call sites?"

   The `export type { ExtractedRouterInclude, ExtractedRouterImport,
   ExtractedRouterModuleAlias } from '../route-extractors/...'` block
   in parse-worker.ts is gone. The remaining `import type {…}` is
   purely local — used only to type the corresponding fields on
   `ParseWorkerResult` below — and the leading comment now says so
   explicitly ("this file does NOT re-export them"). The
   `extractFastAPIRouterBindings` symbol is also no longer re-exported
   from parse-worker.ts; it's still imported here so the worker entry
   can call it per file, but downstream consumers must reach it via
   `route-extractors/fastapi-router-bindings` directly.

   Call sites updated:
     - `gitnexus/src/core/ingestion/parsing-processor.ts`
     - `gitnexus/src/core/ingestion/pipeline-phases/parse-impl.ts`

   Both files now `import type { ExtractedRouterInclude,
   ExtractedRouterImport, ExtractedRouterModuleAlias }` directly from
   `route-extractors/fastapi-router-bindings.js`. The worker types
   they still need (`ParseWorkerResult`, `ExtractedToolDef`, etc.)
   keep coming from `workers/parse-worker.js`.

   The unit + integration tests already imported from the new path,
   so no test changes were required.

2. **`gitnexus/src/core/ingestion/parsing-processor.ts:168`** —
   suggested simplification:

       for (const item of result.routerIncludes ?? []) allRouterIncludes.push(item);
       for (const item of result.routerImports ?? []) allRouterImports.push(item);
       for (const item of result.routerModuleAliases ?? []) allRouterModuleAliases.push(item);

   Applied verbatim. Replaces the previous `if (result.…) for …`
   guards. The cache-compat semantics are unchanged — historical
   parse-cache entries that lack these fields still load cleanly,
   the new form just spells the fallback inline.

No behavior change, no tests touched, no public API change.

Verification

  • `npx tsc --noEmit`: pass
  • `npx tsc` (dist rebuild): pass
  • PR-touched test suites (6 files / 117 cases): all green
  • `npx prettier --check` on touched files: pass
  • `npx eslint` on touched files: 0 errors

Refs PR #1877.

* refactor(ingestion): hoist fastapi-router-bindings type imports to top of parse-worker.ts

Move the `import type { ExtractedRouterInclude, ExtractedRouterImport,
ExtractedRouterModuleAlias }` block to the top of the file with the
other type imports, and drop the comment that previously sat next to
ExtractedDecoratorRoute.

---------

Co-authored-by: henry <zhangwei2017@unipus.cn>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-28 19:04:19 +01:00
azizur100389andGergő Magyar 97c1f85e87 refactor(cpp): Use function-type ADL entities (#1822)
* fix(cpp): use function-type ADL entities

* test(hooks): stabilize concurrency burst reporting

* Fix C++ return type capture subtag handling

* Harden C++ function-type ADL extraction

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-28 17:18:19 +01:00
11fc43b425 feat(impact): per-symbol processes field on byDepth items (#1867)
* feat(impact): per-symbol processes field on byDepth items

Today `impact` returns aggregated `affected_processes` at the top level
but the per-symbol `byDepth` items don't say which processes each caller
participates in. Consumers planning a deploy want to know if a given
caller is hit by a daily cron, a webhook, or a user-facing route - each
is a different deploy-risk profile - and that information requires a
follow-up cypher query per symbol today.

This change attaches `processes: [...]` to every `byDepth[depth][i]`
item, listing the processes that symbol participates in:

  byDepth: {
    "1": [
      {
        depth: 1,
        id: "Function:src/foo.ts:doStuff",
        name: "doStuff",
        ...
        processes: [
          { id: "proc:cron_daily", label: "Daily cron",
            processType: "cron", step: 12 }
        ]
      }
    ]
  }

The list is empty for symbols not in any process. Additive change, no
breaking modifications to existing fields.

Implementation:
- A second chunked Cypher pass runs after the existing per-process
  aggregation pass, returning per-(symbol, process) rows. Same chunk
  size and MAX_CHUNKS as the aggregation pass, so worst-case adds 10
  extra round-trips bounded by the same env var.
- The enrichment pass is skipped entirely when `affectedProcesses.length
  === 0` (nothing to enrich) or `summaryOnly === true` (byDepth not
  returned anyway).
- The aggregation query is unchanged - the new query has a distinct
  RETURN shape (`RETURN s.id AS sid, ...`) so an existing unit test that
  counts STEP_IN_PROCESS chunks was narrowed to match only the
  aggregation pattern.

Tests:
- New: byDepth items always have a `processes` field (default empty
  when no STEP_IN_PROCESS edges exist).
- New: when STEP_IN_PROCESS rows exist, the matching byDepth item
  carries the right `{id, label, processType, step}` entry.
- Updated: impact-batching-grouping test mock narrowed to count only
  aggregation chunks (the new per-symbol pass is covered separately).

* style: apply prettier to gitnexus/src/mcp/local/local-backend.ts

Pure line-wrap fix flagged by quality / format CI on PR #1867. Zero
semantic change: prettier broke a chained .slice().map() across three
lines instead of one. No test changes, no logic changes.

* fix(impact): address PR review findings on per-symbol process enrichment

- byDepth.processes doc now states each item carries processes (Finding 1)
- move per-symbol STEP_IN_PROCESS enrichment post-pagination so symbols
  beyond the pre-pagination cap no longer get false-empty processes:[]
  (Finding 2); hoist CHUNK_SIZE/MAX_CHUNKS to function scope so the
  post-pagination pass can reference them
- dedup per-symbol query with DISTINCT + MIN(r.step) per (symbol,process)
  pair (Finding 3)
- suppress the per-symbol pass under summaryOnly, incl. impactByUid group
  fan-out, plus a test asserting the query never fires (Findings 4, 6)

* fix(impact): address second-round review findings A-E

Finding A (blocker): impactByUid passed summaryOnly:true, which drops the
entire byDepth field. cross-impact.ts reads fan.byDepth to build the group
by_depth output, so cross-repo by_depth was always {}. Replace with a new
skipPerSymbolEnrichment option on _runImpactBFS that suppresses only the
per-symbol STEP_IN_PROCESS pass while preserving byDepth.

Finding B+D (blocker): rewrite the byDepth.processes tool description. Drop
the stale "enrichment cap" wording (no longer true post-pagination), document
the {id,label,processType,step} entry shape, and tell agents to cross-check
affected_processes when partial:true.

Finding C: bound the post-pagination per-symbol enrichment loop to
MAX_CHUNKS*CHUNK_SIZE page IDs and surface partial:true when capped, so a
large page cannot trigger unbounded DB round-trips (DoD 2.6).

Finding E: add a test exercising the real impactByUid -> _runImpactBFS path
asserting byDepth survives and the per-symbol query never fires.

---------

Co-authored-by: scotjelinski <58397194+scotjelinski@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-28 16:15:37 +01:00
50715e3894 chore(deps)(deps-dev): bump @playwright/test in /gitnexus-web (#1860)
Bumps [@playwright/test](https://github.com/microsoft/playwright) from 1.58.2 to 1.60.0.
- [Release notes](https://github.com/microsoft/playwright/releases)
- [Commits](https://github.com/microsoft/playwright/compare/v1.58.2...v1.60.0)

---
updated-dependencies:
- dependency-name: "@playwright/test"
  dependency-version: 1.60.0
  dependency-type: direct:development
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-28 07:51:05 +01:00
dependabot[bot]andGergő Magyar ca95df6316 chore(deps): bump github/codeql-action from 4.35.4 to 4.35.5 (#1866)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4.35.4 to 4.35.5.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/68bde559dea0fdcac2102bfdf6230c5f70eb485e...9e0d7b8d25671d64c341c19c0152d693099fb5ba)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.35.5
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-28 06:46:18 +01:00
dependabot[bot]andGergő Magyar 9d609cc386 chore(deps)(deps): bump axios from 1.16.0 to 1.16.1 in /gitnexus-web (#1864)
Bumps [axios](https://github.com/axios/axios) from 1.16.0 to 1.16.1.
- [Release notes](https://github.com/axios/axios/releases)
- [Changelog](https://github.com/axios/axios/blob/v1.x/CHANGELOG.md)
- [Commits](https://github.com/axios/axios/compare/v1.16.0...v1.16.1)

---
updated-dependencies:
- dependency-name: axios
  dependency-version: 1.16.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-28 06:45:30 +01:00
dependabot[bot]andGergő Magyar 128a199970 chore(deps)(deps-dev): bump @types/node in /gitnexus-web (#1863)
Bumps [@types/node](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/node) from 25.6.0 to 25.9.1.
- [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases)
- [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/node)

---
updated-dependencies:
- dependency-name: "@types/node"
  dependency-version: 25.9.1
  dependency-type: direct:development
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-28 06:45:18 +01:00
dependabot[bot]andGergő Magyar 76409783aa chore(deps)(deps): bump @langchain/langgraph in /gitnexus-web (#1861)
Bumps [@langchain/langgraph](https://github.com/langchain-ai/langgraphjs/tree/HEAD/libs/langgraph-core) from 1.2.9 to 1.3.2.
- [Release notes](https://github.com/langchain-ai/langgraphjs/releases)
- [Changelog](https://github.com/langchain-ai/langgraphjs/blob/main/libs/langgraph-core/CHANGELOG.md)
- [Commits](https://github.com/langchain-ai/langgraphjs/commits/@langchain/langgraph@1.3.2/libs/langgraph-core)

---
updated-dependencies:
- dependency-name: "@langchain/langgraph"
  dependency-version: 1.3.2
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-28 06:45:01 +01:00
Gergő Magyar 99168be773 feat(ingestion): trace indirect call patterns — FastAPI Depends() and frontend HTTP consumers (#1852) 2026-05-28 05:33:30 +01:00
henry201605andhenry d9d6318b64 feat(group): add Kotlin Spring HTTP consumer extraction (#1855)
* feat(group): add Kotlin Spring HTTP consumer extraction

Follow-up to #1849 (Kotlin providers). Extends `http-patterns/kotlin.ts`
with three call-site patterns common in Kotlin Spring projects:

  - RestTemplate: `restTemplate.getForObject("/x", ...)` and the
    full verb family (getForObject/getForEntity → GET,
    postForObject/postForEntity → POST, put → PUT, delete → DELETE,
    patchForObject → PATCH). Mirrors the Java plugin's
    `REST_TEMPLATE_TO_HTTP` map so polyglot repos coalesce on a
    single contract id.

  - WebClient short form: `webClient.get().uri("/x")` and the
    `.post()` / `.put()` / `.delete()` / `.patch()` siblings. The
    chain parses as two nested `call_expression` nodes; the query
    anchors on the outer `.uri(...)` and walks one level inward
    to constrain the verb.

  - OkHttp: `Request.Builder().url("/x")`. Kotlin parses
    `Request.Builder()` as a `call_expression` whose callee is a
    `navigation_expression` (not Java's `object_creation_expression`),
    so the query shape differs from `java.ts` but the receiver/method
    constraints (`Request` / `Builder` / `url`) and emitted
    contract format match.

Out of scope: `webClient.method(HttpMethod.X).uri("/y")` long form.
The verb sits on a sibling `call_expression` two hops away, so it
needs a walk-up helper rather than a flat tree-sitter query. A
dedicated anti-overreach test pins the current behavior so a future
short-form change can't accidentally start matching the long form.

Receiver name constraints (`#eq? @obj "restTemplate"`,
`#eq? @cls "Request"`) match the Java plugin's heuristic — a project
that aliases the receiver under a different name won't be picked up.
This trade-off keeps false-positive rates low and is documented in
the file header.

Tests: 5 new cases under `consumer extraction — fetch patterns`,
gated by tree-sitter-kotlin grammar availability.

  positive (3)
   - RestTemplate verbs (5 calls × 5 verbs)
   - WebClient short-form verbs (5 calls × 5 verbs)
   - OkHttp Request.Builder().url("/x")
  anti-regression (2)
   - WebClient long form `.method(HttpMethod.X)` produces no
     consumer (deferred-feature pin)
   - non-restTemplate receiver does not match (receiver-name pin)

Reverse-validated: removing the `(#eq? @obj "restTemplate")`
constraint causes the receiver-name anti-regression test to fail.

Local validation:
  - test/unit/group/http-route-extractor.test.ts: 59/59 ✅
  - test/unit/group: 539/539 ✅
  - npm run format:check: clean ✅

* test(group): pin Kotlin OkHttp POST-chain heuristic-default GET behavior

Address Claude review on PR #1855 (Finding 1).

The OkHttp query in `kotlin.ts:OK_HTTP_PATTERNS` matches the
`.url("/x")` sub-expression of a builder chain, but the verb is
encoded on a separate sibling call (`.post(body)` / `.delete()` /
...). The query intentionally does not walk the chain to recover
the verb — it emits `method: 'GET'` for every match, mirroring the
Java plugin's `OK_HTTP_PATTERNS` (java.ts).

Concretely: `Request.Builder().url("/x").post(body).build()` becomes
`http::GET::/x`, not `http::POST::/x`. This is an already-accepted
Java parity heuristic, but it was untested on the Kotlin side.

This commit:
  - Adds an anti-overreach test pinning the current behavior:
      * exactly one consumer is emitted with method=GET
      * no second http::POST::/x consumer appears
  - Documents the limitation in kotlin.ts as a "Known limitation"
    block tied to the test, so a future verb-walk implementation
    has to update the comment in lockstep with the assertion.

Rationale for not implementing verb-walk in this PR:
  - Verb-walk requires walking sibling call_expression nodes (the
    `.post(body)` chain), which is the same shape as the
    deferred WebClient long-form work
  - Java has the same limitation in production today; fixing only
    Kotlin would create polyglot drift
  - A coordinated future PR can add verb-walk to both plugins at
    once and update both comments + the pin tests together

Finding 2 (silent test-skip when tree-sitter-kotlin grammar is
unavailable) is intentionally NOT addressed here — same gating
pattern was accepted in #1849 for Provider tests, and a coordinated
follow-up should add a CI sentinel covering both Provider and
Consumer suites in one place.

Local validation:
  - test/unit/group/http-route-extractor.test.ts: 60/60 ✅
  - test/unit/group: 540/540 ✅
  - npm run format:check: clean ✅

---------

Co-authored-by: henry <zhangwei2017@unipus.cn>
2026-05-27 21:32:24 +01:00
henry201605andhenry 46eb0ebf56 feat(group): add Kotlin Spring HTTP route extraction (named + positional) (#1849)
* feat(group): add Kotlin Spring HTTP route extraction (named + positional)

Mirror the Java Spring named-argument fix for Kotlin Spring Boot
controllers. Adds a new `http-patterns/kotlin.ts` plugin behind the
optional `tree-sitter-kotlin` grammar, registered for `.kt`/`.kts`.

Both annotation forms produce providers:
  @RequestMapping("/api")          / @GetMapping("/users")
  @RequestMapping(path = "/api")   / @GetMapping(value = "/users")
  @RequestMapping(value = "/api")  / @GetMapping(path = "/users")

The Kotlin AST (fwcd/tree-sitter-kotlin) shares one node type
(`value_argument`) for positional and named forms, so the queries
are split:
  - positional: anchors `string_literal` as the first named child
    of `value_argument` via the immediate-child anchor `.`
  - named: explicitly captures `simple_identifier` and constrains
    it to `^(path|value)$` via `#match?`, mirroring the same
    safety bar enforced by `http-patterns/java.ts` and
    `topic-patterns/java.ts`. Without this constraint the query
    would also capture non-route attributes like `produces`,
    `consumes`, `headers`, `name`, `params`.

`tree-sitter-kotlin` is an optionalDependency (parser-loader.ts,
parse-worker.ts pattern). When the native binding is unavailable
the plugin exports `null` and `index.ts` skips registering
`.kt`/`.kts` so the orchestrator stays healthy.

Scope: providers only. Consumer detection (RestTemplate, WebClient,
OkHttp) on Kotlin call-site ASTs differs enough from Java's
`method_invocation` shape to warrant a separate, focused PR.

Tests: 11 new cases under `provider extraction — source-scan
fallback (Strategy B)`, gated by the kotlin grammar availability.

  positive (8)
   - class @RequestMapping("/api/v1") (positional)
   - class @RequestMapping(path = "/api/v2")
   - class @RequestMapping(value = "/orders")
   - method @GetMapping(value = "/users")
   - method @GetMapping(path = "/users")
   - method @PostMapping(path = "/users")
   - mixed: class named-arg + method positional
   - mixed: class positional + method named-arg
  anti-regression (3)
   - @GetMapping(produces = "application/json") emits no provider
   - @GetMapping(name = "x", value = "/users") emits exactly one provider
   - @RequestMapping(path = "/api", name = "myApi") prefix stays /api

Reverse-validated: removing the `(#match? @key "^(path|value)$")`
constraint causes precisely the 3 anti-regression tests to fail.

Local validation:
  - test/unit/group/http-route-extractor.test.ts: 54/54
  - test/unit/group: 534/534
  - npx tsc --noEmit: clean (modulo the pre-existing TS2339 in
    user-defined-conversions.ts merged from main, unrelated)

* style(test): apply prettier line wrapping to long itKotlin titles

---------

Co-authored-by: henry <zhangwei2017@unipus.cn>
2026-05-27 09:33:52 +01:00
eeea46466b fix(group): handle named annotation args in Java Spring route extraction (#1834)
* fix(group): handle named annotation args in Java Spring route extraction

The Java HTTP plugin only matched positional `@RequestMapping("/path")`
syntax for class-level prefixes and method-level routes. Named argument
forms (`path = "/path"` and `value = "/path"`) produce an
`element_value_pair` AST node that the tree-sitter queries did not cover,
causing the class prefix to be lost and named-arg method routes to be
missed entirely during cross-repo contract extraction.

Add a second pattern to both SPRING_CLASS_PREFIX_PATTERNS and
SPRING_METHOD_ROUTE_PATTERNS matching the element_value_pair structure.

* fix(group): constrain Spring named-arg query to path/value keys + add regression tests

Address Claude review on PR #1834. The named-argument patterns added
in 8b6fa6e used `value: (string_literal)` (a tree-sitter field
selector for the right-hand side of element_value_pair), which matched
ANY annotation member with a string value — not just `path`/`value`.

Concrete fallout (without this fix):
  @GetMapping(produces = "application/json") → bogus http::GET::/application/json
  @GetMapping(name = "listUsers", value = "/users") → extra http::GET::/listUsers
  @RequestMapping(headers = "X-Foo=bar", path = "/api") → class prefix
    could be set to "X-Foo=bar" because prefixByClassId.set runs per
    match in document order, so the LAST element_value_pair wins.

The sibling topic-patterns/java.ts already demonstrates the correct
shape: constrain the `key:` field to the route member names.

This commit:
  - Adds `key: (identifier) @key (#match? @key "^(path|value)$")` to
    both SPRING_CLASS_PREFIX_PATTERNS and SPRING_METHOD_ROUTE_PATTERNS
    named-arg queries.
  - Adds 9 regression tests under
    `provider extraction — source-scan fallback (Strategy B)`:
      * @RequestMapping(path = "/api/v3") class prefix
      * @RequestMapping(value = "/orders") class prefix
      * @GetMapping(value = "/users") method route
      * @PostMapping(path = "/users") method route
      * mixed: class named-arg + method positional
      * mixed: class positional + method named-arg
      * @GetMapping(produces = "application/json") → no provider emitted
      * @GetMapping(name = "listUsers", value = "/users") → exactly one
        provider with path "/users", no /listUsers route
      * @RequestMapping(path = "/api", name = "myApi") → prefix is /api,
        not myApi (verifies the class-prefix overwrite scenario)

Tests: 42/42 pass in http-route-extractor.test.ts;
       522/522 pass under test/unit/group;
       npx tsc --noEmit clean.

* test(group): add @GetMapping(path = ...) case to match review checklist verbatim

Claude review on PR #1834 explicitly asked for the method-level
`@GetMapping(path = "/users")` case. The previous commit covered it
indirectly by exercising path= on @PostMapping (the Spring method
annotations share the same query, so any verb proves the path= field
is matched). Add a dedicated GET+path= test so the reviewer's
checklist is satisfied 1:1, and keep the POST+path= case as a bonus
verb-coverage test.

Tests: 43/43 pass in http-route-extractor.test.ts.

---------

Co-authored-by: henry <zhangwei2017@unipus.cn>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-27 07:35:46 +01:00
dependabot[bot]andGergő Magyar ca3e1755c2 chore(deps)(deps): bump lru-cache from 11.4.0 to 11.5.0 in /gitnexus (#1844)
Bumps [lru-cache](https://github.com/isaacs/node-lru-cache) from 11.4.0 to 11.5.0.
- [Changelog](https://github.com/isaacs/node-lru-cache/blob/main/CHANGELOG.md)
- [Commits](https://github.com/isaacs/node-lru-cache/compare/v11.4.0...v11.5.0)

---
updated-dependencies:
- dependency-name: lru-cache
  dependency-version: 11.5.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-27 06:45:18 +01:00
dependabot[bot]andGergő Magyar 6acdc49f06 chore(deps)(deps-dev): bump @types/node in /gitnexus (#1845)
Bumps [@types/node](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/node) from 25.9.0 to 25.9.1.
- [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases)
- [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/node)

---
updated-dependencies:
- dependency-name: "@types/node"
  dependency-version: 25.9.1
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-27 06:44:41 +01:00
azizur100389 b1445daf04 feat(cpp): rank user-defined conversions (#1829) 2026-05-27 06:36:23 +01:00
178 changed files with 14350 additions and 677 deletions
+43
View File
@@ -0,0 +1,43 @@
# GitNexus PR Reviewer Swarm — Claude Code adapter
This is the **Claude Code** entrypoint for the cross-CLI GitNexus PR reviewer swarm. The
review logic itself is CLI-neutral and lives in **[`pr-swarm-review/`](../pr-swarm-review/README.md)**
— that README is the canonical guide and covers every CLI (Claude Code, Gemini, Copilot,
Cursor, Codex, and any AGENTS.md-aware agent).
## Invocation (Claude Code)
```
/gitnexus-pr-swarm-review <PR URL or PR number>
```
Runs in **Swarm mode**: the coordinator skill dispatches the seven `gitnexus-*` subagents in
parallel (lanes 1–2 first, 3–6 in parallel, lane 7 last as a hard gate).
## Files in this adapter
| File | Role |
|------|------|
| `.claude/skills/gitnexus-pr-swarm-review/SKILL.md` | Coordinator — runs Swarm mode per `pr-swarm-review/orchestration.md` |
| `.claude/agents/gitnexus-*.md` | Seven thin subagent wrappers; each reads its canonical persona in `pr-swarm-review/personas/` |
Each subagent keeps valid Claude Code frontmatter (model, tools, etc.); the mechanical
verifier lanes (`test-ci-verifier`, `branch-hygiene-reviewer`) run on Haiku, the analytical
lanes on Sonnet.
## Key properties
- **Read-only.** Tools limited to Read/Grep/Glob/Bash, and every persona enforces an
explicit permitted/prohibited Bash list. No agent edits files, commits, or posts.
- **Evidence-grounded**; **missing visibility becomes verification work**; **manually invoked.**
## Editing
Edit review behavior in the canonical files under `pr-swarm-review/` (orchestration +
personas), **not** in these wrappers. After adding or editing files in `.claude/agents/`,
restart Claude Code so it reloads the agent definitions.
## Relationship to `/gitnexus-pr-review`
Coexists with the single-agent `/gitnexus-pr-review` skill (a linear checklist using GitNexus
MCP tools). This swarm is the multi-persona deep production-readiness review.
@@ -0,0 +1,24 @@
---
name: gitnexus-branch-hygiene-reviewer
description: "GitNexus branch hygiene and mergeability reviewer. Use to classify merge state, conflicts, stale branches, merge-from-main commits, unrelated churn, mixed domains, and whether rebase or split is required."
tools:
- Read
- Grep
- Glob
- Bash
model: claude-haiku-4-5-20251001
maxTurns: 30
---
# GitNexus Branch Hygiene & Mergeability Reviewer
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
**`pr-swarm-review/personas/02-branch-hygiene-reviewer.md`**
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
## Rules (always enforced)
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
@@ -0,0 +1,24 @@
---
name: gitnexus-docs-dod-reviewer
description: "GitNexus docs and Definition-of-Done reviewer. Use to translate repo guidance, linked issues, changed domains, docs requirements, release notes, and acceptance criteria into a PR-specific DoD."
tools:
- Read
- Grep
- Glob
- Bash
model: claude-sonnet-4-6
maxTurns: 30
---
# GitNexus Docs & Definition-of-Done Reviewer
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
**`pr-swarm-review/personas/06-docs-dod-reviewer.md`**
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
## Rules (always enforced)
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
@@ -0,0 +1,24 @@
---
name: gitnexus-pr-facts-historian
description: "GitNexus PR facts and repository-history investigator. Use to gather PR identity, visible GitHub state, changed files, commits, linked issues, related PRs, historical fixes, regressions, stale follow-ups, and missing visibility."
tools:
- Read
- Grep
- Glob
- Bash
model: claude-sonnet-4-6
maxTurns: 40
---
# GitNexus PR Facts & Repository-History Investigator
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
**`pr-swarm-review/personas/01-pr-facts-historian.md`**
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
## Rules (always enforced)
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
+24
View File
@@ -0,0 +1,24 @@
---
name: gitnexus-risk-architect
description: "GitNexus production-risk reviewer. Use for risk-model-first review of changed files, runtime behavior, multi-domain changes, user impact, failure modes, compatibility, and merge-blocking risk."
tools:
- Read
- Grep
- Glob
- Bash
model: claude-sonnet-4-6
maxTurns: 40
---
# GitNexus Production-Risk Architect
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
**`pr-swarm-review/personas/03-risk-architect.md`**
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
## Rules (always enforced)
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
@@ -0,0 +1,24 @@
---
name: gitnexus-security-boundary-reviewer
description: "GitNexus security and trust-boundary reviewer. Use for auth, permissions, secrets, injection, unsafe parsing, external input handling, hidden Unicode, YAML/Docker/workflow risks, and suspicious non-ASCII hygiene."
tools:
- Read
- Grep
- Glob
- Bash
model: claude-sonnet-4-6
maxTurns: 35
---
# GitNexus Security & Trust-Boundary Reviewer
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
**`pr-swarm-review/personas/05-security-boundary-reviewer.md`**
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
## Rules (always enforced)
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
@@ -0,0 +1,24 @@
---
name: gitnexus-synthesis-critic
description: "GitNexus final review synthesis critic. Use to check whether the final PR review is evidence-grounded, risk-prioritized, GitNexus-specific, non-generic, and follows required verdict rules."
tools:
- Read
- Grep
- Glob
- Bash
model: claude-sonnet-4-6
maxTurns: 25
---
# GitNexus Final-Review Synthesis Critic
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
**`pr-swarm-review/personas/07-synthesis-critic.md`**
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
## Rules (always enforced)
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
@@ -0,0 +1,24 @@
---
name: gitnexus-test-ci-verifier
description: "GitNexus test and CI reviewer. Use to verify whether changed behavior is covered by targeted tests, whether CI actually runs those tests, and whether workflow changes weaken validation."
tools:
- Read
- Grep
- Glob
- Bash
model: claude-haiku-4-5-20251001
maxTurns: 35
---
# GitNexus Test & CI Verifier
Your complete operating spec — role, what to inspect, classifications, and the required output sections — lives in the canonical, CLI-neutral persona file:
**`pr-swarm-review/personas/04-test-ci-verifier.md`**
Read that file now with the Read tool and follow it exactly. It is the single source of truth shared across all AI CLIs; this subagent only adapts it to Claude Code. The orchestration contract (lane order, Swarm vs Solo execution, output structure) is in `pr-swarm-review/orchestration.md`.
## Rules (always enforced)
- **Do not edit files.** You are read-only.
- **Bash is read-only.** Permitted: `git log`, `git diff`, `git show`, `git grep`, `git ls-files`, `gh pr view`, `gh pr diff`, `gh pr checks`, `gh issue view`, and inspection tools (`grep`, `cat`, `find`, `ls`). Prohibited: any command that writes files, modifies git state (`git commit`, `git add`, `git checkout -- <path>`), posts to GitHub (`gh pr comment`, `gh pr review`, `gh issue comment`), installs packages, or runs arbitrary scripts.
@@ -0,0 +1,31 @@
---
name: gitnexus-pr-swarm-review
description: "Run a GitNexus production-readiness pull request review using a coordinated reviewer swarm."
---
# GitNexus PR Swarm Review (Claude Code adapter)
Use this skill to review a GitNexus pull request and produce a production-readiness review.
```
/gitnexus-pr-swarm-review <PR URL or PR number>
```
You are the **swarm coordinator**. The full review contract — lanes, dependencies,
classifications, output structure, finding format, hidden-Unicode checks, and behavior
rules — is the canonical, CLI-neutral spec:
**`pr-swarm-review/orchestration.md`** — read it now and follow it.
This adapter only pins the Claude Code specifics:
- **Run in Swarm mode.** Dispatch each lane as its own subagent via the Agent tool. The
seven subagents are the project agents named `gitnexus-*` (one per persona); each reads
its canonical persona under `pr-swarm-review/personas/`. Run lanes 1–2 first, lanes 3–6
in parallel after, and lane 7 last on the draft.
- **Lane 7 is a hard gate.** Do not emit the final review while the synthesis critic's
"Required corrections before posting" section is non-empty — revise and re-run it.
- Stay **read-only**: investigate and report; never edit, commit, or post.
Do not flatten the review into a generic checklist; delegate to the subagents and
synthesize per `orchestration.md`.
@@ -0,0 +1,17 @@
# GitNexus PR Swarm Review
You are the GitNexus PR review coordinator. Review the pull request named after this command
(a PR URL or number for `https://github.com/abhigyanpatwari/GitNexus`). If none was given,
ask for one.
Read `pr-swarm-review/orchestration.md` in this repository and follow it exactly — it is the
canonical, CLI-neutral review contract (lanes, classifications, output structure, finding
format, hidden-Unicode checks, behavior rules).
Run in **Solo mode**: you are a single agent, so perform all seven lanes yourself in
dependency order, adopting each persona in `pr-swarm-review/personas/0N-*.md` in turn
(lanes 1–2 first, then 3–6, then lane 7). Keep every lane's findings in context. Lane 7
(synthesis critic) is a hard gate: do not emit the final review until its "Required
corrections before posting" section is empty.
Stay strictly read-only: investigate and report; never edit files, commit, or post to GitHub.
@@ -0,0 +1,19 @@
description = "GitNexus production-readiness PR swarm review (Solo mode)"
prompt = """
You are the GitNexus PR review coordinator. Review this pull request: {{args}}
(a PR URL or number for https://github.com/abhigyanpatwari/GitNexus). If no target was
given, ask for one.
Read `pr-swarm-review/orchestration.md` in this repository and follow it exactly. It is the
canonical, CLI-neutral review contract (lanes, classifications, output structure, finding
format, hidden-Unicode checks, behavior rules).
Run in **Solo mode**: you are a single agent, so perform all seven lanes yourself in
dependency order, adopting each persona in `pr-swarm-review/personas/0N-*.md` in turn
(lanes 1-2 first, then 3-6, then lane 7). Keep every lane's findings in context. Lane 7
(synthesis critic) is a hard gate: do not emit the final review until its "Required
corrections before posting" section is empty — revise and re-run it otherwise.
Stay strictly read-only: investigate and report; never edit files, commit, or post to GitHub.
"""
@@ -0,0 +1,19 @@
---
description: 'GitNexus production-readiness PR swarm review (Solo mode)'
mode: 'agent'
---
You are the GitNexus PR review coordinator. Review the pull request the user names (a PR URL
or number for `https://github.com/abhigyanpatwari/GitNexus`). If none was given, ask for one.
Read `pr-swarm-review/orchestration.md` in this repository and follow it exactly — it is the
canonical, CLI-neutral review contract (lanes, classifications, output structure, finding
format, hidden-Unicode checks, behavior rules).
Run in **Solo mode**: you are a single agent, so perform all seven lanes yourself in
dependency order, adopting each persona in `pr-swarm-review/personas/0N-*.md` in turn
(lanes 1–2 first, then 3–6, then lane 7). Keep every lane's findings in context. Lane 7
(synthesis critic) is a hard gate: do not emit the final review until its "Required
corrections before posting" section is empty.
Stay strictly read-only: investigate and report; never edit files, commit, or post to GitHub.
+2 -2
View File
@@ -48,7 +48,7 @@ jobs:
persist-credentials: false
- name: Initialize CodeQL
uses: github/codeql-action/init@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
uses: github/codeql-action/init@9e0d7b8d25671d64c341c19c0152d693099fb5ba # v4.35.5
with:
languages: ${{ matrix.language }}
queries: security-and-quality
@@ -69,6 +69,6 @@ jobs:
- '**/test/fixtures/**'
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
uses: github/codeql-action/analyze@9e0d7b8d25671d64c341c19c0152d693099fb5ba # v4.35.5
with:
category: '/language:${{ matrix.language }}'
+1 -1
View File
@@ -53,6 +53,6 @@ jobs:
retention-days: 5
- name: Upload to Security tab
uses: github/codeql-action/upload-sarif@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
uses: github/codeql-action/upload-sarif@9e0d7b8d25671d64c341c19c0152d693099fb5ba # v4.35.5
with:
sarif_file: results.sarif
+1 -1
View File
@@ -76,7 +76,7 @@ jobs:
exit-code: '0'
- name: Upload to Security tab
uses: github/codeql-action/upload-sarif@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
uses: github/codeql-action/upload-sarif@9e0d7b8d25671d64c341c19c0152d693099fb5ba # v4.35.5
with:
sarif_file: trivy-${{ matrix.image.name }}.sarif
category: trivy-${{ matrix.image.name }}
+1 -1
View File
@@ -76,7 +76,7 @@ jobs:
continue-on-error: true
- name: Upload SARIF
uses: github/codeql-action/upload-sarif@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
uses: github/codeql-action/upload-sarif@9e0d7b8d25671d64c341c19c0152d693099fb5ba # v4.35.5
with:
sarif_file: zizmor.sarif
category: zizmor
+4 -2
View File
@@ -91,11 +91,13 @@ gitnexus/vendor/**/node_modules/
.claude-flow/
.claude/agents/
.claude/agents/*
!.claude/agents/gitnexus-*.md
.claude/commands/
.claude/helpers
.claude/skills/
.claude/skills/*
!.claude/skills/gitnexus/
!.claude/skills/gitnexus-pr-swarm-review/
.history/
+12
View File
@@ -44,6 +44,18 @@ Commands and gotchas live under **Repo reference** below and in **[CONTRIBUTING.
- **Cursor:** `.cursor/index.mdc` (always-on); `.cursor/rules/*.mdc` (glob-scoped). Legacy `.cursorrules` deprecated.
- **GitNexus:** skills in `.claude/skills/gitnexus/`; MCP rules in `gitnexus:start` block below.
## PR Swarm Review (cross-CLI)
To run a production-readiness review of a GitNexus pull request from **any** AI CLI, follow
the canonical, CLI-neutral spec **[`pr-swarm-review/orchestration.md`](pr-swarm-review/orchestration.md)**
(seven read-only review personas under `pr-swarm-review/personas/`). It defines two
execution modes with the same output contract: **Swarm mode** (parallel subagents, e.g.
Claude Code) and **Solo mode** (one agent runs all lanes sequentially — Codex, Gemini,
Cursor, Copilot, or any agent reading this file). Per-CLI entrypoints are thin wrappers
listed in [`pr-swarm-review/README.md`](pr-swarm-review/README.md); edit review logic only
in the canonical files, never in the wrappers. The review is read-only — it never edits,
commits, or posts.
## Changelog
| Date | Version | Change |
+8
View File
@@ -660,6 +660,14 @@ UPSTREAM (what depends on this):
Options: `maxDepth`, `minConfidence`, `relationTypes` (`CALLS`, `IMPORTS`, `EXTENDS`, `IMPLEMENTS`), `includeTests`, `limit` (max symbols per depth, default 100), `offset` (pagination start per depth), `summaryOnly` (counts and risk only, omits symbol list)
**Disambiguation** — when several symbols share the target name, `impact` returns a ranked `ambiguous` candidate list instead of guessing. Narrow it with `target_uid` (exact, zero-ambiguity), `file_path`, or `kind` (`Function`, `Class`, `Method`, …). From the CLI these are `--uid`, `--file`, and `--kind`, matching `gitnexus context`:
```bash
gitnexus impact get_embeddings # → ambiguous: lists ranked candidates
gitnexus impact get_embeddings --file src/embed.py # → resolves to the one in that file
gitnexus impact get_embeddings --uid "Function:src/embed.py:get_embeddings" # exact
```
### Process-Grouped Search
```
@@ -53,6 +53,10 @@ export interface SymbolDefinition {
* `ScopeResolver.constraintCompatibility` hook during overload narrowing.
* Absent for symbols that have no constraints (the common case). */
templateConstraints?: unknown;
/** True when the producing language marked this callable as explicit.
* Currently used by C++ overload ranking to exclude explicit constructors
* from implicit user-defined conversion candidates. */
isExplicit?: boolean;
/** Links Method/Constructor/Property to owning Class/Struct/Trait nodeId */
ownerId?: string;
}
+82 -51
View File
@@ -11,12 +11,12 @@
"@langchain/anthropic": "^1.3.29",
"@langchain/core": "^1.1.44",
"@langchain/google-genai": "^2.1.30",
"@langchain/langgraph": "^1.2.9",
"@langchain/langgraph": "^1.3.2",
"@langchain/ollama": "^1.2.6",
"@langchain/openai": "^1.4.5",
"@sigma/edge-curve": "^3.1.0",
"@tailwindcss/vite": "^4.3.0",
"axios": "^1.16.0",
"axios": "^1.16.1",
"d3": "^7.9.0",
"dompurify": "^3.4.3",
"gitnexus-shared": "file:../gitnexus-shared",
@@ -48,12 +48,12 @@
},
"devDependencies": {
"@babel/types": "^7.29.0",
"@playwright/test": "^1.58.2",
"@playwright/test": "^1.60.0",
"@testing-library/jest-dom": "^6.9.1",
"@testing-library/react": "^16.3.2",
"@testing-library/user-event": "^14.6.1",
"@types/dompurify": "^3.2.0",
"@types/node": "^25.6.0",
"@types/node": "^25.9.1",
"@types/react": "^19.2.14",
"@types/react-dom": "^19.2.3",
"@types/react-syntax-highlighter": "^15.5.13",
@@ -1393,13 +1393,14 @@
}
},
"node_modules/@langchain/langgraph": {
"version": "1.2.9",
"resolved": "https://registry.npmjs.org/@langchain/langgraph/-/langgraph-1.2.9.tgz",
"integrity": "sha512-3c7BtGycHC2v9p6w/Hv8L7kEl1YnZYOQTDJtmAp3knk6JOedO7d2bYP3y0SRyhv5orUEGf/KGvx8ZsB/ideP7g==",
"version": "1.3.2",
"resolved": "https://registry.npmjs.org/@langchain/langgraph/-/langgraph-1.3.2.tgz",
"integrity": "sha512-SL7Ktsr681R7da+1b2MVOWEbaCoFJOXEJPTGOjg4JIG4C7quWbTYC8DzxhcCxte6D/8cGp0rYDBnbKLXEpNqlA==",
"license": "MIT",
"dependencies": {
"@langchain/langgraph-checkpoint": "^1.0.1",
"@langchain/langgraph-sdk": "~1.8.9",
"@langchain/langgraph-checkpoint": "^1.0.2",
"@langchain/langgraph-sdk": "~1.9.4",
"@langchain/protocol": "^0.0.15",
"@standard-schema/spec": "1.1.0",
"uuid": "^10.0.0"
},
@@ -1407,7 +1408,7 @@
"node": ">=18"
},
"peerDependencies": {
"@langchain/core": "^1.1.40",
"@langchain/core": "^1.1.44",
"zod": "^3.25.32 || ^4.2.0",
"zod-to-json-schema": "^3.x"
},
@@ -1418,9 +1419,9 @@
}
},
"node_modules/@langchain/langgraph-checkpoint": {
"version": "1.0.1",
"resolved": "https://registry.npmjs.org/@langchain/langgraph-checkpoint/-/langgraph-checkpoint-1.0.1.tgz",
"integrity": "sha512-HM0cJLRpIsSlWBQ/xuDC67l52SqZ62Bh2Y61DX+Xorqwoh5e1KxYvfCD7GnSTbWWhjBOutvnR0vPhu4orFkZfw==",
"version": "1.0.2",
"resolved": "https://registry.npmjs.org/@langchain/langgraph-checkpoint/-/langgraph-checkpoint-1.0.2.tgz",
"integrity": "sha512-F4E5Tr0nt8FGghgdscJtHw+ABzChOHeI80R7Y1pjIHdiJom6c2ieo76vL+FWiny80JmoGqhrVAEIWrw0cXKPxg==",
"license": "MIT",
"dependencies": {
"uuid": "^10.0.0"
@@ -1429,7 +1430,7 @@
"node": ">=18"
},
"peerDependencies": {
"@langchain/core": "^1.0.1"
"@langchain/core": "^1.1.44"
}
},
"node_modules/@langchain/langgraph-checkpoint/node_modules/uuid": {
@@ -1446,27 +1447,25 @@
}
},
"node_modules/@langchain/langgraph-sdk": {
"version": "1.8.10",
"resolved": "https://registry.npmjs.org/@langchain/langgraph-sdk/-/langgraph-sdk-1.8.10.tgz",
"integrity": "sha512-wrB3rkRw5KAmsqezwvKP3midT4qJrV6Hj9XJMYo+cbvXC4HYpSAmyY/VriSyeTFRbLG/OP/pY2Yz+9Z54nSaXQ==",
"version": "1.9.9",
"resolved": "https://registry.npmjs.org/@langchain/langgraph-sdk/-/langgraph-sdk-1.9.9.tgz",
"integrity": "sha512-aiWHbmqxWj5sAMwFsaB3eSGQvKpMbUKTlt9zbAC0T7IiFqDYUWi9gJUGsTdvJutAfB3P/NzC4s8ETUtUQEUlYg==",
"license": "MIT",
"dependencies": {
"@langchain/protocol": "^0.0.15",
"@types/json-schema": "^7.0.15",
"p-queue": "^9.0.1",
"p-retry": "^7.1.1",
"uuid": "^13.0.0"
},
"peerDependencies": {
"@langchain/core": "^1.1.16",
"@langchain/core": "^1.1.44",
"react": "^18 || ^19",
"react-dom": "^18 || ^19",
"svelte": "^4.0.0 || ^5.0.0",
"vue": "^3.0.0"
},
"peerDependenciesMeta": {
"@langchain/core": {
"optional": true
},
"react": {
"optional": true
},
@@ -1488,9 +1487,9 @@
"license": "MIT"
},
"node_modules/@langchain/langgraph-sdk/node_modules/p-queue": {
"version": "9.2.0",
"resolved": "https://registry.npmjs.org/p-queue/-/p-queue-9.2.0.tgz",
"integrity": "sha512-dWgLE8AH0HjQ9fe74pUkKkvzzYT18Inp4zra3lKHnnwqGvcfcUBrvF2EAVX+envufDNBOzpPq/IBUONDbI7+3g==",
"version": "9.3.0",
"resolved": "https://registry.npmjs.org/p-queue/-/p-queue-9.3.0.tgz",
"integrity": "sha512-7NED7xhQ74Ngp4JP/2e0VZHp7vSWfJfqeiR92jPgxsz6m0Se4P03YoTKa9dDXyZ3r6P616gUXttrB6nnHYKang==",
"license": "MIT",
"dependencies": {
"eventemitter3": "^5.0.4",
@@ -1516,9 +1515,9 @@
}
},
"node_modules/@langchain/langgraph-sdk/node_modules/uuid": {
"version": "13.0.1",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-13.0.1.tgz",
"integrity": "sha512-9ezox2roIft6ExBVTVqibSd5dc5/47Sw/uY6b4SjQUT2TzQ0tltNquWA46y4xPQmdZYqvnio22SgWd41M86+jw==",
"version": "13.0.2",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-13.0.2.tgz",
"integrity": "sha512-vzi9uRZ926x4XV73S/4qQaTwPXM2JBj6/6lI/byHH1jOpCzb0zDbfytgA9LcN/hzb2l7WQSQnxITOVx5un/wGw==",
"funding": [
"https://github.com/sponsors/broofa",
"https://github.com/sponsors/ctavan"
@@ -1587,6 +1586,12 @@
"@langchain/core": "^1.1.42"
}
},
"node_modules/@langchain/protocol": {
"version": "0.0.15",
"resolved": "https://registry.npmjs.org/@langchain/protocol/-/protocol-0.0.15.tgz",
"integrity": "sha512-MllvbpMjqHevUm+v94M422mH7XKN+wGCvJRBVROTWBotEDOATYB4Ktk2UheYP859y9o2LlhtPek5t1T9eyfAbQ==",
"license": "MIT"
},
"node_modules/@mapbox/node-pre-gyp": {
"version": "2.0.3",
"resolved": "https://registry.npmjs.org/@mapbox/node-pre-gyp/-/node-pre-gyp-2.0.3.tgz",
@@ -1684,13 +1689,13 @@
}
},
"node_modules/@playwright/test": {
"version": "1.58.2",
"resolved": "https://registry.npmjs.org/@playwright/test/-/test-1.58.2.tgz",
"integrity": "sha512-akea+6bHYBBfA9uQqSYmlJXn61cTa+jbO87xVLCWbTqbWadRVmhxlXATaOjOgcBaWU4ePo0wB41KMFv3o35IXA==",
"version": "1.60.0",
"resolved": "https://registry.npmjs.org/@playwright/test/-/test-1.60.0.tgz",
"integrity": "sha512-O71yZIbAh/PxDMNGns37GHBIfrVkEVyn+AXyIa5dOTfb4/xNvRWV+Vv/NMbNCtODB/pO7vLlF2OTmMVLhmr7Ag==",
"dev": true,
"license": "Apache-2.0",
"dependencies": {
"playwright": "1.58.2"
"playwright": "1.60.0"
},
"bin": {
"playwright": "cli.js"
@@ -2854,13 +2859,13 @@
"license": "MIT"
},
"node_modules/@types/node": {
"version": "25.6.0",
"resolved": "https://registry.npmjs.org/@types/node/-/node-25.6.0.tgz",
"integrity": "sha512-+qIYRKdNYJwY3vRCZMdJbPLJAtGjQBudzZzdzwQYkEPQd+PJGixUL5QfvCLDaULoLv+RhT3LDkwEfKaAkgSmNQ==",
"version": "25.9.1",
"resolved": "https://registry.npmjs.org/@types/node/-/node-25.9.1.tgz",
"integrity": "sha512-xfrlY7UD5rMJk3ZVJP8BNzS28J36YJg+xp+LPXV1TdWxr8uMH5A860QNxYDGQe/ylDSgjxE52Q9VnO7p75tJxg==",
"devOptional": true,
"license": "MIT",
"dependencies": {
"undici-types": "~7.19.0"
"undici-types": ">=7.24.0 <7.24.7"
}
},
"node_modules/@types/prismjs": {
@@ -3414,16 +3419,42 @@
"license": "MIT"
},
"node_modules/axios": {
"version": "1.16.0",
"resolved": "https://registry.npmjs.org/axios/-/axios-1.16.0.tgz",
"integrity": "sha512-6hp5CwvTPlN2A31g5dxnwAX0orzM7pmCRDLnZSX772mv8WDqICwFjowHuPs04Mc8deIld1+ejhtaMn5vp6b+1w==",
"version": "1.16.1",
"resolved": "https://registry.npmjs.org/axios/-/axios-1.16.1.tgz",
"integrity": "sha512-caYkukvroVPO8KrzuJEb50Hm07KwfBZPEC3VeFHTsqWHvKTsy54hjJz9BS/cdaypROE2rH6xvm9mHX4fgWkr3A==",
"license": "MIT",
"dependencies": {
"follow-redirects": "^1.16.0",
"form-data": "^4.0.5",
"https-proxy-agent": "^5.0.1",
"proxy-from-env": "^2.1.0"
}
},
"node_modules/axios/node_modules/agent-base": {
"version": "6.0.2",
"resolved": "https://registry.npmjs.org/agent-base/-/agent-base-6.0.2.tgz",
"integrity": "sha512-RZNwNclF7+MS/8bDg70amg32dyeZGZxiDuQmZxKLAlQjr3jGyLx+4Kkk58UO7D2QdgFIQCovuSuZESne6RG6XQ==",
"license": "MIT",
"dependencies": {
"debug": "4"
},
"engines": {
"node": ">= 6.0.0"
}
},
"node_modules/axios/node_modules/https-proxy-agent": {
"version": "5.0.1",
"resolved": "https://registry.npmjs.org/https-proxy-agent/-/https-proxy-agent-5.0.1.tgz",
"integrity": "sha512-dFcAjpTQFgoLMzC2VwU+C/CbS7uRL0lWmxDITmqm7C+7F0Odmj6s9l6alZc6AELXhrnggM2CeWSXHGOdX2YtwA==",
"license": "MIT",
"dependencies": {
"agent-base": "6",
"debug": "4"
},
"engines": {
"node": ">= 6"
}
},
"node_modules/bail": {
"version": "2.0.2",
"resolved": "https://registry.npmjs.org/bail/-/bail-2.0.2.tgz",
@@ -5444,9 +5475,9 @@
}
},
"node_modules/is-network-error": {
"version": "1.3.1",
"resolved": "https://registry.npmjs.org/is-network-error/-/is-network-error-1.3.1.tgz",
"integrity": "sha512-6QCxa49rQbmUWLfk0nuGqzql9U8uaV2H6279bRErPBHe/109hCzsLUBUHfbEtvLIHBd6hyXbgedBSHevm43Edw==",
"version": "1.3.2",
"resolved": "https://registry.npmjs.org/is-network-error/-/is-network-error-1.3.2.tgz",
"integrity": "sha512-PhBY86zaxNZUuWP6h13Vu5oFe0XY6/UlKzQnYFELzGVHygP3MxmvTfYSG7GN3aIab/iWudSMgjSnG9Dq+nHrgA==",
"license": "MIT",
"engines": {
"node": ">=16"
@@ -7572,13 +7603,13 @@
}
},
"node_modules/playwright": {
"version": "1.58.2",
"resolved": "https://registry.npmjs.org/playwright/-/playwright-1.58.2.tgz",
"integrity": "sha512-vA30H8Nvkq/cPBnNw4Q8TWz1EJyqgpuinBcHET0YVJVFldr8JDNiU9LaWAE1KqSkRYazuaBhTpB5ZzShOezQ6A==",
"version": "1.60.0",
"resolved": "https://registry.npmjs.org/playwright/-/playwright-1.60.0.tgz",
"integrity": "sha512-hheHdokM8cdqCb0lcE3s+zT4t4W+vvjpGxsZlDnikarzx8tSzMebh3UiFtgqwFwnTnjYQcsyMF8ei2mCO/tpeA==",
"dev": true,
"license": "Apache-2.0",
"dependencies": {
"playwright-core": "1.58.2"
"playwright-core": "1.60.0"
},
"bin": {
"playwright": "cli.js"
@@ -7591,9 +7622,9 @@
}
},
"node_modules/playwright-core": {
"version": "1.58.2",
"resolved": "https://registry.npmjs.org/playwright-core/-/playwright-core-1.58.2.tgz",
"integrity": "sha512-yZkEtftgwS8CsfYo7nm0KE8jsvm6i/PTgVtB8DL726wNf6H2IMsDuxCpJj59KDaxCtSnrWan2AeDqM7JBaultg==",
"version": "1.60.0",
"resolved": "https://registry.npmjs.org/playwright-core/-/playwright-core-1.60.0.tgz",
"integrity": "sha512-9bW6zvX/m0lEbgTKJ6YppOKx8H3VOPBMOCFh2irXFOT4BbHgrx5hPjwJYLT40Lu+4qtD36qKc/Hn56StUW57IA==",
"dev": true,
"license": "Apache-2.0",
"bin": {
@@ -8567,9 +8598,9 @@
}
},
"node_modules/undici-types": {
"version": "7.19.2",
"resolved": "https://registry.npmjs.org/undici-types/-/undici-types-7.19.2.tgz",
"integrity": "sha512-qYVnV5OEm2AW8cJMCpdV20CDyaN3g0AjDlOGf1OW4iaDEx8MwdtChUp4zu4H0VP3nDRF/8RKWH+IPp9uW0YGZg==",
"version": "7.24.6",
"resolved": "https://registry.npmjs.org/undici-types/-/undici-types-7.24.6.tgz",
"integrity": "sha512-WRNW+sJgj5OBN4/0JpHFqtqzhpbnV0GuB+OozA9gCL7a993SmU+1JBZCzLNxYsbMfIeDL+lTsphD5jN5N+n0zg==",
"devOptional": true,
"license": "MIT"
},
+4 -4
View File
@@ -21,12 +21,12 @@
"@langchain/anthropic": "^1.3.29",
"@langchain/core": "^1.1.44",
"@langchain/google-genai": "^2.1.30",
"@langchain/langgraph": "^1.2.9",
"@langchain/langgraph": "^1.3.2",
"@langchain/ollama": "^1.2.6",
"@langchain/openai": "^1.4.5",
"@sigma/edge-curve": "^3.1.0",
"@tailwindcss/vite": "^4.3.0",
"axios": "^1.16.0",
"axios": "^1.16.1",
"d3": "^7.9.0",
"dompurify": "^3.4.3",
"gitnexus-shared": "file:../gitnexus-shared",
@@ -58,12 +58,12 @@
},
"devDependencies": {
"@babel/types": "^7.29.0",
"@playwright/test": "^1.58.2",
"@playwright/test": "^1.60.0",
"@testing-library/jest-dom": "^6.9.1",
"@testing-library/react": "^16.3.2",
"@testing-library/user-event": "^14.6.1",
"@types/dompurify": "^3.2.0",
"@types/node": "^25.6.0",
"@types/node": "^25.9.1",
"@types/react": "^19.2.14",
"@types/react-dom": "^19.2.3",
"@types/react-syntax-highlighter": "^15.5.13",
+7
View File
@@ -170,6 +170,13 @@ gitnexus clean --all --force # Delete all indexes
gitnexus wiki [path] # Generate LLM-powered docs from knowledge graph
gitnexus wiki --model <model> # Wiki with custom LLM model (default: gpt-4o-mini)
# Direct graph queries — the same tools the MCP server exposes, no MCP daemon needed
gitnexus query "<concept>" # Process-grouped hybrid search
gitnexus context <symbol> [--uid <uid> | --file <path>] # 360° symbol view; flags disambiguate a shared name
gitnexus impact <symbol> [--uid <uid> | --file <path> | --kind <kind>] # Blast radius; flags disambiguate a shared name
gitnexus detect-changes # Map the working-tree diff to affected symbols and execution flows
gitnexus cypher "<query>" # Run a raw Cypher query against the knowledge graph
# Repository groups (multi-repo / monorepo service tracking)
gitnexus group create <name> # Create a repository group
gitnexus group add <group> <groupPath> <registryName> # Add a repo to a group. <groupPath> is a hierarchy path (e.g. hr/hiring/backend); <registryName> is the repo's name from the registry (see `gitnexus list`)
+10 -9
View File
@@ -28,6 +28,7 @@
"jsonc-parser": "^3.3.1",
"lru-cache": "^11.0.0",
"mnemonist": "^0.40.3",
"node-addon-api": "8.8.0",
"onnxruntime-node": "^1.24.0",
"pandemonium": "^2.4.0",
"pino": "^10.3.1",
@@ -1792,9 +1793,9 @@
"license": "MIT"
},
"node_modules/@types/node": {
"version": "25.9.0",
"resolved": "https://registry.npmjs.org/@types/node/-/node-25.9.0.tgz",
"integrity": "sha512-AOQwYUNolgy3VosiRqXrACUXTN8nJUtPl7FJXMqZVyxiiCLhQuG3jXKvCS1ALr+Y2OmZhzzLVlYPEqJaiqkaJQ==",
"version": "25.9.1",
"resolved": "https://registry.npmjs.org/@types/node/-/node-25.9.1.tgz",
"integrity": "sha512-xfrlY7UD5rMJk3ZVJP8BNzS28J36YJg+xp+LPXV1TdWxr8uMH5A860QNxYDGQe/ylDSgjxE52Q9VnO7p75tJxg==",
"license": "MIT",
"dependencies": {
"undici-types": ">=7.24.0 <7.24.7"
@@ -3594,9 +3595,9 @@
"license": "Apache-2.0"
},
"node_modules/lru-cache": {
"version": "11.4.0",
"resolved": "https://registry.npmjs.org/lru-cache/-/lru-cache-11.4.0.tgz",
"integrity": "sha512-W+R+kFL4HgVxONq2bhXPi3bGpzGe/yEhVOp233qw9wCRtgncJ15P3bC+e4zZMu4Cq7d+WAJjXGW0uUkifhcatA==",
"version": "11.5.0",
"resolved": "https://registry.npmjs.org/lru-cache/-/lru-cache-11.5.0.tgz",
"integrity": "sha512-5YgH9UJd7wVb9hIouI2adWpgqrrICkt070Dnj8EUY1+B4B2P9eRLPAkAAo6NICA7CEhOIeBHl46u9zSNpNu7zA==",
"license": "BlueOak-1.0.0",
"engines": {
"node": "20 || >=22"
@@ -3799,9 +3800,9 @@
}
},
"node_modules/node-addon-api": {
"version": "8.7.0",
"resolved": "https://registry.npmjs.org/node-addon-api/-/node-addon-api-8.7.0.tgz",
"integrity": "sha512-9MdFxmkKaOYVTV+XVRG8ArDwwQ77XIgIPyKASB1k3JPq3M8fGQQQE3YpMOrKm6g//Ktx8ivZr8xo1Qmtqub+GA==",
"version": "8.8.0",
"resolved": "https://registry.npmjs.org/node-addon-api/-/node-addon-api-8.8.0.tgz",
"integrity": "sha512-c5Ko1fZJIJmzhFIkhRN76WTq+fC6tWnGy9CXA0fA+XygsWZmEwG8vmbkNqxMyoaa0Tin4djul49NzdVcJJcjeA==",
"license": "MIT",
"engines": {
"node": "^18 || ^20 || >= 21"
+11
View File
@@ -1096,6 +1096,17 @@ const analyzeCommandImpl = async (inputPath?: string, options?: AnalyzeOptions):
);
console.log(` ${repoPath}`);
// Persistent (non-scrolling) warning when FTS indexing was skipped — the
// progress-bar log() that fired mid-run has already scrolled away, so the
// degraded-search state must also appear in the final summary (#1161).
if (result.ftsSkipped) {
console.log(
`\n Warning: full-text/BM25 search is disabled — the LadybugDB FTS extension was unavailable.\n` +
` Install it once with network access (GITNEXUS_LBUG_EXTENSION_INSTALL=auto) then rerun, or\n` +
` run \`gitnexus analyze --repair-fts\` when connected. Run \`gitnexus doctor\` for details.`,
);
}
try {
await fs.access(getGlobalRegistryPath());
} catch {
+12
View File
@@ -2,6 +2,7 @@ import { getRuntimeCapabilities, getRuntimeFingerprint } from '../core/platform/
import { resolveEmbeddingConfig } from '../core/embeddings/config.js';
import { isHttpMode } from '../core/embeddings/http-client.js';
import { checkLbugNative } from '../core/lbug/native-check.js';
import { getExtensionInstallPolicy } from '../core/lbug/extension-loader.js';
import { t } from './i18n/index.js';
function isCombiningMark(codePoint: number): boolean {
@@ -74,6 +75,17 @@ export const doctorCommand = async () => {
console.log(` ${label('doctor.labels.fullTextSearch', 18)}${capabilities.fts}`);
console.log(` ${label('doctor.labels.vectorIndex', 18)}${capabilities.vector}`);
console.log(` ${label('doctor.labels.semanticMode', 18)}${capabilities.semanticMode}`);
// Surface the optional-extension install policy so offline users can see
// whether analyze/query will reach the network (extension.ladybugdb.com).
// Literal label (like the 'native' line) to avoid adding i18n keys.
const installPolicy = getExtensionInstallPolicy();
const policyHint =
installPolicy === 'load-only'
? ' (offline; load only, no network install)'
: installPolicy === 'never'
? ' (optional extensions disabled)'
: ' (installs missing extensions over network)';
console.log(` ${padDisplayEnd('Ext install:', 18)}${installPolicy}${policyHint}`);
console.log(
` ${label('doctor.labels.exactScanLimit', 18)}${t('doctor.chunks', { count: capabilities.exactScanLimit })}`,
);
+3
View File
@@ -101,6 +101,9 @@ const OPTION_DESCRIPTION_KEYS = {
'context|--content': 'help.option.content',
'impact|-d, --direction <dir>': 'help.option.impact.direction',
'impact|-r, --repo <name>': 'help.option.repo.target',
'impact|-u, --uid <uid>': 'help.option.context.uid',
'impact|-f, --file <path>': 'help.option.context.file',
'impact|--kind <kind>': 'help.option.impact.kind',
'impact|--depth <n>': 'help.option.impact.depth',
'impact|--include-tests': 'help.option.impact.includeTests',
'impact|--limit <n>': 'help.option.impact.limit',
+6 -1
View File
@@ -43,8 +43,11 @@ export const en = {
'tool.noIndexed': 'GitNexus: No indexed repositories found. Run: gitnexus analyze',
'tool.usage.query': 'Usage: gitnexus query <search_query>',
'tool.usage.context': 'Usage: gitnexus context <symbol_name> [--uid <uid>] [--file <path>]',
'tool.usage.impact': 'Usage: gitnexus impact <symbol_name> [--direction upstream|downstream]',
'tool.usage.impact':
'Usage: gitnexus impact <symbol_name> [--uid <uid>] [--file <path>] [--kind <kind>] [--direction upstream|downstream]',
'tool.usage.cypher': 'Usage: gitnexus cypher <cypher_query>',
'tool.warn.unknownKind':
"--kind '{{kind}}' is not a known symbol kind (e.g. Function, Class, Method); it will not narrow the result.",
'tool.detectChanges.noChanges': 'No changes detected.',
'tool.detectChanges.changesSummary': 'Changes: {{files}} files, {{symbols}} symbols',
'tool.detectChanges.affectedProcesses': 'Affected processes: {{count}}',
@@ -213,6 +216,8 @@ export const en = {
'help.option.repo.target': 'Target repository',
'help.option.context.uid': 'Direct symbol UID (zero-ambiguity lookup)',
'help.option.context.file': 'File path to disambiguate common names',
'help.option.impact.kind':
'Kind filter to disambiguate common names (e.g. Function, Class, Method)',
'help.option.impact.direction': 'upstream (dependants) or downstream (dependencies)',
'help.option.impact.depth': 'Max relationship depth (default: 3)',
'help.option.impact.includeTests': 'Include test files in results',
+5 -1
View File
@@ -47,8 +47,11 @@ export const zhCN = {
'tool.noIndexed': 'GitNexus:未找到已索引仓库。请运行:gitnexus analyze',
'tool.usage.query': '用法:gitnexus query <搜索词>',
'tool.usage.context': '用法:gitnexus context <符号名> [--uid <uid>] [--file <路径>]',
'tool.usage.impact': '用法:gitnexus impact <符号名> [--direction upstream|downstream]',
'tool.usage.impact':
'用法:gitnexus impact <符号名> [--uid <uid>] [--file <路径>] [--kind <类型>] [--direction upstream|downstream]',
'tool.usage.cypher': '用法:gitnexus cypher <Cypher 查询>',
'tool.warn.unknownKind':
"--kind '{{kind}}' 不是已知的符号类型(如 Function、Class、Method),不会用于缩小结果范围。",
'tool.detectChanges.noChanges': '未检测到变更。',
'tool.detectChanges.changesSummary': '变更:{{files}} 个文件,{{symbols}} 个符号',
'tool.detectChanges.affectedProcesses': '受影响流程:{{count}}',
@@ -199,6 +202,7 @@ export const zhCN = {
'help.option.repo.target': '目标仓库',
'help.option.context.uid': '直接符号 UID(零歧义查找)',
'help.option.context.file': '用于消除常见名称歧义的文件路径',
'help.option.impact.kind': '用于消除常见名称歧义的类型过滤(如 Function、Class、Method)',
'help.option.impact.direction': 'upstream(依赖它的项)或 downstream(它依赖的项)',
'help.option.impact.depth': '最大关系遍历深度(默认:3)',
'help.option.impact.includeTests': '在结果中包含测试文件',
+7 -1
View File
@@ -219,10 +219,16 @@ program
.action(createLbugLazyAction(() => import('./tool.js'), 'contextCommand'));
program
.command('impact <target>')
.command('impact [target]')
.description('Blast radius analysis: what breaks if you change a symbol')
.option('-d, --direction <dir>', 'upstream (dependants) or downstream (dependencies)', 'upstream')
.option('-r, --repo <name>', 'Target repository')
.option('-u, --uid <uid>', 'Direct symbol UID (zero-ambiguity lookup)')
.option('-f, --file <path>', 'File path to disambiguate common names')
.option(
'--kind <kind>',
'Kind filter to disambiguate common names (e.g. Function, Class, Method)',
)
.option('--depth <n>', 'Max relationship depth (default: 3)')
.option('--include-tests', 'Include test files in results')
.option('--limit <n>', 'Max symbols per depth level (default: 100)')
+31 -5
View File
@@ -16,8 +16,8 @@
*/
import { writeSync } from 'node:fs';
import { LocalBackend } from '../mcp/local/local-backend.js';
import { cliErrorKey } from './cli-message.js';
import { LocalBackend, VALID_NODE_LABELS } from '../mcp/local/local-backend.js';
import { cliErrorKey, cliWarnKey } from './cli-message.js';
import { formatDetectChangesResult } from './detect-changes-format.js';
let _backend: LocalBackend | null = null;
@@ -94,6 +94,11 @@ export async function contextCommand(
content?: boolean;
},
): Promise<void> {
// Reject a `--`-prefixed uid swallowed from a following flag (see impactCommand).
if (options?.uid?.startsWith('--')) {
cliErrorKey('tool.usage.context');
process.exit(1);
}
if (!name?.trim() && !options?.uid) {
cliErrorKey('tool.usage.context');
process.exit(1);
@@ -111,10 +116,13 @@ export async function contextCommand(
}
export async function impactCommand(
target: string,
target?: string,
options?: {
direction?: string;
repo?: string;
uid?: string;
file?: string;
kind?: string;
depth?: string;
includeTests?: boolean;
limit?: string;
@@ -122,10 +130,25 @@ export async function impactCommand(
summaryOnly?: boolean;
},
): Promise<void> {
if (!target?.trim()) {
// A `--`-prefixed uid means Commander swallowed a following flag as the uid
// value (e.g. `impact --uid --file x` → uid === '--file'). Reject it rather
// than forwarding a garbage uid that would silently resolve to not-found.
if (options?.uid?.startsWith('--')) {
cliErrorKey('tool.usage.impact');
process.exit(1);
}
// Target is an optional positional: a uid alone is enough to resolve (parity
// with `context [name]`). Only error when neither a target nor a uid is given.
if (!target?.trim() && !options?.uid) {
cliErrorKey('tool.usage.impact');
process.exit(1);
}
// Soft-validate --kind: an unknown kind is a no-op hint (the backend scores
// it but it matches nothing), so warn and proceed rather than rejecting —
// parity with the lenient MCP surface and forward-compatible with new labels.
if (options?.kind && !VALID_NODE_LABELS.has(options.kind)) {
cliWarnKey('tool.warn.unknownKind', { kind: options.kind });
}
try {
const backend = await getBackend();
@@ -134,7 +157,10 @@ export async function impactCommand(
const parsedLimit = Number.isFinite(rawLimit) ? rawLimit : undefined;
const parsedOffset = Number.isFinite(rawOffset) ? rawOffset : undefined;
const result = await backend.callTool('impact', {
target,
target: target || undefined,
target_uid: options?.uid,
file_path: options?.file,
kind: options?.kind,
direction: options?.direction || 'upstream',
maxDepth: options?.depth ? parseInt(options.depth, 10) : undefined,
includeTests: options?.includeTests ?? false,
@@ -43,20 +43,38 @@ import {
STALE_HASH_SENTINEL,
} from '../lbug/schema.js';
import { loadVectorExtension } from '../lbug/lbug-adapter.js';
import type { ExtensionInstallPolicy } from '../lbug/extension-loader.js';
import { getExactScanLimit } from '../platform/capabilities.js';
import { logger } from '../logger.js';
const isDev = process.env.NODE_ENV === 'development';
const vectorUnavailableMessage =
'VECTOR extension is unavailable for this LadybugDB runtime; semantic search will use exact scan when embeddings exist.';
'VECTOR extension unavailable; semantic embeddings fall back to exact scan. ' +
'To enable vector search, install it once with network access ' +
'(GITNEXUS_LBUG_EXTENSION_INSTALL=auto), or pre-install it for offline use. ' +
'Set GITNEXUS_LBUG_EXTENSION_INSTALL=never to skip installs and silence this.';
/**
* Resolve the extension-install policy for the embedding WRITE path (analyze).
*
* Generating embeddings is an explicit opt-in to a feature that requires the
* VECTOR extension, so when the operator has NOT pinned a policy we default to
* `auto` (one bounded, out-of-process INSTALL) — matching the documented
* "auto = default for analyze" intent in extension-loader.ts. An explicit
* GITNEXUS_LBUG_EXTENSION_INSTALL=load-only|never|auto always wins, so an
* offline or locked-down operator is never silently forced onto the network
* (the #1153 regression caused by hard-coding `auto` here). Read on every call
* (not memoized) so test env stubbing works.
*/
export const resolveEmbeddingInstallPolicy = (): ExtensionInstallPolicy => {
const raw = process.env.GITNEXUS_LBUG_EXTENSION_INSTALL;
if (raw === 'load-only' || raw === 'never' || raw === 'auto') return raw;
return 'auto';
};
const ensureVectorExtensionAvailable = async (): Promise<boolean> => {
const vectorReady = await loadVectorExtension();
if (!vectorReady) {
return false;
}
return true;
return loadVectorExtension(undefined, { policy: resolveEmbeddingInstallPolicy() });
};
/**
* Bump this when the embedding text template changes in a way that should
@@ -257,7 +275,7 @@ export const runEmbeddingPipeline = async (
try {
const vectorAvailable = await ensureVectorExtensionAvailable();
if (!vectorAvailable && isDev) {
if (!vectorAvailable) {
logger.warn(vectorUnavailableMessage);
}
@@ -584,7 +602,11 @@ export const semanticSearch = async (
string,
{ distance: number; chunkIndex: number; startLine: number; endLine: number }
>();
if (await loadVectorExtension()) {
// Query/read path: NEVER spawn a network INSTALL on a user query. If the
// VECTOR extension was not pre-installed, fall back to exact scan rather than
// blocking the query on a download (offline-first; see extension-loader.ts
// "load-only" — used by all serve/MCP query paths).
if (await loadVectorExtension(undefined, { policy: 'load-only' })) {
try {
bestChunks = await collectBestChunks(k, async (fetchLimit) => {
const vectorQuery = `
@@ -188,6 +188,16 @@ function makeContract(
export interface ProtoServiceInfo {
package: string;
/**
* Optional. Value of `option java_package = "..."` declared in the
* same `.proto` file, when present and different from `package`.
* Empty string when the option is absent or equals `package`. Used by
* `detectionToContract()` to translate a Java import path back to the
* proto package whenever the proto explicitly publishes its generated
* Java code under a different namespace (a common pattern in
* Google-style protobuf projects).
*/
javaPackage: string;
serviceName: string;
methods: string[];
protoPath: string;
@@ -207,6 +217,19 @@ function extractProtoImports(content: string): string[] {
return imports;
}
/**
* Extract `option java_package = "..."` from a `.proto` file, if any.
* The Java code generator places generated `XxxGrpc.java` classes under
* this package (instead of the proto `package` declaration) when the
* option is set. Real-world projects (Google Cloud Java APIs, internal
* shaded SDKs) routinely use this to publish their Java artifacts under
* a corporate namespace different from the wire-protocol package.
*/
function extractJavaPackageOption(content: string): string {
const m = content.match(/^\s*option\s+java_package\s*=\s*"([\w.]+)"\s*;/m);
return m?.[1] ?? '';
}
function longestSharedSegmentRun(aPath: string, bPath: string): number {
const a = aPath.split('/').filter(Boolean);
const b = bPath.split('/').filter(Boolean);
@@ -228,8 +251,18 @@ function longestSharedSegmentRun(aPath: string, bPath: string): number {
async function buildProtoContext(repoPath: string): Promise<{
packagesByProto: Map<string, string>;
servicesByName: Map<string, ProtoServiceInfo[]>;
/**
* Reverse index: `option java_package` value → ProtoServiceInfo[]
* declared in `.proto` files that ship under that Java namespace.
* Only populated when `java_package` is set AND differs from
* `package`. Lets `detectionToContract()` translate an import-derived
* Java package back to its source proto package whenever the proto
* is in the same repository.
*/
servicesByJavaPackage: Map<string, ProtoServiceInfo[]>;
}> {
const servicesByName = new Map<string, ProtoServiceInfo[]>();
const servicesByJavaPackage = new Map<string, ProtoServiceInfo[]>();
// `.gitnexusignore` / `.gitignore` honoured via the shared IgnoreService —
// see `filesystem-walker.ts` for the canonical pattern. Replaces a
// hardcoded `[node_modules, .git, vendor]` array; those names plus the
@@ -292,6 +325,13 @@ async function buildProtoContext(repoPath: string): Promise<{
const content = contents.get(normalizedRel);
if (!content) continue;
const pkg = resolvePackage(normalizedRel);
const javaPkgOption = extractJavaPackageOption(content);
// Only retain `javaPackage` when it actively diverges from `pkg`.
// When equal (or absent), the import-derived path produces the
// same FQN as the proto-derived path, so no translation is needed
// and we keep the field empty to avoid populating the reverse
// index with redundant entries.
const javaPackage = javaPkgOption && javaPkgOption !== pkg ? javaPkgOption : '';
const serviceBlocks = extractServiceBlocks(content);
for (const block of serviceBlocks) {
@@ -303,6 +343,7 @@ async function buildProtoContext(repoPath: string): Promise<{
}
const info: ProtoServiceInfo = {
package: pkg,
javaPackage,
serviceName: block.name,
methods,
protoPath: normalizedRel,
@@ -310,10 +351,16 @@ async function buildProtoContext(repoPath: string): Promise<{
const existing = servicesByName.get(block.name) ?? [];
existing.push(info);
servicesByName.set(block.name, existing);
if (javaPackage) {
const byJava = servicesByJavaPackage.get(javaPackage) ?? [];
byJava.push(info);
servicesByJavaPackage.set(javaPackage, byJava);
}
}
}
return { packagesByProto, servicesByName };
return { packagesByProto, servicesByName, servicesByJavaPackage };
}
export async function buildProtoMap(repoPath: string): Promise<Map<string, ProtoServiceInfo[]>> {
@@ -377,6 +424,7 @@ export class GrpcExtractor implements ContractExtractor {
const out: ExtractedContract[] = [];
const protoContext = await buildProtoContext(repoPath);
const protoMap = protoContext.servicesByName;
const javaPackageMap = protoContext.servicesByJavaPackage;
// ─── Proto files — definitive provider source ─────────────────
// When tree-sitter-proto is available, .proto files are handled by
@@ -435,7 +483,7 @@ export class GrpcExtractor implements ContractExtractor {
continue;
}
for (const d of detections) {
const contract = this.detectionToContract(d, rel, protoMap);
const contract = this.detectionToContract(d, rel, protoMap, javaPackageMap);
if (contract) out.push(contract);
}
}
@@ -449,12 +497,163 @@ export class GrpcExtractor implements ContractExtractor {
* either a service-level (`grpc::pkg.Svc/*`) or method-level
* (`grpc::pkg.Svc/Method`) contract id, and selecting confidence
* based on whether the proto map had an entry.
*
* Resolution order for the package prefix:
*
* 1. **Java-package translation** (when detection
* supplied a `protoPackage` from a Java import).
* A `.proto` in the SAME repo may set `option
* java_package = "..."` to publish its generated
* Java classes under a namespace different from
* the proto `package`. Real-world projects (e.g.
* Google Cloud Java APIs) routinely do this.
* When the import-derived package matches that
* `java_package` value, translate back to the
* proto `package` so the resulting contract id
* is wire-correct rather than Java-namespace.
*
* 2. **Per-repo proto map check** (when the same
* service name has `.proto` candidates in this
* repo). The proto file is the authoritative
* source. If the proto's `package` agrees with
* the import's `protoPackage`, both paths produce
* the same FQN — emit it. If they DISAGREE (e.g.
* a typo'd Java import, or a mismatched
* java_package the reverse index didn't catch),
* trust the proto map and warn — the import
* MUST NOT silently overwrite an authoritative
* proto package.
*
* 3. **Import-derived FQN fallback** (when neither
* a `java_package` translation nor a proto map
* candidate exists in this repo). Typical for the
* "client-jar" pattern, where a consumer repo
* depends on a published stub jar and never
* carries the originating `.proto`. Use the
* import path verbatim as the proto package. Note
* the known limitation: when the published proto
* sets `option java_package` differing from
* `package`, the resulting FQN reflects the Java
* namespace rather than the proto namespace and
* will not match a provider repo's contract id —
* we cannot translate without sight of the proto.
*
* 4. **Per-repo proto map (no import)** — the legacy
* path. Used when the plugin didn't supply
* `protoPackage` (no import statement, wildcard
* import only, or non-Java languages that haven't
* been retrofitted yet).
*
* 5. **Short-name fallback** — when none of the
* above resolves a package, emit a service-only
* short-name contract id (`grpc::Svc/*`),
* preserving the pre-fix behaviour.
*/
private detectionToContract(
d: GrpcDetection,
filePath: string,
protoMap: Map<string, ProtoServiceInfo[]>,
javaPackageMap: Map<string, ProtoServiceInfo[]>,
): ExtractedContract | null {
if (d.protoPackage) {
// Step 1: java_package translation. The import-derived package
// may be the `option java_package` value of a `.proto` in the
// SAME repo. Look it up and, if found for the same service name,
// use the underlying proto `package` to build a wire-correct
// contract id.
const javaCandidates = javaPackageMap.get(d.protoPackage) ?? [];
const javaTranslated = javaCandidates.find((p) => p.serviceName === d.serviceName);
if (javaTranslated) {
const cid = d.methodName
? contractId(javaTranslated.package, d.serviceName, d.methodName)
: serviceContractId(javaTranslated.package, d.serviceName);
const meta: Record<string, unknown> = {
service: d.serviceName,
source: d.source,
package: javaTranslated.package,
protoPackageSource: 'import-translated',
};
if (d.methodName) meta.method = d.methodName;
return makeContract(cid, d.role, filePath, d.symbolName, d.confidenceWithProto, meta);
}
// Step 2: proto map cross-check. When this repo also carries a
// `.proto` defining the same short service name, the proto is
// authoritative and decides the package. The import is only used
// to disambiguate among same-short-name candidates when the
// resolution heuristic can't pick a unique winner on path alone.
const candidates = protoMap.get(d.serviceName) ?? [];
if (candidates.length > 0) {
const proto = resolveProtoConflict(d.serviceName, filePath, candidates);
if (proto === null) {
// Ambiguous proto resolution; resolveProtoConflict already warned.
return null;
}
const protoPkg = proto.package;
if (protoPkg === d.protoPackage) {
// Both paths agree.
const cid = d.methodName
? contractId(protoPkg, d.serviceName, d.methodName)
: serviceContractId(protoPkg, d.serviceName);
const meta: Record<string, unknown> = {
service: d.serviceName,
source: d.source,
package: protoPkg,
protoPackageSource: 'import',
};
if (d.methodName) meta.method = d.methodName;
return makeContract(cid, d.role, filePath, d.symbolName, d.confidenceWithProto, meta);
}
// Disagreement. Trust the proto file and emit a warning so
// operators can investigate the import. This protects against
// the symmetric Finding 2 case: a stale or typo'd Java import
// silently corrupting the contract id of a service whose
// `.proto` lives in the same repo.
logger.warn(
`[grpc-extractor] Java import package "${d.protoPackage}" for service ` +
`"${d.serviceName}" disagrees with local proto package "${protoPkg}" at ` +
`${filePath}; using proto package as authoritative source`,
);
const cid = d.methodName
? contractId(protoPkg, d.serviceName, d.methodName)
: serviceContractId(protoPkg, d.serviceName);
const meta: Record<string, unknown> = {
service: d.serviceName,
source: d.source,
package: protoPkg,
protoPackageSource: 'proto-override',
importPackage: d.protoPackage,
};
if (d.methodName) meta.method = d.methodName;
return makeContract(cid, d.role, filePath, d.symbolName, d.confidenceWithProto, meta);
}
// Step 3: import-derived fallback. No `.proto` in this repo
// names the service, and no `java_package` reverse-lookup
// matched. Emit the FQN with the import-derived package. This
// is the typical client-jar consumer path.
//
// Known limitation: when the published proto sets
// `option java_package` to a value that differs from
// `package`, this path produces a contract id that reflects
// the Java namespace, not the proto namespace, and will not
// match a provider repo. Resolving that case requires
// group-level proto knowledge, which is intentionally out of
// scope for this fix.
const cid = d.methodName
? contractId(d.protoPackage, d.serviceName, d.methodName)
: serviceContractId(d.protoPackage, d.serviceName);
const meta: Record<string, unknown> = {
service: d.serviceName,
source: d.source,
package: d.protoPackage,
protoPackageSource: 'import',
};
if (d.methodName) meta.method = d.methodName;
return makeContract(cid, d.role, filePath, d.symbolName, d.confidenceWithProto, meta);
}
// Steps 4 + 5: legacy per-repo proto map resolution (no import).
const candidates = protoMap.get(d.serviceName) ?? [];
const proto = resolveProtoConflict(d.serviceName, filePath, candidates);
// If there were proto candidates but resolution was ambiguous, skip
@@ -78,6 +78,33 @@ const STUB_PATTERNS = compilePatterns({
],
} satisfies LanguagePatterns<Record<string, never>>);
// `import <pkg>.<XxxGrpc>;` — captures the proto package of the
// imported gRPC class (e.g. `cn.unipus.ucf.admin.proto.client.service`
// for `import cn.unipus.ucf.admin.proto.client.service.ContentRpcServiceGrpc`).
// Used by `scan` to build a per-file `XxxGrpc → fullPackage` map so
// consumer-side detections can carry a fully-qualified contract id
// even when the consumer repo does not contain any `.proto` files.
//
// `import static …` is excluded by tree-sitter shape: the `name:`
// field is only present on the non-static form. `import w.x.*;` is
// also excluded for the same reason — wildcard imports have an
// `asterisk` child instead of a named identifier.
const GRPC_CLASS_IMPORT_PATTERNS = compilePatterns({
name: 'java-grpc-class-import',
language: Java,
patterns: [
{
meta: {},
query: `
(import_declaration
(scoped_identifier
scope: (_) @import_pkg
name: (identifier) @import_name (#match? @import_name "Grpc$")))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
/**
* Check whether a `class_declaration` node has a `@GrpcService`
* annotation in its modifiers list. In tree-sitter-java, class-level
@@ -118,6 +145,39 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
const out: GrpcDetection[] = [];
const emittedClassIds = new Set<number>();
// ─── Build per-file gRPC class import map ───────────────────────
// Maps `XxxGrpc` (short class name) → fully-qualified proto package
// (e.g. `cn.unipus.ucf.admin.proto.client.service`). Used below to
// tag both provider and consumer detections with a `protoPackage`
// so the orchestrator can build a fully-qualified contract id
// without depending on the current repo carrying any `.proto`
// files. This is the key fix for client-jar consumer repos.
//
// Same-short-name disambiguation: when two distinct `import` lines
// bring different `XxxGrpc` classes from different packages into
// the same file (rare for grpc — the second import would be a
// compile error in Java), the last one wins. Java's compiler
// forbids that case so we don't bother modelling it.
const grpcClassImports = new Map<string, string>();
for (const match of runCompiledPatterns(GRPC_CLASS_IMPORT_PATTERNS, tree)) {
const pkgNode = match.captures.import_pkg;
const nameNode = match.captures.import_name;
if (!pkgNode || !nameNode) continue;
grpcClassImports.set(nameNode.text, pkgNode.text);
}
/**
* Resolve the fully-qualified proto package for a short service
* name in this file. Looks up `<serviceName>Grpc` in the import
* map; returns `undefined` when the class is referenced via a
* fully-qualified name on every call site (no import line) or
* when only a wildcard import is present. The orchestrator falls
* back to the per-repo proto map in that case, preserving the
* pre-fix behaviour.
*/
const protoPackageFor = (serviceName: string): string | undefined =>
grpcClassImports.get(`${serviceName}Grpc`);
// ─── Providers: scoped form (`...Grpc.XxxImplBase`) ─────────────
for (const match of runCompiledPatterns(SCOPED_IMPL_BASE_PATTERNS, tree)) {
const classNode = match.captures.class;
@@ -127,6 +187,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
if (!serviceName) continue;
emittedClassIds.add(classNode.id);
const annotated = hasGrpcServiceAnnotation(classNode);
const protoPackage = protoPackageFor(serviceName);
out.push({
role: 'provider',
serviceName,
@@ -134,6 +195,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
source: annotated ? 'java_grpc_service' : 'java_impl_base',
confidenceWithProto: 0.8,
confidenceWithoutProto: 0.65,
...(protoPackage ? { protoPackage } : {}),
});
}
@@ -147,6 +209,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
if (!serviceName) continue;
emittedClassIds.add(classNode.id);
const annotated = hasGrpcServiceAnnotation(classNode);
const protoPackage = protoPackageFor(serviceName);
out.push({
role: 'provider',
serviceName,
@@ -154,6 +217,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
source: annotated ? 'java_grpc_service' : 'java_impl_base',
confidenceWithProto: 0.8,
confidenceWithoutProto: 0.65,
...(protoPackage ? { protoPackage } : {}),
});
}
@@ -164,6 +228,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
const grpcMatch = GRPC_SUFFIX_RE.exec(grpcClsNode.text);
if (!grpcMatch) continue;
const serviceName = grpcMatch[1];
const protoPackage = protoPackageFor(serviceName);
out.push({
role: 'consumer',
serviceName,
@@ -171,6 +236,7 @@ export const JAVA_GRPC_PLUGIN: GrpcLanguagePlugin = {
source: 'java_stub',
confidenceWithProto: 0.75,
confidenceWithoutProto: 0.55,
...(protoPackage ? { protoPackage } : {}),
});
}
@@ -86,16 +86,33 @@ const NEW_QUALIFIED_CTOR_SPEC: PatternSpec<Record<string, never>> = {
// proto loader). Matches either a bare call or an `obj.loadPackageDefinition(...)`
// call. Plugin gates the qualified-constructor consumer on this —
// structural check avoids materializing `tree.rootNode.text` for every file.
const LOAD_PACKAGE_DEFINITION_SPEC: PatternSpec<Record<string, never>> = {
meta: {},
query: `
(call_expression
function: [
(identifier) @fn (#eq? @fn "loadPackageDefinition")
(member_expression property: (property_identifier) @fn (#eq? @fn "loadPackageDefinition"))
])
`,
};
//
// These are TWO separate specs, NOT one `function: [ (identifier) ... (member_expression) ... ]`
// alternation. Under the pinned tree-sitter@0.21.1 binding a top-level alternation
// whose branches reuse the same capture name (`@fn`) collapses to one pattern with
// a shared predicate bucket; the second branch's `@fn` is left unbound and its
// `#eq?` is never enforced, so the member-expression branch would match EVERY
// `obj.method(...)` call (e.g. `console.log(...)`) — turning this gate always-on
// and emitting spurious qualified-constructor consumers. Two specs compile to two
// queries with independent predicate buckets; `runCompiledPatterns` concatenates
// their matches, so the `.length > 0` gate still means "either form is present".
const LOAD_PACKAGE_DEFINITION_SPECS: PatternSpec<Record<string, never>>[] = [
{
meta: {},
query: `
(call_expression
function: (identifier) @fn (#eq? @fn "loadPackageDefinition"))
`,
},
{
meta: {},
query: `
(call_expression
function: (member_expression
property: (property_identifier) @fn (#eq? @fn "loadPackageDefinition")))
`,
},
];
interface NodeGrpcPatternBundle {
grpcMethod: CompiledPatterns<Record<string, never>>;
@@ -107,11 +124,14 @@ interface NodeGrpcPatternBundle {
}
function compileBundle(language: unknown, name: string): NodeGrpcPatternBundle {
const mk = (spec: PatternSpec<Record<string, never>>, suffix: string) =>
const mk = (
spec: PatternSpec<Record<string, never>> | PatternSpec<Record<string, never>>[],
suffix: string,
) =>
compilePatterns({
name: `${name}-${suffix}`,
language,
patterns: [spec],
patterns: Array.isArray(spec) ? spec : [spec],
} satisfies LanguagePatterns<Record<string, never>>);
return {
grpcMethod: mk(GRPC_METHOD_SPEC, 'grpc-method'),
@@ -119,7 +139,7 @@ function compileBundle(language: unknown, name: string): NodeGrpcPatternBundle {
getService: mk(GET_SERVICE_SPEC, 'get-service'),
newSimpleCtor: mk(NEW_SIMPLE_CTOR_SPEC, 'new-simple-ctor'),
newQualifiedCtor: mk(NEW_QUALIFIED_CTOR_SPEC, 'new-qualified-ctor'),
loadPackageDefinition: mk(LOAD_PACKAGE_DEFINITION_SPEC, 'load-package-definition'),
loadPackageDefinition: mk(LOAD_PACKAGE_DEFINITION_SPECS, 'load-package-definition'),
};
}
@@ -36,6 +36,18 @@ export interface GrpcDetection {
confidenceWithProto: number;
/** Confidence when the proto map has no entry. */
confidenceWithoutProto: number;
/**
* Optional. Fully-qualified proto package the detection's service
* belongs to (e.g. `cn.unipus.ucf.admin.proto.client.service`),
* derived directly from the source file's import statements when
* available. When set, the orchestrator uses this package to build
* the contract id INSTEAD of consulting the per-repo proto map —
* letting consumer repos that don't carry `.proto` files (the
* client-jar architecture used by most Java gRPC microservices)
* still emit a fully-qualified contract id that matches the
* provider repo's contract id verbatim.
*/
protoPackage?: string;
}
/**
@@ -2,12 +2,19 @@ import * as path from 'node:path';
import { isBladeTemplateFilename } from 'gitnexus-shared';
import type { HttpLanguagePlugin } from './types.js';
import { JAVA_HTTP_PLUGIN } from './java.js';
import { KOTLIN_HTTP_PLUGIN } from './kotlin.js';
import { GO_HTTP_PLUGIN } from './go.js';
import { PYTHON_HTTP_PLUGIN } from './python.js';
import { PHP_HTTP_PLUGIN } from './php.js';
import { JAVASCRIPT_HTTP_PLUGIN, TYPESCRIPT_HTTP_PLUGIN, TSX_HTTP_PLUGIN } from './node.js';
export type { HttpDetection, HttpLanguagePlugin, HttpRole } from './types.js';
export type {
HttpDetection,
HttpFileDetections,
HttpLanguagePlugin,
HttpRole,
HttpScanInput,
} from './types.js';
/**
* File-extension → HTTP language plugin registry. The top-level
@@ -18,6 +25,11 @@ export type { HttpDetection, HttpLanguagePlugin, HttpRole } from './types.js';
* new language, drop a `http-patterns/<lang>.ts` that exports a
* `HttpLanguagePlugin`, import it here and register the extension(s).
* No edits to `http-route-extractor.ts` are required.
*
* Optional grammar plugins (e.g. `kotlin.ts`, which depends on the
* optionalDependency `tree-sitter-kotlin`) export `null` when the
* native binding is unavailable; we skip registration in that case so
* a missing optional grammar never crashes the orchestrator.
*/
const REGISTRY: Record<string, HttpLanguagePlugin> = {
'.java': JAVA_HTTP_PLUGIN,
@@ -30,16 +42,26 @@ const REGISTRY: Record<string, HttpLanguagePlugin> = {
'.tsx': TSX_HTTP_PLUGIN,
};
if (KOTLIN_HTTP_PLUGIN) {
REGISTRY['.kt'] = KOTLIN_HTTP_PLUGIN;
REGISTRY['.kts'] = KOTLIN_HTTP_PLUGIN;
}
/**
* Glob for files worth scanning for HTTP routes. Kept alongside the
* registry so adding a new language widens the glob in one edit.
*
* `.kt`/`.kts` are always present in the glob even when the optional
* `tree-sitter-kotlin` grammar isn't installed — `getPluginForFile`
* will return `undefined` for those files in that case, so the
* orchestrator simply skips them at scan time without erroring.
*
* `.vue` / `.svelte` files are intentionally omitted for the source-scan
* path — they need their own grammar-aware extraction and the existing
* regex fallback for them was never very accurate. The graph-assisted
* Strategy A still handles them via the ingestion pipeline.
*/
export const HTTP_SCAN_GLOB = '**/*.{ts,tsx,js,jsx,java,go,py,php}';
export const HTTP_SCAN_GLOB = '**/*.{ts,tsx,js,jsx,java,kt,kts,go,py,php}';
/**
* Return the HTTP plugin registered for the given file's extension,
@@ -6,19 +6,31 @@ import {
unquoteLiteral,
type LanguagePatterns,
} from '../tree-sitter-scanner.js';
import type { HttpDetection, HttpLanguagePlugin } from './types.js';
import type {
HttpDetection,
HttpFileDetections,
HttpLanguagePlugin,
HttpScanInput,
} from './types.js';
/**
* Java HTTP plugin. Handles:
* - Spring `@RequestMapping` class prefixes + `@(Get|Post|...)Mapping` method annotations
* - Spring `RestTemplate.getForObject/...`, `WebClient.method(HttpMethod.X, ...)`
* - Spring `RestTemplate.getForObject/...`, `exchange(...)`
* - Spring `WebClient.method(HttpMethod.X, ...)`, `WebClient.get().uri(...)`
* - OkHttp `new Request.Builder().url("...")`
* - OpenFeign interfaces with Spring MVC method annotations or
* native `@RequestLine("METHOD /path")` annotations
* - Java / Apache HttpClient literal request construction
*
* The plugin runs two pattern bundles: one to collect class-level
* `@RequestMapping` prefixes keyed by the enclosing class node, and a
* second to match method-level annotations. The `scan` function walks
* up from each matched annotation to find its enclosing class and
* combines the prefix with the method path.
* Every route-defining annotation (class/interface `@RequestMapping`
* prefixes, `@FeignClient(path)` prefixes, `@(Get|...)Mapping` method
* routes and native `@RequestLine`s) is matched by a single consolidated
* query (`JAVA_ROUTE_ANNOTATION_PATTERNS`) in one pass via
* `scanRouteAnnotations`. The `scan` function then walks up from each
* matched method to its enclosing class/interface to combine the prefix
* with the method path. Call-site consumers (RestTemplate, WebClient,
* OkHttp, Java/Apache HttpClient) keep their own focused queries.
*/
const METHOD_ANNOTATION_TO_HTTP: Record<string, string> = {
@@ -29,49 +41,178 @@ const METHOD_ANNOTATION_TO_HTTP: Record<string, string> = {
PatchMapping: 'PATCH',
};
// ─── Provider: Spring class-level @RequestMapping prefix ──────────────
const SPRING_CLASS_PREFIX_PATTERNS = compilePatterns({
name: 'java-spring-class-prefix',
// Each route-defining annotation has two AST shapes — a positional argument
// and a named one — that must both be matched:
// @RequestMapping("/api") → (annotation_argument_list (string_literal))
// @RequestMapping(path = "/api") → (annotation_argument_list (element_value_pair key:(identifier) value:(string_literal)))
// @RequestMapping(value = "/api") → same as above
// For named arguments only the route member keys (`path`/`value`) carry a URL;
// non-route attributes (`produces`, `consumes`, `headers`, `name`, `params`)
// would otherwise be mis-extracted (e.g. `produces = "application/json"` would
// corrupt every route). That key filtering is done in `isRouteMemberKey`, and
// all of these annotations are matched by the one `JAVA_ROUTE_ANNOTATION_PATTERNS`
// query below (see its header for why the filtering lives in JS, not the query).
interface SpringRouteBinding {
method: string;
path: string;
}
interface SpringMethodInfo {
name: string;
routes: SpringRouteBinding[];
}
interface SpringTypeInfo {
filePath: string;
kind: 'class' | 'interface';
name: string;
classPrefix: string;
implementedInterfaces: string[];
isController: boolean;
methods: SpringMethodInfo[];
}
// ─── Route-defining annotations (one generic query, one pass) ─────────
// Every Java route-mapper annotation shares one shape: an annotation carrying a
// single string argument — positional `"..."` or named `key = "..."` — on a
// class, interface, or method. This SINGLE query matches that shape generically;
// `scanRouteAnnotations` then reads the annotation NAME (`@ann`) and declaration
// kind (`@node.type`) in its for-loop to decide what each match means. Adding a
// new framework annotation that follows this single-string-argument shape is a
// change to that loop (and the lookup maps), not to this query. Annotations with
// a different argument shape — e.g. an array value `@RequestMapping({"/a","/b"})`
// — are out of scope here (as they were for the prior queries) and would need a
// new branch.
//
// Captures (shared across all branches; intentionally framework-agnostic):
// @ann → the annotation name identifier (RequestMapping, GetMapping, RequestLine, …)
// @node → the enclosing declaration (class_declaration | interface_declaration | method_declaration)
// @value → the string-literal argument
// @key → the named-argument member key (absent for the positional shape)
// @member → the method name (method_declaration branches only)
//
// The query carries NO `#eq?` / `#match?` predicates. Under the pinned
// tree-sitter 0.21.x binding a top-level `[ ... ]` alternation compiles to one
// pattern whose text predicates share a single bucket keyed by capture name, and
// a `#match?` against a capture absent from the matched branch evaluates FALSE —
// silently dropping sibling-branch matches. Keeping the query predicate-free
// sidesteps that hazard entirely; all name/key discrimination lives in the
// for-loop, where it reads as straight-line code.
const JAVA_ROUTE_ANNOTATION_PATTERNS = compilePatterns({
name: 'java-route-annotation',
language: Java,
patterns: [
{
meta: {},
query: `
(class_declaration
(modifiers
(annotation
name: (identifier) @ann (#eq? @ann "RequestMapping")
arguments: (annotation_argument_list (string_literal) @prefix)))) @class
[
(class_declaration
(modifiers
(annotation
name: (identifier) @ann
arguments: (annotation_argument_list (string_literal) @value)))) @node
(interface_declaration
(modifiers
(annotation
name: (identifier) @ann
arguments: (annotation_argument_list (string_literal) @value)))) @node
(class_declaration
(modifiers
(annotation
name: (identifier) @ann
arguments: (annotation_argument_list
(element_value_pair
key: (identifier) @key
value: (string_literal) @value))))) @node
(interface_declaration
(modifiers
(annotation
name: (identifier) @ann
arguments: (annotation_argument_list
(element_value_pair
key: (identifier) @key
value: (string_literal) @value))))) @node
(method_declaration
(modifiers
(annotation
name: (identifier) @ann
arguments: (annotation_argument_list (string_literal) @value)))
name: (identifier) @member) @node
(method_declaration
(modifiers
(annotation
name: (identifier) @ann
arguments: (annotation_argument_list
(element_value_pair
key: (identifier) @key
value: (string_literal) @value))))
name: (identifier) @member) @node
]
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Provider: Spring @(Get|Post|...)Mapping method annotations ───────
const SPRING_METHOD_ROUTE_PATTERNS = compilePatterns({
name: 'java-spring-method-route',
const SPRING_TYPE_DECLARATION_PATTERNS = compilePatterns({
name: 'java-spring-type-declaration',
language: Java,
patterns: [
{
meta: {},
query: `
(method_declaration
(modifiers
(annotation
name: (identifier) @ann (#match? @ann "^(Get|Post|Put|Delete|Patch)Mapping$")
arguments: (annotation_argument_list (string_literal) @path)))
name: (identifier) @method_name) @method
[
(class_declaration name: (identifier) @type_name) @type
(interface_declaration name: (identifier) @type_name) @type
]
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: OpenFeign `@RequestLine("METHOD /path")` parsing ───────
// OpenFeign's native annotation pairs an HTTP method and path in a single
// string literal — see https://github.com/OpenFeign/feign#interface-annotations.
// It is method-level only and is mutually exclusive with Spring MVC
// `@GetMapping` / `@PostMapping` etc. on the same method (mixing them
// requires a different Feign Contract — they are not combined). The match
// itself comes from `JAVA_ROUTE_ANNOTATION_PATTERNS`; this regex splits the
// verb from the path of the captured literal.
//
// Examples:
// @RequestLine("GET /users/{id}")
// @RequestLine("POST /users?status=active")
const REQUEST_LINE_VERB_RE = /^\s*(GET|POST|PUT|DELETE|PATCH|HEAD|OPTIONS)\s+(\S.*?)\s*$/i;
/**
* Parse a Feign `@RequestLine` value into a method + path pair.
*
* `@RequestLine("METHOD /path[?query]")` packs both fields in one string;
* the query portion is dropped because contract IDs are method+path only
* (consistent with how other consumers like RestTemplate/WebClient drop
* query strings when their values are inline literals).
*
* Returns null if the value is not a recognized HTTP verb followed by a
* path beginning with `/`.
*/
function parseRequestLine(raw: string): { method: string; path: string } | null {
const match = REQUEST_LINE_VERB_RE.exec(raw);
if (!match) return null;
const [, verb, rest] = match;
if (typeof verb !== 'string' || typeof rest !== 'string') return null;
const queryIdx = rest.indexOf('?');
const pathOnly = (queryIdx >= 0 ? rest.slice(0, queryIdx) : rest).trim();
if (!pathOnly.startsWith('/')) return null;
return { method: verb.toUpperCase(), path: pathOnly };
}
// ─── Consumer: Spring RestTemplate (object-named + method-named) ──────
// RestTemplate.getForObject / getForEntity → GET
// RestTemplate.postForObject / postForEntity → POST
// RestTemplate.put → PUT
// RestTemplate.delete → DELETE
// RestTemplate.patchForObject → PATCH
// Source-scan only: receiver must be named exactly `restTemplate`.
// Fields, `this.restTemplate`, aliases, and other injection names are deferred.
const REST_TEMPLATE_TO_HTTP: Record<string, string> = {
getForObject: 'GET',
getForEntity: 'GET',
@@ -102,22 +243,48 @@ const REST_TEMPLATE_PATTERNS = compilePatterns({
],
} satisfies LanguagePatterns<RestTemplateMeta>);
// ─── Consumer: Spring WebClient — webClient.method(HttpMethod.X, "path") ─
const WEB_CLIENT_PATTERNS = compilePatterns({
name: 'java-web-client',
const REST_TEMPLATE_EXCHANGE_PATTERNS = compilePatterns({
name: 'java-rest-template-exchange',
language: Java,
patterns: [
{
meta: { framework: 'spring-rest-template' },
query: `
(method_invocation
object: (identifier) @obj (#eq? @obj "restTemplate")
name: (identifier) @method (#eq? @method "exchange")
arguments: (argument_list
. (string_literal) @path
(field_access
object: (identifier) @httpMethodCls (#eq? @httpMethodCls "HttpMethod")
field: (identifier) @http_method)))
`,
},
],
} satisfies LanguagePatterns<RestTemplateMeta>);
const WEB_CLIENT_SHORT_TO_HTTP: Record<string, string> = {
get: 'GET',
post: 'POST',
put: 'PUT',
delete: 'DELETE',
patch: 'PATCH',
};
const WEB_CLIENT_SHORT_FORM_PATTERNS = compilePatterns({
name: 'java-web-client-short-form',
language: Java,
patterns: [
{
meta: {},
query: `
(method_invocation
object: (identifier) @obj (#eq? @obj "webClient")
name: (identifier) @method (#eq? @method "method")
arguments: (argument_list
(field_access
object: (identifier) @httpMethodCls (#eq? @httpMethodCls "HttpMethod")
field: (identifier) @http_method)
(string_literal) @path))
object: (method_invocation
object: (identifier) @obj (#eq? @obj "webClient")
name: (identifier) @verb (#match? @verb "^(get|post|put|delete|patch)$")
arguments: (argument_list))
name: (identifier) @uri_method (#eq? @uri_method "uri")
arguments: (argument_list . (string_literal) @path))
`,
},
],
@@ -144,10 +311,58 @@ const OK_HTTP_PATTERNS = compilePatterns({
],
} satisfies LanguagePatterns<Record<string, never>>);
const JAVA_HTTP_CLIENT_PATTERNS = compilePatterns({
name: 'java-http-client',
language: Java,
patterns: [
{
meta: {},
query: `
(method_invocation
object: (method_invocation
object: (method_invocation
object: (identifier) @builderCls (#eq? @builderCls "HttpRequest")
name: (identifier) @newBuilder (#eq? @newBuilder "newBuilder")
arguments: (argument_list))
name: (identifier) @uri_method (#eq? @uri_method "uri")
arguments: (argument_list
(method_invocation
object: (identifier) @uriCls (#eq? @uriCls "URI")
name: (identifier) @create (#eq? @create "create")
arguments: (argument_list . (string_literal) @path))))
name: (identifier) @http_method (#match? @http_method "^(GET|POST|PUT|DELETE)$"))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
const APACHE_HTTP_CLIENT_TO_HTTP: Record<string, string> = {
HttpGet: 'GET',
HttpPost: 'POST',
HttpPut: 'PUT',
HttpDelete: 'DELETE',
HttpPatch: 'PATCH',
};
const APACHE_HTTP_CLIENT_PATTERNS = compilePatterns({
name: 'java-apache-http-client',
language: Java,
patterns: [
{
meta: {},
query: `
(object_creation_expression
type: (type_identifier) @type (#match? @type "^Http(Get|Post|Put|Delete|Patch)$")
arguments: (argument_list . (string_literal) @path))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
/**
* Find the nearest enclosing class_declaration ancestor for a node, or
* null if the node is top-level. Tree-sitter's SyntaxNode.parent walks
* one level at a time.
* Find the nearest enclosing class/interface declaration ancestor for
* a node, or null if the node is top-level. Tree-sitter's
* SyntaxNode.parent walks one level at a time.
*/
function findEnclosingClass(node: Parser.SyntaxNode): Parser.SyntaxNode | null {
let cur: Parser.SyntaxNode | null = node.parent;
@@ -158,6 +373,15 @@ function findEnclosingClass(node: Parser.SyntaxNode): Parser.SyntaxNode | null {
return null;
}
function findEnclosingInterface(node: Parser.SyntaxNode): Parser.SyntaxNode | null {
let cur: Parser.SyntaxNode | null = node.parent;
while (cur) {
if (cur.type === 'interface_declaration') return cur;
cur = cur.parent;
}
return null;
}
/**
* Join a class-level prefix and a method-level path into a single URL
* path. Mirrors the semantics of the original regex implementation:
@@ -171,45 +395,356 @@ function joinPath(prefix: string, methodPath: string): string {
return `/${cleanPrefix}/${cleanSub}`;
}
function getNodeName(node: Parser.SyntaxNode): string | null {
return node.childForFieldName('name')?.text ?? null;
}
function hasAnnotation(node: Parser.SyntaxNode, names: string | readonly string[]): boolean {
const modifiers = node.namedChildren.find((child) => child.type === 'modifiers');
if (!modifiers) return false;
const allowed = new Set(typeof names === 'string' ? [names] : names);
const stack = [...modifiers.namedChildren];
while (stack.length > 0) {
const cur = stack.pop()!;
const annotationName = cur.childForFieldName('name')?.text ?? '';
const simpleName = annotationName.split('.').pop() ?? annotationName;
if (
(cur.type === 'annotation' || cur.type === 'marker_annotation') &&
(allowed.has(annotationName) || allowed.has(simpleName))
) {
return true;
}
stack.push(...cur.namedChildren);
}
return false;
}
/**
* A named annotation argument contributes a route only when its member key is
* `path` or `value`; a positional argument (no key node) always qualifies.
* This is the JS-side replacement for the in-query `^(path|value)$` filter and
* drops Spring's non-route string attributes (`produces`, `consumes`,
* `headers`, `name`, `params`) that would otherwise be mis-read as routes.
*/
function isRouteMemberKey(keyNode: Parser.SyntaxNode | undefined): boolean {
if (!keyNode) return true;
return keyNode.text === 'path' || keyNode.text === 'value';
}
interface MethodRouteAnnotation {
methodNode: Parser.SyntaxNode;
methodName: string | null;
httpMethod: string;
rawPath: string;
}
interface RequestLineAnnotation {
methodNode: Parser.SyntaxNode;
methodName: string | null;
parsed: { method: string; path: string };
}
interface RouteAnnotationScan {
/** Spring `@RequestMapping` URL prefix per class/interface node id (last write wins). */
prefixByTypeId: Map<number, string>;
/** OpenFeign interface prefix per interface node id; `@FeignClient(path)` wins over `@RequestMapping`. */
feignPrefixByInterfaceId: Map<number, string>;
/** One entry per resolved Spring `@(Get|...)Mapping` route — a method with N mappings yields N entries. */
methodRoutes: MethodRouteAnnotation[];
/** One entry per OpenFeign `@RequestLine` whose value parses to a verb + path. */
requestLines: RequestLineAnnotation[];
}
/**
* Resolve every Java route-defining annotation in a single tree-sitter pass.
*
* The generic `JAVA_ROUTE_ANNOTATION_PATTERNS` query yields one match per
* annotation-carrying-a-string-argument on any class / interface / method. This
* loop reads the annotation name and declaration kind to decide what each match
* means, ignoring annotations it does not recognise. The HTTP verb map
* (`METHOD_ANNOTATION_TO_HTTP`) and the `path`/`value` key filter
* (`isRouteMemberKey`) live here rather than in the query (see its header).
*/
function scanRouteAnnotations(tree: Parser.Tree): RouteAnnotationScan {
const matches = runCompiledPatterns(JAVA_ROUTE_ANNOTATION_PATTERNS, tree);
// The two prefix maps intentionally diverge for the same interface node:
// `prefixByTypeId` feeds the Spring *provider* path (class prefix +
// collectSpringTypes cross-file inheritance), while `feignPrefixByInterfaceId`
// feeds the OpenFeign *consumer* path in scan(). An interface carrying both
// `@RequestMapping` and `@FeignClient(path)` lands a different value in each.
const prefixByTypeId = new Map<number, string>();
const feignPrefixByInterfaceId = new Map<number, string>();
const methodRoutes: MethodRouteAnnotation[] = [];
const requestLines: RequestLineAnnotation[] = [];
// Interface `@RequestMapping` prefixes rank below `@FeignClient(path)`;
// collect them and apply only after the FeignClient pass below.
const interfaceRequestMappingPrefixes: Array<{ id: number; prefix: string }> = [];
for (const { captures } of matches) {
const annNode = captures.ann;
const node = captures.node;
const valueNode = captures.value;
if (!annNode || !node || !valueNode) continue;
const ann = annNode.text;
const keyNode = captures.key; // undefined for the positional shape
if (node.type === 'method_declaration') {
// Method-level: a Spring `@(Get|...)Mapping` route, or native `@RequestLine`.
const httpMethod = METHOD_ANNOTATION_TO_HTTP[ann];
if (httpMethod) {
if (!isRouteMemberKey(keyNode)) continue;
const rawPath = unquoteLiteral(valueNode.text);
if (rawPath !== null) {
methodRoutes.push({
methodNode: node,
methodName: captures.member?.text ?? null,
httpMethod,
rawPath,
});
}
} else if (ann === 'RequestLine') {
// Feign packs verb + path in one literal; its only named argument is `value`.
if (keyNode && keyNode.text !== 'value') continue;
const raw = unquoteLiteral(valueNode.text);
const parsed = raw !== null ? parseRequestLine(raw) : null;
if (parsed) {
requestLines.push({
methodNode: node,
methodName: captures.member?.text ?? null,
parsed,
});
}
}
continue;
}
// Type-level (class or interface): a Spring `@RequestMapping` URL prefix, or
// — on an interface — an OpenFeign `@FeignClient(path = "...")` prefix.
if (ann === 'RequestMapping') {
if (!isRouteMemberKey(keyNode)) continue;
const prefix = unquoteLiteral(valueNode.text);
if (prefix !== null) {
prefixByTypeId.set(node.id, prefix);
if (node.type === 'interface_declaration') {
interfaceRequestMappingPrefixes.push({ id: node.id, prefix });
}
}
} else if (ann === 'FeignClient' && node.type === 'interface_declaration') {
// Feign's `name`/`value` identify a service, not a path — only `path` is a prefix.
if (!keyNode || keyNode.text !== 'path') continue;
const prefix = unquoteLiteral(valueNode.text);
if (prefix !== null && !feignPrefixByInterfaceId.has(node.id)) {
feignPrefixByInterfaceId.set(node.id, prefix);
}
}
}
for (const { id, prefix } of interfaceRequestMappingPrefixes) {
if (!feignPrefixByInterfaceId.has(id)) feignPrefixByInterfaceId.set(id, prefix);
}
return { prefixByTypeId, feignPrefixByInterfaceId, methodRoutes, requestLines };
}
function collectDirectMethods(typeNode: Parser.SyntaxNode): Parser.SyntaxNode[] {
const out: Parser.SyntaxNode[] = [];
const visit = (node: Parser.SyntaxNode): void => {
for (const child of node.namedChildren) {
if (child.type === 'method_declaration') {
out.push(child);
continue;
}
if (
child !== typeNode &&
(child.type === 'class_declaration' || child.type === 'interface_declaration')
) {
continue;
}
visit(child);
}
};
visit(typeNode);
return out;
}
function collectImplementedInterfaces(typeNode: Parser.SyntaxNode): string[] {
const interfacesNode = typeNode.childForFieldName('interfaces');
if (!interfacesNode) return [];
const out: string[] = [];
const visit = (node: Parser.SyntaxNode): void => {
if (node.type === 'type_identifier' || node.type === 'scoped_type_identifier') {
out.push(node.text.split('.').pop() ?? node.text);
return;
}
for (const child of node.namedChildren) visit(child);
};
visit(interfacesNode);
return out;
}
function collectSpringTypes(filePath: string, tree: Parser.Tree): SpringTypeInfo[] {
const { prefixByTypeId, methodRoutes } = scanRouteAnnotations(tree);
const routesByMethodId = new Map<number, SpringRouteBinding[]>();
for (const route of methodRoutes) {
const routes = routesByMethodId.get(route.methodNode.id) ?? [];
routes.push({ method: route.httpMethod, path: route.rawPath });
routesByMethodId.set(route.methodNode.id, routes);
}
const out: SpringTypeInfo[] = [];
for (const match of runCompiledPatterns(SPRING_TYPE_DECLARATION_PATTERNS, tree)) {
const typeNode = match.captures.type;
const typeNameNode = match.captures.type_name;
if (!typeNode || !typeNameNode) continue;
const kind = typeNode.type === 'interface_declaration' ? 'interface' : 'class';
const methods = collectDirectMethods(typeNode)
.map((methodNode) => ({
name: getNodeName(methodNode),
routes: routesByMethodId.get(methodNode.id) ?? [],
}))
.filter((method): method is SpringMethodInfo => method.name !== null);
out.push({
filePath,
kind,
name: typeNameNode.text,
classPrefix: prefixByTypeId.get(typeNode.id) ?? '',
implementedInterfaces: kind === 'class' ? collectImplementedInterfaces(typeNode) : [],
isController: kind === 'class' && hasAnnotation(typeNode, ['RestController', 'Controller']),
methods,
});
}
return out;
}
function scanSpringProject(files: readonly HttpScanInput[]): HttpFileDetections[] {
const types = files.flatMap((file) => collectSpringTypes(file.filePath, file.tree));
const interfaceRoutes = new Map<string, Map<string, SpringRouteBinding[]> | null>();
for (const type of types) {
if (type.kind !== 'interface') continue;
if (interfaceRoutes.has(type.name)) {
interfaceRoutes.set(type.name, null);
continue;
}
const methodMap = new Map<string, SpringRouteBinding[]>();
for (const method of type.methods) {
const routes = method.routes.map((route) => ({
method: route.method,
path: type.classPrefix ? joinPath(type.classPrefix, route.path) : route.path,
}));
if (routes.length > 0) methodMap.set(method.name, routes);
}
interfaceRoutes.set(type.name, methodMap);
}
const detectionsByFile = new Map<string, HttpDetection[]>();
for (const type of types) {
if (type.kind !== 'class' || !type.isController) continue;
for (const method of type.methods) {
if (method.routes.length > 0) continue;
const inheritedRoutes = type.implementedInterfaces.flatMap((interfaceName) => {
const routeMap = interfaceRoutes.get(interfaceName);
if (!routeMap) return [];
const routes = routeMap.get(method.name) ?? [];
return routes.map((route) => ({
method: route.method,
path: joinPath(type.classPrefix, route.path),
}));
});
for (const route of inheritedRoutes) {
const detections = detectionsByFile.get(type.filePath) ?? [];
detections.push({
role: 'provider',
framework: 'spring',
method: route.method,
path: route.path,
name: method.name,
confidence: 0.8,
});
detectionsByFile.set(type.filePath, detections);
}
}
}
return [...detectionsByFile.entries()].map(([filePath, detections]) => ({
filePath,
detections,
}));
}
export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
name: 'java-http',
language: Java,
scan(tree) {
const out: HttpDetection[] = [];
// ─── Providers: Spring class prefix + method annotations ────────
const prefixByClassId = new Map<number, string>();
for (const match of runCompiledPatterns(SPRING_CLASS_PREFIX_PATTERNS, tree)) {
const prefixNode = match.captures.prefix;
const classNode = match.captures.class;
if (!prefixNode || !classNode) continue;
const prefix = unquoteLiteral(prefixNode.text);
if (prefix !== null) prefixByClassId.set(classNode.id, prefix);
}
// ─── Spring providers + OpenFeign consumers (one query pass) ────
// `scanRouteAnnotations` resolves every route-defining annotation —
// class/interface prefixes, method `@(Get|...)Mapping`s and native
// `@RequestLine`s — from a single `matches()` pass over the tree.
const { prefixByTypeId, feignPrefixByInterfaceId, methodRoutes, requestLines } =
scanRouteAnnotations(tree);
for (const match of runCompiledPatterns(SPRING_METHOD_ROUTE_PATTERNS, tree)) {
const annNode = match.captures.ann;
const pathNode = match.captures.path;
const nameNode = match.captures.method_name;
const methodNode = match.captures.method;
if (!annNode || !pathNode || !methodNode) continue;
const httpMethod = METHOD_ANNOTATION_TO_HTTP[annNode.text];
if (!httpMethod) continue;
const rawPath = unquoteLiteral(pathNode.text);
if (rawPath === null) continue;
const enclosingClass = findEnclosingClass(methodNode);
const prefix = enclosingClass ? (prefixByClassId.get(enclosingClass.id) ?? '') : '';
const fullPath = joinPath(prefix, rawPath);
// A `@(Get|...)Mapping` inside a `@FeignClient` interface is an OpenFeign
// *consumer* (it describes a remote call); the same annotation inside a
// class is a Spring *provider*. A mapping on a non-Feign interface has no
// enclosing class and is dropped here — interface→controller inheritance is
// handled by `scanProject`.
for (const route of methodRoutes) {
const enclosingInterface = findEnclosingInterface(route.methodNode);
if (enclosingInterface && hasAnnotation(enclosingInterface, 'FeignClient')) {
const prefix = feignPrefixByInterfaceId.get(enclosingInterface.id) ?? '';
out.push({
role: 'consumer',
framework: 'openfeign',
method: route.httpMethod,
path: joinPath(prefix, route.rawPath),
name: route.methodName,
confidence: 0.7,
});
continue;
}
const enclosingClass = findEnclosingClass(route.methodNode);
if (!enclosingClass) continue;
const prefix = prefixByTypeId.get(enclosingClass.id) ?? '';
out.push({
role: 'provider',
framework: 'spring',
method: httpMethod,
path: fullPath,
name: nameNode?.text ?? null,
method: route.httpMethod,
path: joinPath(prefix, route.rawPath),
name: route.methodName,
confidence: 0.8,
});
}
// Native OpenFeign `@RequestLine("METHOD /path")`. Method-level only and
// always declared on an interface (Feign builds a proxy from the interface).
// We do NOT require an enclosing `@FeignClient`: `@RequestLine` is a core
// `feign.*` annotation used with `Feign.builder()`, whereas `@FeignClient`
// is the Spring Cloud variant that uses Spring MVC annotations instead — the
// two are effectively mutually exclusive, so requiring `@FeignClient` here
// would miss the annotation's primary use. The `RequestLine` name is itself
// a strong, framework-specific signal, so a structural interface check is
// enough to keep false positives away. A `@FeignClient(path=...)` prefix is
// still applied when present (rare, but harmless).
for (const requestLine of requestLines) {
const enclosingInterface = findEnclosingInterface(requestLine.methodNode);
if (!enclosingInterface) continue;
const prefix = feignPrefixByInterfaceId.get(enclosingInterface.id) ?? '';
out.push({
role: 'consumer',
framework: 'openfeign',
method: requestLine.parsed.method,
path: joinPath(prefix, requestLine.parsed.path),
name: requestLine.methodName,
confidence: 0.75,
});
}
// ─── Consumers: RestTemplate ────────────────────────────────────
for (const match of runCompiledPatterns(REST_TEMPLATE_PATTERNS, tree)) {
const methodNode = match.captures.method;
@@ -229,8 +764,7 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
});
}
// ─── Consumers: WebClient.method(HttpMethod.X, "path") ──────────
for (const match of runCompiledPatterns(WEB_CLIENT_PATTERNS, tree)) {
for (const match of runCompiledPatterns(REST_TEMPLATE_EXCHANGE_PATTERNS, tree)) {
const httpMethodNode = match.captures.http_method;
const pathNode = match.captures.path;
if (!httpMethodNode || !pathNode) continue;
@@ -238,7 +772,7 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'spring-web-client',
framework: 'spring-rest-template',
method: httpMethodNode.text.toUpperCase(),
path,
name: null,
@@ -246,6 +780,28 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
});
}
// ─── Consumers: WebClient.get().uri("path") short form ─────────
// Source-scan only: receiver must be named exactly `webClient`.
// The real long-form chain `webClient.method(HttpMethod.X).uri("/x")`
// needs multi-hop chain analysis and is intentionally deferred.
for (const match of runCompiledPatterns(WEB_CLIENT_SHORT_FORM_PATTERNS, tree)) {
const verbNode = match.captures.verb;
const pathNode = match.captures.path;
if (!verbNode || !pathNode) continue;
const httpMethod = WEB_CLIENT_SHORT_TO_HTTP[verbNode.text];
if (!httpMethod) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'spring-web-client',
method: httpMethod,
path,
name: null,
confidence: 0.7,
});
}
// ─── Consumers: OkHttp Request.Builder().url("path") ────────────
for (const match of runCompiledPatterns(OK_HTTP_PATTERNS, tree)) {
const pathNode = match.captures.path;
@@ -262,6 +818,45 @@ export const JAVA_HTTP_PLUGIN: HttpLanguagePlugin = {
});
}
// ─── Consumers: Java HttpClient request builder ─────────────────
// Java's builder exposes GET/POST/PUT/DELETE helpers. PATCH uses
// `.method("PATCH", body)`, which is intentionally deferred.
for (const match of runCompiledPatterns(JAVA_HTTP_CLIENT_PATTERNS, tree)) {
const httpMethodNode = match.captures.http_method;
const pathNode = match.captures.path;
if (!httpMethodNode || !pathNode) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'java-http-client',
method: httpMethodNode.text.toUpperCase(),
path,
name: null,
confidence: 0.65,
});
}
// ─── Consumers: Apache HttpClient request constructors ──────────
for (const match of runCompiledPatterns(APACHE_HTTP_CLIENT_PATTERNS, tree)) {
const typeNode = match.captures.type;
const pathNode = match.captures.path;
if (!typeNode || !pathNode) continue;
const httpMethod = APACHE_HTTP_CLIENT_TO_HTTP[typeNode.text];
if (!httpMethod) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'apache-http-client',
method: httpMethod,
path,
name: null,
confidence: 0.65,
});
}
return out;
},
scanProject: scanSpringProject,
};
@@ -0,0 +1,564 @@
import Parser from 'tree-sitter';
import { createRequire } from 'node:module';
import {
compilePatterns,
runCompiledPatterns,
unquoteLiteral,
type LanguagePatterns,
} from '../tree-sitter-scanner.js';
import type { HttpDetection, HttpLanguagePlugin } from './types.js';
/**
* Kotlin HTTP plugin (Spring providers + consumers).
*
* **Providers** (#1849) — Spring `@RequestMapping` class prefixes and
* `@(Get|Post|...)Mapping` method annotations on Kotlin Spring Boot
* controllers. Both positional shorthand (`@GetMapping("/x")`) and
* named annotation arguments (`@GetMapping(value = "/x")` and
* `@GetMapping(path = "/x")`) are supported.
*
* **Consumers** — four call-site patterns common in Kotlin
* Spring projects:
*
* 1. `restTemplate.getForObject("/x", ...)` and friends (#1855)
* 2. `webClient.get().uri("/x")` — short form (#1855)
* 3. `Request.Builder().url("/x")` — OkHttp (#1855)
* 4. `webClient.method(HttpMethod.X).uri("/y")` — long form (this PR)
*
* The long form puts the verb on a sibling `call_expression` two hops
* away from the path. Rather than introducing imperative walk-up logic,
* we use a single deeper tree-sitter query that matches the full chain
* structurally — see `WEB_CLIENT_LONG_PATTERNS` below. The verb is
* captured directly as the `simple_identifier` of `HttpMethod.X`, so
* variable-bound verbs (`val verb = HttpMethod.PATCH; webClient.method(verb)...`)
* are intentionally NOT picked up — those need a graph-aware resolver
* and are out of scope for source-scan.
*
* tree-sitter-kotlin (fwcd) AST shapes used here:
* class_declaration
* modifiers
* annotation
* constructor_invocation
* user_type → type_identifier ← annotation name
* value_arguments
* value_argument
* (simple_identifier "=")? ← absent for positional, present for named
* string_literal
* type_identifier ← class name
*
* Consumer call shape (Kotlin chains everything via `navigation_expression`):
* call_expression ← outer `.uri("/x")` or `.url("/x")`
* navigation_expression
* call_expression ← inner `.get()` / `Request.Builder()` / `restTemplate.x`
* navigation_expression
* simple_identifier ← receiver: `webClient` / `Request` / `restTemplate`
* navigation_suffix ← `.method` / `.Builder` / `.getForObject`
* call_suffix (value_arguments)
* navigation_suffix ← `.uri` / `.url`
* call_suffix
* value_arguments
* value_argument
* string_literal ← the path
*
* tree-sitter-kotlin is an optional npm dependency — when its native
* binding is unavailable the plugin gracefully exports `null` and
* `http-patterns/index.ts` skips registration for `.kt`/`.kts` files.
*/
const _require = createRequire(import.meta.url);
/** Loaded lazily; null when the grammar binding isn't installed. */
let Kotlin: unknown | null = null;
try {
Kotlin = _require('tree-sitter-kotlin');
} catch {
Kotlin = null;
}
const METHOD_ANNOTATION_TO_HTTP: Record<string, string> = {
GetMapping: 'GET',
PostMapping: 'POST',
PutMapping: 'PUT',
DeleteMapping: 'DELETE',
PatchMapping: 'PATCH',
};
/**
* RestTemplate method-name → HTTP verb. Mirrors the Java plugin's
* `REST_TEMPLATE_TO_HTTP` (java.ts) so a polyglot repo emits the
* same contract IDs from .java and .kt sources.
*/
const REST_TEMPLATE_TO_HTTP: Record<string, string> = {
getForObject: 'GET',
getForEntity: 'GET',
postForObject: 'POST',
postForEntity: 'POST',
put: 'PUT',
delete: 'DELETE',
patchForObject: 'PATCH',
};
/**
* WebClient short-form verb → HTTP verb. The reactive WebClient API
* exposes `.get()`, `.post()`, `.put()`, `.delete()`, `.patch()` as
* one-liners that return a `RequestHeadersUriSpec` whose `.uri(...)`
* carries the path. We capture both pieces in a single query (see
* `WEB_CLIENT_SHORT_PATTERNS` below) and translate the verb here.
*/
const WEB_CLIENT_SHORT_TO_HTTP: Record<string, string> = {
get: 'GET',
post: 'POST',
put: 'PUT',
delete: 'DELETE',
patch: 'PATCH',
};
/**
* Allowed HTTP verbs for the WebClient long-form path
* `webClient.method(HttpMethod.X).uri("/y")`. Compiled once at module
* load (instead of inside the scan loop) per maintainer feedback on
* PR #1884. Mirrors the keys of `WEB_CLIENT_SHORT_TO_HTTP` above —
* keeping HEAD/OPTIONS/TRACE intentionally excluded for symmetry
* with the short form and the Java plugin.
*/
const WEB_CLIENT_LONG_VERB_RE = /^(GET|POST|PUT|DELETE|PATCH)$/;
/**
* Build the plugin only if the Kotlin grammar is available. Compiling
* the queries against a null grammar would throw at module load time
* and abort the whole http-route-extractor module.
*/
function buildKotlinPlugin(language: unknown): HttpLanguagePlugin {
// ─── Provider: Spring class-level @RequestMapping prefix ──────────────
// Two patterns mirror the Java plugin's positional vs named split:
// @RequestMapping("/api") → value_argument has string_literal as its first named child
// @RequestMapping(path = "/api") → value_argument has [simple_identifier @key, string_literal]
// @RequestMapping(value = "/api") → same as above, with key="value"
//
// Tree-sitter-kotlin grammar (fwcd 0.3.8) does NOT have a separate
// node for named arguments — both positional and named forms share
// `value_argument`. The positional pattern uses the immediate-child
// anchor `.` so it only matches when the string_literal is the FIRST
// named child (i.e. no preceding simple_identifier "=" prefix). The
// named pattern explicitly captures the simple_identifier and uses
// `#match?` to restrict it to `path`/`value`, matching the same
// safety bar that the Java plugin enforces (see java.ts and the
// sibling topic-patterns/java.ts for the analogous constraint).
//
// Without the `key:` constraint the named query would also capture
// unrelated attributes like `produces`, `consumes`, `headers`,
// `name`, `params` — emitting bogus route contracts (a regression
// identical to the one Claude flagged on PR #1834 for Java).
const SPRING_CLASS_PREFIX_PATTERNS = compilePatterns({
name: 'kotlin-spring-class-prefix',
language,
patterns: [
{
meta: {},
query: `
(class_declaration
(modifiers
(annotation
(constructor_invocation
(user_type (type_identifier) @ann (#eq? @ann "RequestMapping"))
(value_arguments
(value_argument . (string_literal) @prefix)))))
(type_identifier) @cls) @class
`,
},
{
meta: {},
query: `
(class_declaration
(modifiers
(annotation
(constructor_invocation
(user_type (type_identifier) @ann (#eq? @ann "RequestMapping"))
(value_arguments
(value_argument
(simple_identifier) @key (#match? @key "^(path|value)$")
(string_literal) @prefix)))))
(type_identifier) @cls) @class
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Provider: Spring @(Get|Post|...)Mapping method annotations ───────
// Same dual-pattern positional/named approach. The Kotlin AST puts the
// function name (`simple_identifier`) outside the `modifiers` subtree,
// so we capture it from `function_declaration` directly.
const SPRING_METHOD_ROUTE_PATTERNS = compilePatterns({
name: 'kotlin-spring-method-route',
language,
patterns: [
{
meta: {},
query: `
(function_declaration
(modifiers
(annotation
(constructor_invocation
(user_type (type_identifier) @ann (#match? @ann "^(Get|Post|Put|Delete|Patch)Mapping$"))
(value_arguments
(value_argument . (string_literal) @path)))))
(simple_identifier) @method_name) @method
`,
},
{
meta: {},
query: `
(function_declaration
(modifiers
(annotation
(constructor_invocation
(user_type (type_identifier) @ann (#match? @ann "^(Get|Post|Put|Delete|Patch)Mapping$"))
(value_arguments
(value_argument
(simple_identifier) @key (#match? @key "^(path|value)$")
(string_literal) @path)))))
(simple_identifier) @method_name) @method
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: Spring RestTemplate ────────────────────────────────────
// Kotlin call-site shape mirrors the Java plugin's
// `REST_TEMPLATE_PATTERNS`, but goes through tree-sitter-kotlin's
// `navigation_expression` instead of Java's `method_invocation`:
//
// restTemplate.getForObject("/x", User::class.java)
//
// becomes
//
// call_expression
// navigation_expression
// simple_identifier "restTemplate"
// navigation_suffix → simple_identifier "getForObject"
// call_suffix
// value_arguments
// value_argument . string_literal "/x" ← captured
// value_argument User::class.java
//
// The receiver name is constrained to `restTemplate` (#eq? @obj),
// matching the Java plugin's heuristic. This means a non-conventional
// field name (e.g. `userServiceTemplate`) will not be picked up;
// that's the same trade-off already accepted on the Java side.
const REST_TEMPLATE_PATTERNS = compilePatterns({
name: 'kotlin-rest-template',
language,
patterns: [
{
meta: {},
query: `
(call_expression
(navigation_expression
(simple_identifier) @obj (#eq? @obj "restTemplate")
(navigation_suffix (simple_identifier) @method))
(call_suffix
(value_arguments . (value_argument . (string_literal) @path))))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: Spring WebClient (short form) ──────────────────────────
// Reactive WebClient exposes one-liner verb helpers:
//
// webClient.get().uri("/x").retrieve().awaitBody<T>()
// webClient.post().uri("/x")...
//
// The chain `webClient.get().uri("/x")` parses as two nested
// `call_expression` nodes — the OUTER call is `.uri("/x")` and the
// INNER call is `webClient.get()`. We anchor on the outer call and
// require:
// - inner receiver is `webClient`
// - inner suffix is one of the HTTP verbs (#match?)
// - outer suffix is exactly `uri`
// - outer call's first value_argument is a string literal
//
// The long-form `webClient.method(HttpMethod.GET).uri("/x")` chain
// uses an extra navigation hop and an enum field access — handled
// by `WEB_CLIENT_LONG_PATTERNS` below, separately so each query is
// straightforward to reason about.
const WEB_CLIENT_SHORT_PATTERNS = compilePatterns({
name: 'kotlin-web-client-short',
language,
patterns: [
{
meta: {},
query: `
(call_expression
(navigation_expression
(call_expression
(navigation_expression
(simple_identifier) @obj (#eq? @obj "webClient")
(navigation_suffix
(simple_identifier) @verb (#match? @verb "^(get|post|put|delete|patch)$")))
(call_suffix (value_arguments)))
(navigation_suffix (simple_identifier) @uri (#eq? @uri "uri")))
(call_suffix
(value_arguments . (value_argument . (string_literal) @path))))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: Spring WebClient (long form) ───────────────────────────
// The fluent long form passes the verb as a `HttpMethod.X` enum field
// access through `.method(...)`, then carries the path on a separate
// `.uri(...)` hop further down the chain:
//
// webClient.method(HttpMethod.GET).uri("/x").retrieve().awaitBody<T>()
//
// Compared to the short form there are two extra structural hops:
// - the inner `.method(...)` `call_expression` has a `value_argument`
// whose payload is itself a `navigation_expression` (HttpMethod → .GET)
// - the outer `.uri(...)` is reached via one more
// `navigation_expression` wrapping that inner call
//
// We capture the verb at the `simple_identifier` under `HttpMethod`'s
// `navigation_suffix`. That `simple_identifier` is the literal field
// name (`GET`, `POST`, ...) used in source — Kotlin enum fields by
// convention are upper-case, matching `HttpMethod` from
// `org.springframework.http`. We forward the captured text as-is.
//
// Variable-bound verbs (`val verb = HttpMethod.PATCH; webClient.method(verb)...`)
// do NOT match — they fail the `(navigation_expression ...)` shape
// because the value_argument carries a bare `simple_identifier` instead
// of a `HttpMethod.X` field access. This is intentional: source-scan
// can't follow the binding without graph context. Pinned by an
// anti-overreach test in the consumer suite.
const WEB_CLIENT_LONG_PATTERNS = compilePatterns({
name: 'kotlin-web-client-long',
language,
patterns: [
{
meta: {},
query: `
(call_expression
(navigation_expression
(call_expression
(navigation_expression
(simple_identifier) @obj (#eq? @obj "webClient")
(navigation_suffix
(simple_identifier) @method_call (#eq? @method_call "method")))
(call_suffix
(value_arguments
. (value_argument
(navigation_expression
(simple_identifier) @httpMethodCls (#eq? @httpMethodCls "HttpMethod")
(navigation_suffix (simple_identifier) @verb))))))
(navigation_suffix (simple_identifier) @uri (#eq? @uri "uri")))
(call_suffix
(value_arguments . (value_argument . (string_literal) @path))))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: OkHttp Request.Builder().url("/x") ─────────────────────
// Kotlin parses `Request.Builder()` as a `call_expression` whose
// callee is a `navigation_expression` (Request → .Builder), NOT as
// Java's `object_creation_expression`. The chain `.url("/x")` then
// wraps that in another `call_expression`. The query mirrors Java's
// `OK_HTTP_PATTERNS` (java.ts) but adapts the node types.
//
// Receiver `Request` is constrained by name (#eq? @cls); a project
// that imports OkHttp's `Request` under an alias (`import okhttp3.Request as OkRequest`)
// would not be picked up — this matches the Java plugin's heuristic.
//
// **Known limitation — verb defaults to GET.** OkHttp encodes the
// verb on a *sibling* call further down the builder chain (e.g.
// `.post(body)` / `.get()` / `.delete()`), not on `.url(...)` itself.
// This query intentionally does not walk the chain to recover the
// verb — it emits `method: 'GET'` for every match, mirroring
// `java.ts:OK_HTTP_PATTERNS`. So a `Request.Builder().url("/x").post(body).build()`
// call becomes `http::GET::/x`, not `http::POST::/x`. This is the
// same trade-off Java has accepted; pinned by an anti-overreach
// test in `http-route-extractor.test.ts` so a future verb-walk
// implementation has to update this comment in lockstep.
const OK_HTTP_PATTERNS = compilePatterns({
name: 'kotlin-okhttp',
language,
patterns: [
{
meta: {},
query: `
(call_expression
(navigation_expression
(call_expression
(navigation_expression
(simple_identifier) @cls (#eq? @cls "Request")
(navigation_suffix (simple_identifier) @builder (#eq? @builder "Builder")))
(call_suffix (value_arguments)))
(navigation_suffix (simple_identifier) @method (#eq? @method "url")))
(call_suffix
(value_arguments . (value_argument . (string_literal) @path))))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
/**
* Find the nearest enclosing class_declaration ancestor for a node, or
* null if the node is top-level. Mirrors the Java plugin's helper.
*/
function findEnclosingClass(node: Parser.SyntaxNode): Parser.SyntaxNode | null {
let cur: Parser.SyntaxNode | null = node.parent;
while (cur) {
if (cur.type === 'class_declaration') return cur;
cur = cur.parent;
}
return null;
}
/**
* Join a class-level prefix and a method-level path. Identical
* semantics to the Java plugin: strip leading/trailing slashes on
* the prefix, strip leading slashes on the method path, ensure a
* single slash between them.
*/
function joinPath(prefix: string, methodPath: string): string {
const cleanPrefix = prefix.replace(/^\/+/, '').replace(/\/+$/, '');
const cleanSub = methodPath.replace(/^\/+/, '');
if (!cleanPrefix) return `/${cleanSub}`;
return `/${cleanPrefix}/${cleanSub}`;
}
return {
name: 'kotlin-http',
language,
scan(tree) {
const out: HttpDetection[] = [];
// ─── Class prefixes ─────────────────────────────────────────────
const prefixByClassId = new Map<number, string>();
for (const match of runCompiledPatterns(SPRING_CLASS_PREFIX_PATTERNS, tree)) {
const prefixNode = match.captures.prefix;
const classNode = match.captures.class;
if (!prefixNode || !classNode) continue;
const prefix = unquoteLiteral(prefixNode.text);
if (prefix !== null) prefixByClassId.set(classNode.id, prefix);
}
// ─── Method routes ──────────────────────────────────────────────
for (const match of runCompiledPatterns(SPRING_METHOD_ROUTE_PATTERNS, tree)) {
const annNode = match.captures.ann;
const pathNode = match.captures.path;
const nameNode = match.captures.method_name;
const methodNode = match.captures.method;
if (!annNode || !pathNode || !methodNode) continue;
const httpMethod = METHOD_ANNOTATION_TO_HTTP[annNode.text];
if (!httpMethod) continue;
const rawPath = unquoteLiteral(pathNode.text);
if (rawPath === null) continue;
const enclosingClass = findEnclosingClass(methodNode);
const prefix = enclosingClass ? (prefixByClassId.get(enclosingClass.id) ?? '') : '';
const fullPath = joinPath(prefix, rawPath);
out.push({
role: 'provider',
framework: 'spring',
method: httpMethod,
path: fullPath,
name: nameNode?.text ?? null,
confidence: 0.8,
});
}
// ─── Consumers: RestTemplate ────────────────────────────────────
for (const match of runCompiledPatterns(REST_TEMPLATE_PATTERNS, tree)) {
const methodNode = match.captures.method;
const pathNode = match.captures.path;
if (!methodNode || !pathNode) continue;
const httpMethod = REST_TEMPLATE_TO_HTTP[methodNode.text];
if (!httpMethod) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'spring-rest-template',
method: httpMethod,
path,
name: null,
confidence: 0.7,
});
}
// ─── Consumers: WebClient short form (.get()/.post()/etc → .uri) ─
for (const match of runCompiledPatterns(WEB_CLIENT_SHORT_PATTERNS, tree)) {
const verbNode = match.captures.verb;
const pathNode = match.captures.path;
if (!verbNode || !pathNode) continue;
const httpMethod = WEB_CLIENT_SHORT_TO_HTTP[verbNode.text];
if (!httpMethod) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'spring-web-client',
method: httpMethod,
path,
name: null,
confidence: 0.7,
});
}
// ─── Consumers: WebClient long form (.method(HttpMethod.X) → .uri) ─
for (const match of runCompiledPatterns(WEB_CLIENT_LONG_PATTERNS, tree)) {
const verbNode = match.captures.verb;
const pathNode = match.captures.path;
if (!verbNode || !pathNode) continue;
// The captured text is the literal `HttpMethod.X` field name.
// Spring's `org.springframework.http.HttpMethod` defines GET,
// POST, PUT, DELETE, PATCH, HEAD, OPTIONS, TRACE — we only
// emit for the five verbs we already handle elsewhere, so
// exotic ones are silently skipped (consistent with the
// short form's WEB_CLIENT_SHORT_TO_HTTP guard). The accepted
// verb regex is hoisted to module scope (see
// `WEB_CLIENT_LONG_VERB_RE` near the top of this file).
const verbText = verbNode.text;
if (!WEB_CLIENT_LONG_VERB_RE.test(verbText)) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'spring-web-client',
method: verbText,
path,
name: null,
confidence: 0.7,
});
}
// ─── Consumers: OkHttp Request.Builder().url("path") ────────────
for (const match of runCompiledPatterns(OK_HTTP_PATTERNS, tree)) {
const pathNode = match.captures.path;
if (!pathNode) continue;
const path = unquoteLiteral(pathNode.text);
if (path === null) continue;
out.push({
role: 'consumer',
framework: 'okhttp',
method: 'GET',
path,
name: null,
confidence: 0.7,
});
}
return out;
},
};
}
/**
* The exported plugin is `null` when tree-sitter-kotlin's native
* binding is unavailable. `http-patterns/index.ts` checks for null
* before registering `.kt`/`.kts` so missing optional grammars never
* crash the orchestrator.
*/
export const KOTLIN_HTTP_PLUGIN: HttpLanguagePlugin | null = Kotlin
? buildKotlinPlugin(Kotlin)
: null;
@@ -6,7 +6,7 @@ import {
unquoteLiteral,
type LanguagePatterns,
} from '../tree-sitter-scanner.js';
import type { HttpDetection, HttpLanguagePlugin } from './types.js';
import type { HttpDetection, HttpLanguagePlugin, RepoContext } from './types.js';
/**
* Python HTTP plugin. Handles:
@@ -29,9 +29,13 @@ const FASTAPI_VERBS: Record<string, string> = {
patch: 'PATCH',
};
// ─── Provider: FastAPI @app.get/... ──────────────────────────────────
const FASTAPI_PATTERNS = compilePatterns({
name: 'python-fastapi',
// ─── Provider: FastAPI @app.<verb> / @router.<verb> ──────────────────
// Two separate patterns so we can tag detections by decorator object.
// Only `@router.*` detections participate in `include_router(prefix=)`
// path-prefix joining (see `PythonRepoContext` + `joinPrefix`); `@app.*`
// routes already carry their final path verbatim.
const FASTAPI_APP_PATTERNS = compilePatterns({
name: 'python-fastapi-app',
language: Python,
patterns: [
{
@@ -48,6 +52,138 @@ const FASTAPI_PATTERNS = compilePatterns({
],
} satisfies LanguagePatterns<Record<string, never>>);
const FASTAPI_ROUTER_PATTERNS = compilePatterns({
name: 'python-fastapi-router',
language: Python,
patterns: [
{
meta: {},
query: `
(decorator
(call
function: (attribute
object: (identifier) @obj (#eq? @obj "router")
attribute: (identifier) @method (#match? @method "^(get|post|put|delete|patch)$"))
arguments: (argument_list . (string) @path)))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── include_router(<router_obj>, prefix='/x') across the repo ────────
// Two shapes are common:
// app.include_router(assistant.router, prefix='/ai')
// app.include_router(my_router, prefix='/ai')
// The first names the originating module via `<module>.router`; the second
// references a name imported into the host file. We capture both.
const INCLUDE_ROUTER_ATTR_PATTERNS = compilePatterns({
name: 'python-fastapi-include-router-attr',
language: Python,
patterns: [
{
meta: {},
// Match any `<host>.include_router(<module>.router, ..., prefix='/x')`
// call. We deliberately do NOT pin `<host>` to the literal name `app`
// — production code routinely uses `api`, `application`, `asgi_app`,
// etc. The shape (`include_router` invoked with a router argument and
// a `prefix=` keyword) is specific enough on its own; restricting the
// host produces false negatives without removing meaningful false
// positives.
query: `
(call
function: (attribute
attribute: (identifier) @incl (#eq? @incl "include_router"))
arguments: (argument_list
(attribute
object: (identifier) @router_module
attribute: (identifier) @router_attr (#eq? @router_attr "router"))
(keyword_argument
name: (identifier) @kw (#eq? @kw "prefix")
value: (string) @prefix)))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
const INCLUDE_ROUTER_NAME_PATTERNS = compilePatterns({
name: 'python-fastapi-include-router-name',
language: Python,
patterns: [
{
meta: {},
// Same `<host>` rationale as INCLUDE_ROUTER_ATTR_PATTERNS — see above.
query: `
(call
function: (attribute
attribute: (identifier) @incl (#eq? @incl "include_router"))
arguments: (argument_list
(identifier) @router_name
(keyword_argument
name: (identifier) @kw (#eq? @kw "prefix")
value: (string) @prefix)))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// `from .api.assistant import router` style — used together with
// INCLUDE_ROUTER_NAME so we can map a local name back to its module
// path, then back to the file the router was declared in.
const FROM_IMPORT_ROUTER_PATTERNS = compilePatterns({
name: 'python-fastapi-from-import-router',
language: Python,
patterns: [
{
meta: {},
query: `
(import_from_statement
module_name: (_) @module
name: (dotted_name (identifier) @imported (#eq? @imported "router")))
`,
},
{
meta: {},
query: `
(import_from_statement
module_name: (_) @module
name: (aliased_import
name: (dotted_name (identifier) @imported (#eq? @imported "router"))
alias: (identifier) @alias))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// `from api import users` / `from api import users as u` — module-level
// imports where the imported name is itself the module that owns
// `<name>.router`. Lets Shape A (`<host>.include_router(<name>.router, …)`)
// look up the full package path of `<name>` and pin the prefix onto the
// exact file (`api/users.py`) rather than every file basenamed `users.py`.
const FROM_IMPORT_MODULE_PATTERNS = compilePatterns({
name: 'python-fastapi-from-import-module',
language: Python,
patterns: [
{
meta: {},
query: `
(import_from_statement
module_name: (_) @module
name: (dotted_name (identifier) @imported))
`,
},
{
meta: {},
query: `
(import_from_statement
module_name: (_) @module
name: (aliased_import
name: (dotted_name (identifier) @imported)
alias: (identifier) @alias))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: requests.get/post/... ──────────────────────────────────
const REQUESTS_VERB_PATTERNS = compilePatterns({
name: 'python-requests-verb',
@@ -447,15 +583,228 @@ const HTTPX_ASYNC_CLIENT_GENERIC_PATTERNS = compilePatterns({
],
} satisfies LanguagePatterns<Record<string, never>>);
// ─── prepareRepo: build router-module → prefix list map ─────────────
//
// FastAPI splits route declarations across files: handler decorators
// live in `api/<feature>.py` while `app.include_router(<x>.router,
// prefix='/ai')` lives in `main.py`. A per-file plugin scan therefore
// can't see the prefix that ought to be applied. We resolve this by
// running a one-shot pre-pass over the repo: for every file that
// hosts an `app.include_router(...)` we record the module the router
// came from (either via `module.router` attribute access, or via a
// local name resolved through a `from <module> import router` import)
// together with the prefix string. At scan time the python plugin
// looks up the current file's module key in this map and joins each
// prefix with each `@router.<verb>` decorator's path.
//
// Multiple prefixes for the same module are kept and emitted as
// separate detections — this matches FastAPI's behaviour when one
// router is mounted under several prefixes.
//
// Module keying is two-tiered to avoid prefix bleed between same-named
// files in different packages (e.g. `api/users.py` vs `admin/users.py`):
// • short key — file basename without `.py` (`users`)
// • long key — `<parent-dir>/<basename>` (`api/users`)
// The pre-pass records prefixes against the long key whenever the import
// site supplies enough context (`from api.users import router as ...` →
// long key `api/users`); otherwise it falls back to the short key.
// At scan time the file's own long key is consulted first; only when no
// long-key entry targets this file do we look up the short key. This
// preserves the previous coarse-grained behaviour where context is
// missing while delivering precision wherever the import statement
// gives us a multi-segment module path.
interface PythonRepoContext {
/** `<parent>/<stem>` → set of prefixes (precise, package-aware) */
prefixesByLongKey: Map<string, Set<string>>;
/** stem only → set of prefixes (basename fallback, may collide) */
prefixesByShortKey: Map<string, Set<string>>;
}
/** Strip `.py` and return the bare basename (e.g. `api/users.py` → `users`). */
function fileShortKey(rel: string): string {
const normalized = rel.replace(/\\/g, '/');
const slash = normalized.lastIndexOf('/');
const file = slash >= 0 ? normalized.slice(slash + 1) : normalized;
return file.endsWith('.py') ? file.slice(0, -3) : file;
}
/**
* Long key for a `.py` file: parent directory + stem, joined with `/`.
* Files at the repo root return the empty string (no parent), in which
* case callers should fall back to the short key.
*/
function fileLongKey(rel: string): string {
const normalized = rel.replace(/\\/g, '/');
const noExt = normalized.endsWith('.py') ? normalized.slice(0, -3) : normalized;
const lastSlash = noExt.lastIndexOf('/');
if (lastSlash < 0) return '';
const beforeLast = noExt.slice(0, lastSlash);
const stem = noExt.slice(lastSlash + 1);
const prevSlash = beforeLast.lastIndexOf('/');
const parent = prevSlash >= 0 ? beforeLast.slice(prevSlash + 1) : beforeLast;
return `${parent}/${stem}`;
}
/** Last `.`-separated segment of a (possibly relative) module path. */
function lastSegmentOfDotted(text: string): string {
const stripped = text.replace(/^\.+/, '');
if (!stripped) return '';
const dot = stripped.lastIndexOf('.');
return dot >= 0 ? stripped.slice(dot + 1) : stripped;
}
/**
* Last two `.`-separated segments of a (possibly relative) module path
* joined with `/`, e.g. `api.users` → `api/users`. Single-segment paths
* and pure-dot inputs return the empty string; callers should fall back
* to the short key in that case.
*/
function lastTwoSegmentsAsLongKey(text: string): string {
const stripped = text.replace(/^\.+/, '');
if (!stripped) return '';
const last = stripped.lastIndexOf('.');
if (last <= 0) return '';
const beforeLast = stripped.slice(0, last);
const stem = stripped.slice(last + 1);
const prev = beforeLast.lastIndexOf('.');
const parent = prev >= 0 ? beforeLast.slice(prev + 1) : beforeLast;
return `${parent}/${stem}`;
}
function recordPrefix(target: Map<string, Set<string>>, key: string, prefix: string): void {
const set = target.get(key) ?? new Set<string>();
set.add(prefix);
target.set(key, set);
}
function buildPythonRepoContext(
files: string[],
parser: Parser,
readFile: (rel: string) => string | null,
parseSource: (parser: Parser, src: string) => Parser.Tree | null,
): PythonRepoContext {
const prefixesByLongKey = new Map<string, Set<string>>();
const prefixesByShortKey = new Map<string, Set<string>>();
// Pre-pass over .py files. We deliberately run this even on files
// that don't contain `include_router` — the cost of an extra parse
// is bounded by the file count, and detecting `include_router`
// beforehand would require its own grep/scan.
for (const rel of files) {
if (!rel.endsWith('.py')) continue;
const src = readFile(rel);
if (!src) continue;
if (!src.includes('include_router')) continue;
parser.setLanguage(Python);
const tree = parseSource(parser, src);
if (!tree) continue;
// Local name → (short, long) map for the current file, populated
// from `from <module> import router [as <alias>]` statements. The
// alias (or 'router' when there is no alias) is the local name
// we'll later see passed to `<host>.include_router`.
interface LocalImport {
moduleShort: string;
moduleLong: string;
}
const localNameToModule = new Map<string, LocalImport>();
for (const m of runCompiledPatterns(FROM_IMPORT_ROUTER_PATTERNS, tree)) {
const moduleNode = m.captures.module;
const aliasNode = m.captures.alias;
const importedNode = m.captures.imported;
if (!moduleNode || !importedNode) continue;
const localName = aliasNode?.text ?? importedNode.text;
const moduleShort = lastSegmentOfDotted(moduleNode.text);
if (!moduleShort) continue;
const moduleLong = lastTwoSegmentsAsLongKey(moduleNode.text);
localNameToModule.set(localName, { moduleShort, moduleLong });
}
// Module-alias map: name imported from a multi-segment package →
// long key. Lets Shape A look up the precise file for `<name>.router`
// even when `<name>` collides with another package's basename.
const localNameToModuleAlias = new Map<string, string>();
for (const m of runCompiledPatterns(FROM_IMPORT_MODULE_PATTERNS, tree)) {
const moduleNode = m.captures.module;
const importedNode = m.captures.imported;
const aliasNode = m.captures.alias;
if (!moduleNode || !importedNode) continue;
// Skip the `router` shape — already handled by FROM_IMPORT_ROUTER_PATTERNS
// above and stored under its router-aware semantics.
if (importedNode.text === 'router') continue;
const moduleLong = lastTwoSegmentsAsLongKey(`${moduleNode.text}.${importedNode.text}`);
if (!moduleLong) continue;
const localName = aliasNode?.text ?? importedNode.text;
localNameToModuleAlias.set(localName, moduleLong);
}
// Shape A: `<host>.include_router(<module>.router, prefix='/x')`.
// The call site gives us only a short module name. We promote to a
// long key when the same file imports `<module>` via either
// `from <pkg> import <module>` (recorded in `localNameToModuleAlias`
// — the typical pattern) or, less commonly, a router-aware import
// statement. Only fall back to the basename short key when neither
// alias is available.
for (const m of runCompiledPatterns(INCLUDE_ROUTER_ATTR_PATTERNS, tree)) {
const modNode = m.captures.router_module;
const prefixNode = m.captures.prefix;
if (!modNode || !prefixNode) continue;
const prefix = unquoteLiteral(prefixNode.text);
if (prefix === null) continue;
const moduleShort = modNode.text;
const aliasLong = localNameToModuleAlias.get(moduleShort);
const sameFileImport = localNameToModule.get(moduleShort);
const longKey = aliasLong ?? sameFileImport?.moduleLong;
if (longKey) {
recordPrefix(prefixesByLongKey, longKey, prefix);
} else {
recordPrefix(prefixesByShortKey, moduleShort, prefix);
}
}
// Shape B: `<host>.include_router(my_router, prefix='/x')` — resolve
// `my_router` via the import map built above. Whenever the import
// statement supplied a multi-segment module path the long key is
// recorded, eliminating cross-package collisions.
for (const m of runCompiledPatterns(INCLUDE_ROUTER_NAME_PATTERNS, tree)) {
const nameNode = m.captures.router_name;
const prefixNode = m.captures.prefix;
if (!nameNode || !prefixNode) continue;
const localImp = localNameToModule.get(nameNode.text);
if (!localImp) continue;
const prefix = unquoteLiteral(prefixNode.text);
if (prefix === null) continue;
if (localImp.moduleLong) {
recordPrefix(prefixesByLongKey, localImp.moduleLong, prefix);
} else {
recordPrefix(prefixesByShortKey, localImp.moduleShort, prefix);
}
}
}
return { prefixesByLongKey, prefixesByShortKey };
}
function joinPrefix(prefix: string, route: string): string {
// Mirror FastAPI's path joining: trim trailing slash off prefix,
// ensure exactly one leading slash on the result.
const p = prefix.replace(/\/+$/, '');
const r = route.startsWith('/') ? route : `/${route}`;
return `${p}${r}`;
}
export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
name: 'python-http',
language: Python,
scan(tree) {
prepareRepo({ files, parser, readFile, parseSource }): RepoContext {
return buildPythonRepoContext(files, parser, readFile, parseSource);
},
scan(tree, repoContext, fileRel) {
const out: HttpDetection[] = [];
const httpxAsyncClients = collectHttpxAsyncClients(tree);
const ctx = repoContext as PythonRepoContext | undefined;
// Providers: FastAPI
for (const match of runCompiledPatterns(FASTAPI_PATTERNS, tree)) {
// Providers: FastAPI @app.<verb>("/path") — already absolute path.
for (const match of runCompiledPatterns(FASTAPI_APP_PATTERNS, tree)) {
const methodNode = match.captures.method;
const pathNode = match.captures.path;
if (!methodNode || !pathNode) continue;
@@ -473,6 +822,47 @@ export const PYTHON_HTTP_PLUGIN: HttpLanguagePlugin = {
});
}
// Providers: FastAPI @router.<verb>("/path") — must be joined
// with the prefix(es) declared at the include_router site. When
// no prefix is found we still emit the unprefixed path so this
// change is strictly additive vs. the prior @app-only behaviour;
// when the same router is mounted under multiple prefixes we emit
// one detection per prefix.
for (const match of runCompiledPatterns(FASTAPI_ROUTER_PATTERNS, tree)) {
const methodNode = match.captures.method;
const pathNode = match.captures.path;
if (!methodNode || !pathNode) continue;
const httpMethod = FASTAPI_VERBS[methodNode.text];
if (!httpMethod) continue;
const rawPath = unquoteLiteral(pathNode.text);
if (rawPath === null) continue;
// Long key first (precise, package-aware), short key as fallback.
// Mirrors the ingestion-side resolution in parse-impl.ts so the
// graph nodes and group contracts agree on which prefix applies.
const longKey = fileRel ? fileLongKey(fileRel) : '';
const longPrefixes = longKey ? ctx?.prefixesByLongKey.get(longKey) : undefined;
const shortKey = fileRel ? fileShortKey(fileRel) : '';
const shortPrefixes =
longPrefixes || !shortKey ? undefined : ctx?.prefixesByShortKey.get(shortKey);
const prefixSet = longPrefixes ?? shortPrefixes;
const paths =
prefixSet && prefixSet.size > 0
? [...prefixSet].map((p) => joinPrefix(p, rawPath))
: [rawPath];
for (const p of paths) {
out.push({
role: 'provider',
framework: 'fastapi',
method: httpMethod,
path: p,
name: null,
confidence: 0.8,
});
}
}
// Consumers: requests.<verb>
for (const match of runCompiledPatterns(REQUESTS_VERB_PATTERNS, tree)) {
const methodNode = match.captures.method;
@@ -40,6 +40,16 @@ export interface HttpDetection {
confidence: number;
}
export interface HttpScanInput {
filePath: string;
tree: Parser.Tree;
}
export interface HttpFileDetections {
filePath: string;
detections: HttpDetection[];
}
/**
* One language-scoped HTTP plugin. The plugin owns the tree-sitter
* grammar and the `scan` function that translates a parsed tree into
@@ -51,15 +61,54 @@ export interface HttpDetection {
* `LanguagePatterns.language` in `tree-sitter-scanner.ts` — the
* grammar modules export different shapes.
*/
/**
* Per-repo state a plugin can build during a `prepareRepo` pass before
* any per-file `scan` is invoked. The orchestrator threads this opaque
* value back into each `scan` call so plugins can resolve cross-file
* facts (e.g. FastAPI `app.include_router(prefix=...)` mappings live
* in `main.py` but apply to handlers declared in `api/*.py`).
*
* Plugins that have no cross-file state can omit `prepareRepo` and
* receive `undefined`.
*/
export type RepoContext = unknown;
export interface HttpLanguagePlugin {
/** Human-readable plugin name for diagnostics. */
name: string;
/** tree-sitter grammar object (passed to the shared parser). */
language: unknown;
/**
* Optional pre-pass: walk the relevant files in the repo and produce
* an opaque context that `scan` can use to resolve cross-file facts.
* Implementations must not throw — return undefined on any error so
* the orchestrator falls back to context-less scanning.
*/
prepareRepo?(args: {
repoPath: string;
files: string[];
parser: Parser;
readFile: (rel: string) => string | null;
parseSource: (parser: Parser, src: string) => Parser.Tree | null;
}): RepoContext | undefined;
/**
* Scan a parsed tree and return zero or more HTTP detections. Plugins
* must not throw — they should swallow per-match errors so a single
* malformed construct does not abort the whole file.
*
* `repoContext` is whatever the plugin's `prepareRepo` produced (or
* `undefined` if there is no `prepareRepo`).
*
* `fileRel` is the repo-relative path of the file being scanned;
* plugins that resolve cross-file facts (e.g. FastAPI router prefix
* joining) need it to key into `repoContext`. Optional so existing
* single-file plugins can keep their unary `scan(tree)` shape.
*/
scan(tree: Parser.Tree): HttpDetection[];
scan(tree: Parser.Tree, repoContext?: RepoContext, fileRel?: string): HttpDetection[];
/**
* Optional project-level scan hook for language rules that require
* multiple files, such as Java controllers inheriting Spring mappings
* from annotated interfaces.
*/
scanProject?(files: readonly HttpScanInput[]): HttpFileDetections[];
}
@@ -6,7 +6,13 @@ import type { ContractExtractor, CypherExecutor } from '../contract-extractor.js
import type { ExtractedContract, RepoHandle } from '../types.js';
import { readSafe } from './fs-utils.js';
import { parseSourceSafe } from '../../tree-sitter/safe-parse.js';
import { getPluginForFile, HTTP_SCAN_GLOB, type HttpDetection } from './http-patterns/index.js';
import {
getPluginForFile,
HTTP_SCAN_GLOB,
type HttpDetection,
type HttpLanguagePlugin,
type HttpScanInput,
} from './http-patterns/index.js';
/**
* Language-agnostic orchestrator for HTTP route (provider + consumer)
@@ -160,31 +166,85 @@ export class HttpRouteExtractor implements ContractExtractor {
// both graph-assisted enrichment and source-scan emission.
const parser = new Parser();
const cachedDetections = new Map<string, HttpDetection[]>();
const getDetections = (rel: string): HttpDetection[] => {
const cached = cachedDetections.get(rel);
if (cached) return cached;
const cachedInputs = new Map<
string,
{ plugin: HttpLanguagePlugin; input: HttpScanInput; repoContext: unknown } | null
>();
const projectDetections = new Map<string, HttpDetection[]>();
let projectScanComplete = false;
// Per-plugin cross-file context (e.g. Python's FastAPI router →
// include_router(prefix=...) map). Built lazily on first
// `getDetections` call for a file the plugin handles, scoped to the
// file list returned by `getScannedFiles`. Stored by plugin name so
// a repo with multiple languages keeps each plugin's context
// independent.
const repoContextByPlugin = new Map<string, unknown>();
const ensureRepoContext = async (
plugin: ReturnType<typeof getPluginForFile>,
): Promise<unknown> => {
if (!plugin || typeof plugin.prepareRepo !== 'function') return undefined;
if (repoContextByPlugin.has(plugin.name)) return repoContextByPlugin.get(plugin.name);
try {
const ctx = plugin.prepareRepo({
repoPath,
files: await getScannedFiles(),
parser,
readFile: (rel) => readSafe(repoPath, rel),
parseSource: (p, src) => parseSourceSafe(p, src),
});
repoContextByPlugin.set(plugin.name, ctx);
return ctx;
} catch {
repoContextByPlugin.set(plugin.name, undefined);
return undefined;
}
};
const getScanInput = async (
rel: string,
): Promise<{
plugin: HttpLanguagePlugin;
input: HttpScanInput;
repoContext: unknown;
} | null> => {
if (cachedInputs.has(rel)) return cachedInputs.get(rel) ?? null;
const plugin = getPluginForFile(rel);
if (!plugin) {
cachedDetections.set(rel, []);
return [];
cachedInputs.set(rel, null);
return null;
}
const repoContext = await ensureRepoContext(plugin);
const content = readSafe(repoPath, rel);
if (!content) {
cachedDetections.set(rel, []);
return [];
cachedInputs.set(rel, null);
return null;
}
try {
parser.setLanguage(plugin.language);
const tree = parseSourceSafe(parser, content);
const detections = plugin.scan(tree);
cachedDetections.set(rel, detections);
return detections;
const input = { filePath: rel, tree };
const item = { plugin, input, repoContext };
cachedInputs.set(rel, item);
return item;
} catch {
cachedDetections.set(rel, []);
return [];
cachedInputs.set(rel, null);
return null;
}
};
const getDetections = async (rel: string): Promise<HttpDetection[]> => {
const cached = cachedDetections.get(rel);
if (cached) return cached;
const scanInput = await getScanInput(rel);
const ownDetections = scanInput
? scanInput.plugin.scan(scanInput.input.tree, scanInput.repoContext, rel)
: [];
const detections = [...ownDetections, ...(projectDetections.get(rel) ?? [])];
cachedDetections.set(rel, detections);
return detections;
};
// Glob the source-scan file list at most once per extract() —
// both provider and consumer fallback paths share the same list.
let scannedFiles: string[] | null = null;
@@ -194,20 +254,46 @@ export class HttpRouteExtractor implements ContractExtractor {
return scannedFiles;
};
const collectProjectDetections = async (files: string[]): Promise<void> => {
if (projectScanComplete) return;
projectScanComplete = true;
const byPlugin = new Map<HttpLanguagePlugin, HttpScanInput[]>();
for (const rel of files) {
const scanInput = await getScanInput(rel);
if (!scanInput?.plugin.scanProject) continue;
const items = byPlugin.get(scanInput.plugin) ?? [];
items.push(scanInput.input);
byPlugin.set(scanInput.plugin, items);
}
for (const [plugin, inputs] of byPlugin) {
const results = plugin.scanProject?.(inputs) ?? [];
for (const result of results) {
const existing = projectDetections.get(result.filePath) ?? [];
projectDetections.set(result.filePath, [...existing, ...result.detections]);
}
}
cachedDetections.clear();
};
const files = await getScannedFiles();
await collectProjectDetections(files);
const graphProviders =
dbExecutor != null ? await this.extractProvidersGraph(dbExecutor, getDetections) : [];
// Source scan always runs to capture routes in languages/files not covered
// by graph edges; the glob and per-file parse results are cached above.
const providers = this.mergeGraphAndSourceContracts(
graphProviders,
this.extractProvidersSourceScan(await getScannedFiles(), getDetections),
await this.extractProvidersSourceScan(files, getDetections),
);
const graphConsumers =
dbExecutor != null ? await this.extractConsumersGraph(dbExecutor, getDetections) : [];
const consumers = this.mergeGraphAndSourceContracts(
graphConsumers,
this.extractConsumersSourceScan(await getScannedFiles(), getDetections),
await this.extractConsumersSourceScan(files, getDetections),
);
return [...providers, ...consumers];
@@ -232,7 +318,7 @@ export class HttpRouteExtractor implements ContractExtractor {
private async extractProvidersGraph(
db: CypherExecutor,
getDetections: (rel: string) => HttpDetection[],
getDetections: (rel: string) => Promise<HttpDetection[]>,
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
let rows: Record<string, unknown>[];
@@ -254,7 +340,7 @@ export class HttpRouteExtractor implements ContractExtractor {
// helpers — tree-sitter gives both pieces of information
// structurally. Always run the lookup: even when method is set by
// `methodFromRouteReason`, we still need the handler name.
const detections = filePath ? getDetections(filePath) : [];
const detections = filePath ? await getDetections(filePath) : [];
const providerDetections = detections.filter((d) => d.role === 'provider');
let handlerName: string | null = null;
const normalizedRoute = normalizeHttpPath(routePath);
@@ -331,13 +417,13 @@ export class HttpRouteExtractor implements ContractExtractor {
// ─── Source-scan providers ─────────────────────────────────────────
private extractProvidersSourceScan(
private async extractProvidersSourceScan(
files: string[],
getDetections: (rel: string) => HttpDetection[],
): ExtractedContract[] {
getDetections: (rel: string) => Promise<HttpDetection[]>,
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
for (const rel of files) {
const detections = getDetections(rel);
const detections = await getDetections(rel);
for (const d of detections) {
if (d.role !== 'provider') continue;
const pathNorm = normalizeHttpPath(d.path);
@@ -366,7 +452,7 @@ export class HttpRouteExtractor implements ContractExtractor {
private async extractConsumersGraph(
db: CypherExecutor,
getDetections: (rel: string) => HttpDetection[],
getDetections: (rel: string) => Promise<HttpDetection[]>,
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
let rows: Record<string, unknown>[];
@@ -382,7 +468,7 @@ export class HttpRouteExtractor implements ContractExtractor {
let method = 'GET';
// Prefer the plugin's detected method if we can find a matching
// fetch/axios call in the same file.
const detections = filePath ? getDetections(filePath) : [];
const detections = filePath ? await getDetections(filePath) : [];
// Symmetric to the provider path: if multiple consumer calls in
// the same file share the same normalized path (e.g. a GET
// fetch AND a POST fetch to `/api/orders`), `.find()` silently
@@ -436,13 +522,13 @@ export class HttpRouteExtractor implements ContractExtractor {
// ─── Source-scan consumers ─────────────────────────────────────────
private extractConsumersSourceScan(
private async extractConsumersSourceScan(
files: string[],
getDetections: (rel: string) => HttpDetection[],
): ExtractedContract[] {
getDetections: (rel: string) => Promise<HttpDetection[]>,
): Promise<ExtractedContract[]> {
const out: ExtractedContract[] = [];
for (const rel of files) {
const detections = getDetections(rel);
const detections = await getDetections(rel);
for (const d of detections) {
if (d.role !== 'consumer') continue;
const pathNorm = normalizeConsumerPath(d.path);
+16 -1
View File
@@ -150,9 +150,24 @@ export const processCobol = (
const entry = copybookMap.get(name.toUpperCase());
return entry ? entry.path : null;
};
// Memoize preprocessed copybook content for the duration of this
// processCobol call. A single copybook is COPYed by many programs (and at
// many COPY sites within a program); without this cache
// preprocessCobolSource would re-run once per COPY site —
// O(programs × copybooks) preprocessing passes over the same content.
// Keyed by the resolved copybook path. REPLACING is applied later by the
// expander on the returned (pre-REPLACING) content (see
// cobol-copy-expander.ts readFile→applyReplacing), so caching the
// pre-REPLACING preprocessed text here is safe and per-call-scoped.
const preprocessedCopyCache = new Map<string, string>();
const readCopy = (copyPath: string): string | null => {
const cached = preprocessedCopyCache.get(copyPath);
if (cached !== undefined) return cached;
const content = copybookByPath.get(copyPath);
return content ? preprocessCobolSource(content) : null;
if (!content) return null; // preserves original falsy→null (missing/empty)
const preprocessed = preprocessCobolSource(content);
preprocessedCopyCache.set(copyPath, preprocessed);
return preprocessed;
};
// Track module names for cross-program CALL resolution
@@ -0,0 +1,154 @@
/**
* Pure predicates gating C# `using` suffix-fallback resolution so BCL usings
* (e.g. `System.Threading.Tasks`) can't match a coincidentally-named local
* file (#1881).
*
* Lives in the shared `ingestion/` layer — NOT under `languages/csharp/` — so
* BOTH the registry-primary scope resolver (`languages/csharp/import-target.ts`)
* and the legacy DAG resolver (`import-resolvers/csharp.ts`) can import it
* without an `import-resolvers/ -> languages/` dependency inversion (#5).
*/
import type { CSharpNamespaceEvidence } from './language-config.js';
/**
* Top-level namespace segments that clearly belong to the BCL / runtime / a
* ubiquitous third-party package — i.e. roots a normal repo does NOT declare.
* These stay gated even when the namespace scan is truncated, so a single
* unreadable file / capped subtree can't silently re-enable BCL→local suffix
* matches repo-wide (#1881). A repo that legitimately declares one of these
* roots is still allowed via the alignment escape hatch below.
*/
const CSHARP_EXTERNAL_ROOTS: ReadonlySet<string> = new Set([
// .NET BCL / runtime
'System',
'Microsoft',
'Windows',
'Mono',
// ubiquitous third-party NuGet roots
'Newtonsoft',
'Serilog',
'AutoMapper',
'MediatR',
'Polly',
'FluentValidation',
'Grpc',
'Google',
'Azure',
'Amazon',
'AWSSDK',
// common test frameworks
'Xunit',
'NUnit',
'Moq',
'FluentAssertions',
'NSubstitute',
'Shouldly',
]);
/** Whether `targetRaw`'s top-level segment is a clearly-external root. */
function isExternalRoot(targetRaw: string): boolean {
const dot = targetRaw.indexOf('.');
const top = dot === -1 ? targetRaw : targetRaw.slice(0, dot);
return CSHARP_EXTERNAL_ROOTS.has(top);
}
/**
* Whether the unanchored suffix fallback may run for `targetRaw`.
*
* Fails OPEN when the namespace scan was truncated (large repos must not
* silently lose legitimate edges, #1881 #11) and when no evidence was
* threaded at all (preserves legacy permissive behavior). The truncation
* fail-open is carved out for clearly-external roots (BCL / well-known
* packages) that the repo does not declare, so one incomplete scan can't
* re-open the #1881 hole repo-wide. Otherwise defers to
* {@link importAlignsWithDeclaredNamespaces}.
*/
export function csharpSuffixFallbackAllowed(
targetRaw: string,
evidence: CSharpNamespaceEvidence | undefined,
): boolean {
if (evidence === undefined) return true;
if (evidence.truncated) {
// Keep clearly-external roots blocked through truncation UNLESS the repo
// actually declares an aligning namespace (the alignment check is the
// escape hatch — a repo that declares `namespace System;` still resolves).
if (
isExternalRoot(targetRaw) &&
!importAlignsWithDeclaredNamespaces(
targetRaw,
evidence.declaredNamespaces,
evidence.rootNamespaces,
)
) {
return false;
}
return true;
}
return importAlignsWithDeclaredNamespaces(
targetRaw,
evidence.declaredNamespaces,
evidence.rootNamespaces,
);
}
/** True when `targetRaw` plausibly refers to a namespace declared in-repo. */
export function importAlignsWithDeclaredNamespaces(
targetRaw: string,
declaredNamespaces: ReadonlySet<string> | undefined,
rootNamespaces?: ReadonlySet<string>,
): boolean {
if (declaredNamespaces === undefined || declaredNamespaces.size === 0) return false;
// Exact: the import IS a declared in-repo namespace.
if (declaredNamespaces.has(targetRaw)) return true;
// Child-of: the import's IMMEDIATE parent namespace is declared in-repo.
// Anchoring on the direct parent — not "any declared prefix" — is what stops
// a declared BCL prefix from green-lighting an unrelated BCL using: a repo
// that declares `namespace System;` must NOT make `using
// System.Threading.Tasks;` resolve to a coincidental local `Tasks.cs`,
// because the import's parent `System.Threading` is not itself declared
// (#1881). The case this still allows is a type / `using static` import under
// a declared namespace laid out without its full path on disk, e.g.
// `using static MyApp.Utils.Logger;` when `MyApp.Utils` is declared.
const lastDot = targetRaw.lastIndexOf('.');
if (lastDot > 0 && declaredNamespaces.has(targetRaw.slice(0, lastDot))) return true;
// Ancestor-of: the import is a strict prefix of some declared namespace
// (e.g. `using MyApp;` when `MyApp.Models` is declared). Only honored when
// the import also sits at or above an in-repo root namespace, so a BCL prefix
// can't qualify merely because a file declares something deeper under it
// (e.g. `System.Threading.Tasks.Extensions`) (#1881).
const childPrefix = targetRaw + '.';
for (const ns of declaredNamespaces) {
if (ns.startsWith(childPrefix)) {
return isAtOrAboveInRepoRoot(targetRaw, declaredNamespaces, rootNamespaces);
}
}
return false;
}
function isAtOrAboveInRepoRoot(
targetRaw: string,
declaredNamespaces: ReadonlySet<string>,
rootNamespaces: ReadonlySet<string> | undefined,
): boolean {
const descendantPrefix = targetRaw + '.';
if (rootNamespaces !== undefined && rootNamespaces.size > 0) {
for (const root of rootNamespaces) {
// targetRaw equals a root, or is an ancestor of one (e.g. `using MyApp;`
// for csproj RootNamespace `MyApp.Core`).
if (root === targetRaw || root.startsWith(descendantPrefix)) return true;
}
return false;
}
// No explicit roots (e.g. no csproj): treat the top-level segment of each
// declared namespace as the implied root.
for (const ns of declaredNamespaces) {
const dot = ns.indexOf('.');
const top = dot === -1 ? ns : ns.slice(0, dot);
if (top === targetRaw) return true;
}
return false;
}
@@ -7,27 +7,45 @@ import { SupportedLanguages } from 'gitnexus-shared';
import type { ImportResolutionConfig, ImportResolverStrategy } from '../types.js';
import { createStandardStrategy } from '../standard.js';
import { resolveCSharpImportInternal, resolveCSharpNamespaceDir } from '../csharp.js';
import { csharpSuffixFallbackAllowed } from '../../csharp-namespace-gate.js';
/** C# namespace-based resolution strategy via .csproj configs. */
export const csharpNamespaceStrategy: ImportResolverStrategy = (rawImportPath, _filePath, ctx) => {
const csharpConfigs = ctx.configs.csharpConfigs;
if (csharpConfigs.length > 0) {
const resolvedFiles = resolveCSharpImportInternal(
rawImportPath,
csharpConfigs,
ctx.normalizedFileList,
ctx.allFileList,
ctx.index,
);
if (resolvedFiles.length > 1) {
const dirSuffix = resolveCSharpNamespaceDir(rawImportPath, csharpConfigs);
if (dirSuffix) {
return { kind: 'package', files: resolvedFiles, dirSuffix };
}
const evidence = ctx.configs.csharpNamespaces;
if (csharpConfigs.length === 0) {
// No csproj → there's no namespace→directory mapping to apply, so the
// generic strategy would normally take over. But that generic suffix match
// is UNGATED: it re-introduces the BCL→local spurious match the #1881 gate
// exists to stop. Mirror the registry leg's no-csproj path — defer to the
// generic strategy ONLY for imports that align with an in-repo declared
// namespace; for everything else (BCL usings) return an authoritative empty
// result that STOPS the chain (#2 parity). With no evidence threaded the
// gate fails open, so behavior is unchanged when the scan didn't run.
if (!csharpSuffixFallbackAllowed(rawImportPath, evidence)) {
return { kind: 'files', files: [] };
}
if (resolvedFiles.length > 0) return { kind: 'files', files: resolvedFiles };
return null;
}
return null;
const resolvedFiles = resolveCSharpImportInternal(
rawImportPath,
csharpConfigs,
ctx.normalizedFileList,
ctx.allFileList,
ctx.index,
evidence,
);
if (resolvedFiles.length > 1) {
const dirSuffix = resolveCSharpNamespaceDir(rawImportPath, csharpConfigs);
if (dirSuffix) {
return { kind: 'package', files: resolvedFiles, dirSuffix };
}
}
// Authoritative once csproj configs exist: return even an empty result to
// STOP the chain, so the generic suffix fallback can't re-introduce the
// gated BCL→local match this resolver just suppressed (#1881).
return { kind: 'files', files: resolvedFiles };
};
export const csharpImportConfig: ImportResolutionConfig = {
@@ -7,11 +7,16 @@
import type { SuffixIndex } from './utils.js';
import { suffixResolve } from './utils.js';
import type { CSharpProjectConfig } from '../language-config.js';
import type { CSharpProjectConfig, CSharpNamespaceEvidence } from '../language-config.js';
import { csharpSuffixFallbackAllowed } from '../csharp-namespace-gate.js';
/**
* Resolve a C# using-directive import path to matching .cs files (low-level helper).
* Tries single-file match first, then directory match for namespace imports.
*
* The final unanchored suffix fallback is gated on `evidence` so BCL usings
* (e.g. `System.Threading.Tasks`) can't match a coincidentally-named local
* file (#1881). When `evidence` is omitted the fallback stays permissive.
*/
export function resolveCSharpImportInternal(
importPath: string,
@@ -19,6 +24,7 @@ export function resolveCSharpImportInternal(
normalizedFileList: string[],
allFileList: string[],
index?: SuffixIndex,
evidence?: CSharpNamespaceEvidence,
): string[] {
const namespacePath = importPath.replace(/\./g, '/');
const results: string[] = [];
@@ -86,7 +92,11 @@ export function resolveCSharpImportInternal(
}
}
// Fallback: suffix matching without namespace stripping (single file)
// Fallback: suffix matching without namespace stripping (single file).
// Gated on in-repo declared-namespace evidence (#1881).
if (!csharpSuffixFallbackAllowed(importPath, evidence)) {
return [];
}
const pathParts = namespacePath.split('/').filter(Boolean);
const fallback = suffixResolve(pathParts, normalizedFileList, allFileList, index);
return fallback ? [fallback] : [];
@@ -8,6 +8,7 @@ import type {
TsconfigPaths,
GoModuleConfig,
CSharpProjectConfig,
CSharpNamespaceEvidence,
ComposerConfig,
} from '../language-config.js';
import type { SwiftPackageConfig } from '../language-config.js';
@@ -32,6 +33,8 @@ export interface ImportConfigs {
composerConfig: ComposerConfig | null;
swiftPackageConfig: SwiftPackageConfig | null;
csharpConfigs: CSharpProjectConfig[];
/** In-repo namespace evidence gating C# suffix-fallback resolution (#1881). */
csharpNamespaces?: CSharpNamespaceEvidence;
}
/** Pre-built lookup structures for import resolution. Build once, reuse across chunks. */
+286 -43
View File
@@ -1,6 +1,9 @@
import fs from 'fs/promises';
import { createReadStream } from 'fs';
import { createInterface } from 'readline';
import path from 'path';
import type { ImportConfigs } from './import-resolvers/types.js';
import type { CsharpStructureLineScanner } from './languages/csharp/namespace-siblings.js';
import { isDev } from './utils/env.js';
@@ -40,6 +43,44 @@ export interface CSharpProjectConfig {
projectDir: string;
}
/**
* Declared-namespace evidence used to gate C# suffix-fallback resolution so
* BCL usings (e.g. `System.Threading.Tasks`) can't match a coincidentally-
* named local file (#1881).
*/
export interface CSharpNamespaceEvidence {
/** Every `namespace X.Y` declared in-repo (scan may be capped — see `truncated`). */
readonly declaredNamespaces?: ReadonlySet<string>;
/** csproj RootNamespace values plus the top-level segment of each declared
* namespace — the anchor set for the parent-namespace gate direction. */
readonly rootNamespaces?: ReadonlySet<string>;
/** True when the BFS hit its dir/depth cap, so the namespace set may be
* incomplete; the gate fails open (allows) in that case. */
readonly truncated?: boolean;
}
/** Result of a single BFS over a repo collecting both csproj configs and
* declared `.cs` namespaces (one disk traversal — see `scanCSharpProject`). */
export interface CSharpProjectScan {
readonly configs: CSharpProjectConfig[];
readonly declaredNamespaces: ReadonlySet<string>;
readonly rootNamespaces: ReadonlySet<string>;
readonly truncated: boolean;
}
/** Project the one-pass {@link CSharpProjectScan} into the
* {@link CSharpNamespaceEvidence} both import-resolution legs thread to the
* #1881 gate — one shape, two carriers (`ImportConfigs.csharpNamespaces` for
* the legacy DAG, `CsharpResolutionConfig.namespaces` for the scope resolver).
* Keeps the field mapping in one place so the two carriers can't drift. */
export function csharpScanToEvidence(scan: CSharpProjectScan): CSharpNamespaceEvidence {
return {
declaredNamespaces: scan.declaredNamespaces,
rootNamespaces: scan.rootNamespaces,
truncated: scan.truncated,
};
}
/** Swift Package Manager module config */
export interface SwiftPackageConfig {
/** Map of target name -> source directory path (e.g., "SiuperModel" -> "Package/Sources/SiuperModel") */
@@ -141,58 +182,258 @@ export async function loadComposerConfig(repoRoot: string): Promise<ComposerConf
}
}
/**
* Parse .csproj files to extract RootNamespace.
* Scans the repo root for .csproj files and returns configs for each.
*/
export async function loadCSharpProjectConfig(repoRoot: string): Promise<CSharpProjectConfig[]> {
const configs: CSharpProjectConfig[] = [];
// BFS scan for .csproj files up to 5 levels deep, cap at 100 dirs to avoid runaway scanning
const scanQueue: { dir: string; depth: number }[] = [{ dir: repoRoot, depth: 0 }];
const maxDepth = 5;
const maxDirs = 100;
let dirsScanned = 0;
// BFS bounds shared by the C# project/namespace scan. Sized to comfortably
// exceed normal C# repos so `truncated` stays the rare exception it was meant
// to be: a too-low cap trips `truncated=true` on ordinary repos, which makes
// `csharpSuffixFallbackAllowed` fail OPEN for every import and silently
// disables the #1881 gate. Truncation remains the safety valve for genuinely
// pathological trees (deep generated output, huge monorepos).
const CSHARP_SCAN_MAX_DEPTH = 24;
const CSHARP_SCAN_MAX_DIRS = 20000;
// Bound on in-flight file reads per directory so a directory with thousands of
// `.cs` files can't exhaust file descriptors / spike memory. Mirrors the
// Phase-1 walker's `READ_CONCURRENCY` (see `filesystem-walker.ts`).
const CSHARP_SCAN_READ_CONCURRENCY = 32;
const CSHARP_SCAN_SKIP_DIRS = new Set(['node_modules', '.git', 'bin', 'obj']);
const CSHARP_ROOT_NAMESPACE_RE = /<RootNamespace>\s*([^<]+)\s*<\/RootNamespace>/;
while (scanQueue.length > 0 && dirsScanned < maxDirs) {
// Declared `namespace` names are extracted with the comment/string-aware
// scanner shared with the scope-resolution namespace-siblings pass
// (`extractCsharpStructureViaScanner`), not a bare regex: a regex matches
// `namespace` inside comments and string literals, seeding the #1881 gate
// with phantom namespaces. Imported lazily (and memoized) so the always-on
// `loadImportConfigs` path — every repo, every language — doesn't eagerly
// pull tree-sitter-c-sharp in via `namespace-siblings.ts` → `query.ts`.
let csharpScannerFactoryPromise: Promise<() => CsharpStructureLineScanner> | undefined;
function getCsharpStructureScannerFactory(): Promise<() => CsharpStructureLineScanner> {
if (csharpScannerFactoryPromise === undefined) {
csharpScannerFactoryPromise = import('./languages/csharp/namespace-siblings.js').then(
(mod) => mod.createCsharpStructureScanner,
);
}
return csharpScannerFactoryPromise;
}
/**
* Single BFS over a repo that collects BOTH .csproj configs and the set of
* `namespace` declarations from `.cs` files.
*
* The csproj walk is cheap (a handful of project files); the namespace scan
* is NOT — it opens and reads every `.cs` file in the repo to collect its
* `namespace` declarations. That `.cs` read cost is the price of the #1881
* gate, not a saving: collapsing the csproj and namespace walks into one BFS
* avoids a second directory traversal, but the per-file `.cs` reads are new
* work this scan introduces. Reads within a directory are issued in bounded
* windows (see below); directories are still visited breadth-first.
*/
export async function scanCSharpProject(repoRoot: string): Promise<CSharpProjectScan> {
const configs: CSharpProjectConfig[] = [];
const declaredNamespaces = new Set<string>();
const rootNamespaces = new Set<string>();
const scanQueue: { dir: string; depth: number }[] = [{ dir: repoRoot, depth: 0 }];
let dirsScanned = 0;
let truncated = false;
while (scanQueue.length > 0) {
if (dirsScanned >= CSHARP_SCAN_MAX_DIRS) {
truncated = true;
break;
}
const { dir, depth } = scanQueue.shift()!;
dirsScanned++;
let entries: import('fs').Dirent[];
try {
const entries = await fs.readdir(dir, { withFileTypes: true });
for (const entry of entries) {
if (entry.isDirectory() && depth < maxDepth) {
// Skip common non-project directories
if (
entry.name === 'node_modules' ||
entry.name === '.git' ||
entry.name === 'bin' ||
entry.name === 'obj'
)
continue;
entries = await fs.readdir(dir, { withFileTypes: true });
} catch {
// Unreadable directory → its `.cs` namespaces are missed, so the scan is
// incomplete. Mark truncated so the #1881 gate fails OPEN (allows the
// suffix fallback) rather than wrongly blocking an import whose declaring
// namespace lived in the unread subtree (#5).
truncated = true;
continue;
}
// Collect read targets, then issue them in bounded windows (rather than all
// at once) so a directory with thousands of `.cs` files can't exhaust file
// descriptors / spike memory. csproj reads keep entry order (config
// precedence matters); `.cs` namespace results land in shared Sets where
// order is irrelevant.
const csprojNames: string[] = [];
const csNames: string[] = [];
for (const entry of entries) {
if (entry.isDirectory()) {
if (CSHARP_SCAN_SKIP_DIRS.has(entry.name)) continue;
if (depth < CSHARP_SCAN_MAX_DEPTH) {
scanQueue.push({ dir: path.join(dir, entry.name), depth: depth + 1 });
} else {
truncated = true; // a real subtree was pruned at the depth cap
}
if (entry.isFile() && entry.name.endsWith('.csproj')) {
try {
const csprojPath = path.join(dir, entry.name);
const content = await fs.readFile(csprojPath, 'utf-8');
const nsMatch = content.match(/<RootNamespace>\s*([^<]+)\s*<\/RootNamespace>/);
const rootNamespace = nsMatch ? nsMatch[1].trim() : entry.name.replace(/\.csproj$/, '');
const projectDir = path.relative(repoRoot, dir).replace(/\\/g, '/');
configs.push({ rootNamespace, projectDir });
if (isDev) {
logger.info(
`📦 Loaded C# project: ${entry.name} (namespace: ${rootNamespace}, dir: ${projectDir})`,
);
}
} catch {
// Can't read .csproj
}
continue;
}
if (!entry.isFile()) continue;
if (entry.name.endsWith('.csproj')) {
csprojNames.push(entry.name);
} else if (entry.name.endsWith('.cs')) {
csNames.push(entry.name);
}
}
for (let i = 0; i < csprojNames.length; i += CSHARP_SCAN_READ_CONCURRENCY) {
const batch = csprojNames.slice(i, i + CSHARP_SCAN_READ_CONCURRENCY);
const settled = await Promise.allSettled(
batch.map((name) => readCsprojConfig(path.join(dir, name), name, repoRoot, dir)),
);
for (const r of settled) {
const config = r.status === 'fulfilled' ? r.value : null;
if (config) {
configs.push(config);
rootNamespaces.add(config.rootNamespace);
}
}
} catch {
// Can't read directory
}
for (let i = 0; i < csNames.length; i += CSHARP_SCAN_READ_CONCURRENCY) {
const batch = csNames.slice(i, i + CSHARP_SCAN_READ_CONCURRENCY);
const settled = await Promise.allSettled(
batch.map((name) =>
collectDeclaredNamespaces(path.join(dir, name), declaredNamespaces, rootNamespaces),
),
);
// A `.cs` that was unreadable (or whose read/scan unexpectedly rejected)
// leaves its namespaces uncollected → mark truncated to fail the #1881
// gate OPEN rather than wrongly suppress an import. The scan streams each
// file, so file size no longer trips truncation.
for (const r of settled) {
if (r.status !== 'fulfilled' || r.value === 'truncated') truncated = true;
}
}
}
return configs;
if (truncated) {
// Surface the fail-open so an incomplete scan (dir/depth cap, or an
// unreadable directory or `.cs` file) silently disabling the #1881 gate
// repo-wide is observable (#4) rather than a mystery edge regression.
logger.warn(
`[csharp] namespace scan of ${repoRoot} truncated (dir cap ${CSHARP_SCAN_MAX_DIRS}, depth cap ${CSHARP_SCAN_MAX_DEPTH}, an unreadable directory, or an unreadable .cs file); the #1881 suffix-fallback gate fails open for unmatched usings`,
);
}
return { configs, declaredNamespaces, rootNamespaces, truncated };
}
// Generous soft budget for locating `<RootNamespace>`: a real .csproj declares
// it in the first PropertyGroup near the top, so this is only reached by a
// pathological project file with a huge leading ItemGroup and no early
// RootNamespace. On hit we OMIT the config rather than guess a root (Codex F4).
const CSPROJ_ROOT_SCAN_MAX_BYTES = 4 * 1024 * 1024;
// Overlap kept across stream chunks so a `<RootNamespace>` tag straddling a
// chunk boundary is still matched (the tag + a short namespace value fit well
// within this window).
const CSPROJ_TAG_OVERLAP = 512;
/**
* Stream a `.csproj` just far enough to find `<RootNamespace>`, in constant
* memory and without a stat-then-read filesystem race. Returns the namespace
* when found; otherwise `rootNamespace: null` with `capHit` distinguishing a
* genuine read-to-EOF absence (`false`) from "not found within the soft budget"
* (`true`) — so the caller never synthesizes a wrong filename root for a late
* tag (Codex F4).
*/
async function findCsprojRootNamespace(
csprojPath: string,
): Promise<{ rootNamespace: string | null; capHit: boolean }> {
const stream = createReadStream(csprojPath, { encoding: 'utf-8' });
let window = '';
let bytesRead = 0;
try {
for await (const chunk of stream) {
const text = chunk as string;
bytesRead += text.length;
window =
(window.length > CSPROJ_TAG_OVERLAP ? window.slice(-CSPROJ_TAG_OVERLAP) : window) + text;
const match = window.match(CSHARP_ROOT_NAMESPACE_RE);
if (match) {
stream.destroy();
return { rootNamespace: match[1]!.trim(), capHit: false };
}
if (bytesRead >= CSPROJ_ROOT_SCAN_MAX_BYTES) {
stream.destroy();
return { rootNamespace: null, capHit: true };
}
}
} catch {
// Unreadable .csproj: don't guess a filename root either — omit the config.
return { rootNamespace: null, capHit: true };
}
return { rootNamespace: null, capHit: false }; // read to EOF, tag genuinely absent
}
async function readCsprojConfig(
csprojPath: string,
fileName: string,
repoRoot: string,
dir: string,
): Promise<CSharpProjectConfig | null> {
const { rootNamespace: found, capHit } = await findCsprojRootNamespace(csprojPath);
// A late `<RootNamespace>` we couldn't reach (capHit) or an unreadable file
// must NOT synthesize a filename root — a wrong authoritative root would make
// imports under the real root resolve to nothing and suppress the fallback
// (Codex F4). Omit the config so the no-csproj fallback stays available. Only
// fall back to the filename on a genuine read-to-EOF absence of the tag.
if (capHit) return null;
const rootNamespace = found ?? fileName.replace(/\.csproj$/, '');
const projectDir = path.relative(repoRoot, dir).replace(/\\/g, '/');
if (isDev) {
logger.info(
`📦 Loaded C# project: ${fileName} (namespace: ${rootNamespace}, dir: ${projectDir})`,
);
}
return { rootNamespace, projectDir };
}
/**
* Stream one `.cs` file line-by-line and collect its declared `namespace` names
* into the shared Sets.
*
* Streaming (rather than reading the whole file into a string) keeps memory
* constant regardless of file size, so a large generated `.cs` (`*.g.cs`, EF /
* gRPC output) is fully scanned instead of skipped by a per-file size cap —
* which would otherwise trip `truncated` and disable the #1881 gate repo-wide.
* Only the cheap line scan streams here; the tree-sitter PARSE path keeps its
* own size cap.
*
* Returns `'truncated'` when the file could not be read, so the caller marks the
* scan truncated and the #1881 gate fails OPEN rather than wrongly suppress an
* import declared in the unread file. Returns `'ok'` on a complete read.
*/
async function collectDeclaredNamespaces(
filePath: string,
declaredNamespaces: Set<string>,
rootNamespaces: Set<string>,
): Promise<'ok' | 'truncated'> {
const createScanner = await getCsharpStructureScannerFactory();
const scanner = createScanner();
try {
// `crlfDelay: Infinity` treats every `\r\n` as a single break; the line
// scanner is terminator-agnostic, so a streamed scan yields the same
// namespaces as scanning the whole file content at once.
const lines = createInterface({
input: createReadStream(filePath, { encoding: 'utf-8' }),
crlfDelay: Infinity,
});
for await (const line of lines) {
scanner.pushLine(line);
}
} catch {
return 'truncated'; // unreadable source → signal truncation (fail open)
}
const structure = scanner.result();
for (const ns of structure.namespaces) {
declaredNamespaces.add(ns);
const dot = ns.indexOf('.');
rootNamespaces.add(dot === -1 ? ns : ns.slice(0, dot));
}
// A declaration the scanner could not fully capture (Codex F3) means the
// collected namespaces are an incomplete picture of this file — treat it like
// a truncated read so the #1881 gate fails OPEN rather than over-block an
// import whose namespace was dropped.
return structure.incomplete ? 'truncated' : 'ok';
}
export async function loadSwiftPackageConfig(repoRoot: string): Promise<SwiftPackageConfig | null> {
@@ -231,11 +472,13 @@ export async function loadSwiftPackageConfig(repoRoot: string): Promise<SwiftPac
/** Load all language-specific configs once for an ingestion run. */
export async function loadImportConfigs(repoRoot: string): Promise<ImportConfigs> {
const csharpScan = await scanCSharpProject(repoRoot);
return {
tsconfigPaths: await loadTsconfigPaths(repoRoot),
goModule: await loadGoModulePath(repoRoot),
composerConfig: await loadComposerConfig(repoRoot),
swiftPackageConfig: await loadSwiftPackageConfig(repoRoot),
csharpConfigs: await loadCSharpProjectConfig(repoRoot),
csharpConfigs: csharpScan.configs,
csharpNamespaces: csharpScanToEvidence(csharpScan),
};
}
@@ -80,7 +80,11 @@ export function emitCobolScopeCaptures(
: rangeOf(startLine, startCol, endLine, endCol);
const grouped: Record<string, Capture> = {
'@scope.module': capture('@scope.module', nameRange, name),
'@scope.module': capture(
'@scope.module',
rangeOf(startLine, startCol, endLine, endCol),
name,
),
'@declaration.program': capture(
'@declaration.program',
rangeOf(startLine, startCol, endLine, endCol),
@@ -118,7 +122,11 @@ export function emitCobolScopeCaptures(
: rangeOf(startLine, startCol, endLine, endCol);
const grouped: Record<string, Capture> = {
'@scope.module': capture('@scope.module', nameRange, prog.name),
'@scope.module': capture(
'@scope.module',
rangeOf(startLine, startCol, endLine, endCol),
prog.name,
),
'@declaration.program': capture(
'@declaration.program',
rangeOf(startLine, startCol, endLine, endCol),
+134 -46
View File
@@ -24,22 +24,18 @@
* V2 additionally walks class ancestors (via MRO), so base-class enclosing
* namespaces also contribute associated namespaces.
*
* **GitNexus approximation (not strict ISO C++ ADL):** passing a qualified
* function reference like `utils::worker` contributes `utils` to the associated
* set, enabling resolution of unqualified calls like `with_callback(utils::worker)`
* to `utils::with_callback`. Under ISO C++ `[basic.lookup.argdep]`, associated
* entities for function-type arguments come from the **parameter types and return
* type** of each function in the overload set — NOT the function's enclosing
* namespace. For `void worker()`, the standard-compliant associated set is empty.
* GitNexus instead contributes the enclosing namespace of any Function/Method
* def whose simple name matches, because it enables the dominant real-world ADL
* pattern at reasonable precision cost.
* Function-reference arguments follow ISO C++ `[basic.lookup.argdep]`:
* associated entities come from the parameter types and return type of each
* referenced function in the overload set, not from the function's enclosing
* namespace. For `void worker()`, the associated set is empty. For
* `void worker(api::Token)` or `api::Token make_token()`, `api` is associated
* through `Token`.
*
* For qualified refs (e.g. `utils::worker`) the namespace is confirmed via a
* workspace lookup (only contributed when a Function/Method named `worker` exists
* in `utils`). For unqualified refs the workspace is searched for any Function
* def with that simple name. Locally-declared function-pointer variables
* (e.g. `void (*g)()`) and function parameters are excluded from this path.
* For qualified refs (e.g. `utils::worker`) the workspace lookup is restricted
* to functions/methods named `worker` in `utils`; for unqualified refs the
* workspace is searched for matching functions/methods by simple name. Locally
* declared function-pointer variables and function parameters are excluded
* from this path.
*
* ADL candidates are merged with ordinary unqualified-lookup candidates
* in the free-call fallback before overload narrowing.
@@ -70,6 +66,7 @@
import type { ParsedFile, ScopeId, SymbolDefinition } from 'gitnexus-shared';
import type { ScopeResolutionIndexes } from '../../model/scope-resolution-indexes.js';
import { normalizeCppParamType } from './arity-metadata.js';
import { isCppInlineNamespaceScope } from './inline-namespaces.js';
/**
@@ -97,11 +94,8 @@ export interface CppAdlArgInfo {
/** When set, the arg is a potential free-function reference (not a locally-
* declared function-pointer variable or function parameter). Contains the
* identifier text as written in source (e.g. `"utils::worker"` or
* `"worker"`). GitNexus approximation: the function's enclosing namespace
* is contributed to the ADL associated set. For qualified refs a workspace
* lookup confirms a Function/Method with that simple name exists in the
* namespace before contributing; for unqualified refs every namespace
* containing a matching Function/Method def is contributed. */
* `"worker"`). Resolution contributes associated namespaces from each
* referenced Function/Method def's parameter and return types. */
readonly functionRefText?: string;
}
@@ -207,7 +201,12 @@ export function pickCppAdlCandidates(
for (const arg of args) {
collectAssociatedNamespacesForAdlArg(arg, scopes, associatedNamespaces);
if (arg.functionRefText !== undefined) {
collectFunctionRefNamespaces(arg.functionRefText, parsedFiles, associatedNamespaces);
collectFunctionTypeAssociatedNamespaces(
arg.functionRefText,
scopes,
parsedFiles,
associatedNamespaces,
);
}
}
if (associatedNamespaces.size === 0) return undefined;
@@ -472,23 +471,12 @@ function findCppClassDefBySimpleName(
}
/**
* Contribute associated namespaces for a function-reference argument.
*
* - **Qualified refs** (`utils::worker`, `outer::inner::fn`): the namespace
* is extracted from the qualifier text (converting `::` to `.` for dot-joined
* QName matching). A workspace lookup then **verifies** that a Function or
* Method def named `worker` (the simple name after the last `::`) actually
* exists in the extracted namespace. This prevents false positives from
* namespace-qualified variables, enum values, and static data members, which
* also produce `qualified_identifier` AST nodes in tree-sitter-cpp (the
* AST node type alone does not distinguish functions from non-function names).
* - **Unqualified refs** (`worker`): the workspace is searched for any
* Function/Method def whose simple name matches. Every distinct enclosing
* namespace found is added — overloads across the same namespace produce
* a single entry; GitNexus does not select a specific overload at this stage.
* Contribute associated namespaces for a function-reference argument by walking
* the referenced overload set's parameter and return types.
*/
function collectFunctionRefNamespaces(
function collectFunctionTypeAssociatedNamespaces(
refText: string,
scopes: ScopeResolutionIndexes,
parsedFiles: readonly ParsedFile[],
out: Set<string>,
): void {
@@ -511,30 +499,130 @@ function collectFunctionRefNamespaces(
for (const def of scope.ownedDefs) {
if (def.type !== 'Function' && def.type !== 'Method') continue;
const simple = def.qualifiedName?.split('.').pop() ?? def.qualifiedName ?? '';
if (simple === simpleName) {
out.add(nsText);
return; // Namespace confirmed; no need to scan further files.
}
if (simple === simpleName) collectAssociatedNamespacesForFunctionDef(def, scopes, out);
}
}
}
return;
}
// Unqualified: search all namespace scopes for a Function def with this
// simple name and contribute its enclosing namespace.
// Unqualified function references are approximated workspace-wide, matching
// the previous V1 lookup scope. The stricter part of this PR is what each
// overload contributes: only namespaces from parameter/return types, never
// the function's own enclosing namespace.
for (const parsed of parsedFiles) {
const scopesById = new Map<ScopeId, (typeof parsed.scopes)[number]>();
for (const sc of parsed.scopes) scopesById.set(sc.id, sc);
for (const scope of parsed.scopes) {
if (scope.kind !== 'Namespace') continue;
for (const def of scope.ownedDefs) {
if (def.type !== 'Function' && def.type !== 'Method') continue;
const simple = def.qualifiedName?.split('.').pop() ?? def.qualifiedName ?? '';
if (simple !== refText) continue;
const nsQName = computeNamespaceQName(scope, scopesById);
if (nsQName !== '') out.add(nsQName);
collectAssociatedNamespacesForFunctionDef(def, scopes, out);
}
}
}
}
function collectAssociatedNamespacesForFunctionDef(
def: SymbolDefinition,
scopes: ScopeResolutionIndexes,
out: Set<string>,
): void {
const parameterTypes = def.parameterTypeClasses?.map((typeClass) => typeClass.base);
for (const paramType of parameterTypes ?? def.parameterTypes ?? []) {
collectAssociatedNamespacesForFunctionTypeText(paramType, scopes, out);
}
if (def.returnType !== undefined) {
collectAssociatedNamespacesForFunctionTypeText(def.returnType, scopes, out);
}
}
function collectAssociatedNamespacesForFunctionTypeText(
typeText: string,
scopes: ScopeResolutionIndexes,
out: Set<string>,
): void {
for (const token of extractCppTypeNameTokens(typeText)) {
if (isIgnoredCppAdlNamespace(token.namespaceName)) continue;
addAssociatedNamespaceForClassName(token.simpleName, scopes, out);
if (token.namespaceName !== '') out.add(token.namespaceName);
}
}
function extractCppTypeNameTokens(typeText: string): readonly {
readonly simpleName: string;
readonly namespaceName: string;
}[] {
const cleaned = normalizeCppParamType(typeText);
if (cleaned === '' || isPrimitiveCppAdlType(cleaned)) return [];
const out: { simpleName: string; namespaceName: string }[] = [];
const seen = new Set<string>();
const tokenSource = typeText.includes('<') ? `${cleaned} ${typeText}` : cleaned;
for (const rawToken of tokenSource.match(/[A-Za-z_]\w*(?:::[A-Za-z_]\w*)*/g) ?? []) {
if (isPrimitiveCppAdlType(rawToken)) continue;
const segments = rawToken.split('::').filter((part) => part.length > 0);
const simpleName = segments.at(-1) ?? '';
if (simpleName === '' || isPrimitiveCppAdlType(simpleName)) continue;
const namespaceName = segments.length > 1 ? segments.slice(0, -1).join('.') : '';
const key = `${namespaceName}\0${simpleName}`;
if (seen.has(key)) continue;
seen.add(key);
out.push({
simpleName,
namespaceName,
});
}
return out;
}
const CPP_ADL_PRIMITIVE_OR_KEYWORD_TYPES = new Set<string>([
'alignas',
'alignof',
'auto',
'bool',
'char',
'char8_t',
'char16_t',
'char32_t',
'class',
'const',
'consteval',
'constexpr',
'constinit',
'decltype',
'double',
'enum',
'explicit',
'extern',
'float',
'inline',
'int',
'long',
'mutable',
'noexcept',
'null',
'register',
'short',
'signed',
'static',
'string',
'struct',
'template',
'thread_local',
'typename',
'union',
'unknown',
'unsigned',
'void',
'volatile',
'wchar_t',
'...',
]);
function isPrimitiveCppAdlType(typeText: string): boolean {
return CPP_ADL_PRIMITIVE_OR_KEYWORD_TYPES.has(typeText);
}
function isIgnoredCppAdlNamespace(namespaceName: string): boolean {
return namespaceName === 'std' || namespaceName.startsWith('std.');
}
@@ -126,6 +126,21 @@ export function emitCppScopeCaptures(
JSON.stringify(arity.parameterTypeClasses),
);
}
const returnType = extractCppDeclarationReturnType(fnNode);
if (returnType !== undefined) {
grouped['@declaration.return-type'] = syntheticCapture(
'@declaration.return-type',
fnNode,
returnType,
);
}
if (hasExplicitSpecifier(fnNode)) {
grouped['@declaration.is-explicit'] = syntheticCapture(
'@declaration.is-explicit',
fnNode,
'true',
);
}
// Detect static storage class (file-local linkage)
if (hasStaticStorageClass(fnNode)) {
@@ -410,6 +425,30 @@ export function emitCppScopeCaptures(
return out;
}
function extractCppDeclarationReturnType(fnNode: SyntaxNode): string | undefined {
const typeNode = fnNode.childForFieldName('type');
if (typeNode === null) return undefined;
const funcDeclarator = findFunctionDeclarator(fnNode);
if (funcDeclarator !== null && isCppUnsupportedReturnTypeDeclarator(funcDeclarator)) {
return undefined;
}
const typeText = typeNode.text.trim();
if (typeText !== 'auto') return typeText.length > 0 ? typeText : undefined;
if (funcDeclarator === null) return typeText;
for (let i = 0; i < funcDeclarator.namedChildCount; i++) {
const child = funcDeclarator.namedChild(i);
if (child?.type !== 'trailing_return_type') continue;
const typeDesc = child.firstNamedChild;
return typeDesc?.text.trim() || typeText;
}
return typeText;
}
function isCppUnsupportedReturnTypeDeclarator(funcDeclarator: SyntaxNode): boolean {
const text = funcDeclarator.text;
return /\boperator\b/.test(text) || /(^|[(:\s])~\s*[A-Za-z_]\w*/.test(text);
}
/**
* Walk every C++ class/struct base clause and emit `@reference.inherits`
* captures for each base so scope resolution can resolve them into EXTENDS
@@ -1542,6 +1581,20 @@ function extractDeclaratorLeafName(node: SyntaxNode): string | null {
return null;
}
/**
* Check if a C++ declaration has an `explicit` specifier. Tree-sitter-cpp
* exposes `explicit` as a direct keyword child on constructor declarations in
* current grammar builds; the bounded text prefix keeps this resilient across
* small grammar shape differences without scanning whole function bodies.
*/
function hasExplicitSpecifier(node: SyntaxNode): boolean {
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (child !== null && child.text === 'explicit') return true;
}
return /\bexplicit\b/.test(node.text.slice(0, 128));
}
/**
* Check if a C++ function_definition or declaration has `static` storage class.
*/
@@ -12,7 +12,8 @@
* - rank 2: standard conversion (arithmetic, nullptr -> T*, T* -> bool,
* T* -> void*)
* - rank 3: nullptr -> bool (kept worse than nullptr -> T*)
* - rank 4: ellipsis conversion (worst viable)
* - rank 4: user-defined conversion (one-step, conservative)
* - rank 5: ellipsis conversion (worst viable)
* - Infinity: mismatch (string -> int, user types, unsupported shapes)
*
* This function is intentionally C++-specific. Other languages may define
@@ -20,6 +21,7 @@
*/
import type { ParameterTypeClass } from 'gitnexus-shared';
import { hasCppUserDefinedConversion } from './user-defined-conversions.js';
/** Set of normalized arithmetic types that support implicit conversion. */
const ARITHMETIC = new Set(['int', 'double', 'char', 'bool']);
@@ -34,7 +36,8 @@ const INTEGRAL_PROMOTION = new Map([
* Return the conversion rank from `argType` to `paramType`.
*
* @returns 0 for exact match, 1 for integral promotion, 2 for standard
* conversion, 3 for nullptr -> bool, 4 for ellipsis, Infinity
* conversion, 3 for nullptr -> bool, 4 for user-defined conversion,
* 5 for ellipsis, Infinity
* for mismatch.
*/
export function cppConversionRank(
@@ -46,13 +49,14 @@ export function cppConversionRank(
if (argType === paramType) {
return exactShapeCompatible(argTypeClass, paramTypeClass) ? 0 : Infinity;
}
if (paramType === '...') return 4;
if (paramType === '...') return 5;
if (INTEGRAL_PROMOTION.get(argType) === paramType) return 1;
if (ARITHMETIC.has(argType) && ARITHMETIC.has(paramType)) return 2;
if (argType === 'null' && isPointer(paramTypeClass)) return 2;
if (argType === 'null' && paramType === 'bool') return 3;
if (isPointer(argTypeClass) && paramType === 'bool') return 2;
if (isPointer(argTypeClass) && isPointer(paramTypeClass) && paramType === 'void') return 2;
if (hasCppUserDefinedConversion(argType, paramType)) return 4;
return Infinity;
}
@@ -225,6 +225,12 @@ const CPP_SCOPE_QUERY = `
declarator: (function_declarator
declarator: (field_identifier) @declaration.name))) @declaration.method
;; Constructor prototype in class body: User(int id);
(field_declaration_list
(declaration
declarator: (function_declarator
declarator: (identifier) @declaration.name)) @declaration.method)
;; Method prototype with reference return: User& getRef();
(field_declaration
declarator: (reference_declarator
@@ -34,6 +34,10 @@ import {
} from './inline-namespaces.js';
import { populateCppRangeBindings } from './range-bindings.js';
import { cppConstraintCompatibility } from './constraint-filter.js';
import {
clearCppUserDefinedConversions,
populateCppUserDefinedConversions,
} from './user-defined-conversions.js';
/**
* C++ `ScopeResolver` registered in `SCOPE_RESOLVERS` and consumed by
@@ -61,6 +65,7 @@ export const cppScopeResolver: ScopeResolver = {
clearCppDependentBases();
clearCppAdlState();
clearCppInlineNamespaces();
clearCppUserDefinedConversions();
return scanCppHeaderFiles(repoPath);
},
@@ -110,6 +115,10 @@ export const cppScopeResolver: ScopeResolver = {
// by ADL (U2 of plan 2026-05-13-001) to identify each argument type's
// associated namespace for Koenig lookup.
populateCppAssociatedNamespaces(parsed);
// Build conservative one-step user-defined conversion facts for
// overload ranking (#1631): implicit converting constructors only,
// with no chaining or conversion-operator handling.
populateCppUserDefinedConversions(parsed);
},
// Resolve recorded template-class → dependent-base simple names to
@@ -0,0 +1,140 @@
import type { ParsedFile, SymbolDefinition } from 'gitnexus-shared';
import type { ScopeId } from 'gitnexus-shared';
import { normalizeCppParamType } from './arity-metadata.js';
const userDefinedConversions = new Set<string>();
const pendingUserDefinedConversions: PendingUserDefinedConversion[] = [];
const classIdentitiesBySimpleName = new Map<string, Set<string>>();
interface PendingUserDefinedConversion {
readonly argType: string;
readonly paramType: string;
readonly ownerClassName: string;
}
export function clearCppUserDefinedConversions(): void {
userDefinedConversions.clear();
pendingUserDefinedConversions.length = 0;
classIdentitiesBySimpleName.clear();
}
export function hasCppUserDefinedConversion(argType: string, paramType: string): boolean {
return userDefinedConversions.has(conversionKey(argType, paramType));
}
export function populateCppUserDefinedConversions(parsed: ParsedFile): void {
const scopesById = new Map<ScopeId, (typeof parsed.scopes)[number]>();
for (const scope of parsed.scopes) scopesById.set(scope.id, scope);
for (const classScope of parsed.scopes) {
if (classScope.kind !== 'Class') continue;
const classDef = classScope.ownedDefs.find(isClassLike);
if (classDef !== undefined) recordClassIdentity(classDef);
}
for (const classScope of parsed.scopes) {
if (classScope.kind !== 'Class') continue;
const classDef = classScope.ownedDefs.find(isClassLike);
if (classDef === undefined) continue;
const className = normalizedSimpleName(classDef);
if (className === '') continue;
const methodDefs = collectClassMethodDefs(classScope.id, parsed, scopesById);
for (const def of methodDefs) {
const simpleName = simpleNameOf(def);
if (simpleName === className && def.parameterTypes?.length === 1) {
if (def.isExplicit === true) continue;
registerPendingCppUserDefinedConversion(def.parameterTypes[0], className, className);
}
}
}
rebuildCppUserDefinedConversions();
}
export function registerCppUserDefinedConversion(argType: string, paramType: string): void {
if (argType === '' || paramType === '') return;
if (argType === paramType) return;
userDefinedConversions.add(conversionKey(argType, paramType));
}
function collectClassMethodDefs(
classScopeId: ScopeId,
parsed: ParsedFile,
scopesById: ReadonlyMap<ScopeId, (typeof parsed.scopes)[number]>,
): SymbolDefinition[] {
const methods: SymbolDefinition[] = [];
const classScope = scopesById.get(classScopeId);
if (classScope === undefined) return methods;
for (const def of classScope.ownedDefs) {
if (isCallableMember(def)) methods.push(def);
}
for (const scope of parsed.scopes) {
if (scope.parent !== classScopeId) continue;
if (scope.kind === 'Class') continue;
for (const def of scope.ownedDefs) {
if (isCallableMember(def)) methods.push(def);
}
}
return methods;
}
function conversionKey(argType: string, paramType: string): string {
return `${argType}\0${paramType}`;
}
function registerPendingCppUserDefinedConversion(
argType: string,
paramType: string,
ownerClassName: string,
): void {
if (argType === '' || paramType === '') return;
if (argType === paramType) return;
pendingUserDefinedConversions.push({ argType, paramType, ownerClassName });
}
function rebuildCppUserDefinedConversions(): void {
userDefinedConversions.clear();
for (const conversion of pendingUserDefinedConversions) {
if (isAmbiguousClassName(conversion.ownerClassName)) continue;
userDefinedConversions.add(conversionKey(conversion.argType, conversion.paramType));
}
}
function recordClassIdentity(def: SymbolDefinition): void {
const simpleName = normalizedSimpleName(def);
if (simpleName === '') return;
const identities = classIdentitiesBySimpleName.get(simpleName) ?? new Set<string>();
identities.add(normalizedQualifiedClassName(def));
classIdentitiesBySimpleName.set(simpleName, identities);
}
function isAmbiguousClassName(simpleName: string): boolean {
return (classIdentitiesBySimpleName.get(simpleName)?.size ?? 0) > 1;
}
function normalizedQualifiedClassName(def: SymbolDefinition): string {
const qualifiedName = def.qualifiedName ?? simpleNameOf(def);
if (qualifiedName === '' || !qualifiedName.includes('.')) return `${def.filePath}:${def.nodeId}`;
return qualifiedName
.split('.')
.map((part) => normalizeCppParamType(part))
.join('.');
}
function normalizedSimpleName(def: SymbolDefinition): string {
return normalizeCppParamType(simpleNameOf(def));
}
function simpleNameOf(def: SymbolDefinition): string {
return def.qualifiedName?.split('.').pop() ?? def.qualifiedName ?? '';
}
function isClassLike(def: SymbolDefinition): boolean {
return def.type === 'Class' || def.type === 'Struct' || def.type === 'Interface';
}
function isCallableMember(def: SymbolDefinition): boolean {
return def.type === 'Method' || def.type === 'Constructor';
}
@@ -9,30 +9,112 @@
* match. Cross-file partial-class aggregation runs at graph-bridge
* time (Unit 6) via `populateOwners`.
*
* The legacy csproj-based `resolveCSharpImportInternal` needs config
* objects the scope-resolver doesn't carry; the Unit 7 parity gate
* will surface cases where the suffix-match diverges from the
* namespace-based resolver and we'll adjust the contract if needed.
* When `.csproj` configs are available, consults the legacy
* namespace-directory resolver first. Both that resolver's suffix
* fallback and the progressive prefix stripping below are gated on
* declared in-repo namespaces so BCL usings like `System.Threading.Tasks`
* cannot spuriously match a local `Tasks.cs` (#1881).
*
* Returning `null` lets the finalize algorithm mark the edge as
* `linkStatus: 'unresolved'`.
*/
import type { ParsedImport, WorkspaceIndex } from 'gitnexus-shared';
import type { CSharpProjectConfig, CSharpNamespaceEvidence } from '../../language-config.js';
import { resolveCSharpImportInternal } from '../../import-resolvers/csharp.js';
import { buildSuffixIndex, type SuffixIndex } from '../../import-resolvers/utils.js';
import { csharpSuffixFallbackAllowed } from '../../csharp-namespace-gate.js';
export interface CsharpResolveContext {
readonly fromFile: string;
readonly allFilePaths: ReadonlySet<string>;
readonly csharpConfigs?: readonly CSharpProjectConfig[];
readonly namespaces?: CSharpNamespaceEvidence;
}
/** Normalized file list + suffix index, built once per workspace `allFilePaths`. */
interface WorkspaceFileIndex {
readonly normalized: string[];
readonly all: string[];
readonly index: SuffixIndex;
}
// Memoize on Set identity: the orchestrator passes the SAME `allFilePaths`
// Set through every `resolveImportTarget` call in a pass, so this rebuilds
// the normalized list + suffix index once instead of once per import (#1881 #2).
const workspaceFileIndexCache = new WeakMap<ReadonlySet<string>, WorkspaceFileIndex>();
function getWorkspaceFileIndex(allFilePaths: ReadonlySet<string>): WorkspaceFileIndex {
const cached = workspaceFileIndexCache.get(allFilePaths);
if (cached) return cached;
const all = [...allFilePaths];
const normalized = all.map((f) => f.replace(/\\/g, '/'));
const built: WorkspaceFileIndex = { normalized, all, index: buildSuffixIndex(normalized, all) };
workspaceFileIndexCache.set(allFilePaths, built);
return built;
}
export function resolveCsharpImportTarget(
parsedImport: ParsedImport,
workspaceIndex: WorkspaceIndex,
): string | null {
// WorkspaceIndex is `unknown` in the shared contract (Ring 1
// placeholder). The scope-resolution orchestrator hands us a
// CsharpResolveContext-shaped object; narrow structurally rather
// than via a cast chain so unexpected shapes return null cleanly.
const ctx = narrowContext(workspaceIndex);
if (ctx === null) return null;
if (parsedImport.kind === 'dynamic-unresolved') return null;
if (parsedImport.targetRaw === null || parsedImport.targetRaw === '') return null;
const targetRaw = parsedImport.targetRaw;
const evidence = ctx.namespaces;
const csharpConfigs = ctx.csharpConfigs ?? [];
if (csharpConfigs.length > 0) {
const { normalized, all, index } = getWorkspaceFileIndex(ctx.allFilePaths);
const fromCsproj = resolveCSharpImportInternal(
targetRaw,
[...csharpConfigs],
normalized,
all,
index,
evidence,
);
if (fromCsproj.length > 0) return fromCsproj[0]!;
// csproj configs are authoritative: mirror legacy `configs/csharp.ts`,
// which returns an empty result to STOP the chain. Falling through to the
// ungated `resolveDirectMatch` would re-introduce the BCL→local match the
// internal resolver's gate just suppressed (#1881 parity, #2).
return null;
}
// Namespace path: `System.Collections.Generic` → `System/Collections/Generic`.
const pathLike = targetRaw.replace(/\./g, '/');
// Gate the WHOLE no-csproj path on declared in-repo namespaces — the direct
// path/suffix match INCLUDED — so a BCL using can't resolve to a
// coincidentally path-aligned local file (e.g. `Legacy/System/Threading/
// Tasks.cs` satisfying `using System.Threading.Tasks;`). Running the gate
// before `resolveDirectMatch` mirrors the legacy leg's gate-first ordering
// (`import-resolvers/configs/csharp.ts`), so the two legs are equivalent
// (#1881 parity, Codex F2). The gate keeps its fail-open for
// undefined/truncated evidence, so legitimate edges in unscanned repos are
// unaffected.
if (!csharpSuffixFallbackAllowed(targetRaw, evidence)) {
return null;
}
// Exact file / nested-suffix / namespace-dir direct-child match.
const direct = resolveDirectMatch(ctx.allFilePaths, pathLike);
if (direct !== null) return direct;
// Progressive prefix stripping — mirrors csproj's root-namespace mapping
// without the csproj.
return resolveByProgressiveStripping(ctx.allFilePaths, pathLike);
}
/**
* `WorkspaceIndex` is an opaque `unknown` placeholder in the shared contract;
* the orchestrator hands us a `CsharpResolveContext`-shaped object. Narrow
* structurally rather than via a cast chain so unexpected shapes fail cleanly.
*/
function narrowContext(workspaceIndex: WorkspaceIndex): CsharpResolveContext | null {
const ctx = workspaceIndex as CsharpResolveContext | undefined;
if (
ctx === undefined ||
@@ -41,90 +123,78 @@ export function resolveCsharpImportTarget(
) {
return null;
}
if (parsedImport.kind === 'dynamic-unresolved') return null;
if (parsedImport.targetRaw === null || parsedImport.targetRaw === '') return null;
return ctx;
}
// Namespace path: `System.Collections.Generic` → `System/Collections/Generic`.
const pathLike = parsedImport.targetRaw.replace(/\./g, '/');
const suffix = `/${pathLike}`;
// Exact file match: `System/Collections/Generic.cs` (rare but legal).
// Suffix match for nested layouts: `src/lib/System/Collections/Generic.cs`.
// Directory match: first `.cs` file directly inside the namespace dir
// (e.g. `System/Collections/Generic/List.cs` matches namespace Generic).
let exactFile: string | null = null;
/**
* First-pass resolution against the full namespace path:
* exact whole-path file > nested suffix file > first `.cs` directly inside
* the namespace directory.
*/
function resolveDirectMatch(allFilePaths: ReadonlySet<string>, pathLike: string): string | null {
const exactName = `${pathLike}.cs`;
const nestedSuffix = `/${exactName}`;
let suffixFile: string | null = null;
let directoryChild: string | null = null;
const dirPrefix = `${pathLike}/`;
const suffixDirPrefix = `/${dirPrefix}`;
for (const raw of ctx.allFilePaths) {
for (const raw of allFilePaths) {
const f = raw.replace(/\\/g, '/');
if (!f.endsWith('.cs')) continue;
if (f === `${pathLike}.cs`) {
exactFile = raw;
break;
}
if (suffixFile === null && f.endsWith(`${suffix}.cs`)) {
suffixFile = raw;
}
if (directoryChild === null) {
// Namespace-to-directory match: pick the first `.cs` directly in
// the namespace dir (not nested deeper). Legacy resolver emits
// all of them; we take one so the scope-resolver contract stays
// single-target.
const atRoot = f.startsWith(dirPrefix);
const atNested = f.includes(suffixDirPrefix);
if (atRoot || atNested) {
const idx = atRoot ? 0 : f.indexOf(suffixDirPrefix) + 1;
const after = f.slice(idx + dirPrefix.length);
if (after.length > 0 && !after.includes('/')) {
directoryChild = raw;
}
}
}
if (f === exactName) return raw; // exact whole-path match wins
if (suffixFile === null && f.endsWith(nestedSuffix)) suffixFile = raw;
}
if (exactFile !== null) return exactFile;
if (suffixFile !== null) return suffixFile;
if (directoryChild !== null) return directoryChild;
return findDirectChild(allFilePaths, pathLike);
}
// Progressive prefix stripping — mirrors csproj's root-namespace
// mapping without the csproj. `using CrossFile.Models;` in a repo
// laid out `Models/User.cs` (no `CrossFile/` prefix) works because
// the legacy resolver consults csproj; the scope-resolver layer
// doesn't have csproj, so we try each suffix of the namespace path
// against `.cs` files and directories.
//
// Also handles `using static CrossFile.Models.UserFactory;` —
// strip the leading segment, try `Models/UserFactory.cs`; strip
// two, try `UserFactory.cs`.
/**
* First `.cs` file that lives directly inside the namespace directory
* `dirSegment` (at repo root or nested under a project prefix), not deeper.
* The legacy resolver emits all of them; the scope-resolver contract is
* single-target so we take one.
*/
function findDirectChild(allFilePaths: ReadonlySet<string>, dirSegment: string): string | null {
const dirPrefix = `${dirSegment}/`;
const nestedDirPrefix = `/${dirPrefix}`;
for (const raw of allFilePaths) {
const f = raw.replace(/\\/g, '/');
if (!f.endsWith('.cs')) continue;
const atRoot = f.startsWith(dirPrefix);
const atNested = f.includes(nestedDirPrefix);
if (!atRoot && !atNested) continue;
const idx = atRoot ? 0 : f.indexOf(nestedDirPrefix) + 1;
const after = f.slice(idx + dirPrefix.length);
if (after.length > 0 && !after.includes('/')) return raw;
}
return null;
}
/**
* Try each suffix of the namespace path against `.cs` files and directories,
* stripping leading segments one at a time. Models `using CrossFile.Models;`
* resolving to `Models/User.cs` in a repo laid out without the `CrossFile/`
* prefix (the scope-resolver layer has no csproj to consult).
*/
function resolveByProgressiveStripping(
allFilePaths: ReadonlySet<string>,
pathLike: string,
): string | null {
const segments = pathLike.split('/').filter(Boolean);
for (let skip = 1; skip < segments.length; skip++) {
const tail = segments.slice(skip).join('/');
if (tail === '') continue;
const tailFile = `${tail}.cs`;
const tailSuffix = `/${tailFile}`;
const tailDir = `${tail}/`;
const tailSuffixDir = `/${tailDir}`;
let tailDirectChild: string | null = null;
for (const raw of ctx.allFilePaths) {
let tailFileMatch: string | null = null;
for (const raw of allFilePaths) {
const f = raw.replace(/\\/g, '/');
if (!f.endsWith('.cs')) continue;
if (f === tailFile) return raw;
if (f.endsWith(tailSuffix)) return raw;
if (tailDirectChild === null) {
const atRoot = f.startsWith(tailDir);
const atNested = f.includes(tailSuffixDir);
if (atRoot || atNested) {
const idx = atRoot ? 0 : f.indexOf(tailSuffixDir) + 1;
const after = f.slice(idx + tailDir.length);
if (after.length > 0 && !after.includes('/')) tailDirectChild = raw;
}
if (f === tailFile || f.endsWith(tailSuffix)) {
tailFileMatch = raw;
break;
}
}
if (tailDirectChild !== null) return tailDirectChild;
if (tailFileMatch !== null) return tailFileMatch;
const child = findDirectChild(allFilePaths, tail);
if (child !== null) return child;
}
return null;
}
@@ -28,17 +28,19 @@
* aliased `using static X = Y.Z;`, attributed namespace declarations,
* and preprocessor-guarded declarations correctly because the
* tree-sitter grammar parses them as real nodes (not textual
* coincidences).
* coincidences). When the orchestrator's `treeCache` has no Tree for a
* file — the worker path, where native Trees can't cross MessageChannels
* — `extractFileStructure` falls back to a line scanner rather than
* re-parsing every file from scratch (that re-parse dominated worker-mode
* scope-resolution time). See `extractCsharpStructureViaScanner`.
*/
import type { SyntaxNode } from 'tree-sitter';
import type { BindingRef, ParsedFile, Scope, ScopeId, SymbolDefinition } from 'gitnexus-shared';
import type { ScopeResolutionIndexes } from '../../model/scope-resolution-indexes.js';
import { getCsharpParser } from './query.js';
import { getTreeSitterBufferSize } from '../../constants.js';
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
interface CsharpFileStructure {
export interface CsharpFileStructure {
/** Declared namespace names in file source order. Empty array means
* the file has no `namespace X;` / `namespace X { }` declaration
* and sits in the default (global) namespace. */
@@ -46,20 +48,253 @@ interface CsharpFileStructure {
/** Dotted paths from `using static X.Y.Z;` (including
* `global using static` and aliased `using static A = X.Y.Z;`). */
readonly usingStaticPaths: readonly string[];
/** True when the scanner saw a `namespace` / `using static` declaration it
* could not fully capture (keyword not at line start, split across lines, or
* an unparseable identifier form). Callers feeding the #1881 gate must treat
* this like a truncated scan and fail OPEN, since a dropped namespace would
* otherwise over-block a legitimate import (Codex F3). Absent/false on a
* cleanly-scanned file. */
readonly incomplete?: boolean;
}
/** Build a structural view of a C# file by walking the tree-sitter
* AST. Prefers `cachedTree` (handed in via `treeCache`) so we don't
* re-parse files the orchestrator already parsed for `extractParsedFile`;
* falls back to a fresh parse on cache miss. Parser singleton is
* shared across calls. */
// A dotted C# namespace identifier: each segment is an optional verbatim `@`
// followed by a Unicode letter/`_` and Unicode letters/digits/`_`. The `u` flag
// makes the classes Unicode-aware so `namespace Café.Models;` is captured (the
// old ASCII `[A-Za-z…]` truncated it). The `@` markers are stripped from the
// capture so it matches the tree-sitter AST's `name` text.
const CS_NS_IDENT = String.raw`@?[\p{L}_][\p{L}\p{N}_]*(?:\.@?[\p{L}_][\p{L}\p{N}_]*)*`;
// Line-anchored matchers for the worker-path fallback (see
// `extractCsharpStructureViaScanner`). Anchored at line start (after
// indentation); the scanner additionally tracks block-comment / string
// state across lines so a keyword at the start of a line inside one of
// those regions is skipped.
const CS_NAMESPACE_RE = new RegExp(String.raw`^[ \t]*namespace[ \t]+(${CS_NS_IDENT})`, 'u');
// `global using static`, plain `using static`, and the aliased
// `using static Alias = NS.Type;` form (the AST keeps the RHS path, so
// the optional `Alias =` is skipped and only the dotted path captured).
const CS_USING_STATIC_RE = new RegExp(
String.raw`^[ \t]*(?:global[ \t]+)?using[ \t]+static[ \t]+(?:@?[\p{L}_][\p{L}\p{N}_]*[ \t]*=[ \t]*)?(${CS_NS_IDENT})`,
'u',
);
// Incompleteness detectors — used ONLY when the precise matchers above failed,
// to flag a declaration the scanner could not capture (so the file fails the
// #1881 gate OPEN instead of silently dropping the namespace). Kept
// high-precision so ordinary files never trip them (which would wrongly disable
// the gate repo-wide):
// - `…_BARE`: the keyword alone on a line (the name is on the next line).
// - `…_AT_START`: a line-start declaration the precise matcher couldn't parse.
// - `CS_NAMESPACE_AFTER_CODE`: a `namespace` keyword right after a `}`/`;`/`{`/`]`
// (real code, NOT a `//` comment), i.e. not at line start.
const CS_NAMESPACE_BARE = /^[ \t]*namespace[ \t]*\r?$/;
const CS_USING_STATIC_BARE = /^[ \t]*(?:global[ \t]+)?using[ \t]+static[ \t]*\r?$/;
const CS_NAMESPACE_AT_START = /^[ \t]*namespace[ \t]+\S/;
const CS_USING_STATIC_AT_START = /^[ \t]*(?:global[ \t]+)?using[ \t]+static[ \t]+\S/;
const CS_NAMESPACE_AFTER_CODE = /[}\];{][ \t]*namespace[ \t]+@?[\p{L}_]/u;
/** Whether a `code`-state line declares a namespace / using-static the precise
* matchers could not capture — see the detectors above. */
function looksLikeUncapturedDeclaration(line: string): boolean {
return (
CS_NAMESPACE_BARE.test(line) ||
CS_USING_STATIC_BARE.test(line) ||
CS_NAMESPACE_AT_START.test(line) ||
CS_USING_STATIC_AT_START.test(line) ||
CS_NAMESPACE_AFTER_CODE.test(line)
);
}
/** Multi-line lexical state carried line-to-line by the scanner. */
type CsScanState = 'code' | 'block' | 'verbatim' | 'raw';
/** Advance the scanner's lexical state across one line, consuming block
* comments (slash-star), line comments (`//`), single-line regular /
* interpolated strings, verbatim strings (`@"…"`), and raw string literals
* (`"""…"""`, fence length tracked in `rawFence`). Returns the state and
* raw-fence length in effect at the START of the next line. Single-line
* strings and `//` comments resolve back to `code` before end of line; only
* block comments and multi-line strings carry state forward. */
function advanceCsScanState(
line: string,
state: CsScanState,
rawFence: number,
): [CsScanState, number] {
const n = line.length;
let i = 0;
while (i < n) {
if (state === 'block') {
const end = line.indexOf('*/', i);
if (end === -1) return ['block', rawFence];
i = end + 2;
state = 'code';
} else if (state === 'verbatim') {
// Ends at a `"` that is not doubled (`""` is an escaped quote).
while (i < n) {
if (line[i] === '"') {
if (line[i + 1] === '"') {
i += 2;
continue;
}
break;
}
i++;
}
if (i >= n) return ['verbatim', rawFence];
i += 1;
state = 'code';
} else if (state === 'raw') {
// Ends at a run of `"` at least `rawFence` long.
let closed = false;
while (i < n) {
if (line[i] === '"') {
let k = i;
while (k < n && line[k] === '"') k++;
if (k - i >= rawFence) {
i = k;
state = 'code';
rawFence = 0;
closed = true;
break;
}
i = k;
} else {
i++;
}
}
if (!closed) return ['raw', rawFence];
} else {
const c = line[i];
const next = line[i + 1];
if (c === '/' && next === '/') return ['code', rawFence]; // line comment to EOL
if (c === '/' && next === '*') {
state = 'block';
i += 2;
} else if (c === '@' && next === '"') {
state = 'verbatim';
i += 2;
} else if ((c === '$' && next === '@') || (c === '@' && next === '$')) {
if (line[i + 2] === '"') {
state = 'verbatim'; // interpolated verbatim ($@"…" / @$"…")
i += 3;
} else {
i++;
}
} else if (c === '"') {
let k = i;
while (k < n && line[k] === '"') k++;
const run = k - i;
if (run >= 3) {
state = 'raw';
rawFence = run;
i = k;
} else if (run === 2) {
i = k; // "" — empty string
} else {
// single-line regular / interpolated string; consume to closer
let j = i + 1;
while (j < n) {
if (line[j] === '\\') {
j += 2;
continue;
}
if (line[j] === '"') break;
j++;
}
i = j >= n ? n : j + 1;
}
} else {
i++;
}
}
}
return [state, rawFence];
}
/** Line-scanner used when no cached tree is available (worker-parsed files
* can't transfer native tree-sitter Trees across MessageChannels, so
* `treeCache` is empty for them). Re-parsing every C# file here with
* tree-sitter was the dominant scope-resolution cost on large worker-mode
* runs — for a multi-thousand-file solution this loop alone re-parsed the
* whole repo a second time. The scanner extracts the same `namespaces` /
* `usingStaticPaths` the AST walk produces for line-anchored declarations,
* while tracking block-comment and string state across lines (via
* `advanceCsScanState`) so a `namespace` / `using static` keyword at the
* start of a line inside a block comment, verbatim string, or raw string
* literal is NOT mistaken for a declaration. The remaining trade-off vs the
* AST is a declaration whose keyword is not at the start of a code line
* (split across lines, or sharing a line with a comment/string closer).
* Mirrors PHP's `extractNamespaceViaScanner` (issue #1741). */
/** Incremental form of {@link extractCsharpStructureViaScanner}: feed lines one
* at a time via `pushLine` (in source order), then read the accumulated
* structure with `result()`. Lets a caller stream a file off disk
* (`createReadStream` + `readline`) and scan it for `namespace` / `using
* static` declarations in CONSTANT memory rather than buffering the whole file
* into a string — the line splitting and per-line matching are identical, so a
* streamed scan yields the same result as scanning the full content. The line
* terminator must be stripped (as `readline` does, or `String.split('\n')`); a
* trailing `\r` on a CRLF line is inert to both the matchers and the lexer. */
export interface CsharpStructureLineScanner {
pushLine(line: string): void;
result(): CsharpFileStructure;
}
/** Create a fresh stateful line scanner — see {@link CsharpStructureLineScanner}. */
export function createCsharpStructureScanner(): CsharpStructureLineScanner {
const namespaces: string[] = [];
const usingStaticPaths: string[] = [];
let incomplete = false;
let state: CsScanState = 'code';
let rawFence = 0;
return {
pushLine(line: string): void {
// Only match when the line START is real code — keywords reached while
// inside a block comment / multi-line string are skipped.
if (state === 'code') {
const ns = CS_NAMESPACE_RE.exec(line);
if (ns !== null) {
namespaces.push(ns[1]!.replace(/@/g, ''));
} else {
const us = CS_USING_STATIC_RE.exec(line);
if (us !== null) {
usingStaticPaths.push(us[1]!.replace(/@/g, ''));
} else if (looksLikeUncapturedDeclaration(line)) {
// A declaration the precise matchers couldn't capture → mark the
// file incomplete so the #1881 gate fails OPEN (Codex F3).
incomplete = true;
}
}
}
[state, rawFence] = advanceCsScanState(line, state, rawFence);
},
result(): CsharpFileStructure {
return incomplete
? { namespaces, usingStaticPaths, incomplete }
: { namespaces, usingStaticPaths };
},
};
}
export function extractCsharpStructureViaScanner(content: string): CsharpFileStructure {
const scanner = createCsharpStructureScanner();
for (const line of content.split('\n')) scanner.pushLine(line);
return scanner.result();
}
/** Build a structural view of a C# file. Prefers `cachedTree` (handed in
* via `treeCache`) and walks the tree-sitter AST — the authoritative
* path that sees `global using static`, aliased `using static X = Y.Z;`,
* attributed namespace declarations, and preprocessor-guarded nodes
* correctly. On cache miss (worker-parsed files, whose native Trees
* can't cross MessageChannels) it falls back to the line scanner instead
* of a fresh tree-sitter parse — the parse here dominated worker-mode
* scope-resolution time. Parser singleton is shared across calls. */
function extractFileStructure(content: string, cachedTree: unknown): CsharpFileStructure {
if (!cachedTree) {
return extractCsharpStructureViaScanner(content);
}
type CsharpTree = ReturnType<ReturnType<typeof getCsharpParser>['parse']>;
const tree =
(cachedTree as CsharpTree | undefined) ??
parseSourceSafe(getCsharpParser(), content, undefined, {
bufferSize: getTreeSitterBufferSize(content),
});
const tree = cachedTree as CsharpTree;
const namespaces: string[] = [];
const usingStaticPaths: string[] = [];
@@ -277,11 +512,17 @@ export function populateCsharpNamespaceSiblings(
// scope, so `Record(...)` (without `Logger.` qualifier) resolves
// to `Logger.Record`. AST walk above captured these (including
// `global using static` and aliased forms).
// Pre-index files by path once: the member-injection lookup below would
// otherwise be an O(files) scan per `using static` import.
const fileByPath = new Map<string, ParsedFile>(parsedFiles.map((p) => [p.filePath, p]));
for (const parsed of parsedFiles) {
const struct = structureByFile.get(parsed.filePath);
if (struct === undefined) continue;
const moduleScope = parsed.scopes.find((s) => s.kind === 'Module');
if (moduleScope === undefined) continue;
// Per-file de-dup sets keyed by simple name, seeded lazily from the
// augmentation bucket — replaces the per-member O(A) `.some` scan below.
const seenByName = new Map<string, Set<string>>();
for (const fullPath of struct.usingStaticPaths) {
const lastDot = fullPath.lastIndexOf('.');
@@ -302,7 +543,7 @@ export function populateCsharpNamespaceSiblings(
// Inject the class's member methods into the importer's module
// scope. `memberByOwner` wasn't built yet here, so we walk the
// file's localDefs to find members with `ownerId === targetDef.nodeId`.
const targetFile = parsedFiles.find((p) => p.filePath === targetDef.filePath);
const targetFile = fileByPath.get(targetDef.filePath);
if (targetFile === undefined) continue;
for (const memberDef of targetFile.localDefs) {
if ((memberDef as { ownerId?: string }).ownerId !== targetDef.nodeId) continue;
@@ -316,7 +557,14 @@ export function populateCsharpNamespaceSiblings(
// `lookupBindingsAt`, which fans out across `bindings` +
// `bindingAugmentations`.
const bucketArr = getAugmentationBucket(augmentations, moduleScope.id, simpleName);
if (bucketArr.some((b) => b.def.nodeId === memberDef.nodeId)) continue;
let seen = seenByName.get(simpleName);
if (seen === undefined) {
seen = new Set<string>();
for (const b of bucketArr) seen.add(b.def.nodeId);
seenByName.set(simpleName, seen);
}
if (seen.has(memberDef.nodeId)) continue;
seen.add(memberDef.nodeId);
bucketArr.push({ def: memberDef, origin: 'import' });
}
}
@@ -332,6 +580,9 @@ export function populateCsharpNamespaceSiblings(
for (const parsed of parsedFiles) {
const moduleScope = parsed.scopes.find((s) => s.kind === 'Module');
if (moduleScope === undefined) continue;
// Per-file de-dup sets keyed by simple name, seeded lazily from the
// augmentation bucket — replaces the per-def O(A) `.some` scan below.
const seenByName = new Map<string, Set<string>>();
for (const imp of parsed.parsedImports) {
if (imp.kind !== 'namespace') continue;
const targetNs = imp.targetRaw;
@@ -344,41 +595,113 @@ export function populateCsharpNamespaceSiblings(
const simpleName = q.includes('.') ? q.slice(q.lastIndexOf('.') + 1) : q;
if (simpleName === '') continue;
const bucketArr = getAugmentationBucket(augmentations, moduleScope.id, simpleName);
if (bucketArr.some((b) => b.def.nodeId === def.nodeId)) continue;
let seen = seenByName.get(simpleName);
if (seen === undefined) {
seen = new Set<string>();
for (const b of bucketArr) seen.add(b.def.nodeId);
seenByName.set(simpleName, seen);
}
if (seen.has(def.nodeId)) continue;
seen.add(def.nodeId);
bucketArr.push({ def, origin: 'namespace' });
}
}
}
for (const [, bucket] of buckets) {
// De-dup by (nodeId, filePath) across multiple declarations (e.g.
// partial classes declaring the same name in two files — we take
// both and leave de-dup to downstream consumers of bindings).
// Workspace-level binding channel for global-namespace types (see the
// global fast-path below). `lookupBindingsAt` consults this as a third
// source after finalized + per-scope augmented bindings. Its inner arrays
// are mutable by contract (append-only, like `bindingAugmentations` — see
// the ScopeResolutionIndexes doc + validateBindingsImmutability), so the
// ReadonlyMap→Map cast is localized to this one line and all writes go
// through `getWorkspaceBucket`.
const workspace = indexes.workspaceFqnBindings as Map<string, BindingRef[]>;
for (const [nsName, bucket] of buckets) {
// Group sibling defs by simple name. Append in place — the previous
// `[...prev, def]` copy made this O(D²) per bucket, which on the
// global (`''`) namespace bucket of a large Unity solution (tens of
// thousands of type defs) was a primary slowness/OOM source. We keep
// every declaration (e.g. partial classes across files) and leave
// de-dup to downstream consumers.
const defsByName = new Map<string, SymbolDefinition[]>();
for (const def of bucket.classDefs) {
// Simple name = last segment of qualifiedName (e.g. `App.User` → `User`).
const q = def.qualifiedName ?? '';
const key = q.includes('.') ? q.slice(q.lastIndexOf('.') + 1) : q;
if (key === '') continue;
const arr = [...(defsByName.get(key) ?? [])];
let arr = defsByName.get(key);
if (arr === undefined) {
arr = [];
defsByName.set(key, arr);
}
arr.push(def);
defsByName.set(key, arr);
}
// Global-namespace fast path (Unity OOM guard). Types declared in the
// default (global) namespace are visible from EVERY file in C# — the
// global namespace is always implicitly in scope — so one workspace-
// level entry per simple name is both semantically correct and O(D)
// instead of the O(S·D) per-scope augmentation that materialized
// billions of BindingRefs on large Unity solutions (tens of thousands
// of global types × tens of thousands of scopes). `walkScopeChain`
// checks local `scope.bindings` first, so local declarations still
// shadow these workspace entries; a file resolving its own global type
// hits the local binding before this map. Dedup by `def.nodeId` keeps
// partial-class / duplicate declarations from double-emitting.
if (nsName === '') {
for (const [name, defs] of defsByName) {
const bucket = getWorkspaceBucket(workspace, name);
const seen = new Set<string>();
for (const b of bucket) seen.add(b.def.nodeId);
for (const def of defs) {
if (seen.has(def.nodeId)) continue; // dedup by nodeId (keeps partials, drops re-emits)
seen.add(def.nodeId);
bucket.push({ def, origin: 'namespace' });
}
}
continue;
}
// Pre-index the first scope per file once (O(S)) instead of an
// O(S) `.find` re-run for every (scope, name) pair, which made the
// injection loop O(S²·D) and was the dominant cost on large buckets.
// Multiple scopes share a filePath (Module + Namespace); the local
// shadow check only needs that file's lexical `Scope.bindings`, which
// is identical regardless of which of those scopes we read.
const firstScopeByFile = new Map<string, Scope>();
for (const s of bucket.scopes) {
if (!firstScopeByFile.has(s.filePath)) firstScopeByFile.set(s.filePath, s.scope);
}
for (const { scopeId, filePath } of bucket.scopes) {
const localScope = firstScopeByFile.get(filePath);
for (const [name, defs] of defsByName) {
// Skip names already present locally — `origin: 'local'` in
// scope.bindings would naturally shadow the cross-file
// namespace entry, but we also keep this index lean.
const local = bucket.scopes.find((s) => s.filePath === filePath)?.scope.bindings.get(name);
const local = localScope?.bindings.get(name);
if (local !== undefined && local.some((b) => b.origin === 'local')) continue;
let bucketArr: BindingRef[] | null = null;
// Bind the augmentation bucket and its seeded de-dup set together
// under one nullable lifecycle, so neither needs a non-null
// assertion (they are always set or unset as a pair). Stays lazy:
// nothing is allocated for a name with no cross-file defs.
let inject: { bucket: BindingRef[]; seen: Set<string> } | null = null;
for (const def of defs) {
if (def.filePath === filePath) continue; // don't self-reference
if (bucketArr === null) bucketArr = getAugmentationBucket(augmentations, scopeId, name);
if (bucketArr.some((b) => b.def.nodeId === def.nodeId)) continue;
bucketArr.push({ def, origin: 'namespace' });
if (inject === null) {
const bucket = getAugmentationBucket(augmentations, scopeId, name);
// Seed the de-dup set from any entries an earlier pass
// (using-static / cross-namespace imports) already added,
// replacing the per-def O(A) `.some` scan.
const seen = new Set<string>();
for (const b of bucket) seen.add(b.def.nodeId);
inject = { bucket, seen };
}
if (inject.seen.has(def.nodeId)) continue;
inject.seen.add(def.nodeId);
inject.bucket.push({ def, origin: 'namespace' });
}
}
}
@@ -409,6 +732,22 @@ function getAugmentationBucket(
return bucketArr;
}
/** Get-or-create a mutable inner bucket inside the `workspaceFqnBindings`
* channel (the scope-independent third channel; see
* `ScopeResolutionIndexes.workspaceFqnBindings`). Like
* `getAugmentationBucket`, the inner arrays are mutable by contract —
* callers `push` directly. Keeping the get-or-create here means the one
* ReadonlyMap→Map cast at the call site is the only place the mutable
* view is taken. */
function getWorkspaceBucket(workspace: Map<string, BindingRef[]>, name: string): BindingRef[] {
let bucketArr = workspace.get(name);
if (bucketArr === undefined) {
bucketArr = [];
workspace.set(name, bucketArr);
}
return bucketArr;
}
function isTypeDef(def: SymbolDefinition): boolean {
return (
def.type === 'Class' ||
@@ -0,0 +1,30 @@
/**
* Per-workspace config for C# scope-resolution import targeting.
*
* Loaded once per analyze pass via `csharpScopeResolver.loadResolutionConfig`
* and threaded into `resolveCsharpImportTarget`. The pure gate predicates live
* in `../../csharp-namespace-gate.ts` (shared with the legacy DAG resolver).
*/
import {
scanCSharpProject,
csharpScanToEvidence,
type CSharpProjectConfig,
type CSharpNamespaceEvidence,
} from '../../language-config.js';
export interface CsharpResolutionConfig {
readonly csharpConfigs: readonly CSharpProjectConfig[];
/** In-repo declared-namespace evidence gating suffix-fallback resolution (#1881). */
readonly namespaces?: CSharpNamespaceEvidence;
}
export async function loadCsharpResolutionConfig(
repoRoot: string,
): Promise<CsharpResolutionConfig> {
const scan = await scanCSharpProject(repoRoot);
return {
csharpConfigs: scan.configs,
namespaces: csharpScanToEvidence(scan),
};
}
@@ -19,6 +19,7 @@ import {
type CsharpResolveContext,
} from './index.js';
import { populateCsharpNamespaceSiblings } from './namespace-siblings.js';
import { loadCsharpResolutionConfig, type CsharpResolutionConfig } from './resolution-config.js';
import { unwrapCsharpCollectionAccessor } from './accessor-unwrap.js';
const csharpScopeResolver: ScopeResolver = {
@@ -26,8 +27,16 @@ const csharpScopeResolver: ScopeResolver = {
languageProvider: csharpProvider,
importEdgeReason: 'csharp-scope: using',
resolveImportTarget: (targetRaw, fromFile, allFilePaths) => {
const ws: CsharpResolveContext = { fromFile, allFilePaths };
loadResolutionConfig: (repoPath) => loadCsharpResolutionConfig(repoPath),
resolveImportTarget: (targetRaw, fromFile, allFilePaths, resolutionConfig) => {
const config = resolutionConfig as CsharpResolutionConfig | undefined;
const ws: CsharpResolveContext = {
fromFile,
allFilePaths,
csharpConfigs: config?.csharpConfigs,
namespaces: config?.namespaces,
};
// `WorkspaceIndex` is an opaque `unknown` placeholder in the
// shared contract, so `ws` passes structurally without a cast.
return resolveCsharpImportTarget(
@@ -39,6 +39,56 @@ import {
interpretGoTypeBinding,
} from './go/index.js';
const GO_BUILT_INS: ReadonlySet<string> = new Set([
// built-in functions
'make',
'new',
'len',
'cap',
'append',
'copy',
'delete',
'close',
'panic',
'recover',
'print',
'println',
'complex',
'real',
'imag',
'clear',
'min',
'max',
// built-in types
'error',
'bool',
'string',
'int',
'int8',
'int16',
'int32',
'int64',
'uint',
'uint8',
'uint16',
'uint32',
'uint64',
'uintptr',
'float32',
'float64',
'complex64',
'complex128',
'byte',
'rune',
'any',
'comparable',
// built-in values
'true',
'false',
'nil',
'iota',
]);
export const goProvider = defineLanguage({
id: SupportedLanguages.Go,
extensions: ['.go'],
@@ -92,6 +142,7 @@ export const goProvider = defineLanguage({
variableExtractor: createVariableExtractor(goVariableConfig),
classExtractor: createClassExtractor(goClassConfig),
heritageExtractor: createHeritageExtractor(goHeritageConfig),
builtInNames: GO_BUILT_INS,
// ── RFC #909 Ring 3: scope-based resolution hooks ──────────
emitScopeCaptures: emitGoScopeCaptures,
@@ -1,10 +1,5 @@
import type { Capture, CaptureMatch } from 'gitnexus-shared';
import {
findNodeAtRange,
nodeToCapture,
syntheticCapture,
type SyntaxNode,
} from '../../utils/ast-helpers.js';
import { nodeToCapture, syntheticCapture, type SyntaxNode } from '../../utils/ast-helpers.js';
import { getGoParser, getGoScopeQuery } from './query.js';
import { recordGoCacheHit, recordGoCacheMiss } from './cache-stats.js';
import { computeGoCallArity, computeGoDeclarationArity } from './arity-metadata.js';
@@ -34,18 +29,29 @@ export function emitGoScopeCaptures(
for (const m of rawMatches) {
const grouped: Record<string, Capture> = {};
// Parallel tag -> captured SyntaxNode map. The tree-sitter query already
// hands us the matched node as `c.node`; keeping it here lets us derive the
// anchor/relative node by walking LOCALLY (parent chain / own subtree)
// instead of re-walking from tree.rootNode (the O(matches x rootChildren)
// hotpath that made #1848's 250-struct DAO file take ~10s). The captured
// node either IS the node the old findNodeAtRange re-derived, or is a close
// relative reachable by a bounded local walk.
const nodeMap: Record<string, SyntaxNode> = {};
for (const c of m.captures) {
const tag = '@' + c.name;
if (tag.startsWith('@_')) continue; // skip anonymous captures
grouped[tag] = nodeToCapture(tag, c.node);
nodeMap[tag] = c.node;
}
if (Object.keys(grouped).length === 0) continue;
if (grouped['@import.statement'] !== undefined) {
const anchor = grouped['@import.statement']!;
const importNode =
findNodeAtRange(tree.rootNode, anchor.range, 'import_declaration') ??
findNodeAtRange(tree.rootNode, anchor.range, 'import_spec');
// The captured node is the `import_spec`; the original code preferred its
// enclosing `import_declaration` ONLY when that ancestor shares the exact
// same range (which never happens — the declaration always includes the
// `import` keyword prefix — so it falls back to the import_spec itself).
// Replicate that exactly via a local ancestor walk, never from root.
const importNode = resolveImportNode(nodeMap['@import.statement']!);
if (importNode !== null) {
out.push(...splitGoImportStatement(importNode));
continue;
@@ -53,23 +59,33 @@ export function emitGoScopeCaptures(
}
if (grouped['@scope.function'] !== undefined) {
const scopeCap = grouped['@scope.function']!;
// @scope.function captures function_declaration | method_declaration |
// func_literal. The original looked for a function_declaration or
// method_declaration at the captured range; the captured node IS that
// node for the first two, and a func_literal never coincides in range
// with either, so the lookup yields null for func_literal.
const scopeNode = nodeMap['@scope.function']!;
const fnNode =
findNodeAtRange(tree.rootNode, scopeCap.range, 'function_declaration') ??
findNodeAtRange(tree.rootNode, scopeCap.range, 'method_declaration');
scopeNode.type === 'function_declaration' || scopeNode.type === 'method_declaration'
? scopeNode
: null;
if (fnNode !== null) {
const receiver = synthesizeGoReceiverBinding(fnNode);
if (receiver !== null) out.push(receiver);
}
}
if (isRawMultiAssignTypeBinding(tree.rootNode, grouped)) continue;
if (isRawMultiAssignTypeBinding(nodeMap)) continue;
const declAnchor = grouped['@declaration.function'] ?? grouped['@declaration.method'];
if (declAnchor !== undefined) {
const declAnchorNode = nodeMap['@declaration.function'] ?? nodeMap['@declaration.method'];
if (declAnchorNode !== undefined) {
// @declaration.function / @declaration.method are captured directly on
// the function_declaration / method_declaration node.
const fnNode =
findNodeAtRange(tree.rootNode, declAnchor.range, 'function_declaration') ??
findNodeAtRange(tree.rootNode, declAnchor.range, 'method_declaration');
declAnchorNode.type === 'function_declaration' ||
declAnchorNode.type === 'method_declaration'
? declAnchorNode
: null;
if (fnNode !== null) {
const arity = computeGoDeclarationArity(fnNode);
if (arity.parameterCount !== undefined) {
@@ -98,15 +114,15 @@ export function emitGoScopeCaptures(
continue;
}
const callAnchor =
grouped['@reference.call.free'] ??
grouped['@reference.call.member'] ??
grouped['@reference.call.constructor'];
if (callAnchor !== undefined && grouped['@reference.arity'] === undefined) {
const callNode =
findNodeAtRange(tree.rootNode, callAnchor.range, 'call_expression') ??
findNodeAtRange(tree.rootNode, callAnchor.range, 'composite_literal');
if (callNode !== null) {
// @reference.call.free / .member are captured on the call_expression;
// @reference.call.constructor on the composite_literal. The captured node
// IS the node the old findNodeAtRange re-derived for each, so use it.
const callNode =
nodeMap['@reference.call.free'] ??
nodeMap['@reference.call.member'] ??
nodeMap['@reference.call.constructor'];
if (callNode !== undefined && grouped['@reference.arity'] === undefined) {
if (callNode.type === 'call_expression' || callNode.type === 'composite_literal') {
grouped['@reference.arity'] = syntheticCapture(
'@reference.arity',
callNode,
@@ -146,18 +162,56 @@ export function emitGoScopeCaptures(
return out;
}
function isRawMultiAssignTypeBinding(
rootNode: SyntaxNode,
grouped: Record<string, Capture>,
): boolean {
/**
* Resolve the node passed to `splitGoImportStatement` for an @import.statement
* match. The capture is on the `import_spec`; the original preferred an
* `import_declaration` at the SAME range, else the import_spec. An
* import_declaration always includes the `import` keyword and so never shares
* the spec's exact range — the only candidate is an ancestor, and it can only
* match when ranges coincide. Walk the parent chain (bounded, local) for an
* import_declaration whose range equals the spec's; otherwise return the spec.
*/
function resolveImportNode(importSpec: SyntaxNode): SyntaxNode {
let current: SyntaxNode | null = importSpec.parent;
while (current !== null) {
if (current.type === 'import_declaration') {
if (nodeRangeEquals(current, importSpec)) return current;
break;
}
// import_spec is nested at most under import_declaration ->
// import_spec_list -> import_spec; stop once we leave the import subtree.
if (current.type !== 'import_spec_list') break;
current = current.parent;
}
return importSpec;
}
/** True iff two nodes occupy the exact same source range. */
function nodeRangeEquals(a: SyntaxNode, b: SyntaxNode): boolean {
return (
a.startPosition.row === b.startPosition.row &&
a.startPosition.column === b.startPosition.column &&
a.endPosition.row === b.endPosition.row &&
a.endPosition.column === b.endPosition.column
);
}
function isRawMultiAssignTypeBinding(nodeMap: Record<string, SyntaxNode>): boolean {
const anchor =
grouped['@type-binding.constructor'] ??
grouped['@type-binding.call-return'] ??
grouped['@type-binding.assertion'];
nodeMap['@type-binding.constructor'] ??
nodeMap['@type-binding.call-return'] ??
nodeMap['@type-binding.assertion'];
if (anchor === undefined) return false;
const node = findNodeAtRange(rootNode, anchor.range, 'short_var_declaration');
if (node === null) return false;
// These tags are captured directly ON the short_var_declaration, so the
// captured node IS what the original findNodeAtRange(root, range,
// 'short_var_declaration') re-derived. The var_declaration (var-form)
// variants — @type-binding.assertion (`var x = e.(T)`) and
// @type-binding.call-return (`var x = Func()`) — anchor on a var_declaration
// instead; the old range+type lookup found no short_var_declaration at that
// range and returned null -> false, which this type guard reproduces exactly.
if (anchor.type !== 'short_var_declaration') return false;
const node = anchor;
const lhs = node.childForFieldName('left');
const rhs = node.childForFieldName('right');
if (lhs === null) return false;
@@ -38,6 +38,7 @@ import { splitImportStatement } from '../typescript/import-decomposer.js';
import { getJsParser, getJsScopeQuery, jsCachedTreeMatchesGrammar } from './query.js';
import { computeTsArityMetadata } from '../typescript/arity-metadata.js';
import { synthesizeTsReceiverBinding } from '../typescript/receiver-binding.js';
import { isArrayMethodCallbackArrow } from '../typescript/array-callback.js';
import { getTreeSitterBufferSize } from '../../constants.js';
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
@@ -640,6 +641,21 @@ export function emitJsScopeCaptures(
}
}
// #1876: drop @declaration.function for array higher-order-method
// callbacks (`const x = arr.map(a => …)`). The HOC-wrapped-arrow
// pattern matches them, but the binding holds a value, not a callable.
// The binding keeps its separate @declaration.const / .variable match,
// and the arrow's own @scope.function match (a different pattern) is
// untouched, so inner-call attribution falls through to the enclosing
// scope instead of a phantom Function.
const fnDeclAnchor = grouped['@declaration.function'];
if (fnDeclAnchor !== undefined) {
const arrowNode = findFunctionNode(tree.rootNode, fnDeclAnchor.range);
if (arrowNode !== null && isArrayMethodCallbackArrow(arrowNode)) {
continue;
}
}
// Synthesize arity metadata on function-like declarations.
const declAnchor = pickFirstDefined(grouped, FUNCTION_DECL_TAGS);
if (declAnchor !== undefined) {
@@ -148,6 +148,12 @@ const JAVASCRIPT_SCOPE_QUERY = `
;; HOC-wrapped variable declarations: const X = HOC((args) => { ... }).
;; Covers React.forwardRef, memo, useCallback, useMemo, observer,
;; debounce, and any user-defined HOC factory.
;;
;; #1876: this shape also matches array higher-order-method callbacks
;; (const x = arr.map(a => ...)), where x is a value, not a function.
;; Those are filtered out emit-side in captures.ts via
;; isArrayMethodCallbackArrow (member-expression callee whose property
;; is a known Array method), so only the @declaration.const survives.
(lexical_declaration
(variable_declarator
name: (identifier) @declaration.name
@@ -21,6 +21,7 @@ import { findNodeAtRange, nodeToCapture, syntheticCapture } from '../../utils/as
import { splitImportStatement } from './import-decomposer.js';
import { getPythonParser, getPythonScopeQuery } from './query.js';
import { synthesizeReceiverTypeBinding } from './receiver-binding.js';
import { synthesizeDependsReferences } from './depends-references.js';
import { computePythonArityMetadata } from './arity-metadata.js';
import { recordCacheHit, recordCacheMiss } from './cache-stats.js';
import { getTreeSitterBufferSize } from '../../constants.js';
@@ -98,6 +99,7 @@ export function emitPythonScopeCaptures(
if (fnNode !== null) {
const synth = synthesizeReceiverTypeBinding(fnNode);
if (synth !== null) out.push(synth);
for (const depRef of synthesizeDependsReferences(fnNode)) out.push(depRef);
}
continue;
}
@@ -0,0 +1,72 @@
/**
* Synthesize `@reference.call.free` captures for FastAPI `Depends(callable)`
* parameter defaults.
*
* `Depends(get_db)` passes `get_db` as a callable that the DI framework
* calls on every request. The route handler is functionally a caller of
* the dependency — impact analysis needs that edge.
*
* Tree-sitter can't express "the first argument of a call named Depends
* inside a parameter default" in a single static query, so we synthesize
* reference captures in code, mirroring the receiver-binding pattern.
*/
import type { CaptureMatch } from 'gitnexus-shared';
import { nodeToCapture, type SyntaxNode } from '../../utils/ast-helpers.js';
/**
* Inspect a `function_definition` node's parameters for `Depends(callable)`
* defaults. Returns one `@reference.call.free` CaptureMatch per dependency.
*/
export function synthesizeDependsReferences(fnNode: SyntaxNode): readonly CaptureMatch[] {
const params = fnNode.childForFieldName('parameters');
if (params === null) return [];
const results: CaptureMatch[] = [];
for (let i = 0; i < params.namedChildCount; i++) {
const param = params.namedChild(i);
if (param === null) continue;
if (param.type !== 'typed_default_parameter' && param.type !== 'default_parameter') {
continue;
}
const defaultValue = param.childForFieldName('value') ?? param.childForFieldName('default');
if (defaultValue === null) continue;
const callNode = defaultValue.type === 'call' ? defaultValue : null;
if (callNode === null) continue;
const fnIdent = callNode.childForFieldName('function');
if (fnIdent === null || fnIdent.type !== 'identifier' || fnIdent.text !== 'Depends') continue;
const args = callNode.childForFieldName('arguments');
if (args === null || args.namedChildCount === 0) continue;
const firstArg = args.namedChild(0);
if (firstArg === null) continue;
if (firstArg.type === 'identifier') {
results.push({
'@reference.call.free': nodeToCapture('@reference.call.free', firstArg),
'@reference.name': nodeToCapture('@reference.name', firstArg),
});
continue;
}
if (firstArg.type === 'attribute') {
const attrName = firstArg.childForFieldName('attribute');
const obj = firstArg.childForFieldName('object');
if (attrName !== null && obj !== null) {
results.push({
'@reference.call.member': nodeToCapture('@reference.call.member', attrName),
'@reference.name': nodeToCapture('@reference.name', attrName),
'@reference.receiver': nodeToCapture('@reference.receiver', obj),
});
}
}
}
return results;
}
@@ -0,0 +1,98 @@
/**
* Array higher-order-method callback detection (issue #1876).
*
* The HOC-wrapped-arrow declaration pattern in the JS/TS scope queries
* (`const X = call((args) => …)`) was added for React idioms
* (`forwardRef` / `memo` / `useCallback`). It has the same AST shape as
* an array higher-order-method call (`const x = arr.map(a => …)`), so
* those callbacks also match and produce a spurious `@declaration.function`
* named after the binding — duplicating the `@declaration.const` /
* `@declaration.variable` def that the same binding already gets.
*
* For an array-method callback the binding holds a *value* (the method's
* result), not a callable, so the `Function` def is semantically wrong.
* `isArrayMethodCallbackArrow` lets the emitter (`captures.ts`) drop that
* `@declaration.function` match, leaving only the value def.
*
* Shared by both the JavaScript and TypeScript capture emitters — the
* relevant grammar nodes (`arrow_function`, `function_expression`,
* `arguments`, `call_expression`, `member_expression`,
* `property_identifier`) are identical across `tree-sitter-javascript`
* and `tree-sitter-typescript`.
*
* Pure given the input node. No I/O, no globals.
*/
import type { SyntaxNode } from '../../utils/ast-helpers.js';
/**
* Array prototype higher-order methods whose result is a value, not a
* function. A callback passed to one of these is an anonymous callback,
* never a top-level function definition. Identifier-callee HOCs
* (`forwardRef(...)`, `useCallback(...)`, custom factories) are
* deliberately NOT listed — they keep their `Function` classification.
*
* Trade-off (unchanged from before #1876): a custom *fluent-API* member
* call with a callback whose method name is not in this set
* (`qb.where(x => …)`) still classifies as `Function`. There is no clean
* syntactic line beyond the well-known Array surface, so the set is
* intentionally closed and easy to extend.
*
* Receiver-blind, by design: the match keys on the method NAME only, never
* the receiver type (tree-sitter has no type information here). So an in-set
* name on a NON-array receiver — `Map`/`Set` `.forEach`, an RxJS
* `observable.map(…)`, a query builder `.sort(…)`, a lodash chain
* `.filter(…)` — is ALSO treated as a callback and has its
* `@declaration.function` dropped. This is an accepted limitation, not a
* regression: those bindings hold the call's *result value*, not a callable,
* so a value def is the correct classification anyway. The only genuine loss
* is a bespoke DSL whose in-set-named method returns something callable —
* rare enough to accept rather than guard with type inference. Pinned by the
* "in-set method on a non-array receiver" case in `*-captures.test.ts`.
*/
export const ARRAY_CALLBACK_METHODS: ReadonlySet<string> = new Set([
'map',
'filter',
'find',
'findIndex',
'findLast',
'findLastIndex',
'forEach',
'reduce',
'reduceRight',
'some',
'every',
'flatMap',
'sort',
]);
/**
* True when `node` (an `arrow_function` / `function_expression`) is the
* callback argument of an array higher-order-method call, i.e. the
* enclosing call's callee is a `member_expression` whose property is one
* of {@link ARRAY_CALLBACK_METHODS}.
*
* Returns false for direct assignments (`const fn = () => {}` — parent is
* `variable_declarator`, not `arguments`) and for identifier-callee HOCs
* (`forwardRef(() => …)` — callee is an `identifier`, not a
* `member_expression`), so neither is ever suppressed.
*
* Intentional non-suppressing gaps (preserve current behavior, no
* regression): parenthesized callee `(arr.map)(cb)` (`parenthesized_expression`)
* and computed callee `arr['map'](cb)` (`subscript_expression`).
*/
export function isArrayMethodCallbackArrow(node: SyntaxNode): boolean {
const args = node.parent;
if (args === null || args.type !== 'arguments') return false;
const call = args.parent;
if (call === null || call.type !== 'call_expression') return false;
const callee = call.childForFieldName('function');
if (callee === null || callee.type !== 'member_expression') return false;
const property = callee.childForFieldName('property');
if (property === null || property.type !== 'property_identifier') return false;
return ARRAY_CALLBACK_METHODS.has(property.text);
}
@@ -37,6 +37,7 @@ import { getTsParser, getTsScopeQuery, tsCachedTreeMatchesGrammar } from './quer
import { recordCacheHit, recordCacheMiss } from './cache-stats.js';
import { synthesizeTsReceiverBinding } from './receiver-binding.js';
import { computeTsArityMetadata } from './arity-metadata.js';
import { isArrayMethodCallbackArrow } from './array-callback.js';
import { getTreeSitterBufferSize } from '../../constants.js';
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
@@ -252,6 +253,25 @@ export function emitTsScopeCaptures(
}
}
// #1876: drop @declaration.function for array higher-order-method
// callbacks (`const x = arr.map(a => …)`). The HOC-wrapped-arrow
// pattern matches them, but the binding holds a value, not a callable.
// The binding keeps its separate @declaration.const / .variable match,
// and the arrow's own @scope.function match (a different pattern) is
// untouched, so inner-call attribution falls through to the enclosing
// scope instead of a phantom Function.
const fnDeclAnchor = grouped['@declaration.function'];
if (fnDeclAnchor !== undefined) {
const arrowNode = findFunctionNode(
tree.rootNode,
fnDeclAnchor.range,
groupedNodes['@declaration.function'],
);
if (arrowNode !== null && isArrayMethodCallbackArrow(arrowNode)) {
continue;
}
}
// Synthesize arity metadata on function-like declaration anchors
// before pushing the match. The registry uses these to narrow
// overloads — TypeScript supports overload signatures via
@@ -250,20 +250,22 @@ const TYPESCRIPT_SCOPE_QUERY = `
;; that promotes the binding to the parent scope (where \`const X\`
;; lives).
;;
;; Trade-off — chained array-method form: \`const x = arr.find((y) => p(y))\`
;; has the same syntactic shape and would also match, naming the
;; \`.find\` callback as \`x\`. The resulting \`Function:x\` is mostly
;; harmless: \`x\` is consumed as a value (\`if (x) { ... }\`), never
;; invoked as a function, so it gets zero incoming \`CALLS\` edges. The
;; one outgoing edge \`Function:x → p\` is a minor mis-attribution that
;; could in principle be fixed by adding a \`function: [(identifier)
;; (member_expression)]\` predicate that excludes property-identifiers
;; matching a known array-method blocklist (\`map\` / \`filter\` / \`find\`
;; / \`reduce\` / \`forEach\` / \`some\` / \`every\`). We don't do that here
;; because (a) the false-positive cost is negligible, (b) the blocklist
;; would need maintenance, and (c) any user-defined fluent-API method
;; with a callback argument would still false-positive — there's no
;; clean syntactic line.
;; #1876 — chained array-method form: \`const x = arr.find((y) => p(y))\`
;; has the same syntactic shape and matches here too, naming the
;; \`.find\` callback as \`x\`. Because \`x\` holds a value (the method
;; result), not a callable, the spurious \`Function:x\` def is dropped
;; emit-side in captures.ts: \`isArrayMethodCallbackArrow\` skips any
;; \`@declaration.function\` whose enclosing call has a member-expression
;; callee with a known Array-method property (\`ARRAY_CALLBACK_METHODS\`:
;; \`map\` / \`filter\` / \`find\` / \`reduce\` / \`forEach\` / \`some\` /
;; \`every\` / …). Only the \`@declaration.variable\` survives, so the
;; binding is a single value def and calls inside the callback attribute
;; to the enclosing scope rather than \`Function:x\`.
;;
;; Residual (intentional): a user-defined fluent-API method with a
;; callback (\`qb.where(x => …)\`) is NOT in the blocklist and still
;; classifies as \`Function\` — there's no clean syntactic line beyond
;; the well-known Array surface, so the set is closed and easy to extend.
;;
;; Trade-off — multi-arrow arguments: \`const x = call(arrow1, arrow2)\`
;; would emit TWO matches with the same name \`x\`. tree-sitter-query
@@ -8,6 +8,7 @@ import type {
MethodVisibility,
} from '../../method-types.js';
import { hasKeyword } from '../../field-extractors/configs/helpers.js';
import { classifyCppParameterType } from '../../languages/cpp/arity-metadata.js';
import { extractSimpleTypeName } from '../../type-extractors/shared.js';
import type { SyntaxNode } from '../../utils/ast-helpers.js';
@@ -149,6 +150,11 @@ function extractCppParameters(node: SyntaxNode): ParameterInfo[] {
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null,
rawType: typeNode?.text?.trim() ?? null,
typeClass: classifyCppParameterType(
typeNode?.text?.trim() ?? 'unknown',
declNode?.text,
param.text,
),
isOptional: false,
isVariadic: false,
});
@@ -164,6 +170,11 @@ function extractCppParameters(node: SyntaxNode): ParameterInfo[] {
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null,
rawType: typeNode?.text?.trim() ?? null,
typeClass: classifyCppParameterType(
typeNode?.text?.trim() ?? 'unknown',
declNode?.text,
param.text,
),
isOptional: true,
isVariadic: false,
});
@@ -180,6 +191,11 @@ function extractCppParameters(node: SyntaxNode): ParameterInfo[] {
? (extractSimpleTypeName(typeNode) ?? typeNode.text?.trim() ?? null)
: null,
rawType: typeNode?.text?.trim() ?? null,
typeClass: classifyCppParameterType(
typeNode?.text?.trim() ?? 'unknown',
declNode?.text,
param.text,
),
isOptional: false,
isVariadic: true,
});
+2 -1
View File
@@ -1,6 +1,6 @@
// gitnexus/src/core/ingestion/method-types.ts
import type { SupportedLanguages } from 'gitnexus-shared';
import type { ParameterTypeClass, SupportedLanguages } from 'gitnexus-shared';
import type { FieldVisibility } from './field-types.js';
import type { SyntaxNode } from './utils/ast-helpers.js';
@@ -14,6 +14,7 @@ export interface ParameterInfo {
* Used by typeTagForId for overload disambiguation where generic args matter.
* Falls back to `type` when not set. */
rawType?: string | null;
typeClass?: ParameterTypeClass;
isOptional: boolean;
isVariadic: boolean;
}
@@ -77,11 +77,15 @@ export interface ScopeResolutionIndexes {
* are returned first and win duplicate `def.nodeId` metadata, with
* unique augmentations appended after. See I8. */
readonly bindingAugmentations: ReadonlyMap<ScopeId, ReadonlyMap<string, readonly BindingRef[]>>;
/** Workspace-level FQN binding lookup. Populated by PHP namespace-
* siblings Step 3b as a shared map instead of per-scope duplication.
* Consulted by `lookupBindingsAt` as a third source after finalized
* and per-scope augmented bindings. Keys are backslash-separated FQNs
* (e.g. `App\Models\User`). */
/** Workspace-level binding lookup, shared instead of per-scope
* duplication. Consulted by `lookupBindingsAt` as a third source after
* finalized and per-scope augmented bindings. Language-specific
* namespace-sibling hooks populate it with disjoint key formats that
* never collide — e.g. backslash-separated FQNs (`App\Models\User`) for
* backslash-namespace languages, and bare simple names (`User`) for
* global-/default-namespace types that are visible from every file. The
* shared map gives those workspace-wide names one entry each instead of
* O(scopes × defs) per-scope augmentation. */
readonly workspaceFqnBindings: ReadonlyMap<string, readonly BindingRef[]>;
/** Pre-resolution usage facts; consumed by the resolution phase. */
readonly referenceSites: readonly ReferenceSite[];
@@ -30,6 +30,7 @@ import {
typeTagForId,
constTagForId,
buildCollisionGroups,
parameterShapeIdTag,
} from './utils/method-props.js';
import {
extractTemplateArguments,
@@ -53,7 +54,13 @@ import type {
FileConstructorBindings,
FileScopeBindings,
ExtractedORMQuery,
FetchWrapperDef,
} from './workers/parse-worker.js';
import type {
ExtractedRouterImport,
ExtractedRouterInclude,
ExtractedRouterModuleAlias,
} from './route-extractors/fastapi-router-bindings.js';
import {
getTreeSitterBufferSize,
getTreeSitterContentByteLength,
@@ -69,7 +76,11 @@ export interface WorkerExtractedData {
heritage: ExtractedHeritage[];
routes: ExtractedRoute[];
fetchCalls: ExtractedFetchCall[];
fetchWrapperDefs: FetchWrapperDef[];
decoratorRoutes: ExtractedDecoratorRoute[];
routerIncludes: ExtractedRouterInclude[];
routerImports: ExtractedRouterImport[];
routerModuleAliases: ExtractedRouterModuleAlias[];
toolDefs: ExtractedToolDef[];
ormQueries: ExtractedORMQuery[];
constructorBindings: FileConstructorBindings[];
@@ -110,7 +121,11 @@ export const mergeChunkResults = (
const allHeritage: ExtractedHeritage[] = [];
const allRoutes: ExtractedRoute[] = [];
const allFetchCalls: ExtractedFetchCall[] = [];
const allFetchWrapperDefs: FetchWrapperDef[] = [];
const allDecoratorRoutes: ExtractedDecoratorRoute[] = [];
const allRouterIncludes: ExtractedRouterInclude[] = [];
const allRouterImports: ExtractedRouterImport[] = [];
const allRouterModuleAliases: ExtractedRouterModuleAlias[] = [];
const allToolDefs: ExtractedToolDef[] = [];
const allORMQueries: ExtractedORMQuery[] = [];
const allConstructorBindings: FileConstructorBindings[] = [];
@@ -147,7 +162,11 @@ export const mergeChunkResults = (
for (const item of result.heritage) allHeritage.push(item);
for (const item of result.routes) allRoutes.push(item);
for (const item of result.fetchCalls) allFetchCalls.push(item);
for (const item of result.fetchWrapperDefs ?? []) allFetchWrapperDefs.push(item);
for (const item of result.decoratorRoutes) allDecoratorRoutes.push(item);
for (const item of result.routerIncludes ?? []) allRouterIncludes.push(item);
for (const item of result.routerImports ?? []) allRouterImports.push(item);
for (const item of result.routerModuleAliases ?? []) allRouterModuleAliases.push(item);
for (const item of result.toolDefs) allToolDefs.push(item);
if (result.ormQueries) for (const item of result.ormQueries) allORMQueries.push(item);
for (const item of result.constructorBindings) allConstructorBindings.push(item);
@@ -163,7 +182,11 @@ export const mergeChunkResults = (
heritage: allHeritage,
routes: allRoutes,
fetchCalls: allFetchCalls,
fetchWrapperDefs: allFetchWrapperDefs,
decoratorRoutes: allDecoratorRoutes,
routerIncludes: allRouterIncludes,
routerImports: allRouterImports,
routerModuleAliases: allRouterModuleAliases,
toolDefs: allToolDefs,
ormQueries: allORMQueries,
constructorBindings: allConstructorBindings,
@@ -203,7 +226,11 @@ const processParsingWithWorkers = async (
heritage: [],
routes: [],
fetchCalls: [],
fetchWrapperDefs: [],
decoratorRoutes: [],
routerIncludes: [],
routerImports: [],
routerModuleAliases: [],
toolDefs: [],
ormQueries: [],
constructorBindings: [],
@@ -639,6 +666,13 @@ const processParsingSequential = async (
cached.groups,
);
}
const parameterShapeTag =
nodeLabel === 'Function' || nodeLabel === 'Method'
? parameterShapeIdTag(
methodProps.parameterTypes as string[] | undefined,
methodProps.parameterTypeClasses as ParameterTypeClass[] | undefined,
)
: '';
const classTemplateArguments =
extractedClassSymbol?.templateArguments ??
provider.classExtractor?.extractTemplateArgumentsFromCapture?.({
@@ -691,7 +725,7 @@ const processParsingSequential = async (
}
const nodeId = generateId(
nodeLabel,
`${file.path}:${qualifiedName}${classTemplateTag}${arityTag}${constraintsTag}`,
`${file.path}:${qualifiedName}${classTemplateTag}${arityTag}${constraintsTag}${parameterShapeTag}`,
);
const classNodeForSymbol = definitionNodeForRange || definitionNode || nameNode;
const qualifiedTypeName =
@@ -61,7 +61,13 @@ import type {
ExtractedRoute,
ExtractedToolDef,
FileConstructorBindings,
FetchWrapperDef,
} from '../workers/parse-worker.js';
import type {
ExtractedRouterImport,
ExtractedRouterInclude,
ExtractedRouterModuleAlias,
} from '../route-extractors/fastapi-router-bindings.js';
import type { ExtractedHeritage } from '../model/heritage-map.js';
import type { KnowledgeGraph } from '../../graph/types.js';
import type { PipelineOptions } from '../pipeline.js';
@@ -141,6 +147,7 @@ export async function runChunkedParseAndResolve(
): Promise<{
exportedTypeMap: ExportedTypeMap;
allFetchCalls: ExtractedFetchCall[];
allFetchWrapperDefs: FetchWrapperDef[];
allExtractedRoutes: ExtractedRoute[];
allDecoratorRoutes: ExtractedDecoratorRoute[];
allToolDefs: ExtractedToolDef[];
@@ -352,8 +359,12 @@ export async function runChunkedParseAndResolve(
// it, and later wildcard chunks re-run it themselves.
let hasSynthesized = false;
const allFetchCalls: ExtractedFetchCall[] = [];
const allFetchWrapperDefs: FetchWrapperDef[] = [];
const allExtractedRoutes: ExtractedRoute[] = [];
const allDecoratorRoutes: ExtractedDecoratorRoute[] = [];
const allRouterIncludes: ExtractedRouterInclude[] = [];
const allRouterImports: ExtractedRouterImport[] = [];
const allRouterModuleAliases: ExtractedRouterModuleAlias[] = [];
const allToolDefs: ExtractedToolDef[] = [];
const allORMQueries: ExtractedORMQuery[] = [];
const deferredWorkerCalls: ExtractedCall[] = [];
@@ -663,12 +674,24 @@ export async function runChunkedParseAndResolve(
if (chunkWorkerData.fetchCalls?.length) {
for (const item of chunkWorkerData.fetchCalls) allFetchCalls.push(item);
}
if (chunkWorkerData.fetchWrapperDefs?.length) {
for (const item of chunkWorkerData.fetchWrapperDefs) allFetchWrapperDefs.push(item);
}
if (chunkWorkerData.routes?.length) {
for (const item of chunkWorkerData.routes) allExtractedRoutes.push(item);
}
if (chunkWorkerData.decoratorRoutes?.length) {
for (const item of chunkWorkerData.decoratorRoutes) allDecoratorRoutes.push(item);
}
if (chunkWorkerData.routerIncludes?.length) {
for (const item of chunkWorkerData.routerIncludes) allRouterIncludes.push(item);
}
if (chunkWorkerData.routerImports?.length) {
for (const item of chunkWorkerData.routerImports) allRouterImports.push(item);
}
if (chunkWorkerData.routerModuleAliases?.length) {
for (const item of chunkWorkerData.routerModuleAliases) allRouterModuleAliases.push(item);
}
if (chunkWorkerData.toolDefs?.length) {
for (const item of chunkWorkerData.toolDefs) allToolDefs.push(item);
}
@@ -1079,9 +1102,161 @@ export async function runChunkedParseAndResolve(
importCtx.index = EMPTY_INDEX;
importCtx.normalizedFileList = [];
// FastAPI router-prefix resolution (cross-file).
//
// Workers emit two kinds of records per Python file:
// • `routerIncludes` — every `app.include_router(<routerExpr>, prefix='/x')`
// site, where `routerExpr` is either `<module>.router` (Shape A) or a
// bare local name (Shape B).
// • `routerImports` — every `from <module> import router [as <alias>]`,
// mapping a local name to a module key (the basename of the source
// module). These let us resolve Shape-B router includes back to the
// module that defines the router.
//
// We build `module-basename → Set<prefix>` and then walk
// `allDecoratorRoutes`: any decorator route emitted from a `router.<verb>`
// decorator inherits its file-basename's prefix. When a router is mounted
// under multiple prefixes we duplicate the route entry, mirroring FastAPI's
// runtime behaviour.
if (allRouterIncludes.length > 0 && allDecoratorRoutes.length > 0) {
// Group `routerImports` by file so we can resolve Shape-B locals against
// imports declared in the SAME file as the include_router call. We carry
// both the short module key (file basename) and, when available, the long
// key (`<dir>/<basename>`) so cross-package same-name modules don't blur
// their prefixes together. `routerModuleAliases` lifts the same long-key
// information for Shape-A includes whose receiving module was imported
// via `from <pkg> import <module>`.
interface LocalImport {
moduleKey: string;
moduleKeyLong: string | undefined;
}
const importsByFile = new Map<string, Map<string, LocalImport>>();
for (const imp of allRouterImports) {
let m = importsByFile.get(imp.filePath);
if (!m) {
m = new Map();
importsByFile.set(imp.filePath, m);
}
m.set(imp.localName, {
moduleKey: imp.moduleKey,
moduleKeyLong: imp.moduleKeyLong,
});
}
// Module-alias map keyed by file: `localName` (the imported module
// identifier in this file) → long key. Shape-A receivers like
// `users.router` are matched against this map; the long key, when
// present, scopes the prefix to the precise source file.
const moduleAliasesByFile = new Map<string, Map<string, string>>();
for (const alias of allRouterModuleAliases) {
let m = moduleAliasesByFile.get(alias.filePath);
if (!m) {
m = new Map();
moduleAliasesByFile.set(alias.filePath, m);
}
m.set(alias.localName, alias.moduleKeyLong);
}
// Two parallel maps: long-key (precise) and short-key (basename
// fallback). Long-key entries are preferred when the file's own long
// key matches; short-key entries match any file with that basename and
// remain the fallback when no long key is known (e.g. Shape A includes
// without a corresponding import statement).
const prefixesByLongKey = new Map<string, Set<string>>();
const prefixesByShortKey = new Map<string, Set<string>>();
const recordPrefix = (target: Map<string, Set<string>>, key: string, prefix: string): void => {
let set = target.get(key);
if (!set) {
set = new Set();
target.set(key, set);
}
set.add(prefix);
};
for (const inc of allRouterIncludes) {
// Shape A: `<module>.router`. The worker emits `routerExpr` already
// including `.router`, so split it back. We only know a short module
// key here — the call site doesn't carry the dotted package path. If
// the same file imports `<module>` via `from <pkg> import <module>`
// (recorded in `allRouterModuleAliases`) we promote to a long key.
const dotIdx = inc.routerExpr.indexOf('.router');
if (dotIdx > 0) {
const moduleShort = inc.routerExpr.slice(0, dotIdx);
const aliasLong = moduleAliasesByFile.get(inc.filePath)?.get(moduleShort);
if (aliasLong) {
recordPrefix(prefixesByLongKey, aliasLong, inc.prefix);
} else {
recordPrefix(prefixesByShortKey, moduleShort, inc.prefix);
}
continue;
}
// Shape B: bare local name. Resolve through this file's imports. The
// import line gives us a long key whenever the module path was multi-
// segment, so cross-package collisions are eliminated for Shape B.
const localImp = importsByFile.get(inc.filePath)?.get(inc.routerExpr);
if (!localImp) continue;
if (localImp.moduleKeyLong) {
recordPrefix(prefixesByLongKey, localImp.moduleKeyLong, inc.prefix);
} else {
recordPrefix(prefixesByShortKey, localImp.moduleKey, inc.prefix);
}
}
if (prefixesByLongKey.size > 0 || prefixesByShortKey.size > 0) {
const fileLongKey = (rel: string): string => {
// Strip `.py`, then take the last two path segments. `api/users.py`
// → `api/users`. Files at the repo root return the empty string,
// which can never match a long-key entry (those always include a
// parent directory) and so fall through to the short-key lookup.
const noExt = rel.endsWith('.py') ? rel.slice(0, -3) : rel;
const lastSlash = noExt.lastIndexOf('/');
if (lastSlash < 0) return '';
const beforeLast = noExt.slice(0, lastSlash);
const stem = noExt.slice(lastSlash + 1);
const prevSlash = beforeLast.lastIndexOf('/');
const parent = prevSlash >= 0 ? beforeLast.slice(prevSlash + 1) : beforeLast;
return `${parent}/${stem}`;
};
const fileShortKey = (rel: string): string => {
const slash = rel.lastIndexOf('/');
const file = slash >= 0 ? rel.slice(slash + 1) : rel;
return file.endsWith('.py') ? file.slice(0, -3) : file;
};
const expanded: ExtractedDecoratorRoute[] = [];
for (const dr of allDecoratorRoutes) {
if (dr.decoratorReceiver !== 'router' || !dr.filePath.endsWith('.py')) {
expanded.push(dr);
continue;
}
// Long-key lookup first; only fall back to the short key when no
// long-key prefix targets this file. This avoids prefix leakage
// between e.g. `api/users.py` and `admin/users.py`.
const longKey = fileLongKey(dr.filePath);
const longPrefixes = longKey ? prefixesByLongKey.get(longKey) : undefined;
const shortPrefixes = longPrefixes
? undefined
: prefixesByShortKey.get(fileShortKey(dr.filePath));
const prefixes = longPrefixes ?? shortPrefixes;
if (!prefixes || prefixes.size === 0) {
expanded.push(dr);
continue;
}
for (const prefix of prefixes) {
expanded.push({ ...dr, prefix });
}
}
allDecoratorRoutes.length = 0;
for (const dr of expanded) allDecoratorRoutes.push(dr);
}
}
return {
exportedTypeMap,
allFetchCalls,
allFetchWrapperDefs,
allExtractedRoutes,
allDecoratorRoutes,
allToolDefs,
@@ -27,6 +27,7 @@ import type {
ExtractedDecoratorRoute,
ExtractedToolDef,
ExtractedORMQuery,
FetchWrapperDef,
} from '../workers/parse-worker.js';
import type { createResolutionContext } from '../model/resolution-context.js';
import { runChunkedParseAndResolve } from './parse-impl.js';
@@ -45,6 +46,7 @@ export interface ParseOutput {
*/
readonly exportedTypeMap: ReadonlyMap<string, ReadonlyMap<string, string>>;
readonly allFetchCalls: readonly ExtractedFetchCall[];
readonly allFetchWrapperDefs: readonly FetchWrapperDef[];
readonly allExtractedRoutes: readonly ExtractedRoute[];
readonly allDecoratorRoutes: readonly ExtractedDecoratorRoute[];
readonly allToolDefs: readonly ExtractedToolDef[];
@@ -131,6 +131,10 @@ export function normalizeExtractedRoutePath(routePath: string, prefix: string |
return joined.replace(/\/+/g, '/') || '/';
}
function escapeRegex(s: string): string {
return s.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
}
export const routesPhase: PipelinePhase<RoutesOutput> = {
name: 'routes',
deps: ['parse'],
@@ -142,6 +146,7 @@ export const routesPhase: PipelinePhase<RoutesOutput> = {
const {
allPaths,
allFetchCalls: parseFetchCalls,
allFetchWrapperDefs,
allExtractedRoutes,
allDecoratorRoutes,
} = getPhaseOutput<ParseOutput>(deps, 'parse');
@@ -193,7 +198,6 @@ export const routesPhase: PipelinePhase<RoutesOutput> = {
}
}
const ensureSlash = (path: string) => (path.startsWith('/') ? path : '/' + path);
let duplicateRoutes = 0;
const namedRouteRegistry = new Map<string, string>();
const addRoute = (url: string, entry: RouteEntry) => {
@@ -215,7 +219,8 @@ export const routesPhase: PipelinePhase<RoutesOutput> = {
}
}
for (const dr of allDecoratorRoutes) {
addRoute(ensureSlash(dr.routePath), {
const url = normalizeExtractedRoutePath(dr.routePath, dr.prefix ?? null);
addRoute(url, {
filePath: dr.filePath,
source: `decorator-${dr.decoratorName}`,
});
@@ -357,6 +362,35 @@ export const routesPhase: PipelinePhase<RoutesOutput> = {
}
}
// ── Cross-file fetch wrapper consumer extraction ──
// When the parse phase discovered functions that internally call fetch(),
// scan JS/TS consumer files for calls to those wrapper functions with
// URL-like string arguments and add them to allFetchCalls so
// processNextjsFetchRoutes can create FETCHES edges.
if (allFetchWrapperDefs && allFetchWrapperDefs.length > 0 && routeRegistry.size > 0) {
const wrapperNames = new Set(allFetchWrapperDefs.map((d) => d.functionName));
const jsFiles = allPaths.filter((p) => /\.[jt]sx?$/.test(p));
if (jsFiles.length > 0 && wrapperNames.size > 0) {
const jsContents = await readFileContents(ctx.repoPath, jsFiles);
for (const [filePath, content] of jsContents) {
for (const name of wrapperNames) {
const regex = new RegExp(
`\\b${escapeRegex(name)}\\s*\\(\\s*['"\`](/[^'"\`\\s)]+)['"\`]`,
'g',
);
let match;
while ((match = regex.exec(content)) !== null) {
allFetchCalls.push({
filePath,
fetchURL: match[1],
lineNumber: content.substring(0, match.index).split('\n').length,
});
}
}
}
}
}
if (routeRegistry.size > 0 && allFetchCalls.length > 0) {
const routeURLToFile = new Map<string, string>();
for (const [url, entry] of routeRegistry) routeURLToFile.set(url, entry.filePath);
@@ -81,6 +81,7 @@ export const MIGRATED_LANGUAGES: ReadonlySet<SupportedLanguages> = new Set<Suppo
SupportedLanguages.Java,
SupportedLanguages.Rust,
SupportedLanguages.Ruby,
SupportedLanguages.Cobol,
]);
/**
@@ -0,0 +1,275 @@
/**
* FastAPI router-prefix detection — pure functions, no worker thread.
*
* NOT A WORKER. This module exports plain synchronous functions; it
* does not import `worker_threads`, does not call `parentPort`, and
* is not a new worker entry point. It lives next to the other route
* extractors (expo, nextjs, php, laravel) for that reason.
*
* The implementation was historically inlined in `workers/parse-worker.ts`,
* but parse-worker.ts is itself the worker entry point and cannot be
* loaded from the main thread (see the same constraint used by
* `test/unit/call-attribution-issue-1166.test.ts`). Splitting the pure
* extraction here lets unit tests import the function directly without
* booting a worker, satisfying DoD §2.7.
*
* Worker phase is per-file, so the heavy cross-file resolution lives in
* `pipeline-phases/parse-impl.ts`. Here we only extract two raw record
* kinds and let the pipeline aggregate them across files:
*
* • {@link ExtractedRouterInclude} — every
* `<host>.include_router(<routerExpr>, prefix='/x')` site, where
* `<routerExpr>` is either `<module>.router` (Shape A) or a bare
* local name (Shape B). `<host>` is intentionally unconstrained:
* production code uses `app`, `api`, `application`, `asgi_app`,
* etc., and the call shape (`include_router` invoked with a
* `prefix=` keyword) is specific enough on its own.
*
* • {@link ExtractedRouterImport} — every
* `from <module> import router [as <alias>]`, captured for both
* absolute and relative module paths (`from .calls import …`).
* parse-impl uses the imports to resolve Shape-B local names back
* to the file that declares the router.
*
* Module keying is two-tiered to avoid prefix bleed between same-named
* files in different packages (e.g. `api/users.py` vs `admin/users.py`):
*
* • short key — basename without `.py` (`users`)
* • long key — `<parent-dir>/<basename>` (`api/users`)
*
* Imports always carry the short key and, when the module path was
* multi-segment, also the long key. parse-impl matches against the
* long key first and falls back to the short key, so cross-package
* collisions are eliminated for Shape B and minimised for Shape A.
*
* The functions in this module are pure (no Worker / parentPort
* dependency) so they can be unit-tested directly without booting a
* worker thread.
*/
/**
* One `<host>.include_router(<routerExpr>, prefix='/x')` site.
*
* `routerExpr` is the raw text of the first argument — either
* `<module>.router` (Shape A) or a bare local name (Shape B).
* parse-impl resolves Shape B against {@link ExtractedRouterImport}
* records emitted by the same file.
*/
export interface ExtractedRouterInclude {
filePath: string;
routerExpr: string;
prefix: string;
lineNumber: number;
}
/**
* One `from <module> import router [as <alias>]` discovered in a
* Python file.
*
* `moduleKey` is the short key (last `.`-segment of the module path,
* e.g. `api.users` → `users`). `moduleKeyLong` is the long key (last
* two segments joined with `/`, e.g. `api/users`); it is the empty
* string / undefined when the import is single-segment (e.g.
* `from users import router`) or pure-dots (e.g. `from . import
* router`). The long key, when present, gives parse-impl a precise
* way to bind a Shape-B `include_router` call to exactly one Python
* file even when other packages contain a same-named module.
*/
export interface ExtractedRouterImport {
filePath: string;
localName: string;
moduleKey: string;
moduleKeyLong?: string;
}
/**
* One `from <package> import <module>` discovered in a Python file
* where `<module>` is later used as a Shape-A include receiver
* (`<host>.include_router(<module>.router, prefix='/x')`). Without
* this record parse-impl would have to fall back to the short key
* `<module>`, which collides between e.g. `api/users.py` and
* `admin/users.py`. The record carries the long key
* (`<package>/<module>`) so parse-impl can pin the prefix onto the
* exact source file.
*
* Only emitted when the import path was multi-segment (a single
* `from users import users` would yield no long key). All fields
* carry the same module-key semantics as
* {@link ExtractedRouterImport}.
*/
export interface ExtractedRouterModuleAlias {
filePath: string;
/** Local name in the importing file (== imported name or its alias). */
localName: string;
/** Long key (`<parent>/<stem>`) — non-empty for every emitted record. */
moduleKeyLong: string;
}
// `<host>.include_router(<module>.router, ..., prefix='/x')` (Shape A).
// `<host>` is left unrestricted — common production names include
// `app`, `api`, `application`, `asgi_app`. Pinning to the literal
// `app` would silently drop these.
const INCLUDE_ROUTER_ATTR_RE =
/\b(?:[A-Za-z_][\w.]*)\.include_router\s*\(\s*([A-Za-z_][\w]*)\.router\b[^)]*?\bprefix\s*=\s*(['"])([^'"]*)\2/g;
// `<host>.include_router(<local_name>, ..., prefix='/x')` (Shape B).
const INCLUDE_ROUTER_NAME_RE =
/\b(?:[A-Za-z_][\w.]*)\.include_router\s*\(\s*([A-Za-z_][\w]*)\b[^)]*?\bprefix\s*=\s*(['"])([^'"]*)\2/g;
// Module path: a sequence of dots (`.`, `..`, `...`) for "current
// package" imports, OR an optional leading-dot prefix followed by a
// dotted identifier (`api.users`, `.api.users`, `..siblings.users`).
// The latter is the common case and the only one we can map back to
// a module stem.
const FROM_IMPORT_ROUTER_RE = /^\s*from\s+(\.+|\.*[A-Za-z_][\w.]*)\s+import\s+([^#\n]+)/gm;
/**
* Last `.`-separated segment of a (possibly relative) Python module
* path. Strips any leading dots first so `from .api.assistant import
* …` and `from api.assistant import …` both yield `assistant`.
* Pure-dot inputs (`.`, `..`) have no segment and return the empty
* string; callers should skip empty results.
*/
export function lastDottedSegment(text: string): string {
const stripped = text.replace(/^\.+/, '');
if (!stripped) return '';
const dot = stripped.lastIndexOf('.');
return dot >= 0 ? stripped.slice(dot + 1) : stripped;
}
/**
* Last two `.`-separated segments of a (possibly relative) module
* path joined with `/`, e.g. `api.users` → `api/users`. Mirrors the
* long-key shape used for files (`api/users.py` → `api/users`).
* Returns the empty string when no parent segment is available
* (single-segment imports or pure dots); callers should fall back
* to the short key in that case.
*/
export function lastTwoSegmentsAsPath(text: string): string {
const stripped = text.replace(/^\.+/, '');
if (!stripped) return '';
const last = stripped.lastIndexOf('.');
if (last <= 0) return '';
const beforeLast = stripped.slice(0, last);
const stem = stripped.slice(last + 1);
const prev = beforeLast.lastIndexOf('.');
const parent = prev >= 0 ? beforeLast.slice(prev + 1) : beforeLast;
return `${parent}/${stem}`;
}
/**
* Scan a single Python file's source text for FastAPI router
* `include_router` sites and `from <module> import router` imports,
* appending raw records to the supplied collectors.
*
* `outModuleAliases` is optional: when supplied, every multi-segment
* `from <pkg> import <name>` (other than `router` itself) is recorded
* as a module alias so parse-impl can pin Shape-A
* `<name>.include_router(...)` calls onto the exact module file. When
* omitted, the function preserves the pre-existing behaviour and
* skips the alias collection — this keeps the function signature
* back-compat with older callers (and the parse-cache replay path).
*/
export function extractFastAPIRouterBindings(
filePath: string,
content: string,
outIncludes: ExtractedRouterInclude[],
outImports: ExtractedRouterImport[],
outModuleAliases?: ExtractedRouterModuleAlias[],
): void {
if (!content.includes('include_router') && !content.includes('router')) return;
// `from <module> import router [as <alias>]`. We capture every name
// in the import list. `router` (with or without an `as` alias) maps
// to outImports; every other name lands in outModuleAliases when a
// long key is available, so Shape-A `<name>.router` includes can be
// pinned to the exact module file.
if (content.includes(' import ')) {
FROM_IMPORT_ROUTER_RE.lastIndex = 0;
let m: RegExpExecArray | null;
while ((m = FROM_IMPORT_ROUTER_RE.exec(content)) !== null) {
const moduleText = m[1];
const importList = m[2];
const moduleShort = lastDottedSegment(moduleText);
if (!moduleShort) continue;
// Long key for the imported MODULE itself (used by router
// imports — `from api.users import router` sets
// `moduleKeyLong = api/users`).
const moduleLong = lastTwoSegmentsAsPath(moduleText);
// Strip surrounding parens / trailing whitespace; split on
// commas. (Multiline import groups already have their newlines
// present in the captured list.)
const cleaned = importList.replace(/[()]/g, '').trim();
for (const rawPart of cleaned.split(',')) {
const part = rawPart.trim();
if (!part) continue;
// `router` or `router as foo` → ExtractedRouterImport.
const routerAlias = /^router(?:\s+as\s+([A-Za-z_]\w*))?$/.exec(part);
if (routerAlias) {
const localName = routerAlias[1] ?? 'router';
outImports.push({
filePath,
localName,
moduleKey: moduleShort,
...(moduleLong ? { moduleKeyLong: moduleLong } : {}),
});
continue;
}
// Any other `<name>` or `<name> as <alias>` — recorded as a
// module alias so parse-impl can pin Shape-A includes. The
// long key here is computed against the IMPORTED MODULE PATH
// (`<moduleText>.<name>`), not the package path that `<name>`
// was imported FROM. `from api import users` therefore yields
// `api/users`, the same long key as the file it points at.
if (!outModuleAliases) continue;
const otherAlias = /^([A-Za-z_]\w*)(?:\s+as\s+([A-Za-z_]\w*))?$/.exec(part);
if (!otherAlias) continue;
const importedName = otherAlias[1];
const localName = otherAlias[2] ?? importedName;
const aliasLong = lastTwoSegmentsAsPath(`${moduleText}.${importedName}`);
if (!aliasLong) continue;
outModuleAliases.push({
filePath,
localName,
moduleKeyLong: aliasLong,
});
}
}
}
if (!content.includes('include_router')) return;
// Shape A: `<host>.include_router(<module>.router, prefix='/x')`.
INCLUDE_ROUTER_ATTR_RE.lastIndex = 0;
let m: RegExpExecArray | null;
while ((m = INCLUDE_ROUTER_ATTR_RE.exec(content)) !== null) {
outIncludes.push({
filePath,
routerExpr: `${m[1]}.router`,
prefix: m[3],
lineNumber: content.substring(0, m.index).split('\n').length,
});
}
// Shape B: `<host>.include_router(my_router, prefix='/x')`.
// Resolution to a module key happens in parse-impl using
// outImports from the same file.
INCLUDE_ROUTER_NAME_RE.lastIndex = 0;
while ((m = INCLUDE_ROUTER_NAME_RE.exec(content)) !== null) {
// Skip cases that already matched Shape A — INCLUDE_ROUTER_NAME_RE
// is intentionally permissive and would re-capture `<mod>.router`
// as the bare name `mod`. Discriminate by re-checking the
// immediate source around the captured argument position.
const argStart = m.index + m[0].indexOf(m[1]);
const dotProbe = content.slice(argStart + m[1].length, argStart + m[1].length + 8);
if (/^\s*\.\s*router/.test(dotProbe)) continue;
outIncludes.push({
filePath,
routerExpr: m[1],
prefix: m[3],
lineNumber: content.substring(0, m.index).split('\n').length,
});
}
}
@@ -574,6 +574,7 @@ function buildDefFromDeclarationMatch(
const declaredType = match['@declaration.field-type']?.text;
const returnType = match['@declaration.return-type']?.text;
const templateConstraints = parseJsonCapture(match['@declaration.template-constraints']);
const isExplicit = parseBooleanCapture(match['@declaration.is-explicit']);
return {
nodeId: makeDefId(filePath, anchor.range, type, nameCap.text),
@@ -588,6 +589,7 @@ function buildDefFromDeclarationMatch(
...(returnType !== undefined ? { returnType } : {}),
...(templateArguments !== undefined ? { templateArguments } : {}),
...(templateConstraints !== undefined ? { templateConstraints } : {}),
...(isExplicit === true ? { isExplicit: true } : {}),
};
}
@@ -610,6 +612,13 @@ function parseIntCapture(cap: { readonly text: string } | undefined): number | u
return Number.isFinite(n) ? n : undefined;
}
function parseBooleanCapture(cap: { readonly text: string } | undefined): boolean | undefined {
if (cap === undefined) return undefined;
if (cap.text === 'true') return true;
if (cap.text === 'false') return false;
return undefined;
}
function parseJsonParameterTypeClassesCapture(
cap: { readonly text: string } | undefined,
): ParameterTypeClass[] | undefined {
@@ -741,6 +750,61 @@ function normalizeNodeLabel(kindStr: string): SymbolDefinition['type'] | undefin
}
}
/** Function-like labels: callable defs that must keep incoming CALLS edges. */
const NODE_BEARING_FUNCTION_LABELS: ReadonlySet<SymbolDefinition['type']> = new Set([
'Function',
'Method',
'Constructor',
]);
/** Value labels: non-callable bindings (a `const`/`let`/`var` holds a value). */
const NODE_BEARING_VALUE_LABELS: ReadonlySet<SymbolDefinition['type']> = new Set([
'Const',
'Variable',
]);
/**
* Collapse rule for the deferred node-creation migration (#1876).
*
* When graph-node creation moves from the legacy DAG onto the
* registry-primary path, a single source binding can carry more than one
* `SymbolDefinition` for the same name in the same scope — e.g. a direct
* arrow `const fn = () => {}` is classified BOTH as a `Function` (the
* arrow) and a `Variable` (the binding). Emitting one graph node per def
* would reproduce exactly the duplicate-node bug this issue tracks.
*
* `selectNodeBearingDef` picks the ONE def that should bear the graph node
* for such a binding group:
*
* 1. a function-like def (`Function` / `Method` / `Constructor`) if any —
* the binding is callable and must keep incoming `CALLS` edges;
* 2. otherwise a value def (`Const` / `Variable`) — the binding holds a
* value (e.g. an array-method result after the U1/U2 narrowing);
* 3. otherwise the first def — deterministic fallback for label sets this
* rule does not rank.
*
* INPUT CONTRACT: `group` must be the defs bound to ONE name within ONE
* scope (a binding group). It deliberately does NOT dedup by range —
* `SymbolDefinition` carries no range and `makeDefId` encodes only the
* start position, so containment is uncomputable here; the caller forms the
* group (e.g. from a scope's `ownedDefs` keyed by name) before calling.
*
* Pure. No production call site yet — this dead export is intentional and
* tracked by #1876 (the deferred node-creation migration); it is the
* executable contract that follow-up will consume, pinned today by the
* scope-extractor unit test.
*/
export function selectNodeBearingDef(
group: readonly SymbolDefinition[],
): SymbolDefinition | undefined {
if (group.length === 0) return undefined;
const functionLike = group.find((def) => NODE_BEARING_FUNCTION_LABELS.has(def.type));
if (functionLike !== undefined) return functionLike;
const value = group.find((def) => NODE_BEARING_VALUE_LABELS.has(def.type));
if (value !== undefined) return value;
return group[0];
}
function makeDefId(
filePath: string,
range: Range,
@@ -1078,7 +1142,9 @@ const KNOWN_SUB_TAGS: ReadonlySet<string> = new Set<string>([
'@declaration.required-parameter-count',
'@declaration.parameter-types',
'@declaration.parameter-type-classes',
'@declaration.return-type',
'@declaration.template-constraints',
'@declaration.is-explicit',
]);
/**
@@ -17,11 +17,12 @@
* migrate.
*/
import type { NodeLabel, ScopeId, SymbolDefinition } from 'gitnexus-shared';
import type { NodeLabel, ParameterTypeClass, ScopeId, SymbolDefinition } from 'gitnexus-shared';
import type { ScopeResolutionIndexes } from '../../model/scope-resolution-indexes.js';
import { generateId } from '../../../../lib/utils.js';
import { qualifiedKey, simpleKey, type GraphNodeLookup } from '../graph-bridge/node-lookup.js';
import { templateConstraintsIdTag } from '../../utils/template-arguments.js';
import { parameterShapeIdTag } from '../../utils/method-props.js';
/**
* Labels that may legitimately ANCHOR a CALLS/ACCESSES edge as the
* source ("caller"). A Variable / Property can be the TARGET of an
@@ -76,6 +77,7 @@ export function resolveDefGraphId(
qualifiedName?: string;
type?: NodeLabel;
parameterTypes?: readonly string[];
parameterTypeClasses?: readonly ParameterTypeClass[];
templateArguments?: readonly string[];
templateConstraints?: unknown;
},
@@ -102,11 +104,23 @@ export function resolveDefGraphId(
const cHit = nodeLookup.get(cKey);
if (cHit !== undefined) return cHit;
}
if (
(def.type === 'Function' || def.type === 'Method') &&
def.parameterTypes !== undefined &&
def.parameterTypeClasses !== undefined
) {
const shapeTag = parameterShapeIdTag(def.parameterTypes, def.parameterTypeClasses);
if (shapeTag !== '') {
const shapeKey = qualifiedKey(filePath, def.type, `${qn}${shapeTag}`);
const shapeHit = nodeLookup.get(shapeKey);
if (shapeHit !== undefined) return shapeHit;
}
}
// Overload disambiguation: when the def carries parameter types,
// try the parameter-typed key first so same-name same-arity
// overloads route to their distinct graph nodes.
if (
def.type === 'Method' &&
(def.type === 'Function' || def.type === 'Method') &&
def.parameterTypes !== undefined &&
def.parameterTypes.length > 0
) {
@@ -18,9 +18,10 @@
* format that downstream consumers (queries, edges, MCP) expect.
*/
import type { NodeLabel } from 'gitnexus-shared';
import type { NodeLabel, ParameterTypeClass } from 'gitnexus-shared';
import type { KnowledgeGraph } from '../../../graph/types.js';
import { templateConstraintsIdTag } from '../../utils/template-arguments.js';
import { parameterShapeIdTag } from '../../utils/method-props.js';
export type GraphNodeLookup = ReadonlyMap<string, string>;
@@ -42,6 +43,10 @@ function parseQualifiedFromId(id: string, label: NodeLabel, filePath: string): s
return hash === -1 ? suffix : suffix.slice(0, hash);
}
function stripCallableDisambiguatorTags(qualifiedName: string): string {
return qualifiedName.replace(/~shape:.*$/, '').replace(/~c:[a-z0-9]+$/, '');
}
/**
* Build a qualified-key string in a separate keyspace from simple-key
* strings. Prefix `<q>` can't appear in a valid filePath on any OS, so
@@ -84,7 +89,8 @@ export function buildGraphNodeLookup(graph: KnowledgeGraph): GraphNodeLookup {
const qualified =
props.qualifiedName ?? parseQualifiedFromId(node.id, node.label, props.filePath);
if (qualified !== undefined && qualified.length > 0) {
const qKey = qualifiedKey(props.filePath, node.label, qualified);
const keyQualified = stripCallableDisambiguatorTags(qualified);
const qKey = qualifiedKey(props.filePath, node.label, keyQualified);
if (!lookup.has(qKey)) lookup.set(qKey, node.id);
// Overload-disambiguating key: include parameter types so two
// same-arity overloads (e.g. `Lookup(int)` vs `Lookup(string)`)
@@ -93,10 +99,25 @@ export function buildGraphNodeLookup(graph: KnowledgeGraph): GraphNodeLookup {
// a parameter-types-suffixed key so resolveDefGraphId can find
// the right overload by matching its def's parameterTypes.
const pTypes = (props as { parameterTypes?: readonly string[] }).parameterTypes;
if (pTypes !== undefined && pTypes.length > 0 && node.label === 'Method') {
const pKey = qualifiedKey(props.filePath, node.label, `${qualified}~${pTypes.join(',')}`);
if (
pTypes !== undefined &&
pTypes.length > 0 &&
(node.label === 'Function' || node.label === 'Method')
) {
const pKey = qualifiedKey(
props.filePath,
node.label,
`${keyQualified}~${pTypes.join(',')}`,
);
// Each overload is unique — set unconditionally.
lookup.set(pKey, node.id);
if (!lookup.has(pKey)) lookup.set(pKey, node.id);
}
const pClasses = (props as { parameterTypeClasses?: readonly ParameterTypeClass[] })
.parameterTypeClasses;
const shapeTag = parameterShapeIdTag(pTypes, pClasses);
if (shapeTag !== '' && (node.label === 'Function' || node.label === 'Method')) {
const shapeKey = qualifiedKey(props.filePath, node.label, `${keyQualified}${shapeTag}`);
if (!lookup.has(shapeKey)) lookup.set(shapeKey, node.id);
}
// SFINAE / `requires`-clause disambiguation (issue #1579) — register
// a constraint-fingerprinted key so resolveDefGraphId can locate the
@@ -109,7 +130,7 @@ export function buildGraphNodeLookup(graph: KnowledgeGraph): GraphNodeLookup {
const cKey = qualifiedKey(
props.filePath,
node.label,
`${qualified}${templateConstraintsIdTag(tConstraints)}`,
`${keyQualified}${templateConstraintsIdTag(tConstraints)}`,
);
lookup.set(cKey, node.id);
}
@@ -125,7 +146,7 @@ export function buildGraphNodeLookup(graph: KnowledgeGraph): GraphNodeLookup {
const tKey = qualifiedKey(
props.filePath,
node.label,
`${qualified}~${props.templateArguments.join(',')}`,
`${keyQualified}~${props.templateArguments.join(',')}`,
);
if (!lookup.has(tKey)) lookup.set(tKey, node.id);
}
@@ -35,6 +35,11 @@
* candidates whose template constraints provably fail at the
* call site. Three-valued; `'unknown'` keeps the candidate
* (monotonicity).
* 4d. Conservative C++ template partial-order approximation. When
* template-placeholder overloads remain tied, prefer a candidate
* whose parameter shape is more specialized for the observed
* argument shape (`T*` over `T`, `const T&` over `T`). Unknown or
* incomparable shapes are left ambiguous.
* 5. Empty input returns empty output.
*/
@@ -205,6 +210,15 @@ export function narrowOverloadCandidates(
});
}
if (result.length > 1 && argTypes !== undefined && argTypes.length > 0) {
const partiallyOrdered = rankByTemplatePartialOrdering(
result,
argTypes,
hookCtx?.argumentTypeClasses,
);
if (partiallyOrdered !== undefined) result = partiallyOrdered;
}
return result;
}
@@ -330,6 +344,109 @@ function pairwiseCompare(a: readonly number[], b: readonly number[]): -1 | 0 | 1
return 0;
}
/**
* Closed-table approximation of C++ function-template partial ordering.
*
* Full `[temp.func.order]` requires template argument deduction. GitNexus
* keeps this graph-safe by recognizing only syntactic placeholder shapes
* that the C++ parameter sidecar already preserves:
* - `T*` is more specialized than `T` for pointer arguments.
*
* Anything with unknown argument shape, non-template parameter spelling, or
* incomparable specialized shapes stays ambiguous so callers suppress. The
* placeholder detector is intentionally narrow: lowercase template parameters
* are left ambiguous rather than guessed.
*/
function rankByTemplatePartialOrdering(
candidates: readonly SymbolDefinition[],
argTypes: readonly string[],
argTypeClasses?: readonly ParameterTypeClass[],
): readonly SymbolDefinition[] | undefined {
if (argTypeClasses === undefined) return undefined;
const viable: Array<{ def: SymbolDefinition; ranks: number[] }> = [];
for (const def of candidates) {
const params = def.parameterTypes;
const paramClasses = def.parameterTypeClasses;
if (params === undefined || paramClasses === undefined) continue;
const ranks: number[] = [];
let sawTemplateSlot = false;
let ok = true;
for (let i = 0; i < argTypes.length; i++) {
const paramType = parameterTypeAt(params, i);
const paramClass = parameterTypeClassAt(paramClasses, i);
const argClass = argTypeClasses[i];
if (paramType === undefined || paramClass === undefined || argClass === undefined) {
ok = false;
break;
}
const rank = templatePartialOrderSlotRank(paramType, paramClass, argClass);
if (rank === undefined) {
ok = false;
break;
}
sawTemplateSlot ||= isTemplatePlaceholder(paramType);
ranks.push(rank);
}
if (ok && sawTemplateSlot) viable.push({ def, ranks });
}
if (viable.length === 0) return undefined;
if (viable.length !== candidates.length) return [];
if (viable.length <= 1) return viable.map((v) => v.def);
const dominated = new Set<number>();
for (let i = 0; i < viable.length; i++) {
if (dominated.has(i)) continue;
for (let j = i + 1; j < viable.length; j++) {
if (dominated.has(j)) continue;
const cmp = compareSpecializationRanks(viable[i].ranks, viable[j].ranks);
if (cmp < 0) dominated.add(j);
else if (cmp > 0) dominated.add(i);
}
}
return viable.filter((_, idx) => !dominated.has(idx)).map((v) => v.def);
}
function templatePartialOrderSlotRank(
paramType: string,
paramClass: ParameterTypeClass,
argClass: ParameterTypeClass,
): number | undefined {
if (!isTemplatePlaceholder(paramType)) return undefined;
if (argClass.indirection === 'unknown' || paramClass.indirection === 'unknown') {
return undefined;
}
if (isPointerShape(paramClass)) {
return isPointerShape(argClass) ? 3 : undefined;
}
if (paramClass.indirection === 'value') return 1;
return undefined;
}
function isTemplatePlaceholder(typeName: string): boolean {
return /^[A-Z]\w*$/.test(typeName);
}
/**
* Higher specialization rank is better. Returns -1 when `a` dominates `b`,
* +1 when `b` dominates `a`, and 0 for ties / incomparable vectors.
*/
function compareSpecializationRanks(a: readonly number[], b: readonly number[]): -1 | 0 | 1 {
let aBetter = false;
let bBetter = false;
const len = Math.min(a.length, b.length);
for (let i = 0; i < len; i++) {
if (a[i] > b[i]) aBetter = true;
else if (b[i] > a[i]) bBetter = true;
if (aBetter && bBetter) return 0;
}
if (aBetter && !bBetter) return -1;
if (bBetter && !aBetter) return 1;
return 0;
}
/**
* Detect when >1 candidate share identical `parameterTypes` after the
* per-language normalizer has collapsed distinct underlying types. This
@@ -116,6 +116,20 @@ export const scopeResolutionPhase: PipelinePhase<ScopeResolutionOutput> = {
preExtractedByPath.set(pf.filePath, pf);
}
// Drop pre-extracted entries for standalone providers — these
// languages are skipped by the canonical guard below (line 164)
// and never consume preExtractedByPath, so holding onto their
// entries leaks memory until the cleanup loop at 262-264 which
// also never runs for skipped providers.
for (const [path] of preExtractedByPath) {
const lang = getLanguageFromFilename(path);
if (lang === null) continue;
const provider = SCOPE_RESOLVERS.get(lang);
if (provider?.languageProvider.parseStrategy === 'standalone') {
preExtractedByPath.delete(path);
}
}
let totalFiles = 0;
let totalImports = 0;
let totalRefs = 0;
@@ -158,6 +172,14 @@ export const scopeResolutionPhase: PipelinePhase<ScopeResolutionOutput> = {
for (const [lang, provider] of SCOPE_RESOLVERS) {
if (!isRegistryPrimary(lang)) continue;
// Standalone providers (COBOL, JCL) don't emit graph edges yet
// through the scope-resolution path. This is the canonical guard:
// runScopeResolution is never called for standalone providers, which
// keeps cobolPhase as the sole IMPORTS edge producer. Keep this guard
// in sync with any additional standalone providers added to
// SCOPE_RESOLVERS.
if (provider.languageProvider.parseStrategy === 'standalone') continue;
const langFiles = scannedFiles.filter((f) => getLanguageFromFilename(f.path) === lang);
if (langFiles.length === 0) continue;
@@ -1,6 +1,6 @@
/**
* Dev-mode runtime validator for the two-channel binding lifecycle
* (Contract Invariant I8 in `contract/scope-resolver.ts`).
* Dev-mode runtime validator for the post-finalize binding-channel
* lifecycle (Contract Invariant I8 in `contract/scope-resolver.ts`).
*
* The two channels:
* - `indexes.bindings` — finalize-output channel. After
@@ -74,5 +74,21 @@ export function validateBindingsImmutability(
}
}
// Third channel: `workspaceFqnBindings` (scope-independent, shared map
// populated by language namespace-sibling hooks — PHP FQN keys, C#
// global-namespace simple names). Like bindingAugmentations its inner
// arrays are mutable by contract (hooks `push()` directly), so freezing
// one is the same defect as freezing an augmentation bucket.
for (const [name, bucket] of indexes.workspaceFqnBindings) {
if (Object.isFrozen(bucket)) {
onWarn(
`binding-immutability: indexes.workspaceFqnBindings[${name}] is FROZEN — ` +
`the workspace channel is mutable by contract; freezing it defeats the ` +
`append-only purpose. See ScopeResolver Invariant I8.`,
);
violations++;
}
}
return violations;
}
@@ -99,6 +99,14 @@ const EMPTY_NAMES: Iterable<string> = Object.freeze([]) as readonly string[];
* Fast paths (zero allocation) when at most one channel is populated:
* returns the underlying `Map.keys()` iterator directly. Only when both
* channels carry names do we materialize a `Set` for deduplication.
*
* Scope: enumerates only the per-scope `bindings` and `bindingAugmentations`
* channels. It deliberately EXCLUDES the scope-independent
* `workspaceFqnBindings` channel (PHP FQN keys, C# global-namespace simple
* names). `lookupBindingsAt` consults that third channel when resolving a
* specific name, but name *enumeration* here does not — those names apply at
* every scope and would flood per-scope callers. Callers that need
* workspace-level names must read `workspaceFqnBindings` directly.
*/
export function namesAtScope(scopeId: ScopeId, scopes: ScopeResolutionIndexes): Iterable<string> {
const finalized = scopes.bindings.get(scopeId);
@@ -241,6 +241,12 @@ export const TYPESCRIPT_QUERIES = `
[(string (string_fragment) @route.url)
(template_string) @route.template_url])) @route.fetch
; Custom fetch wrappers: apiFetch('/path'), fetchJSON('/api/data'), httpGet('/users'), etc.
(call_expression
function: (identifier) @_wrapper_fn (#match? @_wrapper_fn "^(api(Fetch|Get|Post|Put|Delete|Patch|Request)|fetch(API|JSON|Data|Endpoint|Resource|Url)|http(Fetch|Get|Post|Put|Delete|Patch|Request))$")
arguments: (arguments
(string (string_fragment) @route.url))) @route.fetch
; axios.get/post/put/delete/patch('/path'), $.get/post/ajax({url:'/path'})
(call_expression
function: (member_expression
@@ -434,6 +440,12 @@ export const JAVASCRIPT_QUERIES = `
[(string (string_fragment) @route.url)
(template_string) @route.template_url])) @route.fetch
; Custom fetch wrappers: apiFetch('/path'), fetchJSON('/api/data'), httpGet('/users'), etc.
(call_expression
function: (identifier) @_wrapper_fn (#match? @_wrapper_fn "^(api(Fetch|Get|Post|Put|Delete|Patch|Request)|fetch(API|JSON|Data|Endpoint|Resource|Url)|http(Fetch|Get|Post|Put|Delete|Patch|Request))$")
arguments: (arguments
(string (string_fragment) @route.url))) @route.fetch
; axios.get/post, $.get/post/ajax
(call_expression
function: (member_expression
@@ -1,5 +1,5 @@
import type { MethodInfo } from '../method-types.js';
import { SupportedLanguages } from 'gitnexus-shared';
import { SupportedLanguages, type ParameterTypeClass } from 'gitnexus-shared';
/** Languages where class overload signatures are declaration-only contracts
* that should collapse to the implementation body's node ID. */
@@ -139,13 +139,58 @@ export function constTagForId(
return '';
}
/**
* Disambiguate function-template overloads whose normalized parameter types
* intentionally collapse to the same placeholder token (`T`, `U`, ...), but
* whose C++ sidecar shape is semantically different (`T` vs `T*` / `T&`).
*
* Kept intentionally narrow: concrete types already use the existing raw-type
* overload tag, and non-template languages should not acquire sidecar-shaped
* IDs.
*/
export function parameterShapeIdTag(
parameterTypes?: readonly string[],
parameterTypeClasses?: readonly ParameterTypeClass[],
): string {
if (
parameterTypes === undefined ||
parameterTypeClasses === undefined ||
parameterTypes.length === 0
) {
return '';
}
let hasTemplatePlaceholder = false;
let hasDisambiguatingShape = false;
const parts: string[] = [];
for (let i = 0; i < parameterTypes.length; i++) {
const type = parameterTypes[i];
const typeClass = parameterTypeClasses[i];
if (typeClass === undefined) return '';
if (/^[A-Z]\w*$/.test(type)) hasTemplatePlaceholder = true;
if (
typeClass.indirection !== 'value' ||
typeClass.pointerDepth > 0 ||
(typeClass.cv !== 'none' && typeClass.cv !== 'unknown')
) {
hasDisambiguatingShape = true;
}
parts.push(
`${type}:${typeClass.cv}:${typeClass.indirection}:${typeClass.pointerDepth.toString()}`,
);
}
if (!hasTemplatePlaceholder || !hasDisambiguatingShape) return '';
return `~shape:${parts.join('|')}`;
}
/** Convert MethodInfo from methodExtractor into flat properties for a graph node. */
export function buildMethodProps(info: MethodInfo): Record<string, unknown> {
const types: string[] = [];
const typeClasses: ParameterTypeClass[] = [];
let optionalCount = 0;
let hasVariadic = false;
for (const p of info.parameters) {
if (p.type !== null) types.push(p.type);
if (p.typeClass !== undefined) typeClasses.push(p.typeClass);
if (p.isOptional) optionalCount++;
if (p.isVariadic) hasVariadic = true;
}
@@ -155,6 +200,9 @@ export function buildMethodProps(info: MethodInfo): Record<string, unknown> {
? { requiredParameterCount: info.parameters.length - optionalCount }
: {}),
...(types.length > 0 ? { parameterTypes: types } : {}),
...(typeClasses.length === info.parameters.length && typeClasses.length > 0
? { parameterTypeClasses: typeClasses }
: {}),
returnType: info.returnType ?? undefined,
visibility: info.visibility,
isStatic: info.isStatic,
@@ -23,6 +23,11 @@ import {
import { parseSourceSafe } from '../../tree-sitter/safe-parse.js';
import type { SymbolTableReader } from '../model/symbol-table.js';
import type { ExtractedHeritage } from '../model/heritage-map.js';
import type {
ExtractedRouterInclude,
ExtractedRouterImport,
ExtractedRouterModuleAlias,
} from '../route-extractors/fastapi-router-bindings.js';
/** Language grammar type accepted by Parser.setLanguage(). */
type TreeSitterLanguage = Parameters<typeof Parser.prototype.setLanguage>[0];
@@ -80,6 +85,7 @@ import {
typeTagForId,
constTagForId,
buildCollisionGroups,
parameterShapeIdTag,
} from '../utils/method-props.js';
import { extractTemplateArguments, templateArgumentsIdTag } from '../utils/template-arguments.js';
import type { LanguageProvider } from '../language-provider.js';
@@ -198,12 +204,30 @@ export interface ExtractedFetchCall {
lineNumber: number;
}
export interface FetchWrapperDef {
filePath: string;
functionName: string;
}
export interface ExtractedDecoratorRoute {
filePath: string;
routePath: string;
httpMethod: string;
decoratorName: string;
lineNumber: number;
/**
* Decorator receiver identifier (e.g. `router` for `@router.get(...)`,
* `app` for `@app.get(...)`). Used by parse-impl to decide which routes
* participate in `include_router(prefix=...)` joining.
*/
decoratorReceiver?: string;
/**
* FastAPI `app.include_router(prefix='/x')` prefix that applies to
* this route. Filled by parse-impl after cross-file aggregation; the
* routes phase joins it via `normalizeExtractedRoutePath`. `null` /
* absent ⇒ no prefix applies.
*/
prefix?: string | null;
}
export interface ExtractedToolDef {
@@ -268,7 +292,20 @@ export interface ParseWorkerResult {
heritage: ExtractedHeritage[];
routes: ExtractedRoute[];
fetchCalls: ExtractedFetchCall[];
fetchWrapperDefs: FetchWrapperDef[];
decoratorRoutes: ExtractedDecoratorRoute[];
routerIncludes: ExtractedRouterInclude[];
routerImports: ExtractedRouterImport[];
/**
* Optional. `from <pkg> import <module>` records from Python files
* where `<module>` is later used as a Shape-A include receiver
* (`<host>.include_router(<module>.router, prefix='/x')`). parse-impl
* uses these to promote Shape-A short-key entries to long keys, so
* same-named modules in different packages don't share prefixes.
* Optional for cache backward compatibility (older cache entries
* predate the field; consumers must guard with `if (… ?? [])`).
*/
routerModuleAliases?: ExtractedRouterModuleAlias[];
toolDefs: ExtractedToolDef[];
ormQueries: ExtractedORMQuery[];
constructorBindings: FileConstructorBindings[];
@@ -732,7 +769,11 @@ const processBatch = (
heritage: [],
routes: [],
fetchCalls: [],
fetchWrapperDefs: [],
decoratorRoutes: [],
routerIncludes: [],
routerImports: [],
routerModuleAliases: [],
toolDefs: [],
ormQueries: [],
constructorBindings: [],
@@ -772,9 +813,34 @@ const processBatch = (
for (const [language, langFiles] of byLanguage) {
const provider = getProvider(language);
const queryString = provider.treeSitterQueries;
if (!queryString) continue;
// Track if we need to handle tsx separately
if (!queryString) {
// Standalone providers (regex-based, no tree-sitter) that implement
// emitScopeCaptures feed into the scope-resolution pipeline via
// extractParsedFile directly — no tree-sitter involved.
if (provider.emitScopeCaptures) {
for (const file of langFiles) {
const parsedFile = extractParsedFile(
provider,
file.content,
file.path,
(message) => {
if (parentPort) {
parentPort.postMessage({ type: 'warning', message });
} else {
logger.warn(message);
}
},
undefined, // no cachedTree for standalone providers
);
if (parsedFile !== undefined) {
result.parsedFiles.push(parsedFile);
result.fileCount++;
onFileProcessed?.();
}
}
}
continue;
}
const tsxFiles: ParseWorkerInput[] = [];
const regularFiles: ParseWorkerInput[] = [];
@@ -842,6 +908,23 @@ const EXPRESS_ROUTE_METHODS = new Set([
'route',
]);
/**
* Walk a tree-sitter AST subtree looking for a call to the global `fetch()` function.
* Returns `true` if found within `maxDepth` levels of nesting — keeps the check
* lightweight so it doesn't slow down parse-worker on large function bodies.
*/
const checkForFetchCall = (node: SyntaxNode, depth = 0, maxDepth = 5): boolean => {
if (depth > maxDepth) return false;
if (node.type === 'call_expression') {
const fn = node.childForFieldName('function');
if (fn?.type === 'identifier' && fn.text === 'fetch') return true;
}
for (let i = 0; i < node.childCount; i++) {
if (checkForFetchCall(node.child(i)!, depth + 1, maxDepth)) return true;
}
return false;
};
// HTTP client methods that are ONLY used by clients, not Express route registration.
// Methods like get/post/put/delete/patch overlap with Express — those are captured by
// the express_route handler as route definitions, not consumers. The fetch() global
@@ -944,6 +1027,18 @@ export function extractORMQueries(
}
}
// ============================================================================
// FastAPI router prefix detection (Python)
// ============================================================================
//
// The extraction lives in `../route-extractors/fastapi-router-bindings`
// (a pure-function module — NOT a worker, no `worker_threads`, no
// `parentPort`). It's imported here only so the worker entry can call it
// per file; this module does not re-export it. Downstream consumers
// import the function and its types directly from `route-extractors/`.
import { extractFastAPIRouterBindings } from '../route-extractors/fastapi-router-bindings.js';
const processFileGroup = (
files: ParseWorkerInput[],
language: SupportedLanguages,
@@ -1176,6 +1271,7 @@ const processFileGroup = (
if (captureMap['decorator'] && captureMap['decorator.name']) {
const decoratorName = captureMap['decorator.name'].text;
const decoratorArg = captureMap['decorator.arg']?.text;
const decoratorReceiver = captureMap['decorator.receiver']?.text;
const decoratorNode = captureMap['decorator'];
// Store by the decorator's end line — the definition follows immediately after
fileDecorators.set(decoratorNode.endPosition.row, {
@@ -1195,6 +1291,7 @@ const processFileGroup = (
httpMethod,
decoratorName,
lineNumber: decoratorNode.startPosition.row + lineOffset,
...(decoratorReceiver ? { decoratorReceiver } : {}),
});
}
// MCP/RPC tool detection: @mcp.tool(), @app.tool(), @server.tool()
@@ -1718,6 +1815,13 @@ const processFileGroup = (
);
arityTag += constTagForId(defMethodMap, nodeName, arityForId, defMethodInfo, groups);
}
const parameterShapeTag =
nodeLabel === 'Function' || nodeLabel === 'Method'
? parameterShapeIdTag(
methodProps.parameterTypes as string[] | undefined,
methodProps.parameterTypeClasses as ParameterTypeClass[] | undefined,
)
: '';
const classTemplateArguments =
extractedClassSymbol?.templateArguments ??
provider.classExtractor?.extractTemplateArgumentsFromCapture?.({
@@ -1741,7 +1845,7 @@ const processFileGroup = (
: '';
const nodeId = generateId(
nodeLabel,
`${file.path}:${qualifiedName}${classTemplateTag}${arityTag}`,
`${file.path}:${qualifiedName}${classTemplateTag}${arityTag}${parameterShapeTag}`,
);
const classNodeForSymbol = definitionNode || nameNode;
const qualifiedTypeName =
@@ -1944,6 +2048,21 @@ const processFileGroup = (
: '',
});
}
// ── Fetch wrapper detection: record functions that call fetch() internally ──
if (
nodeLabel === 'Function' &&
definitionNode &&
nameNode &&
(language === SupportedLanguages.TypeScript || language === SupportedLanguages.JavaScript)
) {
if (checkForFetchCall(definitionNode)) {
result.fetchWrapperDefs.push({
filePath: file.path,
functionName: nameNode.text,
});
}
}
}
// Extract framework routes via provider detection (e.g., Laravel routes.php)
@@ -1955,6 +2074,20 @@ const processFileGroup = (
// Extract ORM queries (Prisma, Supabase)
extractORMQueries(file.path, parseContent, result.ormQueries);
// Extract FastAPI include_router(prefix=...) and `from <mod> import router`
// sites. parse-impl aggregates these into a per-module prefix map and
// injects the resolved prefix onto each ExtractedDecoratorRoute that
// came from a `@router.<verb>` decorator. Python-only.
if (language === SupportedLanguages.Python) {
extractFastAPIRouterBindings(
file.path,
parseContent,
result.routerIncludes,
result.routerImports,
(result.routerModuleAliases ??= []),
);
}
// Vue: emit CALLS edges for components used in <template>
if (language === SupportedLanguages.Vue) {
const templateComponents = extractTemplateComponents(file.content);
@@ -1985,7 +2118,11 @@ let accumulated: ParseWorkerResult = {
heritage: [],
routes: [],
fetchCalls: [],
fetchWrapperDefs: [],
decoratorRoutes: [],
routerIncludes: [],
routerImports: [],
routerModuleAliases: [],
toolDefs: [],
ormQueries: [],
constructorBindings: [],
@@ -2013,7 +2150,14 @@ const mergeResult = (target: ParseWorkerResult, src: ParseWorkerResult) => {
appendAll(target.heritage, src.heritage);
appendAll(target.routes, src.routes);
appendAll(target.fetchCalls, src.fetchCalls);
appendAll(target.fetchWrapperDefs, src.fetchWrapperDefs);
appendAll(target.decoratorRoutes, src.decoratorRoutes);
if (src.routerIncludes) appendAll(target.routerIncludes, src.routerIncludes);
if (src.routerImports) appendAll(target.routerImports, src.routerImports);
if (src.routerModuleAliases) {
target.routerModuleAliases ??= [];
appendAll(target.routerModuleAliases, src.routerModuleAliases);
}
appendAll(target.toolDefs, src.toolDefs);
appendAll(target.ormQueries, src.ormQueries);
appendAll(target.constructorBindings, src.constructorBindings);
@@ -2104,7 +2248,11 @@ parentPort!.on('message', (msg: WorkerIncomingMessage) => {
heritage: [],
routes: [],
fetchCalls: [],
fetchWrapperDefs: [],
decoratorRoutes: [],
routerIncludes: [],
routerImports: [],
routerModuleAliases: [],
toolDefs: [],
ormQueries: [],
constructorBindings: [],
+23 -1
View File
@@ -51,6 +51,28 @@ const alreadyAvailable = (message: string): boolean =>
message.includes('already exists');
const resolvePolicyFromEnv = (): ExtensionInstallPolicy => {
const raw = process.env.GITNEXUS_LBUG_EXTENSION_INSTALL;
if (raw === 'load-only' || raw === 'never' || raw === 'auto') return raw;
return 'load-only';
};
export const getExtensionInstallPolicy = (): ExtensionInstallPolicy => resolvePolicyFromEnv();
/**
* Install policy for the **analyze (write) path**.
*
* The global default (`resolvePolicyFromEnv`) is `load-only` so serve/query
* read paths never require outbound network access (PR #1161, offline-first).
* The analyze path is different: it owns building the search indexes, so it
* defaults to `auto` — LOAD the extension if present, otherwise attempt one
* bounded out-of-process INSTALL. This keeps FTS symmetric with the
* VECTOR/embeddings path (which already defaults to `auto`) and matches the
* #726 contract. An explicit `GITNEXUS_LBUG_EXTENSION_INSTALL` value still
* wins, so operators can force `load-only`/`never` for fully offline analyze;
* `auto` LOADs-first, so offline machines still degrade gracefully when the
* INSTALL cannot reach the network.
*/
export const resolveAnalyzeInstallPolicy = (): ExtensionInstallPolicy => {
const raw = process.env.GITNEXUS_LBUG_EXTENSION_INSTALL;
if (raw === 'load-only' || raw === 'never' || raw === 'auto') return raw;
return 'auto';
@@ -148,7 +170,7 @@ export const installDuckDbExtensionOutOfProcess = async (
* subsequent analyze or query calls.
*
* Policy precedence (most specific wins):
* per-call `opts.policy` → constructor `options.policy` → env → `auto`
* per-call `opts.policy` → constructor `options.policy` → env → `load-only`
*/
export class ExtensionManager {
private readonly capabilities = new Map<string, ExtensionCapability>();
+61 -15
View File
@@ -24,8 +24,10 @@ import {
deleteNodesForFile,
deleteAllCommunitiesAndProcesses,
queryImporters,
loadFTSExtension,
} from './lbug/lbug-adapter.js';
import { createSearchFTSIndexes, verifySearchFTSIndexes } from './search/fts-indexes.js';
import { resolveAnalyzeInstallPolicy } from './lbug/extension-loader.js';
import {
startWalCheckpointDriver,
type WalCheckpointDriver,
@@ -144,8 +146,26 @@ export interface AnalyzeResult {
pipelineResult?: any;
/** True when analyze only repaired FTS indexes and skipped pipeline re-analysis. */
ftsRepairedOnly?: boolean;
/**
* True when the FTS extension was unavailable so search-index creation was
* skipped (offline-first degradation). The graph is fully queryable; only
* full-text/BM25 search is disabled. Lets callers (CLI summary, server) and
* the persisted meta surface the degraded state instead of reporting healthy.
*/
ftsSkipped?: boolean;
}
/**
* Logged when the optional FTS extension cannot be loaded or installed during
* a full analyze. Kept as a named constant so the env-var/command guidance
* stays in one place (mirrors the VECTOR message in embedding-pipeline.ts).
*/
const FTS_UNAVAILABLE_MESSAGE =
'FTS extension unavailable; skipping search-index creation. ' +
'Full-text/BM25 search will be disabled until the LadybugDB FTS extension is ' +
'installed once with network access (GITNEXUS_LBUG_EXTENSION_INSTALL=auto) or ' +
'pre-installed for offline use. Run `gitnexus doctor` for details.';
// Re-export the pure flag-derivation helper so external callers (and tests)
// keep importing from this module's stable surface.
export { deriveEmbeddingMode, DEFAULT_EMBEDDING_NODE_LIMIT } from './embedding-mode.js';
@@ -684,23 +704,41 @@ export async function runFullAnalysis(
}
// ── Phase 3: FTS (85–90%) ─────────────────────────────────────────
// The analyze (write) path owns building the search indexes, so it uses
// the `auto` install policy (LOAD-first, then one bounded INSTALL) —
// symmetric with the VECTOR/embeddings path below and consistent with the
// #726 contract. The global `load-only` default (PR #1161) governs the
// serve/query read paths, not this one. When the extension still cannot be
// loaded (genuinely offline + not pre-installed, or policy forced to
// load-only/never), degrade gracefully — exactly like the VECTOR path — so
// analyze still produces a fully queryable graph; only full-text/BM25
// search falls back. `--repair-fts` (whose sole job is FTS) still fails
// loudly on its own path above.
progress('fts', 85, 'Creating search indexes...');
await createSearchFTSIndexes({
onIndexStart: options.verbose
? (table, indexName) => log(`FTS: creating ${table}.${indexName}`)
: undefined,
onIndexReady: options.verbose
? (table, indexName) => log(`FTS: ready ${table}.${indexName}`)
: undefined,
const ftsAvailable = await loadFTSExtension(undefined, {
policy: resolveAnalyzeInstallPolicy(),
});
const missingIndexNames = await verifySearchFTSIndexes(executeQuery);
if (missingIndexNames.length > 0) {
throw new Error(
`FTS verification failed - missing indexes after analyze: ${missingIndexNames.join(', ')}. ` +
'Check FTS extension availability, then retry `gitnexus analyze --force` for a full rebuild.',
);
if (ftsAvailable) {
await createSearchFTSIndexes({
onIndexStart: options.verbose
? (table, indexName) => log(`FTS: creating ${table}.${indexName}`)
: undefined,
onIndexReady: options.verbose
? (table, indexName) => log(`FTS: ready ${table}.${indexName}`)
: undefined,
});
const missingIndexNames = await verifySearchFTSIndexes(executeQuery);
if (missingIndexNames.length > 0) {
throw new Error(
`FTS verification failed - missing indexes after analyze: ${missingIndexNames.join(', ')}. ` +
'Check FTS extension availability, then retry `gitnexus analyze --force` for a full rebuild.',
);
}
progress('fts', 90, 'Search indexes ready');
} else {
log(FTS_UNAVAILABLE_MESSAGE);
progress('fts', 90, 'Search indexes skipped (FTS unavailable)');
}
progress('fts', 90, 'Search indexes ready');
// ── Phase 3.5: Re-insert cached embeddings ────────────────────────
// Runs on BOTH the full-rebuild path and the incremental path:
@@ -889,7 +927,14 @@ export async function runFullAnalysis(
},
capabilities: {
graph: { provider: 'ladybugdb', status: runtimeCapabilities.graph },
fts: { provider: 'ladybugdb-fts', status: runtimeCapabilities.fts },
// Reflect what this analyze run actually produced: when the FTS
// extension was unavailable the indexes were skipped, so record
// 'unavailable' rather than the static runtime default. Keeps
// meta.json / `gitnexus doctor` honest about degraded search.
fts: {
provider: 'ladybugdb-fts',
status: ftsAvailable ? runtimeCapabilities.fts : 'unavailable',
},
vectorSearch: {
provider: effectiveSemanticMode === 'vector-index' ? 'ladybugdb-vector' : 'exact-scan',
status: embeddingCount > 0 ? effectiveSemanticMode : 'unavailable',
@@ -989,6 +1034,7 @@ export async function runFullAnalysis(
repoPath,
stats: meta.stats,
pipelineResult,
ftsSkipped: !ftsAvailable,
};
} catch (err) {
// Ensure LadybugDB is closed even on error. Stop the driver first
+124 -13
View File
@@ -2907,9 +2907,11 @@ export class LocalBackend {
limit?: number;
offset?: number;
summaryOnly?: boolean;
skipPerSymbolEnrichment?: boolean;
},
): Promise<any> {
const { maxDepth, relationTypes, includeTests, minConfidence } = opts;
const skipPerSymbolEnrichment = opts.skipPerSymbolEnrichment ?? false;
const hasExplicitLimit = typeof opts.limit === 'number' && Number.isFinite(opts.limit);
const paginationLimit = hasExplicitLimit
? Math.max(1, Math.min(Math.trunc(opts.limit!), 10000))
@@ -2918,8 +2920,14 @@ export class LocalBackend {
typeof opts.offset === 'number' && Number.isFinite(opts.offset) ? opts.offset : 0;
const paginationOffset = Math.max(0, Math.trunc(rawOffset));
const summaryOnly = opts.summaryOnly ?? false;
const relTypeFilter = relationTypes.map((t) => `'${t}'`).join(', ');
const confidenceFilter = minConfidence > 0 ? ` AND r.confidence >= ${minConfidence}` : '';
// Bind the BFS frontier query's filters as parameters (#1907 review F5):
// node ids and relation types as bound lists, the confidence floor as a
// bound number — no string interpolation reaches the query text. Preserve
// the original "no confidence clause when minConfidence <= 0" behavior: an
// unconditional `>= 0` would wrongly exclude NULL-confidence edges that the
// unfiltered query includes.
const safeMinConfidence = Number.isFinite(minConfidence) ? minConfidence : 0;
const confidenceFilter = safeMinConfidence > 0 ? ' AND r.confidence >= $minConfidence' : '';
const symId = sym.id || sym[0];
@@ -3007,15 +3015,19 @@ export class LocalBackend {
for (let depth = 1; depth <= maxDepth && frontier.length > 0; depth++) {
const nextFrontier: string[] = [];
// Batch frontier nodes into a single Cypher query per depth level
const idList = frontier.map((id) => `'${id.replace(/'/g, "''")}'`).join(', ');
// Batch frontier nodes into a single Cypher query per depth level.
// ids/types/confidence are bound parameters (see above) — no interpolation.
const query =
direction === 'upstream'
? `MATCH (caller)-[r:CodeRelation]->(n) WHERE n.id IN [${idList}] AND r.type IN [${relTypeFilter}]${confidenceFilter} RETURN n.id AS sourceId, caller.id AS id, caller.name AS name, labels(caller)[0] AS type, caller.filePath AS filePath, r.type AS relType, r.confidence AS confidence`
: `MATCH (n)-[r:CodeRelation]->(callee) WHERE n.id IN [${idList}] AND r.type IN [${relTypeFilter}]${confidenceFilter} RETURN n.id AS sourceId, callee.id AS id, callee.name AS name, labels(callee)[0] AS type, callee.filePath AS filePath, r.type AS relType, r.confidence AS confidence`;
? `MATCH (caller)-[r:CodeRelation]->(n) WHERE n.id IN $frontierIds AND r.type IN $relTypes${confidenceFilter} RETURN n.id AS sourceId, caller.id AS id, caller.name AS name, labels(caller)[0] AS type, caller.filePath AS filePath, r.type AS relType, r.confidence AS confidence`
: `MATCH (n)-[r:CodeRelation]->(callee) WHERE n.id IN $frontierIds AND r.type IN $relTypes${confidenceFilter} RETURN n.id AS sourceId, callee.id AS id, callee.name AS name, labels(callee)[0] AS type, callee.filePath AS filePath, r.type AS relType, r.confidence AS confidence`;
try {
const related = await executeQuery(repo.id, query);
const related = await executeParameterized(repo.id, query, {
frontierIds: frontier,
relTypes: relationTypes,
...(safeMinConfidence > 0 ? { minConfidence: safeMinConfidence } : {}),
});
for (const rel of related) {
const relId = rel.id || rel[1];
@@ -3066,13 +3078,25 @@ export class LocalBackend {
const directCount = (grouped[1] || []).length;
let affectedProcesses: any[] = [];
let affectedModules: any[] = [];
// Per-symbol process membership: maps impacted symbol id -> list of processes
// it participates in. Populated by a second chunked Cypher pass below when
// any process is affected at all. Surfaced as `processes: [...]` on each
// byDepth item so consumers can tell which caller belongs to which cron/
// webhook/route without a follow-up query.
const perSymbolProcesses = new Map<
string,
Array<{ id: string; label: string; processType: string; step: number }>
>();
// Chunking bounds for batched DB round-trips. Declared at function scope so
// both the in-block enrichment passes and the post-pagination per-symbol
// process enrichment can reference them.
const CHUNK_SIZE = 100;
// Max number of chunks to process to avoid unbounded DB round-trips.
// Configurable via env IMPACT_MAX_CHUNKS, default 10 => max items = 1000
const MAX_CHUNKS = parseInt(process.env.IMPACT_MAX_CHUNKS || '10', 10);
if (impacted.length > 0) {
const CHUNK_SIZE = 100;
// Max number of chunks to process to avoid unbounded DB round-trips.
// Configurable via env IMPACT_MAX_CHUNKS, default 10 => max items = 1000
const MAX_CHUNKS = parseInt(process.env.IMPACT_MAX_CHUNKS || '10', 10);
// ── Process enrichment: batched chunking (bounded by MAX_CHUNKS) ─
// Uses merged Cypher query (WITH + OPTIONAL MATCH) to fetch
// process + entry point info in 1 round-trip per chunk. Converted to
@@ -3218,6 +3242,10 @@ export class LocalBackend {
}))
.sort((a, b) => b.total_hits - a.total_hits);
// Per-symbol process membership is populated post-pagination (see below)
// so it covers exactly the symbols returned in byDepth, not a pre-capped
// flat slice that could miss depth-2+ symbols when depth-1 is large.
// ── Module enrichment: use same cap as process enrichment and parameterized queries
const maxItems = Math.min(impacted.length, MAX_CHUNKS * CHUNK_SIZE);
const cappedImpacted = impacted.slice(0, maxItems);
@@ -3360,7 +3388,7 @@ export class LocalBackend {
return base;
}
// Apply limit/offset pagination per depth level
// Apply limit/offset pagination per depth level.
const paginatedGrouped: Record<number, any[]> = {};
let anyTruncated = false;
for (const [depth, items] of Object.entries(grouped)) {
@@ -3372,8 +3400,82 @@ export class LocalBackend {
}
}
// ── Per-symbol process membership enrichment (post-pagination) ───────
// Runs after paginatedGrouped is built so we enrich only the IDs that
// actually appear in the response. This eliminates the false-empty
// processes:[] case where a depth-2+ symbol's flat position in `impacted`
// exceeded MAX_CHUNKS*CHUNK_SIZE even though it is returned by byDepth.
// Also uses DISTINCT + MIN(r.step) per (symbol, process) pair to avoid
// duplicate entries when a symbol has multiple STEP_IN_PROCESS edges.
// Skipped entirely when `skipPerSymbolEnrichment` is set (group cross-repo
// fan-out, which consumes byDepth but not byDepth[].processes); the
// attach-loop below still stamps an empty processes:[] for shape stability.
let perSymbolEnrichmentCapped = false;
if (affectedProcesses.length > 0 && !skipPerSymbolEnrichment) {
// Collect unique IDs from the paginated result in one pass.
const pageIds = new Set<string>();
for (const items of Object.values(paginatedGrouped)) {
for (const it of items) {
const id = String(it.id ?? '');
if (id) pageIds.add(id);
}
}
// Bound the enrichment to the same ceiling as the aggregation pass
// (MAX_CHUNKS * CHUNK_SIZE) so a large paginated page cannot trigger
// unbounded DB round-trips (DoD 2.6). When capped, mark the result
// partial so callers know some returned symbols may carry an empty
// processes:[] that is a cap artifact, not a true absence.
const maxPageIds = MAX_CHUNKS * CHUNK_SIZE;
let pageIdArr = Array.from(pageIds);
if (pageIdArr.length > maxPageIds) {
pageIdArr = pageIdArr.slice(0, maxPageIds);
perSymbolEnrichmentCapped = true;
}
for (let i = 0; i < pageIdArr.length; i += CHUNK_SIZE) {
const chunkIds = pageIdArr.slice(i, i + CHUNK_SIZE);
try {
const rows = await executeParameterized(
repo.id,
`
MATCH (s)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
WHERE s.id IN $ids
RETURN s.id AS sid, p.id AS pid, p.heuristicLabel AS pName,
p.processType AS pType, MIN(r.step) AS step
`,
{ ids: chunkIds },
).catch(() => []);
for (const row of rows) {
const sid = row.sid ?? row[0];
if (!sid) continue;
const procEntry = {
id: String(row.pid ?? row[1] ?? ''),
label: String(row.pName ?? row[2] ?? ''),
processType: String(row.pType ?? row[3] ?? ''),
step: Number(row.step ?? row[4] ?? -1),
};
const list = perSymbolProcesses.get(String(sid));
if (list) list.push(procEntry);
else perSymbolProcesses.set(String(sid), [procEntry]);
}
} catch (e) {
logQueryError('impact:per-symbol-process-chunk', e);
}
}
}
// Attach processes field to each paginated item.
for (const items of Object.values(paginatedGrouped)) {
for (const it of items) {
it.processes = perSymbolProcesses.get(String(it.id)) ?? [];
}
}
return {
...base,
// Surface partial if the per-symbol enrichment was capped, even when the
// BFS traversal itself completed — some returned symbols may carry an
// empty processes:[] that is a cap artifact rather than a true absence.
...(perSymbolEnrichmentCapped && { partial: true }),
...(anyTruncated && {
pagination: {
...(Number.isFinite(paginationLimit) && { limit: paginationLimit }),
@@ -3467,11 +3569,20 @@ export class LocalBackend {
];
try {
// skipPerSymbolEnrichment suppresses ONLY the per-symbol STEP_IN_PROCESS
// enrichment pass while preserving byDepth. Group-mode cross-repo fan-out
// may fan across many repos; the per-symbol pass adds up to MAX_CHUNKS
// extra round-trips per repo, which is unacceptable at group scale. But
// cross-impact fan-out DOES consume byDepth (cross-impact.ts reads
// fan.byDepth to populate group by_depth), so summaryOnly would wrongly
// drop it. Group callers do not consume byDepth[].processes, so skipping
// only that enrichment is the correct, targeted suppression.
return await this._runImpactBFS(repo, sym, symType, dir, {
maxDepth: opts.maxDepth,
relationTypes,
includeTests: opts.includeTests,
minConfidence: opts.minConfidence,
skipPerSymbolEnrichment: true,
});
} catch {
return null;
+50 -8
View File
@@ -284,6 +284,37 @@ Follow these steps:
/**
* Start the MCP server on stdio transport (for CLI use).
*/
/** Force-exit fallback budget if graceful shutdown cleanup hangs. */
const SHUTDOWN_FORCE_EXIT_MS = 5_000;
/** Conventional 128 + signal-number exit codes for graceful termination. */
export const SHUTDOWN_EXIT_CODES = { SIGINT: 130, SIGTERM: 143 } as const;
type SignalRegistrar = (
event: 'SIGINT' | 'SIGTERM',
listener: (...args: unknown[]) => void,
) => void;
/**
* Wire SIGINT/SIGTERM to a graceful shutdown using NUMERIC exit codes.
*
* Node invokes signal listeners with the signal NAME string as the first
* argument, so registering an `(exitCode = 0) => process.exit(exitCode)`
* shutdown directly passes `'SIGTERM'` into `process.exit()` and crashes with
* `ERR_INVALID_ARG_TYPE` (#1132). These wrappers discard the signal argument
* and pass the conventional 128+signal code instead. `on` is injectable so the
* mapping can be unit-tested without touching the real process.
*/
export function installSignalShutdown(
shutdown: (exitCode?: number) => unknown,
on: SignalRegistrar = (event, listener) => {
process.on(event, listener);
},
): void {
on('SIGINT', () => void shutdown(SHUTDOWN_EXIT_CODES.SIGINT));
on('SIGTERM', () => void shutdown(SHUTDOWN_EXIT_CODES.SIGTERM));
}
export async function startMCPServer(backend: LocalBackend): Promise<void> {
const server = createMCPServer(backend);
@@ -321,6 +352,11 @@ export async function startMCPServer(backend: LocalBackend): Promise<void> {
const shutdown = async (exitCode = 0) => {
if (shuttingDown) return;
shuttingDown = true;
// Safety net: if backend.disconnect()/server.close() hangs, still exit so a
// SIGINT/SIGTERM reliably terminates the process. Unref'd so the timer alone
// never keeps the event loop alive.
const forceExit = setTimeout(() => process.exit(exitCode), SHUTDOWN_FORCE_EXIT_MS);
forceExit.unref();
try {
await backend.disconnect();
} catch {}
@@ -329,12 +365,16 @@ export async function startMCPServer(backend: LocalBackend): Promise<void> {
} catch {}
const { flushLoggerSync } = await import('../core/logger.js');
flushLoggerSync();
clearTimeout(forceExit);
process.exit(exitCode);
};
// Handle graceful shutdown
process.on('SIGINT', shutdown);
process.on('SIGTERM', shutdown);
// Handle graceful shutdown. Node invokes signal listeners with the signal
// NAME (e.g. 'SIGTERM') as the first argument; registering `shutdown`
// directly passed that string to process.exit() and crashed with
// ERR_INVALID_ARG_TYPE (#1132). Map each signal to its conventional
// 128+signal exit code instead.
installSignalShutdown(shutdown);
// Log crashes to stderr so they aren't silently lost.
// uncaughtException is fatal — shut down.
@@ -342,14 +382,16 @@ export async function startMCPServer(backend: LocalBackend): Promise<void> {
// killing the server for one missed catch would be worse than logging it.
process.on('uncaughtException', (err) => {
process.stderr.write(`GitNexus MCP uncaughtException: ${err?.stack || err}\n`);
shutdown(1);
void shutdown(1);
});
process.on('unhandledRejection', (reason: any) => {
process.stderr.write(`GitNexus MCP unhandledRejection: ${reason?.stack || reason}\n`);
});
// Handle stdio errors — stdin close means the parent process is gone
process.stdin.on('end', shutdown);
process.stdin.on('error', () => shutdown());
process.stdout.on('error', () => shutdown());
// Handle stdio errors — stdin close means the parent process is gone.
// Wrap so the event payload (e.g. an Error for 'error') can never reach
// process.exit() as a non-numeric exit code, and void the returned promise.
process.stdin.on('end', () => void shutdown(0));
process.stdin.on('error', () => void shutdown(0));
process.stdout.on('error', () => void shutdown(0));
}
+1 -1
View File
@@ -336,7 +336,7 @@ Output includes:
- summary: direct callers, processes affected, modules affected
- affected_processes: which execution flows break and at which step
- affected_modules: which functional areas are hit (direct vs indirect)
- byDepth: affected symbols grouped by traversal depth (paginated by limit/offset; omitted when summaryOnly:true — use byDepthCounts for totals per depth, pagination object when truncated)
- byDepth: affected symbols grouped by traversal depth (paginated by limit/offset; omitted when summaryOnly:true — use byDepthCounts for totals per depth, pagination object when truncated). Each item includes a processes:[{id,label,processType,step}] field listing the execution flows that symbol participates in. Empty when the symbol has no process membership. Can ALSO be empty when partial:true is set — either the process-aggregation pass hit its cap before detecting affected processes, or per-symbol enrichment was capped on a very large page. When partial:true, do NOT treat processes:[] as proof of no participation; cross-check the top-level affected_processes list.
Depth groups:
- d=1: WILL BREAK (direct callers/importers)
+1 -1
View File
@@ -44,7 +44,7 @@ import type { ParseWorkerResult } from '../core/ingestion/workers/parse-worker.j
* On version mismatch, `loadParseCache` returns an empty cache and the
* next save overwrites the on-disk file with the new version baked in.
*/
const SCHEMA_BUMP = 1;
const SCHEMA_BUMP = 2;
const GITNEXUS_PKG_VERSION = (() => {
try {
// package.json sits at gitnexus/package.json — two levels up from
@@ -0,0 +1,14 @@
from fastapi import APIRouter
# Same module name as `api/users.py`. Before the long-key fix the
# basename `users` collided across packages, leaking `/users` (the
# prefix mounted on `api/users.py`) onto these admin routes.
# parse-impl now keys prefixes by `<dir>/<stem>` whenever the import
# statement carried enough context, so this file's `@router.get`
# routes must NOT be prefixed with `/users`.
router = APIRouter()
@router.get("/audit")
def audit():
return []
@@ -0,0 +1,8 @@
from fastapi import APIRouter
router = APIRouter()
@router.get("/list")
def list_calls():
return []
+13
View File
@@ -0,0 +1,13 @@
from fastapi import APIRouter
router = APIRouter()
@router.get("/list")
def list_users():
return []
@router.post("/create")
def create_user(payload):
return {"ok": True}
+13
View File
@@ -0,0 +1,13 @@
from fastapi import FastAPI
from api import users
from api.calls import router as calls_router
from .relative import router as rel_router
# Hostname is `application`, NOT `app` — exercises the unrestricted-host
# path through both the parse-worker regex and the group-layer
# tree-sitter pattern. Pinning to literal `app` would silently drop
# every prefix here.
application = FastAPI()
application.include_router(users.router, prefix="/users", tags=["users"])
application.include_router(calls_router, prefix="/calls")
application.include_router(rel_router, prefix="/rel")
+8
View File
@@ -0,0 +1,8 @@
from fastapi import APIRouter
router = APIRouter()
@router.get("/info")
def info():
return {}
@@ -0,0 +1,362 @@
{
"go-aliased-package-import/internal/util/log.go": {
"captureGroups": 4,
"digest": "77be0bd9a9df464f42a6cefd0c16064d4ef97e79f5f61faa8df6069989cec9f6"
},
"go-aliased-package-import/main.go": {
"captureGroups": 7,
"digest": "3eb2e6d441dadede554b271b7a87b7b9ca557ab418bfb9e8524e6ace4c8fc547"
},
"go-ambiguous/internal/models/handler.go": {
"captureGroups": 9,
"digest": "619c516a5791095bc6380de62fa861364b3f9480f2ec87d4499c2c998f228713"
},
"go-ambiguous/internal/other/handler.go": {
"captureGroups": 9,
"digest": "08a61721581c4f17741ef0c4ee1c8945235ee7f2aca7d88d2fe122071cd62e6f"
},
"go-ambiguous/internal/services/user.go": {
"captureGroups": 8,
"digest": "802b81a07c64c01f2381cf33d41cdb1a58f535c0007b5ae151c38cda87f28321"
},
"go-assignment-chain/cmd/main.go": {
"captureGroups": 50,
"digest": "47ba5fd2ea96ee202b3a3de5c0db75ea18d0889064ed2138d62c67594f7d22e9"
},
"go-assignment-chain/models/repo.go": {
"captureGroups": 8,
"digest": "1b3acbf48751105d056253488f34159266bdb0b229248c85065e927b72bf4801"
},
"go-assignment-chain/models/user.go": {
"captureGroups": 8,
"digest": "457c76ebf1c86efac7a7d9f384a8e49acd5649c006e6cc8aa29a59c36f0b102a"
},
"go-call-result-binding/cmd/main.go": {
"captureGroups": 16,
"digest": "83d611dcee826ec848a0b3fd5a18a98103896d17c353d8d63cd15551039818e9"
},
"go-call-result-binding/models/user.go": {
"captureGroups": 10,
"digest": "5c31199e148340ee32dd48e32a00003e094be963b1287b9b40c6b4effca12175"
},
"go-calls/cmd/main.go": {
"captureGroups": 6,
"digest": "deaea55087c2652aa2d39fefa30048d5393cec40f7af1ab13921331b035030d6"
},
"go-calls/internal/onearg/log.go": {
"captureGroups": 6,
"digest": "a101eafe9f08396bb176cf3cadd3960d752802048e47c0ac9e9ace07244a46fb"
},
"go-calls/internal/zeroarg/log.go": {
"captureGroups": 5,
"digest": "7b322767a38298de8c6ce99fa6b1d69dda8bbb79704053644bfaa7d2e4473fdb"
},
"go-chain-call/cmd/main.go": {
"captureGroups": 20,
"digest": "5a8d7de8ae87887902d16f28cb70d31014c96ab5d2eb63cbd6c8be27650fa9a1"
},
"go-chain-call/models/repo.go": {
"captureGroups": 10,
"digest": "ac4799aae638d528c5c7c01c8b9734fc61e995596a12b12e790dcffab97b4029"
},
"go-chain-call/models/user.go": {
"captureGroups": 10,
"digest": "5c31199e148340ee32dd48e32a00003e094be963b1287b9b40c6b4effca12175"
},
"go-child-extends-parent/models/child.go": {
"captureGroups": 3,
"digest": "6fd9fe7b82066f82a93bf5e5024ddf89382091ec04e55648845c8295f13bd412"
},
"go-child-extends-parent/models/parent.go": {
"captureGroups": 8,
"digest": "454a724f571a5aede89d7657f76ed8e3c9b19bfc18d524141d1b646005f98e49"
},
"go-child-extends-parent/services/app.go": {
"captureGroups": 10,
"digest": "05f5df0369c90bc6e0da5a8a76e0661be36cddbe913ab2f83da2ee3733a46ec2"
},
"go-cmd-helper/cmd/server/internal/config/config.go": {
"captureGroups": 5,
"digest": "b10874198d380b0a186fb1e1ee8cedac644eb3b6b624398110a660183e59a0b5"
},
"go-cmd-helper/cmd/server/main.go": {
"captureGroups": 7,
"digest": "a1f9453bd71926d60e3f148f43b9af813cbd1cccc11b323896a55bdb443f8931"
},
"go-constructor-type-inference/cmd/main.go": {
"captureGroups": 15,
"digest": "4c496bf5ebaebae8c7480b1826b3285752a79ce50562756ded3b7fe0c6d6b325"
},
"go-constructor-type-inference/models/repo.go": {
"captureGroups": 8,
"digest": "1b3acbf48751105d056253488f34159266bdb0b229248c85065e927b72bf4801"
},
"go-constructor-type-inference/models/user.go": {
"captureGroups": 8,
"digest": "457c76ebf1c86efac7a7d9f384a8e49acd5649c006e6cc8aa29a59c36f0b102a"
},
"go-deep-field-chain/cmd/main.go": {
"captureGroups": 13,
"digest": "ca9cc7ae0f75928b1ea338f42e58cf02502e0c93ce4ea8867258c90543f79d16"
},
"go-deep-field-chain/models/models.go": {
"captureGroups": 33,
"digest": "0ad1df946f58e446a8dbd8ef13041b6e0177f53293946b955acf3c46684e4295"
},
"go-field-types/cmd/main.go": {
"captureGroups": 9,
"digest": "3bc98cf5640568cdff3202f2594fb39d563ee595560cb28d1adb31f264670ceb"
},
"go-field-types/models/models.go": {
"captureGroups": 22,
"digest": "fbdbf74d927c0ed07f4820c190fc6f2039b74ddead8abb45fc238812dfa4d4ef"
},
"go-for-call-expr/cmd/main.go": {
"captureGroups": 27,
"digest": "95217f85260d85baeb57638fff470c0a358be92cc49a5f918b084d809f47ba2a"
},
"go-for-call-expr/models/repo.go": {
"captureGroups": 14,
"digest": "312e59c1402cb83a4ce51c27866fde6c597f7121098b57ff1d3bd4dc6363834a"
},
"go-for-call-expr/models/user.go": {
"captureGroups": 14,
"digest": "c442a26c4051c2b7426850a381137507a147c21d631d6535d71ef1035b1506fb"
},
"go-inc-dec-write-access/main.go": {
"captureGroups": 21,
"digest": "0414398239624e44b1589f6a68f2c19636bf4b3ce74990cee1bd7cbeeb0591e6"
},
"go-local-shadow/cmd/main.go": {
"captureGroups": 12,
"digest": "1cda9982d9b4e4878208894523f0a9c7a806c6f08a5582d4771205716a1b2733"
},
"go-local-shadow/internal/utils/utils.go": {
"captureGroups": 6,
"digest": "51e941fa4c7765efc6e472d15a9d7ea31f59b67a6b606be868f4af47e5f54cb1"
},
"go-make-builtin/main.go": {
"captureGroups": 18,
"digest": "7c56328d8416338ae0075ae7dd9669b16f2035aa353fac7ee1e46109a58e80a6"
},
"go-make-builtin/models.go": {
"captureGroups": 15,
"digest": "4d5628f66471f1ad1180a47b797195506d190ddc84f61f6c57f2e22afd754c1e"
},
"go-map-range/main.go": {
"captureGroups": 10,
"digest": "7f3580a0e7858e6eb3176e0b5e1bec4867a2d5d07f2b216a04a3de79ff96170b"
},
"go-map-range/models/repo.go": {
"captureGroups": 9,
"digest": "6cbc4422fb287007735b4c58b5e9c84bbd6b8f6082e3b4d5fe48c306e6acc75b"
},
"go-map-range/models/user.go": {
"captureGroups": 9,
"digest": "7e4dbc05ad1de859cd3103c27cddb60566d95a8a8354c752e8c75f84d7ca055c"
},
"go-member-calls/cmd/main.go": {
"captureGroups": 11,
"digest": "e46f6d1dff39943f8f89c05e0d28f61f8471cdc729c91f01ca693a38a244b8ba"
},
"go-member-calls/models/user.go": {
"captureGroups": 8,
"digest": "457c76ebf1c86efac7a7d9f384a8e49acd5649c006e6cc8aa29a59c36f0b102a"
},
"go-method-chain-binding/cmd/main.go": {
"captureGroups": 21,
"digest": "0fff0df5dd77e13e0b9e62dd7f38efc038d33ffd5d891f684289eccbc46ff583"
},
"go-method-chain-binding/models/user.go": {
"captureGroups": 24,
"digest": "504ea598176dd8e01d759cc54e012735c746f14ec3a660af2c362fc356326f65"
},
"go-method-enrichment/animal.go": {
"captureGroups": 15,
"digest": "ac0933f59d4a88a25629d02f308c66847fd34ceba5831660fc8bff07109b7708"
},
"go-method-enrichment/app.go": {
"captureGroups": 15,
"digest": "ac0bdc2e6daf7d4e28fd255fe2edd143e8fab7ff316f4e32e2ca04a3b33f71c2"
},
"go-mixed-chain/cmd/main.go": {
"captureGroups": 20,
"digest": "3662b803da4f4fc1ac0552a45d3d262bd0db57f97b056614caacea8ad856250e"
},
"go-mixed-chain/models/models.go": {
"captureGroups": 41,
"digest": "dc2cedaefcd73faf13a9f704a4be8f0bd4140c608cd86480e1f4400fec49dbd1"
},
"go-multi-assign/app.go": {
"captureGroups": 16,
"digest": "66958a3e86caa54aedd795227dac274cfb3f4b39ab98964e8a7bcbdfc8a08ca6"
},
"go-multi-assign/models.go": {
"captureGroups": 19,
"digest": "a4d43cf2cd2f7bdbc750a7ee1f5e9c7611e5d1e25f1a46386d0eeb9f73bb7846"
},
"go-multi-return-inference/cmd/main.go": {
"captureGroups": 36,
"digest": "6229b69cfd15bf1465d770d97486522faac5a91cf8828c424c8c8c989b310224"
},
"go-multi-return-inference/models/repo.go": {
"captureGroups": 10,
"digest": "ac4799aae638d528c5c7c01c8b9734fc61e995596a12b12e790dcffab97b4029"
},
"go-multi-return-inference/models/user.go": {
"captureGroups": 10,
"digest": "5c31199e148340ee32dd48e32a00003e094be963b1287b9b40c6b4effca12175"
},
"go-new-builtin/main.go": {
"captureGroups": 12,
"digest": "1177b99217a42a28b0d768d29e1df4198f0d5ca5f1c2b584d528f5f02354b38f"
},
"go-new-builtin/models.go": {
"captureGroups": 17,
"digest": "40c9f89942406caf2610f0d1954b66e905c39dd5d617d7359e572445d586c15d"
},
"go-nullable-receiver/cmd/main.go": {
"captureGroups": 27,
"digest": "7a93b339f8c68ed3882802da21343c1d6b5231ae5cd8dd13eb4c05e2e7716320"
},
"go-nullable-receiver/models/repo.go": {
"captureGroups": 8,
"digest": "1b3acbf48751105d056253488f34159266bdb0b229248c85065e927b72bf4801"
},
"go-nullable-receiver/models/user.go": {
"captureGroups": 8,
"digest": "457c76ebf1c86efac7a7d9f384a8e49acd5649c006e6cc8aa29a59c36f0b102a"
},
"go-parent-resolution/models/base.go": {
"captureGroups": 8,
"digest": "2afaeb50d544a55fe437ef20e2c0de92152d2ba62f2693c329255787bb3d0a02"
},
"go-parent-resolution/models/user.go": {
"captureGroups": 8,
"digest": "c76ba16343dd94024fdaac10fa7640536966e530d675c0084434f42edb5b5f15"
},
"go-pkg/cmd/main.go": {
"captureGroups": 14,
"digest": "a7781d23802e876b9ca2951b7ebb317c5df5adf390ae9a57e8e2906af34f3255"
},
"go-pkg/internal/auth/service.go": {
"captureGroups": 17,
"digest": "2204643b50f486423ee7a5877b2bab7d6334cbe62b14b4435fe4ba8a6465ce92"
},
"go-pkg/internal/models/admin.go": {
"captureGroups": 13,
"digest": "1a5ec9fd5e752adcfec91cd03b3c2a67c124852228a527002237ec51cd4b39b7"
},
"go-pkg/internal/models/repository.go": {
"captureGroups": 3,
"digest": "7de6e11a3cf9c89afa9d89fe37dab85b208699f20105fbade223e4f78a63b15c"
},
"go-pkg/internal/models/user.go": {
"captureGroups": 13,
"digest": "e56fcea1c473866ed06fc0262702e54c72016a9556cc6337330c03cc1f638fe1"
},
"go-pointer-constructor-inference/cmd/main.go": {
"captureGroups": 15,
"digest": "39f9030e909a37f0e13724f50673dd4a72ba240fb6ac263192ee4feadbc88adf"
},
"go-pointer-constructor-inference/models/repo.go": {
"captureGroups": 10,
"digest": "ac4799aae638d528c5c7c01c8b9734fc61e995596a12b12e790dcffab97b4029"
},
"go-pointer-constructor-inference/models/user.go": {
"captureGroups": 10,
"digest": "5c31199e148340ee32dd48e32a00003e094be963b1287b9b40c6b4effca12175"
},
"go-receiver-method-free-call/example.go": {
"captureGroups": 8,
"digest": "2a3c26672d3b997bdc39644361c550f8cf0749489f0945d210e2fb7f3bca9383"
},
"go-receiver-method-free-call/util.go": {
"captureGroups": 4,
"digest": "0ac9740c13c851ca101e074bd16422f9da45b8c27db4befbc0c5e7479c2ccdda"
},
"go-receiver-resolution/cmd/main.go": {
"captureGroups": 13,
"digest": "e92c59312a46972a6083ba489888b550cdd024b809acbbf367ad025efb844a5c"
},
"go-receiver-resolution/models/repo.go": {
"captureGroups": 8,
"digest": "1b3acbf48751105d056253488f34159266bdb0b229248c85065e927b72bf4801"
},
"go-receiver-resolution/models/user.go": {
"captureGroups": 8,
"digest": "457c76ebf1c86efac7a7d9f384a8e49acd5649c006e6cc8aa29a59c36f0b102a"
},
"go-return-type-inference/cmd/main.go": {
"captureGroups": 39,
"digest": "ac8ca3dc1fb7fb4d7a1f947a77e93db890f04d328135dc514f36ba5cd04bc3e1"
},
"go-return-type-inference/models/repo.go": {
"captureGroups": 16,
"digest": "8c500db82093b4622acbca734f3040e6176000a9d9a4843e2a81e3f2b3daee7c"
},
"go-return-type-inference/models/user.go": {
"captureGroups": 16,
"digest": "a70b19b9c46a02003f26d0b70fa5368983a791583ec3776e983c733a9283eefa"
},
"go-same-package-factory/main.go": {
"captureGroups": 14,
"digest": "505d4d279615f3c99457d8cb0f29fdf946b547c6b1fda6d43aa6d4302dd61809"
},
"go-same-package-factory/repo.go": {
"captureGroups": 8,
"digest": "f8bb213588f517e166b421f19d72cde981153464cc3f3168499769e778c4a22e"
},
"go-same-package-factory/user.go": {
"captureGroups": 8,
"digest": "4daaad60f35d4519a24cd067baf4cd4bdc421e3dd80b9110a8dfaa8a9e75d9eb"
},
"go-split-method-owner/main.go": {
"captureGroups": 9,
"digest": "e9c105208ad6ef2f087f758972475aa403294e886751d75113b4ca45fa4e0b8a"
},
"go-split-method-owner/repo.go": {
"captureGroups": 8,
"digest": "f8bb213588f517e166b421f19d72cde981153464cc3f3168499769e778c4a22e"
},
"go-split-method-owner/save.go": {
"captureGroups": 6,
"digest": "6e0e1b5521a2fad1cd00252e771926d5caa7930440278968f7f8607821ad339c"
},
"go-split-method-owner/user.go": {
"captureGroups": 3,
"digest": "827dc0208b47776976313a4560fbd250ab7fa521ae46581213a79fc15fd52ce8"
},
"go-struct-literals/app.go": {
"captureGroups": 10,
"digest": "97ec29ec0e0dc1f23a804ff6808023a32eafdeb00f194d3bea9b50ed54c7f258"
},
"go-struct-literals/user.go": {
"captureGroups": 10,
"digest": "89b791a9d150fe341924d54d3173d9f35f72be899d0e7ac03a35d32119b94f8f"
},
"go-type-assertion/main.go": {
"captureGroups": 11,
"digest": "b8bd327d3965531802a93bc01e3c24969a027540397a1635119ef0b7da46c348"
},
"go-type-assertion/models.go": {
"captureGroups": 17,
"digest": "3780f8f7c145a15f849ae6958e492db0044e825b81aa8dfd21b9873759fad65e"
},
"go-variadic-resolution/cmd/main.go": {
"captureGroups": 6,
"digest": "443f9736b9c67df30fe97d8353ddc2584de4c88ba698e62878683d5d77362d89"
},
"go-variadic-resolution/internal/logger/logger.go": {
"captureGroups": 4,
"digest": "83b987f527f95e360966793148d07e569fa02fdc5e4febc09daa51cfce16b114"
},
"go-write-access/main.go": {
"captureGroups": 20,
"digest": "92fdad46b6933fd78dcbb19677a720f037306a4f03fedcca9156b54c1c9bf994"
},
"synthetic:dao-20": {
"captureGroups": 481,
"digest": "1698b5dd78c8094f251b10ab8cacebfbf453f38eddf7233fbc7964e32f04ceeb"
}
}
@@ -0,0 +1,7 @@
#include "lib.h"
namespace caller {
void run() {
run_callback(utils::make_token);
}
}

Some files were not shown because too many files have changed in this diff Show More